Stefan Wuchty

dblp:60/840 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0001-8916-6522ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021
YearPublicationVenuePosition
2025 Graph neural network integrated with pretrained protein language model for predicting human-virus protein-protein interactions
abstract
The systematic identification of human-virus protein-protein interactions (PPIs) is a critical step toward elucidating the underlying mechanisms of viral infection, directly informing the development of targeted interventions against existing and emerging viral threats. In this work, we presented DeepGNHV, an end-to-end framework that integrated a pretrained protein language model with structural features derived from AlphaFold2 and leveraged graph attention networks to predict human-virus PPIs. In comparison to other state-of-the-art approaches, DeepGNHV exhibited superior predictive performance, especially when applied to viral proteins absent from the training process, indicating its strong generalization capability for detecting newly emerging virus-related PPIs. We further demonstrated DeepGNHV's robustness across diverse perturbations and its practical application under high-confidence thresholds. Additionally, we conducted extensive predictions of human-HPV PPIs, which were supported by multiple lines of evidence and identified several host factors that specifically interact with high-risk HPV. To further explore the biological significance of DeepGNHV, we provided a case study to pinpoint specific residues that play critical roles in facilitating the corresponding PPIs. The source code of DeepGNHV and related data is publicly available on GitHub (https://github.com/bioboy0415/DeepGNHV).
Linyang Jiang, Xiaodi Yang, Xiaokun Guo, Dianke Li, Stefan Wuchty, Wenyu Shi, Ziding Zhang
Briefings Bioinform.6
2024 Improving Audience Ratings Prediction of Japanese TV Dramas Using Knowledge-Based Embeddings
abstract
Accurate prediction of television (TV) audience ratings is a crucial technology for advertisers and broadcast-ers, which influences advertising costs and indicates program popularity. As for Japanese TV dramas, previous studies have examined factors available prior to broadcast, such as cast com-position, metadata variables (e.g., scheduled time slot, broadcasting station, episode duration), and indicators of actor popularity. However, such studies have been limited in their evaluation of contextual information related to these factors and have not considered additional modalities such as text data. Here, we propose the use of knowledge-based embeddings derived from Japanese Wikipedia to improve audience ratings prediction. The proposed bag-of-entities method is applied to represent dramas in terms of cast, production team, and plot synopsis elements, and is evaluated against conventional metadata features. In an ablation study, supervised forecasting models using Support Vector Machine (SVM) are trained on combinations of the feature groups. These are devised and tested on viewership data collected from the Audience Rating TV (ARTV) database for 574 Japanese TV dramas aired between 2003 and 2021. Our experimental results indicate that a SVM model conditioned on the full feature set outperforms baseline methods, achieving a 77.4 % F1-score. After stratifying predictions based on time slot popularity, the approach is found to be effective for predicting audience ratings in challenging, high-popularity time slots, with improvements in recognition rate up to 19.1 %.
Jerry Bonnell, Stefan Wuchty, Mitsunori Ogihara
ICMLA2
2024 Multi-modal features-based human-herpesvirus protein-protein interaction prediction by using LightGBM
abstract
The identification of human-herpesvirus protein-protein interactions (PPIs) is an essential and important entry point to understand the mechanisms of viral infection, especially in malignant tumor patients with common herpesvirus infection. While natural language processing (NLP)-based embedding techniques have emerged as powerful approaches, the application of multi-modal embedding feature fusion to predict human-herpesvirus PPIs is still limited. Here, we established a multi-modal embedding feature fusion-based LightGBM method to predict human-herpesvirus PPIs. In particular, we applied document and graph embedding approaches to represent sequence, network and function modal features of human and herpesviral proteins. Training our LightGBM models through our compiled non-rigorous and rigorous benchmarking datasets, we obtained significantly better performance compared to individual-modal features. Furthermore, our model outperformed traditional feature encodings-based machine learning methods and state-of-the-art deep learning-based methods using various benchmarking datasets. In a transfer learning step, we show that our model that was trained on human-herpesvirus PPI dataset without cytomegalovirus data can reliably predict human-cytomegalovirus PPIs, indicating that our method can comprehensively capture multi-modal fusion features of protein interactions across various herpesvirus subtypes. The implementation of our method is available at https://github.com/XiaodiYangpku/MultimodalPPI/.
Xiaodi Yang, Stefan Wuchty, Zeyin Liang, Ziding Zhang, Yujun Dong
Briefings Bioinform.2
2023 SGPPI: structure-aware prediction of protein-protein interactions in rigorous conditions with graph convolutional network
abstract
While deep learning (DL)-based models have emerged as powerful approaches to predict protein-protein interactions (PPIs), the reliance on explicit similarity measures (e.g. sequence similarity and network neighborhood) to known interacting proteins makes these methods ineffective in dealing with novel proteins. The advent of AlphaFold2 presents a significant opportunity and also a challenge to predict PPIs in a straightforward way based on monomer structures while controlling bias from protein sequences. In this work, we established Structure and Graph-based Predictions of Protein Interactions (SGPPI), a structure-based DL framework for predicting PPIs, using the graph convolutional network. In particular, SGPPI focused on protein patches on the protein-protein binding interfaces and extracted the structural, geometric and evolutionary features from the residue contact map to predict PPIs. We demonstrated that our model outperforms traditional machine learning methods and state-of-the-art DL-based methods using non-representation-bias benchmark datasets. Moreover, our model trained on human dataset can be reliably transferred to predict yeast PPIs, indicating that SGPPI can capture converging structural features of protein interactions across various species. The implementation of SGPPI is available at https://github.com/emerson106/SGPPI.
Stefan Wuchty, Ziding Zhang
Briefings Bioinform.2
2021 Multi-scale Convolutional Neural Networks for the Prediction of Human-virus Protein Interactions
abstract
Allowing the prediction of human-virus protein-protein interactions (PPI), our algorithm is based on a Siamese Convolutional Neural Network architecture (CNN), accounting for pre-acquired protein evolutionary profiles (i e PSSM) as input In combinations with a multilayer perceptron, we evaluate our model on a variety of human-virus PPI datasets and compare its results with traditional machine learning frameworks, a deep learning architecture and several other human-virus PPI prediction methods, showing superior performance Furthermore, we propose two transfer learning methods, allowing the reliable prediction of interactions in cross-viral settings, where we train our system with PPIs in a source human-virus domain and predict interactions in a target human-virus domain Notable, we observed that our transfer learning approaches allowed the reliable prediction of PPIs in relatively less investigated human-virus domains, such as Dengue, Zika and SARS-CoV-2 © 2021 by SCITEPRESS - Science and Technology Publications, Lda
Xiaodi Yang, Ziding Zhang, Stefan Wuchty
ICAART (2)3
2021 HVIDB: a comprehensive database for human-virus protein-protein interactions
abstract
While leading to millions of people's deaths every year the treatment of viral infectious diseases remains a huge public health challenge.Therefore, an in-depth understanding of human-virus protein-protein interactions (PPIs) as the molecular interface between a virus and its host cell is of paramount importance to obtain new insights into the pathogenesis of viral infections and development of antiviral therapeutic treatments. However, current human-virus PPI database resources are incomplete, lack annotation and usually do not provide the opportunity to computationally predict human-virus PPIs. Here, we present the Human-Virus Interaction DataBase (HVIDB, http://zzdlab.com/hvidb/) that provides comprehensively annotated human-virus PPI data as well as seamlessly integrates online PPI prediction tools. Currently, HVIDB highlights 48 643 experimentally verified human-virus PPIs covering 35 virus families, 6633 virally targeted host complexes, 3572 host dependency/restriction factors as well as 911 experimentally verified/predicted 3D complex structures of human-virus PPIs. Furthermore, our database resource provides tissue-specific expression profiles of 6790 human genes that are targeted by viruses and 129 Gene Expression Omnibus series of differentially expressed genes post-viral infections. Based on these multifaceted and annotated data, our database allows the users to easily obtain reliable information about PPIs of various human viruses and conduct an in-depth analysis of their inherent biological significance. In particular, HVIDB also integrates well-performing machine learning models to predict interactions between the human host and viral proteins that are based on (i) sequence embedding techniques, (ii) interolog mapping and (iii) domain-domain interaction inference. We anticipate that HVIDB will serve as a one-stop knowledge base to further guide hypothesis-driven experimental efforts to investigate human-virus relationships.
Xiaodi Yang, Xianyi Lian, Stefan Wuchty, Ziding Zhang
Briefings Bioinform.4
2021 Transfer learning via multi-scale convolutional neural layers for human-virus protein-protein interaction prediction
abstract
MOTIVATION: To complement experimental efforts, machine learning-based computational methods are playing an increasingly important role to predict human-virus protein-protein interactions (PPIs). Furthermore, transfer learning can effectively apply prior knowledge obtained from a large source dataset/task to a small target dataset/task, improving prediction performance. RESULTS: To predict interactions between human and viral proteins, we combine evolutionary sequence profile features with a Siamese convolutional neural network (CNN) architecture and a multi-layer perceptron. Our architecture outperforms various feature encodings-based machine learning and state-of-the-art prediction methods. As our main contribution, we introduce two transfer learning methods (i.e. 'frozen' type and 'fine-tuning' type) that reliably predict interactions in a target human-virus domain based on training in a source human-virus domain, by retraining CNN layers. Finally, we utilize the 'frozen' type transfer learning approach to predict human-SARS-CoV-2 PPIs, indicating that our predictions are topologically and functionally similar to experimentally known interactions. AVAILABILITY AND IMPLEMENTATION: The source codes and datasets are available at https://github.com/XiaodiYangCAU/TransPPI/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaodi Yang, Xianyi Lian, Stefan Wuchty, Ziding Zhang
Bioinform.4
2020 Stable Feature Selection for Gene Expression using Enhanced Binary Particle Swarm Optimization
Hassen Dhrif, Stefan Wuchty
ICAART (2)2
2020 Gene Subset Selection for Transfer Learning using Bilevel Particle Swarm Optimization
abstract
In classification problems that involve multiple sources, data distributions may vary. Therefore, knowledge of dimensions that differ in the source and target data is important to reduce the distance between domains, allowing accurate transfer knowledge. Here, we present a novel method to identify (in)variant genes between source and target datasets and integrate such results to simultaneously reduce the variance between two distributions while optimizing the size and classification error of the selected subset. In particular, we use an evolutionary computation particle swarm optimization algorithm to implement such a bilevel multi-objective programming approach, allowing us to solve a gene subset selection problem.
Hassen Dhrif, Verónica Bolón-Canedo, Stefan Wuchty
ICMLA3
2019 A stable hybrid method for feature subset selection using particle swarm optimization with local search
abstract
The determination of a small set of biomarkers to make a diagnostic call can be formulated as a feature subset selection (FSS) problem to find a small set of genes with high relevance for the underlying classification task and low mutual redundancy. However, repeated application of a heuristic, evolutionary FSS technique usually fails to produce consistent results. Here, we introduce COMB-PSO-LS, a novel hybrid (wrapper-filter) FSS algorithm based on Particle Swarm Optimization (PSO) that features a local search strategy to select the least dependent and most relevant feature subsets. In particular, we employ a Randomized Dependence Coefficient (RDC)-based filter technique to guide the search process of the particle swarm, allowing the selection of highly relevant and consistent features. Classifying cancer samples through patient gene expression profiles, we found that COMB-PSO-LS provides highly stable and non-redundant gene subsets that are relevant for the classification process, outperforming standard PSO methods.
Hassen Dhrif, Luis Gonzalo Sánchez Giraldo, Miroslav Kubat, Stefan Wuchty
GECCO4
2017 Bacterial protein meta-interactomes predict cross-species interactions and protein function
abstract
BACKGROUND: Protein-protein interactions (PPIs) can offer compelling evidence for protein function, especially when viewed in the context of proteome-wide interactomes. Bacteria have been popular subjects of interactome studies: more than six different bacterial species have been the subjects of comprehensive interactome studies while several more have had substantial segments of their proteomes screened for interactions. The protein interactomes of several bacterial species have been completed, including several from prominent human pathogens. The availability of interactome data has brought challenges, as these large data sets are difficult to compare across species, limiting their usefulness for broad studies of microbial genetics and evolution. RESULTS: In this study, we use more than 52,000 unique protein-protein interactions (PPIs) across 349 different bacterial species and strains to determine their conservation across data sets and taxonomic groups. When proteins are collapsed into orthologous groups (OGs) the resulting meta-interactome still includes more than 43,000 interactions, about 14,000 of which involve proteins of unknown function. While conserved interactions provide support for protein function in their respective species data, we found only 429 PPIs (~1% of the available data) conserved in two or more species, rendering any cross-species interactome comparison immediately useful. The meta-interactome serves as a model for predicting interactions, protein functions, and even full interactome sizes for species with limited to no experimentally observed PPI, including Bacillus subtilis and Salmonella enterica which are predicted to have up to 18,000 and 31,000 PPIs, respectively. CONCLUSIONS: In the course of this work, we have assembled cross-species interactome comparisons that will allow interactomics researchers to anticipate the structures of yet-unexplored microbial interactomes and to focus on well-conserved yet uncharacterized interactors for further study. Such conserved interactions should provide evidence for important but yet-uncharacterized aspects of bacterial physiology and may provide targets for anti-microbial therapies.
J. Harry Caufield, Christopher Wimble, Shary Semarjit, Stefan Wuchty, Peter Uetz
BMC Bioinform.4
2015 The road to linking genomics and proteomics of pathogenic bacteria: from binary protein complexes to interaction pathways
abstract
The availability of fully sequenced genomes of many bacterial organisms has enabled mapping networks of binary protein interactions that form the basic building blocks of molecular pathways and dynamic assemblies defining all cellular activities. Few proteome-scale studies have been reported for pathogenic bacteria though, suggesting that a systems-wide network analysis of binary interaction partners could reveal groups of proteins that coordinate to achieve specific biological tasks important to pathogenesis and provide a functional map useful to the discovery of new antibiotics, vaccines, and diagnostic tools. We performed a comprehensive proteomics analysis of the pathogenic bacterium Yersinia pestis and analytically identified more than 74,000 binary interactions. Using a library of biotinylated recombinant proteins to probe a planar microarray comprised of immobilized proteins that represented approximately 85% (3,552 proteins) of the Y. pestis proteome, we measured protein-protein interactions by fluorescence intensity of the laser-scanned microarrays. We obtained kinetic interaction data for >1,600 binary complexes by microarray-based, surface plasmon resonance imaging, and identified several high-affinity ( K D ~ nM) interactions. We applied a machine learning algorithm that used previously reported experimental protein-protein interactions from Escherichia coli as a training set in order to extract E. coli -like interactions from the Y. pestis dataset. The node degree distribution of the resulting network, comprised of 2344 interactions between 314 proteins, approximates a power-law distribution typical of scale-free networks. Functional annotation clustering of proteins within the network revealed statistically enriched complexes and pathways involved in diverse biological processes. Among the more notable protein assemblies identified were components of the RNA polymerase enzyme and ribosomes. Small modules of proteins related to various metabolic pathways, as well as previously reported interactions involved in homologous recombination and fatty acid biosynthesis, were also present in the network. Two highly interconnected network sub-regions contained a large percentage of proteins with functions linked to transcription and translation. We have systematically identified and analyzed thousands of direct binary protein interactions within Y. pestis . This new benchmark data set will serve as a critical tool for the analysis of protein interaction networks functioning within an important human pathogen.
Sarah L. Keasey, Mohan Natesan, Christine Pugh, Teddy Kamata, Stefan Wuchty, Robert G. Ulrich
BMC Bioinform.5
2015 Essentiality and centrality in protein interaction networks revisited
abstract
BACKGROUND: Minimum dominating sets (MDSet) of protein interaction networks allow the control of underlying protein interaction networks through their topological placement. While essential proteins are enriched in MDSets, we hypothesize that the statistical properties of biological functions of essential genes are enhanced when we focus on essential MDSet proteins (e-MDSet). RESULTS: Here, we determined minimum dominating sets of proteins (MDSet) in interaction networks of E. coli, S. cerevisiae and H. sapiens, defined as subsets of proteins whereby each remaining protein can be reached by a single interaction. We compared several topological and functional parameters of essential, MDSet, and essential MDSet (e-MDSet) proteins. In particular, we observed that their topological placement allowed e-MDSet proteins to provide a positive correlation between degree and lethality, connect more protein complexes, and have a stronger impact on network resilience than essential proteins alone. In comparison to essential proteins we further found that interactions between e-MDSet proteins appeared more frequently within complexes, while interactions of e-MDSet proteins between complexes were depleted. Finally, these e-MDSet proteins classified into functional groupings that play a central role in survival and adaptability. CONCLUSIONS: The determination of e-MDSet of an organism highlights a set of proteins that enhances the enrichment signals of biological functions of essential proteins. As a consequence, we surmise that e-MDSets may provide a new method of evaluating the core proteins of an organism.
Sawsan Khuri, Stefan Wuchty
BMC Bioinform.2
2013 Important miRs of Pathways in Different Tumor Types
abstract
We computationally determined miRs that are significantly connected to molecular pathways by utilizing gene expression profiles in different cancer types such as glioblastomas, ovarian and breast cancers. Specifically, we assumed that the knowledge of physical interactions between miRs and genes indicated subsets of important miRs (IM) that significantly contributed to the regression of pathway-specific enrichment scores. Despite the different nature of the considered cancer types, we found strongly overlapping sets of IMs. Furthermore, IMs that were important for many pathways were enriched with literature-curated cancer and differentially expressed miRs. Such sets of IMs also coincided well with clusters of miRs that were experimentally indicated in numerous other cancer types. In particular, we focused on an overlapping set of 99 overall important miRs (OIM) that were found in glioblastomas, ovarian and breast cancers simultaneously. Notably, we observed that interactions between OIMs and leading edge genes of differentially expressed pathways were characterized by considerable changes in their expression correlations. Such gains/losses of miR and gene expression correlation indicated miR/gene pairs that may play a causal role in the underlying cancers.
Stefan Wuchty, Dolores Arjona, Peter O. Bauer
PLoS Comput. Biol.1
2011 Identifying Causal Genes and Dysregulated Pathways in Complex Diseases
abstract
In complex diseases, various combinations of genomic perturbations often lead to the same phenotype. On a molecular level, combinations of genomic perturbations are assumed to dys-regulate the same cellular pathways. Such a pathway-centric perspective is fundamental to understanding the mechanisms of complex diseases and the identification of potential drug targets. In order to provide an integrated perspective on complex disease mechanisms, we developed a novel computational method to simultaneously identify causal genes and dys-regulated pathways. First, we identified a representative set of genes that are differentially expressed in cancer compared to non-tumor control cases. Assuming that disease-associated gene expression changes are caused by genomic alterations, we determined potential paths from such genomic causes to target genes through a network of molecular interactions. Applying our method to sets of genomic alterations and gene expression profiles of 158 Glioblastoma multiforme (GBM) patients we uncovered candidate causal genes and causal paths that are potentially responsible for the altered expression of disease genes. We discovered a set of putative causal genes that potentially play a role in the disease. Combining an expression Quantitative Trait Loci (eQTL) analysis with pathway information, our approach allowed us not only to identify potential causal genes but also to find intermediate nodes and pathways mediating the information flow between causal and target genes. Our results indicate that different genomic perturbations indeed dys-regulate the same functional pathways, supporting a pathway-centric perspective of cancer. While copy number alterations and gene expression data of glioblastoma patients provided opportunities to test our approach, our method can be applied to any disease system where genetic variations play a fundamental causal role.
Yoo-Ah Kim, Stefan Wuchty, Teresa M. Przytycka
PLoS Comput. Biol.2
2010 Simultaneous Identification of Causal Genes and Dys-Regulated Pathways in Complex Diseases
Yoo-Ah Kim, Stefan Wuchty, Teresa M. Przytycka
RECOMB2
2010 FastMEDUSA: a parallelized tool to infer gene regulatory networks
abstract
MOTIVATION: In order to construct gene regulatory networks of higher organisms from gene expression and promoter sequence data efficiently, we developed FastMEDUSA. In this parallelized version of the regulatory network-modeling tool MEDUSA, expression and sequence data are shared among a user-defined number of processors on a single multi-core machine or cluster. Our results show that FastMEDUSA allows a more efficient utilization of computational resources. While the determination of a regulatory network of brain tumor in Homo sapiens takes 12 days with MEDUSA, FastMEDUSA obtained the same results in 6 h by utilizing 100 processors. AVAILABILITY: Source code and documentation of FastMEDUSA are available at https://wiki.nci.nih.gov/display/NOBbioinf/FastMEDUSA
Serdar Bozdag, Stefan Wuchty, Howard A. Fine
Bioinform.3
2010 Gene pathways and subnetworks distinguish between major glioma subtypes and elucidate potential underlying biology
Stefan Wuchty, Jennifer Walling, Susie Ahn, Martha Quezado, Carl Oberholtzer, Jean-Claude Zenklusen, Howard A. Fine
J. Biomed. Informatics1
2009 Graph theoretical approach to study eQTL: a case study of Plasmodium falciparum
abstract
MOTIVATION: Analysis of expression quantitative trait loci (eQTL) significantly contributes to the determination of gene regulation programs. However, the discovery and analysis of associations of gene expression levels and their underlying sequence polymorphisms continue to pose many challenges. Methods are limited in their ability to illuminate the full structure of the eQTL data. Most rely on an exhaustive, genome scale search that considers all possible locus-gene pairs and tests the linkage between each locus and gene. RESULT: To analyze eQTLs in a more comprehensive and efficient way, we developed the Graph based eQTL Decomposition method (GeD) that allows us to model genotype and expression data using an eQTL association graph. Through graph-based heuristics, GeD identifies dense subgraphs in the eQTL association graph. By identifying eQTL association cliques that expose the hidden structure of genotype and expression data, GeD effectively filters out most locus-gene pairs that are unlikely to have significant linkage. We apply GeD on eQTL data from Plasmodium falciparum, the human malaria parasite, and show that GeD reveals the structure of the relationship between all loci and all genes on a whole genome level. Furthermore, GeD allows us to uncover additional eQTLs with lower FDR, providing an important complement to traditional eQTL analysis methods.
Stefan Wuchty, Michael T. Ferdig, Teresa M. Przytycka
Bioinform.2
2007 Predicting Protein-Protein Interactions from Protein Domains Using a Set Cover Approach
abstract
One goal of contemporary proteome research is the elucidation of cellular protein interactions. Based on currently available protein-protein interaction and domain data, we introduce a novel method, Maximum Specificity Set Cover (MSSC), for the prediction of protein-protein interactions. In our approach, we map the relationship between interactions of proteins and their corresponding domain architectures to a generalized weighted set cover problem. The application of a greedy algorithm provides sets of domain interactions which explain the presence of protein interactions to the largest degree of specificity. Utilizing domain and protein interaction data of S. cerevisiae, MSSC enables prediction of previously unknown protein interactions, links that are well supported by a high tendency of coexpression and functional homogeneity of the corresponding proteins. Focusing on concrete examples, we show that MSSC reliably predicts protein interactions in well-studied molecular systems, such as the 26S proteasome and RNA polymerase II of S. cerevisiae. We also show that the quality of the predictions is comparable to the Maximum Likelihood Estimation while MSSC is faster. This new algorithm and all data sets used are accessible through a Web portal at http://ppi.cse.nd.edu.
Chengbang Huang, Faruck Morcos, Simon P. Kanaan, Stefan Wuchty, Danny Ziyi Chen, Jesús A. Izaguirre
IEEE ACM Trans. Comput. Biol. Bioinform.4