Ziding Zhang

dblp:83/7105 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-9296-571XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Graph neural network integrated with pretrained protein language model for predicting human-virus protein-protein interactions
abstract
The systematic identification of human-virus protein-protein interactions (PPIs) is a critical step toward elucidating the underlying mechanisms of viral infection, directly informing the development of targeted interventions against existing and emerging viral threats. In this work, we presented DeepGNHV, an end-to-end framework that integrated a pretrained protein language model with structural features derived from AlphaFold2 and leveraged graph attention networks to predict human-virus PPIs. In comparison to other state-of-the-art approaches, DeepGNHV exhibited superior predictive performance, especially when applied to viral proteins absent from the training process, indicating its strong generalization capability for detecting newly emerging virus-related PPIs. We further demonstrated DeepGNHV's robustness across diverse perturbations and its practical application under high-confidence thresholds. Additionally, we conducted extensive predictions of human-HPV PPIs, which were supported by multiple lines of evidence and identified several host factors that specifically interact with high-risk HPV. To further explore the biological significance of DeepGNHV, we provided a case study to pinpoint specific residues that play critical roles in facilitating the corresponding PPIs. The source code of DeepGNHV and related data is publicly available on GitHub (https://github.com/bioboy0415/DeepGNHV).
Linyang Jiang, Xiaodi Yang, Xiaokun Guo, Dianke Li, Stefan Wuchty, Wenyu Shi, Ziding Zhang
Briefings Bioinform.8
2025 OMetaNet: an efficient hybrid deep learning model based on multimodal data fusion and contrastive learning for predicting 2'-O-methylation sites in human RNA
abstract
BACKGROUND: Accurately identifying RNA 2'-O-methylation (2OM) sites is a crucial step in gaining an in-depth understanding of RNA regulatory mechanisms. Although there are currently multiple prediction tools available, they still suffer from limited prediction accuracy and an inability to fully capture the associations between sequences and sites. RESULTS: This study constructs a novel low-redundancy dataset and innovatively proposes the KN-PairMatrix encoding scheme, effectively addressing the research gap in sequence-site association analysis. Based on this foundation, we developed the deep learning framework OMetaNet, which integrates residual and downsampling-optimized CNN modules, Mamba network, and a proprietary cross-modal interactive fusion module. The framework incorporates a contrastive learning-driven adaptive hybrid loss function. Employing a progressive feature disentanglement strategy, it enhances the learning capability for 2OM site-specific patterns. Independent evaluation results demonstrate that OMetaNet significantly outperforms existing methods in predicting 2OM sites across all four nucleotide types. CONCLUSIONS: We proposed a novel computational model, OMetaNet. Its unique design structure may potentially reshape the paradigm of transcriptome analysis, open up new directions for extracting modification site information, and show significant potential in biomarker research and cross-species generalization studies.
Yiyu Lin, Sen Yang 0017, Ziding Zhang
BMC Bioinform.4
2024 Multi-modal features-based human-herpesvirus protein-protein interaction prediction by using LightGBM
abstract
The identification of human-herpesvirus protein-protein interactions (PPIs) is an essential and important entry point to understand the mechanisms of viral infection, especially in malignant tumor patients with common herpesvirus infection. While natural language processing (NLP)-based embedding techniques have emerged as powerful approaches, the application of multi-modal embedding feature fusion to predict human-herpesvirus PPIs is still limited. Here, we established a multi-modal embedding feature fusion-based LightGBM method to predict human-herpesvirus PPIs. In particular, we applied document and graph embedding approaches to represent sequence, network and function modal features of human and herpesviral proteins. Training our LightGBM models through our compiled non-rigorous and rigorous benchmarking datasets, we obtained significantly better performance compared to individual-modal features. Furthermore, our model outperformed traditional feature encodings-based machine learning methods and state-of-the-art deep learning-based methods using various benchmarking datasets. In a transfer learning step, we show that our model that was trained on human-herpesvirus PPI dataset without cytomegalovirus data can reliably predict human-cytomegalovirus PPIs, indicating that our method can comprehensively capture multi-modal fusion features of protein interactions across various herpesvirus subtypes. The implementation of our method is available at https://github.com/XiaodiYangpku/MultimodalPPI/.
Xiaodi Yang, Stefan Wuchty, Zeyin Liang, Ziding Zhang, Yujun Dong
Briefings Bioinform.7
2023 SGPPI: structure-aware prediction of protein-protein interactions in rigorous conditions with graph convolutional network
abstract
While deep learning (DL)-based models have emerged as powerful approaches to predict protein-protein interactions (PPIs), the reliance on explicit similarity measures (e.g. sequence similarity and network neighborhood) to known interacting proteins makes these methods ineffective in dealing with novel proteins. The advent of AlphaFold2 presents a significant opportunity and also a challenge to predict PPIs in a straightforward way based on monomer structures while controlling bias from protein sequences. In this work, we established Structure and Graph-based Predictions of Protein Interactions (SGPPI), a structure-based DL framework for predicting PPIs, using the graph convolutional network. In particular, SGPPI focused on protein patches on the protein-protein binding interfaces and extracted the structural, geometric and evolutionary features from the residue contact map to predict PPIs. We demonstrated that our model outperforms traditional machine learning methods and state-of-the-art DL-based methods using non-representation-bias benchmark datasets. Moreover, our model trained on human dataset can be reliably transferred to predict yeast PPIs, indicating that SGPPI can capture converging structural features of protein interactions across various species. The implementation of SGPPI is available at https://github.com/emerson106/SGPPI.
Stefan Wuchty, Ziding Zhang
Briefings Bioinform.4
2021 Multi-scale Convolutional Neural Networks for the Prediction of Human-virus Protein Interactions
abstract
Allowing the prediction of human-virus protein-protein interactions (PPI), our algorithm is based on a Siamese Convolutional Neural Network architecture (CNN), accounting for pre-acquired protein evolutionary profiles (i e PSSM) as input In combinations with a multilayer perceptron, we evaluate our model on a variety of human-virus PPI datasets and compare its results with traditional machine learning frameworks, a deep learning architecture and several other human-virus PPI prediction methods, showing superior performance Furthermore, we propose two transfer learning methods, allowing the reliable prediction of interactions in cross-viral settings, where we train our system with PPIs in a source human-virus domain and predict interactions in a target human-virus domain Notable, we observed that our transfer learning approaches allowed the reliable prediction of PPIs in relatively less investigated human-virus domains, such as Dengue, Zika and SARS-CoV-2 © 2021 by SCITEPRESS - Science and Technology Publications, Lda
Xiaodi Yang, Ziding Zhang, Stefan Wuchty
ICAART (2)2
2021 Current status and future perspectives of computational studies on human-virus protein-protein interactions
abstract
The protein-protein interactions (PPIs) between human and viruses mediate viral infection and host immunity processes. Therefore, the study of human-virus PPIs can help us understand the principles of human-virus relationships and can thus guide the development of highly effective drugs to break the transmission of viral infectious diseases. Recent years have witnessed the rapid accumulation of experimentally identified human-virus PPI data, which provides an unprecedented opportunity for bioinformatics studies revolving around human-virus PPIs. In this article, we provide a comprehensive overview of computational studies on human-virus PPIs, especially focusing on the method development for human-virus PPI predictions. We briefly introduce the experimental detection methods and existing database resources of human-virus PPIs, and then discuss the research progress in the development of computational prediction methods. In particular, we elaborate the machine learning-based prediction methods and highlight the need to embrace state-of-the-art deep-learning algorithms and new feature engineering techniques (e.g. the protein embedding technique derived from natural language processing). To further advance the understanding in this research topic, we also outline the practical applications of the human-virus interactome in fundamental biological discovery and new antiviral therapy development.
Xianyi Lian, Xiaodi Yang, Ziding Zhang
Briefings Bioinform.4
2021 HVIDB: a comprehensive database for human-virus protein-protein interactions
abstract
While leading to millions of people's deaths every year the treatment of viral infectious diseases remains a huge public health challenge.Therefore, an in-depth understanding of human-virus protein-protein interactions (PPIs) as the molecular interface between a virus and its host cell is of paramount importance to obtain new insights into the pathogenesis of viral infections and development of antiviral therapeutic treatments. However, current human-virus PPI database resources are incomplete, lack annotation and usually do not provide the opportunity to computationally predict human-virus PPIs. Here, we present the Human-Virus Interaction DataBase (HVIDB, http://zzdlab.com/hvidb/) that provides comprehensively annotated human-virus PPI data as well as seamlessly integrates online PPI prediction tools. Currently, HVIDB highlights 48 643 experimentally verified human-virus PPIs covering 35 virus families, 6633 virally targeted host complexes, 3572 host dependency/restriction factors as well as 911 experimentally verified/predicted 3D complex structures of human-virus PPIs. Furthermore, our database resource provides tissue-specific expression profiles of 6790 human genes that are targeted by viruses and 129 Gene Expression Omnibus series of differentially expressed genes post-viral infections. Based on these multifaceted and annotated data, our database allows the users to easily obtain reliable information about PPIs of various human viruses and conduct an in-depth analysis of their inherent biological significance. In particular, HVIDB also integrates well-performing machine learning models to predict interactions between the human host and viral proteins that are based on (i) sequence embedding techniques, (ii) interolog mapping and (iii) domain-domain interaction inference. We anticipate that HVIDB will serve as a one-stop knowledge base to further guide hypothesis-driven experimental efforts to investigate human-virus relationships.
Xiaodi Yang, Xianyi Lian, Stefan Wuchty, Ziding Zhang
Briefings Bioinform.6
2021 Transfer learning via multi-scale convolutional neural layers for human-virus protein-protein interaction prediction
abstract
MOTIVATION: To complement experimental efforts, machine learning-based computational methods are playing an increasingly important role to predict human-virus protein-protein interactions (PPIs). Furthermore, transfer learning can effectively apply prior knowledge obtained from a large source dataset/task to a small target dataset/task, improving prediction performance. RESULTS: To predict interactions between human and viral proteins, we combine evolutionary sequence profile features with a Siamese convolutional neural network (CNN) architecture and a multi-layer perceptron. Our architecture outperforms various feature encodings-based machine learning and state-of-the-art prediction methods. As our main contribution, we introduce two transfer learning methods (i.e. 'frozen' type and 'fine-tuning' type) that reliably predict interactions in a target human-virus domain based on training in a source human-virus domain, by retraining CNN layers. Finally, we utilize the 'frozen' type transfer learning approach to predict human-SARS-CoV-2 PPIs, indicating that our predictions are topologically and functionally similar to experimentally known interactions. AVAILABILITY AND IMPLEMENTATION: The source codes and datasets are available at https://github.com/XiaodiYangCAU/TransPPI/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaodi Yang, Xianyi Lian, Stefan Wuchty, Ziding Zhang
Bioinform.5
2019 Prediction of protein-protein interactions between fungus (Magnaporthe grisea) and rice (Oryza sativa L.)
abstract
Rice blast disease caused by the fungus Magnaporthe grisea (M. grisea) is one of the most serious diseases for the cultivated rice Oryza sativa (O. sativa). A key factor causing rice blast disease and defense might be protein-protein interactions (PPIs) between rice and fungus. In this research, we have developed a computational pipeline to predict PPIs between blast fungus and rice. After cross-prediction by interolog-based and domain-based method, we achieved 532 potential PPIs between 27 fungus proteins and 236 rice proteins. Accuracy of jackknife test, 10-fold cross-validation test and independent test for these PPIs were 90.43, 93.85 and 84.67%, respectively, by using support vector machine classification method. Meanwhile, the pathogenic genes of blast fungus were enriched in the predicted PPIs network when compared with 1000 random interaction networks. The rice regulatory network was downloaded and divided into 228 subnetworks with over six nodes, and the top seven subnetworks affected by blast fungus through PPIs were investigated. The results indicated that 34 upregulated and 12 downregulated master regulators in rice interacting with the fungus proteins in response to the infection of blast fungus. The common master regulators in rice in response to the infection of M. grisea, Xanthomonas oryzae pv.oryzae and rice stripe virus were analyzed. The ubiquitin proteasome pathway was the common pathway in rice regulated by these three pathogens, while apoptosis signaling pathway was induced by fungus and bacteria. In summary, the results in this article provide insight into the process of blast fungus infection.
Shiwei Ma, Andrew P. Harrison, Wei Liu 0072, Shoukai Lin, Ziding Zhang, Yufang Ai, Huaqin He
Briefings Bioinform.8
2019 Critical assessment and performance improvement of plant-pathogen protein-protein interaction prediction methods
abstract
The identification of plant-pathogen protein-protein interactions (PPIs) is an attractive and challenging research topic for deciphering the complex molecular mechanism of plant immunity and pathogen infection. Considering that the experimental identification of plant-pathogen PPIs is time-consuming and labor-intensive, computational methods are emerging as an important strategy to complement the experimental methods. In this work, we first evaluated the performance of traditional computational methods such as interolog, domain-domain interaction and domain-motif interaction in predicting known plant-pathogen PPIs. Owing to the low sensitivity of the traditional methods, we utilized Random Forest to build an inter-species PPI prediction model based on multiple sequence encodings and novel network attributes in the established plant PPI network. Critical assessment of the features demonstrated that the integration of sequence information and network attributes resulted in significant and robust performance improvement. Additionally, we also discussed the influence of Gene Ontology and gene expression information on the prediction performance. The Web server implementing the integrated prediction method, named InterSPPI, has been made freely available at http://systbio.cau.edu.cn/intersppi/index.php. InterSPPI could achieve a reasonably high accuracy with a precision of 73.8% and a recall of 76.6% in the independent test. To examine the applicability of InterSPPI, we also conducted cross-species and proteome-wide plant-pathogen PPI prediction tests. Taken together, we hope this work can provide a comprehensive understanding of the current status of plant-pathogen PPI predictions, and the proposed InterSPPI can become a useful tool to accelerate the exploration of plant-pathogen interactions.
Huaqin He, Ziding Zhang
Briefings Bioinform.5
2015 Towards more accurate prediction of ubiquitination sites: a comprehensive review of current methods, tools and features
abstract
Protein ubiquitination is one of the most important reversible post-translational modifications (PTMs). In many biochemical, pathological and pharmaceutical studies on understanding the function of proteins in biological processes, identification of ubiquitination sites is an important first step. However, experimental approaches for identifying ubiquitination sites are often expensive, labor-intensive and time-consuming, partly due to the dynamics and reversibility of ubiquitination. In silico prediction of ubiquitination sites is potentially a useful strategy for whole proteome annotation. A number of bioinformatics approaches and tools have recently been developed for predicting protein ubiquitination sites. However, these tools have different methodologies, prediction algorithms, functionality and features, which complicate their utility and application. The purpose of this review is to aid users in selecting appropriate tools for specific analyses and circumstances. We first compared five popular webservers and standalone software options, assessing their performance on four up-to-date ubiquitination benchmark datasets from Saccharomyces cerevisiae, Homo sapiens, Mus musculus and Arabidopsis thaliana. We then discussed and summarized these tools to guide users in choosing among the tools efficiently and rapidly. Finally, we assessed the importance of features of existing tools for ubiquitination site prediction, ranking them by performance. We also discussed the features that make noticeable contributions to species-specific ubiquitination site prediction.
Zhen Chen 0009, Ziding Zhang, Jiangning Song
Briefings Bioinform.3
2011 Outer membrane proteins can be simply identified using secondary structure element alignment
abstract
BACKGROUND: Outer membrane proteins (OMPs) are frequently found in the outer membranes of gram-negative bacteria, mitochondria and chloroplasts and have been found to play diverse functional roles. Computational discrimination of OMPs from globular proteins and other types of membrane proteins is helpful to accelerate new genome annotation and drug discovery. RESULTS: Based on the observation that almost all OMPs consist of antiparallel β-strands in a barrel shape and that their secondary structure arrangements differ from those of other types of proteins, we propose a simple method called SSEA-OMP to identify OMPs using secondary structure element alignment. Through intensive benchmark experiments, the proposed SSEA-OMP method is better than some well-established OMP detection methods. CONCLUSIONS: The major advantage of SSEA-OMP is its good prediction performance considering its simplicity. The web server implements the method is freely accessible at http://protein.cau.edu.cn/SSEA-OMP/index.html.
Ren-Xiang Yan, Zhen Chen 0009, Ziding Zhang
BMC Bioinform.3
2009 DescFold: A web server for protein fold recognition
abstract
BACKGROUND: Machine learning-based methods have been proven to be powerful in developing new fold recognition tools. In our previous work [Zhang, Kochhar and Grigorov (2005) Protein Science, 14: 431-444], a machine learning-based method called DescFold was established by using Support Vector Machines (SVMs) to combine the following four descriptors: a profile-sequence-alignment-based descriptor using Psi-blast e-values and bit scores, a sequence-profile-alignment-based descriptor using Rps-blast e-values and bit scores, a descriptor based on secondary structure element alignment (SSEA), and a descriptor based on the occurrence of PROSITE functional motifs. In this work, we focus on the improvement of DescFold by incorporating more powerful descriptors and setting up a user-friendly web server. RESULTS: In seeking more powerful descriptors, the profile-profile alignment score generated from the COMPASS algorithm was first considered as a new descriptor (i.e., PPA). When considering a profile-profile alignment between two proteins in the context of fold recognition, one protein is regarded as a template (i.e., its 3D structure is known). Instead of a sequence profile derived from a Psi-blast search, a structure-seeded profile for the template protein was generated by searching its structural neighbors with the assistance of the TM-align structural alignment algorithm. Moreover, the COMPASS algorithm was used again to derive a profile-structural-profile-alignment-based descriptor (i.e., PSPA). We trained and tested the new DescFold in a total of 1,835 highly diverse proteins extracted from the SCOP 1.73 version. When the PPA and PSPA descriptors were introduced, the new DescFold boosts the performance of fold recognition substantially. Using the SCOP_1.73_40% dataset as the fold library, the DescFold web server based on the trained SVM models was further constructed. To provide a large-scale test for the new DescFold, a stringent test set of 1,866 proteins were selected from the SCOP 1.75 version. At a less than 5% false positive rate control, the new DescFold is able to correctly recognize structural homologs at the fold level for nearly 46% test proteins. Additionally, we also benchmarked the DescFold method against several well-established fold recognition algorithms through the LiveBench targets and Lindahl dataset. CONCLUSIONS: The new DescFold method was intensively benchmarked to have very competitive performance compared with some well-established fold recognition methods, suggesting that it can serve as a useful tool to assist in template-based protein structure prediction. The DescFold server is freely accessible at http://202.112.170.199/DescFold/index.html.
Ren-Xiang Yan, Jing-Na Si, Ziding Zhang
BMC Bioinform.4
2008 Prediction of mucin-type O-glycosylation sites in mammalian proteins using the composition of k-spaced amino acid pairs
abstract
BACKGROUND: As one of the most common protein post-translational modifications, glycosylation is involved in a variety of important biological processes. Computational identification of glycosylation sites in protein sequences becomes increasingly important in the post-genomic era. A new encoding scheme was employed to improve the prediction of mucin-type O-glycosylation sites in mammalian proteins. RESULTS: A new protein bioinformatics tool, CKSAAP_OGlySite, was developed to predict mucin-type O-glycosylation serine/threonine (S/T) sites in mammalian proteins. Using the composition of k-spaced amino acid pairs (CKSAAP) based encoding scheme, the proposed method was trained and tested in a new and stringent O-glycosylation dataset with the assistance of Support Vector Machine (SVM). When the ratio of O-glycosylation to non-glycosylation sites in training datasets was set as 1:1, 10-fold cross-validation tests showed that the proposed method yielded a high accuracy of 83.1% and 81.4% in predicting O-glycosylated S and T sites, respectively. Based on the same datasets, CKSAAP_OGlySite resulted in a higher accuracy than the conventional binary encoding based method (about +5.0%). When trained and tested in 1:5 datasets, the CKSAAP encoding showed a more significant improvement than the binary encoding. We also merged the training datasets of S and T sites and integrated the prediction of S and T sites into one single predictor (i.e. S+T predictor). Either in 1:1 or 1:5 datasets, the performance of this S+T predictor was always slightly better than those predictors where S and T sites were independently predicted, suggesting that the molecular recognition of O-glycosylated S/T sites seems to be similar and the increase of the S+T predictor's accuracy may be a result of expanded training datasets. Moreover, CKSAAP_OGlySite was also shown to have better performance when benchmarked against two existing predictors. CONCLUSION: Because of CKSAAP encoding's ability of reflecting characteristics of the sequences surrounding mucin-type O-glycosylation sites, CKSAAP_ OGlySite has been proved more powerful than the conventional binary encoding based method. This suggests that it can be used as a competitive mucin-type O-glycosylation site predictor to the biological community. CKSAAP_OGlySite is now available at http://bioinformatics.cau.edu.cn/zzd_lab/CKSAAP_OGlySite/.
Yong-Zi Chen, Yu-Rong Tang, Zhi-Ya Sheng, Ziding Zhang
BMC Bioinform.4