VLDB 2026 Research / reviewers in the wild / expert
Xiaolei Zhu 0001
dblp:04/5901-1
· DBLP profile ↗
13ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-1967-2806ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive dynamic training heterogeneous graph attention neural network for microRNA-disease association prediction and analysis frameworkabstractAccurately predicting microRNA-disease associations (MDAs) is crucial for identifying biomarkers and therapeutic targets in complex diseases. However, experimental validation is costly, and existing computational methods, particularly Graph Neural Networks (GNNs), are limited by data sparsity. In such sparse graphs, the lack of topological connections hinders effective message passing, leading to poor performance for isolated nodes and in cold-start scenarios. To address these challenges, we propose the adaptive dynamic heterogeneous graph attention neural network (ADHGMDA), a framework for MDAs prediction and analysis. First, to resolve connectivity issues, we construct a dual-layer heterogeneous graph enhanced with virtual nodes derived via K-means clustering. These virtual nodes act as semantic bridges, integrating isolated microRNAs and diseases into the network. Second, we introduce a dynamic feature learning mechanism that simulates data sparsity during training, forcing the model to learn robust representations for cold-start prediction. Third, to mitigate the over-smoothing common in deep GNNs, we design a gradient conflict-aware multi-task optimization strategy with dynamic weight adaptation. Furthermore, addressing the issue that existing benchmarks rely on outdated data, we reconstructed a benchmark dataset based on the latest databases, integrated with visualization tools. Experimental results demonstrate that ADHGMDA significantly outperforms seven state-of-the-art methods, achieving area under the receiver operating characteristic curve scores of 0.9698 and 0.9730. In case studies on herpes simplex and interstitial nephritis, the model exhibits excellent performance in uncovering potential pathogenic pathways. Meanwhile, extensive validation experiments confirm that it has good robustness against label noise, class imbalance, and the cold-start problem of isolated nodes. Yinbo Liu, Sijian Wen, Yue Yi, Yongmei Michelle Wang, Xiaolei Zhu 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Prediction of circRNA-drug sensitivity using random auto-encoders and multi-layer heterogeneous graph transformers
Yinbo Liu, Xinxin Ren, Xiaolei Zhu 0001 |
Appl. Intell. | 5 |
| 2025 | Predicting nucleic acid binding sites by attention map-guided graph convolutional network with protein language embeddings and physicochemical informationabstractProtein-nucleic acid binding sites play a crucial role in biological processes such as gene expression, signal transduction, replication, and transcription. In recent years, with the development of artificial intelligence, protein language models, graph neural networks, and transformer architectures have been adopted to develop both structure-based and sequence-based predictive models. Structure-based methods benefit from the spatial relationship between residues and have shown promising performance. However, structure-based information requires 3D protein structures, which is a challenge for large-scale protein sequence spaces. To address this limitation, researchers have attempted to use predicted protein structure information to guide binding site prediction. While this strategy has improved accuracy, it still depends on the quality of structure predictions. Thus, some studies have returned to prediction methods based solely on protein sequences, particularly those using protein language models, which have greatly enhanced the prediction accuracy. This paper proposes a novel protein-nucleic acid binding site prediction framework, ATtention Maps and Graph convolutional neural networks to predict nucleic acid-protein Binding sites (ATMGBs), which first fuses protein language embeddings with physicochemical properties to obtain multiview information, then leverages the attention map of a protein language model to simulate the relationship between residues, and then utilizes graph convolutional networks for enhancing the feature representations for final prediction. ATMGBs was evaluated on several different independent test sets. The results indicate that the proposed approach significantly improves sequence-based prediction performance, even achieving prediction accuracy comparable to structure-based frameworks. The dataset and code used in this study are available at https://github.com/lixiangli01/ATMGBs. Xiang Li 0207, Xiaolei Zhu 0001 |
Briefings Bioinform. | 3 |
| 2025 | AVPpred-BWR: antiviral peptides prediction via biological words representationabstractMOTIVATION: Antiviral peptides (AVPs) are short chains of amino acids, showing great potential as antiviral drugs. The traditional wisdom (e.g. wet experiments) for identifying the AVPs is time-consuming and laborious, while cutting-edge computational methods are less accurate to predict them. RESULTS: In this article, we propose an AVPs prediction model via biological words representation, dubbed AVPpred-BWR. Based on the fact that the secondary structures of AVPs mainly consist of α-helix and loop, we explore the biological words of 1mer (corresponding to loops) and 4mer (4 continuous residues, corresponding to α-helix). That is, the peptides sequences are decomposed into biological words, and then the concealed sequential information is represented by training the Word2Vec models. Moreover, in order to extract multi-scale features, we leverage a CNN-Transformer framework to process the embeddings of 1mer and 4mer generated by Word2Vec models. To the best of our knowledge, this is the first time to realize the word segmentation of protein primary structure sequences based on the regularity of protein secondary structure. AVPpred-BWR illustrates clear improvements over its competitors on the independent test set (e.g. improvements of 4.6% and 11.0% for AUROC and MCC, respectively, compared to UniDL4BioPep). AVAILABILITY AND IMPLEMENTATION: AVPpred-BWR is publicly available at: https://github.com/zyweizm/AVPpred-BWR or https://zenodo.org/records/14880447 (doi: 10.5281/zenodo.14880447). Zhuoyu Wei, Yongqi Shen, Xiang Tang, Youyi Song, Mingqiang Wei, Xiaolei Zhu 0001 |
Bioinform. | 8 |
| 2025 | DTI-MHAPR: optimized drug-target interaction prediction via PCA-enhanced features and heterogeneous graph attention networksabstractDrug-target interactions (DTIs) are pivotal in drug discovery and development, and their accurate identification can significantly expedite the process. Numerous DTI prediction methods have emerged, yet many fail to fully harness the feature information of drugs and targets or address the issue of feature redundancy. We aim to refine DTI prediction accuracy by eliminating redundant features and capitalizing on the node topological structure to enhance feature extraction. To achieve this, we introduce a PCA-augmented multi-layer heterogeneous graph-based network that concentrates on key features throughout the encoding-decoding phase. Our approach initiates with the construction of a heterogeneous graph from various similarity metrics, which is then encoded via a graph neural network. We concatenate and integrate the resultant representation vectors to merge multi-level information. Subsequently, principal component analysis is applied to distill the most informative features, with the random forest algorithm employed for the final decoding of the integrated data. Our method outperforms six baseline models in terms of accuracy, as demonstrated by extensive experimentation. Comprehensive ablation studies, visualization of results, and in-depth case analyses further validate our framework's efficacy and interpretability, providing a novel tool for drug discovery that integrates multimodal features. Yinbo Liu, Sijian Wen, Xiaolei Zhu 0001, Yongmei Michelle Wang |
BMC Bioinform. | 5 |
| 2023 | miRNA-Disease Association Prediction based on Heterogeneous Graph Transformer with Multi-view similarity and Random Auto-encoderabstractMicroRNAs (miRNAs) are a class of short non-coding single-stranded RNA molecules that play a key role in gene expression regulation. Understanding the association between miRNAs and diseases is crucial for disease diagnosis and treatment. Although the wet experimental methods can be used to determine the associations, they are both laborious and expensive. In this paper, we propose a novel computational method called TWMHGT for predicting the associations between miRNAs and diseases based a two-way Multi-layer Heterogeneous Graph Transformer (MHGT) framework. For the first way, multi-view similarity of miRNAs and diseases is used as the input encodings for MHGT, and for the second way, random auto-encoders is used to generate the input encodings. In each MHGT way, the encodings of each layer are concatenated to obtain comprehensive embeddings of miRNAs and diseases, thereby integrating multiple high-level information. Then, the attention mechanisms are used to fuse the embeddings generated from the two-way MHGT. During the decoding stage, the fused embeddings are decoded to obtain the predicted association matrix by using matrix multiplication. Our model was benchmarked on two datasets and the 5-fold and 10-fold cross-validation results show that TWMHGT outperforms the state-of-the-art methods in terms of AUC, AUPR, accuracy, sensitivity, and specificity. Furthermore, we conducted case studies on three different diseases to validate the predictive performance of TWMHGT. The results show excellent performance in all cases, indicating the potential of TWMHGT in discovering novel miRNA-disease associations. Yinbo Liu, Xiaodi Yan, Xinxin Ren, Gang-Ao Wang, Xiaolei Zhu 0001 |
BIBM | 8 |
| 2022 | BERT-Kcr: prediction of lysine crotonylation sites by a transfer learning method with pre-trained BERT modelsabstractMOTIVATION: As one of the most important post-translational modifications (PTMs), protein lysine crotonylation (Kcr) has attracted wide attention, which involves in important physiological activities, such as cell differentiation and metabolism. However, experimental methods are expensive and time-consuming for Kcr identification. Instead, computational methods can predict Kcr sites in silico with high efficiency and low cost. RESULTS: In this study, we proposed a novel predictor, BERT-Kcr, for protein Kcr sites prediction, which was developed by using a transfer learning method with pre-trained bidirectional encoder representations from transformers (BERT) models. These models were originally used for natural language processing (NLP) tasks, such as sentence classification. Here, we transferred each amino acid into a word as the input information to the pre-trained BERT model. The features encoded by BERT were extracted and then fed to a BiLSTM network to build our final model. Compared with the models built by other machine learning and deep learning classifiers, BERT-Kcr achieved the best performance with AUROC of 0.983 for 10-fold cross validation. Further evaluation on the independent test set indicates that BERT-Kcr outperforms the state-of-the-art model Deep-Kcr with an improvement of about 5% for AUROC. The results of our experiment indicate that the direct use of sequence information and advanced pre-trained models of NLP could be an effective way for identifying PTM sites of proteins. AVAILABILITY AND IMPLEMENTATION: The BERT-Kcr model is publicly available on http://zhulab.org.cn/BERT-Kcr_models/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yanhua Qiao, Xiaolei Zhu 0001, Haipeng Gong |
Bioinform. | 2 |
| 2020 | m5CPred-SVM: a novel method for predicting m5C sites of RNAabstractBACKGROUND: As one of the most common post-transcriptional modifications (PTCM) in RNA, 5-cytosine-methylation plays important roles in many biological functions such as RNA metabolism and cell fate decision. Through accurate identification of 5-methylcytosine (m5C) sites on RNA, researchers can better understand the exact role of 5-cytosine-methylation in these biological functions. In recent years, computational methods of predicting m5C sites have attracted lots of interests because of its efficiency and low-cost. However, both the accuracy and efficiency of these methods are not satisfactory yet and need further improvement. RESULTS: In this work, we have developed a new computational method, m5CPred-SVM, to identify m5C sites in three species, H. sapiens, M. musculus and A. thaliana. To build this model, we first collected benchmark datasets following three recently published methods. Then, six types of sequence-based features were generated based on RNA segments and the sequential forward feature selection strategy was used to obtain the optimal feature subset. After that, the performance of models based on different learning algorithms were compared, and the model based on the support vector machine provided the highest prediction accuracy. Finally, our proposed method, m5CPred-SVM was compared with several existing methods, and the result showed that m5CPred-SVM offered substantially higher prediction accuracy than previously published methods. It is expected that our method, m5CPred-SVM, can become a useful tool for accurate identification of m5C sites. CONCLUSION: In this study, by introducing position-specific propensity related features, we built a new model, m5CPred-SVM, to predict RNA m5C sites of three different species. The result shows that our model outperformed the existing state-of-art models. Our model is available for users through a web server at https://zhulab.ahu.edu.cn/m5CPred-SVM . Yi Xiong 0002, Yinbo Liu, Shoudong Bi, Xiaolei Zhu 0001 |
BMC Bioinform. | 6 |
| 2020 | iPNHOT: a knowledge-based approach for identifying protein-nucleic acid interaction hot spotsabstractAbstract Background The interaction between proteins and nucleic acids plays pivotal roles in various biological processes such as transcription, translation, and gene regulation. Hot spots are a small set of residues that contribute most to the binding affinity of a protein-nucleic acid interaction. Compared to the extensive studies of the hot spots on protein-protein interfaces, the hot spot residues within protein-nucleic acids interfaces remain less well-studied, in part because mutagenesis data for protein-nucleic acids interaction are not as abundant as that for protein-protein interactions. Results In this study, we built a new computational model, iPNHOT, to effectively predict hot spot residues on protein-nucleic acids interfaces. One training data set and an independent test set were collected from dbAMEPNI and some recent literature, respectively. To build our model, we generated 97 different sequential and structural features and used a two-step strategy to select the relevant features. The final model was built based only on 7 features using a support vector machine (SVM). The features include two unique features such as ∆SASsa 1/2 and esp3, which are newly proposed in this study. Based on the cross validation results, our model gave F1 score and AUROC as 0.725 and 0.807 on the subset collected from ProNIT, respectively, compared to 0.407 and 0.670 of mCSM-NA, a state-of-the art model to predict the thermodynamic effects of protein-nucleic acid interaction. The iPNHOT model was further tested on the independent test set, which showed that our model outperformed other methods. Conclusion In this study, by collecting data from a recently published database dbAMEPNI, we proposed a new model, iPNHOT, to predict hotspots on both protein-DNA and protein-RNA interfaces. The results show that our model outperforms the existing state-of-art models. Our model is available for users through a webserver: http://zhulab.ahu.edu.cn/iPNHOT/ . Xiaolei Zhu 0001, Yi Xiong 0002, Julie C. Mitchell |
BMC Bioinform. | 1 |
| 2018 | PseUI: Pseudouridine sites identification based on RNA sequence informationabstractBACKGROUND: Pseudouridylation is the most prevalent type of posttranscriptional modification in various stable RNAs of all organisms, which significantly affects many cellular processes that are regulated by RNA. Thus, accurate identification of pseudouridine (Ψ) sites in RNA will be of great benefit for understanding these cellular processes. Due to the low efficiency and high cost of current available experimental methods, it is highly desirable to develop computational methods for accurately and efficiently detecting Ψ sites in RNA sequences. However, the predictive accuracy of existing computational methods is not satisfactory and still needs improvement. RESULTS: In this study, we developed a new model, PseUI, for Ψ sites identification in three species, which are H. sapiens, S. cerevisiae, and M. musculus. Firstly, five different kinds of features including nucleotide composition (NC), dinucleotide composition (DC), pseudo dinucleotide composition (pseDNC), position-specific nucleotide propensity (PSNP), and position-specific dinucleotide propensity (PSDP) were generated based on RNA segments. Then, a sequential forward feature selection strategy was used to gain an effective feature subset with a compact representation but discriminative prediction power. Based on the selected feature subsets, we built our model by using a support vector machine (SVM). Finally, the generalization of our model was validated by both the jackknife test and independent validation tests on the benchmark datasets. The experimental results showed that our model is more accurate and stable than the previously published models. We have also provided a user-friendly web server for our model at http://zhulab.ahu.edu.cn/PseUI , and a brief instruction for the web server is provided in this paper. By using this instruction, the academic users can conveniently get their desired results without complicated calculations. CONCLUSION: In this study, we proposed a new predictor, PseUI, to detect Ψ sites in RNA sequences. It is shown that our model outperformed the existing state-of-art models. It is expected that our model, PseUI, will become a useful tool for accurate identification of RNA Ψ sites. Zizheng Zhang, Bei Huang, Xiaolei Zhu 0001, Yi Xiong 0002 |
BMC Bioinform. | 5 |
| 2018 | Protein-protein interface hot spots prediction based on a hybrid feature selection strategyabstractBACKGROUND: Hot spots are interface residues that contribute most binding affinity to protein-protein interaction. A compact and relevant feature subset is important for building machine learning methods to predict hot spots on protein-protein interfaces. Although different methods have been used to detect the relevant feature subset from a variety of features related to interface residues, it is still a challenge to detect the optimal feature subset for building the final model. RESULTS: In this study, three different feature selection methods were compared to propose a new hybrid feature selection strategy. This new strategy was proved to effectively reduce the feature space when we were building the prediction models for identifying hotspot residues. It was tested on eighty-two features, both conventional and newly proposed. According to the strategy, combining the feature subsets selected by decision tree and mRMR (maximum Relevance Minimum Redundancy) individually, we were able to build a model with 6 features by using a PSFS (Pseudo Sequential Forward Selection) process. Compared with other state-of-art methods for the independent test set, our model had shown better or comparable predictive performances (with F-measure 0.622 and recall 0.821). Analysis of the 6 features confirmed that our newly proposed feature CNSV_REL1 was important for our model. The analysis also showed that the complementarity between features should be considered as an important aspect when conducting the feature selection. CONCLUSION: In this study, most important of all, a new strategy for feature selection was proposed and proved to be effective in selecting the optimal feature subset for building prediction models, which can be used to predict hot spot residues on protein-protein interfaces. Moreover, two aspects, the generalization of the single feature and the complementarity between features, were proved to be of great importance and should be considered in feature selection methods. Finally, our newly proposed feature CNSV_REL1 had been proved an alternative and effective feature in predicting hot spots by our study. Our model is available for users through a webserver: http://zhulab.ahu.edu.cn/iPPHOT/ . Yanhua Qiao, Yi Xiong 0002, Xiaolei Zhu 0001, Peng Chen 0001 |
BMC Bioinform. | 4 |
| 2016 | DBSI server: DNA binding site identifierabstractUNLABELLED: : Protein-nucleic acid interactions are among the most important intermolecular interactions in the regulation of cellular events. Identifying residues involved in these interactions from protein structure alone is an important challenge. Here we introduce the webserver interface to DNA Binding Site Identifier (DBSI), a powerful structure-based SVM model for the prediction and visualization of DNA binding sites on protein structures. DBSI has been shown to be a top-performing model to predict DNA binding sites on the surface of a protein or peptide and shows promise in predicting RNA binding sites. AVAILABILITY AND IMPLEMENTATION: Server is available at http://dbsi.mitchell-lab.org CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shravan Sukumar, Xiaolei Zhu 0001, Spencer S. Ericksen, Julie C. Mitchell |
Bioinform. | 2 |
| 2015 | Large-scale binding ligand prediction by improved patch-based method Patch-Surfer2.0abstractMOTIVATION: Ligand binding is a key aspect of the function of many proteins. Thus, binding ligand prediction provides important insight in understanding the biological function of proteins. Binding ligand prediction is also useful for drug design and examining potential drug side effects. RESULTS: We present a computational method named Patch-Surfer2.0, which predicts binding ligands for a protein pocket. By representing and comparing pockets at the level of small local surface patches that characterize physicochemical properties of the local regions, the method can identify binding pockets of the same ligand even if they do not share globally similar shapes. Properties of local patches are represented by an efficient mathematical representation, 3D Zernike Descriptor. Patch-Surfer2.0 has significant technical improvements over our previous prototype, which includes a new feature that captures approximate patch position with a geodesic distance histogram. Moreover, we constructed a large comprehensive database of ligand binding pockets that will be searched against by a query. The benchmark shows better performance of Patch-Surfer2.0 over existing methods. AVAILABILITY AND IMPLEMENTATION: http://kiharalab.org/patchsurfer2.0/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaolei Zhu 0001, Yi Xiong 0002, Daisuke Kihara |
Bioinform. | 1 |