EDBT 2026 Demo / reviewers in the wild / expert
Yannan Bin
dblp:203/2071
· DBLP profile ↗
21ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0001-6122-5930ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pepxml: ESM2-based extreme multilabel classification of pathogen-targeted antimicrobial peptidesabstractIn recent years, antimicrobial peptides (AMPs) have attracted interest as potential peptide antibiotic due to their broad-spectrum antibacterial activity and high target specificity. However, existing research on AMP prediction mainly focuses on their functional properties, such as antibacterial, antiviral, and anticancer. This emphasis has created a significant gap in identifying AMPs that specifically target pathogens. Given the large variety of pathogens and the sparsity and imbalance of labels, it is challenging to determine which specific pathogens AMPs can effective against. To address this issue, we present PepXML, a large language model-based tool for extreme multilabel classification of pathogen-targeted AMPs. Our first step involved constructing a benchmark dataset of AMPs and their corresponding targeted pathogens, sourced from public databases. In PepXML, the peptides are embedded using ESM2. Further, clustering on a specifically designed label co-occurrence graph and hard negative sampling were employed to address challenges on data sparsity and label imbalance. To validate the reliability of our predictive results, we conducted molecular docking studies focused on peptide-bilayer membrane interactions and performed molecular dynamics simulations to elucidate the mechanisms of peptide-pathogen interactions. We anticipate that PepXML will be a valuable resource for advancing peptide-based therapeutics. The data and Python codes of the PepXML model are available at https://github.com/YannanBin/PepXML.git. Yannan Bin, Daijun Zhang, Zhiyang Hu, Chun-Gui Xu, Yansen Su |
Briefings Bioinform. | 1 |
| 2025 | AEPMA: peptide-microbe association prediction based on autoevolutionary heterogeneous graph learningabstractThe inappropriate use of antibiotics has precipitated the emergence of multidrug-resistant bacteria, prompting significant interest in antimicrobial peptides (AMPs) as potential alternatives to traditional antibiotics. Given the prohibitive costs and time-consuming nature of biological experiments, computational methods provide an efficient alternative for the development of AMP-based drugs. However, existing computational studies primarily focus on identifying AMPs with antimicrobial activity, lacking a targeted identification of AMPs against specific microbial species. To address this gap, we propose a peptide-microbe association (PMA) prediction framework, termed AEPMA, which is constructed based on an autoevolutionary heterogeneous graph. Within AEPMA, we construct an innovative peptide-microbe-disease network (PMDHAN). Furthermore, we design an autoevolutionary information aggregation mechanism that facilitates the representation learning of the heterogeneous graph. This model automatically aggregates semantic information within the heterogeneous network while thoroughly accounting for the spatiotemporal dependencies and heterogeneous interactions in the PMDHAN. Experiments conducted on one peptide-microbe and three drug-microbe association datasets demonstrate that the performance of AEPMA outperforms five state-of-the-art methods, demonstrating its robust modeling capability and exceptional generalization ability. In addition, this study identifies a novel anti-Staphylococcus aureus peptide and an anti-Escherichia coli peptide, thereby contributing valuable information for the development of antimicrobial drugs and strategies for mitigating antibiotic resistance. Zhiyang Hu, Linqiang Pan, Daijun Zhang, Yannan Bin, Yansen Su |
Briefings Bioinform. | 4 |
| 2024 | AMGDTI: drug-target interaction prediction based on adaptive meta-graph learning in heterogeneous networkabstractPrediction of drug-target interactions (DTIs) is essential in medicine field, since it benefits the identification of molecular structures potentially interacting with drugs and facilitates the discovery and reposition of drugs. Recently, much attention has been attracted to network representation learning to learn rich information from heterogeneous data. Although network representation learning algorithms have achieved success in predicting DTI, several manually designed meta-graphs limit the capability of extracting complex semantic information. To address the problem, we introduce an adaptive meta-graph-based method, termed AMGDTI, for DTI prediction. In the proposed AMGDTI, the semantic information is automatically aggregated from a heterogeneous network by training an adaptive meta-graph, thereby achieving efficient information integration without requiring domain knowledge. The effectiveness of the proposed AMGDTI is verified on two benchmark datasets. Experimental results demonstrate that the AMGDTI method overall outperforms eight state-of-the-art methods in predicting DTI and achieves the accurate identification of novel DTIs. It is also verified that the adaptive meta-graph exhibits flexibility and effectively captures complex fine-grained semantic information, enabling the learning of intricate heterogeneous network topology and the inference of potential drug-target relationship. Yansen Su, Zhiyang Hu, Fei Wang 0095, Yannan Bin, Chun-Hou Zheng 0001, Haitao Li 0004, Xiangxiang Zeng |
Briefings Bioinform. | 4 |
| 2023 | Generative Adversarial Network-Based Data Augmentation Method for Anti-coronavirus Peptides Prediction
Jiliang Xu, Chun-Gui Xu, Yonghui He, Yannan Bin, Chun-Hou Zheng 0001 |
ICIC (3) | 5 |
| 2023 | FFMAVP: a new classifier based on feature fusion and multitask learning for identifying antiviral peptides and their subclassesabstractAntiviral peptides (AVPs) are widely found in animals and plants, with high specificity and strong sensitivity to drug-resistant viruses. However, due to the great heterogeneity of different viruses, most of the AVPs have specific antiviral activities. Therefore, it is necessary to identify the specific activities of AVPs on virus types. Most existing studies only identify AVPs, with only a few studies identifying subclasses by training multiple binary classifiers. We develop a two-stage prediction tool named FFMAVP that can simultaneously predict AVPs and their subclasses. In the first stage, we identify whether a peptide is AVP or not. In the second stage, we predict the six virus families and eight species specifically targeted by AVPs based on two multiclass tasks. Specifically, the feature extraction module in the two-stage task of FFMAVP adopts the same neural network structure, in which one branch extracts features based on amino acid feature descriptors and the other branch extracts sequence features. Then, the two types of features are fused for the following task. Considering the correlation between the two tasks of the second stage, a multitask learning model is constructed to improve the effectiveness of the two multiclass tasks. In addition, to improve the effectiveness of the second stage, the network parameters trained through the first-stage data are used to initialize the network parameters in the second stage. As a demonstration, the cross-validation results, independent test results and visualization results show that FFMAVP achieves great advantages in both stages. Weiling Hu, Pi-Jing Wei, Yun Ding, Yannan Bin, Chun-Hou Zheng 0001 |
Briefings Bioinform. | 5 |
| 2023 | Deep learning-based multi-functional therapeutic peptides prediction with a multi-label focal dice loss functionabstractMOTIVATION: With the great number of peptide sequences produced in the postgenomic era, it is highly desirable to identify the various functions of therapeutic peptides quickly. Furthermore, it is a great challenge to predict accurate multi-functional therapeutic peptides (MFTP) via sequence-based computational tools. RESULTS: Here, we propose a novel multi-label-based method, named ETFC, to predict 21 categories of therapeutic peptides. The method utilizes a deep learning-based model architecture, which consists of four blocks: embedding, text convolutional neural network, feed-forward network, and classification blocks. This method also adopts an imbalanced learning strategy with a novel multi-label focal dice loss function. multi-label focal dice loss is applied in the ETFC method to solve the inherent imbalance problem in the multi-label dataset and achieve competitive performance. The experimental results state that the ETFC method is significantly better than the existing methods for MFTP prediction. With the established framework, we use the teacher-student-based knowledge distillation to obtain the attention weight from the self-attention mechanism in the MFTP prediction and quantify their contributions toward each of the investigated activities. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available via: https://github.com/xialab-ahu/ETFC. Henghui Fan, Wenhui Yan, Yannan Bin, Junfeng Xia |
Bioinform. | 5 |
| 2023 | PhaGAA: an integrated web server platform for phage genome annotation and analysisabstractMOTIVATION: Phage genome annotation plays a key role in the design of phage therapy. To date, there have been various genome annotation tools for phages, but most of these tools focus on mono-functional annotation and have complex operational processes. Accordingly, comprehensive and user-friendly platforms for phage genome annotation are needed. RESULTS: Here, we propose PhaGAA, an online integrated platform for phage genome annotation and analysis. By incorporating several annotation tools, PhaGAA is constructed to annotate the prophage genome at DNA and protein levels and provide the analytical results. Furthermore, PhaGAA could mine and annotate phage genomes from bacterial genome or metagenome. In summary, PhaGAA will be a useful resource for experimental biologists and help advance the phage synthetic biology in basic and application research. AVAILABILITY AND IMPLEMENTATION: PhaGAA is freely available at http://phage.xialab.info/. Qingrui Liu, Jiliang Xu, Junyin Zhang, Minfeng Xiao, Yannan Bin, Junfeng Xia |
Bioinform. | 8 |
| 2023 | PACVP: Prediction of Anti-Coronavirus Peptides Using a Stacking Learning Strategy With Effective Feature RepresentationabstractDue to the global outbreak of COVID-19 and its variants, antiviral peptides with anti-coronavirus activity (ACVPs) represent a promising new drug candidate for the treatment of coronavirus infection. At present, several computational tools have been developed to identify ACVPs, but the overall prediction performance is still not enough to meet the actual therapeutic application. In this study, we constructed an efficient and reliable prediction model PACVP (Prediction of Anti-CoronaVirus Peptides) for identifying ACVPs based on effective feature representation and a two-layer stacking learning framework. In the first layer, we use nine feature encoding methods with different feature representation angles to characterize the rich sequence information and fuse them into a feature matrix. Secondly, data normalization and unbalanced data processing are carried out. Next, 12 baseline models are constructed by combining three feature selection methods and four machine learning classification algorithms. In the second layer, we input the optimal probability features into the logistic regression algorithm (LR) to train the final model PACVP. The experiments show that PACVP achieves favorable prediction performance on independent test dataset, with ACC of 0.9208 and AUC of 0.9465. We hope that PACVP will become a useful method for identifying, annotating and characterizing novel ACVPs. Shouzhi Chen, Yanhong Liao, Jianping Zhao 0001, Yannan Bin, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | An Ensemble Framework Integrating Whole Slide Pathological Images and miRNA Data to Predict Radiosensitivity of Breast Cancer Patients
Wenhui Yan, Mengmeng Han, Junfeng Xia, Yannan Bin |
ICIC (2) | 7 |
| 2022 | NeuroPred-CLQ: incorporating deep temporal convolutional networks and multi-head attention mechanism to predict neuropeptidesabstractNeuropeptides (NPs) are a particular class of informative substances in the immune system and physiological regulation. They play a crucial role in regulating physiological functions in various biological growth and developmental stages. In addition, NPs are crucial for developing new drugs for the treatment of neurological diseases. With the development of molecular biology techniques, some data-driven tools have emerged to predict NPs. However, it is necessary to improve the predictive performance of these tools for NPs. In this study, we developed a deep learning model (NeuroPred-CLQ) based on the temporal convolutional network (TCN) and multi-head attention mechanism to identify NPs effectively and translate the internal relationships of peptide sequences into numerical features by the Word2vec algorithm. The experimental results show that NeuroPred-CLQ learns data information effectively, achieving 93.6% accuracy and 98.8% AUC on the independent test set. The model has better performance in identifying NPs than the state-of-the-art predictors. Visualization of features using t-distribution random neighbor embedding shows that the NeuroPred-CLQ can clearly distinguish the positive NPs from the negative ones. We believe the NeuroPred-CLQ can facilitate drug development and clinical trial studies to treat neurological disorders. Shouzhi Chen, Jianping Zhao 0001, Yannan Bin, Chun-Hou Zheng 0001 |
Briefings Bioinform. | 4 |
| 2022 | Identifying multi-functional bioactive peptide functions using multi-label deep learningabstractThe bioactive peptide has wide functions, such as lowering blood glucose levels and reducing inflammation. Meanwhile, computational methods such as machine learning are becoming more and more important for peptide functions prediction. Most of the previous studies concentrate on the single-functional bioactive peptides prediction. However, the number of multi-functional peptides is on the increase; therefore, novel computational methods are needed. In this study, we develop a method MLBP (Multi-Label deep learning approach for determining the multi-functionalities of Bioactive Peptides), which can predict multiple functions including anti-cancer, anti-diabetic, anti-hypertensive, anti-inflammatory and anti-microbial simultaneously. MLBP model takes the peptide sequence vector as input to replace the biological and physiochemical features used in other peptides predictors. Using the embedding layer, the dense continuous feature vector is learnt from the sequence vector. Then, we extract convolution features from the feature vector through the convolutional neural network layer and combine with the bidirectional gated recurrent unit layer to improve the prediction performance. The 5-fold cross-validation experiments are conducted on the training dataset, and the results show that Accuracy and Absolute true are 0.695 and 0.685, respectively. On the test dataset, Accuracy and Absolute true of MLBP are 0.709 and 0.697, with 5.0 and 4.7% higher than those of the suboptimum method, respectively. The results indicate MLBP has superior prediction performance on the multi-functional peptides identification. MLBP is available at https://github.com/xialab-ahu/MLBP and http://bioinfo.ahu.edu.cn/MLBP/. Wending Tang, Ruyu Dai, Wenhui Yan, Yannan Bin, En-Hua Xia, Junfeng Xia |
Briefings Bioinform. | 5 |
| 2022 | PrMFTP: Multi-functional therapeutic peptides prediction based on multi-head self-attention mechanism and class weight optimizationabstractPrediction of therapeutic peptide is a significant step for the discovery of promising therapeutic drugs. Most of the existing studies have focused on the mono-functional therapeutic peptide prediction. However, the number of multi-functional therapeutic peptides (MFTP) is growing rapidly, which requires new computational schemes to be proposed to facilitate MFTP discovery. In this study, based on multi-head self-attention mechanism and class weight optimization algorithm, we propose a novel model called PrMFTP for MFTP prediction. PrMFTP exploits multi-scale convolutional neural network, bi-directional long short-term memory, and multi-head self-attention mechanisms to fully extract and learn informative features of peptide sequence to predict MFTP. In addition, we design a class weight optimization scheme to address the problem of label imbalanced data. Comprehensive evaluation demonstrate that PrMFTP is superior to other state-of-the-art computational methods for predicting MFTP. We provide a user-friendly web server of PrMFTP, which is available at http://bioinfo.ahu.edu.cn/PrMFTP. Wenhui Yan, Wending Tang, Yannan Bin, Junfeng Xia |
PLoS Comput. Biol. | 4 |
| 2022 | DPProm: A Two-Layer Predictor for Identifying Promoters and Their Types on Phage Genome Using Deep LearningabstractWith the number of phage genomes increasing, it is urgent to develop new bioinformatics methods for phage genome annotation. Promoter, a DNA region, is important for gene transcriptional regulation. In the era of post-genomics, the availability of data makes it possible to establish computational models for promoter identification with robustness. In this work, we introduce DPProm, a two-layer model composed of DPProm-1L and DPProm-2L, to predict promoters and their types for phages. On the first layer, as a dual-channel deep neural network ensemble method fusing multi-view features (sequence feature and handcrafted feature), the model DPProm-1L is proposed to identify whether a DNA sequence is a promoter or non-promoter. The sequence feature is extracted with convolutional neural network (CNN). And the handcrafted feature is the combination of free energy, GC content, cumulative skew, and Z curve features. On the second layer, DPProm-2L based on CNN is trained to predict the promoters' types (host or phage). For the realization of prediction on the whole genomes, the model DPProm, combines with a novel sequence data processing workflow, which contains sliding window and merging sequences modules. Experimental results show that DPProm outperforms the state-of-the-art methods, and decreases the false positive rate effectively on whole genome prediction. Furthermore, we provide a user-friendly web at http://bioinfo.ahu.edu.cn/DPProm. We expect that DPProm can serve as a useful tool for identification of promoters and their types. Junyin Zhang, Minfeng Xiao, Junfeng Xia, Yannan Bin |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | An improved DNA-binding hot spot residues prediction method by exploring interfacial neighbor propertiesabstractBACKGROUND: DNA-binding hot spots are dominant and fundamental residues that contribute most of the binding free energy yet accounting for a small portion of protein-DNA interfaces. As experimental methods for identifying hot spots are time-consuming and costly, high-efficiency computational approaches are emerging as alternative pathways to experimental methods. RESULTS: Herein, we present a new computational method, termed inpPDH, for hot spot prediction. To improve the prediction performance, we extract hybrid features which incorporate traditional features and new interfacial neighbor properties. To remove redundant and irrelevant features, feature selection is employed using a two-step feature selection strategy. Finally, a subset of 7 optimal features are chosen to construct the predictor using support vector machine. The results on the benchmark dataset show that this proposed method yields significantly better prediction accuracy than those previously published methods in the literature. Moreover, a user-friendly web server for inpPDH is well established and is freely available at http://bioinfo.ahu.edu.cn/inpPDH . CONCLUSIONS: We have developed an accurate improved prediction model, inpPDH, for hot spot residues in protein-DNA binding interfaces by given the structure of a protein-DNA complex. Moreover, we identify a comprehensive and useful feature subset including the proposed interfacial neighbor features that has an important strength for identifying hot spot residues. Our results indicate that these features are more effective than the conventional features considered previously, and that the combination of interfacial neighbor features and traditional features may support the creation of a discriminative feature set for efficient prediction of hot spot residues in protein-DNA complexes. Menglu Li, Yannan Bin, Junfeng Xia |
BMC Bioinform. | 7 |
| 2021 | A Deep Learning-Based Method for Identification of Bacteriophage-Host InteractionabstractMulti-drug resistance (MDR) has become one of the greatest threats to human health worldwide, and novel treatment methods of infections caused by MDR bacteria are urgently needed. Phage therapy is a promising alternative to solve this problem, to which the key is correctly matching target pathogenic bacteria with the corresponding therapeutic phage. Deep learning is powerful for mining complex patterns to generate accurate predictions. In this study, we develop PredPHI (Predicting Phage-Host Interactions), a deep learning-based tool capable of predicting the host of phages from sequence data. We collect >3000 phage-host pairs along with their protein sequences from PhagesDB and GenBank databases and extract a set of features. Then we select high-quality negative samples based on the K-Means clustering method and construct a balanced training set. Finally, we employ a deep convolutional neural network to build the predictive model. The results indicate that PredPHI can achieve a predictive performance of 81 percent in terms of the area under the receiver operating characteristic curve on the test set, and the clustering-based method is significantly more robust than that based on randomly selecting negative samples. These results highlight that PredPHI is a useful and accurate tool for identifying phage-host interactions from sequence data. Menglu Li, Yanan Wang 0003, Fuyi Li, Yun Zhao 0004, Yannan Bin, Alexander Ian Smith, Geoffrey I. Webb, Jian Li 0052, Jiangning Song, Junfeng Xia |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2020 | Prediction of hot spots in protein-DNA binding interfaces based on supervised isometric feature mapping and extreme gradient boostingabstractBACKGROUND: Identification of hot spots in protein-DNA interfaces provides crucial information for the research on protein-DNA interaction and drug design. As experimental methods for determining hot spots are time-consuming, labor-intensive and expensive, there is a need for developing reliable computational method to predict hot spots on a large scale. RESULTS: Here, we proposed a new method named sxPDH based on supervised isometric feature mapping (S-ISOMAP) and extreme gradient boosting (XGBoost) to predict hot spots in protein-DNA complexes. We obtained 114 features from a combination of the protein sequence, structure, network and solvent accessible information, and systematically assessed various feature selection methods and feature dimensionality reduction methods based on manifold learning. The results show that the S-ISOMAP method is superior to other feature selection or manifold learning methods. XGBoost was then used to develop hot spots prediction model sxPDH based on the three dimensionality-reduced features obtained from S-ISOMAP. CONCLUSION: Our method sxPDH boosts prediction performance using S-ISOMAP and XGBoost. The AUC of the model is 0.773, and the F1 score is 0.713. Experimental results on benchmark dataset indicate that sxPDH can achieve generally better performance in predicting hot spots compared to the state-of-the-art methods. Yannan Bin, Junfeng Xia |
BMC Bioinform. | 4 |
| 2019 | Distinguishing Driver Missense Mutations from Benign Polymorphisms in Breast Cancer
Ruoqing Xu, Yannan Bin |
ICIC (2) | 3 |
| 2018 | Further Evidence for Role of Promoter Polymorphisms in TNF Gene in Alzheimer's Disease
Yannan Bin, Ling Shu, Qizhi Zhu, Huanhuan Zhu, Junfeng Xia |
ICIC (2) | 1 |
| 2018 | Nucleotide-Based Significance of Somatic Synonymous Mutations for Pan-Cancer
Yannan Bin, Qizhi Zhu, Pengbo Wen, Junfeng Xia |
ICIC (2) | 1 |
| 2018 | Computational Prediction of Driver Missense Mutations in Melanoma
Junfeng Xia, Yannan Bin, Di Zhang 0006 |
ICIC (2) | 5 |
| 2017 | Investigating Alzheimer's Disease Candidate Genes Based on Combined Network Using Subnetwork Extraction Algorithms
Di Zhang 0006, Yannan Bin, Junfeng Xia |
ICIC (2) | 5 |