EDBT 2026 Demo / reviewers in the wild / expert
Yun Zuo 0001
dblp:15/10851-1
· DBLP profile ↗
18ranked-venue papers
8as first author
16since 2021 · last 2026
0009-0009-3877-8102ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FKSUDDAPre: A drug-disease association prediction framework based on F-TEST feature selection and AMDKSU resampling with interpretability analysisabstractIn drug discovery and therapeutic research, the prediction of drug-disease associations (DDAs) holds significant scientific and clinical value. Drug molecules exert their effects by precisely identifying disease-related biological targets, systematically modulating the entire pharmacological process from absorption, distribution, and metabolism to final efficacy. Accurate prediction of drug-disease associations not only facilitates an in-depth understanding of molecular mechanisms of drug action but also provides critical theoretical foundations for drug repositioning and personalized medicine. While traditional prediction methods based on in vitro experiments and clinical statistics yield reliable results, they suffer from inherent drawbacks such as long development cycles, substantial resource consumption, and low throughput. In contrast, emerging machine learning techniques offer a promising solution to these bottlenecks, enabling the intelligent and efficient discovery of potential drug-disease association networks and significantly improving drug development efficiency. However, it is noteworthy that existing machine learning methods still face significant challenges in practical applications: the complexity of feature construction raises the threshold for data processing; data sparsity constrains the depth of information mining; and the pervasive issue of sample imbalance poses a severe challenge to the model's predictive accuracy and generalization performance. In this study, we developed an efficient and accurate framework for drug-disease association prediction named FKSUDDAPre. The model employs a multi-modal feature fusion strategy: on one hand, it leverages an ensemble of Mol2vec and K- BERT to deeply capture the semantic features of drug molecular fingerprints; on the other hand, it integrates Medical Subject Headings (MeSH) with DeepWalk to effectively reduce the dimensionality of disease features while preserving their relational structure. To address the class imbalance problem, FKSUDDAPre designed an optimization algorithm called AMDKSU, which combined clustering with an improved distance metric strategy, significantly enhancing the discriminative power of the sample set. For data processing, F-test was employed for feature importance ranking, effectively reducing data dimensionality and improving model generalization. For the predictive architecture, FKSUDDAPre proposed a novel ensemble framework composed of XGBoost, Decision Tree, Random Forest, and HyperFast. By employing a dynamic weight allocation strategy, this ensemble effectively harnesses the complementary strengths of these models to achieve significantly enhanced predictive performance. Rigorous validation demonstrated the system's outstanding performance across multiple evaluation metrics, with an average AUC of 0.9725, improving the AUC by approximately 3.88% compared to the best-performing baseline model. In the prediction of Alzheimer's disease and Parkinson's disease, 80% and 60% of the top 10 candidate drugs recommended by FKSUDDAPre, respectively, had been confirmed by literature, demonstrating the model's good practical application potential. Furthermore, we conducted a LIME-based feature importance analysis on the model's predictions, visualizing the correlations between features and the target variable to demonstrate the model's interpretability. A cross-platform, user-friendly visualization tool had also been developed using the PyQt5 framework. Yun Zuo 0001, Ge Hua, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
PLoS Comput. Biol. | 1 |
| 2025 | Catching mRNA's Hiddens Marks: A Dual-Path Network by Contrastive Learning for N4-acetylcytidine PredictionabstractN4-acetylcytidine (ac4C) is a crucial RNA modification associated with mRNA stability and translational efficiency. Accurate identification of ac4C sites is essential for understanding their regulatory functions. However, experimental detection remains expensive and labor-intensive. At the same time, existing computational models suffer from limited generalization and insufficient feature discrimination, especially in distinguishing subtle nucleotide patterns. In this work, we propose a deep learning model named SNN-ac4C, which is based on a contrastive learning-based neural network. The model integrates a dual-path structure that combines BiLSTM and Multi-Head Self-Attention (MHSA) for capturing long-range dependencies and global context, while using CNN to extract local biological sequence features. The contrastive learning module further enhances the discriminative ability of ac4C and Non-ac4C sites by increasing the separation between positive and negative samples. Experiments on the test set confirm the effectiveness of SNN-ac4C, which achieves an accuracy (ACC) of 84.60% and a Matthews Correlation Coefficient (MCC) of 0.6934. Compared with NBCR-ac4C, the current state-of-the-art model, SNN-ac4C improves ACC and MCC by 1.09% and 0.0228, respectively. The source code and relevant supplementary are publicly available at https://github.com/2103374200/SNN. Wenying He, Haolu Zhou, Yun Zuo 0001, Yude Bai, Fei Guo 0001 |
ECAI | 4 |
| 2025 | DCFICSH: A Dual-Channel Fusion Model Combining Multi-Modal Data for Identifying Cell-Specific Silencers and Their Strength in the Human Genome
Jingdong Yuan, Qinqin Zhu, Haolu Zhou, Yun Zuo 0001, Yude Bai, Wenying He |
ICIC (19) | 5 |
| 2025 | MlyPredCSED: based on extreme point deviation compensated clustering combined with cross-scale convolutional neural networks to predict multiple lysine sites in humanabstractIn post-translational modification, covalent bonds on lysine and attached chemical groups significantly change proteins' physical and chemical properties. They shape protein structures, enhance function and stability, and are vital for physiological processes, affecting health and disease through mechanisms like gene expression, signal transduction, protein degradation, and cell metabolism. Although lysine (K) modification sites are considered among the most common types of post-translational modifications in proteins, research on K-PTMs has largely overlooked the synergistic effects between different modifications and lacked the techniques to address the problem of sample imbalance. Based on this, the Extreme Point Deviation Compensated Clustering (EPDCC) Undersampling algorithm was proposed in this study and combined with Cross-Scale Convolutional Neural Networks (CSCNNs) to develop a novel computational tool, MlyPredCSED, for simultaneously predicting multiple lysine modification sites. MlyPredCSED employs Multi-Label Position-Specific Triad Amino Acid Propensity and the physicochemical properties of amino acids to enhance the richness of sequence information. To address the challenge of sample imbalance, the innovative EPDCC Undersampling technique was introduced to adjust the majority class samples. The model's training and testing phase relies on the advanced CSCNN framework. MlyPredCSED, through cross-validation and testing, outperformed existing models, especially in complex categories with multiple modification sites. This research not only provides an efficient method for the identification of lysine modification sites but also demonstrates its value in biological research and drug development. To facilitate efficient use of MlyPredCSED by researchers, we have specifically developed an accessible free web tool: http://www.mlypredcsed.com. Yun Zuo 0001, Xingze Fang, Jiankang Chen, Jiayi Ji, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng, Hongwei Yin, Anjing Zhao |
Briefings Bioinform. | 1 |
| 2025 | HyperACP: A cutting-edge hybrid framework for anticancer peptide classification via scalable feature extraction and adaptive neighbor-based synthesisabstractCancer remains a major contributor to global mortality, constituting a significant and escalating threat to human health. Anticancer peptides (ACPs) have emerged as promising therapeutic agents due to their specific mechanisms of action, pronounced tumor-targeting capability, and low toxicity. Nevertheless, traditional approaches for ACP identification are constrained by their reliance on shallow, hand-crafted sequence features, which fail to capture deeper semantic and structural characteristics. Moreover, such models exhibit limited robustness and interpretability when confronted with practical challenges such as severe class imbalance. To address these limitations, this study proposes HyperACP, an innovative framework for ACP recognition that integrates deep representation learning, adaptive sampling, and mechanistic interpretability. The framework leverages the ESMC protein language model to extract comprehensive sequence features and employs a novel adaptive algorithm, ANBS, to mitigate class imbalance at the decision boundary. For enhanced model transparency, SHAP-Res is incorporated to elucidate the contributions of individual residues to the final predictions. Comprehensive evaluations demonstrate that HyperACP consistently outperforms state-of-the-art methods across multiple datasets and validation protocols-including 10-fold cross-validation and independent test sets-according to metrics such as Accuracy (ACC), Sensitivity (SN), Specificity (SP), Matthews Correlation Coefficient (MCC), and Area Under the Curve (AUC). Furthermore, the model yields biologically interpretable results, pinpointing key residues (K, L, F, G) known to play pivotal roles in anticancer activity. These findings provide not only a robust predictive tool (available at www.hyperacp.com) but also novel insights into the structure-function relationships underlying ACPs. Bangyi Zhang, Yun Zuo 0001, Jiayue Liu, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
PLoS Comput. Biol. | 2 |
| 2025 | CATransUnetLBP: Accurate Prediction of Protein-Ligand Binding Pockets Using a Hybrid NetworkabstractThe development of intelligent methods capable of predicting protein-ligand binding sites has become a popular research field. Recently, deep learning based methods have been proposed as a promising solution for this task. However, some limitations still exist. For example, the network structure is not optimized for predicting protein binding pockets, which limits the model's capabilities. To address the aforementioned challenges, a novel method called CATransUnetLPB is proposed, in which a new network structure named CATransUnet is designed. The proposed CATransUnet combines CNN and Transformer models to accurately segment binding pocket regions from protein 3D structures. It outperforms existing representative methods on three test sets, demonstrating the effectiveness of optimizing the deep network model for detecting protein ligand binding pockets. Furthermore, we conduct thorough analysis on applying data augmentation to protein data structure and confirm that such technique can enhance the model's generalization ability, thereby ensuring good performance on new protein structures. Moreover, experiments show that the predicted binding pockets from our model can complement the results obtained from other methods. This suggests that integrating our method with existing approaches could further improve the prediction of protein-ligand binding pockets. Cheng Cai, Zhaohong Deng, Andong Li, Yun Zuo 0001, Haoran Chen 0003, Zhisheng Wei, Xiaoyong Pan, Hong-Bin Shen, Dongjun Yu |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | DMMAFS: Protein Function Prediction Based on Multi-Modal Multi-Attention Fusion FeaturesabstractIntelligent prediction of protein function is more efficient and less resource-consuming and has achieved significant progress in recent years. However, most of the current methods are performed solely based on the sequence information of proteins. These methods overlook information of other modalities that the proteins themselves possess, which makes it difficult to achieve the desired predicted results. Furthermore, a few existing methods based on multiple modal information fuse them in a simple splicing manner and fail to fully exploit the complementary relation between different modalities. To address the above-mentioned challenges, we propose Multi-modal Multi-attention fusion Features (DMMAFS), a method based on deep learning, to predict protein function. On the one hand, DMMAFS gains the semantic information embedded in the sequence itself through the self-attention learning of the sequence. On the other hand, DMMAFS employs the 3D structural information of proteins to compensate for the sequence information. Particularly, a S-C cross-modal cross-attention fusion network module is proposed that not only optimizes the weights of the semantic information but also efficiently fuses the sequence features with the structural information, thus avoiding the simple splicing of different modal features. Our experimental results demonstrate that the proposed DMMAFS outperforms the state-of-the-art methods in protein function prediction. Liangwen He, Zhaohong Deng, Fuping Hu, Yun Zuo 0001, Haoran Chen 0003, Xiaoyong Pan, Zhisheng Wei, Hong-Bin Shen, Dongjun Yu, Jing Wu 0030 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | MBSCLoc: Multi-Label Subcellular Localization Predict Based on Cluster Balanced Subspace Partitioning Method and Multi-Class Contrastive Representation LearningabstractmRNA subcellular localization is a prevalent and essential mechanism that precisely regulates protein translation and significantly impacts various cellular processes. mRNA subcellular localization has advanced the understanding of mRNA function, yet existing methods face limitations, including imbalanced data, suboptimal model performance, and inadequate generalization, particularly in multi-label localization scenarios where solutions are scarce. This study introduces MBSCLoc, a predictor for mRNA multi-label subcellular localization. MBSCLoc predicts mRNA locations across multiple cellular compartments simultaneously, overcoming challenges like single-location prediction, incomplete feature extraction, and imbalanced data. MBSCLoc leverages UTR-LM model for feature extraction, followed by multi-class contrastive representation learning and Clustering Balanced Subspace Partitioning to construct balanced subspaces. It then optimizes sample distribution to tackle severe data imbalance and uses multiple XGBoost classifiers, integrated through voting, to enhance accuracy and generalization. Five-fold cross-validation and independent testing results show that MBSCLoc significantly outperforms other methods. Additionally, MBSCLoc offers superior pixel-level interpretability, strongly supporting mRNA multi-label subcellular localization research. Crucially, the importance of the 5' UTR and 3' UTR regions has been preliminarily confirmed using traditional biological analysis and Tree-SHAP, with most mRNA sequences showing significant relevance in these regions, especially the 3' UTR where about 80% of specific sites reach peak significance. Bangyi Zhang, Yun Zuo 0001, Zhiqiang Dai, Sifan Zhu, Zhaohong Deng |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | LM-TLIPs: Integrating Large Model and Transfer Learning Technology for Precise Identification of Phosphorylation Sites in SARS-CoV-2abstractIn recent years, the rapid spread of SARS-CoV-2 has triggered a global health crisis and socio-economic challenges. As a crucial post-translational modification, phosphorylation plays a vital role in the regulation of cellular functions. Given its close relationship with SARS-CoV-2 infection, accurately identifying virus-induce phosphorylation sites is essential for understanding the molecular mechanisms of viral infection and its impact on host cells. Although the development of various computational tools for predicting phosphorylation sites, these tools have several shortcomings, such as insufficient data and limited model generalization ability, which limit their effectiveness in practical applications. To overcome these limitations, this study proposes a novel method for predicting SARS-CoV-2 phosphorylation sites, LM-TLIPs, based on the latest technology. This method uses the most advanced large model technology ESM-2 to extract information from S/T sites and Y sites; by fine-tuning the large model and introducing transfer learning technology, it addresses the challenge of accurately predicting Y sites due to insufficient data in this study. Independent testing on S/T sites(Acc:0.8309, Sn:0.8443, Sp:0.8174, MCC:0.6620, AUC:0.8993) and Y sites(Acc:0.9048, Sn:0.9524, Sp:0.8571, MCC:0.8132, AUC:0.9388) has validated that LM-TLIPs outperforms existing optimal prediction tools, demonstrating its superior ability in identifying phosphorylation sites. Furthermore, we conducted an exhaustive interpretability analysis based on attention weight heatmaps and feature importance ranking to enhance the transparency and confidence of prediction results. Yun Zuo 0001, Minquan Wan, Xinyue Shao, Dandan Qiao, Bulanni Xiong, Zhaohong Deng |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | MSlocPRED: deep transfer learning-based identification of multi-label mRNA subcellular localizationabstractSubcellular localization of messenger ribonucleic acid (mRNA) is a universal mechanism for precise and efficient control of the translation process. Although many computational methods have been constructed by researchers for predicting mRNA subcellular localization, very few of these computational methods have been designed to predict subcellular localization with multiple localization annotations, and their generalization performance could be improved. In this study, the prediction model MSlocPRED was constructed to identify multi-label mRNA subcellular localization. First, the preprocessed Dataset 1 and Dataset 2 are transformed into the form of images. The proposed MDNDO-SMDU resampling technique is then used to balance the number of samples in each category in the training dataset. Finally, deep transfer learning was used to construct the predictive model MSlocPRED to identify subcellular localization for 16 classes (Dataset 1) and 18 classes (Dataset 2). The results of comparative tests of different resampling techniques show that the resampling technique proposed in this study is more effective in preprocessing for subcellular localization. The prediction results of the datasets constructed by intercepting different NC end (Both the 5' and 3' untranslated regions that flank the protein-coding sequence and influence mRNA function without encoding proteins themselves.) lengths show that for Dataset 1 and Dataset 2, the prediction performance is best when the NC end is intercepted by 35 nucleotides, respectively. The results of both independent testing and five-fold cross-validation comparisons with established prediction tools show that MSlocPRED is significantly better than established tools for identifying multi-label mRNA subcellular localization. Additionally, to understand how the MSlocPRED model works during the prediction process, SHapley Additive exPlanations was used to explain it. The predictive model and associated datasets are available on the following github: https://github.com/ZBYnb1/MSlocPRED/tree/main. Yun Zuo 0001, Bangyi Zhang, Wenying He, Yue Bi, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
Briefings Bioinform. | 1 |
| 2024 | MINDG: a drug-target interaction prediction method based on an integrated learning algorithmabstractMOTIVATION: Drug-target interaction (DTI) prediction refers to the prediction of whether a given drug molecule will bind to a specific target and thus exert a targeted therapeutic effect. Although intelligent computational approaches for drug target prediction have received much attention and made many advances, they are still a challenging task that requires further research. The main challenges are manifested as follows: (i) most graph neural network-based methods only consider the information of the first-order neighboring nodes (drug and target) in the graph, without learning deeper and richer structural features from the higher-order neighboring nodes. (ii) Existing methods do not consider both the sequence and structural features of drugs and targets, and each method is independent of each other, and cannot combine the advantages of sequence and structural features to improve the interactive learning effect. RESULTS: To address the above challenges, a Multi-view Integrated learning Network that integrates Deep learning and Graph Learning (MINDG) is proposed in this study, which consists of the following parts: (i) a mixed deep network is used to extract sequence features of drugs and targets, (ii) a higher-order graph attention convolutional network is proposed to better extract and capture structural features, and (iii) a multi-view adaptive integrated decision module is used to improve and complement the initial prediction results of the above two networks to enhance the prediction performance. We evaluate MINDG on two dataset and show it improved DTI prediction performance compared to state-of-the-art baselines. AVAILABILITY AND IMPLEMENTATION: https://github.com/jnuaipr/MINDG. Hailong Yang 0001, Yun Zuo 0001, Zhaohong Deng, Xiaoyong Pan, Hong-Bin Shen, Kup-Sze Choi, Dongjun Yu |
Bioinform. | 3 |
| 2024 | PreMLS: The undersampling technique based on ClusterCentroids to predict multiple lysine sitesabstractThe translated protein undergoes a specific modification process, which involves the formation of covalent bonds on lysine residues and the attachment of small chemical moieties. The protein's fundamental physicochemical properties undergo a significant alteration. The change significantly alters the proteins' 3D structure and activity, enabling them to modulate key physiological processes. The modulation encompasses inhibiting cancer cell growth, delaying ovarian aging, regulating metabolic diseases, and ameliorating depression. Consequently, the identification and comprehension of post-translational lysine modifications hold substantial value in the realms of biological research and drug development. Post-translational modifications (PTMs) at lysine (K) sites are among the most common protein modifications. However, research on K-PTMs has been largely centered on identifying individual modification types, with a relative scarcity of balanced data analysis techniques. In this study, a classification system is developed for the prediction of concurrent multiple modifications at a single lysine residue. Initially, a well-established multi-label position-specific triad amino acid propensity algorithm is utilized for feature encoding. Subsequently, PreMLS: a novel ClusterCentroids undersampling algorithm based on MiniBatchKmeans was introduced to eliminate redundant or similar major class samples, thereby mitigating the issue of class imbalance. A convolutional neural network architecture was specifically constructed for the analysis of biological sequences to predict multiple lysine modification sites. The model, evaluated through five-fold cross-validation and independent testing, was found to significantly outperform existing models such as iMul-kSite and predML-Site. The results presented here aid in prioritizing potential lysine modification sites, facilitating subsequent biological assays and advancing pharmaceutical research. To enhance accessibility, an open-access predictive script has been crafted for the multi-label predictive model developed in this study. Yun Zuo 0001, Xingze Fang, Jiayong Wan, Wenying He, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
PLoS Comput. Biol. | 1 |
| 2023 | MLNGCF: circRNA-disease associations prediction with multilayer attention neural graph-based collaborative filteringabstractMOTIVATION: CircRNAs play a critical regulatory role in physiological processes, and the abnormal expression of circRNAs can mediate the processes of diseases. Therefore, exploring circRNAs-disease associations is gradually becoming an important area of research. Due to the high cost of validating circRNA-disease associations using traditional wet-lab experiments, novel computational methods based on machine learning are gaining more and more attention in this field. However, current computational methods suffer to insufficient consideration of latent features in circRNA-disease interactions. RESULTS: In this study, a multilayer attention neural graph-based collaborative filtering (MLNGCF) is proposed. MLNGCF first enhances multiple biological information with autoencoder as the initial features of circRNAs and diseases. Then, by constructing a central network of different diseases and circRNAs, a multilayer cooperative attention-based message propagation is performed on the central network to obtain the high-order features of circRNAs and diseases. A neural network-based collaborative filtering is constructed to predict the unknown circRNA-disease associations and update the model parameters. Experiments on the benchmark datasets demonstrate that MLNGCF outperforms state-of-the-art methods, and the prediction results are supported by the literature in the case studies. AVAILABILITY AND IMPLEMENTATION: The source codes and benchmark datasets of MLNGCF are available at https://github.com/ABard0/MLNGCF. Qunzhuo Wu, Zhaohong Deng, Wei Zhang 0221, Xiaoyong Pan, Kup-Sze Choi, Yun Zuo 0001, Hong-Bin Shen, Dongjun Yu |
Bioinform. | 6 |
| 2022 | MLysPRED: graph-based multi-view clustering and multi-dimensional normal distribution resampling techniques to predict multiple lysine sitesabstractPosttranslational modification of lysine residues, K-PTM, is one of the most popular PTMs. Some lysine residues in proteins can be continuously or cascaded covalently modified, such as acetylation, crotonylation, methylation and succinylation modification. The covalent modification of lysine residues may have some special functions in basic research and drug development. Although many computational methods have been developed to predict lysine PTMs, up to now, the K-PTM prediction methods have been modeled and learned a single class of K-PTM modification. In view of this, this study aims to fill this gap by building a multi-label computational model that can be directly used to predict multiple K-PTMs in proteins. In this study, a multi-label prediction model, MLysPRED, is proposed to identify multiple lysine sites using features generated from human protein sequences. In MLysPRED, three kinds of multi-label sequence encoding algorithms (MLDBPB, MLPSDAAP, MLPSTAAP) are proposed and combined with three encoding strategies (CHHAA, DR and Kmer) to convert preprocessed lysine sequences into effective numerical features. A multidimensional normal distribution oversampling technique and graph-based multi-view clustering under-sampling algorithm were first proposed and incorporated to reduce the proportion of the original training samples, and multi-label nearest neighbor algorithm is used for classification. It is observed that MLysPRED achieved an Aiming of 92.21%, Coverage of 94.98%, Accuracy of 89.63%, Absolute-True of 81.46% and Absolute-False of 0.0682 on the independent datasets. Additionally, comparison of results with five existing predictors also indicated that MLysPRED is very promising and encouraging to predict multiple K-PTMs in proteins. For the convenience of the experimental scientists, 'MLysPRED' has been deployed as a user-friendly web-server at http://47.100.136.41:8181. Yun Zuo 0001, Xiangxiang Zeng, Qiang Zhang 0026, Xiangrong Liu |
Briefings Bioinform. | 1 |
| 2021 | Pm6 A: an Integrated Classification Algorithm for 2021 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) Identifying m6 A SitesabstractAs a major RNA methylation modification, $\mathrm{N}^{6}_{-}$ methyladenosine (m6A) affects the occurrence and development of various human cancers through a variety of mechanisms. It has been reported that m6A RNA methylation involves different physiological and pathological processes. Therefore, the detection of m6A is helpful to reveal its biological function. Due to the high cost and time-consuming of high-throughput sequencing and the inaccurate sites identified, computational tools are needed to guide the accurate prediction of m6A modified sites and help reduce the costs associated with high-throughput sequencing. In this study, an integrated classification algorithm, Pm6A, is proposed to identify S. cerevisiae m6A sites using features generated from RNA sequences. In Pm6A, six sequence encoding schemes (pseudo dinucleotide composition, dinucleotide-based auto covariance, dinucleotide-based cross covariance, dinucleotide-based auto-cross covariance, mismatch and subsequence) are used for feature extraction, the VIBES ensemble algorithm based on four base classifiers, namely nearest neighbor, support vector machine, discriminant analysis and artificial neural network is used for classification, the optimized forward search algorithm is used to find the optimal parameters of model. The results of 10-fold cross-validation show that the proposed approach achieved better specificity (at Dataset 1) and better accuracy (at Dataset 2) than other methods. It is expected that Pm6A will be a useful tool for predicting m6A sites. Yun Zuo 0001, Xiangrong Liu, Xiangxiang Zeng |
BIBM | 1 |
| 2021 | CarSite-II: an integrated classification algorithm for identifying carbonylated sites based on K-means similarity-based undersampling and synthetic minority oversampling techniquesabstractBACKGROUND: Carbonylation is a non-enzymatic irreversible protein post-translational modification, and refers to the side chain of amino acid residues being attacked by reactive oxygen species and finally converted into carbonyl products. Studies have shown that protein carbonylation caused by reactive oxygen species is involved in the etiology and pathophysiological processes of aging, neurodegenerative diseases, inflammation, diabetes, amyotrophic lateral sclerosis, Huntington's disease, and tumor. Current experimental approaches used to predict carbonylation sites are expensive, time-consuming, and limited in protein processing abilities. Computational prediction of the carbonylation residue location in protein post-translational modifications enhances the functional characterization of proteins. RESULTS: In this study, an integrated classifier algorithm, CarSite-II, was developed to identify K, P, R, and T carbonylated sites. The resampling method K-means similarity-based undersampling and the synthetic minority oversampling technique (SMOTE-KSU) were incorporated to balance the proportions of K, P, R, and T carbonylated training samples. Next, the integrated classifier system Rotation Forest uses "support vector machine" subclassifications to divide three types of feature spaces into several subsets. CarSite-II gained Matthew's correlation coefficient (MCC) values of 0.2287/0.3125/0.2787/0.2814, False Positive rate values of 0.2628/0.1084/0.1383/0.1313, False Negative rate values of 0.2252/0.0205/0.0976/0.0608 for K/P/R/T carbonylation sites by tenfold cross-validation, respectively. On our independent test dataset, CarSite-II yield MCC values of 0.6358/0.2910/0.4629/0.3685, False Positive rate values of 0.0165/0.0203/0.0188/0.0094, False Negative rate values of 0.1026/0.1875/0.2037/0.3333 for K/P/R/T carbonylation sites. The results show that CarSite-II achieves remarkably better performance than all currently available prediction tools. CONCLUSION: The related results revealed that CarSite-II achieved better performance than the currently available five programs, and revealed the usefulness of the SMOTE-KSU resampling approach and integration algorithm. For the convenience of experimental scientists, the web tool of CarSite-II is available in http://47.100.136.41:8081/. Yun Zuo 0001, Jianyuan Lin, Xiangxiang Zeng, Quan Zou 0001, Xiangrong Liu |
BMC Bioinform. | 1 |
| 2019 | A Deep Neural Network for Antimicrobial Peptide RecognitionabstractWith the widespread use of antibiotics, many bacteria have developed resistance. Antimicrobial peptides have broad applications in medicine because of their high antibacterial activity. In this paper, a neural network model is introduced to recognize and detect antimicrobial peptides. Our model consists of an embedded, convolutional, bidirectional LSTM, and full connection layers. The embedded layer is used to code different amino acid residues into different vectors. The convolutional layer and bidirectional LSTM extract peptide amino acid residue sequence information. The full connection layer maps the sequence information linearly to the interval from 0 to 1, as the peptide for the probability of antimicrobial peptides. Training and testing on several different datasets reveal that our model performs better than other proposed models. Jianyuan Lin, Xiangxiang Zeng, Yun Zuo 0001, Ying Ju 0002, Xiangrong Liu |
BIBM | 3 |
| 2018 | O-GlcNAcPRED-II: an integrated classification algorithm for identifying O-GlcNAcylation sites based on fuzzy undersampling and a K-means PCA oversampling techniqueabstractMotivation: Protein O-GlcNAcylation (O-GlcNAc) is an important post-translational modification of serine (S)/threonine (T) residues that involves multiple molecular and cellular processes. Recent studies have suggested that abnormal O-G1cNAcylation causes many diseases, such as cancer and various neurodegenerative diseases. With the available protein O-G1cNAcylation sites experimentally verified, it is highly desired to develop automated methods to rapidly and effectively identify O-GlcNAcylation sites. Although some computational methods have been proposed, their performance has been unsatisfactory, particularly in terms of prediction sensitivity. Results: In this study, we developed an ensemble model O-GlcNAcPRED-II to identify potential O-GlcNAcylation sites. A K-means principal component analysis oversampling technique (KPCA) and fuzzy undersampling method (FUS) were first proposed and incorporated to reduce the proportion of the original positive and negative training samples. Then, rotation forest, a type of classifier-integrated system, was adopted to divide the eight types of feature space into several subsets using four sub-classifiers: random forest, k-nearest neighbour, naive Bayesian and support vector machine. We observed that O-GlcNAcPRED-II achieved a sensitivity of 81.05%, specificity of 95.91%, accuracy of 91.43% and Matthew's correlation coefficient of 0.7928 for five-fold cross-validation run 10 times. Additionally, the results obtained by O-GlcNAcPRED-II on two independent datasets also indicated that the proposed predictor outperformed five published prediction tools. Availability and implementation: http://121.42.167.206/OGlcPred/. Supplementary information: Supplementary data are available at Bioinformatics online. Cangzhi Jia, Yun Zuo 0001, Quan Zou 0001 |
Bioinform. | 2 |