EDBT 2026 Demo / reviewers in the wild / expert
Leyi Wei
dblp:145/6788
· DBLP profile ↗
97ranked-venue papers
15as first author
69since 2021 · last 2026
0000-0003-1444-190XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 74 · 8 first-author · 56 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRACE: Transformation-Aware Graph Refinement for Reaction Condition PredictionabstractIdentifying suitable reaction conditions is critical for chemical synthesis, as they directly affect yield, selectivity, and transformation feasibility. While recent methods have shown promising results, most approaches either encode reactants and products independently or rely on rule-based reaction graphs, both of which constrain the ability of the model to capture condition-relevant structural transformations. In this work, we propose TRACE, a transformation-aware graph refinement framework for reaction condition prediction. TRACE constructs atom-level joint graphs that integrate both reactant and product structures to represent condition-relevant transformations. A structure-aware encoder enriches atom features with local chemical context, followed by a dynamic interaction refinement module that adaptively infers task-specific edges. To further guide the model toward condition-relevant patterns, a mechanism regularized graph encoder incorporates reaction center information, enabling more accurate modeling of transformation mechanisms. Experiments on benchmark datasets show that TRACE achieves state-of-the-art performance across multiple condition types. The integration of transformation-aware refinement leads to improvements in prediction accuracy and generalization, while maintaining robust performance in challenging and realistic synthesis planning scenarios. Yujie Chen 0002, Tengfei Ma 0002, Yuansheng Liu, Leyi Wei, Dong-Sheng Cao 0001, Xiangxiang Zeng |
AAAI | 4 |
| 2026 | Quantum computing applications in drug discoveryabstractIn early drug discovery, virtual screening based on deep learning, virtual screening based on molecular docking, and molecular dynamics are three widely used computational strategies, but they always face a trade-off between throughput, search stability, and physical fidelity. This article discusses how quantum computing can be integrated into these processes under the constraints of Noisy Intermediate-Scale Quantum (NISQ). At present, the most realistic role of quantum computing is not the complete replacement of classical processes, but modular coprocessing for selected decision-sensitive subroutines. In the screening of deep learning, quantum modules are mainly inserted into selected components of the model. In predictive models, they are used to enhance representation learning or feature extraction. In generative models, they serve as priors or generators. In docking screening, quantum integration is suitable for specific substeps such as site recognition, pose search, and flexible docking. In molecular dynamics, representative examples include ground state ab initio molecular dynamics, annealer-based trajectory propagation, and excited state molecular dynamics, while most large-scale sampling is still done by classical methods. The actual problem in these scenarios is not whether the quantum module can be inserted, but whether it can provide repeatable gains related to decision-making under the constraints of actual running time and resources. Therefore, we emphasize strong classical baselines, reliable ranking and calibration, transparent resource reporting, and evaluation at downstream decision points as key criteria for assessing progress in the near term. Leyi Wei, Henry H. Y. Tong, Quan Zou 0001 |
Briefings Bioinform. | 2 |
| 2026 | HKD-CPI: high-order knowledge distillation enhanced inductive compound-protein interaction predictionabstractMOTIVATION: Accurately identifying compound-protein interactions (CPIs) is critical for accelerating drug discovery. Recent deep learning methods have achieved impressive results, yet they primarily focus on local structures and neighborhood information, often overlooking high-order interaction patterns shared among similar molecules. RESULTS: In this paper, we propose HKD-CPI, a high-order knowledge-enhanced inductive framework designed to improve generalization to unseen compound-protein pairs. Specifically, HKD-CPI introduces a molecular graph tokenization mechanism that aligns compound molecular graph features with token embeddings from sequence-pretrained large language models (LLMs), effectively infusing sequence-derived semantics into structural representations. To capture shared interaction patterns among functionally similar biomolecules, we construct a hypergraph-based representation to model high-order relationships between feature-similar compound/protein groups and their binding partners. Furthermore, a knowledge distillation strategy is further adopted to transfer high-order interaction knowledge from the hypergraph to a lightweight student model, enabling efficient and robust CPI prediction. Extensive experiments demonstrate that HKD-CPI outperforms existing state-of-the-art methods in inductive CPI prediction tasks. In particular, it achieves an average improvement of 4.94% in AUROC and 3.64% in AUPRC over the best-performing baseline across five benchmark datasets. AVAILABILITY AND IMPLEMENTATION: Our code and data are available at https://github.com/Hezy618/HKD-CPI. Zhongyu He, Xiangrong Liu, Yinghui Jiang, Junlin Xu, Shuting Jin, Leyi Wei, Youyu Wang |
Bioinform. | 7 |
| 2026 | CNNCaps-DBP: Leveraging protein language models with attention-augmented convolution for DNA-binding protein prediction
Ziyuan Yan, Aoyun Geng, Yazi Li, Jiajing Wang, Junlin Xu, Yajie Meng, Leyi Wei, Quan Zou 0001, Feifei Cui |
Neural Networks | 7 |
| 2026 | MFDL-DDI: An effective deep learning-based framework for predicting drug-drug interactions through multimodal information fusion
Yazi Li, Shuting Jin, Junlin Xu, Yajie Meng, Leyi Wei, Xin Gao 0001, Feifei Cui |
Pattern Recognit. | 6 |
| 2026 | METRON: Metabolic Dynamic Perception Kolmogorov-Arnold Network for Biological Age EstimationabstractBiological age is a more direct reflection of physiological status than chronological age, serving as a vital measure to evaluate health risks and aging interventions. While steroid metabolomics offers rich information for exploring aging mechanisms, the complex and nonlinear interactions within metabolic networks remain challenging in modeling. Here, we propose and describe METRON as a deep learning framework to predict biological ages from steroid metabolomics. Specifically, a Metabolite Interaction Perception Module (MIPM) is proposed to capture the interactions. Subsequently, a Group-Rational Kolmogorov-Arnold Network is also integrated to capture intricate dependencies and enhance the representation capability. We demonstrate that METRON achieves promising performance as compared to other machine learning and deep learning methods. Beyond performance, METRON offers interpretability by recovering the established markers such as Dehydroepiandrosterone (DHEA) and identifying 17-hydroxyprogesterone (17-OH-P4) as the key signature linked to hypothalamic-pituitary-adrenal axis dynamics. These results support the capacity of METRON not only to estimate biological age but also to uncover underappreciated metabolic drivers behind aging. Zhongshen Li, Jixiang Yu, Shen You, Hao Liu 0072, Luyang Cai, Yuxuan Deng, Leyi Wei, Junkai Ji, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2026 | DeepNhKcr: Explainable Deep Learning Framework for the Prediction of Crotonylation Sites of Non-Histone Lysine in Plants Based on Pre-Trained Protein Language ModelabstractLysine crotonylation (Kcr) is an important protein modification occurring after translation in biology, serving an essential function in a range of biological processes in both plants and animals, including the regulation of gene expression, the maintenance of cellular metabolic balance, and the enhancement of photosynthesis. Exploring the detection of Kcr sites is essential for uncovering their biological functions. Nonetheless, conventional experimental approaches for detection are often time-consuming, expensive, and hindered by various technical constraints, making the precise identification of Kcr sites a significant challenge. This study seeks to develop a computational approach for the rapid and accurate prediction of Kcr sites in plant non-histone proteins. We introduce a novel deep learning framework named DeepNhKcr, which integrates the protein language model (ESM2) with a bidirectional long short-term memory (BiLSTM) network. To address the challenge of data imbalance, the model replaces the conventional cross-entropy loss with the focal loss function. In addition, DeepNhKcr combines advanced deep learning approaches with traditional protein encoding strategies to enable effective feature extraction and integration. This method not only significantly boosts the accuracy of predicting Kcr sites in non-histone proteins of plants. but also provides interpretability, shedding light on the potential links between key sequence characteristics and their biological roles. DeepNhKcr delivers outstanding results, surpassing existing machine learning and deep learning models, and demonstrating excellent performance in both five-fold cross-validation and independent test experiments. Moreover, the model integrates interpretability analysis techniques to investigate the connections between important sequence features and their biological roles. DeepNhKcr acts as a powerful method for detecting Kcr sites in plant non-histone proteins and is anticipated to greatly advance future studies in plant Kcr site prediction. Zhenjie Luo, Aoyun Geng, Junlin Xu, Yajie Meng, Shankai Yan, Leyi Wei, Qingchen Zhang 0001, Quan Zou 0001, Feifei Cui |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2026 | A Multi-Modal Contrastive Learning Framework for Cyclic Peptide Permeability PredictionabstractCyclic peptides represent a rapidly growing class of therapeutics, yet their development is often hindered by the challenge of predicting cell membrane permeability, a critical determinant of drug efficacy. Existing computational methods often struggle to integrate the diverse structural information inherent in these complex molecules, resulting in suboptimal predictive accuracy. Here, we introduce MCPerm, a multi-modal deep learning framework that synergistically integrates 1D SMILES, 2D topological, and 3D geometric information through a novel modality share and contrastive learning strategy to accurately predict cyclic peptide permeability. MCPerm fine-tunes a pretrained peptide language model for SMILES encoding and uses a parameter-sharing graph transformer for structural representation, while a dual contrastive learning mechanism enforces representational consistency both within and between modalities. On the benchmark PAMPA dataset, MCPerm achieves state-of-the-art performance, significantly outperforming leading methods. We further demonstrate its robustness and competitive transferability across three independent assays (Caco-2, MDCK, and RRCK). Our work presents a robust in silico framework that holds potential to accelerate the rational design and discovery of cell-permeable cyclic peptide drugs. Furthermore, to move beyond predictive accuracy, we introduced an attention-based visualization analysis. The results demonstrate that our model is not a "black box"; it has learned key chemical principles governing cyclic peptide permeability. Shuwen Xiong, Feifei Cui, Rao Zeng, Ran Su, Leyi Wei |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2026 | DeepR2OM: Accurate Recognition for RNA 2′-O-Methylation Sites in Human Genome Using Deep Learningabstract2'-O-methylation (2OM) of ribose is a widespread RNA modification that significantly impacts RNA stability, structure, and function. Accurately predicting 2OM sites is crucial for understanding RNA's biological functions and related pathologies. Traditional detection methods pose challenges such as resource intensiveness, potential RNA sample damage, and high costs. However, recent advancements in machine learning, particularly deep learning techniques, offer rapid and cost-effective prediction solutions. In this study, we introduce DeepR2OM, a novel method integrating feature selection and deep learning for 2OM sites prediction. DeepR2OM encodes sequences using eight RNA descriptors, employs feature selection algorithms to reduce dimensions, and then utilizes a deep learning network for training. After evaluating various deep learning architectures, we selected Convolutional Neural Network (CNN), Multi-Head Self-Attention mechanism, and Deep Neural Network (DNN) as our final prediction models. Experimental results demonstrate DeepR2OM's effectiveness, achieving 87.1% accuracy (ACC), 85.5% recall rate (Recall), 87.9% precision (PRE), and a Matthews correlation coefficient (MCC) of 75.7% on an independent test set. This tool serves as a valuable resource for exploring the functional and bioinformatic aspects of 2OM sites. Shun Gao, Ziyuan Yan, Feifei Cui, Leyi Wei, Qingchen Zhang 0001, Quan Zou 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2026 | KG-CMI: Knowledge Graph Enhanced Cross-Mamba Interaction for Medical Visual Question Answering
Xianyao Zheng, Hui Cui 0002, Changming Sun, Xiangyu Li 0004, Ran Su, Leyi Wei, Qiangguo Jin |
IEEE Trans. Ind. Informatics | 7 |
| 2026 | PGST: A Prototype-Guided Parameter-Efficient Network for Spatial Transcriptomics PredictionabstractSpatial transcriptomics (ST) aims to decode spatially resolved gene expression patterns while preserving tissue morphology. Current methods tend to use lower-cost deep learning approaches for gene expression prediction, yet face severe challenges. First, existing methods fail to give sufficient consideration to the spatial specificity of positional encoding inherent in ST; second, they neglect to leverage spatially coherent co-expression patterns across different domains; third, their reliance on linearly weighted aggregation induces vulnerability to noise and distribution shifts; and finally, these architectures exhibit limited parameter efficiency. To address these issues, we introduce prototype-guided network for spatial transcriptomics (PGST), which includes four parts: (1) oriented signal propagation through polar embedding strategy for spatial transcriptomics (PEST); (2) prototype-guided aggregation for global co-feature preservation; (3) global consistency enforcement via shared decoder with reconstruction loss; and (4) lightweight architectural design. Our framework integrates contrastive learning with graph neural networks to balance local-global spatial dependencies and cross-modal consistency. Experimental results on multiple datasets from ST demonstrate the superior performance of our PGST model than existing methods. Yuan He 0016, Kaimiao Hu, Changming Sun, Leyi Wei, Ran Su |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | ADSA-Net: Addressing Intra- and Inter-Class Variabilities for Severity Assessment of Atopic DermatitisabstractAtopic dermatitis (AD) is a chronic inflammatory skin disorder characterized by recurrent itching, erythema, dryness, and eczematous lesions. Automated AD severity assessment is crucial for cost-effective and precision clinical decision-making but remains challenging. This is due to the subtle contrast variations between key dermatological signs and significant variations in lesion sizes across patients and disease stages. To address these issues, we propose ADSA-Net, which is designed to handle both intra- and inter-class variabilities. ADSA-Net first extracts multi-scale texture-aware features to effectively model variations in lesion size and texture. It then leverages contrastive learning to enhance intra- and inter-class differentiation, strengthening model's discriminatory ability for samples that are difficult to distinguish. Finally, ADSA-Net refines the learning process by leveraging a dynamic feature pool of correctly classified samples to guide the calibration of misclassified instances, enhancing overall accuracy. We further establish a dataset for AD severity assessment. Comprehensive experiments on this dataset show that ADSA-Net significantly outperforms existing state-of-the-art methods. Qiangguo Jin, Xurong Chen, Hui Cui 0002, Changming Sun, Youpeng Deng, Cong Cong 0001, Yuqi Fang, Ran Su, Leyi Wei |
BIBM | 9 |
| 2025 | Iterative clustering algorithm G-DESC-E and pan-cancer key gene analysis based on single-cell sequencing dataabstractSingle-cell sequencing technology has profoundly revolutionized the field of cancer genomics, enabling researchers to explore gene expression profiles at the resolution of individual cells. Despite its extensive applications in the study of cancer gene states, pan-cancer analyses remain relatively underexplored. In this study, we propose the G-DESC-E algorithm, which effectively distinguishes dimensionality-reduced data through a grid-based approach, filters out outliers during the preprocessing phase, and employs the Louvain algorithm for prescreening cluster centroids as initial clusters. We construct an objective function by integrating label entropy with the Kullback-Leibler divergence formula, achieving final clustering results through iterative optimization. Our findings demonstrate the effectiveness of the G-DESC-E algorithm in enhancing clustering accuracy. By applying our methodology to real-world datasets, we illustrate its capability to identify critical transcriptional features associated with distinct cancer subtypes. Coupled with clustering visualization and gene ontology analysis, we identify over thirty genes potentially related to cancer occurrence and progression. The algorithm and research framework presented in this study pave the way for new directions in clinical research by applying single-cell sequencing technology to the analysis of key genes within the realm of pan-cancer analysis for the first time. This approach offers valuable insights that can inform further clinical investigations. Ke Wu 0023, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su |
Briefings Bioinform. | 6 |
| 2025 | MCAMEF-BERT: an efficient deep learning method for RNA N7-methylguanosine site prediction via multi-branch feature integrationabstractAccurate identification of N7-methylguanosine (m7G) modification sites plays a critical role in uncovering the regulatory mechanisms of various biological processes, including human development, tumor initiation, and progression. However, existing prediction methods still suffer from limited representational power, redundant feature fusion, insufficient utilization of biological prior knowledge, and poor interpretability. In this study, we propose a novel deep learning model named MCAMEF-BERT. This model adopts a parallel architecture that integrates both a DNABERT-2-based pretrained model branch and multiple traditional feature encoding branches, enabling comprehensive multi-perspective sequence feature extraction. To address the redundancy issue in feature fusion, we introduce a multi-channel attention module. Our model demonstrates superior accuracy and effectiveness on datasets from m7GHub, outperforming other state-of-the-art classifiers. Furthermore, we validate the interpretability of MCAMEF-BERT through in silico saturation mutagenesis experiments, and confirm its robustness in motif recognition. Moreover, its generalization capability is validated across diverse RNA modification site prediction tasks. Junlei Yu, Wenjia Gao, Siqi Chen 0001, Ronglin Lu, Jianbo Qiao, Junru Jin, Leyi Wei, Feifei Cui, Xinbo Jiang, Zhongmin Yan |
Briefings Bioinform. | 7 |
| 2025 | Synergizing multimodal data and fingerprint space exploration for mechanism of action predictionabstractMOTIVATION: Effective computational methods for predicting the mechanism of action (MoA) of compounds are essential in drug discovery. Current MoA prediction models mainly utilize the structural information of compounds. However, high-throughput screening technologies have generated more targeted cell perturbation data for MoA prediction, a factor frequently disregarded by the majority of current approaches. Moreover, exploring the commonalities and specificities among different fingerprint representations remains challenging. RESULTS: In this paper, we propose IFMoAP, a model integrating cell perturbation image and fingerprint data for MoA prediction. Firstly, we modify the Res-Net to accommodate the feature extraction of five-channel cell perturbation images and establish a granularity-level attention mechanism to combine coarse- and fine-grained features. To learn both common and specific fingerprint features, we introduce an FP-CS module, projecting four fingerprint embeddings into distinct spaces and incorporating two loss functions for effective learning. Finally, we construct two independent classifiers based on image and fingerprint features for prediction and for weighting the two prediction scores. Experimental results demonstrate that our model achieves highest accuracy of 0.941 when using multimodal data. The comparison with other methods and explorations further highlights the superiority of our proposed model and the complementary characteristics of multimodal data. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/ s1mplehu/IFMoAP. The raw image data of Cell Painting can be accessed from Figshare (https://doi.org/10.17044/scilifelab.21378906). Kaimiao Hu, Jianguo Wei, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su |
Bioinform. | 5 |
| 2025 | Molecular pretraining models towards molecular property prediction
Jianbo Qiao, Wenjia Gao, Junru Jin, Balachandran Manavalan, Leyi Wei |
Sci. China Inf. Sci. | 7 |
| 2025 | PKDF-Net: Anticancer peptide prediction via a prior-knowledge-aware dual-path feature-entangled network
Qiangguo Jin, Ankang Wu, Leyi Wei, Hui Cui 0002, Ping Xuan, Xikang Feng, Ran Su |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Iterative pseudo-labeling based adaptive copy-paste supervision for semi-supervised tumor segmentation
Qiangguo Jin, Hui Cui 0002, Changming Sun, Yimiao He, Ping Xuan, Cong Cong 0001, Leyi Wei, Ran Su |
Knowl. Based Syst. | 9 |
| 2025 | Elucidating spatiotemporal chromatin dynamics with multi-stage differential variations from Hi-CabstractHigh-throughput sequencing such as Hi-C captures spatiotemporal chromatin interactions, revealing the intricate interplays within transcriptional regulation and chromatin dynamics during cellular reprogramming and developmental processes. However, the forecast on chromatin dynamics in successive developmental stages remains challenging due to the inherent complexity of spatial and temporal patterns in Hi-C data across developmental stages. Towards such a direction, we present StarMie, a deep learning framework that integrates the spatiotemporal-aware module and the multi-stage differential variation module to predict high-throughput chromatin interactions in next developmental stages. Our comprehensive evaluation demonstrates that StarMie outperforms existing methods and sufficiently captures discriminative spatial and temporal dependencies as well as inter-stage-level variations of chromatin interactions. Moreover, the dual importance of both spatial and temporal information in Hi-C data is observed in parameter analysis. Ablation studies also confirm the essential role of each component in StarMie. Furthermore, five cross-species case studies support StarMie's cross-species generalizability and its capability to extract universal chromatin interaction patterns in different developmental stages. In-depth analysis demonstrates that StarMie uncovers conserved genomic logic in cardiac development and disease. Overall, this work paves a new approach for exploring genome reprogramming and development through predictive modeling of Hi-C dynamics. Zhongshen Li, Jixiang Yu, Shen You, Leyi Wei, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
Knowl. Based Syst. | 4 |
| 2025 | Taco-DDI: accurate prediction of drug-drug interaction events using graph transformer-based architecture and dynamic co-attention matrices
Jianbo Qiao, Junru Jin, Kefei Li, Wenjia Gao, Feifei Cui, Leyi Wei |
Neural Networks | 10 |
| 2025 | MC-MSTLoc: Self-Supervised Pre-Training for Imbalanced Multi-Label Protein Subcellular Localization Prediction Using Immunofluorescence ImagesabstractWith the rapid growth of high-resolution microscopy imaging data, current protein subcellular localization methods often face the problem of imbalanced data with long-tailed distributions in large-scale protein data. To address this challenge, this paper proposes a self-supervised pre-training method called MC-MSTLoc. Aiming to maximize feature consistency and inconsistency of microscopy imaging data, the pre-training scheme is proposed based on contrastive task at scale and view levels, which substantially improves the quality of the learned feature representations. Experimental results on benchmark datasets demonstrate that MC-MSTLoc outperforms existing self-supervised pretraining methods for protein subcellular localization prediction. Model ablation experiments and pretraining effectiveness analysis confirm the method performance. Additionally, model visualization analysis and interpretability experiments demonstrate the crucial role of the method in learning information distribution and patterns of different subcellular locations. Fengsheng Wang, Jianbo Qiao, Leyi Wei |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | BFGTP: A BERT-Guided Two-Stage Molecular Representation Learning Framework for Toxicity PredictionabstractAccurate prediction of molecular toxicity is vital for drug development. Most mainstream methods rely on fingerprints or graph-based feature extraction, the emergence of large language models (LLMs) offers new prospects for molecular representation learning in toxicity prediction. Although several studies attempt to leverage LLMs to integrate molecular sequence data for pretraining molecular representations, certain limitations remain. Current LLM-based approaches usually utilize solely on class embedding features, overlooking the rich information in sequence embedding. Moreover, integrating pre-trained molecular representations with multi-modal molecular data may further enhance performance in toxicity prediction. To address these challenges, we propose BFGTP, a BERT-guided two-stage molecular representation learning framework for toxicity prediction. Firstly, we design independent encoders for molecular descriptions of three modalities, where the fingerprint encoder with dual level attention mechanisms effectively integrates multi-category fingerprints. Then, the two-stage guide strategy is introduced to fully utilize the prior knowledge of LLMs, employing contrastive learning to align and fuse the tri-modal representations and knowledge distillation to align predicted value distributions. BFGTP ultimately combines fingerprint and graph representations to predict molecular toxicity. Experiments on seven toxicity datasets show that BFGTP outperforms baselines, achieving the highest AUC on five datasets and the best average performance across five evaluation metrics. Ablation studies, t-SNE visualization and case study confirm the effectiveness of BFGTP's components and its ability to capture meaningful molecular representations. Kaimiao Hu, Yuan He 0016, Jianguo Wei, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Multi-Modal Deep Representation Learning Accurately Identifies and Interprets Drug-Target InteractionsabstractDeep learning offers efficient solutions for drug-target interaction prediction, but current methods often fail to capture the full complexity of multi-modal data (i.e., sequence, graphs, and three-dimensional structures), limiting both performance and generalization. Here, we present UnitedDTA, a novel explainable deep learning framework capable of integrating multi-modal biomolecule data to improve the binding affinity prediction, especially for novel (unseen) drugs and targets. UnitedDTA enables automatic learning unified discriminative representations from multi-modality data via contrastive learning and cross-attention mechanisms for cross-modality alignment and integration. Comparative results on multiple benchmark datasets show that UnitedDTA significantly outperforms the state-of-the-art drug-target affinity prediction methods and exhibits better generalization ability in predicting unseen drug-target pairs. More importantly, unlike most "black-box" deep learning methods, our well-established model offers better interpretability which enables us to directly infer the important substructures of the drug-target complexes that influence the binding activity, thus providing the insights in unveiling the binding preferences. Moreover, by extending UnitedDTA to other downstream tasks (e.g., molecular property prediction), we showcase the proposed multi-modal representation learning is capable of capturing the latent molecular representations that are closely associated with the molecular property, demonstrating the broad application potential for advancing the drug discovery process. Jiayue Hu, Xiangxiang Zeng, Quan Zou 0001, Ran Su, Leyi Wei |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | A Multi-Objective Comprehensive Framework for Predicting Protein-Peptide Interactions and Binding ResiduesabstractIdentifying protein-peptide interaction pairs and their corresponding binding residues is crucial and can greatly facilitate peptide therapeutic design as well as improve our understanding of protein function mechanisms. Recently, several computational approaches have been proposed to solve the protein-peptide interaction prediction problem. However, most existing prediction methods cannot simultaneously predict protein-peptide interaction pairs and their binding residues directly from the protein and peptide sequences. Here, we developed a Comprehensive Protein-Peptide Interaction prediction Framework (CPPIF), to predict both binary protein-peptide interaction and their binding residues. We also constructed a benchmark dataset containing more than 8900 protein-peptide interacting pairs with non-covalent interactions and their corresponding binding residues to systematically evaluate the performances of existing models. Comprehensive evaluation on the benchmark datasets demonstrated that CPPIF can successfully predict the non-covalent protein-peptide interactions that cannot be effectively captured by previous prediction methods. Moreover, CPPIF outperformed other state-of-the-art methods in predicting binding residues in the peptides and achieved good performance in the identification of important binding residues in the proteins. Ruheng Wang, Xuetong Yang, Leyi Wei |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | GCNLA: Inferring Cell-Cell Interactions From Spatial Transcriptomics With Long Short-Term Memory and Graph Convolutional NetworksabstractSpatial transcriptomics analysis methods offer an opportunity to investigate highly diverse biological tissues. Cell-cell communication is fundamental for maintaining physiological homeostasis in organisms and coordinating complex biological processes. Identifying cell-cell interactions is critical for understanding cellular activities. The interaction of a cell with other cells depends on several factors, and most of the existing methods that consider only gene expression information of neighbouring cells and spatial location information are somewhat limited. In this paper, we propose a network architecture based on graph convolution network and long short-term memory attention module-GCNLA, which contains graph convolution layer, long short-term memory network, attention module, and residual connections. GCNLA not only learns the spatial structure of cells but also captures interaction information between distal cells, the attention module further extracting and enhancing features related to cell-cell interactions. Finally, the inner product decoding calculates the cosine similarity, which is used to infer cell-cell interactions. In addition, GCNLA is capable of reconstructing the complete cell-cell interaction network. The experimental results on seqFISH and MERFISH demonstrate that the GCNLA network structure has better robustness and noise immunity. The potential features learned by GCNLA enable other downstream analyses, including single-cell resolution cell clustering based on spatial information resolving cell heterogeneity. Xiuhao Fu, Zhenjie Luo, Leyi Wei, Jingbing Li, Feifei Cui, Quan Zou 0001, Qingchen Zhang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Guest Editorial: Large Language Models With Applications in Bioinformatics and Biomedicine
Quan Zou 0001, Limin Jiang, Leyi Wei |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Multiview Deep Learning-Based Molecule Design and Structural Optimization Accelerates Inhibitor DiscoverabstractIn this work, we propose MEDICO, a multiview deep generative model for molecule generation, structural optimization, and the SARS-CoV-2 inhibitor discovery. To the best of our knowledge, MEDICO is the first-of-this-kind graph generative model that can generate molecular graphs similar to the structure of targeted molecules, with a multiview representation learning framework to sufficiently and adaptively learn comprehensive structural semantics from targeted molecular topology and geometry. We show that our MEDICO significantly outperforms the state-of-the-art methods in generating valid, novel, and unique molecules under benchmarking comparisons, particularly achieving $\tilde {8}5 \%$ improvement compared with the state-of-the-art methods in terms of validity. Importantly, we showcase that the multiview deep learning model enables us to generate not only the molecules structurally similar to the targeted molecules but also the molecules with desired chemical properties. Moreover, case study results on targeted molecule generation for the SARS-CoV-2 main protease (Mpro) show that we successfully generate new small molecules with desired drug-like properties for the Mpro by integrating molecular docking into our model as a chemical priori, potentially accelerating the de novo design of COVID-19 drugs. Furthermore, we apply MEDICO to the structural optimization of three well-known Mpro inhibitors (N3, 11a, and GC376) and achieve $\tilde {8}8 \%$ improvement compared with the origin inhibitors in their binding affinity to Mpro, demonstrating the application value of our model for the development of therapeutics for SARS-CoV-2 infection. Ruheng Wang, Quan Zou 0001, Xiangxiang Zeng, Ran Su, Leyi Wei |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2024 | MSKI-Net: Towards modality-specific knowledge interaction for glioma survival predictionabstractGliomas hold a prominent position in neurooncology due to their high malignancy and poor survival rates. Accurately predicting the prognosis and survival risk of glioma patients is crucial for clinical treatment. Recent advances in survival prediction methods emphasize the importance of integrating complementary information from diverse modalities while neglecting the significant modality gap between pathological images and genomic data. To address this issue, we propose a modality-specific knowledge interaction network (MSKI-Net), which integrates whole slide images (WSI), RNA-Seq gene expression data, and copy number variation (CNV) data for glioma survival analysis. The MSKI-Net consists of a modality-specific feature enhancement (MSFE) module, a modality-interactive cross-attention (MICA) module, and a modality-specific knowledge-guided representation learning (MSKR) module. The three modules collaborate by complementing modality-specific features with modality-agnostic knowledge to improve the learning capability of MSKI-Net. Furthermore, we construct a dataset named TCGAmm, which combines WSI, RNA-Seq, and CNV data from The Cancer Genome Atlas (TCGA) to address the issue of data scarcity. Extensive experiments demonstrate that MSKI-Net achieves superior performance in predicting the survival risk of glioma cancer. Ran Su, Hui Cui 0002, Ping Xuan, Xikang Feng, Leyi Wei, Qiangguo Jin |
BIBM | 6 |
| 2024 | Location Embedding Based Pairwise Distance Learning for Fine-Grained Diagnosis of Urinary Stones
Qiangguo Jin, Jiapeng Huang, Changming Sun, Hui Cui 0002, Ping Xuan, Ran Su, Leyi Wei, Yu-Jie Wu, Chia-An Wu, Henry Been-Lirn Duh, Yueh-Hsun Lu |
MICCAI (11) | 7 |
| 2024 | CELA-MFP: a contrast-enhanced and label-adaptive framework for multi-functional therapeutic peptides predictionabstractFunctional peptides play crucial roles in various biological processes and hold significant potential in many fields such as drug discovery and biotechnology. Accurately predicting the functions of peptides is essential for understanding their diverse effects and designing peptide-based therapeutics. Here, we propose CELA-MFP, a deep learning framework that incorporates feature Contrastive Enhancement and Label Adaptation for predicting Multi-Functional therapeutic Peptides. CELA-MFP utilizes a protein language model (pLM) to extract features from peptide sequences, which are then fed into a Transformer decoder for function prediction, effectively modeling correlations between different functions. To enhance the representation of each peptide sequence, contrastive learning is employed during training. Experimental results demonstrate that CELA-MFP outperforms state-of-the-art methods on most evaluation metrics for two widely used datasets, MFBP and MFTP. The interpretability of CELA-MFP is demonstrated by visualizing attention patterns in pLM and Transformer decoder. Finally, a user-friendly online server for predicting multi-functional peptides is established as the implementation of the proposed CELA-MFP and can be freely accessed at http://dreamai.cmii.online/CELA-MFP. Yitian Fang, Mingshuang Luo, Zhixiang Ren, Leyi Wei |
Briefings Bioinform. | 4 |
| 2024 | Therapeutic peptides identification via kernel risk sensitive loss-based k-nearest neighbor model and multi-Laplacian regularizationabstractTherapeutic peptides are therapeutic agents synthesized from natural amino acids, which can be used as carriers for precisely transporting drugs and can activate the immune system for preventing and treating various diseases. However, screening therapeutic peptides using biochemical assays is expensive, time-consuming, and limited by experimental conditions and biological samples, and there may be ethical considerations in the clinical stage. In contrast, screening therapeutic peptides using machine learning and computational methods is efficient, automated, and can accurately predict potential therapeutic peptides. In this study, a k-nearest neighbor model based on multi-Laplacian and kernel risk sensitive loss was proposed, which introduces a kernel risk loss function derived from the K-local hyperplane distance nearest neighbor model as well as combining the Laplacian regularization method to predict therapeutic peptides. The findings indicated that the suggested approach achieved satisfactory results and could effectively predict therapeutic peptide sequences. Yijie Ding, Leyi Wei, Xiaoyi Guo, Fengming Ni |
Briefings Bioinform. | 3 |
| 2024 | StructuralDPPIV: a novel deep learning model based on atom structure for predicting dipeptidyl peptidase-IV inhibitory peptidesabstractMOTIVATION: Diabetes is a chronic metabolic disorder that has been a major cause of blindness, kidney failure, heart attacks, stroke, and lower limb amputation across the world. To alleviate the impact of diabetes, researchers have developed the next generation of anti-diabetic drugs, known as dipeptidyl peptidase IV inhibitory peptides (DPP-IV-IPs). However, the discovery of these promising drugs has been restricted due to the lack of effective peptide-mining tools. RESULTS: Here, we presented StructuralDPPIV, a deep learning model designed for DPP-IV-IP identification, which takes advantage of both molecular graph features in amino acid and sequence information. Experimental results on the independent test dataset and two wet experiment datasets show that our model outperforms the other state-of-art methods. Moreover, to better study what StructuralDPPIV learns, we used CAM technology and perturbation experiment to analyze our model, which yielded interpretable insights into the reasoning behind prediction results. AVAILABILITY AND IMPLEMENTATION: The project code is available at https://github.com/WeiLab-BioChem/Structural-DPP-IV. Junru Jin, Zhongshen Li, Mushuang Fan, Sirui Liang, Ran Su, Leyi Wei |
Bioinform. | 8 |
| 2024 | NanoCon: contrastive learning-based deep hybrid network for nanopore methylation detectionabstractMOTIVATION: 5-Methylcytosine (5mC), a fundamental element of DNA methylation in eukaryotes, plays a vital role in gene expression regulation, embryonic development, and other biological processes. Although several computational methods have been proposed for detecting the base modifications in DNA like 5mC sites from Nanopore sequencing data, they face challenges including sensitivity to noise, and ignoring the imbalanced distribution of methylation sites in real-world scenarios. RESULTS: Here, we develop NanoCon, a deep hybrid network coupled with contrastive learning strategy to detect 5mC methylation sites from Nanopore reads. In particular, we adopted a contrastive learning module to alleviate the issues caused by imbalanced data distribution in nanopore sequencing, offering a more accurate and robust detection of 5mC sites. Evaluation results demonstrate that NanoCon outperforms existing methods, highlighting its potential as a valuable tool in genomic sequencing and methylation prediction. In addition, we also verified the effectiveness of our representation learning ability on two datasets by visualizing the dimension reduction of the features of methylation and nonmethylation sites from our NanoCon. Furthermore, cross-species and cross-5mC methylation motifs experiments indicated the robustness and the ability to perform transfer learning of our model. We hope this work can contribute to the community by providing a powerful and reliable solution for 5mC site detection in genomic studies. AVAILABILITY AND IMPLEMENTATION: The project code is available at https://github.com/Challis-yin/NanoCon. Chenglin Yin, Ruheng Wang, Jianbo Qiao, Hongliang Duan, Xinbo Jiang, Saisai Teng, Leyi Wei |
Bioinform. | 8 |
| 2024 | Inter- and intra-uncertainty based feature aggregation model for semi-supervised histopathology image segmentation
Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiangbin Zheng 0001, Leilei Cao, Leyi Wei, Ran Su |
Expert Syst. Appl. | 7 |
| 2024 | Towards retraining-free RNA modification prediction with incremental learning
Jianbo Qiao, Junru Jin, Haoqing Yu, Leyi Wei |
Inf. Sci. | 4 |
| 2023 | Shape-aware contrastive deep supervision for esophageal tumor segmentation from CT scansabstractAccurate tumor segmentation is crucial for esophageal cancer radiotherapy treatment planning. The low contrast among the esophagus, tumors, and surrounding tissues, and irregular tumor shapes limit the performance of automatic segmentation methods. In this paper, we aim to exploit the irregular shapes of tumors to facilitate accurate segmentation. We propose a simple and pluggable shape-aware contrastive deep supervision network (SCDSNet) with shape-aware regularization and voxel-to-voxel contrastive deep supervision. Specifically, the shape-aware regularization with an uncertainty minimization strategy encourages the precise predictions of an additional shape-aware head. The voxel-to-voxel contrastive deep supervision enhances the multi-scale shape-tumor contrast for better voxel-to-voxel prediction of shapes. The proposed method is simple and highly pluggable, which can easily be extended to other frameworks. Further, we establish a large in-house dataset on esophageal cancer to validate the effectiveness of our proposed method. The quantitative and qualitative experimental results demonstrate the effectiveness of SCDSNet on the esophageal cancer dataset. Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiapeng Huang, Ping Xuan, Yiyue Xu, Leilei Cao, Leyi Wei, Ran Su |
BIBM | 9 |
| 2023 | Multi-modality Contrastive Learning for Sarcopenia Screening from Hip X-rays and Clinical Information
Qiangguo Jin, Changjiang Zou, Hui Cui 0002, Changming Sun, Shu-Wei Huang, Yi-Jie Kuo, Ping Xuan, Leilei Cao, Ran Su, Leyi Wei, Henry Been-Lirn Duh, Yu-Pin Chen |
MICCAI (6) | 10 |
| 2023 | AFP-MFL: accurate identification of antifungal peptides using multi-view feature learningabstractRecently, peptide-based drugs have gained unprecedented interest in discovering and developing antifungal drugs due to their high efficacy, broad-spectrum activity, low toxicity and few side effects. However, it is time-consuming and expensive to identify antifungal peptides (AFPs) experimentally. Therefore, computational methods for accurately predicting AFPs are highly required. In this work, we develop AFP-MFL, a novel deep learning model that predicts AFPs only relying on peptide sequences without using any structural information. AFP-MFL first constructs comprehensive feature profiles of AFPs, including contextual semantic information derived from a pre-trained protein language model, evolutionary information, and physicochemical properties. Subsequently, the co-attention mechanism is utilized to integrate contextual semantic information with evolutionary information and physicochemical properties separately. Extensive experiments show that AFP-MFL outperforms state-of-the-art models on four independent test datasets. Furthermore, the SHAP method is employed to explore each feature contribution to the AFPs prediction. Finally, a user-friendly web server of the proposed AFP-MFL is developed and freely accessible at http://inner.wei-group.net/AFPMFL/, which can be considered as a powerful tool for the rapid screening and identification of novel AFPs. Yitian Fang, Lesong Wei, Jie Chen 0001, Leyi Wei |
Briefings Bioinform. | 6 |
| 2023 | CoraL: interpretable contrastive meta-learning for the prediction of cancer-associated ncRNA-encoded small peptidesabstractNcRNA-encoded small peptides (ncPEPs) have recently emerged as promising targets and biomarkers for cancer immunotherapy. Therefore, identifying cancer-associated ncPEPs is crucial for cancer research. In this work, we propose CoraL, a novel supervised contrastive meta-learning framework for predicting cancer-associated ncPEPs. Specifically, the proposed meta-learning strategy enables our model to learn meta-knowledge from different types of peptides and train a promising predictive model even with few labeled samples. The results show that our model is capable of making high-confidence predictions on unseen cancer biomarkers with only five samples, potentially accelerating the discovery of novel cancer biomarkers for immunotherapy. Moreover, our approach remarkably outperforms existing deep learning models on 15 cancer-associated ncPEPs datasets, demonstrating its effectiveness and robustness. Interestingly, our model exhibits outstanding performance when extended for the identification of short open reading frames derived from ncPEPs, demonstrating the strong prediction ability of CoraL at the transcriptome level. Importantly, our feature interpretation analysis discovers unique sequential patterns as the fingerprint for each cancer-associated ncPEPs, revealing the relationship among certain cancer biomarkers that are validated by relevant literature and motif comparison. Overall, we expect CoraL to be a useful tool to decipher the pathogenesis of cancer and provide valuable information for cancer research. The dataset and source code of our proposed method can be found at https://github.com/Johnsunnn/CoraL. Zhongshen Li, Junru Jin, Wentao Long, Haoqing Yu, Xin Gao 0001, Kenta Nakai, Quan Zou 0001, Leyi Wei |
Briefings Bioinform. | 9 |
| 2023 | MPI-VGAE: protein-metabolite enzymatic reaction link learning by variational graph autoencodersabstractEnzymatic reactions are crucial to explore the mechanistic function of metabolites and proteins in cellular processes and to understand the etiology of diseases. The increasing number of interconnected metabolic reactions allows the development of in silico deep learning-based methods to discover new enzymatic reaction links between metabolites and proteins to further expand the landscape of existing metabolite-protein interactome. Computational approaches to predict the enzymatic reaction link by metabolite-protein interaction (MPI) prediction are still very limited. In this study, we developed a Variational Graph Autoencoders (VGAE)-based framework to predict MPI in genome-scale heterogeneous enzymatic reaction networks across ten organisms. By incorporating molecular features of metabolites and proteins as well as neighboring information in the MPI networks, our MPI-VGAE predictor achieved the best predictive performance compared to other machine learning methods. Moreover, when applying the MPI-VGAE framework to reconstruct hundreds of metabolic pathways, functional enzymatic reaction networks and a metabolite-metabolite interaction network, our method showed the most robust performance among all scenarios. To the best of our knowledge, this is the first MPI predictor by VGAE for enzymatic reaction link prediction. Furthermore, we implemented the MPI-VGAE framework to reconstruct the disease-specific MPI network based on the disrupted metabolites and proteins in Alzheimer's disease and colorectal cancer, respectively. A substantial number of novel enzymatic reaction links were identified. We further validated and explored the interactions of these enzymatic reactions using molecular docking. These results highlight the potential of the MPI-VGAE framework for the discovery of novel disease-related enzymatic reactions and facilitate the study of the disrupted metabolisms in diseases. Chuang Yuan, Ranran Chen, Yuying Shi, Tao Zhang 0127, Fuzhong Xue, Gary J. Patti, Leyi Wei, Qingzhen Hou |
Briefings Bioinform. | 9 |
| 2023 | SiameseCPP: a sequence-based Siamese network to predict cell-penetrating peptides by contrastive learningabstractBACKGROUND: Cell-penetrating peptides (CPPs) have received considerable attention as a means of transporting pharmacologically active molecules into living cells without damaging the cell membrane, and thus hold great promise as future therapeutics. Recently, several machine learning-based algorithms have been proposed for predicting CPPs. However, most existing predictive methods do not consider the agreement (disagreement) between similar (dissimilar) CPPs and depend heavily on expert knowledge-based handcrafted features. RESULTS: In this study, we present SiameseCPP, a novel deep learning framework for automated CPPs prediction. SiameseCPP learns discriminative representations of CPPs based on a well-pretrained model and a Siamese neural network consisting of a transformer and gated recurrent units. Contrastive learning is used for the first time to build a CPP predictive model. Comprehensive experiments demonstrate that our proposed SiameseCPP is superior to existing baseline models for predicting CPPs. Moreover, SiameseCPP also achieves good performance on other functional peptide datasets, exhibiting satisfactory generalization ability. Lesong Wei, Xiucai Ye, Saisai Teng, Zhongshen Li, Junru Jin, Min Jae Kim, Tetsuya Sakurai, Li-Zhen Cui 0001, Balachandran Manavalan, Leyi Wei |
Briefings Bioinform. | 12 |
| 2023 | DeepProSite: structure-aware protein binding site prediction using ESMFold and pretrained language modelabstractMOTIVATION: Identifying the functional sites of a protein, such as the binding sites of proteins, peptides, or other biological components, is crucial for understanding related biological processes and drug design. However, existing sequence-based methods have limited predictive accuracy, as they only consider sequence-adjacent contextual features and lack structural information. RESULTS: In this study, DeepProSite is presented as a new framework for identifying protein binding site that utilizes protein structure and sequence information. DeepProSite first generates protein structures from ESMFold and sequence representations from pretrained language models. It then uses Graph Transformer and formulates binding site predictions as graph node classifications. In predicting protein-protein/peptide binding sites, DeepProSite outperforms state-of-the-art sequence- and structure-based methods on most metrics. Moreover, DeepProSite maintains its performance when predicting unbound structures, in contrast to competing structure-based prediction methods. DeepProSite is also extended to the prediction of binding sites for nucleic acids and other ligands, verifying its generalization capability. Finally, an online server for predicting multiple types of residue is established as the implementation of the proposed DeepProSite. AVAILABILITY AND IMPLEMENTATION: The datasets and source codes can be accessed at https://github.com/WeiLab-Biology/DeepProSite. The proposed DeepProSite can be accessed at https://inner.wei-group.net/DeepProSite/. Yitian Fang, Leyi Wei, Qin Ma 0003, Zhixiang Ren, Qianmu Yuan |
Bioinform. | 3 |
| 2023 | ExamPle: explainable deep learning framework for the prediction of plant small secreted peptidesabstractMOTIVATION: Plant Small Secreted Peptides (SSPs) play an important role in plant growth, development, and plant-microbe interactions. Therefore, the identification of SSPs is essential for revealing the functional mechanisms. Over the last few decades, machine learning-based methods have been developed, accelerating the discovery of SSPs to some extent. However, existing methods highly depend on handcrafted feature engineering, which easily ignores the latent feature representations and impacts the predictive performance. RESULTS: Here, we propose ExamPle, a novel deep learning model using Siamese network and multi-view representation for the explainable prediction of the plant SSPs. Benchmarking comparison results show that our ExamPle performs significantly better than existing methods in the prediction of plant SSPs. Also, our model shows excellent feature extraction ability. Importantly, by utilizing in silicomutagenesis experiment, ExamPle can discover sequential characteristics and identify the contribution of each amino acid for the predictions. The key novel principle learned by our model is that the head region of the peptide and some specific sequential patterns are strongly associated with the SSPs' functions. Thus, ExamPle is expected to be a useful tool for predicting plant SSPs and designing effective plant SSPs. AVAILABILITY AND IMPLEMENTATION: Our codes and datasets are available at https://github.com/Johnsunnn/ExamPle. Zhongshen Li, Junru Jin, Wentao Long, Yuanhao Ding, Leyi Wei |
Bioinform. | 7 |
| 2023 | A general hypergraph learning algorithm for drug multi-task predictions in micro-to-macro biomedical networksabstractThe powerful combination of large-scale drug-related interaction networks and deep learning provides new opportunities for accelerating the process of drug discovery. However, chemical structures that play an important role in drug properties and high-order relations that involve a greater number of nodes are not tackled in current biomedical networks. In this study, we present a general hypergraph learning framework, which introduces Drug-Substructures relationship into Molecular interaction Networks to construct the micro-to-macro drug centric heterogeneous network (DSMN), and develop a multi-branches HyperGraph learning model, called HGDrug, for Drug multi-task predictions. HGDrug achieves highly accurate and robust predictions on 4 benchmark tasks (drug-drug, drug-target, drug-disease, and drug-side-effect interactions), outperforming 8 state-of-the-art task specific models and 6 general-purpose conventional models. Experiments analysis verifies the effectiveness and rationality of the HGDrug model architecture as well as the multi-branches setup, and demonstrates that HGDrug is able to capture the relations between drugs associated with the same functional groups. In addition, our proposed drug-substructure interaction networks can help improve the performance of existing network models for drug-related prediction tasks. Shuting Jin, Yinghui Jiang, Leyi Wei, Zhuohang Yu, Xiangxiang Zeng, Xiangrong Liu |
PLoS Comput. Biol. | 6 |
| 2022 | Semi-supervised Histological Image Segmentation via Hierarchical Consistency Enforcement
Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiangbin Zheng 0001, Leyi Wei, Zhenyu Fang, Zhaopeng Meng, Ran Su |
MICCAI (2) | 5 |
| 2022 | Accelerating bioactive peptide discovery via mutual information-based meta-learningabstractRecently, machine learning methods have been developed to identify various peptide bio-activities. However, due to the lack of experimentally validated peptides, machine learning methods cannot provide a sufficiently trained model, easily resulting in poor generalizability. Furthermore, there is no generic computational framework to predict the bioactivities of different peptides. Thus, a natural question is whether we can use limited samples to build an effective predictive model for different kinds of peptides. To address this question, we propose Mutual Information Maximization Meta-Learning (MIMML), a novel meta-learning-based predictive model for bioactive peptide discovery. Using few samples from various functional peptides, MIMML can sufficiently learn the discriminative information amongst various functions and characterize functional differences. Experimental results show excellent performance of MIMML though using far fewer training samples as compared to the state-of-the-art methods. We also decipher the latent relationships among different kinds of functions to understand what meta-model learned to improve a specific task. In summary, this study is a pioneering work in the field of functional peptide mining and provides the first-of-its-kind solution for few-sample learning problems in biological sequence analysis, accelerating the new functional peptide discovery. The source codes and datasets are available on https://github.com/TearsWaiting/MIMML. Junru Jin, Zhongshen Li, Jiaojiao Zhao, Balachandran Manavalan, Ran Su, Xin Gao 0001, Leyi Wei |
Briefings Bioinform. | 9 |
| 2022 | SRDFM: Siamese Response Deep Factorization Machine to improve anti-cancer drug recommendationabstractPredicting the response of cancer patients to a particular treatment is a major goal of modern oncology and an important step toward personalized treatment. In the practical clinics, the clinicians prefer to obtain the most-suited drugs for a particular patient instead of knowing the exact values of drug sensitivity. Instead of predicting the exact value of drug response, we proposed a deep learning-based method, named Siamese Response Deep Factorization Machines (SRDFM) Network, for personalized anti-cancer drug recommendation, which directly ranks the drugs and provides the most effective drugs. A Siamese network (SN), a type of deep learning network that is composed of identical subnetworks that share the same architecture, parameters and weights, was used to measure the relative position (RP) between drugs for each cell line. Through minimizing the difference between the real RP and the predicted RP, an optimal SN model was established to provide the rank for all the candidate drugs. Specifically, the subnetwork in each side of the SN consists of a feature generation level and a predictor construction level. On the feature generation level, both drug property and gene expression, were adopted to build a concatenated feature vector, which even enables the recommendation for newly designed drugs with only chemical property known. Particularly, we developed a response unit here to generate weighted genetic feature vector to simulate the biological interaction mechanism between a specific drug and the genes. For the predictor construction level, we built this level integrating a factorization machine (FM) component with a deep neural network component. The FM can well handle the discrete chemical information and both low-order and high-order feature interactions could be sufficiently learned. Impressively, the SRDFM works well on both single-drug recommendation and synergic drug combination. Experiment result on both single-drug and synergetic drug data sets have shown the efficiency of the SRDFM. The Python implementation for the proposed SRDFM is available at at https://github.com/RanSuLab/SRDFM Contact: [email protected], [email protected] and [email protected]. Ran Su, Guobao Xiao, Leyi Wei |
Briefings Bioinform. | 5 |
| 2022 | Distant metastasis identification based on optimized graph representation of gene interaction patternsabstractMetastasis is a major cause of cancer morbidity and mortality, and most cancer deaths are caused by cancer metastasis rather than by the primary tumor. The prediction of metastasis based on computational methods has not been explored much in the previous research. In this study, we proposed a graph convolutional network embedded with a graph learning (GL) module, named glmGCN, to predict the distant metastasis of cancer. Both the mRNA and lncRNA expressions were used to provide more genetic information than using the mRNA alone and we used them to construct gene interaction graph representation to consider the effect of genetic interaction. Then, the prediction of the cancer metastasis was performed under a GCN framework, which extracted informative and advanced features from the built non-regular graph structures. Particularly, a GL module was embedded in the proposed glmGCN to learn an optimal graph representation of the gene interaction. We firstly constructed the protein-protein interaction network to represent the initial gene(node) relationship graph. Then, through the GL module, a new graph representation was built which optimally learned the gene interaction strength. Finally, the GCN was adopted to identify the distant metastasis cases. It is worth mentioning that the proposed method pays more attentions on the gene-gene relation than the previous GCN-based method, so more accurate prediction performance can be obtained. The glmGCN was trained based on two types of cancer and was further validated using two other cancer types. A series of experiments have shown that the effectiveness of the proposed method. The implementation for the proposed method is available at https://github.com/RanSuLab/Metastasis-glmGCN. Ran Su, Quan Zou 0001, Leyi Wei |
Briefings Bioinform. | 4 |
| 2022 | StackTADB: a stacking-based ensemble learning model for predicting the boundaries of topologically associating domains (TADs) accurately in fruit fliesabstractChromosome is composed of many distinct chromatin domains, referred to variably as topological domains or topologically associating domains (TADs). The domains are stable across different cell types and highly conserved across species, thus these chromatin domains have been considered as the basic units of chromosome folding and regarded as an important secondary structure in chromosome organization. However, the identification of TAD boundaries is still a great challenge due to the high cost and low resolution of Hi-C data or experiments. In this study, we propose a novel ensemble learning framework, termed as StackTADB, for predicting the boundaries of TADs. StackTADB integrates four base classifiers including Random Forest, Logistic Regression, K-NearestNeighbor and Support Vector Machine. From the analysis of a series of examinations on the data set in the previous study, it is concluded that StackTADB has optimal performance in six metrics, AUC, Accuracy, MCC, Precision, Recall and F1 score, and it is superior to the existing methods. In addition, the comparison of the performance of multiple features shows that Kmers-based features play an essential role in predicting TADs boundaries of fruit flies, and we also apply the SHapley Additive exPlanations (SHAP) framework to interpret the predictions of StackTADB to identify the reason why Kmers-based features are vital. The experimental results show that the subsequences matching the BEAF-32 motif play a crucial role in predicting the boundaries of TADs. The source code is freely available at https://github.com/HaoWuLab-Bioinformatics/StackTADB and the webserver of StackTADB is freely available at http://hwtad.sdu.edu.cn:8002/StackTADB. Hao Wu 0062, Zhaoheng Ai, Leyi Wei, Hongming Zhang 0002, Fan Yang 0068, Li-Zhen Cui 0001 |
Briefings Bioinform. | 4 |
| 2022 | Predicting protein-peptide binding residues via interpretable deep learningabstractSUMMARY: Identifying the protein-peptide binding residues is fundamentally important to understand the mechanisms of protein functions and explore drug discovery. Although several computational methods have been developed, most of them highly rely on third-party tools or complex data preprocessing for feature design, easily resulting in low computational efficacy and suffering from low predictive performance. To address the limitations, we propose PepBCL, a novel BERT (Bidirectional Encoder Representation from Transformers) -based contrastive learning framework to predict the protein-peptide binding residues based on protein sequences only. PepBCL is an end-to-end predictive model that is independent of feature engineering. Specifically, we introduce a well pre-trained protein language model that can automatically extract and learn high-latent representations of protein sequences relevant for protein structures and functions. Further, we design a novel contrastive learning module to optimize the feature representations of binding residues underlying the imbalanced dataset. We demonstrate that our proposed method significantly outperforms the state-of-the-art methods under benchmarking comparison, and achieves more robust performance. Moreover, we found that we further improve the performance via the integration of traditional features and our learnt features. Interestingly, the interpretable analysis of our model highlights the flexibility and adaptability of deep learning-based protein language model to capture both conserved and non-conserved sequential characteristics of peptide-binding residues. Finally, to facilitate the use of our method, we establish an online predictive platform as the implementation of the proposed PepBCL, which is now available at http://server.wei-group.net/PepBCL/. AVAILABILITY AND IMPLEMENTATION: https://github.com/Ruheng-W/PepBCL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ruheng Wang, Junru Jin, Quan Zou 0001, Kenta Nakai, Leyi Wei |
Bioinform. | 5 |
| 2022 | Multi-scale deep learning for the imbalanced multi-label protein subcellular localization prediction based on immunohistochemistry imagesabstractMOTIVATION: The development of microscopic imaging techniques enables us to study protein subcellular locations from the tissue level down to the cell level, contributing to the rapid development of image-based protein subcellular location prediction approaches. However, existing methods suffer from intrinsic limitations, such as poor feature representation ability, data imbalanced issue, and multi-label classification problem, greatly impacting the model performance and generalization. RESULTS: In this study, we propose MSTLoc, a novel multi-scale end-to-end deep learning model to identify protein subcellular locations in the imbalanced multi-label immunohistochemistry (IHC) images dataset. In our MSTLoc, we deploy a deep convolution neural network to extract multi-scale features from the IHC images, aggregate the high-level features and low-level features via feature fusion to sufficiently exploit the dependencies amongst various subcellular locations, and utilize Vision Transformer (ViT) to model the relationship amongst the features and enhance the feature representation ability. We demonstrate that the proposed MSTLoc achieves better performance than current state-of-the-art models in multi-label subcellular location prediction. Through feature visualization and interpretation analysis, we demonstrate that as compared with the hand-crafted features, the multi-scale deep features learnt from our model exhibit better ability in capturing discriminative patterns underlying protein subcellular locations, and the features from different scales are complementary for the improvement in performance. Finally, case study results indicate that our MSTLoc can successfully identify some biomarkers from proteins that are closely involved with cancer development. AVAILABILITY AND IMPLEMENTATION: For the convenient use of our method, we establish a user-friendly webserver available at http://server.wei-group.net/MSTLoc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fengsheng Wang, Leyi Wei |
Bioinform. | 2 |
| 2022 | ToxIBTL: prediction of peptide toxicity based on information bottleneck and transfer learningabstractMOTIVATION: Recently, peptides have emerged as a promising class of pharmaceuticals for various diseases treatment poised between traditional small molecule drugs and therapeutic proteins. However, one of the key bottlenecks preventing them from therapeutic peptides is their toxicity toward human cells, and few available algorithms for predicting toxicity are specially designed for short-length peptides. RESULTS: We present ToxIBTL, a novel deep learning framework by utilizing the information bottleneck principle and transfer learning to predict the toxicity of peptides as well as proteins. Specifically, we use evolutionary information and physicochemical properties of peptide sequences and integrate the information bottleneck principle into a feature representation learning scheme, by which relevant information is retained and the redundant information is minimized in the obtained features. Moreover, transfer learning is introduced to transfer the common knowledge contained in proteins to peptides, which aims to improve the feature representation capability. Extensive experimental results demonstrate that ToxIBTL not only achieves a higher prediction performance than state-of-the-art methods on the peptide dataset, but also has a competitive performance on the protein dataset. Furthermore, a user-friendly online web server is established as the implementation of the proposed ToxIBTL. AVAILABILITY AND IMPLEMENTATION: The proposed ToxIBTL and data can be freely accessible at http://server.wei-group.net/ToxIBTL. Our source code is available at https://github.com/WLYLab/ToxIBTL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lesong Wei, Xiucai Ye, Tetsuya Sakurai, Zengchao Mu, Leyi Wei |
Bioinform. | 5 |
| 2022 | EOCSA: Predicting prognosis of Epithelial ovarian cancer with whole slide histopathological images
Tianling Liu, Ran Su, Changming Sun, Xiu-Ting Li, Leyi Wei |
Expert Syst. Appl. | 5 |
| 2022 | A multi-label learning model for predicting drug-induced pathology in multi-organ based on toxicogenomics dataabstractDrug-induced toxicity damages the health and is one of the key factors causing drug withdrawal from the market. It is of great significance to identify drug-induced target-organ toxicity, especially the detailed pathological findings, which are crucial for toxicity assessment, in the early stage of drug development process. A large variety of studies have devoted to identify drug toxicity. However, most of them are limited to single organ or only binary toxicity. Here we proposed a novel multi-label learning model named Att-RethinkNet, for predicting drug-induced pathological findings targeted on liver and kidney based on toxicogenomics data. The Att-RethinkNet is equipped with a memory structure and can effectively use the label association information. Besides, attention mechanism is embedded to focus on the important features and obtain better feature presentation. Our Att-RethinkNet is applicable in multiple organs and takes account the compound type, dose, and administration time, so it is more comprehensive and generalized. And more importantly, it predicts multiple pathological findings at the same time, instead of predicting each pathology separately as the previous model did. To demonstrate the effectiveness of the proposed model, we compared the proposed method with a series of state-of-the-arts methods. Our model shows competitive performance and can predict potential hepatotoxicity and nephrotoxicity in a more accurate and reliable way. The implementation of the proposed method is available at https://github.com/RanSuLab/Drug-Toxicity-Prediction-MultiLabel. Ran Su, Haitang Yang, Leyi Wei, Siqi Chen 0001, Quan Zou 0001 |
PLoS Comput. Biol. | 3 |
| 2022 | Robust Feature Matching for Remote Sensing Image Registration via Guided Hyperplane FittingabstractFeature matching is a fundamental problem in feature-based remote sensing image registration. Due to the ground relief variations and imaging viewpoint changes, remote sensing images often involve local distortions, leading to difficulties in high-accuracy image registration. To address this issue, in this article, we propose a robust feature matching method called First Neighbor Relation Guided (FNRG) for remote sensing image registration via guided hyperplane fitting. The key idea of FNRG is to exploit the first neighbor relation of feature points between two images for seeking consistent seeds in a parameter-free manner. To boost more consistent matches based on the consistent seeds, we formulate the feature matching problem into an affine hyperplane fitting problem by imposing the motion consistency, and then we design a hyperplane updating strategy to refine the fitting model. We also introduce a locality preserving structure-based cost function to promote the matching performance of the hyperplane updating strategy. Our method can mine consistent matches from thousands of putative ones within only a few milliseconds, and it also can handle the data with a large-scale change, rotation, or severe nonrigid deformation. Extensive experiments on the remote sensing image data sets with different types of image transformations show that the proposed method achieves significant superiority over several state-of-the-art methods. Guobao Xiao, Huan Luo 0001, Leyi Wei, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Integrative machine learning framework for the identification of cell-specific enhancers from the human genomeabstractEnhancers are deoxyribonucleic acid (DNA) fragments which when bound by transcription factors enhance the transcription of related genes. Due to its sporadic distribution and similar fractions, identification of enhancers from the human genome seems a daunting task. Compared to the traditional experimental approaches, computational methods with easy-to-use platforms could be efficiently applied to annotate enhancers' functions and physiological roles. In this aspect, several bioinformatics tools have been developed to identify enhancers. Despite their spectacular performances, existing methods have certain drawbacks and limitations, including fixed length of sequences being utilized for model development and cell-specificity negligence. A novel predictor would be beneficial in the context of genome-wide enhancer prediction by addressing the above-mentioned issues. In this study, we constructed new datasets for eight different cell types. Utilizing these data, we proposed an integrative machine learning (ML)-based framework called Enhancer-IF for identifying cell-specific enhancers. Enhancer-IF comprehensively explores a wide range of heterogeneous features with five commonly used ML methods (random forest, extremely randomized tree, multilayer perceptron, support vector machine and extreme gradient boosting). Specifically, these five classifiers were trained with seven encodings and obtained 35 baseline models. The output of these baseline models was integrated and again inputted to five classifiers for the construction of five meta-models. Finally, the integration of five meta-models through ensemble learning improved the model robustness. Our proposed approach showed an excellent prediction performance compared to the baseline models on both training and independent datasets in different cell types, thus highlighting the superiority of our approach in the identification of the enhancers. We assume that Enhancer-IF will be a valuable tool for screening and identifying potential enhancers from the human DNA sequences. Shaherin Basith, Md. Mehedi Hasan 0002, Gwang Lee, Leyi Wei, Balachandran Manavalan |
Briefings Bioinform. | 4 |
| 2021 | PSSP-MVIRT: peptide secondary structure prediction based on a multi-view deep learning architectureabstractThe prediction of peptide secondary structures is fundamentally important to reveal the functional mechanisms of peptides with potential applications as therapeutic molecules. In this study, we propose a multi-view deep learning method named Peptide Secondary Structure Prediction based on Multi-View Information, Restriction and Transfer learning (PSSP-MVIRT) for peptide secondary structure prediction. To sufficiently exploit discriminative information, we introduce a multi-view fusion strategy to integrate different information from multiple perspectives, including sequential information, evolutionary information and hidden state information, respectively, and generate a unified feature space. Moreover, we construct a hybrid network architecture of Convolutional Neural Network and Bi-directional Gated Recurrent Unit to extract global and local features of peptides. Furthermore, we utilize transfer learning to effectively alleviate the lack of training samples (peptides with experimentally validated structures). Comparative results on independent tests demonstrate that our proposed method significantly outperforms state-of-the-art methods. In particular, our method exhibits better performance at the segment level, suggesting the strong ability of our model in capturing local discriminative information. The case study also shows that our PSSP-MVIRT achieves promising and robust performance in the prediction of new peptide secondary structures. Importantly, we establish a webserver to implement the proposed method, which is currently accessible via http://server.malab.cn/PSSP-MVIRT. We expect it can be a useful tool for the researchers of interest, facilitating the wide use of our method. Xiao Cao, Zitan Chen, Lesong Wei, Li-Zhen Cui 0001, Ran Su, Leyi Wei |
Briefings Bioinform. | 10 |
| 2021 | Iterative feature representation algorithm to improve the predictive performance of N7-methylguanosine sitesabstractMOTIVATION: N7-methylguanosine (m7G) is an important epigenetic modification, playing an essential role in gene expression regulation. Therefore, accurate identification of m7G modifications will facilitate revealing and in-depth understanding their potential functional mechanisms. Although high-throughput experimental methods are capable of precisely locating m7G sites, they are still cost ineffective. Therefore, it's necessary to develop new methods to identify m7G sites. RESULTS: In this work, by using the iterative feature representation algorithm, we developed a machine learning based method, namely m7G-IFL, to identify m7G sites. To demonstrate its superiority, m7G-IFL was evaluated and compared with existing predictors. The results demonstrate that our predictor outperforms existing predictors in terms of accuracy for identifying m7G sites. By analyzing and comparing the features used in the predictors, we found that the positive and negative samples in our feature space were more separated than in existing feature space. This result demonstrates that our features extracted more discriminative information via the iterative feature learning process, and thus contributed to the predictive performance improvement. Chichi Dai, Pengmian Feng, Li-Zhen Cui 0001, Ran Su, Wei Chen 0064, Leyi Wei |
Briefings Bioinform. | 6 |
| 2021 | EP3: an ensemble predictor that accurately identifies type III secreted effectorsabstractType III secretion systems (T3SS) can be found in many pathogenic bacteria, such as Dysentery bacillus, Salmonella typhimurium, Vibrio cholera and pathogenic Escherichia coli. The routes of infection of these bacteria include the T3SS transferring a large number of type III secreted effectors (T3SE) into host cells, thereby blocking or adjusting the communication channels of the host cells. Therefore, the accurate identification of T3SEs is the precondition for the further study of pathogenic bacteria. In this article, a new T3SEs ensemble predictor was developed, which can accurately distinguish T3SEs from any unknown protein. In the course of the experiment, methods and models are strictly trained and tested. Compared with other methods, EP3 demonstrates better performance, including the absence of overfitting, strong robustness and powerful predictive ability. EP3 (an ensemble predictor that accurately identifies T3SEs) is designed to simplify the user's (especially nonprofessional users) access to T3SEs for further investigation, which will have a significant impact on understanding the progression of pathogenic bacterial infections. Based on the integrated model that we proposed, a web server had been established to distinguish T3SEs from non-T3SEs, where have EP3_1 and EP3_2. The users can choose the model according to the species of the samples to be tested. Our related tools and data can be accessed through the link http://lab.malab.cn/∼lijing/EP3.html. Leyi Wei, Fei Guo 0001, Quan Zou 0001 |
Briefings Bioinform. | 2 |
| 2021 | Classification and gene selection of triple-negative breast cancer subtype embedding gene connectivity matrix in deep neural networkabstractTriple-negative breast cancer (TNBC) has been a challenging breast cancer subtype for oncological therapy. Normally, it can be classified into different molecular subtypes. Accurate and stable classification of the six subtypes is essential for personalized treatment of TNBC. In this study, we proposed a new framework to distinguish the six subtypes of TNBC, and this is one of the handful studies that completed the classification based on mRNA and long noncoding RNA expression data. Particularly, we developed a gene selection approach named DGGA, which takes correlation information between genes into account in the process of measuring gene importance and then effectively removes redundant genes. A gene scoring approach that combined GeneRank scores with gene importance generated by deep neural network (DNN), taking inter-subtype discrimination and inner-gene correlations into account, was came up to improve gene selection performance. More importantly, we embedded a gene connectivity matrix in the DNN for sparse learning, which takes additional consideration with weight changes during training when obtaining the measurement of the relative importance of each gene. Finally, Genetic Algorithm was used to simulate the natural evolutionary process to search for the optimal subset of TNBC subtype classification. We validated the proposed method through cross-validation, and the results demonstrate that it can use fewer genes to obtain more accurate classification results. The implementation for the proposed method is available at https://github.com/RanSuLab/TNBC. Ran Su, Leyi Wei |
Briefings Bioinform. | 4 |
| 2021 | Protein subcellular localization based on deep image features and criterion learning strategyabstractThe spatial distribution of proteome at subcellular levels provides clues for protein functions, thus is important to human biology and medicine. Imaging-based methods are one of the most important approaches for predicting protein subcellular location. Although deep neural networks have shown impressive performance in a number of imaging tasks, its application to protein subcellular localization has not been sufficiently explored. In this study, we developed a deep imaging-based approach to localize the proteins at subcellular levels. Based on deep image features extracted from convolutional neural networks (CNNs), both single-label and multi-label locations can be accurately predicted. Particularly, the multi-label prediction is quite a challenging task. Here we developed a criterion learning strategy to exploit the label-attribute relevancy and label-label relevancy. A criterion that was used to determine the final label set was automatically obtained during the learning procedure. We concluded an optimal CNN architecture that could give the best results. Besides, experiments show that compared with the hand-crafted features, the deep features present more accurate prediction with less features. The implementation for the proposed method is available at https://github.com/RanSuLab/ProteinSubcellularLocation. Ran Su, Linlin He, Tianling Liu, Xiaofeng Liu 0004, Leyi Wei |
Briefings Bioinform. | 5 |
| 2021 | Predicting drug-induced hepatotoxicity based on biological feature maps and diverse classification strategiesabstractIdentifying hepatotoxicity as early as possible is significant in drug development. In this study, we developed a drug-induced hepatotoxicity prediction model taking account of both the biological context and the computational efficacy based on toxicogenomics data. Specifically, we proposed a novel gene selection algorithm considering gene's participation, named BioCB, to choose the discriminative genes and make more efficient prediction. Then instead of using the raw gene expression levels to characterize each drug, we developed a two-dimensional biological process feature pattern map to represent each drug. Then we employed two strategies to handle the maps and identify the hepatotoxicity, the direct use of maps, named Two-dim branch, and vectorization of maps, named One-dim branch. The two strategies subsequently used the deep convolutional neural networks and LightGBM as predictors, respectively. Additionally, we here for the first time proposed a stacked vectorized gene matrix, which was more predictive than the raw gene matrix. Results validated on both in vivo and in vitro data from two public data sets, the TG-GATES and DrugMatrix, show that the proposed One-dim branch outperforms the deep framework, the Two-dim branch, and has achieved high accuracy and efficiency. The implementation of the proposed method is available at https://github.com/RanSuLab/Hepatotoxicity. Ran Su, Huichen Wu, Leyi Wei |
Briefings Bioinform. | 4 |
| 2021 | Computational prediction and interpretation of cell-specific replication origin sites from multiple eukaryotes by exploiting stacking frameworkabstractOrigins of replication sites (ORIs), which refers to the initiative locations of genomic DNA replication, play essential roles in DNA replication process. Detection of ORIs' distribution in genome scale is one of key steps to in-depth understanding their regulation mechanisms. In this study, we presented a novel machine learning-based approach called Stack-ORI encompassing 10 cell-specific prediction models for identifying ORIs from four different eukaryotic species (Homo sapiens, Mus musculus, Drosophila melanogaster and Arabidopsis thaliana). For each cell-specific model, we employed 12 feature encoding schemes that cover nucleic acid composition, position-specific and physicochemical properties information. The optimal feature set was identified from each encoding individually and developed their respective baseline models using the eXtreme Gradient Boosting (XGBoost) classifier. Subsequently, the predicted scores of 12 baseline models are integrated as a novel feature vector to train XGBoost and develop the final model. Extensive experimental results show that Stack-ORI achieves significantly better performance as compared with their baseline models on both training and independent datasets. Interestingly, Stack-ORI consistently outperforms existing predictor in all cell-specific models, not only on training but also on independent test. Moreover, our novel approach provides necessary interpretations that help understanding model success by leveraging the powerful SHapley Additive exPlanation algorithm, thus underlining the most important feature encoding schemes significant for predicting cell-specific ORIs. Leyi Wei, Adeel Malik, Ran Su, Li-Zhen Cui 0001, Balachandran Manavalan |
Briefings Bioinform. | 1 |
| 2021 | ATSE: a peptide toxicity predictor by exploiting structural and evolutionary information based on graph neural network and attention mechanismabstractMOTIVATION: Peptides have recently emerged as promising therapeutic agents against various diseases. For both research and safety regulation purposes, it is of high importance to develop computational methods to accurately predict the potential toxicity of peptides within the vast number of candidate peptides. RESULTS: In this study, we proposed ATSE, a peptide toxicity predictor by exploiting structural and evolutionary information based on graph neural networks and attention mechanism. More specifically, it consists of four modules: (i) a sequence processing module for converting peptide sequences to molecular graphs and evolutionary profiles, (ii) a feature extraction module designed to learn discriminative features from graph structural information and evolutionary information, (iii) an attention module employed to optimize the features and (iv) an output module determining a peptide as toxic or non-toxic, using optimized features from the attention module. CONCLUSION: Comparative studies demonstrate that the proposed ATSE significantly outperforms all other competing methods. We found that structural information is complementary to the evolutionary information, effectively improving the predictive performance. Importantly, the data-driven features learned by ATSE can be interpreted and visualized, providing additional information for further analysis. Moreover, we present a user-friendly online computational platform that implements the proposed ATSE, which is now available at http://server.malab.cn/ATSE. We expect that it can be a powerful and useful tool for researchers of interest. Lesong Wei, Xiucai Ye, Yuyang Xue, Tetsuya Sakurai, Leyi Wei |
Briefings Bioinform. | 5 |
| 2021 | Learning embedding features based on multisense-scaled attention architecture to improve the predictive performance of anticancer peptidesabstractMOTIVATION: Anticancer peptides (ACPs) have recently emerged as effective anticancer drugs in cancer therapy. Machine learning-based predictors have been developed to identify ACPs and achieve satisfactory performance. However, existing methods suffer from experience-based feature engineering, which not only restricts the representation ability of the models to a certain extent but also lacks adaptivity for different data, limiting the further improvement of the predictive performance and impacting the robustness of the predictive models. To alleviate the above problems, we propose a novel deep-learning-based predictor named ACPred-LAF, in which we propose a novel multisense and multiscaled embedding algorithm to automatically learn and extract context sequential characteristics of ACPs. RESULTS: Through the feature comparative analysis, we demonstrate that our learnable and self-adaptive embedding features are better than hand-crafted features in capturing discriminative information, which can effectively benefit the performance improvement for ACP prediction. In addition, benchmarking comparison results demonstrate that our ACPred-LAF outperforms the state-of-the-art methods both on existing benchmark datasets and our newly constructed dataset. Furthermore, we also prove and validate the robustness of the model via the data interference experiment. To avoid potential evaluation bias, here, we construct a new ACP benchmark dataset named ACP-Mixed by integrating existing datasets. We expect our newly constructed dataset to be a golden standard benchmark dataset in this field. To facilitate the use of our model, we develop a web server as the implementation of ACPred-LAF. AVAILABILITY AND IMPLEMENTATION: Our proposed ACPred-LAF, newly constructed benchmark dataset ACP-Mixed are open source collaborative initiatives available in the GitHub repository (https://github.com/TearsWaiting/ACPred-LAF). Besides, a webserver as the implementation of ACPred-LAF that can be accessed via: http://server.malab.cn/ACPred-LAF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Li-Zhen Cui 0001, Ran Su, Leyi Wei |
Bioinform. | 5 |
| 2021 | iDNA-ABT: advanced deep learning model for detecting DNA methylation with adaptive features and transductive information maximizationabstractMOTIVATION: DNA methylation plays an important role in epigenetic modification, the occurrence, and the development of diseases. Therefore, identification of DNA methylation sites is critical for better understanding and revealing their functional mechanisms. To date, several machine learning and deep learning methods have been developed for the prediction of different DNA methylation types. However, they still highly rely on manual features, which can largely limit the high-latent information extraction. Moreover, most of them are designed for one specific DNA methylation type, and therefore cannot predict multiple methylation sites in multiple species simultaneously. In this study, we propose iDNA-ABT, an advanced deep learning model that utilizes adaptive embedding based on Bidirectional Encoder Representations from Transformers (BERT) together with transductive information maximization (TIM). RESULTS: Benchmark results show that our proposed iDNA-ABT can automatically and adaptively learn the distinguishing features of biological sequences from multiple species, and thus perform significantly better than the state-of-the-art methods in predicting three different DNA methylation types. In addition, TIM loss is proven to be effective in dichotomous tasks via the comparison experiment. Furthermore, we verify that our features have strong adaptability and robustness to different species through comparison of adaptive embedding and six handcrafted feature encodings. Importantly, our model shows great generalization ability in different species, demonstrating that our model can adaptively capture the cross-species differences and improve the predictive performance. For the convenient use of our method, we further established an online webserver as the implementation of the proposed iDNA-ABT. AVAILABILITY AND IMPLEMENTATION: Our proposed iDNA-ABT and data are freely accessible via http://server.wei-group.net/iDNA_ABT and our source codes are available for downloading in the GitHub repository (https://github.com/YUYING07/iDNA_ABT). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Junru Jin, Guobao Xiao, Li-Zhen Cui 0001, Rao Zeng, Leyi Wei |
Bioinform. | 7 |
| 2021 | FEGS: a novel feature extraction model for protein sequences and its applicationsabstractBACKGROUND: Feature extraction of protein sequences is widely used in various research areas related to protein analysis, such as protein similarity analysis and prediction of protein functions or interactions. RESULTS: In this study, we introduce FEGS (Feature Extraction based on Graphical and Statistical features), a novel feature extraction model of protein sequences, by developing a new technique for graphical representation of protein sequences based on the physicochemical properties of amino acids and effectively employing the statistical features of protein sequences. By fusing the graphical and statistical features, FEGS transforms a protein sequence into a 578-dimensional numerical vector. When FEGS is applied to phylogenetic analysis on five protein sequence data sets, its performance is notably better than all of the other compared methods. CONCLUSION: The FEGS method is carefully designed, which is practically powerful for extracting features of protein sequences. The current version of FEGS is developed to be user-friendly and is expected to play a crucial role in the related studies of protein sequence analyses. Zengchao Mu, Ting Yu 0010, Xiaoping Liu 0002, Leyi Wei |
BMC Bioinform. | 5 |
| 2021 | Domain adaptation based self-correction model for COVID-19 infection segmentation in CT images
Qiangguo Jin, Hui Cui 0002, Changming Sun, Zhaopeng Meng, Leyi Wei, Ran Su |
Expert Syst. Appl. | 5 |
| 2021 | Identification of glioblastoma molecular subtype and prognosis based on deep MRI features
Ran Su, Qiangguo Jin, Xiaofeng Liu 0004, Leyi Wei |
Knowl. Based Syst. | 5 |
| 2020 | PRISMOID: a comprehensive 3D structure database for post-translational modifications and mutations with functional impactabstractPost-translational modifications (PTMs) play very important roles in various cell signaling pathways and biological process. Due to PTMs' extremely important roles, many major PTMs have been studied, while the functional and mechanical characterization of major PTMs is well documented in several databases. However, most currently available databases mainly focus on protein sequences, while the real 3D structures of PTMs have been largely ignored. Therefore, studies of PTMs 3D structural signatures have been severely limited by the deficiency of the data. Here, we develop PRISMOID, a novel publicly available and free 3D structure database for a wide range of PTMs. PRISMOID represents an up-to-date and interactive online knowledge base with specific focus on 3D structural contexts of PTMs sites and mutations that occur on PTMs and in the close proximity of PTM sites with functional impact. The first version of PRISMOID encompasses 17 145 non-redundant modification sites on 3919 related protein 3D structure entries pertaining to 37 different types of PTMs. Our entry web page is organized in a comprehensive manner, including detailed PTM annotation on the 3D structure and biological information in terms of mutations affecting PTMs, secondary structure features and per-residue solvent accessibility features of PTM sites, domain context, predicted natively disordered regions and sequence alignments. In addition, high-definition JavaScript packages are employed to enhance information visualization in PRISMOID. PRISMOID equips a variety of interactive and customizable search options and data browsing functions; these capabilities allow users to access data via keyword, ID and advanced options combination search in an efficient and user-friendly way. A download page is also provided to enable users to download the SQL file, computational structural features and PTM sites' data. We anticipate PRISMOID will swiftly become an invaluable online resource, assisting both biologists and bioinformaticians to conduct experiments and develop applications supporting discovery efforts in the sequence-structural-functional relationship of PTMs and providing important insight into mutations and PTM sites interaction mechanisms. The PRISMOID database is freely accessible at http://prismoid.erc.monash.edu/. The database and web interface are implemented in MySQL, JSP, JavaScript and HTML with all major browsers supported. Fuyi Li, Cunshuo Fan, Tatiana T. Marquez-Lago, André Leier, Jerico Revote, Cangzhi Jia, Yan Zhu 0006, Alexander Ian Smith, Geoffrey I. Webb, Quanzhong Liu, Leyi Wei, Jian Li 0052, Jiangning Song |
Briefings Bioinform. | 11 |
| 2020 | CPPred-FL: a sequence-based predictor for large-scale identification of cell-penetrating peptides by feature representation learningabstractCell-penetrating peptides (CPPs) have been shown to be a transport vehicle for delivering cargoes into live cells, offering great potential as future therapeutics. It is essential to identify CPPs for better understanding of their functional mechanisms. Machine learning-based methods have recently emerged as a main approach for computational identification of CPPs. However, one of the main challenges and difficulties is to propose an effective feature representation model that sufficiently exploits the inner difference and relevance between CPPs and non-CPPs, in order to improve the predictive performance. In this paper, we have developed CPPred-FL, a powerful bioinformatics tool for fast, accurate and large-scale identification of CPPs. In our predictor, we introduce a new feature representation learning scheme that enables one to learn feature representations from totally 45 well-trained random forest models with multiple feature descriptors from different perspectives, such as compositional information, position-specific information and physicochemical properties, etc. We integrate class and probabilistic information into our feature representations. To improve the feature representation ability, we further remove redundant and irrelevant features by feature space optimization. Benchmarking experiments showed that CPPred-FL, using 19 informative features only, is able to achieve better performance than the state-of-the-art predictors. We anticipate that CPPred-FL will be a powerful tool for large-scale identification of CPPs, facilitating the characterization of their functional mechanisms and accelerating their applications in clinical therapy. Xiaoli Qiang, Xiucai Ye, Pufeng Du, Ran Su, Leyi Wei |
Briefings Bioinform. | 6 |
| 2020 | ACPred-Fuse: fusing multi-view information improves the prediction of anticancer peptidesabstractFast and accurate identification of the peptides with anticancer activity potential from large-scale proteins is currently a challenging task. In this study, we propose a new machine learning predictor, namely, ACPred-Fuse, that can automatically and accurately predict protein sequences with or without anticancer activity in peptide form. Specifically, we establish a feature representation learning model that can explore class and probabilistic information embedded in anticancer peptides (ACPs) by integrating a total of 29 different sequence-based feature descriptors. In order to make full use of various multiview information, we further fused the class and probabilistic features with handcrafted sequential features and then optimized the representation ability of the multiview features, which are ultimately used as input for training our prediction model. By comparing the multiview features and existing feature descriptors, we demonstrate that the fused multiview features have more discriminative ability to capture the characteristics of ACPs. In addition, the information from different views is complementary for the performance improvement. Finally, our benchmarking comparison results showed that the proposed ACPred-Fuse is more precise and promising in the identification of ACPs than existing predictors. To facilitate the use of the proposed predictor, we built a web server, which is now freely available via http://server.malab.cn/ACPred-Fuse. Bing Rao, Guoying Zhang, Ran Su, Leyi Wei |
Briefings Bioinform. | 5 |
| 2020 | Empirical comparison and analysis of web-based cell-penetrating peptide prediction toolsabstractCell-penetrating peptides (CPPs) facilitate the delivery of therapeutically relevant molecules, including DNA, proteins and oligonucleotides, into cells both in vitro and in vivo. This unique ability explores the possibility of CPPs as therapeutic delivery and its potential applications in clinical therapy. Over the last few decades, a number of machine learning (ML)-based prediction tools have been developed, and some of them are freely available as web portals. However, the predictions produced by various tools are difficult to quantify and compare. In particular, there is no systematic comparison of the web-based prediction tools in performance, especially in practical applications. In this work, we provide a comprehensive review on the biological importance of CPPs, CPP database and existing ML-based methods for CPP prediction. To evaluate current prediction tools, we conducted a comparative study and analyzed a total of 12 models from 6 publicly available CPP prediction tools on 2 benchmark validation sets of CPPs and non-CPPs. Our benchmarking results demonstrated that a model from the KELM-CPPpred, namely KELM-hybrid-AAC, showed a significant improvement in overall performance, when compared to the other 11 prediction models. Moreover, through a length-dependency analysis, we find that existing prediction tools tend to more accurately predict CPPs and non-CPPs with the length of 20-25 residues long than peptides in other length ranges. Ran Su, Quan Zou 0001, Balachandran Manavalan, Leyi Wei |
Briefings Bioinform. | 5 |
| 2020 | MinE-RFE: determine the optimal subset from RFE by minimizing the subset-accuracy-defined energyabstractRecursive feature elimination (RFE), as one of the most popular feature selection algorithms, has been extensively applied to bioinformatics. During the training, a group of candidate subsets are generated by iteratively eliminating the least important features from the original features. However, how to determine the optimal subset from them still remains ambiguous. Among most current studies, either overall accuracy or subset size (SS) is used to select the most predictive features. Using which one or both and how they affect the prediction performance are still open questions. In this study, we proposed MinE-RFE, a novel RFE-based feature selection approach by sufficiently considering the effect of both factors. Subset decision problem was reflected into subset-accuracy space and became an energy-minimization problem. We also provided a mathematical description of the relationship between the overall accuracy and SS using Gaussian Mixture Models together with spline fitting. Besides, we comprehensively reviewed a variety of state-of-the-art applications in bioinformatics using RFE. We compared their approaches of deciding the final subset from all the candidate subsets with MinE-RFE on diverse bioinformatics data sets. Additionally, we also compared MinE-RFE with some well-used feature selection algorithms. The comparative results demonstrate that the proposed approach exhibits the best performance among all the approaches. To facilitate the use of MinE-RFE, we further established a user-friendly web server with the implementation of the proposed approach, which is accessible at http://qgking.wicp.net/MinE/. We expect this web server will be a useful tool for research community. Ran Su, Leyi Wei |
Briefings Bioinform. | 3 |
| 2020 | Meta-GDBP: a high-level stacked regression model to improve anticancer drug response predictionabstractAnticancer drug response prediction plays an important role in personalized medicine. In particular, precisely predicting drug response in specific cancer types and patients is still a challenge problem. Here we propose Meta-GDBP, a novel anticancer drug-response model, which involves two levels. At the first level of Meta-GDBP, we build four optimized base models (BMs) using genetic information, chemical properties and biological context with an ensemble optimization strategy, while at the second level, we construct a weighted model to integrate the four BMs. Notably, the weights of the models are learned upstream, thus the parameter cost is significantly reduced compared to previous methods. We evaluate the Meta-GDBP on Genomics of Drug Sensitivity in Cancer (GDSC) and the Cancer Cell Line Encyclopedia (CCLE) data sets. Benchmarking results demonstrate that compared to other methods, the Meta-GDBP achieves a much higher correlation between the predicted and the observed responses for almost all the drugs. Moreover, we apply the Meta-GDBP to predict the GDSC-missing drug response and use the CCLE-known data to validate the performance. The results show quite a similar tendency between these two response sets. Particularly, we here for the first time introduce a biological context-based frequency matrix (BCFM) to associate the biological context with the drug response. It is encouraging that the proposed BCFM is biologically meaningful and consistent with the reported biological mechanism, further demonstrating its efficacy for predicting drug response. The R implementation for the proposed Meta-GDBP is available at https://github.com/RanSuLab/Meta-GDBP. Ran Su, Guobao Xiao, Leyi Wei |
Briefings Bioinform. | 4 |
| 2020 | Comparative analysis and prediction of quorum-sensing peptides using feature representation learning and machine learning algorithmsabstractQuorum-sensing peptides (QSPs) are the signal molecules that are closely associated with diverse cellular processes, such as cell-cell communication, and gene expression regulation in Gram-positive bacteria. It is therefore of great importance to identify QSPs for better understanding and in-depth revealing of their functional mechanisms in physiological processes. Machine learning algorithms have been developed for this purpose, showing the great potential for the reliable prediction of QSPs. In this study, several sequence-based feature descriptors for peptide representation and machine learning algorithms are comprehensively reviewed, evaluated and compared. To effectively use existing feature descriptors, we used a feature representation learning strategy that automatically learns the most discriminative features from existing feature descriptors in a supervised way. Our results demonstrate that this strategy is capable of effectively capturing the sequence determinants to represent the characteristics of QSPs, thereby contributing to the improved predictive performance. Furthermore, wrapping this feature representation learning strategy, we developed a powerful predictor named QSPred-FL for the detection of QSPs in large-scale proteomic data. Benchmarking results with 10-fold cross validation showed that QSPred-FL is able to achieve better performance as compared to the state-of-the-art predictors. In addition, we have established a user-friendly webserver that implements QSPred-FL, which is currently available at http://server.malab.cn/QSPred-FL. We expect that this tool will be useful for the high-throughput prediction of QSPs and the discovery of important functional mechanisms of QSPs. Leyi Wei, Fuyi Li, Jiangning Song, Ran Su, Quan Zou 0001 |
Briefings Bioinform. | 1 |
| 2020 | Identifying enhancer-promoter interactions with neural network based on pre-trained DNA vectors and attention mechanismabstractMOTIVATION: Identification of enhancer-promoter interactions (EPIs) is of great significance to human development. However, experimental methods to identify EPIs cost too much in terms of time, manpower and money. Therefore, more and more research efforts are focused on developing computational methods to solve this problem. Unfortunately, most existing computational methods require a variety of genomic data, which are not always available, especially for a new cell line. Therefore, it limits the large-scale practical application of methods. As an alternative, computational methods using sequences only have great genome-scale application prospects. RESULTS: In this article, we propose a new deep learning method, namely EPIVAN, that enables predicting long-range EPIs using only genomic sequences. To explore the key sequential characteristics, we first use pre-trained DNA vectors to encode enhancers and promoters; afterwards, we use one-dimensional convolution and gated recurrent unit to extract local and global features; lastly, attention mechanism is used to boost the contribution of key features, further improving the performance of EPIVAN. Benchmarking comparisons on six cell lines show that EPIVAN performs better than state-of-the-art predictors. Moreover, we build a general model, which has transfer ability and can be used to predict EPIs in various cell lines. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at: https://github.com/hzy95/EPIVAN. Zengyan Hong, Xiangxiang Zeng, Leyi Wei, Xiangrong Liu |
Bioinform. | 3 |
| 2020 | Identification of expression signatures for non-small-cell lung carcinoma subtype classificationabstractMOTIVATION: Non-small-cell lung carcinoma (NSCLC) mainly consists of two subtypes: lung squamous cell carcinoma (LUSC) and lung adenocarcinoma (LUAD). It has been reported that the genetic and epigenetic profiles vary strikingly between LUAD and LUSC in the process of tumorigenesis and development. Efficient and precise treatment can be made if subtypes can be identified correctly. Identification of discriminative expression signatures has been explored recently to aid the classification of NSCLC subtypes. RESULTS: In this study, we designed a classification model integrating both mRNA and long non-coding RNA (lncRNA) expression data to effectively classify the subtypes of NSCLC. A gene selection algorithm, named WGRFE, was proposed to identify the most discriminative gene signatures within the recursive feature elimination (RFE) framework. GeneRank scores considering both expression level and correlation, together with the importance generated by classifiers were all taken into account to improve the selection performance. Moreover, a module-based initial filtering of the genes was performed to reduce the computation cost of RFE. We validated the proposed algorithm on The Cancer Genome Atlas (TCGA) dataset. The results demonstrate that the developed approach identified a small number of expression signatures for accurate subtype classification and particularly, we here for the first time show the potential role of LncRNA in building computational NSCLC subtype classification models. AVAILABILITY AND IMPLEMENTATION: The R implementation for the proposed approach is available at https://github.com/RanSuLab/NSCLC-subtype-classification. Ran Su, Xiaofeng Liu 0004, Leyi Wei |
Bioinform. | 4 |
| 2020 | Fusing convolutional neural network features with hand-crafted features for osteoporosis diagnoses
Ran Su, Tianling Liu, Changming Sun, Qiangguo Jin, Rachid Jennane, Leyi Wei |
Neurocomputing | 6 |
| 2019 | LncPred-IEL: A Long Non-coding RNA Prediction Method using Iterative Ensemble LearningabstractA large number of transcripts have been generated by the development of high throughput sequencing technologies. Predicting lncRNA from transcripts is a challenging and important task. In this paper, we propose LncPred-IEL, an iterative ensemble learning long non-coding RNA prediction method. LncPred-IEL not only considers features widely used for the lncRNA prediction, but also take into account sequence-derived features used in the RNA sequence classification, so as to make use of diverse information. LncPred-IEL builds base predictors based on different groups of features, and employs a supervised iterative way to combine base predictors and build ensemble models. Our studies demonstrate that supervised iterative way can learn the representations that help to separate lncRNA and protein-coding transcripts, and further improve the performances. Experiments demonstrate that LncPred-IEL outperforms several state-of-the-art methods when evaluated by 10-fold cross-validation. The capability of LncPred-IEL for the cross-species prediction is also tested. As complementary to wet experiments, LncPred-IEL is a useful computational tool for lncRNA prediction. Yanzhen Xu, Xiaohan Zhao, Shuai Liu 0017, Shichao Liu 0002, Yanqing Niu, Wen Zhang 0008, Leyi Wei |
BIBM | 7 |
| 2019 | mAHTPred: a sequence-based meta-predictor for improving the prediction of anti-hypertensive peptides using effective feature representationabstractMOTIVATION: Cardiovascular disease is the primary cause of death globally accounting for approximately 17.7 million deaths per year. One of the stakes linked with cardiovascular diseases and other complications is hypertension. Naturally derived bioactive peptides with antihypertensive activities serve as promising alternatives to pharmaceutical drugs. So far, there is no comprehensive analysis, assessment of diverse features and implementation of various machine-learning (ML) algorithms applied for antihypertensive peptide (AHTP) model construction. RESULTS: In this study, we utilized six different ML algorithms, namely, Adaboost, extremely randomized tree (ERT), gradient boosting (GB), k-nearest neighbor, random forest (RF) and support vector machine (SVM) using 51 feature descriptors derived from eight different feature encodings for the prediction of AHTPs. While ERT-based trained models performed consistently better than other algorithms regardless of various feature descriptors, we treated them as baseline predictors, whose predicted probability of AHTPs was further used as input features separately for four different ML-algorithms (ERT, GB, RF and SVM) and developed their corresponding meta-predictors using a two-step feature selection protocol. Subsequently, the integration of four meta-predictors through an ensemble learning approach improved the balanced prediction performance and model robustness on the independent dataset. Upon comparison with existing methods, mAHTPred showed superior performance with an overall improvement of approximately 6-7% in both benchmarking and independent datasets. AVAILABILITY AND IMPLEMENTATION: The user-friendly online prediction tool, mAHTPred is freely accessible at http://thegleelab.org/mAHTPred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Balachandran Manavalan, Shaherin Basith, Taehwan Shin, Leyi Wei, Gwang Lee |
Bioinform. | 4 |
| 2019 | Exploring sequence-based features for the improved prediction of DNA N4-methylcytosine sites in multiple speciesabstractMOTIVATION: As one of important epigenetic modifications, DNA N4-methylcytosine (4mC) is recently shown to play crucial roles in restriction-modification systems. For better understanding of their functional mechanisms, it is fundamentally important to identify 4mC modification. Machine learning methods have recently emerged as an effective and efficient approach for the high-throughput identification of 4mC sites, although high predictive error rates are still challenging for existing methods. Therefore, it is highly desirable to develop a computational method to more accurately identify m4C sites. RESULTS: In this study, we propose a machine learning based predictor, namely 4mcPred-SVM, for the genome-wide detection of DNA 4mC sites. In this predictor, we present a new feature representation algorithm that sufficiently exploits sequence-based information. To improve the feature representation ability, we use a two-step feature optimization strategy, thereby obtaining the most representative features. Using the resulting features and Support Vector Machine (SVM), we adaptively train the optimal models for different species. Comparative results on benchmark datasets from six species indicate that our predictor is able to achieve generally better performance in predicting 4mC sites as compared to the state-of-the-art predictors. Importantly, the sequence-based features can reliably and robust predict 4mC sites, facilitating the discovery of potentially important sequence characteristics for the prediction of 4mC sites. AVAILABILITY AND IMPLEMENTATION: The user-friendly webserver that implements the proposed 4mcPred-SVM is well established, and is freely accessible at http://server.malab.cn/4mcPred-SVM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leyi Wei, Shasha Luan, Luis Augusto Eijy Nagai, Ran Su, Quan Zou 0001 |
Bioinform. | 1 |
| 2019 | Iterative feature representations improve N4-methylcytosine site predictionabstractMOTIVATION: Accurate identification of N4-methylcytosine (4mC) modifications in a genome wide can provide insights into their biological functions and mechanisms. Machine learning recently have become effective approaches for computational identification of 4mC sites in genome. Unfortunately, existing methods cannot achieve satisfactory performance, owing to the lack of effective DNA feature representations that are capable to capture the characteristics of 4mC modifications. RESULTS: In this work, we developed a new predictor named 4mcPred-IFL, aiming to identify 4mC sites. To represent and capture discriminative features, we proposed an iterative feature representation algorithm that enables to learn informative features from several sequential models in a supervised iterative mode. Our analysis results showed that the feature representations learnt by our algorithm can capture the discriminative distribution characteristics between 4mC sites and non-4mC sites, enlarging the decision margin between the positives and negatives in feature space. Additionally, by evaluating and comparing our predictor with the state-of-the-art predictors on benchmark datasets, we demonstrate that our predictor can identify 4mC sites more accurately. AVAILABILITY AND IMPLEMENTATION: The user-friendly webserver that implements the proposed 4mcPred-IFL is well established, and is freely accessible at http://server.malab.cn/4mcPred-IFL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leyi Wei, Ran Su, Shasha Luan, Zhijun Liao, Balachandran Manavalan, Quan Zou 0001 |
Bioinform. | 1 |
| 2019 | PEPred-Suite: improved and robust prediction of therapeutic peptides using adaptive feature representation learningabstractMOTIVATION: Prediction of therapeutic peptides is critical for the discovery of novel and efficient peptide-based therapeutics. Computational methods, especially machine learning based methods, have been developed for addressing this need. However, most of existing methods are peptide-specific; currently, there is no generic predictor for multiple peptide types. Moreover, it is still challenging to extract informative feature representations from the perspective of primary sequences. RESULTS: In this study, we have developed PEPred-Suite, a bioinformatics tool for the generic prediction of therapeutic peptides. In PEPred-Suite, we introduce an adaptive feature representation strategy that can learn the most representative features for different peptide types. To be specific, we train diverse sequence-based feature descriptors, integrate the learnt class information into our features, and utilize a two-step feature optimization strategy based on the area under receiver operating characteristic curve to extract the most discriminative features. Using the learnt representative features, we trained eight random forest models for eight different types of functional peptides, respectively. Benchmarking results showed that as compared with existing predictors, PEPred-Suite achieves better and robust performance for different peptides. As far as we know, PEPred-Suite is currently the first tool that is capable of predicting so many peptide types simultaneously. In addition, our work demonstrates that the learnt features can reliably predict different peptides. AVAILABILITY AND IMPLEMENTATION: The user-friendly webserver implementing the proposed PEPred-Suite is freely accessible at http://server.malab.cn/PEPred-Suite. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leyi Wei, Ran Su, Quan Zou 0001 |
Bioinform. | 1 |
| 2019 | Integration of deep feature representations and handcrafted features to improve the prediction of N6-methyladenosine sites
Leyi Wei, Ran Su, Xiu-Ting Li, Quan Zou 0001, Xing Gao 0004 |
Neurocomputing | 1 |
| 2019 | DUNet: A deformable network for retinal vessel segmentation
Qiangguo Jin, Zhaopeng Meng, Tuan D. Pham, Leyi Wei, Ran Su |
Knowl. Based Syst. | 5 |
| 2019 | Developing a Multi-Dose Computational Model for Drug-Induced Hepatotoxicity Prediction Based on Toxicogenomics DataabstractDrug-induced hepatotoxicity may cause acute and chronic liver disease, leading to great concern for patient safety. It is also one of the main reasons for drug withdrawal from the market. Toxicogenomics data has been widely used in hepatotoxicity prediction. In our study, we proposed a multi-dose computational model to predict the drug-induced hepatotoxicity based on gene expression and toxicity data. The dose/concentration information after drug treatment is fully utilized in our study based on the dose-response curve, thus a more informative representative of the dose-response relationship is considered. We also proposed a new feature selection method, named MEMO, which is also one important aspect of our multi-dose model in our study, to deal with the high-dimensional toxicogenomics data. We validated the proposed model using the TG-GATEs, which is a large database recording toxicogenomics data from multiple views. The experimental results show that the drug-induced hepatotoxicity can be predicted with high accuracy and efficiency using the proposed predictive model. Ran Su, Huichen Wu, Xiaofeng Liu 0004, Leyi Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2019 | Fast Prediction of Protein Methylation Sites Using a Sequence-Based Feature Selection TechniqueabstractProtein methylation, an important post-translational modification, plays crucial roles in many cellular processes. The accurate prediction of protein methylation sites is fundamentally important for revealing the molecular mechanisms undergoing methylation. In recent years, computational prediction based on machine learning algorithms has emerged as a powerful and robust approach for identifying methylation sites, and much progress has been made in predictive performance improvement. However, the predictive performance of existing methods is not satisfactory in terms of overall accuracy. Motivated by this, we propose a novel random-forest-based predictor called MePred-RF, integrating several discriminative sequence-based feature descriptors and improving feature representation capability using a powerful feature selection technique. Importantly, unlike other methods based on multiple, complex information inputs, our proposed MePred-RF is based on sequence information alone. Comparative studies on benchmark datasets via vigorous jackknife tests indicate that our proposed MePred-RF method remarkably outperforms other state-of-the-art predictors, leading by a 4.5 percent average in terms of overall accuracy. A user-friendly webserver that implements the proposed method has been established for researchers' convenience, and is now freely available for public use through http://server.malab.cn/MePred-RF. We anticipate our research tool to be useful for the large-scale prediction and analysis of protein methylation sites. Leyi Wei, Pengwei Xing, Gaotao Shi, Zhi-Liang Ji, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | Encoded Texture Features to Characterize Bone Radiograph ImagesabstractOsteoporosis is the most common reason that causes the fracture among the elderly. For the purpose of convenience and safety, 2D texture analysis has been used to diagnose osteoporosis. In this study, a supervised method using proposed texture features to identify osteoporotic cases from healthy was proposed. We designed two groups of new features, Encoded GLCM and Encoded LBP, each of which contains two subgroups through encoding the Gabor and Hessian information into the Gray Level Co-Occurrence Matrix (GLCM) features and Local Binary Patterns (LBP) features respectively. These two groups of features, together with the raw feature group containing the GLCM and LBP features, totally 560 features, were categorized into various groups and used to train the Random Forest classifier. Classification performances using these features were compared inter-and intra-groups/subgroups. And the performance using each individual feature was also provided. We conducted feature selection based on Recursive Feature Elimination (RFE) inside a voting scheme to further increase the efficiency. The inter-and intra-groups/subgroups results indicate that the Encoded GLCM and Encoded LBP, are more discriminative than the raw GLCM and LBP features for the identification of the osteoporosis; The best individual feature is from the Encoded LBP group and can achieve 70% of balanced accuracy; Furthermore, using only ten of the proposed features through feature selection, the balanced accuracy can even be improved from 60% to 71%. This shows that the proposed method is promising to assist the early diagnosis of osteoporosis. Ran Su, Leyi Wei, Xiu-Ting Li, Qiangguo Jin, Wenyuan Tao |
ICPR | 3 |
| 2018 | ACPred-FL: a sequence-based predictor using effective feature representation to improve the prediction of anti-cancer peptidesabstractMotivation: Anti-cancer peptides (ACPs) have recently emerged as promising therapeutic agents for cancer treatment. Due to the avalanche of protein sequence data in the post-genomic era, there is an urgent need to develop automated computational methods to enable fast and accurate identification of novel ACPs within the vast number of candidate proteins and peptides. Results: To address this, we propose a novel predictor named Anti-Cancer peptide Predictor with Feature representation Learning (ACPred-FL) for accurate prediction of ACPs based on sequence information. More specifically, we develop an effective feature representation learning model, with which we can extract and learn a set of informative features from a pool of support vector machine-based models trained using sequence-based feature descriptors. By doing so, the class label information of data samples is fully utilized. To improve the feature representation, we further employ a two-step feature selection technique, resulting in a most informative five-dimensional feature vector for the final peptide representation. Experimental results show that such five features provide the most discriminative power for identifying ACPs than currently available feature descriptors, highlighting the effectiveness of the proposed feature representation learning approach. The developed ACPred-FL method significantly outperforms state-of-the-art methods. Availability and implementation: The web-server of ACPred-FL is available at http://server.malab.cn/ACPred-FL. Supplementary information: Supplementary data are available at Bioinformatics online. Leyi Wei, Huangrong Chen, Jiangning Song, Ran Su |
Bioinform. | 1 |
| 2018 | Prediction of human protein subcellular localization using deep learning
Leyi Wei, Yijie Ding, Ran Su, Jijun Tang, Quan Zou 0001 |
J. Parallel Distributed Comput. | 1 |
| 2017 | A novel hierarchical selective ensemble classifier with bioinformatics application
Leyi Wei, Shixiang Wan, Jiasheng Guo, Kelvin K. L. Wong |
Artif. Intell. Medicine | 1 |
| 2017 | Improved prediction of protein-protein interactions using novel negative samples, features, and an ensemble classifier
Leyi Wei, Pengwei Xing, Jian-Cang Zeng, Jin-Xiu Chen, Ran Su, Fei Guo 0001 |
Artif. Intell. Medicine | 1 |
| 2017 | Local-DPP: An improved DNA-binding protein prediction method by exploring local evolutionary information
Leyi Wei, Jijun Tang, Quan Zou 0001 |
Inf. Sci. | 1 |
| 2016 | mGOF-loc: A novel ensemble learning method for human protein subcellular localization prediction
Leyi Wei, Minghong Liao, Xing Gao 0004 |
Neurocomputing | 1 |
| 2016 | Exploring local discriminative information from evolutionary profiles for cytokine-receptor interaction prediction
Leyi Wei, Xing Gao 0004, Minghong Liao |
Neurocomputing | 1 |
| 2014 | Improved and Promising Identificationof Human MicroRNAs by Incorporatinga High-Quality Negative SetabstractMicroRNA (miRNA) plays an important role as a regulator in biological processes. Identification of (pre-) miRNAs helps in understanding regulatory processes. Machine learning methods have been designed for pre-miRNA identification. However, most of them cannot provide reliable predictive performances on independent testing data sets. We assumed this is because the training sets, especially the negative training sets, are not sufficiently representative. To generate a representative negative set, we proposed a novel negative sample selection technique, and successfully collected negative samples with improved quality. Two recent classifiers rebuilt with the proposed negative set achieved an improvement of ~6 percent in their predictive performance, which confirmed this assumption. Based on the proposed negative set, we constructed a training set, and developed an online system called miRNApre specifically for human pre-miRNA identification. We showed that miRNApre achieved accuracies on updated human and non-human data sets that were 34.3 and 7.6 percent higher than those achieved by current methods. The results suggest that miRNApre is an effective tool for pre-miRNA identification. Additionally, by integrating miRNApre, we developed a miRNA mining tool, mirnaDetect, which can be applied to find potential miRNAs in genome-scale data. MirnaDetect achieved a comparable mining performance on human chromosome 19 data as other existing methods. Leyi Wei, Minghong Liao, Yue Gao 0002, Rongrong Ji, Zengyou He, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |