Lei Deng 0002

dblp:96/755-2 · DBLP profile ↗
← Back
104ranked-venue papers
18as first author
77since 2021 · last 2026
0000-0003-2869-1619ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 102 · 18 first-author · 75 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hierarchical Structure-Property Alignment for Data-Efficient Molecular Generation and Editing
abstract
Property-constrained molecular generation and editing are crucial in AI-driven drug discovery but remain hindered by two factors: (i) capturing the complex relationships between molecular structures and multiple properties remains challenging, and (ii) the narrow coverage and incomplete annotations of molecular properties weaken the effectiveness of property-based models. To tackle these limitations, we propose HSPAG, a data-efficient framework featuring hierarchical structure–property alignment. By treating SMILES and molecular properties as complementary modalities, the model learns their relationships at atom, substructure, and whole-molecule levels. Moreover, we select representative samples through scaffold clustering and hard samples via an auxiliary variational auto-encoder (VAE), substantially reducing the required pre-training data. In addition, we incorporate a property relevance-aware masking mechanism and diversified perturbation strategies to enhance generation quality under sparse annotations. Experiments demonstrate that HSPAG captures fine-grained structure–property relationships and supports controllable generation under multiple property constraints. Two real-world case studies further validate the editing capabilities of HSPAG.
Ziyu Fan, Zhijian Huang 0001, Yahan Li, Yunliang Wang, Zeyu Zhong, Shuhong Liu, Shuning Yang, Shangqian Wu, Min Wu 0008, Lei Deng 0002
AAAI12
2026 Structure-agnostic protein-ligand binding affinity prediction via hierarchical representation alignment
Hongyi Huang, Zhenggang Wang, Lei Deng 0002
Bioinform.8
2026 DeepLMI: deep feature mining with a globally enhanced graph convolutional network for robust lncRNA-miRNA interaction prediction
abstract
MOTIVATION: Interactions between long noncoding RNAs (lncRNAs) and microRNAs (miRNAs) play pivotal roles in gene regulation and disease progression, notably through mechanisms such as competitive miRNA sponging. Accurate identification of lncRNA-miRNA interactions is therefore essential for understanding disease mechanisms and discovering therapeutic targets. However, current knowledge is largely derived from labor-intensive and costly biological experiments, underscoring the need for reliable computational approaches. RESULTS: We propose DeepLMI, a novel deep learning framework for lncRNA-miRNA interaction prediction that integrates deep feature mining with a globally enhanced graph convolutional network. To effectively capture the distinct properties of lncRNAs and miRNAs, DeepLMI employs specialized feature extraction modules: for lncRNAs, we combine sequence pretraining with self-attention mechanisms to learn multiscale semantic representations; for miRNAs, we fuse heterogeneous features through a graph convolutional encoder. To further address the sparsity and structural complexity of known RNA interaction networks, we design a Global-Enhanced Graph Convolutional Network that jointly models local neighborhood information and global topological signals. The embeddings learned for lncRNAs and miRNAs are then integrated to infer interaction probabilities. Extensive experiments across multiple datasets and evaluation settings demonstrate that DeepLMI consistently outperforms existing state-of-the-art methods and exhibits strong robustness, highlighting its potential as a valuable tool for RNA interaction analysis and disease research. AVAILABILITY AND IMPLEMENTATION: The codes and data are publicly available at https://github.com/Hhhzj-7/DeepLMI.
Zhijian Huang 0001, Xianshu Wang, Junheng Wang, Yuanpeng Zhang 0004, Min Wu 0008, Lei Deng 0002
Bioinform.8
2026 MM-GHiNet: A multimodal graph hierarchical network for brain disease diagnosis
Zhijian Huang 0001, Lei Deng 0002
Expert Syst. Appl.4
2026 HHGSynergy: An Adaptive Heterogeneous Hypergraph Representation Learning Method for Anticancer Drug Synergy Prediction
abstract
Compared with monotherapy, combination drug therapy plays a crucial role in clinical treatment. However, the exponential expansion of the drug combination space has rendered traditional exploration methods for synergistic drug combinations inadequate. Recently, numerous efficient and accurate computational approaches have been developed to predict anticancer drug synergy, particularly those leveraging hypergraphs to model the multifaceted relationships between drug combinations and cell lines, which have demonstrated remarkable potential. Nevertheless, existing hypergraph-based methods fail to account for the heterogeneity of anticancer synergy hypergraphs and overlook the underlying similarities among drugs and cell lines, thereby limiting their ability to fully capture the complex interactions between drug combinations and cell lines. To address these limitations, we propose an Adaptive Heterogeneous Hypergraph Representation Learning Method (HHGSynergy) for predicting anticancer drug synergy, enabling more precise identification of synergistic drug combinations. Specifically, our framework first constructs drug/cell line similarity-based synergy hypergraphs based on the foundational anticancer synergy hypergraph, thereby establishing a comprehensive heterogeneous hypergraph. Next, a node importance calculation module is employed to learn both local and global importance weights of nodes, effectively capturing the structural characteristics of the hypergraph. Finally, a type-specific multi-head attention mechanism is utilized to iteratively update node embeddings, adaptively learning the significance of heterogeneous hyperedges. Experimental results demonstrate that HHGSynergy achieves state-of-the-art performance in both classification and regression tasks across diverse experimental scenarios, outperforming existing leading models. Case studies further underscore its potential for discovering novel synergistic drug combinations.
Jinmiao Song, Lei Deng 0002, Qimeng Yang, Qiguo Dai, Shengwei Tian
IEEE Trans. Comput. Biol. Bioinform.3
2026 HECLCDA:CircRNA-Drug Sensitivity Prediction via Heterogeneous Cross-Scale Contrastive Learning
abstract
Circular RNA (circRNA) is a widely distributed class of non-coding RNA molecules that have been shown to play a significant role in cancer development and drug resistance, significantly influencing cellular sensitivity to therapeutic drugs and treatment outcomes. However, traditional biomedical experimental methods are limited by low efficiency and high costs when verifying the association between circular RNA and drug sensitivity. Therefore, developing an efficient and accurate computational method to predict new associations between circRNA and drug sensitivity has become an urgent need in current research. To address this, this study proposes HECLCDA, a novel method based on heterogeneous cross-scale contrastive learning. To construct a comprehensive initial information base for drugs and circRNAs, circRNA gene sequence similarity, drug structural inclusion similarity (SIS), and Gaussian kernel similarity were integrated.Based on the integrated and complete known information of circRNAs and drugs, a heterogeneous graph was built. The model used the Heterogeneous Graph Transformer to extract heterogeneous network topological information, effectively distinguishing the heterogeneity of nodes and edges. The model broke through the information relationship between node attributes and network topology at two scales, and innovatively introduced a cross-scale contrastive learning mechanism in a sparse labeling scenario. Using self-supervised signals, we aimed to enhance the discriminative power of node embeddings and maximize the mutual information between paired nodes at different scales. Cross-validation experiments demonstrated that HECLCDA performs excellently on real data and can efficiently predict drug sensitivity. Additionally, case studies further validate the model's effectiveness in predicting potential circRNA-drug sensitivity associations.
Jinmiao Song, Lei Deng 0002, Qimeng Yang, Qiguo Dai, Shengwei Tian
IEEE Trans. Comput. Biol. Bioinform.3
2025 DSSA: Dual-Stream Synthetic Accessibility Framework for Organic Compounds
abstract
Synthetic accessibility prediction remains a key bottleneck in AI-driven drug discovery, as a large proportion of computationally generated molecules prove infeasible to synthesize. Existing approaches often struggle to distinguish structurally similar compounds with divergent synthetic profiles, limiting their usefulness in practical design pipelines. We present DSSA (Dual-Stream Synthetic Accessibility), a novel architecture that integrates Graph Attention Networks for molecular topology with a Bidirectional GRU for sequential SMILES representations through transformer-based cross-modal fusion. DSSA effectively captures both local structural complexity and global sequential patterns, enabling robust generalization across diverse molecular classes. Cross-modal attention analysis reveals that the model dynamically adapts to molecular complexity, with graph-dominant attention highlighting stereochemical constraints that sequenceonly models overlook. Ablation studies further confirm that cross-modal fusion is essential for achieving balanced structural and sequential reasoning. Collectively, DSSA bridges the gap between computational molecular generation and real-world synthetic feasibility, offering a reliable foundation for data-driven molecular design. Web tool: http://dssa.denglab.org/; code/data: https://github.com/Q-Aljanabi/DSSA.
Qahtan Adnan Aljanabi, Zhijian Huang 0001, Gebremedhin Assefa Girmay, Zhengkang Wang, Yuanpeng Zhang 0004, Lei Deng 0002
BIBM6
2025 DeepBindAffinity: A ResNet-Augmented Hybrid CNN-Transformer for Protein-Ligand Binding Affinity Prediction
abstract
Accurate prediction of protein-ligand binding affinity remains a central challenge in drug discovery, particularly when 3D structural data are unavailable. We introduce DeepBindAffinity, a deep learning framework that predicts binding affinity directly from 1D sequence inputs-protein amino acid sequences, binding pocket residues, and ligand SMILES strings. The model combines a ResNet-augmented CNN with a Transformer encoder to jointly capture local biochemical motifs and long-range dependencies, while a cross-attention module integrates protein and ligand representations in a unified latent space. Trained on a large-scale curated dataset comprising over 285,000 complexes and rigorously benchmarked under 10-fold cross-validation and independent test sets (CASF-2013, CASF-2016, and Core2016), DeepBindAffinity achieves strong generalization and competitive accuracy against structure-based approaches. By relying solely on sequence-level information, it provides a scalable and accessible solution for virtual screening and affinity estimation in early-stage drug discovery where structural data are limited. Code and models are available at: https://github.com/Gere2119/DeepBindAffinity.
Gebremedhin Assefa Girmay, Qahtan Adnan Aljanabi, Lei Deng 0002
BIBM3
2025 MutDTA: Interpretable Transfer Learning for Predicting Mutation Effects on Drug Binding Affinity in Viral Proteins
abstract
Amino acid substitutions in or near the drug-binding pockets of biological targets are critical drivers of drug resistance in viral infections. Despite advances in deep learning-based methods for predicting drug-target binding affinity (DTA) and the effects of amino acid substitutions, their practical application remains limited. This limitation stems from these methods' failure to consider the susceptibility of proteins to mutations under selective pressure and the challenges posed by data scarcity. In this study, we introduce MutDTA, a transfer learning model designed to predict the impact of mutations on DTA and to elucidate resistance mechanisms effectively. MutDTA utilizes protein sequences and drug molecular graphs to embed targets and drugs, respectively, enhancing the model's utility. A cross-attention module not only improves prediction accuracy but also aids in identifying critical interaction sites, thereby enhancing model interpretability. Fine-tuning on the Platinum dataset enables MutDTA to effectively capture mutation-specific insights, overcoming issues of data sparsity. Experimental results from both pre-training and fine-tuning phases show that MutDTA significantly outperforms existing methods, displaying exceptional generalization capabilities in cold-start scenarios. Furthermore, an interpretability analysis involving the HIV-1 protease and the drug TMC114, with and without the V32I mutation, demonstrates MutDTA's precision in pinpointing key sites of resistance. Source code and datasets can be available at https://github.com/altriavin/MutDTA.
Min Wu 0008, Guangdi Li, Zizhang Sheng, Lei Deng 0002
BIBM6
2025 ASRSMA: Atomic-Scale and Structure-Based Modeling for RNA-Small Molecule Binding Affinity Prediction via Contrastive Pretraining
abstract
RNA is intricately involved in aberrant cellular functions and a wide range of disease processes, playing pivotal roles in gene regulation, viral replication, and innate immunity. Consequently, it has emerged as a highly promising therapeutic target. To accelerate the discovery of such drugs, it is essential to develop an effective computational method for predicting RNA-small molecule affinity. Therefore, we propose ASRSMA as an atomiclevel structure-aware model with contrastive pre-training for RNA-small molecule binding affinity prediction. ASRSMA represents RNA and small molecules at atomic resolution, enabling the capture of fine-grained structural features. It incorporates an atom-pair encoding module and an inter-molecular interaction module to model intra- and inter-molecular interactions. In addition, a self-supervised contrastive learning strategy is employed during pre-training, maximising the similarity between different views of the same complex while minimising similarities across different complexes, thereby yielding deep representations with strong generalisation capacity. Experimental results demonstrate that ASRSMA significantly outperforms state-of-the-art baseline models in predicting RNA-small molecule binding affinities.
Yahan Li, Zhijian Huang 0001, Yucheng Wang 0001, Min Wu 0008, Lei Deng 0002
BIBM6
2025 A Modular Framework for Multimorbidity Prediction with MNAR-Aware Imputation and Label-Enhanced Transformers
abstract
We propose a two-stage deep learning framework that separately targets imputation and multimorbidity prediction. For the imputation stage, we develop MPA-SAEI, a Siamese autoencoder-based model enhanced with a missing pattern network and adversarial training. This architecture explicitly estimates conditional missing probabilities and aligns the imputed data distribution with the observed one. The triplet structure helps preserve feature semantics and improves robustness under complex MNAR settings. For the prediction stage, we present ML-LEFTT, a label-enhanced FT-Transformer model for multi-label disease prediction. By incorporating both label embeddings and attention-based interactions between features and labels, it captures inter-label dependencies and label-specific feature relevance more effectively.
Zhenzhen Rao, Lei Deng 0002
BIBM4
2025 TabNet-Based Autoencoder Enhanced with Dynamic Masking for Multiple Imputation of Medical Table Data
abstract
In medical data analysis research, missing data is prevalent and can undermine the reliability of disease diagnosis and clinical decision-making, posing a potential threat to patient health and safety. Therefore, it is imperative to address missing data through reasonable and effective imputation methods. Currently, commonly used deep generative models often struggle to adequately capture the complex interdependencies between features when dealing with imbalanced category distributions and feature sparsity in medical tabular data. Moreover, missing data in real-world scenarios are often a mixture of multiple missing mechanisms, leading to complexity that makes fixed masks unable to dynamically adapt to different missing patterns and data distribution characteristics. Additionally, single impu-tation cannot reflect the inherent uncertainty of missing values, often underestimating variance and misleading subsequent anal-yses. Therefore, this paper proposes a multi-imputation model based on a Tab Net autoencoder(TAEMI) with dynamic mask enhancement for imputing tabular data with complex missing mechanisms. We incorporate shared feature transformations and sparse attention mechanisms into the encoder structure, achieving structured modeling and efficient representation of input data through multi-level feature selection and gradual feature extraction, thereby enhancing the model's ability to capture key features. During the training phase, we employ a dynamic masking enhancement strategy to learn data information under different missing conditions, improving the model's adaptability to complex missing mechanisms. Finally, we enhance the ro-bustness of the results through multiple imputation. imputation experiments on the TCGA dataset and lung cancer nutrition dataset demonstrate that TAEMI exhibits good robustness across various missing data scenarios and shows advantages in high- missing-rate and high-dimensional data scenarios. Additionally, the imputed datasets perform well in downstream prediction tasks.
Zhenzhen Rao, Lei Deng 0002
BIBM4
2025 GroupTransUNet: Group Transformer UNet for Medical Image Segmentation
abstract
Accurate medical image segmentation is a pivotal task in medical image analysis, serving as the foundation for extracting critical clinical information and directly influencing the precision of disease identification, diagnosis, and treatment decision-making. Although Convolutional Neural Networks (CNNs) and Transformers have demonstrated remarkable performance in medical image segmentation, their sensitivity to variations in target size and morphology remains insufficient, resulting in limited capability of single-scale features to comprehensively capture multi-target details. Current approaches also face bottlenecks, including restricted global context modeling capabilities, high computational complexity, and inefficient multilevel feature fusion. To address these challenges in medical image segmentation, this study proposes an innovative architecture named GroupTransUNet. The model implements dual enhancements based on the UNet framework: (i) A novel Next-Generation Transformer Block (NGTB) is designed for the bottleneck layer, which organically integrates Efficient Multi-Head Self-Attention (E-MHSA) with Convolution-Enhanced Multi-Layer Perceptron (MLP) through a complementary mechanism to simultaneously enhance global semantic understanding and local detail characterization. (ii) The Grouped Feature Fusion Module (GFFM) is introduced in skip connections, employing grouped convolutions and multi-dilation rate strategies to construct multi-scale contextual receptive fields while preserving detailed integrity, thereby significantly improving feature fusion efficiency. Experimental results on the Synapse and ACDC datasets validate the effectiveness of our proposed GroupTransUNet. The source code can be obtained at https://anonymous.4open.science/r/GroupTransUNetA3D6.
Yunliang Wang, Ziyu Fan, Zhijian Huang 0001, Jinmiao Song, Lei Deng 0002
BIBM5
2025 MVFDSP: A Multi-View Fusion Framework for Drug Side-Effect Frequency Prediction
abstract
Accurate prediction of drug side effect frequencies is critical for drug safety evaluation and clinical decision-making. Current methods primarily emphasize the associations between drugs and side effects, yet they often neglect the underlying structural and semantic features of both, which limits further advancements in prediction accuracy. In this study, we propose a novel multi-view fusion framework, MVFDSP, which integrates pre-trained molecular representation of 1D and 2D views with graph-based side effect information for side effect frequency prediction. Firstly, we obtain both 1D and 2D molecular representations from the pretrained molecular language model, and combine them using an adaptive fusion strategy. Subsequently, we construct a similarity network based on the side effect frequency matrix using K-Nearest Neighbors (KNN), and incorporate semantic embeddings derived from the terminology system of MedDRA to construct a side effect information graph. A multi-head graph attention network is then employed to capture the multi-dimensional information within this graph, allowing the model to attend to diverse aspects of the semantic and structural relationships among side effects. The final frequency prediction matrix is derived from the inner product between the learned drug and side effect embeddings. Experimental results on the SIDER 4.1 dataset demonstrate that MVFDSP outperforms existing methods, highlighting its effectiveness in capturing complex relationships of drugs and side effects. The code and data are available at https://github.com/Sonder-Echo/MVFDSP.
Zhengkang Wang, Zhijian Huang 0001, Yurong Qian, Yuanpeng Zhang 0004, Yahan Li, Qahtan Adnan Aljanabi, Jinmiao Song, Lei Deng 0002
BIBM8
2025 MVRBind: multi-view learning for RNA-small molecule binding site prediction
abstract
RNA plays a critical role in cellular processes, and its dysregulation is linked to many diseases, positioning RNA-targeted drugs as an important area of research. Accurate prediction of RNA-small molecule binding sites is crucial for advancing RNA-targeted therapies. Although deep learning has shown promise in this area, challenges remain in integrating and processing multi-dimensional data, such as RNA sequences and structural features, particularly given the inherent flexibility of RNA structures. In this study, we present MVRBind, a multi-view graph convolutional network designed to predict RNA-small molecule binding sites. MVRBind generates feature representations of RNA nucleotides across different structural levels. To effectively integrate these features, we developed a multi-view feature fusion module that constructs graphs based on RNA's primary, secondary, and tertiary structural views, enabling the model to capture diverse aspects of RNA structure. In addition, we fuse embeddings from multi-scale to obtain a comprehensive representation of RNA nucleotides, which is then used to predict RNA-small molecule binding sites. Extensive experiments demonstrate that MVRBind consistently outperforms baseline methods in various experimental settings. Our MVRBind shows exceptional performance in predicting binding sites for both the holo and apo forms of RNA, even when RNA adopts multiple conformations. These results suggest that MVRBind offers a robust model for structure-based RNA analysis, contributing toward accurate prediction and analysis of RNA-small molecule binding sites. All datasets and resource codes are available at https://github.com/cschen-y/MVRBind.
Zhijian Huang 0001, Yucheng Wang 0001, Yahan Li, Yaw Sing Tan, Lei Deng 0002, Min Wu 0008
Briefings Bioinform.6
2025 DeepHeteroCDA: circRNA-drug sensitivity associations prediction via multi-scale heterogeneous network and graph attention mechanism
abstract
Drug sensitivity is essential for identifying effective treatments. Meanwhile, circular RNA (circRNA) has potential in disease research and therapy. Uncovering the associations between circRNAs and cellular drug sensitivity is crucial for understanding drug response and resistance mechanisms. In this study, we proposed DeepHeteroCDA, a novel circRNA-drug sensitivity association prediction method based on multi-scale heterogeneous network and graph attention mechanism. We first constructed a heterogeneous graph based on drug-drug similarity, circRNA-circRNA similarity, and known circRNA-drug sensitivity associations. Then, we embedded the 2D structure of drugs into the circRNA-drug sensitivity heterogeneous graph and use graph convolutional networks (GCN) to extract fine-grained embeddings of drug. Finally, by simultaneously updating graph attention network for processing heterogeneous networks and GCN for processing drug structures, we constructed a multi-scale heterogeneous network and use a fully connected layer to predict the circRNA-drug sensitivity associations. Extensive experimental results highlight the superior of DeepHeteroCDA. The visualization experiment shows that DeepHeteroCDA can effectively extract the association information. The case studies demonstrated the effectiveness of our model in identifying potential circRNA-drug sensitivity associations. The source code and dataset are available at https://github.com/Hhhzj-7/DeepHeteroCDA.
Zhijian Huang 0001, Xiaojun Xiao, Ziyu Fan, Yuanpeng Zhang 0004, Lei Deng 0002
Briefings Bioinform.6
2025 Contrastive hypergraph collaborative filtering for transfer RNA-disease association prediction
abstract
Transfer RNAs (tRNAs) play critical roles in the process of protein synthesis by decoding messenger RNA codons into amino acids, which is essential for cellular function across various biological pathways and for maintaining metabolic homeostasis. Available evidence implicates that tRNAs are involved in the progression of diverse diseases, underscoring the importance of accurately predicting tRNA-disease associations to understand disease mechanisms and support precision medicine. However, existing methods often struggle with the complexity and heterogeneity inherent in these associations. To address these challenges, we introduce contrastive hypergraph collaborative filtering (CoHGCL), a prediction framework that integrates hypergraph contrastive learning with collaborative filtering. CoHGCL employs graph attention networks to capture local structural features and random walk with restart algorithms to encode global topological patterns. Subsequently, a node-level contrastive learning mechanism alternates between standard graph and hypergraph representations to enhance multiview feature embeddings. These enriched representations are integrated by a collaborative filtering approach through the utilization of generalized matrix factorization for modeling linear associations and multilayer perceptrons for capturing nonlinear interactions. Extensive experimental results on five-fold cross-validation demonstrate that CoHGCL achieves superior performance compared to existing methods, with an area under the receiver operating characteristic curve of 0.9623, area under the precision-recall curve of 0.9430, outperforming all baselines across all metrics. Furthermore, case studies further confirm CoHGCL's effectiveness in discovering novel and biologically meaningful tRNA-disease associations. The source code and datasets are publicly available at https://github.com/Ouyang-cmd/CoHGCL.
Tianxiang Ouyang, Yuanpeng Zhang 0004, Zhijian Huang 0001, Lei Deng 0002
Briefings Bioinform.4
2025 SeekTP: identifying multifunctional therapeutic peptides via graph attention and deep property integration
abstract
With rapid advancements in biomedical research, the volume of available peptide data has significantly expanded, heightening the need for computational tools capable of efficiently identifying multifunctional therapeutic peptides (MFTPs). These peptides hold great promise as novel therapies across various disease contexts. However, existing computational frameworks face considerable challenges, such as significant class imbalance and difficulties in effectively capturing latent peptide characteristics, leaving substantial room for improvement in predictive capabilities. In this study, we introduce SeekTP, a novel computational framework designed to accurately identify MFTPs from amino acid sequences by integrating diverse predicted structural and sequential properties from multiple categories. The model constructs graph representations from predicted structural, sequential, and embedding information and employs a graph attention network to effectively encode complex interactions. Concurrently, self-attention mechanisms and convolutional neural networks are applied to capture intricate patterns in peptide properties. A feed-forward neural network serves as the final prediction layer. Extensive benchmarking experiments illustrate that SeekTP is able to achieve more delightful manifestations compared with state-of-the-art methods on independent test datasets, explaining its advantages in classification tasks. Additionally, we systematically analyze the discriminative power achieved by various combinations of feature representations and modeling strategies, underscoring the robustness and flexibility of our approach.
Guolun Zhong, Lei Deng 0002
Briefings Bioinform.2
2025 MolCL-SP: a multimodal contrastive learning framework with non-overlapping substructure perturbations for molecular property prediction
abstract
MOTIVATION: Accurate molecular property prediction remains a central challenge in molecular machine learning, critically dependent on comprehensive molecular representation. Existing methods, however, encounter two major limitations: (i) single-modal learning approaches frequently experience representation bottlenecks, whereas multimodal methods often struggle to effectively leverage complementary information without redundancy across modalities; and (ii) conventional data augmentation techniques typically treat atoms as isolated units, neglecting intrinsic dependencies among atoms within molecular substructures. RESULTS: Here, we propose MolCL-SP, a substructure-aware multimodal contrastive learning framework specifically designed for molecular property prediction. Our approach integrates molecular representations derived from three complementary modalities using a Transformer-based encoder, followed by modality-specific reconstruction to organically align and fuse cross-modal information. We also introduce a novel substructure-based non-overlapping perturbation strategy for data augmentation, preserving interpretability and effectively enhancing inter-modal interactions. Extensive experimental evaluations demonstrate that MolCL-SP achieves state-of-the-art performance on benchmark datasets for both 2D and 3D molecular property predictions. Additionally, evaluations on drug-drug interaction prediction tasks highlight the model's strong generalization capabilities. Visualization analyses further indicate that MolCL-SP effectively captures discriminative molecular embeddings even in task-agnostic contexts. Importantly, the model implicitly emphasizes chemically meaningful substructures associated with functional relevance, significantly enhancing interpretability. AVAILABILITY AND IMPLEMENTATION: Codes and materials are available at https://github.com/lylikeeMoon/MolCL-SP.
Lei Deng 0002
Bioinform.2
2025 Precise prediction of hotspot residues in protein-RNA complexes using graph attention networks and pretrained protein language models
abstract
MOTIVATION: Protein-RNA interactions play a pivotal role in biological processes and disease mechanisms, with hotspot residues being critical for targeted drug design. Traditional experimental methods for identifying hotspot residues are often inefficient and expensive. Moreover, many existing prediction methods rely heavily on high-resolution structural data, which may not always be available. Consequently, there is an urgent need for an accurate and efficient sequence-based computational approach for predicting hotspot residues in protein-RNA complexes. RESULTS: In this study, we introduce DeepHotResi, a sequence-based computational method designed to predict hotspot residues in protein-RNA complexes. DeepHotResi leverages a pretrained protein language model to predict protein structure and generate an amino acid contact map. To enhance feature representation, DeepHotResi integrates the Squeeze-and-Excitation (SE) module, which processes diverse amino acid-level features. Next, it constructs an amino acid feature network from the contact map and SE-module-derived features. Finally, DeepHotResi employs a graph attention network to model hotspot residue prediction as a graph node classification task. Experimental results demonstrate that DeepHotResi outperforms state-of-the-art methods, effectively identifying hotspot residues in protein-RNA complexes with superior accuracy on the test set. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available at https://github.com/Q1DT/DeepHotResi.
Zhijian Huang 0001, Yuanpeng Zhang 0004, Ziyu Fan, Yuting Kong, Lei Deng 0002
Bioinform.7
2025 HLN-DDI: hierarchical molecular representation learning with co-attention mechanism for drug-drug interaction prediction
abstract
BACKGROUND: Accurate identification of drug-drug interactions (DDIs) is critical in pharmacology, as DDIs can either enhance therapeutic efficacy or trigger adverse reactions when multiple medications are administered concurrently. Traditional methods for identifying DDIs are labor-intensive and time-consuming, prompting the development of computational alternatives. However, existing computational approaches frequently encounter challenges related to interpretability and struggle to effectively capture the complex, multi-level structures inherent in drug molecules. Specifically, they often fail to adequately analyze substructural components and neglect interactions across hierarchical structural levels, resulting in incomplete molecular representations. RESULTS: In this study, we propose a Hierarchical Learning Network with a co-attention mechanism tailored to molecular structure representation for predicting DDIs, named HLN-DDI. The proposed method advances existing approaches by explicitly encoding motif-level structures and capturing hierarchical molecular representations at atom-level, motif-level, and whole-molecule scales. These hierarchical representations are integrated using a co-attention mechanism and combined with interaction-type information to enhance predictive performance. Comprehensive evaluations demonstrate that HLN-DDI significantly outperforms state-of-the-art methods across multiple benchmark datasets, achieving over 98% accuracy under transductive scenarios and surpassing 99% on various evaluation metrics. Moreover, HLN-DDI achieves a notable accuracy improvement of 2.75% in predicting DDIs involving unseen drugs. Practical assessments with real-world DDI scenarios further validate the efficacy and utility of our proposed model. CONCLUSION: By leveraging hierarchical molecular structures and employing a co-attention mechanism to effectively integrate multi-level representations, HLN-DDI generates comprehensive and precise drug representations, leading to substantially improved predictions of potential drug-drug interactions.
Lei Deng 0002, Zhijian Huang 0001
BMC Bioinform.2
2025 CircGO: Predicting circRNA Functions Through Self-Supervised Learning of Heterogeneous Networks
abstract
Circular RNAs (circRNAs), a class of non-coding RNAs characterized by their covalently closed loop structures, play active roles in diverse physiological processes through interactions with biological macromolecules. Despite the growing discovery of circRNAs enabled by high-throughput technologies, their functional annotations remain largely unexplored. This highlights the need for automated batch annotation methods to unveil the functional roles of circRNAs. In this study, we present a novel approach for predicting Gene Ontology (GO) functions associated with circRNA by leveraging self-supervised pre-training on circRNA-protein heterogeneous network. First, we construct the heterogeneous network by combining circRNA co-expression data, circRNA-protein association data, and protein-protein interaction (PPI) data. Second, we initialize the features and pseudo-labels for nodes using three graph processing methods including walking, aggregation and clustering. The initialized node features and pseudo labels, combined with protein GO annotations, are employed for heterogeneous graph pre-training. During the pre-training, the node features are learned using a heterogeneous graph attention network and the pseudo-labels are updated using the label propagation algorithm (LPA) with an attention mechanism. Finally, the initial node features are combined with those learned during pre-training to predict circRNA GO terms. Evaluation results on the independent test set reveal the superior performance of our method compared with existing approaches. Furthermore, our analysis underscores the importance of network structure and initialization strategies, highlighting the potential benefits of incorporating additional heterogeneous information and association networks.
Zhijian Huang 0001, Rongtao Zheng, Min Wu 0008, Lei Deng 0002
IEEE Trans. Comput. Biol. Bioinform.6
2025 Enhancing Predictions of Drug Solubility Through Multidimensional Structural Characterization Exploitation
abstract
Solubility is not only a significant physical property of molecules but also a vital factor in small-molecule drug development. Determining drug solubility demands stringent equipment, controlled environments, and substantial human and material resources. The accurate prediction of drug solubility using computational methods has long been a goal for researchers. In this study, we introduce MSCSol, a solubility prediction model that integrates multidimensional molecular structure information. We incorporate a graph neural network with geometric vector perceptrons (GVP-GNN) to encode 3D molecular structures, representing spatial arrangement and orientation of atoms, as well as atomic sequences and interactions. We also employ Selective Kernel Convolution combined with Global and Local attention mechanisms to capture molecular features context at different scales. Additionally, various descriptors are calculated to enrich the molecular representation. For the 2D and 3D structural data of molecules, we design different data augmentation strategies to enhance generalization ability and prevent the model from learning irrelevant information. Extensive experiments on benchmark and independent datasets demonstrate MSCSol's superior performance. Ablation studies further confirm the effectiveness of different modules. Interpretability analysis highlights the importance of various atomic groups and substructures for solubility and verifies that our model effectively captures functional molecular structures and higher-order knowledge.
Ziyu Fan, Zhijian Huang 0001, Lei Deng 0002
IEEE J. Biomed. Health Informatics5
2025 AGCLNDA: Enhancing the Prediction of ncRNA-Drug Resistance Association Using Adaptive Graph Contrastive Learning
abstract
Non-coding RNAs (ncRNAs), which do not encode proteins, have been implicated in chemotherapy resistance in cancer treatment. Given the high costs and time requirements of traditional biological experiments, there is an increasing need for computational models to predict ncRNA-drug resistance associations. In this study, we introduce AGCLNDA, an adaptive contrastive learning method designed to uncover these associations. AGCLNDA begins by constructing a bipartite graph from existing ncRNA-drug resistance data. It then utilizes a light graph convolutional network (LightGCN) to learn vector representations for both ncRNAs and drugs. The method assesses resistance association scores through the inner product of these vectors. To tackle data sparsity and noise, AGCLNDA incorporates learnable augmented view generators and denoised view generators, which provide contrastive views for enhanced data augmentation. Comparative experiments demonstrate that AGCLNDA outperforms five other advanced methods. Case studies further validate AGCLNDA as an effective tool for predicting ncRNA-drug resistance associations.
Yanhao Fan, Che Zhang, Zhijian Huang 0001, Lei Deng 0002
IEEE J. Biomed. Health Informatics5
2024 MFF-LncLoc: Subcellular Localization Prediction of lncRNAs Based on Multi-Feature Fusion Using Transformers
abstract
The subcellular localization of long non-coding RNAs (lncRNAs) is fundamental to understanding their functional roles in gene regulation and disease mechanisms. Existing prediction models typically rely on single-source features, which often fail to capture the full complexity of lncRNA localization. To address this limitation, we propose MFF-LncLoc, a Transformer-based model designed to integrate multiple feature types for improved prediction accuracy. MFF-LncLoc processes embedding matrices derived from a non-overlapping trinucleotide approach, leveraging the Transformer’s capacity to capture both contextual and positional information to extract global features. Unlike conventional models that primarily focus on k-mer frequency features, MFF-LncLoc incorporates a diverse set of features, including statistical properties, sequence characteristics, and secondary structure information. These multi-dimensional features are then processed by a Convolutional Neural Network (CNN) to extract local sequence patterns, followed by a fully connected layer for subcellular localization prediction. Ablation studies confirm that the inclusion of multi-feature data significantly enhances model performance. MFF-LncLoc outperforms existing models across several metrics, including accuracy (ACC), macro-recall, macro-F1, AUC, and AUPR, demonstrating that the integration of diverse features offers substantial improvements over traditional single-feature approaches. The source code and dataset are available at https://github.com/ZiyuFanCSU/MFFLncLoc.
Ziyu Fan, Shuning Yang, Lei Deng 0002
BIBM4
2024 LSNSCDA: Unraveling CircRNA-Drug Sensitivity via Local Smoothing Graph Neural Network and Credible Negative Samples
abstract
This study investigates the role of circular RNAs (circRNAs) in drug sensitivity, with a focus on their potential to inform personalized medicine. While current methods for identifying circRNA-drug sensitivity associations are resource-intensive, we propose LSNSCDA, a novel prediction algorithm that integrates Local Smoothing Graph Neural Networks (LS-GNN) and Credible Negative Sampling (CNS) to improve prediction accuracy. Our approach overcomes the challenges of fixed-length propagation in graph neural networks and the unreliability of randomly sampled negative instances. Experimental results show that LSNSCDA outperforms existing models, providing more reliable predictions and valuable insights into cancer treatment. Extensive evaluation confirms the effectiveness of each component of our model, while case studies further demonstrate its practical applicability. The source code and dataset are available at https://github.com/ZiyuFanCSU/LSNSCDA.
Ziyu Fan, Yuanpeng Zhang 0004, Yahan Li, Zeyu Zhong, Lei Deng 0002
BIBM5
2024 Hierarchical Molecular Structure Network with Co-attention Mechanism for Drug-Drug Interaction Prediction
abstract
Drug-drug interactions (DDIs) are crucial in pharmacology, as they can either enhance therapeutic effects or lead to harmful adverse reactions when drugs are co-administered. Accurate prediction of DDIs is essential for ensuring drug safety and efficacy. Traditional DDI prediction methods often rely on manual domain knowledge, which is labor-intensive and time-consuming. Previous computational approaches have struggled with explainability and failed to effectively capture the multilevel structural information of drug molecules, especially when analyzing substructural components. Additionally, the lack of interaction between different structural levels often results in incomplete representations of drug molecules. In this work, we propose a novel Hierarchical Molecular Structure Representation Learning Network based on a Co-attention Mechanism (HMSN-CAM) for DDI prediction. HMSN-CAM encodes motif structures and extracts hierarchical molecular representations at the atom, motif, and molecule levels. These multi-level representations are then integrated using a co-attention mechanism to predict DDIs. Our extensive evaluations demonstrate that HMSN-CAM significantly outperforms state-of-the-art methods across multiple benchmarks. The model achieves over 98% accuracy on two datasets under the transduction setting, with performance metrics exceeding 99% for most evaluation criteria. Notably, HMSN-CAM also improves DDI prediction for pairs involving previously unseen drugs, yielding a 2.75% accuracy improvement over current methods. These results highlight the potential of HMSN-CAM for enhancing the prediction of drug interactions and improving drug safety.
Lei Deng 0002
BIBM2
2024 Enhancing miRNA-Disease Prediction with Attention-Based Multi-View Contrastive Learning
abstract
MicroRNAs (miRNAs) play a critical role in gene expression regulation, and their dysregulation is implicated in diseases such as cancer and cardiovascular disorders. Identifying miRNA-disease associations is therefore crucial for advancing prevention and treatment strategies for these conditions. We introduce MCLAMDA (Multi-view Contrastive Learning with Attention Mechanism for miRNA-Disease Association Prediction), a novel method inspired by contrastive learning and attention mechanisms. MCLAMDA constructs diverse similarity networks for miRNAs and diseases and employs both topological and feature contrastive learning to capture comprehensive node information. An attention mechanism further enhances the model by weighting and aggregating node features based on the relative importance of their neighboring nodes. Our approach outperforms five state-of-the-art methods in predictive accuracy. Case studies provide additional support for MCLAMDA’s robustness and reliable predictive performance. The code and data for MCLAMDA are available at https://github.com/Ouyang-cmd/MCLAMDA.
Tianxiang Ouyang, Zhijian Huang 0001, Lei Deng 0002
BIBM3
2024 DTA-Net: Dual-Task Attention Network for Medical Image Segmentation and Classification
abstract
Medical image segmentation and classification are fundamental tasks in computer-aided diagnosis, where accurate segmentation plays a key role in identifying disease-related features and regions of interest, thus aiding subsequent classification. In this paper, we propose a novel Dual-Task Attention Network (DTA-Net), which simultaneously generates high-quality segmentation masks and performs image classification by incorporating Electronic Medical Records (EMR). The DTA-Net architecture introduces an innovative Shared Proxy Attention Module (SPAM), which leverages a shared mapping function to effectively encode the attention weights of both queries and keys for spatial and channel attention. This dual attention mechanism facilitates complementary learning and interdependence between spatial and channel features. Notably, the proxy projection mechanism within SPAM significantly reduces the computational complexity of spatial attention to a linear level. We validated our approach on three datasets: MSD Prostate for prostate segmentation, the RICORD dataset for lung segmentation, and the iCTCF COVID-19 dataset for both segmentation and classification tasks. Experimental results demonstrate that the proposed network achieves promising performance across all datasets. The source code is available at: https://github.com/HaifengQi/DTA-Net.
Haifeng Qi, Zhijian Huang 0001, Lei Deng 0002
BIBM4
2024 Imputation of incomplete medical data using missing neighborhood perturbation denoising autoencoder
abstract
The problem of missing data in the clinical environment is a common one. However, the presence of these missing values or incorrect processing can affect downstream tasks. Imputation is the most effective way to solve the missing problem. Combining deep generative models and multiple imputation idea provides new ideas for the imputation of data, but some methods do not apply to mixed-type tabular data such as clinical data. The missing types of tabular data in real-world scenarios are varied and unpredictable, so it is necessary to construct a generalized imputation method. In this paper, we propose a multiple imputation architecture for missing neighborhood perturbation denoising autoencoder (MI_NPDAE) for imputing mixed types of incomplete medical tabular data for different missing mechanisms. In MI_NPDAE, the neighborhood information of the data is searched to find the best set of donors to be used as the input layer of the network, and the learning of the neighborhood features of the missing locations is ensured by reconstructing new inputs, while the addition of additive noise exposes the model to different missing mechanisms. The experiments were respectively conducted on a public dataset Breast and one lung cancer nutrition dataset from the Chinese Anti-Cancer Society. The experiments on the datasets show that the proposed model shows better results under different missing mechanisms and can maintain a low error on the comparison experiments on data with different missing ratios. In addition, the imputed results also show significant advantages in downstream prediction tasks.
Sen Fu, Lei Deng 0002
BIBM4
2024 A Self-Attention Synthesizing Model with Privacy-Preserving(ACCT-GAN) for Medical Tabular Data
abstract
Accurate clinical cancer prediction is important for clinical diagnosis. Most of the cancer patient data have high privacy and using clinical patient data directly may create concerns such as privacy leakage, which can be solved by replacing real data with synthetic data for clinical trial modelling. Synthetic data can also fill in the missing, unbalanced clinical datasets. Data augmentation can achieve data synthesis. However, traditional data augmentation models are mostly based on convolutional neural network architectures, which can only capture local dependencies and do not well describe the feature-related information of patients over long survival times. In this paper, we propose a self-attention medical tabular data augmentation model with privacy protection, ACCT-GAN, based on the generative adversarial network model, which uses the AC block generated by combining the ACmix model to reconstruct the discriminator model in the CTAB-GAN+ network, which can effectively solve the limitations of the original convolutional neural network, model the global relationship, and make the generated data retain the feature correlation between the original data better correlations between the original data. Meanwhile, taking into account the privacy of medical data, the DP-SGD algorithm of CTAB-GAN+ is optimised to reduce the privacy risk of synthetic data. The results show that ACCT-GAN synthesizes privacy-preserving data with at least 26.21% higher utility across the dateset of nasopharyngeal cancer patients provided by Xiangya Medical College and learning tasks under different privacy budgets, demonstrating the usability, high quality, and better privacy preservation of the generated data, validating the effectiveness of this paper’s method, and confirming the potential of this paper’s method for cancer datasets.
Wenqin Zou, Sen Fu, Lei Deng 0002
BIBM4
2024 GATBind: Accurate protein-RNA binding sites prediction via graph attention networks with pre-trained language model
abstract
Protein-RNA interactions play a pivotal role in various biological processes, making them essential for discovering novel therapeutic targets. Understanding these interactions is crucial for identifying potential drug targets and designing effective therapeutics. However, traditional experimental approaches are costly and time-consuming, highlighting the need for accurate and efficient computational methods to predict RNA-binding sites. In this study, we propose a novel method called GATBind, designed to predict RNA-binding sites on proteins. GATBind introduces a local environment-aware module that captures the local structural information surrounding target residues. Additionally, we incorporate sequence and structural features, including PSSM, HMM, DSSP, and sequence embeddings generated by the pre-trained protein language model ESMFold-2. These features are integrated and processed through a Graph Attention Network (GAT) module to predict RNA-binding sites accurately. Experimental results demonstrate that GATBind outperforms existing state-of-the-art methods, and case studies further validate its effectiveness in predicting RNA-binding sites.
Shangqian Wu, Zhijian Huang 0001, Lei Deng 0002
BIBM5
2024 ADiffGDA: Exploring Gene-Drug Associations via Adaptive Graph Diffusion Networks
abstract
Exploring gene-drug associations is a key step in identifying new drug candidates, but traditional experimental methods are often expensive and time-consuming. While Graph Neural Network (GNN)-based models have demonstrated effectiveness in association prediction tasks, they face challenges in the information aggregation process. Existing GNN models either treat all nodes uniformly or rely on simple attention mechanisms to assign weights to neighboring nodes, limiting their capacity to capture complex relationships and improve performance. To address these limitations, we propose a novel adaptive graph diffusion network, ADiffGDA, for gene-drug association prediction. The model begins by randomly initializing embeddings for genes and drugs, which are then updated through neighborhood information aggregation. A key feature of ADiffGDA is the incorporation of a heat kernel, enabling each node to dynamically adjust aggregation weights based on its local structure. This approach allows the model to better capture variations in node types and local neighborhood patterns. Through extensive comparative experiments, we show that ADiffGDA outperforms existing state-of-the-art methods. Furthermore, case studies validate its effectiveness as a predictive tool, offering valuable insights for future biological experiments. The code and datasets for ADiffGDA are freely available at https://github.com/one-melon/ADiffGDA.
Che Zhang, Yanhao Fan, Yurong Qian, Lei Deng 0002
BIBM5
2024 HGTRDA: Enhancing Prediction of ncRNA-Mediated Drug Resistance with Hypergraph Transformer
abstract
Exploring the intricate connections between non-coding RNAs (ncRNAs) and drug resistance is crucial for understanding the molecular mechanisms behind drug resistance, identifying novel drug development targets, and uncovering key biomarkers to optimize therapeutic strategies. Traditional biological assays face significant challenges, including high costs and lengthy timelines, prompting the need for advanced computational methods to predict ncRNA-drug resistance associations. In this study, we introduce HGTRDA, a novel computational framework designed to predict potential associations between ncRNAs and drug resistance. HGTRDA leverages LightGCN to generate node representations that capture topological information from the surrounding node neighborhood. These representations are then dynamically optimized using a global hypergraph transformer to model the relationships between ncRNAs and drug resistance. To enhance the quality of the learned embeddings, HGTRDA employs self-supervised learning to fine-tune topology-aware embeddings, reducing the impact of noise and improving representation quality. The final association scores between ncRNAs and drugs are computed using an inner product method. Empirical evaluations on the ncRNADrug database demonstrate that HGTRDA outperforms six contemporary state-of-the-art methods in predicting ncRNA-drug resistance associations. Furthermore, case studies illustrate the practical utility of HGTRDA as a predictive tool in real-world scenarios. The code and dataset for HGTRDA are freely available at https://github.com/one-melon/HGTRDA.
Che Zhang, Ruohui He, Yanhao Fan, Yurong Qian, Lei Deng 0002
BIBM6
2024 SGCLDGA: unveiling drug-gene associations through simple graph contrastive learning
abstract
Drug repurposing offers a viable strategy for discovering new drugs and therapeutic targets through the analysis of drug-gene interactions. However, traditional experimental methods are plagued by their costliness and inefficiency. Despite graph convolutional network (GCN)-based models' state-of-the-art performance in prediction, their reliance on supervised learning makes them vulnerable to data sparsity, a common challenge in drug discovery, further complicating model development. In this study, we propose SGCLDGA, a novel computational model leveraging graph neural networks and contrastive learning to predict unknown drug-gene associations. SGCLDGA employs GCNs to extract vector representations of drugs and genes from the original bipartite graph. Subsequently, singular value decomposition (SVD) is employed to enhance the graph and generate multiple views. The model performs contrastive learning across these views, optimizing vector representations through a contrastive loss function to better distinguish positive and negative samples. The final step involves utilizing inner product calculations to determine association scores between drugs and genes. Experimental results on the DGIdb4.0 dataset demonstrate SGCLDGA's superior performance compared with six state-of-the-art methods. Ablation studies and case analyses validate the significance of contrastive learning and SVD, highlighting SGCLDGA's potential in discovering new drug-gene associations. The code and dataset for SGCLDGA are freely available at https://github.com/one-melon/SGCLDGA.
Yanhao Fan, Che Zhang, Zhijian Huang 0001, Jiameng Xue, Lei Deng 0002
Briefings Bioinform.6
2024 IGCNSDA: unraveling disease-associated snoRNAs with an interpretable graph convolutional network
abstract
Accurately delineating the connection between short nucleolar RNA (snoRNA) and disease is crucial for advancing disease detection and treatment. While traditional biological experimental methods are effective, they are labor-intensive, costly and lack scalability. With the ongoing progress in computer technology, an increasing number of deep learning techniques are being employed to predict snoRNA-disease associations. Nevertheless, the majority of these methods are black-box models, lacking interpretability and the capability to elucidate the snoRNA-disease association mechanism. In this study, we introduce IGCNSDA, an innovative and interpretable graph convolutional network (GCN) approach tailored for the efficient inference of snoRNA-disease associations. IGCNSDA leverages the GCN framework to extract node feature representations of snoRNAs and diseases from the bipartite snoRNA-disease graph. SnoRNAs with high similarity are more likely to be linked to analogous diseases, and vice versa. To facilitate this process, we introduce a subgraph generation algorithm that effectively groups similar snoRNAs and their associated diseases into cohesive subgraphs. Subsequently, we aggregate information from neighboring nodes within these subgraphs, iteratively updating the embeddings of snoRNAs and diseases. The experimental results demonstrate that IGCNSDA outperforms the most recent, highly relevant methods. Additionally, our interpretability analysis provides compelling evidence that IGCNSDA adeptly captures the underlying similarity between snoRNAs and diseases, thus affording researchers enhanced insights into the snoRNA-disease association mechanism. Furthermore, we present illustrative case studies that demonstrate the utility of IGCNSDA as a valuable tool for efficiently predicting potential snoRNA-disease associations. The dataset and source code for IGCNSDA are openly accessible at: https://github.com/altriavin/IGCNSDA.
Dayun Liu, Yuanpeng Zhang 0004, Yihan Dong, Yanhao Fan, Lei Deng 0002
Briefings Bioinform.8
2024 MvMRL: a multi-view molecular representation learning method for molecular property prediction
abstract
Effective molecular representation learning is very important for Artificial Intelligence-driven Drug Design because it affects the accuracy and efficiency of molecular property prediction and other molecular modeling relevant tasks. However, previous molecular representation learning studies often suffer from limitations, such as over-reliance on a single molecular representation, failure to fully capture both local and global information in molecular structure, and ineffective integration of multiscale features from different molecular representations. These limitations restrict the complete and accurate representation of molecular structure and properties, ultimately impacting the accuracy of predicting molecular properties. To this end, we propose a novel multi-view molecular representation learning method called MvMRL, which can incorporate feature information from multiple molecular representations and capture both local and global information from different views well, thus improving molecular property prediction. Specifically, MvMRL consists of four parts: a multiscale CNN-SE Simplified Molecular Input Line Entry System (SMILES) learning component and a multiscale Graph Neural Network encoder to extract local feature information and global feature information from the SMILES view and the molecular graph view, respectively; a Multi-Layer Perceptron network to capture complex non-linear relationship features from the molecular fingerprint view; and a dual cross-attention component to fuse feature information on the multi-views deeply for predicting molecular properties. We evaluate the performance of MvMRL on 11 benchmark datasets, and experimental results show that MvMRL outperforms state-of-the-art methods, indicating its rationality and effectiveness in molecular property prediction. The source code of MvMRL was released in https://github.com/jedison-github/MvMRL.
Yanmei Lin, Yijia Wu, Lei Deng 0002, Mingzhi Liao, Yuzhong Peng
Briefings Bioinform.4
2024 MolMVC: Enhancing molecular representations for drug-related tasks through multi-view contrastive learning
abstract
MOTIVATION: Effective molecular representation is critical in drug development. The complex nature of molecules demands comprehensive multi-view representations, considering 1D, 2D, and 3D aspects, to capture diverse perspectives. Obtaining representations that encompass these varied structures is crucial for a holistic understanding of molecules in drug-related contexts. RESULTS: In this study, we introduce an innovative multi-view contrastive learning framework for molecular representation, denoted as MolMVC. Initially, we use a Transformer encoder to capture 1D sequence information and a Graph Transformer to encode the intricate 2D and 3D structural details of molecules. Our approach incorporates a novel attention-guided augmentation scheme, leveraging prior knowledge to create positive samples tailored to different molecular data views. To align multi-view molecular positive samples effectively in latent space, we introduce an adaptive multi-view contrastive loss (AMCLoss). In particular, we calculate AMCLoss at various levels within the model to effectively capture the hierarchical nature of the molecular information. Eventually, we pre-train the encoders via minimizing AMCLoss to obtain the molecular representation, which can be used for various down-stream tasks. In our experiments, we evaluate the performance of our MolMVC on multiple tasks, including molecular property prediction (MPP), drug-target binding affinity (DTA) prediction and cancer drug response (CDR) prediction. The results demonstrate that the molecular representation learned by our MolMVC can enhance the predictive accuracy on these tasks and also reduce the computational costs. Furthermore, we showcase MolMVC's efficacy in drug repositioning across a spectrum of drug-related applications. AVAILABILITY AND IMPLEMENTATION: The code and pre-trained model are publicly available at https://github.com/Hhhzj-7/MolMVC.
Zhijian Huang 0001, Ziyu Fan, Min Wu 0008, Lei Deng 0002
Bioinform.5
2024 DeepRSMA: a cross-fusion-based deep learning method for RNA-small molecule binding affinity prediction
abstract
MOTIVATION: RNA is implicated in numerous aberrant cellular functions and disease progressions, highlighting the crucial importance of RNA-targeted drugs. To accelerate the discovery of such drugs, it is essential to develop an effective computational method for predicting RNA-small molecule affinity (RSMA). Recently, deep learning-based computational methods have been promising due to their powerful nonlinear modeling ability. However, the leveraging of advanced deep learning methods to mine the diverse information of RNAs, small molecules, and their interaction still remains a great challenge. RESULTS: In this study, we present DeepRSMA, an innovative cross-attention-based deep learning method for RSMA prediction. To effectively capture fine-grained features from RNA and small molecules, we developed nucleotide-level and atomic-level feature extraction modules for RNA and small molecules, respectively. Additionally, we incorporated both sequence and graph views into these modules to capture features from multiple perspectives. Moreover, a transformer-based cross-fusion module is introduced to learn the general patterns of interactions between RNAs and small molecules. To achieve effective RSMA prediction, we integrated the RNA and small molecule representations from the feature extraction and cross-fusion modules. Our results show that DeepRSMA outperforms baseline methods in multiple test settings. The interpretability analysis and the case study on spinal muscular atrophy demonstrate that DeepRSMA has the potential to guide RNA-targeted drug design. AVAILABILITY AND IMPLEMENTATION: The codes and data are publicly available at https://github.com/Hhhzj-7/DeepRSMA.
Zhijian Huang 0001, Yucheng Wang 0001, Yaw Sing Tan, Lei Deng 0002, Min Wu 0008
Bioinform.5
2024 Exploring ncRNA-Drug Sensitivity Associations via Graph Contrastive Learning
abstract
Increasing evidence has shown that noncoding RNAs (ncRNAs) can affect drug efficiency by modulating drug sensitivity genes. Exploring the association between ncRNAs and drug sensitivity is essential for drug discovery and disease prevention. However, traditional biological experiments for identifying ncRNA-drug sensitivity associations are time-consuming and laborious. In this study, we develop a novel graph contrastive learning approach named NDSGCL to predict ncRNA-drug sensitivity. NDSGCL uses graph convolutional networks to learn feature representations of ncRNAs and drugs in ncRNA-drug bipartite graphs. It integrates local structural neighbours and global semantic neighbours to learn a more comprehensive representation by contrastive learning. Specifically, the local structural neighbours aim to capture the higher-order relationship in the ncRNA-drug graph, while the global semantic neighbours are defined based on semantic clusters of the graph that can alleviate the impact of data sparsity. The experimental results show that NDSGCL outperforms basic graph convolutional network methods, existing contrastive learning methods, and state-of-the-art prediction methods. Visualization experiments show that the contrastive objectives of local structural neighbours and global semantic neighbours play a significant role in contrastive learning. Case studies on two drugs show that NDSGCL is an effective tool for predicting ncRNA-drug sensitivity associations.
Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.3
2024 DeepFusionCDR: Employing Multi-Omics Integration and Molecule-Specific Transformers for Enhanced Prediction of Cancer Drug Responses
abstract
Deep learning approaches have demonstrated remarkable potential in predicting cancer drug responses (CDRs), using cell line and drug features. However, existing methods predominantly rely on single-omics data of cell lines, potentially overlooking the complex biological mechanisms governing cell line responses. This paper introduces DeepFusionCDR, a novel approach employing unsupervised contrastive learning to amalgamate multi-omics features, including mutation, transcriptome, methylome, and copy number variation data, from cell lines. Furthermore, we incorporate molecular SMILES-specific transformers to derive drug features from their chemical structures. The unified multi-omics and drug signatures are combined, and a multi-layer perceptron (MLP) is applied to predict IC50 values for cell line-drug pairs. Moreover, this MLP can discern whether a cell line is resistant or sensitive to a particular drug. We assessed DeepFusionCDR's performance on the GDSC dataset and juxtaposed it against cutting-edge methods, demonstrating its superior performance in regression and classification tasks. We also conducted ablation studies and case analyses to exhibit the effectiveness and versatility of our proposed approach. Our results underscore the potential of DeepFusionCDR to enhance CDR predictions by harnessing the power of multi-omics fusion and molecular-specific transformers. The prediction of DeepFusionCDR on TCGA patient data and case study highlight the practical application scenarios of DeepFusionCDR in real-world environments.
Lei Deng 0002
IEEE J. Biomed. Health Informatics4
2024 Tab-Cox: An Interpretable Deep Survival Analysis Model for Patients With Nasopharyngeal Carcinoma Based on TabNet
abstract
The nutritional status of cancer patients is closely associated with the clinical progression of the disease. A survival analysis model combined with a neural network can predict future disease trends in patients, facilitating early prevention and assisting physicians in making diagnoses. However, the complexity of neural networks and their incompatibility with medical tabular data can reduce the interpretability of the model. To address this issue, thr paper propose a novel survival analysis model called Tab-Cox, which combines TabNet and Cox models. This model is specifically designed to predict the survival outcomes of patients with nasopharyngeal carcinoma. The model utilizes TabNet's sequential attention mechanism to extract more interpretable features, providing an interpretable method for identifying disease risk factors. Consequently, the model ensures accurate survival prediction while also making the results more comprehensible for both patients and doctors. The paper tested the efficacy of the model by conducting experiments on various diverse datasets in comparison with other commonly used survival models. The results showed that the proposed model delivered the highest or second-highest accuracy across all datasets. Furthermore, the paper conducted a comparative interpretability analysis against the classical Cox model. In addition and compare the interpretability of the Tab-Cox model with the classical Cox model and discuss the advantages and disadvantages of its interpretability. This demonstrates that Tab-Cox can assist doctors in identifying risk factors that are challenging to capture using artificial methods.
Ruohao Fan, Lei Deng 0002
IEEE J. Biomed. Health Informatics4
2024 AntiViralDL: Computational Antiviral Drug Repurposing Using Graph Neural Network and Self-Supervised Learning
abstract
Viral infections have emerged as significant public health concerns for decades. Antiviral drugs, specifically designed to combat these infections, have the potential to reduce the disease burden substantially. However, traditional drug development methods, based on biological experiments, are resource-intensive, time-consuming, and low efficiency. Therefore, computational approaches for identifying antiviral drugs can enhance drug development efficiency. In this study, we introduce AntiViralDL, a computational framework for predicting virus-drug associations using self-supervised learning. Initially, we construct a reliable virus-drug association dataset by integrating the existing Drugvirus2 database and FDA-approved virus-drug associations. Utilizing these two datasets, we create a virus-drug association bipartite graph and employ the Light Graph Convolutional Network (LightGCN) to learn embedding representations of viruses and drugs. To address the sparsity of virus-drug association pairs, AntiViralDL incorporates contrastive learning to improve prediction accuracy. We implement data augmentation by adding random noise to the embedding representation space of virus and drug nodes, as opposed to traditional edge and node dropout. Finally, we calculate an inner product to predict virus-drug association relationships. Experimental results reveal that AntiViralDL achieves AUC and AUPR values of 0.8450 and 0.8494, respectively, outperforming four benchmarked virus-drug association prediction models. The case study further highlights the efficacy of AntiViralDL in predicting anti-COVID-19 drug candidates.
Guangdi Li, Lei Deng 0002
IEEE J. Biomed. Health Informatics4
2023 Enhancing Protein Solubility Prediction through Pre-trained Language Models and Graph Convolutional Neural Networks
abstract
Achieving optimal protein solubility is pivotal for efficient high-throughput purification, especially in industrial settings. However, conventional experimental techniques for assessing protein solubility in such contexts are not only costly but also time-intensive. Currently, numerous methods are available for predicting protein solubility, yet their effectiveness remains limited. Most of these approaches are predominantly sequence-based, failing to harness the invaluable structural insights inherent in proteins. Addressing these limitations, we introduce PPSol, an innovative protein solubility prediction methodology. Operating on protein sequences, PPSol employs ESM2 to predict protein contact maps, forming the basis for constructing protein graphs. Subsequently, well-established techniques are employed to predict protein feature representations as node features, including the utilization of the Position-Specific Scoring Matrix (PSSM). The resulting graph is fed into a graph convolutional neural network (GCN), enabling the acquisition of spatial structural information from proteins. Concurrently, ESM2-generated features undergo dimensional reduction via fully connected layers, integrating into every layer of the GCN for precise protein solubility prediction. Our approach excels through the fusion of pre-trained protein language models and GCNs, surpassing existing methodologies. Notably, PPSol attains state-of-the-art performance, showcasing a remarkable 2.8% enhancement in AUROC performance com-pared to prior strategies.
Yurong Qian, Zhijian Huang 0001, Xiaojun Xiao, Lei Deng 0002
BIBM5
2023 TGC-ARG: Predicting Antibiotic Resistance through Transformer-based Modeling and Contrastive Learning
abstract
The escalating severity of antibiotic resistance poses substantial challenges across diverse sectors, encompassing everyday life, agriculture, and clinical medical interventions. Conventional methods for investigating antibiotic resistance genes (ARGs), such as culture-based techniques and whole-genome sequencing, often suffer from demands of time, labor, and limited accuracy. Moreover, the fragmented nature of existing datasets hampers a comprehensive analysis of antibiotic resistance gene sequences. In this study, we introduce an innovative computational framework known as TGC-ARG, designed to predict potential ARGs. TGC-ARG harnesses protein sequences as input, retrieves protein structures through SCRATCH-1D, and employs a feature extraction module to deduce feature representations for both protein sequences and structures. Subsequently, we integrate a siamese network to establish a contrastive learning paradigm, thus augmenting the model’s representational capabilities. The resultant sequence embeddings and structure embeddings are merged and directed into a Multilayer Perceptron (MLP) for predicting ARG presence. To assess the performance, we curate a pioneering publicly available dataset named ARSS (Antibiotic Resistance Sequence Statistics). Our extensive comparative experimental outcomes underscore the superiority of our approach over the current state-of-the-art (SOTA) methodology. Furthermore, through comprehensive case analyses, we demonstrate the efficacy of our approach in predicting potential ARGs. The dataset and source code are accessible at https://github.com/angel1gel/TGC-ARG.
Yihan Dong, Zhijian Huang 0001, Lei Deng 0002
BIBM4
2023 CLPiDA: A Contrastive Learning Approach for Predicting Potential PiRNA-Disease Associations
abstract
Piwi-interacting RNAs (piRNAs) function as critical regulators, safeguarding genome stability through mechanisms like transposable element repression and gene stability maintenance, while also being associated with various disease pathways. Developing computationally efficient methods to predict piRNA-disease associations is vital for enhancing disease-specific drug discovery while managing costs. In this study, we present CLPiDA, a novel method for predicting potential piRNA-disease associations. CLPiDA begins by computing gaussian kernel similarities for piRNA-piRNA and disease-disease pairs to establish initial embeddings for piRNAs and diseases. Subsequently, it employs a parameter-sharing online and target network, along with data augmentation techniques, to create a contrastive learning framework. This facilitates the generation of embeddings for piRNAs and diseases using piRNA-disease association pairs. Furthermore, CLPiDA employs a cross-prediction approach to determine association scores for specific piRNAs and diseases. Notably, CLPiDA introduces a novel approach by excluding negative samples, thereby avoiding the introduction of false negatives and enhancing its reliability. Comparative experiments validate CLPiDA’s superiority in terms of performance over existing methods. Case studies underscore CLPiDA’s efficacy as a valuable tool for predicting piRNA-disease associations, providing valuable insights for biological experiments. The data and source code for CLPiDA are available at https://github.com/altriavin/CLPiDA.
Yuanpeng Zhang 0004, Lei Deng 0002
BIBM4
2023 Predicting Associations between circRNAs and Drug Sensitivity using Heterogeneous Graphs and Graph Attention Networks
abstract
Discovering associations between circular RNAs (circRNAs) and cellular drug sensitivity is essential for understanding drug efficacy and therapeutic resistance. Traditional experimental methods to verify such associations are costly and time-consuming. Thus, the development of efficient computational methods for predicting circRNA-drug associations is crucial. In this study, we introduce a novel computational predictor called HETACDA, aimed at predicting potential circRNA-drug sensitivity associations. HETACDA constructs a heterogeneous graph network, incorporating the characteristic structure graph of drugs and circRNAs, along with the circRNA-drug sensitivity topology graph. By employing a graph convolutional network, the drug embedding vector is computed from the molecular structure of the drug. Through a graph attention mechanism, HETACDA assigns distinct attention weights to nodes in order to emphasize the contribution of various neighborhood nodes to the central node. Subsequently, an association score between circRNA-drug sensitivity is predicted using a three-layer fully connected neural network. Extensive experimental comparisons against several state-of-the-art methods highlight the effectiveness of our proposed framework. The availability of our source code and datasets on GitHub (https://github.com/xiaoxiaojun131421/HETACDA) facilitates replication and further research in this area.
Xiaojun Xiao, Yurong Qian, Zhijian Huang 0001, Rongtao Zheng, Lei Deng 0002
BIBM5
2023 PTDA-SWGCL: Predicting tRNA-Disease Associations using Supplementarily Weighted Graph Contrastive Learning
abstract
tRNAs play a pivotal role in protein synthesis by transporting amino acids to the ribosome according to mRNA instructions. These molecules are essential regulators in various biological processes, and their dysregulation is closely linked to human diseases. Predicting associations between tRNAs and diseases is valuable for uncovering biomarkers that aid in disease prevention, detection, prognosis, diagnosis, and treatment. However, experimental validation of such associations is resource-intensive, necessitating the development of robust computational methods. In this study, we propose PTDA-SWGCL, a novel model for predicting potential tRNA-disease associations. PTDA-SWGCL integrates tRNA and disease similarity information derived from Gaussian kernel similarity, sequence similarity, and semantic similarity. It initializes tRNA and disease embeddings using this similarity information and refines them through supplementarily weight and graph comparison learning training on the tRNA-disease association graph. The final association pair prediction is obtained by the inner product of the tRNA and disease embeddings. Experimental results demonstrate that PTDA-SWGCL outperforms state-of-the-art methods. Case studies confirm its effectiveness in predicting tRNA-disease associations. The code and data are available at https://github.com/ZYPssss/PTDA-SWGCL.
Yuanpeng Zhang 0004, Yurong Qian, Xiaojun Xiao, Zhijian Huang 0001, Lei Deng 0002
BIBM6
2023 A comprehensive review and evaluation of graph neural networks for non-coding RNA and complex disease associations
abstract
Non-coding RNAs (ncRNAs) play a critical role in the occurrence and development of numerous human diseases. Consequently, studying the associations between ncRNAs and diseases has garnered significant attention from researchers in recent years. Various computational methods have been proposed to explore ncRNA-disease relationships, with Graph Neural Network (GNN) emerging as a state-of-the-art approach for ncRNA-disease association prediction. In this survey, we present a comprehensive review of GNN-based models for ncRNA-disease associations. Firstly, we provide a detailed introduction to ncRNAs and GNNs. Next, we delve into the motivations behind adopting GNNs for predicting ncRNA-disease associations, focusing on data structure, high-order connectivity in graphs and sparse supervision signals. Subsequently, we analyze the challenges associated with using GNNs in predicting ncRNA-disease associations, covering graph construction, feature propagation and aggregation, and model optimization. We then present a detailed summary and performance evaluation of existing GNN-based models in the context of ncRNA-disease associations. Lastly, we explore potential future research directions in this rapidly evolving field. This survey serves as a valuable resource for researchers interested in leveraging GNNs to uncover the complex relationships between ncRNAs and diseases.
Dayun Liu, Yanhao Fan, Tianxiang Ouyang, Yuanpeng Zhang 0004, Lei Deng 0002
Briefings Bioinform.8
2023 Large-scale predicting protein functions through heterogeneous feature fusion
abstract
As the volume of protein sequence and structure data grows rapidly, the functions of the overwhelming majority of proteins cannot be experimentally determined. Automated annotation of protein function at a large scale is becoming increasingly important. Existing computational prediction methods are typically based on expanding the relatively small number of experimentally determined functions to large collections of proteins with various clues, including sequence homology, protein-protein interaction, gene co-expression, etc. Although there has been some progress in protein function prediction in recent years, the development of accurate and reliable solutions still has a long way to go. Here we exploit AlphaFold predicted three-dimensional structural information, together with other non-structural clues, to develop a large-scale approach termed PredGO to annotate Gene Ontology (GO) functions for proteins. We use a pre-trained language model, geometric vector perceptrons and attention mechanisms to extract heterogeneous features of proteins and fuse these features for function prediction. The computational results demonstrate that the proposed method outperforms other state-of-the-art approaches for predicting GO functions of proteins in terms of both coverage and accuracy. The improvement of coverage is because the number of structures predicted by AlphaFold is greatly increased, and on the other hand, PredGO can extensively use non-structural information for functional prediction. Moreover, we show that over 205 000 ($\sim $100%) entries in UniProt for human are annotated by PredGO, over 186 000 ($\sim $90%) of which are based on predicted structure. The webserver and database are available at http://predgo.denglab.org/.
Rongtao Zheng, Zhijian Huang 0001, Lei Deng 0002
Briefings Bioinform.3
2023 DeepCoVDR: deep transfer learning with graph transformer and cross-attention for predicting COVID-19 drug response
abstract
MOTIVATION: The coronavirus disease 2019 (COVID-19) remains a global public health emergency. Although people, especially those with underlying health conditions, could benefit from several approved COVID-19 therapeutics, the development of effective antiviral COVID-19 drugs is still a very urgent problem. Accurate and robust drug response prediction to a new chemical compound is critical for discovering safe and effective COVID-19 therapeutics. RESULTS: In this study, we propose DeepCoVDR, a novel COVID-19 drug response prediction method based on deep transfer learning with graph transformer and cross-attention. First, we adopt a graph transformer and feed-forward neural network to mine the drug and cell line information. Then, we use a cross-attention module that calculates the interaction between the drug and cell line. After that, DeepCoVDR combines drug and cell line representation and their interaction features to predict drug response. To solve the problem of SARS-CoV-2 data scarcity, we apply transfer learning and use the SARS-CoV-2 dataset to fine-tune the model pretrained on the cancer dataset. The experiments of regression and classification show that DeepCoVDR outperforms baseline methods. We also evaluate DeepCoVDR on the cancer dataset, and the results indicate that our approach has high performance compared with other state-of-the-art methods. Moreover, we use DeepCoVDR to predict COVID-19 drugs from FDA-approved drugs and demonstrate the effectiveness of DeepCoVDR in identifying novel COVID-19 drugs. AVAILABILITY AND IMPLEMENTATION: https://github.com/Hhhzj-7/DeepCoVDR.
Zhijian Huang 0001, Lei Deng 0002
Bioinform.3
2023 GCNPCA: miRNA-Disease Associations Prediction Algorithm Based on Graph Convolutional Neural Networks
abstract
A growing number of studies have confirmed the important role of microRNAs (miRNAs) in human diseases and the aberrant expression of miRNAs affects the onset and progression of human diseases. The discovery of disease-associated miRNAs as new biomarkers promote the progress of disease pathology and clinical medicine. However, only a small proportion of miRNA-disease correlations have been validated by biological experiments. And identifying miRNA-disease associations through biological experiments is both expensive and inefficient. Therefore, it is important to develop efficient and highly accurate computational methods to predict miRNA-disease associations. A miRNA-disease associations prediction algorithm based on Graph Convolutional neural Networks and Principal Component Analysis (GCNPCA) is proposed in this paper. Specifically, the deep topological structure information is extracted from the heterogeneous network composed of miRNA and disease nodes by a Graph Convolutional neural Network (GCN) with an additional attention mechanism. The internal attribute information of the nodes is obtained by the Principal Component Analysis (PCA). Then, the topological structure information and the node attribute information are combined to construct comprehensive feature descriptors. Finally, the Random Forest (RF) is used to train and classify these feature descriptors. In the five-fold cross-validation experiment, the AUC and AUPR for the GCNPCA algorithm are 0.983 and 0.988 respectively.
Jiwen Liu, Zhufang Kuang, Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 HGNNLDA: Predicting lncRNA-Drug Sensitivity Associations via a Dual Channel Hypergraph Neural Network
abstract
Drug sensitivity is critical for enabling personalized treatment. Many studies have shown that long non-coding RNAs (lncRNAs) are closely related to drug sensitivity because lncRNAs can regulate genes related to drug sensitivity to affect drug efficacy. Exploring lncRNA-drug sensitivity associations has important implications for drug development and disease treatment. However, identifying lncRNA-drug sensitivity associations based on traditional biological approaches is small-scale and time-consuming. In this work, we develop a dual-channel hypergraph neural network-based method named HGNNLDA to infer unknown lncRNA-drug sensitivity associations. To our best knowledge, HGNNLDA is the first computational framework to predict lncRNA-drug sensitivity associations. HGNNLDA applies the hypergraph neural network to obtain high-order neighbor information on the lncRNA hypergraph and the drug hypergraph, respectively, and utilizes a joint update mechanism to generate lncRNA embeddings and drug embeddings. In traditional graphs, an edge contains only two nodes. However, hyperedges in hypergraphs can contain any number of nodes and hypergraphs can well describe the higher-order connectivity of the lncRNA-drug bipartite graphs. The comprehensive experimental results show that HGNNLDA significantly outperforms the other six state-of-the-art models. Case studies on two drugs further illustrate that HGNNLDA is an effective tool to predict lncRNA-drug sensitivity associations.
Dayun Liu, Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.7
2023 NGCICM: A Novel Deep Learning-Based Method for Predicting circRNA-miRNA Interactions
abstract
The circRNAs and miRNAs play an important role in the development of human diseases, and they can be widely used as biomarkers of diseases for disease diagnosis. In particular, circRNAs can act as sponge adsorbers for miRNAs and act together in certain diseases. However, the associations between the vast majority of circRNAs and diseases and between miRNAs and diseases remain unclear. Computational-based approaches are urgently needed to discover the unknown interactions between circRNAs and miRNAs. In this paper, we propose a novel deep learning algorithm based on Node2vec and Graph ATtention network (GAT), Conditional Random Field (CRF) layer and Inductive Matrix Completion (IMC) to predict circRNAs and miRNAs interactions (NGCICM). We construct a GAT-based encoder for deep feature learning by fusing the talking-heads attention mechanism and the CRF layer. The IMC-based decoder is also constructed to obtain interaction scores. The Area Under the receiver operating characteristic Curve (AUC) of the NGCICM method is 0.9697, 0.9932 and 0.9980, and the Area Under the Precision-Recall curve (AUPR) is 0.9671, 0.9935 and 0.9981, respectively, using 2-fold, 5-fold and 10-fold Cross-Validation (CV) as the benchmark. The experimental results confirm the effectiveness of the NGCICM algorithm in predicting the interactions between circRNAs and miRNAs.
Zhufang Kuang, Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 Prediction of circRNA-MiRNA Association Using Singular Value Decomposition and Graph Neural Networks
abstract
A large number of experimental studies have shown that circRNAs can act as molecular sponges of microRNAs, interacting with miRNAs to regulate gene expression levels, thereby affecting the development of human diseases. Exploring the potential associations between circRNAs and miRNAs can help understand complex disease mechanisms. Considering that biological experiments are time-consuming and labor-intensive, this study proposes a computational model using a graph neural network and singular value decomposition (CMASG) for circRNA-miRNA association prediction. Specifically, graph neural networks are used to learn nonlinear feature representations of nodes, followed by matrix factorization algorithms to learn linear feature representations of nodes, and then combined feature representations learned from different perspectives. Finally, the lightGBM algorithm was used for circRNA-miRNA association prediction. The proposed CMASG model achieved an AUC value of 0.8804. The experimental results demonstrate the superiority and effectiveness of the CMASG model in predicting circRNA-miRNA association tasks.
Yurong Qian, Shaoqiu Li, Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 DeepFusionGO: Protein function prediction by fusing heterogeneous features through deep learning
abstract
Exploring the functions of proteins is crucial for explaining cellular mechanisms, treating diseases, and developing new drugs. Due to experimental limitations, large-scale identification of protein function remains a challenging task in cell biology. Here we propose DeepFusionGo, a novel protein function prediction method that adopts a graph representation learning approach (GraphSAGE) to extract features from heterogeneous data sources. First, we generate embeddings from protein sequences using the pre-trained protein language model and InterPro domains with scaling gradient. Then we integrate these two embeddings with adaptive feature weights to the PPI graph and use GraphSAGE to generate the representation vector. Finally, we build the classification model to predict protein function based on the concatenated feature vector. The experimental results show that DeepFusionGO outperforms existing state-of-the-art methods, including sequence-based DeepGOPLUS, and PPI-based DeepGraphGO. DeepFusionGO also performs well in difficult protein function prediction. We demonstrate that selecting an appropriate protein features fusion method can improve the prediction performance, and using the PPI network and the protein representation vector obtained from the protein language model through the GraphSAGE algorithm is an effective way to mine potential functional clues. The source code and data sets are available at: https://github.com/Hhhzj-7/DeepFusionGO.
Zhijian Huang 0001, Rongtao Zheng, Lei Deng 0002
BIBM3
2022 inACP: An integrated approach to the prediction of anticancer peptides
abstract
Cancer has become one of the deadliest diseases,which causes catastrophic pressure on healthcare systems and kills massive lives each year. Anticancer peptides (ACPs) have attracted great interest as a new direction for cancer treatment. Thus, developing in silico approaches for ACPs prediction is necessary and urgent. Here, we propose a new computational model called inACP, combining deep sequential representation learning features embedding and predicted protein structure features as input and integrating three machine learning classifiers with an ensemble approach. Comparison results show that inACP has a better capability to recognize ACPs than several state-of-the-art approaches. Datasets and source code are available at https://github.com/lnr3/inACP/.
Guolun Zhong, Lei Deng 0002
BIBM2
2022 Contrastive learning-based computational histopathology predict differential expression of cancer driver genes
abstract
MOTIVATION: Digital pathological analysis is run as the main examination used for cancer diagnosis. Recently, deep learning-driven feature extraction from pathology images is able to detect genetic variations and tumor environment, but few studies focus on differential gene expression in tumor cells. RESULTS: In this paper, we propose a self-supervised contrastive learning framework, HistCode, to infer differential gene expression from whole slide images (WSIs). We leveraged contrastive learning on large-scale unannotated WSIs to derive slide-level histopathological features in latent space, and then transfer it to tumor diagnosis and prediction of differentially expressed cancer driver genes. Our experiments showed that our method outperformed other state-of-the-art models in tumor diagnosis tasks, and also effectively predicted differential gene expression. Interestingly, we found the genes with higher fold change can be more precisely predicted. To intuitively illustrate the ability to extract informative features from pathological images, we spatially visualized the WSIs colored by the attention scores of image tiles. We found that the tumor and necrosis areas were highly consistent with the annotations of experienced pathologists. Moreover, the spatial heatmap generated by lymphocyte-specific gene expression patterns was also consistent with the manually labeled WSIs.
Gongming Zhou, Lei Deng 0002, Dachuan Zhang, Hui Liu 0026
Briefings Bioinform.4
2022 Attention-wise masked graph contrastive learning for predicting molecular property
abstract
MOTIVATION: Accurate and efficient prediction of the molecular property is one of the fundamental problems in drug research and development. Recent advancements in representation learning have been shown to greatly improve the performance of molecular property prediction. However, due to limited labeled data, supervised learning-based molecular representation algorithms can only search limited chemical space and suffer from poor generalizability. RESULTS: In this work, we proposed a self-supervised learning method, ATMOL, for molecular representation learning and properties prediction. We developed a novel molecular graph augmentation strategy, referred to as attention-wise graph masking, to generate challenging positive samples for contrastive learning. We adopted the graph attention network as the molecular graph encoder, and leveraged the learned attention weights as masking guidance to generate molecular augmentation graphs. By minimization of the contrastive loss between original graph and augmented graph, our model can capture important molecular structure and higher order semantic information. Extensive experiments showed that our attention-wise graph mask contrastive learning exhibited state-of-the-art performance in a couple of downstream molecular property prediction tasks. We also verified that our model pretrained on larger scale of unlabeled data improved the generalization of learned molecular representation. Moreover, visualization of the attention heatmaps showed meaningful patterns indicative of atoms and atomic groups important to specific molecular property.
Hui Liu 0026, Yibiao Huang, Lei Deng 0002
Briefings Bioinform.4
2022 TSNAPred: predicting type-specific nucleic acid binding residues via an ensemble approach
abstract
MOTIVATION: The interplay between protein and nucleic acid participates in diverse biological activities. Accurately identifying the interaction between protein and nucleic acid can strengthen the understanding of protein function. However, conventional methods are too time-consuming, and computational methods are type-agnostic predictions. We proposed an ensemble predictor termed TSNAPred and first used it to identify residues that bind to A-DNA, B-DNA, ssDNA, mRNA, tRNA and rRNA. TSNAPred combines LightGBM and capsule network, both learned on the feature derived from protein sequence. TSNAPred utilizes the sliding window technique to extract long-distance dependencies between residues and a weighted ensemble strategy to enhance the prediction performance. The results show that TSNAPred can effectively identify type-specific nucleic acid binding residues in our test set. What is more, it also can discriminate DNA-binding and RNA-binding residues, which has improved 5% to 10% on the AUC value compared with other state-of-the-art methods. The dataset and code of TSNAPred are available at: https://github.com/niewenjuan-csu/TSNAPred.
Wenjuan Nie, Lei Deng 0002
Briefings Bioinform.2
2022 DeepDDS: deep graph neural network with attention mechanism to predict synergistic drug combinations
abstract
MOTIVATION: Drug combination therapy has become an increasingly promising method in the treatment of cancer. However, the number of possible drug combinations is so huge that it is hard to screen synergistic drug combinations through wet-lab experiments. Therefore, computational screening has become an important way to prioritize drug combinations. Graph neural network has recently shown remarkable performance in the prediction of compound-protein interactions, but it has not been applied to the screening of drug combinations. RESULTS: In this paper, we proposed a deep learning model based on graph neural network and attention mechanism to identify drug combinations that can effectively inhibit the viability of specific cancer cells. The feature embeddings of drug molecule structure and gene expression profiles were taken as input to multilayer feedforward neural network to identify the synergistic drug combinations. We compared DeepDDS (Deep Learning for Drug-Drug Synergy prediction) with classical machine learning methods and other deep learning-based methods on benchmark data set, and the leave-one-out experimental results showed that DeepDDS achieved better performance than competitive methods. Also, on an independent test set released by well-known pharmaceutical enterprise AstraZeneca, DeepDDS was superior to competitive methods by more than 16% predictive precision. Furthermore, we explored the interpretability of the graph attention network and found the correlation matrix of atomic features revealed important chemical substructures of drugs. We believed that DeepDDS is an effective tool that prioritized synergistic drug combinations for further wet-lab experiment validation. AVAILABILITY AND IMPLEMENTATION: Source code and data are available at https://github.com/Sinwang404/DeepDDS/tree/master.
Jinxian Wang, Lei Deng 0002, Hui Liu 0026
Briefings Bioinform.4
2022 Computational anti-COVID-19 drug design: progress and challenges
abstract
Vaccines have made gratifying progress in preventing the 2019 coronavirus disease (COVID-19) pandemic. However, the emergence of variants, especially the latest delta variant, has brought considerable challenges to human health. Hence, the development of robust therapeutic approaches, such as anti-COVID-19 drug design, could aid in managing the pandemic more efficiently. Some drug design strategies have been successfully applied during the COVID-19 pandemic to create and validate related lead drugs. The computational drug design methods used for COVID-19 can be roughly divided into (i) structure-based approaches and (ii) artificial intelligence (AI)-based approaches. Structure-based approaches investigate different molecular fragments and functional groups through lead drugs and apply relevant tools to produce antiviral drugs. AI-based approaches usually use end-to-end learning to explore a larger biochemical space to design antiviral drugs. This review provides an overview of the two design strategies of anti-COVID-19 drugs, the advantages and disadvantages of these strategies and discussions of future developments.
Jinxian Wang, Wenjuan Nie, Lei Deng 0002
Briefings Bioinform.5
2022 Graph2MDA: a multi-modal variational graph embedding model for predicting microbe-drug associations
abstract
MOTIVATION: Accumulated clinical studies show that microbes living in humans interact closely with human hosts, and get involved in modulating drug efficacy and drug toxicity. Microbes have become novel targets for the development of antibacterial agents. Therefore, screening of microbe-drug associations can benefit greatly drug research and development. With the increase of microbial genomic and pharmacological datasets, we are greatly motivated to develop an effective computational method to identify new microbe-drug associations. RESULTS: In this article, we proposed a novel method, Graph2MDA, to predict microbe-drug associations by using variational graph autoencoder (VGAE). We constructed multi-modal attributed graphs based on multiple features of microbes and drugs, such as molecular structures, microbe genetic sequences and function annotations. Taking as input the multi-modal attribute graphs, VGAE was trained to learn the informative and interpretable latent representations of each node and the whole graph, and then a deep neural network classifier was used to predict microbe-drug associations. The hyperparameter analysis and model ablation studies showed the sensitivity and robustness of our model. We evaluated our method on three independent datasets and the experimental results showed that our proposed method outperformed six existing state-of-the-art methods. We also explored the meaning of the learned latent representations of drugs and found that the drugs show obvious clustering patterns that are significantly consistent with drug ATC classification. Moreover, we conducted case studies on two microbes and two drugs and found 75-95% predicted associations have been reported in PubMed literature. Our extensive performance evaluations validated the effectiveness of our proposed method. AVAILABILITY AND IMPLEMENTATION: Source codes and preprocessed data are available at https://github.com/moen-hyb/Graph2MDA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lei Deng 0002, Yibiao Huang, Hui Liu 0026
Bioinform.1
2022 MSPCD: predicting circRNA-disease associations via integrating multi-source data and hierarchical neural network
abstract
BACKGROUND: Increasing evidence shows that circRNA plays an essential regulatory role in diseases through interactions with disease-related miRNAs. Identifying circRNA-disease associations is of great significance to precise diagnosis and treatment of diseases. However, the traditional biological experiment is usually time-consuming and expensive. Hence, it is necessary to develop a computational framework to infer unknown associations between circRNA and disease. RESULTS: In this work, we propose an efficient framework called MSPCD to infer unknown circRNA-disease associations. To obtain circRNA similarity and disease similarity accurately, MSPCD first integrates more biological information such as circRNA-miRNA associations, circRNA-gene ontology associations, then extracts circRNA and disease high-order features by the neural network. Finally, MSPCD employs DNN to predict unknown circRNA-disease associations. CONCLUSIONS: Experiment results show that MSPCD achieves a significantly more accurate performance compared with previous state-of-the-art methods on the circFunBase dataset. The case study also demonstrates that MSPCD is a promising tool that can effectively infer unknown circRNA-disease associations.
Lei Deng 0002, Dayun Liu, Yizhan Li, Runqi Wang, Hui Liu 0026
BMC Bioinform.1
2022 Predicting circRNA-drug sensitivity associations via graph attention auto-encoder
abstract
BACKGROUND: Circular RNAs (circRNAs) play essential roles in cancer development and therapy resistance. Many studies have shown that circRNA is closely related to human health. The expression of circRNAs also affects the sensitivity of cells to drugs, thereby significantly affecting the efficacy of drugs. However, traditional biological experiments are time-consuming and expensive to validate drug-related circRNAs. Therefore, it is an important and urgent task to develop an effective computational method for predicting unknown circRNA-drug associations. RESULTS: In this work, we propose a computational framework (GATECDA) based on graph attention auto-encoder to predict circRNA-drug sensitivity associations. In GATECDA, we leverage multiple databases, containing the sequences of host genes of circRNAs, the structure of drugs, and circRNA-drug sensitivity associations. Based on the data, GATECDA employs Graph attention auto-encoder (GATE) to extract the low-dimensional representation of circRNA/drug, effectively retaining critical information in sparse high-dimensional features and realizing the effective fusion of nodes' neighborhood information. Experimental results indicate that GATECDA achieves an average AUC of 89.18% under 10-fold cross-validation. Case studies further show the excellent performance of GATECDA. CONCLUSIONS: Many experimental results and case studies show that our proposed GATECDA method can effectively predict the circRNA-drug sensitivity associations.
Lei Deng 0002, Yurong Qian, Jingpu Zhang
BMC Bioinform.1
2022 MSCNE: Predict miRNA-Disease Associations Using Neural Network Based on Multi-Source Biological Information
abstract
The important role of microRNA (miRNA) in human diseases has been confirmed by some studies. However, only using biological experiments has greater blindness, leading to higher experimental costs. In this paper a high-efficiency algorithm based on a variety of biological source information and applying a combination of a convolutional neural network (CNN) feature extractor and an extreme learning machine (ELM) classifier is proposed. Specifically, the semantic similarity of diseases, the gaussian interaction profile kernel similarity of the four biological information of miRNA, disease, long non-coding RNA (lncRNA) and environmental factors (EFs), and the similarities of miRNAs are fused together. Among them, miRNAs similarity is composed of miRNA target information, sequence information, family information, and function information. Then, the dimensionality of the data set is reduced by the autoencoder (AE). Finally, deep features are extracted through CNN, and then the association between miRNA and disease is predicted by ELM. The experimental results show that the average AUC value based on the multi-biological source information (MSCNE) model is 0.9630, which can reach higher performance than the other classic classifier, feature extractor mentioned and the other existing algorithms. The results show the MSCNE algorithm is effective to predict the correlation of miRNA-disease.
Genwei Han, Zhufang Kuang, Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 MGATMDA: Predicting Microbe-Disease Associations via Multi-Component Graph Attention Network
abstract
Microbes are parasitic in various human body organs and play significant roles in a wide range of diseases. Identifying microbe-disease associations is conducive to the identification of potential drug targets. Considering the high cost and risk of biological experiments, developing computational approaches to explore the relationship between microbes and diseases is an alternative choice. However, most existing methods are based on unreliable or noisy similarity, and the prediction accuracy could be affected. Besides, it is still a great challenge for most previous methods to make predictions for the large-scale dataset. In this work, we develop a multi-component Graph Attention Network (GAT) based framework, termed MGATMDA, for predicting microbe-disease associations. MGATMDA is built on a bipartite graph of microbes and diseases. It contains three essential parts: decomposer, combiner, and predictor. The decomposer first decomposes the edges in the bipartite graph to identify the latent components by node-level attention mechanism. The combiner then recombines these latent components automatically to obtain unified embedding for prediction by component-level attention mechanism. Finally, a fully connected network is used to predict unknown microbes-disease associations. Experimental results showed that our proposed method outperformed eight state-of-the-art methods. Case studies for two common diseases further demonstrated the effectiveness of MGATMDA in predicting potential microbe-disease associations. The codes are available at Github https://github.com/dayunliu/MGATMDA.
Dayun Liu, Qihua He, Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.5
2021 A multi-task graph convolutional network modeling of drug-drug interactions and synergistic efficacy
abstract
Identification of drug-drug interaction(DDI) is critical for safer and more effective drug co-prescription. As wetlab screening assays are time-consuming, labor-intensive and expensive, it is highly desired to develop an effective computational method to predict drug-drug interactions. In this work, we aim to predict of drug-drug interactions and synergistic drug combinations by proposing an end-to-end multi-task learning framework based on graph convolutional network (GCN). Precisely, we first convert the drug into a molecular graph, in which vertices represent atoms and edges represent chemical bonds. Next, the r-radius subgraph method is applied to molecular graph so that a series of subgraphs are produced for each drug. Next, the subgraphs are used as the input of the graph convolutional network to learn the embedding vector. Finally, the pairwise drug embeddings learned by GCN are concatenated as input into a fully-connected layer for predicting drug-drug interaction and synergistic effects. We conducted extensive performance evaluations on different data sets, including benchmark DDI data sets and manually collected drug combination data sets, and the results show that our proposed method is significantly better than the newly proposed methods (DeepCCI) and four typical machine learning methods (FFNN, SVM, RF, AdaBoost). In addition, our case study showed that 11 our of top 20 predicted DDIs have been reported by PubMed literature.
Yuanyuan Deng, Lei Deng 0002, Hui Liu 0026
BIBM3
2021 GCNSDA: Predicting snoRNA-disease associations via graph convolutional network
abstract
Small nucleolar RNAs(snoRNAs) represent an abundant group of noncoding RNAs in the nucleolus of eukaryotes. Recent studies revealed that snoRNAs play a significant role in a wide range of diseases. Identifying snoRNA-disease associations can provide great insights into understanding the disease and treatment and boosting drug discovery development. Traditional methods of using biological experiments are often usually small-scale and time-consuming. Therefore, it is urgent to develop a computational framework to predict snoRNA-disease associations. In this work, we proposed a novel Graph Convolutional Network(GCN) based framework GCNSDA for predicting snoRNA-disease associations. To our best knowledge, GCNSDA is the first framework that uses a graph convolutional network to predict snoRNA-disease associations. GCNSDA is built on the bipartite graph of snoRNAs and diseases; it uses a graph neural network to discover latent factors that cause the association between snoRNAs and diseases and then generate embedded representations of snoRNAs and diseases. Experimental results showed that GCNSDA achieved better performance than the other six state-of-the-art methods. Case study further confirmed the effectiveness of GCNSDA in predicting snoRNA-disease associations.
Dayun Liu, Hanlin Xu, Lei Deng 0002
BIBM6
2021 iPiDA-GBNN: Identification of Piwi-interacting RNA-disease associations based on gradient boosting neural network
abstract
Piwi-interacting RNAs (piRNAs) are a novel class of small non-coding RNAs that interact with the PIWI protein family and are associated with various diseases. Identification of piRNA-disease associations provides candidate piRNA targets for disease treatment and provides promising candidate molecular targets to promote the drug design. However, detection of piRNA-disease associations by biological experiments is often high cost and time-consuming. Therefore, there is an urgent need for reliable computational methods for piRNA-disease associations identification. In this study, we proposed a new computational predictor named iPiDA-GBNN to predict potential piRNA-disease associations. The iPiDA-GBNN presented the piRNA-disease pairs by combining disease similarity information and piRNA similarity information. Disease similarity information includes disease GIP kernel similarity, disease Jaccard similarity, disease semantic similarity, and piRNA similarity includes sequence similarity and GIP kernel similarity. The Stacked auto-encoder(SAE) was then performed on piRNA features to extract the key features. The training datasets consisted of experiments that confirmed positive associations and the same quantity negative associations from the unknown pairs. Finally, the Gradient Boosting Neural Networks(GrownNet) [20] trained with the known piRNA-disease associations and the negative associations to predict new piRNA-disease associations. The experimental results showed that iPiDA-GBNN achieved superior predictive ability compared with other state-of-art predictors. The source code and datasets explored in this work are available at https://github.com/Tracyhuahua666/iPiDA-GBNN.
Yurong Qian, Qihua He, Lei Deng 0002
BIBM3
2021 CMIVGSD: circRNA-miRNA Interaction Prediction Based on Variational Graph Auto-Encoder and Singular Value Decomposition
abstract
A large amount of evidence shows that circular RNAs(circRNAs) participate in transcription and translation regulation and function as “micro RNA(miRNA)-sponges”. Recognizing circRNA-miRNA interaction is helpful to understand the function of circRNAs, especially its role in complex diseases. Obtaining interactive information based on traditional biological experiments is usually small-scale, time-consuming, and laborious. Considering that there are few calculation methods, it is urgent to develop efficient and accurate methods to extract the interaction between circRNA and miRNA. In this work, we proposed a computational framework called CMIVGSD, which uses singular value decomposition and graph variational auto-encoders to predict circRNA-miRNA interaction. To our best knowledge, CMIVGSD is the first calculation framework to predict circRNA-miRNA interaction. CMIVGSD uses the singular value decomposition (SVD) algorithm to obtain linear features from the circRNA-miRNA interaction matrix. We have constructed the similarity networks of circRNA and miRNA, respectively. The graph variational auto-encoder (VGAE) is employed to mine the non-linear features of circRNA-miRNA in similarity networks. Finally, we combine linear and non-linear features and use LightGBM to predict interaction scores. We performed five-fold cross-validation experiments. Experimental results show that our proposed method is better than other methods. The case study further proves the effectiveness of CMIVGSD in predicting circRNA-miRNA interaction.
Yurong Qian, Lei Deng 0002
BIBM6
2021 Accurately Predicting circRNA-disease Associations Using Variational Graph Auto-encoders and LightGBM
abstract
Many studies have shown that circRNAs play essential roles in various biological processes. With the development of technology, the associations between circRNA and diseases have been discovered, and these associations will help diagnose and treat diseases. However, it is time-consuming and costly to detect the associations between circRNAs and diseases with the experimental methods. Therefore, it is necessary to develop a feasible and effective computational method for predicting circRNA-disease associations. In this paper, we propose a new computational framework called VLCDA to identify the potential circRNA-disease associations. Initially, we construct features by fusing circRNA expression profile features and circRNA protein-coding ability features, disease semantic features, circRNA and disease GIP Kernel features, and use VGAE to mine its deep latent features. Finally, we use the fusion features to train the LightGBM classifier and the trained LightGBM to identify the circRNA-disease associations. The main contribution of VLCDA is that we firstly add circRNA protein-coding ability feature to the circRNA-disease association prediction model. In addition, VLCDA uses variational graph auto-encoders to extract the latent features of circRNA-disease associations to improve the prediction model’s accuracy further. VLCDA obtained the area under the ROC curve (AUC) scores of 0.9783 in 5-fold cross-validation. In addition, in the case studies, 16 of the top 20 circRNA-disease associations predicted by VLCDA have been confirmed by relevant literature.
Yurong Qian, Lei Deng 0002
BIBM5
2021 LGCMDS: Predicting miRNA-Drug Sensitivity based on Light Graph Convolution Network
abstract
The research of anticancer drugs has gone through a long process of development, but so far, no drug can cure cancer completely. The drug resistance is one of the main reasons for the failure of cancer treatment. As the relationship between miRNA and cancer is gradually revealed, more and more evidence shows that the sensitivity of cancer cells to anticancer drugs is also affected by miRNA. Research on miRNA-drug sensitivity associations can overcome the challenging clinical situation imposed by drug resistance. However, traditional biological experiments are time-consuming and expensive. Therefore, there is an urgent need to develop a computational method to predict the associations between miRNA and drug sensitivity accurately and efficiently. In this work, we propose a computational method based on simplified GCN to predict the miRNA-drug sensitivity associations, named LGCMDS. We abandon the two common designs in standard GCN-feature transformation and nonlinear activation, only retain the essential component, neighbourhood aggregation, and combine with high-order connectivity in the miRNA-drug graph to effectively integrate the miRNA-drug interactions into the embedding process. The 5-fold cross-validation results show that our proposed method achieves AUC of 0.8872 and AUPR of 0.9026. In comparison with five state-of-the-art models, LGCMDS achieves the best results. In addition, case study for Cisplatin, further proves the effectiveness of LGCMDS in predicting potential miRNA-drug sensitivity associations.
Hanlin Xu, Yizhan Li, Dayun Liu, Lei Deng 0002
BIBM5
2021 SMALF: miRNA-disease associations prediction based on stacked autoencoder and XGBoost
abstract
BACKGROUND: Identifying miRNA and disease associations helps us understand disease mechanisms of action from the molecular level. However, it is usually blind, time-consuming, and small-scale based on biological experiments. Hence, developing computational methods to predict unknown miRNA and disease associations is becoming increasingly important. RESULTS: In this work, we develop a computational framework called SMALF to predict unknown miRNA-disease associations. SMALF first utilizes a stacked autoencoder to learn miRNA latent feature and disease latent feature from the original miRNA-disease association matrix. Then, SMALF obtains the feature vector of representing miRNA-disease by integrating miRNA functional similarity, miRNA latent feature, disease semantic similarity, and disease latent feature. Finally, XGBoost is utilized to predict unknown miRNA-disease associations. We implement cross-validation experiments. Compared with other state-of-the-art methods, SAMLF achieved the best AUC value. We also construct three case studies, including hepatocellular carcinoma, colon cancer, and breast cancer. The results show that 10, 10, and 9 out of the top ten predicted miRNAs are verified in MNDR v3.0 or miRCancer, respectively. CONCLUSION: The comprehensive experimental results demonstrate that SMALF is effective in identifying unknown miRNA-disease associations.
Dayun Liu, Yibiao Huang, Wenjuan Nie, Lei Deng 0002
BMC Bioinform.5
2021 CRPGCN: predicting circRNA-disease associations using graph convolutional network based on heterogeneous network
abstract
BACKGROUND: The existing studies show that circRNAs can be used as a biomarker of diseases and play a prominent role in the treatment and diagnosis of diseases. However, the relationships between the vast majority of circRNAs and diseases are still unclear, and more experiments are needed to study the mechanism of circRNAs. Nowadays, some scholars use the attributes between circRNAs and diseases to study and predict their associations. Nonetheless, most of the existing experimental methods use less information about the attributes of circRNAs, which has a certain impact on the accuracy of the final prediction results. On the other hand, some scholars also apply experimental methods to predict the associations between circRNAs and diseases. But such methods are usually expensive and time-consuming. Based on the above shortcomings, follow-up research is needed to propose a more efficient calculation-based method to predict the associations between circRNAs and diseases. RESULTS: In this study, a novel algorithm (method) is proposed, which is based on the Graph Convolutional Network (GCN) constructed with Random Walk with Restart (RWR) and Principal Component Analysis (PCA) to predict the associations between circRNAs and diseases (CRPGCN). In the construction of CRPGCN, the RWR algorithm is used to improve the similarity associations of the computed nodes with their neighbours. After that, the PCA method is used to dimensionality reduction and extract features, it makes the connection between circRNAs with higher similarity and diseases closer. Finally, The GCN algorithm is used to learn the features between circRNAs and diseases and calculate the final similarity scores, and the learning datas are constructed from the adjacency matrix, similarity matrix and feature matrix as a heterogeneous adjacency matrix and a heterogeneous feature matrix. CONCLUSIONS: After 2-fold cross-validation, 5-fold cross-validation and 10-fold cross-validation, the area under the ROC curve of the CRPGCN is 0.9490, 0.9720 and 0.9722, respectively. The CRPGCN method has a valuable effect in predict the associations between circRNAs and diseases.
Zhufang Kuang, Lei Deng 0002
BMC Bioinform.3
2021 MSCFS: inferring circRNA functional similarity based on multiple data sources
abstract
BACKGROUND: More and more evidence shows that circRNA plays an important role in various biological processes and human health. Therefore, inferring the circRNA's potential functions and obtaining circRNA functional similarity has become more and more significant. However, there is no effective approach to explore the functional similarity of circRNAs. METHODS: In this paper, we propose a new approach, called MSCFS, to calculate the functional similarity of circRNA by integrating multiple data sources. We combine circRNA-disease association, circRNA-gene-Gene Ontology association, and circRNA sequence information to explore the functional similarity of circRNA. Firstly, we employ different learning representation methods from three data sources to establish three circRNA functional similarity networks. Then we integrate the three networks to obtain the final circRNA functional similarity. RESULTS: We utilize circRNA-miRNA association similarity and circRNA co-expression similarity to evaluate the performance of MSCFS. The results show a positive correlation with miRNA association ([Formula: see text]) and circRNA co-expression similarity ([Formula: see text]). Finally, we construct a circRNA functional similarity network and perform case analysis. The result shows our method can be applied to infer new potential functions of circRNA and other associations. CONCLUSIONS: MSCFS combines multiple data sources related to circRNA functions. Correlation analysis and case analyses prove that MSCFS is a useful method to explore circRNA functional similarity.
Liang Shu, Xinxu Yuan, Jingpu Zhang, Lei Deng 0002
BMC Bioinform.5
2021 LDAH2V: Exploring Meta-Paths Across Multiple Networks for lncRNA-Disease Association Prediction
abstract
Accumulating evidence has demonstrated dysfunctions of long non-coding RNAs (lncRNAs) are involved in various complex human diseases. However, even today, the relationships between lncRNAs and diseases remain unknown in most cases. Developing effective computational approaches to identify potential lncRNA-disease associations has become a hot topic. Existing network-based approaches are usually focused on the intrinsic features of lncRNAs and diseases but ignore the heterogeneous information of biological networks. Considering the limitations in previous methods, we propose LDAH2V, an efficient computational framework for predicting potential lncRNA-disease associations. LDAH2V uses the HIN2Vec to calculate the meta-path and feature vector for each lncRNA-disease pair in the heterogeneous information network (HIN), which consists of lncRNA similarity network, disease similarity network, miRNA similarity network, and the associations between them. Then, a Gradient Boosting Tree (GBT) classifier to predict lncRNA-disease associations is built with the feature vectors. The results show that LDAH2V performs significantly better than the four existing state-of-the-art methods and gains an AUC of 0.97 in the 10-fold cross-validation test. Furthermore, case studies of colon cancer and ovarian cancer-related lncRNAs have been confirmed in related databases and medical literature.
Lei Deng 0002, Jingpu Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 Predicting circRNA-disease associations using meta path-based representation learning on heterogenous network
abstract
Circular RNA (circRNA) is a new class of regulatory non-coding RNAs modulating gene expression by acting as a microRNA (miRNA) sponge, RNA binding protein sponge and translational regulator. A increasing number of experimental studies have shown that circRNA plays an important role in the development of diseases, and circRNA biomarkers are helpful for the diagnosis and treatment of various human diseases. There is a pressing demand to establish an effective computational method to identify the associations between circRNAs and diseases. In this paper, we propose a new computational framework for the prediction of the circRNA-disease associations. In particular, we calculated meta path-based feature vectors for each circRNAdisease pair on a heterogeneous information network (HIN) that integrated multiple subnetworks, including circRNA similarity network, disease similarity network, protein similarity network, circRNA-disease associations, circRNA-protein associations and protein-disease associations. A positive-unlabeled learning algorithm was adopted to generate negative samples, and a random forest classifier was trained to predicted circRNA-disease associations. We conducted performance comparison with three popular methods on the CircR2Disease dataset. The experimental results show that our method outperform other existing methods by achieving AUC 0.983 on 5-fold cross-validation.
Lei Deng 0002, Hui Liu 0026
BIBM1
2020 DeepARC: An Attention-based Hybrid Model for Predicting Transcription Factor Binding Sites from Positional Embedded DNA Sequence
abstract
The binding of transcription factors (TFs) to transcription factor binding sites (TFBS) plays a pivotal role in regulating gene expression and evolution. Accurately modeling the specificity of DNA and searching for TFBS helps understand the genome's function and evolution. In recent years, computational identification of TFBS has become an active field of research. Here, we propose DeepARC, an attention-based hybrid approach combining convolutional neural network (CNN) and recurrent neural network (RNN) for predicting TFBS. We employ a position-based embedding strategy to embed a DNA sequence into a matrix with distributed representation contenting the position information and then feed the distributed representations of the sequence into a CNN-BiLSTM-Attention-based framework to classify whether there is a TFBS in a sequence. Take the advantage of the attention mechanism, DeepARC can obtain more valuable information about TFBS and add interpretability to the TFBS search process. Moreover, sufficient experiments prove that DeepARC has better performance than existing predictors. The DeepARC web server is available at http://deeparc.denglab.org.
Lei Deng 0002
BIBM2
2020 Predict the Protein-protein Interaction between Virus and Host through Hybrid Deep Neural Network
abstract
Viral infection has been considered as a threat to human health for many years, where protein-protein interactions (PPIs) between viruses and hosts is involved. Researching the PPI between the virus and the host is conducive to understanding the mechanism of virus infection and the development of new drugs. Currently, most of the existing studies based on sequence only focus on extracting sequence features from original amino acid sequences, whereas the redundancy and noise of the features are neglected.In this paper, we employed Ll-regularized logistic regression to obtain efficacious sequence features related to PPIs without losing accuracy and generalization. A hybrid deep learning framework which combines convolutional neural network together with a long short term memory network to extract more hidden high-level features was designed to extract more latent features. As it is demonstrated in experiments results, the proposed framework is superior to the current advanced framework in both benchmark data and independent testing and is promising for identifying virus-host interactions.
Lei Deng 0002, Jiaojiao Zhao, Jingpu Zhang
BIBM1
2020 Machine learning-based methods and novel data models to predict adverse drug reaction
abstract
Predicting adverse drug reactions (ADRs) plays a critical role in developing new drugs and preventing adverse reactions during the treatment of existing drugs. However, with the rapid progress of machine learning technology, a new situation has been opened up in ADRs prediction. Using appropriate machine learning methods with existing data can achieve high prediction performance, attracting more researchers. This review describes commonly used features (biological, chemical, and phenotypic features) and machine learning algorithms.
Jinxian Wang, Yuanyuan Deng, Liang Shu, Lei Deng 0002
BIBM4
2020 DeepciRGO: functional prediction of circular RNAs through hierarchical deep neural networks using heterogeneous network features
abstract
BACKGROUND: Circular RNAs (circRNAs) are special noncoding RNA molecules with closed loop structures. Compared with the traditional linear RNA, circRNA is more stable and not easily degraded. Many studies have shown that circRNAs are involved in the regulation of various diseases and cancers. Determining the functions of circRNAs in mammalian cells is of great significance for revealing their mechanism of action in physiological and pathological processes, diagnosis and treatment of diseases. However, determining the functions of circRNAs on a large scale is a challenging task because of the high experimental costs. RESULTS: In this paper, we present a hierarchical deep learning model, DeepciRGO, which can effectively predict gene ontology functions of circRNAs. We build a heterogeneous network containing circRNA co-expressions, protein-protein interactions and protein-circRNA interactions. The topology features of proteins and circRNAs are calculated using a novel representation learning approach HIN2Vec across the heterogeneous network. Then, a deep multi-label hierarchical classification model is trained with the topology features to predict the biological process function in the gene ontology for each circRNA. In particular, we manually curated a benchmark dataset containing 185 GO annotations for 62 circRNAs, namely, circRNA2GO-62. The DeepciRGO achieves promising performance on the circRNA2GO-62 dataset with a maximum F-measure of 0.412, a recall score of 0.400, and an accuracy of 0.425, which are significantly better than other state-of-the-art RNA function prediction methods. In addition, we demonstrate the considerable potential of integrating multiple interactions and association networks. CONCLUSIONS: DeepciRGO will be a useful tool for accurately annotating circRNAs. The experimental results show that integrating multi-source data can help to improve the predictive performance of DeepciRGO. Moreover, The model also can combine RNA structure and sequence information to further optimize predictive performance.
Lei Deng 0002, Jingpu Zhang
BMC Bioinform.1
2019 A deep neural network approach using distributed representations of RNA sequence and structure for identifying binding site of RNA-binding proteins
abstract
RNA-binding proteins (RBPs) play a crucial role in the post-transcriptional regulation of RNAs. Identification of RBP binding sites is a key step to understand the biological mechanism of post-transcriptional regulation. Although many computational methods have been developed for predicting RNA-protein binding sites, few study considers the k-mer embedding representation of RNA primary sequence and secondary structure specificities. In this paper, we develop a general deep learning framework, named deepRKE, to predict RNA-protein binding sites. deepRKE takes an unsupervised shallow two-layer neural network to automatically learn the distributed representation of k-mers by taking their neighbor context into account. Compared to conventional k-mers approach, distributed representations effectively detect the latent relationship and similarity between k-mers. The distributed representations of the sequences and secondary structures are fed into CNN convolutional neural network (CNN) and a bidirectional long short term memory network (BLSTM) to discriminate the RBP binding sites from unbound sites. We comprehensively evaluate deepRKE on two large-scale RBP binding sites datasets, and the experimental results show that deepRKE achieves better performance than five competitive methods.
Lei Deng 0002, Youzhi Liu, Yechuan Shi, Hui Liu 0026
BIBM1
2019 DKCirc2GO: Predicting Gene Ontology of circRNAs Using Dual KATZ Approach
abstract
Circular RNAs(circRNAs) have been demonstrated to play significant biological roles in many human biological processes such as competing endogenous RNAs or miRNA sponges, regulating gene transcription, translating proteins and others. Inferring the functions of circRNAs is an important strategy for understanding disease pathogenesis at the molecular level. However, most circRNAs have not been functionally characterized. Developing effective computational approaches to identify potential functions of circRNA has become a hot topic. In this paper, we propose an integrated model, DKCirc2GO, to infer the gene ontology (GO) functions of circRNAs by integrating multiple data sources, including the expression profiles of circRNAs, circRNA-protein associations, protein-protein interactions (PPI), protein-GO associations, GO-GO semantic similarity. The C2P global network is constructed by integrating three heterogeneous networks: circRNA-circRNA similarity network, circRNA-protein association network and protein-protein interaction network. The P2G global network is constructed by integrating three heterogeneous networks: protein-protein interaction network, protein-GO association network and GO-GO semantic similarity network. The KATZ measure is then employed respectively to calculate similarities in the two global network. The circRNAs-protein-GO(cpg) association network is constructed based on the KATZ similarity scores. Finally, we annotate circRNAs with Gene Ontology (GO) terms of their neighboring protein-GO terms. The experimental results show that DKCirc2GO has a significantly better performance than existing state-of-the-art methods in terms of F-max.
Lei Deng 0002, Jingpu Zhang
BIBM1
2019 D2VCB: A Hybrid Deep Neural Network for the Prediction of in-vivo Protein-DNA Binding from Combined DNA Sequence
abstract
Prediction of in-vivo protein-DNA binding is an important, but challenging task in the broad field of computational biology. Although some methods based on deep learning have succeed in modeling in-vivo protein-DNA binding, they often simply extract the sequence features from the original DNA sequence without consideration of other sequence features, such as their reverse, complementary and reverse complementary sequences. Also, one-hot encoding of DNA sequence is vulnerable to the curse of dimensionality, which leads to unwanted equidistance of pairwise sequences. To address these problems, we propose D2VCB (dna2vec, convolution, bi-LSTM), a novel hybrid deep neural network framework using dna2vec to predict in-vivo protein-DNA binding events. We extract input features from DNA original sequences, reverse sequences, complementary and complementary reverse sequences, and then use dna2vec to compute a distributed representation of k-mer. In our D2VCB model, the convolution layer captures motif features, while the recurrent layer captures long-term dependencies among motif features so as to improve prediction accuracy. Our performance comparison experiments show that D2VCB outperforms significantly other existing methods in terms of multiple performance metrics.
Lei Deng 0002, Hui Liu 0026
BIBM1
2019 MADOKA: an ultra-fast approach for large-scale protein structure similarity searching
abstract
BACKGROUND: Protein comparative analysis and similarity searches play essential roles in structural bioinformatics. A couple of algorithms for protein structure alignments have been developed in recent years. However, facing the rapid growth of protein structure data, improving overall comparison performance and running efficiency with massive sequences is still challenging. RESULTS: Here, we propose MADOKA, an ultra-fast approach for massive structural neighbor searching using a novel two-phase algorithm. Initially, we apply a fast alignment between pairwise structures. Then, we employ a score to select pairs with more similarity to carry out a more accurate fragment-based residue-level alignment. MADOKA performs about 6-100 times faster than existing methods, including TM-align and SAL, in massive alignments. Moreover, the quality of structural alignment of MADOKA is better than the existing algorithms in terms of TM-score and number of aligned residues. We also develop a web server to search structural neighbors in PDB database (About 360,000 protein chains in total), as well as additional features such as 3D structure alignment visualization. The MADOKA web server is freely available at: http://madoka.denglab.org/ CONCLUSIONS: MADOKA is an efficient approach to search for protein structure similarity. In addition, we provide a parallel implementation of MADOKA which exploits massive power of multi-core CPUs.
Lei Deng 0002, Guolun Zhong, Chenzhe Liu, Judong Luo, Hui Liu 0026
BMC Bioinform.1
2019 Integrating Multiple Heterogeneous Networks for Novel LncRNA-Disease Association Inference
abstract
Accumulating experimental evidence has indicated that long non-coding RNAs (lncRNAs) are critical for the regulation of cellular biological processes implicated in many human diseases. However, only relatively few experimentally supported lncRNA-disease associations have been reported. Developing effective computational methods to infer lncRNA-disease associations is becoming increasingly important. Current network-based algorithms typically use a network representation to identify novel associations between lncRNAs and diseases. But these methods are concentrated on specific entities of interest (lncRNAs and diseases) and they do not allow to consider networks with more than two types of entities. Considering the limitations in previous computational methods, we develop a new global network-based framework, LncRDNetFlow, to prioritize disease-related lncRNAs. LncRDNetFlow utilizes a flow propagation algorithm to integrate multiple networks based on a variety of biological information including lncRNA similarity, protein-protein interactions, disease similarity, and the associations between them to infer lncRNA-disease associations. We show that LncRDNetFlow performs significantly better than the existing state-of-the-art approaches in cross-validation. To further validate the reproducibility of the performance, we use the proposed method to identify the related lncRNAs for ovarian cancer, glioma, and cervical cancer. The results are encouraging. Many predicted lncRNAs in the top list have been verified by the biological studies.
Jingpu Zhang, Zuping Zhang 0001, Zhigang Chen 0001, Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.4
2019 KATZLGO: Large-Scale Prediction of LncRNA Functions by Using the KATZ Measure Based on Multiple Networks
abstract
Aggregating evidences have shown that long non-coding RNAs (lncRNAs) generally play key roles in cellular biological processes such as epigenetic regulation, gene expression regulation at transcriptional and post-transcriptional levels, cell differentiation, and others. However, most lncRNAs have not been functionally characterized. There is an urgent need to develop computational approaches for function annotation of increasing available lncRNAs. In this article, we propose a global network-based method, KATZLGO, to predict the functions of human lncRNAs at large scale. A global network is constructed by integrating three heterogeneous networks: lncRNA-lncRNA similarity network, lncRNA-protein association network, and protein-protein interaction network. The KATZ measure is then employed to calculate similarities between lncRNAs and proteins in the global network. We annotate lncRNAs with Gene Ontology (GO) terms of their neighboring protein-coding genes based on the KATZ similarity scores. The performance of KATZLGO is evaluated on a manually annotated lncRNA benchmark and a protein-coding gene benchmark with known function annotations. KATZLGO significantly outperforms state-of-the-art computational method both in maximum F-measure and coverage. Furthermore, we apply KATZLGO to predict functions of human lncRNAs and successfully map 12,318 human lncRNA genes to GO terms.
Zuping Zhang 0001, Jingpu Zhang, Yongjun Tang, Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.5
2018 Exploring Disease Similarity by Integrating Multiple Data Sources
Lei Deng 0002, Danyi Ye, Junmin Zhao, Jingpu Zhang
BIBM1
2018 XPredRBR: Accurate and Fast Prediction of RNA-Binding Residues in Proteins Using eXtreme Gradient Boosting
Lei Deng 0002, Zuojin Dong, Hui Liu 0026
ISBRA1
2018 Computational identification of binding energy hot spots in protein-RNA complexes using an ensemble approach
abstract
Motivation: Identifying RNA-binding residues, especially energetically favored hot spots, can provide valuable clues for understanding the mechanisms and functional importance of protein-RNA interactions. Yet, limited availability of experimentally recognized energy hot spots in protein-RNA crystal structures leads to the difficulties in developing empirical identification approaches. Computational prediction of RNA-binding hot spot residues is still in its infant stage. Results: Here, we describe a computational method, PrabHot (Prediction of protein-RNA binding hot spots), that can effectively detect hot spot residues on protein-RNA binding interfaces using an ensemble of conceptually different machine learning classifiers. Residue interaction network features and new solvent exposure characteristics are combined together and selected for classification with the Boruta algorithm. In particular, two new reference datasets (benchmark and independent) have been generated containing 107 hot spots from 47 known protein-RNA complex structures. In 10-fold cross-validation on the training dataset, PrabHot achieves promising performances with an AUC score of 0.86 and a sensitivity of 0.78, which are significantly better than that of the pioneer RNA-binding hot spot prediction method HotSPRing. We also demonstrate the capability of our proposed method on the independent test dataset and gain a competitive advantage as a result. Availability and implementation: The PrabHot webserver is freely available at http://denglab.org/PrabHot/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Yuliang Pan, Weihua Zhan, Lei Deng 0002
Bioinform.4
2018 Ontological function annotation of long non-coding RNAs through hierarchical multi-label classification
abstract
Motivation: Long non-coding RNAs (lncRNAs) are an enormous collection of functional non-coding RNAs. Over the past decades, a large number of novel lncRNA genes have been identified. However, most of the lncRNAs remain function uncharacterized at present. Computational approaches provide a new insight to understand the potential functional implications of lncRNAs. Results: Considering that each lncRNA may have multiple functions and a function may be further specialized into sub-functions, here we describe NeuraNetL2GO, a computational ontological function prediction approach for lncRNAs using hierarchical multi-label classification strategy based on multiple neural networks. The neural networks are incrementally trained level by level, each performing the prediction of gene ontology (GO) terms belonging to a given level. In NeuraNetL2GO, we use topological features of the lncRNA similarity network as the input of the neural networks and employ the output results to annotate the lncRNAs. We show that NeuraNetL2GO achieves the best performance and the overall advantage in maximum F-measure and coverage on the manually annotated lncRNA2GO-55 dataset compared to other state-of-the-art methods. Availability and implementation: The source code and data are available at http://denglab.org/NeuraNetL2GO/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Jingpu Zhang, Zuping Zhang 0001, Lei Deng 0002
Bioinform.5
2018 PDRLGB: precise DNA-binding residue prediction using a light gradient boosting machine
abstract
BACKGROUND: Identifying specific residues for protein-DNA interactions are of considerable importance to better recognize the binding mechanism of protein-DNA complexes. Despite the fact that many computational DNA-binding residue prediction approaches have been developed, there is still significant room for improvement concerning overall performance and availability. RESULTS: Here, we present an efficient approach termed PDRLGB that uses a light gradient boosting machine (LightGBM) to predict binding residues in protein-DNA complexes. Initially, we extract a wide variety of 913 sequence and structure features with a sliding window of 11. Then, we apply the random forest algorithm to sort the features in descending order of importance and obtain the optimal subset of features using incremental feature selection. Based on the selected feature set, we use a light gradient boosting machine to build the prediction model for DNA-binding residues. Our PDRLGB method shows better overall predictive accuracy and relatively less training time than other widely used machine learning (ML) methods such as random forest (RF), Adaboost and support vector machine (SVM). We further compare PDRLGB with various existing approaches on the independent test datasets and show improvement in results over the existing state-of-the-art approaches. CONCLUSIONS: PDRLGB is an efficient approach to predict specific residues for protein-DNA interactions.
Lei Deng 0002, Juan Pan, Wenyi Yang, Chuyao Liu, Hui Liu 0026
BMC Bioinform.1
2017 Combining diffusion and HeteSim features for accurate prediction of protein-lncRNA interactions
abstract
Identifying the interactions between proteins and Long non-coding RNAs (lncRNAs) can provide valuable clues for understanding the mechanisms and physiological functions of lncRNAs. In this work, we propose a computational method, PLIPCOM, which can accurately detect protein-lncRNA interactions by integrating two groups of network features. Low dimensional diffusion characteristics and HeteSim features are combined to build the protein-lncRNA interaction predictor using the Gradient Tree Boosting (GTB) algorithm. In a cross-validation experiment on the benchmark data set, our PLIPCOM method substantially outperformed previous state-of-the-art approaches in predicting the interactions between proteins and lncRNAs.
Junqiang Wang, Weihua Zhan, Lei Deng 0002
BIBM5
2017 BiRWLGO: A global network-based strategy for lncRNA function annotation using bi-random walk
abstract
A large number of long non-coding RNAs (lncRNAs) have been identified over the past decades. Accumulating evidence proves that lncRNAs play key roles in various biological processes. However, the majority of the lncRNAs have not been functionally characterized. The annotation of lncRNA functions has become an area of focus in the fields of biology and bioinformatics. In this paper, we develop a global network-based strategy, BiRWLGO, to predict probable functions for lncRNAs at large scale. In BiRWLGO, we first build a global network consisting of three networks: lncRNA-lncRNA similarity network, lncRNA-protein interaction network and protein-protein interaction network. Then the bi-random walk algorithm is applied to explore similarities between lncRNAs and proteins. The functions of a query lncRNA can be obtained according to the Gene Ontology (GO) terms of its neighboring proteins. We compare the performance of BiRWLGO with other state-of-the-art approaches on a manually annotated lncRNA benchmark with known GO terms. As a result, BiRWLGO achieves the best predictive performance in terms of both maximum F-measure (Fmax) and coverage. Moreover, we demonstrate that integrating the protein-protein interactions can help improve the predictive performance of lncRNA functions.
Jingpu Zhang, Shuai Zou, Lei Deng 0002
BIBM3
2017 A sparse autoencoder-based deep neural network for protein solvent accessibility and contact number prediction
abstract
BACKGROUND: Direct prediction of the three-dimensional (3D) structures of proteins from one-dimensional (1D) sequences is a challenging problem. Significant structural characteristics such as solvent accessibility and contact number are essential for deriving restrains in modeling protein folding and protein 3D structure. Thus, accurately predicting these features is a critical step for 3D protein structure building. RESULTS: In this study, we present DeepSacon, a computational method that can effectively predict protein solvent accessibility and contact number by using a deep neural network, which is built based on stacked autoencoder and a dropout method. The results demonstrate that our proposed DeepSacon achieves a significant improvement in the prediction quality compared with the state-of-the-art methods. We obtain 0.70 three-state accuracy for solvent accessibility, 0.33 15-state accuracy and 0.74 Pearson Correlation Coefficient (PCC) for the contact number on the 5729 monomeric soluble globular protein dataset. We also evaluate the performance on the CASP11 benchmark dataset, DeepSacon achieves 0.68 three-state accuracy and 0.69 PCC for solvent accessibility and contact number, respectively. CONCLUSIONS: We have shown that DeepSacon can reliably predict solvent accessibility and contact number with stacked sparse autoencoder and a dropout approach.
Lei Deng 0002
BMC Bioinform.1
2017 A boosting approach for prediction of protein-RNA binding residues
abstract
BACKGROUND: RNA binding proteins play important roles in post-transcriptional RNA processing and transcriptional regulation. Distinguishing the RNA-binding residues in proteins is crucial for understanding how protein and RNA recognize each other and function together as a complex. RESULTS: We propose PredRBR, an effectively computational approach to predict RNA-binding residues. PredRBR is built with gradient tree boosting and an optimal feature set selected from a large number of sequence and structure characteristics and two categories of structural neighborhood properties. In cross-validation experiments on the RBP170 data set show that PredRBR achieves an overall accuracy of 0.84, a sensitivity of 0.85, MCC of 0.55 and AUC of 0.92, which are significantly better than that of other widely used machine learning algorithms such as Support Vector Machine, Random Forest, and Adaboost. We further calculate the feature importance of different feature categories and find that structural neighborhood characteristics are critical in the recognization of RNA binding residues. Also, PredRBR yields significantly better prediction accuracy on an independent test set (RBP101) in comparison with other state-of-the-art methods. CONCLUSIONS: The superior performance over existing RNA-binding residue prediction methods indicates the importance of the gradient tree boosting algorithm combined with the optimal selected features.
Yongjun Tang, Diwei Liu, Ting Wen, Lei Deng 0002
BMC Bioinform.5
2016 PredRBR: Accurate Prediction of RNA-Binding Residues in proteins using Gradient Tree Boosting
abstract
Prediction of Protein-RNA binding sites is one of the most challenging and intriguing problems in the field of computational biology. Here, we proposed an effectively machine learning algorithm termed PredRBR (Prediction of RNA Binding Residues), using Gradient Tree Boosting algorithm and mRMR-IFS feature selection method in combination with sequence features, structure characteristics and two categories of structural neighborhood feature for prediction of RNA binding sites in proteins. We evaluate PredRBR on the independent test dataset (RBP101), and obtain significant improvement on the prediction performance compared with other state-of-the-art approaches. In addition, we test the variable importance of diverse feature types. The results show that structural neighborhood features play a crucial role in the identification of RNA binding sites.
Diwei Liu, Yongjun Tang, Zhigang Chen 0001, Lei Deng 0002
BIBM5
2016 PredRSA: a gradient boosted regression trees approach for predicting protein solvent accessibility
abstract
BACKGROUND: Protein solvent accessibility prediction is a pivotal intermediate step towards modeling protein tertiary structures directly from one-dimensional sequences. It also plays an important part in identifying protein folds and domains. Although some methods have been presented to the protein solvent accessibility prediction in recent years, the performance is far from satisfactory. In this work, we propose PredRSA, a computational method that can accurately predict relative solvent accessible surface area (RSA) of residues by exploring various local and global sequence features which have been observed to be associated with solvent accessibility. Based on these features, a novel and efficient approach, Gradient Boosted Regression Trees (GBRT), is first adopted to predict RSA. RESULTS: Experimental results obtained from 5-fold cross-validation based on the Manesh-215 dataset show that the mean absolute error (MAE) and the Pearson correlation coefficient (PCC) of PredRSA are 9.0 % and 0.75, respectively, which are better than that of the existing methods. Moreover, we evaluate the performance of PredRSA using an independent test set of 68 proteins. Compared with the state-of-the-art approaches (SPINE-X and ASAquick), PredRSA achieves a significant improvement on the prediction quality. CONCLUSIONS: Our experimental results show that the Gradient Boosted Regression Trees algorithm and the novel feature combination are quite effective in relative solvent accessibility prediction. The proposed PredRSA method could be useful in assisting the prediction of protein structures by applying the predicted RSA as useful restraints.
Diwei Liu, Zhigang Chen 0001, Lei Deng 0002
BMC Bioinform.5
2015 An Integrated Framework for Functional Annotation of Protein Structural Domains
abstract
Structural domains are evolutionary and functional units of proteins and play a critical role in comparative and functional genomics. Computational assignment of domain function with high reliability is essential for understanding whole-protein functions. However, functional annotations are conventionally assigned onto full-length proteins rather than associating specific functions to the individual structural domains. In this article, we present Structural Domain Annotation (SDA), a novel computational approach to predict functions for SCOP structural domains. The SDA method integrates heterogeneous information sources, including structure alignment based protein-SCOP mapping features, InterPro2GO mapping information, PSSM Profiles, and sequence neighborhood features, with a Bayesian network. By large-scale annotating Gene Ontology terms to SCOP domains with SDA, we obtained a database of SCOP domain to Gene Ontology mappings, which contains ~162,000 out of the approximately 166,900 domains in SCOPe 2.03 (>97 percent) and their predicted Gene Ontology functions. We have benchmarked SDA using a single-domain protein dataset and an independent dataset from different species. Comparative studies show that SDA significantly outperforms the existing function prediction methods for structural domains in terms of coverage and maximum F-measure.
Lei Deng 0002, Zhigang Chen 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2014 Structure-Based Prediction of Protein Phosphorylation Sites Using an Ensemble Approach
Weilin Hao, Zhigang Chen 0001, Lei Deng 0002
ICIC (3)4
2014 A Multi-Instance Multi-Label Learning Approach for Protein Domain Annotation
Lei Deng 0002, Zhigang Chen 0001, Diwei Liu
ICIC (3)2
2013 Boosting Prediction Performance of Protein-Protein Interaction Hot Spots by Using Structural Neighborhood Properties - (Extended Abstract)
Lei Deng 0002, Jihong Guan, Xiaoming Wei, Yuan Yi, Shuigeng Zhou
RECOMB1
2009 Prediction of protein-protein interaction sites using an ensemble method
abstract
BACKGROUND: Prediction of protein-protein interaction sites is one of the most challenging and intriguing problems in the field of computational biology. Although much progress has been achieved by using various machine learning methods and a variety of available features, the problem is still far from being solved. RESULTS: In this paper, an ensemble method is proposed, which combines bootstrap resampling technique, SVM-based fusion classifiers and weighted voting strategy, to overcome the imbalanced problem and effectively utilize a wide variety of features. We evaluate the ensemble classifier using a dataset extracted from 99 polypeptide chains with 10-fold cross validation, and get a AUC score of 0.86, with a sensitivity of 0.76 and a specificity of 0.78, which are better than that of the existing methods. To improve the usefulness of the proposed method, two special ensemble classifiers are designed to handle the cases of missing homologues and structural information respectively, and the performance is still encouraging. The robustness of the ensemble method is also evaluated by effectively classifying interaction sites from surface residues as well as from all residues in proteins. Moreover, we demonstrate the applicability of the proposed method to identify interaction sites from the non-structural proteins (NS) of the influenza A virus, which may be utilized as potential drug target sites. CONCLUSION: Our experimental results show that the ensemble classifiers are quite effective in predicting protein interaction sites. The Sub-EnClassifiers with resampling technique can alleviate the imbalanced problem and the combination of Sub-EnClassifiers with a wide variety of feature groups can significantly improve prediction performance.
Lei Deng 0002, Jihong Guan, Qiwen Dong, Shuigeng Zhou
BMC Bioinform.1