Wen Zhang 0008

dblp:43/2368-8 · DBLP profile ↗
← Back
79ranked-venue papers
24as first author
45since 2021 · last 2026
0000-0001-5221-2628ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 61 · 18 first-author · 35 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Revealing Herb-Symptom Associations and Mechanisms of Action in Protein Networks Using Subgraph Matching Learning
abstract
In traditional Chinese medicine, deciphering herb-symptom associations (HSAs) and revealing their mechanisms of action are crucial for bridging traditional knowledge and modern biomedicine. While previous studies have investigated HSAs using protein-protein interaction (PPI)-based network medicine method, they often treat all proteins equally, failing to capture the heterogeneous contributions of individual proteins to HSAs. This limitation hinders their capacity to reveal the mechanisms of action. To address this challenge, we propose a subgraph matching learning method, GraphHSA, for HSA prediction. GraphHSA maps herbs and symptoms onto the PPI network to construct subgraphs. Then, GraphHSA utilizes an attention mechanism to compute the importance of each protein on the subgraph, and weighted aggregate protein information to generate herb/symptom embeddings. Subsequently, these embeddings are combined to model the matching relationship between herb and symptom subgraphs, enabling association prediction. Additionally, a dual-contrastive learning strategy is introduced to generate discriminative representations to enhance prediction. Experiments indicate that GraphHSA not only applies to individual herbs but also extends to compound formulations composed of multiple herbs. By capturing the dynamic interactions among their components, GraphHSA enables the identification of key biological targets and the elucidation of the mechanisms underlying their therapeutic efficacy.
Menglu Li, Yujing Ni, Zhinan Mei, Wen Zhang 0008
IEEE J. Biomed. Health Informatics6
2025 PKAG-DDI: Pairwise Knowledge-Augmented Language Model for Drug-Drug Interaction Event Text Generation
abstract
Drug-drug interactions (DDIs) arise when multiple drugs are administered concurrently. Accurately predicting the specific mechanisms underlying DDIs (named DDI events or DDIEs) is critical for the safe clinical use of drugs. DDIEs are typically represented as textual descriptions. However, most computational methods focus more on predicting the DDIE class label over generating human-readable natural language increasing clinicians’ interpretation costs. Furthermore, current methods overlook the fact that each drug assumes distinct biological functions in a DDI, which, when used as input context, can enhance the understanding of the DDIE process and benefit DDIE generation by the language model (LM). In this work, we propose a novel pairwise knowledge-augmented generative method (termed PKAG-DDI) for DDIE text generation. It consists of a pairwise knowledge selector efficiently injecting structural information between drugs bidirectionally and simultaneously to select pairwise biological functions from the knowledge set, and a pairwise knowledge integration strategy that matches and integrates the selected biological functions into the LM. Experiments on two professional datasets show that PKAG-DDI outperforms existing methods in DDIE text generation, especially in challenging inductive scenarios, indicating its practicality and generalization.
Zhankun Xiong, Feng Huang 0004, Wen Zhang 0008
ACL (1)4
2025 Enhancing Drug-Drug Interaction Prediction via Drug-Centric Hierarchical Augmentation
abstract
Drug-drug interactions (DDIs) can critically affect treatment safety and efficacy, especially when multiple drugs are prescribed concurrently. Such interactions may alter pharmacological activity and complicate therapeutic outcomes. Although graph-based learning methods have advanced DDI prediction, most rely solely on drug-drug networks, overlooking valuable information from auxiliary drug-centric networks. To overcome this limitation, it is crucial to adopt more hierarchical strategies that incorporate both drug-drug and auxiliary drug-centric networks to capture nuanced drug representations. In this research, we construct a hierarchical network to incorporate both drug-drug and auxiliary networks as distinct layers and propose a drug-centric hierarchical augmentation method (DCHA) for DDI prediction. DCHA encompasses three core components: a hierarchical learner, a layer discriminator, and a DDI predictor. The hierarchical learner employs a fusion gate to compute augmented drug representations by integrating core drug representations from the drug-drug network and auxiliary representations from other auxiliary drug-centric networks. The layer discriminator helps the hierarchical learner in capturing auxiliary drug representations. With the help of the hierarchical learner and layer discriminator, the DDI predictor finally augments the performance of DDI prediction. Extensive experimentation demonstrates that DCHA outperforms existing state-of-the-art methods in DDI prediction.
Ziwen Cui, Muhammad Asif Ali, Huan Wang 0005, Ruigang Liu, Wen Zhang 0008, Di Wang 0015
BIBM5
2025 CATSyn: Predicting Synergistic Drug Combinations Through Context-Aware Heterogeneous Graph Convolution Model
abstract
Accurately predicting drug synergy in cancer therapy remains challenging due to the strong dependence of drug effectiveness on cell-line context. Most existing models overlook this variability and fail to fully capture drug-cell line interactions. We present CATSyn, a Context-Aware heTerogeneous graph model for synergistic drug combination prediction. CATSyn introduces a context-aware attention mechanism that dynamically adjusts network weights based on cell-line environments, capturing cellspecific drug effects while maintaining generalization. To model these effects, we construct heterogeneous graphs that integrate drug and cell-line features into composite nodes, supported by universal nodes to share global information. Experiments on benchmark datasets demonstrate that CATSyn achieves state-of-the-art performance in both standard and unseen cell-line settings, highlighting its ability to balance specificity and generalization in synergy prediction.
Biyang Zeng, Shikui Tu, Wen Zhang 0008, Lei Xu 0001
BIBM3
2025 Geometric Heterogeneous Graph Neural Network for Protein-Ligand Binding Affinity Prediction
abstract
Accurately predicting protein-ligand binding affinity (PLA) remains a critical challenge in structure-based drug discovery. Recent advances have focused on using geometry-aware graph neural networks to model the three-dimensional (3D) structure of protein-ligand complexes for PLA prediction. However, they still achieve suboptimal performance due to two potential issues. 1) Representing the protein-ligand complex as a homogeneous graph ignores the inherent difference between intra- and intermolecular interactions, limiting the expressive ability of the models. 2) Given that the geometric complementarity between the ligand and protein binding pocket serves as a fundamental determinant of binding strength, the incomplete exploitation of geometric information constrains the predictive performance. In this study, we propose a novel Geometric Heterogeneous Graph Neural Network (GeoHGN) for PLA prediction. Specifically, we consider complete geometries to characterize the directions of edges in coordinate space through quantum inspired basis functions. To sufficiently incorporate the 3D information and heterogeneous topology of the complexes, we elaborately design a novel heterogeneous directional message passing mechanism (HDMP), which enables the propagation and aggregation of messages from intra- and intermolecular neighbors along with the directional information of linked edges. Extensive benchmarking experiments demonstrate the superiority of GeoHGN in predicting PLA.
Feng Huang 0004, Yuhang Xia, Wen Zhang 0008
CIKM5
2025 3A Multi-Classification Division-Aggregation Framework for Fake News Detection
abstract
Nowadays, as human activities are shifting to social media, fake news detection has been a crucial problem. Existing methods ignore the classification difference in online news and cannot take full advantage of multi-classification knowledges. For example, when coping with a post “A mouse is frightened by a cat,” a model that learns “computer” knowledge tends to misunderstand “mouse” and give a fake label, but a model that learns “animal” knowledge tends to give a true label. Therefore, this research proposes a multi-classification division-aggregation framework to detect fake news, namedCKA, which innovatively learns classification knowledges during training stages and aggregates them during prediction stages. It consists of three main components: a news characterizer, an ensemble coordinator, and a truth predictor. The news characterizer is responsible for extracting news features and obtaining news classifications. Cooperating with the news characterizer, the ensemble coordinator generates classification-specifical models for the maximum reservation of classification knowledges during the training stage, where each classification-specifical model maximizes the detection performance of fake news on corresponding news classifications. Further, to aggregate the classification knowledges during the prediction stage, the truth predictor uses the truth discovery technology to aggregate the predictions from different classification-specifical models based on reliability evaluation of classification-specifical models. Extensive experiments prove that our proposedCKAoutperforms state-of-the-art baselines in fake news detection.
Wen Zhang 0008, Haitao Fu, Huan Wang 0005, Zhiguo Gong, Pan Zhou 0001, Di Wang 0015
IEEE Trans. Big Data1
2025 GNNDRP: Graph Neural Network With Multi-Task Learning for Drug Response Prediction
abstract
Using computational methods to personalize drug response prediction holds great promise to improve cancer therapy. Most existing methods use either biochemical information or response-related networks to predict drug response, nevertheless, the information they considered is not comprehensive. In this study, we present a novel end-to-end deep learning-based method Graph Neural Network with multi-task learning for Drug Response Prediction (GNNDRP). It leverages biochemical features as well as the hidden features from the heterogeneous network which incorporates the known drug-cell line responses, drug similarities, and cell line similarities, to complete the drug response prediction task. Moreover, GNNDRP designs a self-supervised task to enhance the representation capacity from the response network and further improve the model prediction performance. Extensive experiments show that GNNDRP outperforms existing state-of-the-art prediction methods under various experimental settings. The ablation analysis reveals that the biochemical characteristics, response-related network, and our self-supervised strategy can boost the predictive power. Additionally, case studies further validate the effectiveness of GNNDRP in identifying novel drug-cell line responses.
Congzhi Song, Xuan Liu 0010, Zhankun Xiong, Luotao Liu, Wen Zhang 0008
IEEE Trans. Comput. Biol. Bioinform.5
2025 Guest Editorial: The Cutting-Edge Artificial Intelligence Techniques and Their Applications in Drug Discovery
Wen Zhang 0008, Qi Zhao 0010, Yuru Wang
IEEE J. Biomed. Health Informatics1
2024 A Multi-Modal Contrastive Diffusion Model for Therapeutic Peptide Generation
abstract
Therapeutic peptides represent a unique class of pharmaceutical agents crucial for the treatment of human diseases. Recently, deep generative models have exhibited remarkable potential for generating therapeutic peptides, but they only utilize sequence or structure information alone, which hinders the performance in generation. In this study, we propose a Multi-Modal Contrastive Diffusion model (MMCD), fusing both sequence and structure modalities in a diffusion framework to co-generate novel peptide sequences and structures. Specifically, MMCD constructs the sequence-modal and structure-modal diffusion models, respectively, and devises a multi-modal contrastive learning strategy with inter-contrastive and intra-contrastive in each diffusion timestep, aiming to capture the consistency between two modalities and boost model performance. The inter-contrastive aligns sequences and structures of peptides by maximizing the agreement of their embeddings, while the intra-contrastive differentiates therapeutic and non-therapeutic peptides by maximizing the disagreement of their sequence/structure embeddings simultaneously. The extensive experiments demonstrate that MMCD performs better than other state-of-the-art deep generative methods in generating therapeutic peptides across various metrics, including antimicrobial/anticancer score, diversity, and peptide-docking.
Xuan Liu 0010, Feng Huang 0004, Zhankun Xiong, Wen Zhang 0008
AAAI5
2024 Improving Paratope and Epitope Prediction by Multi-Modal Contrastive Learning and Interaction Informativeness Estimation
Wen Zhang 0008
IJCAI3
2024 ZeroDDI: A Zero-Shot Drug-Drug Interaction Event Prediction Method with Semantic Enhanced Learning and Dual-modal Uniform Alignment
Zhankun Xiong, Feng Huang 0004, Xuan Liu 0010, Wen Zhang 0008
IJCAI5
2024 Heterogeneous Causal Metapath Graph Neural Network for Gene-Microbe-Disease Association Prediction
Feng Huang 0004, Luotao Liu, Zhankun Xiong, Yuan Quan, Wen Zhang 0008
IJCAI7
2024 DeepCRBP: improved predicting function of circRNA-RBP binding sites with deep feature learning
Zishan Xu, Linlin Song, Shichao Liu 0002, Wen Zhang 0008
Frontiers Comput. Sci.4
2024 A multi-stream network for retrosynthesis prediction
Qiang Zhang 0031, Juan Liu 0007, Wen Zhang 0008
Frontiers Comput. Sci.3
2024 A Network Enhancement Method to Identify Spurious Drug-Drug Interactions
abstract
As medical safety and drug regulation gain heightened attention, the detection of spurious drug-drug interactions (DDI) has become key in healthcare. Although current research using graph neural networks (GNNs) to predict DDI has shown impressive results, it often fails to account for false DDI in the constructed DDI networks. Such inaccuracies caused by data errors, false alarms, or incorrect drug details can skew the network's structure and hinder the accuracy of GNN-based predictions. To tackle this challenge, we propose ANSM, a network-enhancement method specifically designed to identify and attenuate spurious links between drugs for ensuring the accuracy of DDI networks. ANSM integrates three key components: the feature extractor, the network optimizer, and the discriminative classifier. The feature extractor captures local structural features from drug node pairs, while the network optimizer leverages network information to improve feature extraction and reduce the impact of spurious DDI links. The discriminative classifier then identifies potential spurious links. Experimental results demonstrate that ANSM outperforms state-of-the-art methods in identifying spurious DDI.
Huan Wang 0005, Ziwen Cui, Yinguang Yang, Baijing Wang, Lida Zhu, Wen Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.6
2024 Subgraph-Aware Graph Kernel Neural Network for Link Prediction in Biological Networks
abstract
Identifying links within biological networks is important in various biomedical applications. Recent studies have revealed that each node in a network may play a unique role in different links, but most link prediction methods overlook distinctive node roles, hindering the acquisition of effective link representations. Subgraph-based methods have been introduced as solutions but often ignore shared information among subgraphs. To address these limitations, we propose a Subgraph-aware Graph Kernel Neural Network (SubKNet) for link prediction in biological networks. Specifically, SubKNet extracts a subgraph for each node pair and feeds it into a graph kernel neural network, which decomposes each subgraph into a combination of trainable graph filters with diversity regularization for subgraph-aware representation learning. Additionally, node embeddings of the network are extracted as auxiliary information, aiding in distinguishing node pairs that share the same subgraph. Extensive experiments on five biological networks demonstrate that SubKNet outperforms baselines, including methods especially designed for biological networks and methods adapted to various networks. Further investigations confirm that employing graph filters to subgraphs helps to distinguish node roles in different subgraphs, and the inclusion of diversity regularization further enhances its capacity from diverse perspectives, generating effective link representations that contribute to more accurate link prediction.
Menglu Li, Luotao Liu, Xuan Liu 0010, Wen Zhang 0008
IEEE J. Biomed. Health Informatics5
2023 Multi-Relational Contrastive Learning Graph Neural Network for Drug-Drug Interaction Event Prediction
abstract
Drug-drug interactions (DDIs) could lead to various unexpected adverse consequences, so-called DDI events. Predicting DDI events can reduce the potential risk of combinatorial therapy and improve the safety of medication use, and has attracted much attention in the deep learning community. Recently, graph neural network (GNN)-based models have aroused broad interest and achieved satisfactory results in the DDI event prediction. Most existing GNN-based models ignore either drug structural information or drug interactive information, but both aspects of information are important for DDI event prediction. Furthermore, accurately predicting rare DDI events is hindered by their inadequate labeled instances. In this paper, we propose a new method, Multi-Relational Contrastive learning Graph Neural Network, MRCGNN for brevity, to predict DDI events. Specifically, MRCGNN integrates the two aspects of information by deploying a GNN on the multi-relational DDI event graph attributed with the drug features extracted from drug molecular graphs. Moreover, we implement a multi-relational graph contrastive learning with a designed dual-view negative counterpart augmentation strategy, to capture implicit information about rare DDI events. Extensive experiments on two datasets show that MRCGNN outperforms the state-of-the-art methods. Besides, we observe that MRCGNN achieves satisfactory performance when predicting rare DDI events.
Zhankun Xiong, Shichao Liu 0002, Feng Huang 0004, Xuan Liu 0010, Zhongfei Zhang, Wen Zhang 0008
AAAI7
2023 Causal Intervention for Measuring Confidence in Drug-Target Interaction Prediction
abstract
Identifying and discovering drug-target interactions (DTIs) are vital steps in drug discovery and development. They play a crucial role in assisting scientists in finding new drugs and accelerating the drug development process. Recently, knowledge graph and knowledge graph embedding (KGE) models have made rapid advancements and demonstrated impressive performance in drug discovery. However, such models lack authenticity and accuracy in drug target identification, leading to an increased misjudgment rate and reduced drug development efficiency. To address these issues, we focus on the problem of drug-target interactions, with knowledge mapping as the core technology. Specifically, a causal intervention-based confidence measure is employed to assess the triplet score to improve the accuracy of the drug-target interaction prediction model. Experimental results demonstrate that the developed confidence measurement method based on causal intervention can significantly enhance the accuracy of DTI link prediction, particularly for high-precision models. The predicted results are more valuable in guiding the design and development of subsequent drug development experiments, thereby significantly improving the efficiency of drug development.
Wenting Ye, Chen Li 0027, Wen Zhang 0008, Debo Cheng, Zaiwen Feng
BIBM4
2023 Multi-view Contrastive Learning Hypergraph Neural Network for Drug-Microbe-Disease Association Prediction
abstract
Identifying the potential associations among drugs, microbes and diseases is of great significance in exploring the pathogenesis and improving precision medicine. There are plenty of computational methods for pair-wise association prediction, such as drug-microbe and microbe-disease associations, but few methods focus on the higher-order triple-wise drug-microbe-disease (DMD) associations. Driven by the advancement of hypergraph neural networks (HGNNs), we expect them to fully capture high-order interaction patterns behind the hypergraph formulated by DMD associations and realize sound prediction performance. However, the confirmed DMD associations are insufficient due to the high cost of in vitro screening, which forms a sparse DMD hypergraph and thus brings in suboptimal generalization ability. To mitigate the limitation, we propose a Multi-view Contrastive Learning Hypergraph Neural Network, named MCHNN, for DMD association prediction. We design a novel multi-view contrastive learning on the DMD hypergraph as an auxiliary task, which guides the HGNN to learn more discriminative representations and enhances the generalization ability. Extensive computational experiments show that MCHNN achieves satisfactory performance in DMD association prediction and, more importantly, demonstrate the effectiveness of our devised multi-view contrastive learning on the sparse DMD hypergraph.
Luotao Liu, Feng Huang 0004, Xuan Liu 0010, Zhankun Xiong, Menglu Li, Congzhi Song, Wen Zhang 0008
IJCAI7
2023 HimGNN: a novel hierarchical molecular graph representation learning framework for property prediction
abstract
Accurate prediction of molecular properties is an important topic in drug discovery. Recent works have developed various representation schemes for molecular structures to capture different chemical information in molecules. The atom and motif can be viewed as hierarchical molecular structures that are widely used for learning molecular representations to predict chemical properties. Previous works have attempted to exploit both atom and motif to address the problem of information loss in single representation learning for various tasks. To further fuse such hierarchical information, the correspondence between learned chemical features from different molecular structures should be considered. Herein, we propose a novel framework for molecular property prediction, called hierarchical molecular graph neural networks (HimGNN). HimGNN learns hierarchical topology representations by applying graph neural networks on atom- and motif-based graphs. In order to boost the representational power of the motif feature, we design a Transformer-based local augmentation module to enrich motif features by introducing heterogeneous atom information in motif representation learning. Besides, we focus on the molecular hierarchical relationship and propose a simple yet effective rescaling module, called contextual self-rescaling, that adaptively recalibrates molecular representations by explicitly modelling interdependencies between atom and motif features. Extensive computational experiments demonstrate that HimGNN can achieve promising performances over state-of-the-art baselines on both classification and regression tasks in molecular property prediction.
Shen Han, Haitao Fu, Yuyang Wu, Ganglan Zhao, Feng Huang 0004, Zhongfei Zhang, Shichao Liu 0002, Wen Zhang 0008
Briefings Bioinform.9
2023 A subcomponent-guided deep learning method for interpretable cancer drug response prediction
abstract
Accurate prediction of cancer drug response (CDR) is a longstanding challenge in modern oncology that underpins personalized treatment. Current computational methods implement CDR prediction by modeling responses between entire drugs and cell lines, without the consideration that response outcomes may primarily attribute to a few finer-level 'subcomponents', such as privileged substructures of the drug or gene signatures of the cancer cell, thus producing predictions that are hard to explain. Herein, we present SubCDR, a subcomponent-guided deep learning method for interpretable CDR prediction, to recognize the most relevant subcomponents driving response outcomes. Technically, SubCDR is built upon a line of deep neural networks that enables a set of functional subcomponents to be extracted from each drug and cell line profile, and breaks the CDR prediction down to identifying pairwise interactions between subcomponents. Such a subcomponent interaction form can offer a traceable path to explicitly indicate which subcomponents contribute more to the response outcome. We verify the superiority of SubCDR over state-of-the-art CDR prediction methods through extensive computational experiments on the GDSC dataset. Crucially, we found many predicted cases that demonstrate the strength of SubCDR in finding the key subcomponents driving responses and exploiting these subcomponents to discover new therapeutic drugs. These results suggest that SubCDR will be highly useful for biomedical researchers, particularly in anti-cancer drug design.
Xuan Liu 0010, Wen Zhang 0008
PLoS Comput. Biol.2
2023 DRLM: A Robust Drug Representation Learning Method and its Applications
abstract
Learning representations from data is a fundamental step for machine learning. High-quality and robust drug representations can broaden the understanding of pharmacology, and improve the modeling of multiple drug-related prediction tasks, which further facilitates drug development. Although there are a number of models developed for drug representation learning from various data sources, few researches extract drug representations from gene expression profiles. Since gene expression profiles of drug-treated cells are widely used in clinical diagnosis and therapy, it is believed that leveraging them to eliminate cell specificity can promote drug representation learning. In this paper, we propose a three-stage deep learning method for drug representation learning, named DRLM, which integrates gene expression profiles of drug-related cells and the therapeutic use information of drugs. Firstly, we construct a stacked autoencoder to learn low-dimensional compact drug representations. Secondly, we utilize an iterative clustering module to reduce the negative effects of cell specificity and noise in gene expression profiles on the low-dimensional drug representations. Thirdly, a therapeutic use discriminator is designed to incorporate therapeutic use information into the drug representations. The visualization analysis of drug representations demonstrates DRLM can reduce cell specificity and integrate therapeutic use information effectively. Extensive experiments on three types of prediction tasks are conducted based on different drug representations, and they show that the drug representations learned by DRLM outperform other representations in terms of most metrics. The ablation analysis also demonstrates DRLM's effectiveness of merging the gene expression profiles with the therapeutic use information.
Haitao Fu, Cecheng Zhao, Xiaohui Niu, Wen Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 Enhancing Drug-Drug Interaction Prediction Using Deep Attention Neural Networks
abstract
Drug-drug interactions are one of the main concerns in drug discovery. Accurate prediction of drug-drug interactions plays a key role in increasing the efficiency of drug research and safety when multiple drugs are co-prescribed. With various data sources that describe the relationships and properties between drugs, the comprehensive approach that integrates multiple data sources would be considerably effective in making high-accuracy prediction. In this paper, we propose a Deep Attention Neural Network based Drug-Drug Interaction prediction framework, abbreviated as DANN-DDI, to predict unobserved drug-drug interactions. First, we construct multiple drug feature networks and learn drug representations from these networks using the graph embedding method; then, we concatenate the learned drug embeddings and design an attention neural network to learn representations of drug-drug pairs; finally, we adopt a deep neural network to accurately predict drug-drug interactions. The experimental results demonstrate that our model DANN-DDI has improved prediction performance compared with state-of-the-art methods. Moreover, the proposed model can predict novel drug-drug interactions and drug-drug interaction-associated events.
Shichao Liu 0002, Yang Zhang 0123, Zhongfei Zhang, Wen Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.7
2022 GMFQP: An Ontology-mediated Gut Microbiota Federated Query Platform
abstract
The symbiotic relationship between gut microbes and their hosts has been widely investigated by scientists. Currently, there exists lots of useful data dispersed in different gut microbiota databases. However, it is not easy for biologists to obtain valuable global information directly from these multi-sources and heterogeneous data for further research. Therefore, an ontology-mediated Gut Microbiota Federated Query Platform (GMFQP) is presented in this paper to solve this problem. Firstly, a Gut Microbiota Ontology (GMO) is built to provide a unified domain view for biologists, hiding the heterogeneity of the underlying data sources. Secondly, a three-phased federated query approach including query reasoning, rewriting, and execution based on Datalog+ rules is proposed to transform users’ queries expressed over the GMO to concrete queries executed over the actual data sources. Thirdly, a federated query prototype is implemented based on the approach mentioned above, providing a user-friendly query interface and returning results in both tabular and graphical formats. Several typical biological query examples with performance analysis are illustrated to demonstrate the utility of our approach and the platform.
Yuzhuo Dai, Yue Tang 0005, Wolfgang Mayer, Haiqin Li, Wen Zhang 0008, Zaiwen Feng
BIBM10
2022 Predicting drug transcriptional response similarity using Signed Graph Convolutional Network
abstract
Exploring the transcriptional response after employing chemical compounds assists in treating gene-related diseases and understanding biological activity of compounds. Calculating the similarity of drug transcriptional response can help to discover novel compounds that have the similar biological activity to known drugs for treating the same disease. Considering the transcriptional profiles of compounds are limited and harder to get than the structure of compounds, it is worth modeling the structure-transcriptional response similarity relationship. In this paper, we propose a signed graph convolutional network (SGCN)-based method, namely SGCN-DTRS, to predict drug transcriptional response similarity, which is quantitatively measured by Connectivity Map (CMap) scores, from their structures. SGCNDTRS constructs a CMap signed network from compounds and their CMap scores in the training data, which takes compounds as nodes, molecular structural representations of compounds as the attributes of nodes, and similarity relations between compounds as edges. Then SGCN-DTRS learns the CMap compound embeddings to predict CMap scores of pairwise compounds. Extensive experiments verify the superiority of the proposed method against the compared state-of-the-art methods and reveal that the relational information of the CMap data, which is learned from CMap signed network, is important for the CMap score prediction. SGCN-DTRS can not only work for the compounds in the training set but also is applicable to unseen compounds.
Chengzhi Hong, Xuan Liu 0010, Zhankun Xiong, Wen Zhang 0008
BIBM6
2022 META-DDIE: predicting drug-drug interaction events with few-shot learning
abstract
Drug-drug interactions (DDIs) are one of the major concerns in pharmaceutical research, and a number of computational methods have been developed to predict whether two drugs interact or not. Recently, more attention has been paid to events caused by the DDIs, which is more useful for investigating the mechanism hidden behind the combined drug usage or adverse reactions. However, some rare events may only have few examples, hindering them from being precisely predicted. To address the above issues, we present a few-shot computational method named META-DDIE, which consists of a representation module and a comparing module, to predict DDI events. We collect drug chemical structures and DDIs from DrugBank, and categorize DDI events into hundreds of types using a standard pipeline. META-DDIE uses the structures of drugs as input and learns the interpretable representations of DDIs through the representation module. Then, the model uses the comparing module to predict whether two representations are similar, and finally predicts DDI events with few labeled examples. In the computational experiments, META-DDIE outperforms several baseline methods and especially enhances the predictive capability for rare events. Moreover, META-DDIE helps to identify the key factors that may cause DDI events and reveal the relationship among different events.
Xinran Xu, Shichao Liu 0002, Zhongfei Zhang, Shanfeng Zhu, Wen Zhang 0008
Briefings Bioinform.7
2022 PHIAF: prediction of phage-host interactions with GAN-based data augmentation and sequence-based feature fusion
abstract
Phage therapy has become one of the most promising alternatives to antibiotics in the treatment of bacterial diseases, and identifying phage-host interactions (PHIs) helps to understand the possible mechanism through which a phage infects bacteria to guide the development of phage therapy. Compared with wet experiments, computational methods of identifying PHIs can reduce costs and save time and are more effective and economic. In this paper, we propose a PHI prediction method with a generative adversarial network (GAN)-based data augmentation and sequence-based feature fusion (PHIAF). First, PHIAF applies a GAN-based data augmentation module, which generates pseudo PHIs to alleviate the data scarcity. Second, PHIAF fuses the features originated from DNA and protein sequences for better performance. Third, PHIAF utilizes an attention mechanism to consider different contributions of DNA/protein sequence-derived features, which also provides interpretability of the prediction model. In computational experiments, PHIAF outperforms other state-of-the-art PHI prediction methods when evaluated via 5-fold cross-validation (AUC and AUPR are 0.88 and 0.86, respectively). An ablation study shows that data augmentation, feature fusion and an attention mechanism are all beneficial to improve the prediction performance of PHIAF. Additionally, four new PHIs with the highest PHIAF score in the case study were verified by recent literature. In conclusion, PHIAF is a promising tool to accelerate the exploration of phage therapy.
Menglu Li, Wen Zhang 0008
Briefings Bioinform.2
2022 GraphCDR: a graph neural network method with contrastive learning for cancer drug response prediction
abstract
Predicting the response of a cancer cell line to a therapeutic drug is an important topic in modern oncology that can help personalized treatment for cancers. Although numerous machine learning methods have been developed for cancer drug response (CDR) prediction, integrating diverse information about cancer cell lines, drugs and their known responses still remains a great challenge. In this paper, we propose a graph neural network method with contrastive learning for CDR prediction. GraphCDR constructs a graph neural network based on multi-omics profiles of cancer cell lines, the chemical structure of drugs and known cancer cell line-drug responses for CDR prediction, while a contrastive learning task is presented as a regularizer within a multi-task learning paradigm to enhance the generalization ability. In the computational experiments, GraphCDR outperforms state-of-the-art methods under different experimental configurations, and the ablation study reveals the key components of GraphCDR: biological features, known cancer cell line-drug responses and contrastive learning are important for the high-accuracy CDR prediction. The experimental analyses imply the predictive power of GraphCDR and its potential value in guiding anti-cancer drug selection.
Xuan Liu 0010, Congzhi Song, Feng Huang 0004, Haitao Fu, Wenjie Xiao, Wen Zhang 0008
Briefings Bioinform.6
2022 A heterogeneous network-based method with attentive meta-path extraction for predicting drug-target interactions
abstract
Predicting drug-target interactions (DTIs) is crucial at many phases of drug discovery and repositioning. Many computational methods based on heterogeneous networks (HNs) have proved their potential to predict DTIs by capturing extensive biological knowledge and semantic information from meta-paths. However, existing methods manually customize meta-paths, which is overly dependent on some specific expertise. Such strategy heavily limits the scalability and flexibility of these models, and even affects their predictive performance. To alleviate this limitation, we propose a novel HN-based method with attentive meta-path extraction for DTI prediction, named HampDTI, which is capable of automatically extracting useful meta-paths through a learnable attention mechanism instead of pre-definition based on domain knowledge. Specifically, by scoring multi-hop connections across various relations in the HN with each relation assigned an attention weight, HampDTI constructs a new trainable graph structure, called meta-path graph. Such meta-path graph implicitly measures the importance of every possible meta-path between drugs and targets. To enable HampDTI to extract more diverse meta-paths, we adopt a multi-channel mechanism to generate multiple meta-path graphs. Then, a graph neural network is deployed on the generated meta-path graphs to yield the multi-channel embeddings of drugs and targets. Finally, HampDTI fuses all embeddings from different channels for predicting DTIs. The meta-path graphs are optimized along with the model training such that HampDTI can adaptively extract valuable meta-paths for DTI prediction. The experiments on benchmark datasets not only show the superiority of HampDTI in DTI prediction over several baseline methods, but also, more importantly, demonstrate the effectiveness of the model discovering important meta-paths.
Hongzhun Wang, Feng Huang 0004, Zhankun Xiong, Wen Zhang 0008
Briefings Bioinform.4
2022 SGNNMD: signed graph neural network for predicting deregulation types of miRNA-disease associations
abstract
MiRNAs are a class of small non-coding RNA molecules that play an important role in many biological processes, and determining miRNA-disease associations can benefit drug development and clinical diagnosis. Although great efforts have been made to develop miRNA-disease association prediction methods, few attention has been paid to in-depth classification of miRNA-disease associations, e.g. up/down-regulation of miRNAs in diseases. In this paper, we regard known miRNA-disease associations as a signed bipartite network, which has miRNA nodes, disease nodes and two types of edges representing up/down-regulation of miRNAs in diseases, and propose a signed graph neural network method (SGNNMD) for predicting deregulation types of miRNA-disease associations. SGNNMD extracts subgraphs around miRNA-disease pairs from the signed bipartite network and learns structural features of subgraphs via a labeling algorithm and a neural network, and then combines them with biological features (i.e. miRNA-miRNA functional similarity and disease-disease semantic similarity) to build the prediction model. In the computational experiments, SGNNMD achieves highly competitive performance when compared with several baselines, including the signed graph link prediction methods, multi-relation prediction methods and one existing deregulation type prediction method. Moreover, SGNNMD has good inductive capability and can generalize to miRNAs/diseases unseen during the training.
Guangzhan Zhang, Menglu Li, Xinran Xu, Xuan Liu 0010, Wen Zhang 0008
Briefings Bioinform.6
2022 Predicting cell line-specific synergistic drug combinations through a relational graph convolutional network with attention mechanism
abstract
Identifying synergistic drug combinations (SDCs) is a great challenge due to the combinatorial complexity and the fact that SDC is cell line specific. The existing computational methods either did not consider the cell line specificity of SDC, or did not perform well by building model for each cell line independently. In this paper, we present a novel encoder-decoder network named SDCNet for predicting cell line-specific SDCs. SDCNet learns common patterns across different cell lines as well as cell line-specific features in one model for drug combinations. This is realized by considering the SDC graphs of different cell lines as a relational graph, and constructing a relational graph convolutional network (R-GCN) as the encoder to learn and fuse the deep representations of drugs for different cell lines. An attention mechanism is devised to integrate the drug features from different layers of the R-GCN according to their relative importance so that representation learning is further enhanced. The common patterns are exploited through partial parameter sharing in cell line-specific decoders, which not only reconstruct the known SDCs but also predict new ones for each cell line. Experiments on various datasets demonstrate that SDCNet is superior to state-of-the-art methods and is also robust when generalized to new cell lines that are different from the training ones. Finally, the case study again confirms the effectiveness of our method in predicting novel reliable cell line-specific SDCs.
Peng Zhang 0098, Shikui Tu, Wen Zhang 0008, Lei Xu 0001
Briefings Bioinform.3
2022 MVGCN: data integration through multi-view graph convolutional network for predicting links in biomedical bipartite networks
abstract
MOTIVATION: There are various interaction/association bipartite networks in biomolecular systems. Identifying unobserved links in biomedical bipartite networks helps to understand the underlying molecular mechanisms of human complex diseases and thus benefits the diagnosis and treatment of diseases. Although a great number of computational methods have been proposed to predict links in biomedical bipartite networks, most of them heavily depend on features and structures involving the bioentities in one specific bipartite network, which limits the generalization capacity of applying the models to other bipartite networks. Meanwhile, bioentities usually have multiple features, and how to leverage them has also been challenging. RESULTS: In this study, we propose a novel multi-view graph convolution network (MVGCN) framework for link prediction in biomedical bipartite networks. We first construct a multi-view heterogeneous network (MVHN) by combining the similarity networks with the biomedical bipartite network, and then perform a self-supervised learning strategy on the bipartite network to obtain node attributes as initial embeddings. Further, a neighborhood information aggregation (NIA) layer is designed for iteratively updating the embeddings of nodes by aggregating information from inter- and intra-domain neighbors in every view of the MVHN. Next, we combine embeddings of multiple NIA layers in each view, and integrate multiple views to obtain the final node embeddings, which are then fed into a discriminator to predict the existence of links. Extensive experiments show MVGCN performs better than or on par with baseline methods and has the generalization capacity on six benchmark datasets involving three typical tasks. AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/fuhaitao95/MVGCN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haitao Fu, Feng Huang 0004, Xuan Liu 0010, Wen Zhang 0008
Bioinform.5
2022 Multi-way relation-enhanced hypergraph representation learning for anti-cancer drug synergy prediction
abstract
MOTIVATION: Drug combinations have exhibited promise in treating cancers with less toxicity and fewer adverse reactions. However, in vitro screening of synergistic drug combinations is time-consuming and labor-intensive because of the combinatorial explosion. Although a number of computational methods have been developed for predicting synergistic drug combinations, the multi-way relations between drug combinations and cell lines existing in drug synergy data have not been well exploited. RESULTS: We propose a multi-way relation-enhanced hypergraph representation learning method to predict anti-cancer drug synergy, named HypergraphSynergy. HypergraphSynergy formulates synergistic drug combinations over cancer cell lines as a hypergraph, in which drugs and cell lines are represented by nodes and synergistic drug-drug-cell line triplets are represented by hyperedges, and leverages the biochemical features of drugs and cell lines as node attributes. Then, a hypergraph neural network is designed to learn the embeddings of drugs and cell lines from the hypergraph and predict drug synergy. Moreover, the auxiliary task of reconstructing the similarity networks of drugs and cell lines is considered to enhance the generalization ability of the model. In the computational experiments, HypergraphSynergy outperforms other state-of-the-art synergy prediction methods on two benchmark datasets for both classification and regression tasks and is applicable to unseen drug combinations or cell lines. The studies revealed that the hypergraph formulation allows us to capture and explain complex multi-way relations of drug combinations and cell lines, and also provides a flexible framework to make the best use of diverse information. AVAILABILITY AND IMPLEMENTATION: The source data and codes of HypergraphSynergy can be freely downloaded from https://github.com/liuxuan666/HypergraphSynergy. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xuan Liu 0010, Congzhi Song, Shichao Liu 0002, Menglu Li, Xionghui Zhou, Wen Zhang 0008
Bioinform.6
2022 Hierarchical graph representation learning for the prediction of drug-target binding affinity
Zhaoyang Chu, Feng Huang 0004, Haitao Fu, Yuan Quan, Xionghui Zhou, Shichao Liu 0002, Wen Zhang 0008
Inf. Sci.7
2022 Predicting Coding Potential of RNA Sequences by Solving Local Data Imbalance
abstract
Non-coding RNAs (ncRNAs)play an important role in various biological processes and are associated with diseases. Distinguishing between coding RNAs and ncRNAs, also known as predicting coding potential of RNA sequences, is critical for downstream biological function analysis. Many machine learning-based methods have been proposed for predicting coding potential of RNA sequences. Recent studies reveal that most existing methods have poor performance on RNA sequences with short Open Reading Frames (sORF, ORF length<303nt). In this work, we analyze the distribution of ORF length of RNA sequences, and observe that the number of coding RNAs with sORF is inadequate and coding RNAs with sORF are much less than ncRNAs with sORF. Thus, there exists the problem of local data imbalance in RNA sequences with sORF. We propose a coding potential prediction method CPE-SLDI, which uses data oversampling techniques to augment samples for coding RNAs with sORF so as to alleviate local data imbalance. Compared with existing methods, CPE-SLDI produces the better performances, and studies reveal that data augmentation by various data oversampling techniques can enhance the performance of coding potential prediction, especially for RNA sequences with sORF. The implementation of the proposed method is available at https://github.com/chenxgscuec/CPESLDI.
Xiangan Chen, Shuai Liu 0017, Wen Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 EPIHC: Improving Enhancer-Promoter Interaction Prediction by Using Hybrid Features and Communicative Learning
abstract
Enhancer-promoter interactions (EPIs) regulate the expression of specific genes in cells, which help facilitate understanding of gene regulation, cell differentiation and disease mechanisms. EPI identification approaches through wet experiments are often costly and time-consuming, leading to the design of high-efficiency computational methods is in demand. In this paper, we propose a deep neural network-based method named EPIHC to predict Enhancer-Promoter Interactions with Hybrid features and Communicative learning. EPIHC extracts enhancer and promoter sequence-derived features using convolutional neural networks (CNN), and then we design a communicative learning module to capture the communicative information between enhancer and promoter sequences. Besides, EPIHC takes the genomic features of enhancers and promoters into account, incorporating with the sequence-derived features to predict EPIs. The computational experiments show that EPIHC outperforms the existing state-of-the-art EPI prediction methods on the benchmark datasets and chromosome-split datasets, and the study reveals that the communicative learning module can bring explicit information about EPIs, which is ignored by CNN, and provide explainability about EPIs to some degree. Moreover, we consider two strategies to improve the performances of EPIHC in the cross-cell line prediction, and experimental results show that EPIHC constructed on some cell lines can exhibit good performances for other cell lines. The codes and data are available at https://github.com/BioMedicalBigDataMiningLab/EPIHC.
Shuai Liu 0017, Xinran Xu, Xiaohan Zhao, Shichao Liu 0002, Wen Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.6
2022 A Comprehensive Review of Computational Methods For Drug-Drug Interaction Detection
abstract
The detection of drug-drug interactions (DDIs) is a crucial task for drug safety surveillance, which provides effective and safe co-prescriptions of multiple drugs. Since laboratory researches are often complicated, costly and time-consuming, it's urgent to develop computational approaches to detect drug-drug interactions. In this paper, we conduct a comprehensive review of state-of-the-art computational methods falling into three categories: literature-based extraction methods, machine learning-based prediction methods and pharmacovigilance-based data mining methods. Literature-based extraction methods detect DDIs from published literature using natural language processing techniques; machine learning-based prediction methods build prediction models based on the known DDIs in databases and predict novel ones; pharmacovigilance-based data mining methods usually apply statistical techniques on various electronic data to detect drug-drug interaction signals. We first present the taxonomy of drug-drug interaction detection methods and provide the outlines of three categories of methods. Afterwards, we respectively introduce research backgrounds and data sources of three categories, and illustrate their representative approaches as well as evaluation metrics. Finally, we discuss the current challenges of existing methods and highlight potential opportunities for future directions.
Yang Zhang 0123, Shichao Liu 0002, Wen Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 A Multimodal Framework for Improving in Silico Drug Repositioning With the Prior Knowledge From Knowledge Graphs
abstract
Drug repositioning/repurposing is a very important approach towards identifying novel treatments for diseases in drug discovery. Recently, large-scale biological datasets are increasingly available for pharmaceutical research and promote the development of drug repositioning, but efficiently utilizing these datasets remains challenging. In this paper, we develop a novel multimodal framework, termed GraphPK (Graph-based Prior Knowledge) for improving in silico drug repositioning via using the prior knowledge from a drug knowledge graph. First, we construct a knowledge graph by integrating relevant bio-entities (drugs, diseases, etc.) and associations/interactions among them, and apply the knowledge graph embedding technique to extract prior knowledge of drugs and diseases. Moreover, we make use of the known drug-disease association, and obtain known association-based features from an association bipartite graph through graph embedding, and also take into account biological domain features, i.e., drug chemical structures and disease semantic similarity. Finally, we design a multimodal neural network to combine three types of features from the knowledge graph, the known associations and the biological domain, and build the prediction model for predicting drug-disease associations. Massive experiments show that our method outperforms other state-of-the-art methods in terms of most metrics, and the ablation analysis regarding the three types of features reveals that prior knowledge from knowledge graphs can not only lift the predictive power of in silico drug repositioning, but also enhance the model's robustness to different scenarios. The results of case studies offer support that GraphPK has the potential for actual use.
Zhankun Xiong, Feng Huang 0004, Shichao Liu 0002, Wen Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.5
2021 Predicting Drug-miRNA Resistance with Layer Attention Graph Convolution Network and Multi Channel Feature Extraction
abstract
MicroRNA (miRNA) has became an increasingly important class of attractive drug targets in recent studies. However, there are only few computational tools aiming to predict drugmi-RNA resistance associations. Hence, it is of great significance to develop effective and high accuracy methods for predicting drugmi-RNA resistance associations. In this work, we propose a novel method abbreviated as “DMR-GCN”, which enhances drugmi-RNA resistance interaction prediction by using layer attention graph convolution network and multi channel feature extraction. Specifically, DMR-GCN first constructs a heterogeneous network based on known drug-miRNA interactions, drug-drug similarities and miRNA-miRNA similarities. Secondly, layer attention graph convolution network is used to extract drug representations from the drug molecular graph and the heterogeneous network. We concatenate the extracted representations from molecular graph and heterogeneous network as the drug embedding vectors. Similarly, miRNA representations extracted from the heterogeneous network and the miRNA expression features embedded by MLP are concatenated as the miRNA embedding vectors. Further, we utilize Multi-Layer Perceptron (MLP), Generalized Tensor Factorization (GTF) and Compressed Tensor Network (CTN) to extract node-pair representations from different aspects. Finally, the predictive scores for unobserved drug-miRNA resistance associations are given by a fully connection layer with the integrated embeddings. In the evaluation experiments, DMR-GCN achieves an area under the precision-recall curve of 0.2920 and an area under the receiver-operating characteristic curve of 0.9433, which are better than the state-of-the-art prediction methods. The experimental results demonstrate that layer attention mechanism produces satisfying results for learning representations from graph, and integrating multi channel feature extraction can make further improvements. In conclusion, DMR-GCN is a promising method for predicting drug-miRNA resistance associations.
Haorui Wang, Shahanavaj Khan, Shichao Liu 0002, Fang Zheng 0010, Wen Zhang 0008
BIBM5
2021 A Graph-based Approach for Integrating Biological Heterogeneous Data Based on Connecting Ontology
abstract
Linked Open Data (LOD) is an ongoing effort in the Semantic Web community to build a massive public knowledge graph. The goal is to extend the Web by publishing various open datasets as RDF on the Web and then linking data items to other useful information from different data sources. With linked data, starting from a certain point in the graph, a person or machine can explore the graph to find other related data. In this paper, we develop a novel pipeline for graph-based biological data integration. By using our pipeline, users can easily glue heterogeneous biological ontologies, annotate sources with multiple join tables effectively, obtain a high-quality biological knowledge graph automatically, and enrich the knowledge graph with public biological ontologies finally. We implement a platform that realizes the proposed approach and conduct two case studies to evaluate the effectiveness and efficiency of our approach.
Yue Tang 0005, Linye Li, Peilin Xie, Yuanshuai Gu, Zaiwen Feng, Wen Zhang 0008, Jingbo Xia, Wolfgang Mayer, Guang-Cun He, Keqing He 0002
BIBM11
2021 CSGNN: Contrastive Self-Supervised Graph Neural Network for Molecular Interaction Prediction
abstract
Molecular interactions are significant resources for analyzing sophisticated biological systems. Identification of multifarious molecular interactions attracts increasing attention in biomedicine, bioinformatics, and human healthcare communities. Recently, a plethora of methods have been proposed to reveal molecular interactions in one specific domain. However, existing methods heavily rely on features or structures involving molecules, which limits the capacity of transferring the models to other tasks. Therefore, generalized models for the multifarious molecular interaction prediction (MIP) are in demand. In this paper, we propose a contrastive self-supervised graph neural network (CSGNN) to predict molecular interactions. CSGNN injects a mix-hop neighborhood aggregator into a graph neural network (GNN) to capture high-order dependency in the molecular interaction networks and leverages a contrastive self-supervised learning task as a regularizer within a multi-task learning paradigm to enhance the generalization ability. Experiments on seven molecular interaction networks show that CSGNN outperforms classic and state-of-the-art models. Comprehensive experiments indicate that the mix-hop aggregator and the self-supervised regularizer can effectively facilitate the link inference in multifarious molecular networks.
Chengshuai Zhao, Shuai Liu 0017, Feng Huang 0004, Shichao Liu 0002, Wen Zhang 0008
IJCAI5
2021 Tensor decomposition with relational constraints for predicting multiple types of microRNA-disease associations
abstract
MicroRNAs (miRNAs) play crucial roles in multifarious biological processes associated with human diseases. Identifying potential miRNA-disease associations contributes to understanding the molecular mechanisms of miRNA-related diseases. Most of the existing computational methods mainly focus on predicting whether a miRNA-disease association exists or not. However, the roles of miRNAs in diseases are prominently diverged, for instance, Genetic variants of miRNA (mir-15) may affect the expression level of miRNAs leading to B cell chronic lymphocytic leukemia, while circulating miRNAs (including mir-1246, mir-1307-3p, etc.) have potentials to detecting breast cancer in the early stage. In this paper, we aim to predict multi-type miRNA-disease associations instead of taking them as binary. To this end, we innovatively represent miRNA-disease-type triples as a tensor and introduce tensor decomposition methods to solve the prediction task. Experimental results on two widely-adopted miRNA-disease datasets: HMDD v2.0 and HMDD v3.2 show that tensor decomposition methods improve a recent baseline in a large scale (up to $38\%$ in Top-1F1). We then propose a novel method, Tensor Decomposition with Relational Constraints (TDRC), which incorporates biological features as relational constraints to further the existing tensor decomposition methods. Compared with two existing tensor decomposition methods, TDRC can produce better performance while being more efficient.
Feng Huang 0004, Xiang Yue, Zhankun Xiong, Zhouxin Yu, Shichao Liu 0002, Wen Zhang 0008
Briefings Bioinform.6
2021 ADEIP: an integrated platform of age-dependent expression and immune profiles across human tissues
abstract
Gene expression and immune status in human tissues are changed with aging. There is a need to develop a comprehensive platform to explore the dynamics of age-related gene expression and immune profiles across tissues in genome-wide studies. Here, we collected RNA-Seq datasets from GTEx project, containing 16 704 samples from 30 major tissues in six age groups ranging from 20 to 79 years old. Dynamic gene expression along with aging were depicted and gene set enrichment analysis was performed among those age groups. Genes from 34 known immune function categories and immune cell compositions were investigated and compared among different age groups. Finally, we integrated all the results and developed a platform named ADEIP (http://gb.whu.edu.cn/ADEIP or http://geneyun.net/ADEIP), integrating the age-dependent gene expression and immune profiles across tissues. To demonstrate the usage of ADEIP, we applied two datasets: severe acute respiratory syndrome coronavirus 2 and human mesenchymal stem cells-assoicated genes. We also included the expression and immune dynamics of these genes in the platform. Collectively, ADEIP is a powerful platform for studying age-related immune regulation in organogenesis and other infectious or genetic diseases.
Xuan Liu 0010, Wenbo Chen 0006, Liuping Chang, Haidong Ye, Wen Zhang 0008, Zhiqiang Dong, Leng Han, Chunjiang He
Briefings Bioinform.10
2021 Predicting drug-disease associations through layer attention graph convolutional network
abstract
BACKGROUND: Determining drug-disease associations is an integral part in the process of drug development. However, the identification of drug-disease associations through wet experiments is costly and inefficient. Hence, the development of efficient and high-accuracy computational methods for predicting drug-disease associations is of great significance. RESULTS: In this paper, we propose a novel computational method named as layer attention graph convolutional network (LAGCN) for the drug-disease association prediction. Specifically, LAGCN first integrates the known drug-disease associations, drug-drug similarities and disease-disease similarities into a heterogeneous network, and applies the graph convolution operation to the network to learn the embeddings of drugs and diseases. Second, LAGCN combines the embeddings from multiple graph convolution layers using an attention mechanism. Third, the unobserved drug-disease associations are scored based on the integrated embeddings. Evaluated by 5-fold cross-validations, LAGCN achieves an area under the precision-recall curve of 0.3168 and an area under the receiver-operating characteristic curve of 0.8750, which are better than the results of existing state-of-the-art prediction methods and baseline methods. The case study shows that LAGCN can discover novel associations that are not curated in our dataset. CONCLUSION: LAGCN is a useful tool for predicting drug-disease associations. This study reveals that embeddings from different convolution layers can reflect the proximities of different orders, and combining the embeddings by the attention mechanism can improve the prediction performances.
Zhouxin Yu, Feng Huang 0004, Xiaohan Zhao, Wenjie Xiao, Wen Zhang 0008
Briefings Bioinform.5
2021 A Fast Linear Neighborhood Similarity-Based Network Link Inference Method to Predict MicroRNA-Disease Associations
abstract
Increasing evidences revealed that microRNAs (miRNAs) play critical roles in important biological processes. The identification of disease-related miRNAs is critical to understand the molecular mechanisms of human diseases. Most existing computational methods require diverse features to predict miRNA-disease associations. However, diverse features are not available for all miRNAs or diseases. In addition, most methods can't predict links for miRNAs or diseases without association information. In this paper, we propose a fast linear neighborhood similarity-based network link inference method, named FLNSNLI, to predict miRNA-disease associations. First, known miRNA-disease associations are formulated as a bipartite network, and miRNAs (or diseases) are expressed as association profiles. Second, miRNA-miRNA similarity and disease-disease similarity are calculated by fast linear neighborhood similarity measure and association profiles. Third, the label propagation algorithm is respectively implemented on two sides to score candidate miRNA-disease associations. Finally, FLNSNLI adopts the weighted average strategy and makes predictions. Moreover, we develop a link complementing approach, and extend FLNSNLI to predict links for miRNAs (or diseases) without known associations. In computational experiments, FLNSNLI produces high-accuracy performances, and outperforms other state-of-the-art methods. More importantly, FLNSNLI requires less information but performs well. Case studies on three popular diseases show that FLNSNLI is useful for the microRNA-disease association prediction.
Wen Zhang 0008, Zhishuai Li, Wenzheng Guo, Weitai Yang, Feng Huang 0004
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 A multimodal deep learning framework for predicting drug-drug interaction events
abstract
MOTIVATION: Drug-drug interactions (DDIs) are one of the major concerns in pharmaceutical research. Many machine learning based methods have been proposed for the DDI prediction, but most of them predict whether two drugs interact or not. The studies revealed that DDIs could cause different subsequent events, and predicting DDI-associated events is more useful for investigating the mechanism hidden behind the combined drug usage or adverse reactions. RESULTS: In this article, we collect DDIs from DrugBank database, and extract 65 categories of DDI events by dependency analysis and events trimming. We propose a multimodal deep learning framework named DDIMDL that combines diverse drug features with deep learning to build a model for predicting DDI-associated events. DDIMDL first constructs deep neural network (DNN)-based sub-models, respectively, using four types of drug features: chemical substructures, targets, enzymes and pathways, and then adopts a joint DNN framework to combine the sub-models to learn cross-modality representations of drug-drug pairs and predict DDI events. In computational experiments, DDIMDL produces high-accuracy performances and has high efficiency. Moreover, DDIMDL outperforms state-of-the-art DDI event prediction methods and baseline methods. Among all the features of drugs, the chemical substructures seem to be the most informative. With the combination of substructures, targets and enzymes, DDIMDL achieves an accuracy of 0.8852 and an area under the precision-recall curve of 0.9208. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/YifanDengWHU/DDIMDL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xinran Xu, Jingbo Xia, Wen Zhang 0008, Shichao Liu 0002
Bioinform.5
2020 Graph embedding on biomedical networks: methods, applications and evaluations
abstract
MOTIVATION: Graph embedding learning that aims to automatically learn low-dimensional node representations, has drawn increasing attention in recent years. To date, most recent graph embedding methods are evaluated on social and information networks and are not comprehensively studied on biomedical networks under systematic experiments and analyses. On the other hand, for a variety of biomedical network analysis tasks, traditional techniques such as matrix factorization (which can be seen as a type of graph embedding methods) have shown promising results, and hence there is a need to systematically evaluate the more recent graph embedding methods (e.g. random walk-based and neural network-based) in terms of their usability and potential to further the state-of-the-art. RESULTS: We select 11 representative graph embedding methods and conduct a systematic comparison on 3 important biomedical link prediction tasks: drug-disease association (DDA) prediction, drug-drug interaction (DDI) prediction, protein-protein interaction (PPI) prediction; and 2 node classification tasks: medical term semantic type classification, protein function prediction. Our experimental results demonstrate that the recent graph embedding methods achieve promising results and deserve more attention in the future biomedical graph analysis. Compared with three state-of-the-art methods for DDAs, DDIs and protein function predictions, the recent graph embedding methods achieve competitive performance without using any biological features and the learned embeddings can be treated as complementary representations for the biological features. By summarizing the experimental results, we provide general guidelines for properly selecting graph embedding methods and setting their hyper-parameters for different biomedical tasks. AVAILABILITY AND IMPLEMENTATION: As part of our contributions in the paper, we develop an easy-to-use Python package with detailed instructions, BioNEV, available at: https://github.com/xiangyue9607/BioNEV, including all source code and datasets, to facilitate studying various graph embedding methods on biomedical tasks. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiang Yue, Zhen Wang 0041, Jingong Huang, Srinivasan Parthasarathy 0001, Soheil Moosavinasab, Yungui Huang, Simon M. Lin, Wen Zhang 0008, Ping Zhang 0016, Huan Sun 0001
Bioinform.8
2019 Structural Network Embedding using Multi-modal Deep Auto-encoders for Predicting Drug-drug Interactions
abstract
Predicting drug-drug interactions (DDIs) is crucial for patient safety and public health. The existing DDI prediction methods mainly fall into three categories: knowledge-based, similarity-based and network-based. Most recently, studies have demonstrated that integrating heterogeneous drug features is significantly important for developing high-accuracy prediction models, but it also brings many new challenges, i.e. heterogeneous properties, non-linear relations and incomplete data. In this paper, we propose a multi-modal deep auto-encoders based drug representation learning method for the DDI prediction, abbreviated as DDI-MDAE. The proposed method learns unified representations of drugs simultaneously from multiple drug feature networks using multi-modal deep auto-encoders. Then we adopt several operators on the learned drug embeddings to represent drug-drug pairs, and utilize the random forest to train models for the DDI prediction. Experimental results show that DDI-MDAE effectively learns the representations of drugs by fusing diverse information, and outperforms the other state-of-the-art benchmark methods. More importantly, DDI-MDAE works even for drugs without any known interaction.
Shichao Liu 0002, Yi-Ping Phoebe Chen, Wen Zhang 0008
BIBM5
2019 Detection of Cell Types from Single-cell RNA-seq Data using Similarity via Kernel Preserving Learning Embedding
abstract
The recent advances in single-cell sequencing techniques allow us to study biological issues on cell levels. Detecting cell types from scRNA-seq data analysis is important and meaningful. However, high-level noise and the nonlinearity and sparsity of scRNA-seq data are great challenges. In this paper, we propose a cell-type detection algorithm preserving the overall cell relations named POCR to analyze scRNA-seq data. POCR utilizes a kernel embedding similarity measure to calculate cell-to-cell similarity, by minimizing the reconstruction error of a kernel matrix, rather than the reconstruction error of the original data adopted by other similarity metrics. According to the scale of scRNA-seq datasets, we select Gaussian kernel or linear kernel to calculate the embedding. We then adopt spectral clustering to detect the cell types based on the learned cell-to-cell similarity. The results are further visualized to demonstrate the effectiveness of the cell-type detection algorithm POCR. Further analysis shows that the learned similarity could improve the clustering and visualization of cell types in scRNA-seq data. Our proposed algorithm is compared with five other state-of-the-art cell subtype detection methods. The effectiveness of the algorithms is evaluated by two criteria: ARI and NMI. The experiments show that POCR achieves accurate and robust performance across different scRNA-seq data. Our python implementation of POCR is available at https://github.com/ZeMing-Liu/POCR.
Zeming Liu, Chengzhi Hong, Yi-Ping Phoebe Chen, Shichao Liu 0002, Wen Zhang 0008
BIBM7
2019 Predicting gene-disease associations from the heterogeneous network using graph embedding
abstract
The discovery of gene-disease associations is important for the prevention, diagnosis and treatment of diseases. The studies on gene-disease associations have produced diverse data, which can facilitate the gene-disease association prediction. Integrating diverse information is critical for developing high-accuracy prediction models. In this paper, we propose a heterogeneous network-based method that enhances gene-disease association prediction by using graph embedding and ensemble learning, abbreviated as “HNEEM”. A heterogeneous network is constructed based on gene-disease associations, gene-chemical associations and disease-chemical associations, to combine diverse information. The network uses genes, diseases and chemicals as nodes, and uses their associations as edges. The graph embedding methods are utilized to extract representation vectors of nodes in the heterogeneous network, and the feature vectors of genes and diseases are merged to represent gene-disease pairs, and the random forest is employed to build the prediction model based on gene-disease pairs. We consider six types of graph embedding methods, and take the individual graph embedding method-generated features to build prediction models and use them as base predictors, and then combine base predictors to develop the ensemble learning model HNEEM. We comprehensively compare different graph embedding methods, and results demonstrate that the graph embedding methods produce satisfying results in the gene-disease association prediction, and integrating different graph embedding methods can make further improvements. In computational experiments, HNEEM produces better results compared to the state-of-the-art gene-disease perdition methods, and HNEEM is robust to the data richness as well. Moreover, the usefulness of the proposed method HNEEM is validated by the case studies. In conclusion, HNEEM is a promising method for predicting gene-disease associations.
Xiaochan Wang, Yuchong Gong, Jing Yi, Wen Zhang 0008
BIBM4
2019 LncPred-IEL: A Long Non-coding RNA Prediction Method using Iterative Ensemble Learning
abstract
A large number of transcripts have been generated by the development of high throughput sequencing technologies. Predicting lncRNA from transcripts is a challenging and important task. In this paper, we propose LncPred-IEL, an iterative ensemble learning long non-coding RNA prediction method. LncPred-IEL not only considers features widely used for the lncRNA prediction, but also take into account sequence-derived features used in the RNA sequence classification, so as to make use of diverse information. LncPred-IEL builds base predictors based on different groups of features, and employs a supervised iterative way to combine base predictors and build ensemble models. Our studies demonstrate that supervised iterative way can learn the representations that help to separate lncRNA and protein-coding transcripts, and further improve the performances. Experiments demonstrate that LncPred-IEL outperforms several state-of-the-art methods when evaluated by 10-fold cross-validation. The capability of LncPred-IEL for the cross-species prediction is also tested. As complementary to wet experiments, LncPred-IEL is a useful computational tool for lncRNA prediction.
Yanzhen Xu, Xiaohan Zhao, Shuai Liu 0017, Shichao Liu 0002, Yanqing Niu, Wen Zhang 0008, Leyi Wei
BIBM6
2019 LncRNA-miRNA interaction prediction from the heterogeneous network through graph embedding ensemble learning
abstract
LncRNA-miRNA interactions play crucial roles in gene regulatory networks and can reveal functions of lncRNAs and miRNAs. Although several methods have been proposed to infer interactions on the lncRNA-miRNA interaction network, few attentions have been paid to fully exploiting the structure of lncRNA-miRNA interaction network. In this paper, we propose a Graph Embedding Ensemble Learning method (abbreviated as “GEEL”) to predict lncRNA-miRNA interactions. First, we collect lncRNA sequences and miRNA sequences to calculate lncRNA-lncRNA sequence similarity and miRNA-miRNA sequence similarity, and then we combine them with the known lncRNA-miRNA interactions to construct a heterogeneous network, which takes lncRNAs and miRNAs as nodes. We adopt graph embedding methods to learn representations of lncRNAs and miRNAs from the heterogeneous network, and then merge the representations of lncRNAs and miRNAs to represent the lncRNA-miRNA pairs. Random forest classifiers are built based on the merged representations to predict lncRNA-miRNA interactions. We consider five different graph embedding methods, and evaluate the corresponding models. Further, we use individual graph embedding method-based model as base predictors and build a high-level ensemble model. The experimental results show that GEEL achieves AUPR score of 0.7004 and AUC score of 0.9537, and outperforms base predictors and other state-of-the-art methods. In conclusion, GEEL is an effective tool for lncRNA-miRNA interaction prediction.
Shuang Zhou 0009, Xiang Yue, Xinran Xu, Shichao Liu 0002, Wen Zhang 0008, Yanqing Niu
BIBM5
2019 Efficient Network Representations Learning: An Edge-Centric Perspective
Shichao Liu 0002, Shuangfei Zhai, Lida Zhu, Fuxi Zhu, Zhongfei Zhang, Wen Zhang 0008
KSEM (2)6
2019 A network embedding-based multiple information integration method for the MiRNA-disease association prediction
abstract
BACKGROUND: MiRNAs play significant roles in many fundamental and important biological processes, and predicting potential miRNA-disease associations makes contributions to understanding the molecular mechanism of human diseases. Existing state-of-the-art methods make use of miRNA-target associations, miRNA-family associations, miRNA functional similarity, disease semantic similarity and known miRNA-disease associations, but the known miRNA-disease associations are not well exploited. RESULTS: In this paper, a network embedding-based multiple information integration method (NEMII) is proposed for the miRNA-disease association prediction. First, known miRNA-disease associations are formulated as a bipartite network, and the network embedding method Structural Deep Network Embedding (SDNE) is adopted to learn embeddings of nodes in the bipartite network. Second, the embedding representations of miRNAs and diseases are combined with biological features about miRNAs and diseases (miRNA-family associations and disease semantic similarities) to represent miRNA-disease pairs. Third, the prediction models are constructed based on the miRNA-disease pairs by using the random forest. In computational experiments, NEMII achieves high-accuracy performances and outperforms other state-of-the-art methods: GRNMF, NTSMDA and PBMDA. The usefulness of NEMII is further validated by case studies. The studies demonstrate the great potential of network embedding method for the miRNA-disease association prediction, and SDNE outperforms other popular network embedding methods: DeepWalk, High-Order Proximity preserved Embedding (HOPE) and Laplacian Eigenmaps (LE). CONCLUSION: We propose a new method, named NEMII, for predicting miRNA-disease associations, which has great potential to benefit the field of miRNA-disease association prediction.
Yuchong Gong, Yanqing Niu, Wen Zhang 0008, Xiaohong Li 0003
BMC Bioinform.3
2019 SFLLN: A sparse feature learning ensemble method with linear neighborhood regularization for predicting drug-drug interactions
Wen Zhang 0008, Kanghong Jing, Feng Huang 0004, Yanlin Chen 0002, Bolin Li
Inf. Sci.1
2018 HNGRNMF: Heterogeneous Network-based Graph Regularized Nonnegative Matrix Factorization for predicting events of microbe-disease associations
Wen Zhang 0008, Xiaoting Lu, Weitai Yang, Feng Huang 0004, Binlu Wang, Qi Zhao 0010
BIBM1
2018 Prediction of Drug-Disease Associations and Their Effects by Signed Network-Based Nonnegative Matrix Factorization
Wen Zhang 0008, Feng Huang 0004, Xiang Yue, Xiaoting Lu, Weitai Yang, Zhishuai Li
BIBM1
2018 Sequence-derived linear neighborhood propagation method for predicting lncRNA-miRNA interactions
Wen Zhang 0008, Guifeng Tang, Siman Wang, Yanlin Chen 0002, Shuang Zhou 0008, Xiaohong Li 0003
BIBM1
2018 Sequence-based bacterial small RNAs prediction using ensemble learning strategies
abstract
BACKGROUND: Bacterial small non-coding RNAs (sRNAs) have emerged as important elements in diverse physiological processes, including growth, development, cell proliferation, differentiation, metabolic reactions and carbon metabolism, and attract great attention. Accurate prediction of sRNAs is important and challenging, and helps to explore functions and mechanism of sRNAs. RESULTS: In this paper, we utilize a variety of sRNA sequence-derived features to develop ensemble learning methods for the sRNA prediction. First, we compile a balanced dataset and four imbalanced datasets. Then, we investigate various sRNA sequence-derived features, such as spectrum profile, mismatch profile, reverse compliment k-mer and pseudo nucleotide composition. Finally, we consider two ensemble learning strategies to integrate all features for building ensemble learning models for the sRNA prediction. One is the weighted average ensemble method (WAEM), which uses the linear weighted sum of outputs from the individual feature-based predictors to predict sRNAs. The other is the neural network ensemble method (NNEM), which trains a deep neural network by combining diverse features. In the computational experiments, we evaluate our methods on these five datasets by using 5-fold cross validation. WAEM and NNEM can produce better results than existing state-of-the-art sRNA prediction methods. CONCLUSIONS: WAEM and NNEM have great potential for the sRNA prediction, and are helpful for understanding the biological mechanism of bacteria.
Guifeng Tang, Jingwen Shi, Wenjian Wu, Xiang Yue, Wen Zhang 0008
BMC Bioinform.5
2018 Predicting drug-disease associations by using similarity constrained matrix factorization
abstract
BACKGROUND: Drug-disease associations provide important information for the drug discovery. Wet experiments that identify drug-disease associations are time-consuming and expensive. However, many drug-disease associations are still unobserved or unknown. The development of computational methods for predicting unobserved drug-disease associations is an important and urgent task. RESULTS: In this paper, we proposed a similarity constrained matrix factorization method for the drug-disease association prediction (SCMFDD), which makes use of known drug-disease associations, drug features and disease semantic information. SCMFDD projects the drug-disease association relationship into two low-rank spaces, which uncover latent features for drugs and diseases, and then introduces drug feature-based similarities and disease semantic similarity as constraints for drugs and diseases in low-rank spaces. Different from the classic matrix factorization technique, SCMFDD takes the biological context of the problem into account. In computational experiments, the proposed method can produce high-accuracy performances on benchmark datasets, and outperform existing state-of-the-art prediction methods when evaluated by five-fold cross validation and independent testing. CONCLUSION: We developed a user-friendly web server by using known associations collected from the CTD database, available at http://www.bioinfotech.cn/SCMFDD/ . The case studies show that the server can find out novel associations, which are not included in the CTD database.
Wen Zhang 0008, Xiang Yue, Weiran Lin, Wenjian Wu, Ruoqi Liu, Feng Huang 0004
BMC Bioinform.1
2018 Feature-derived graph regularized matrix factorization for predicting drug side effects
Wen Zhang 0008, Yanlin Chen 0002, Wenjian Wu, Wei Wang 0121, Xiaohong Li 0003
Neurocomputing1
2018 The linear neighborhood propagation method for predicting long non-coding RNA-protein interactions
Wen Zhang 0008, Qianlong Qu, Yunqiu Zhang, Wei Wang 0121
Neurocomputing1
2018 Manifold regularized matrix factorization for drug-drug interaction prediction
Wen Zhang 0008, Yanlin Chen 0002, Dingfang Li, Xiang Yue
J. Biomed. Informatics1
2018 SFPEL-LPI: Sequence-based feature projection ensemble learning for predicting LncRNA-protein interactions
abstract
LncRNA-protein interactions play important roles in post-transcriptional gene regulation, poly-adenylation, splicing and translation. Identification of lncRNA-protein interactions helps to understand lncRNA-related activities. Existing computational methods utilize multiple lncRNA features or multiple protein features to predict lncRNA-protein interactions, but features are not available for all lncRNAs or proteins; most of existing methods are not capable of predicting interacting proteins (or lncRNAs) for new lncRNAs (or proteins), which don't have known interactions. In this paper, we propose the sequence-based feature projection ensemble learning method, "SFPEL-LPI", to predict lncRNA-protein interactions. First, SFPEL-LPI extracts lncRNA sequence-based features and protein sequence-based features. Second, SFPEL-LPI calculates multiple lncRNA-lncRNA similarities and protein-protein similarities by using lncRNA sequences, protein sequences and known lncRNA-protein interactions. Then, SFPEL-LPI combines multiple similarities and multiple features with a feature projection ensemble learning frame. In computational experiments, SFPEL-LPI accurately predicts lncRNA-protein associations and outperforms other state-of-the-art methods. More importantly, SFPEL-LPI can be applied to new lncRNAs (or proteins). The case studies demonstrate that our method can find out novel lncRNA-protein interactions, which are confirmed by literature. Finally, we construct a user-friendly web server, available at http://www.bioinfotech.cn/SFPEL-LPI/.
Wen Zhang 0008, Xiang Yue, Guifeng Tang, Wenjian Wu, Feng Huang 0004
PLoS Comput. Biol.1
2017 Predicting small RNAs in bacteria via sequence learning ensemble method
abstract
Bacterial small non-coding RNAs (sRNAs) play important roles in various physiological processes, and predicting sRNAs is an important task. In this paper, we develop a computational method for the sRNA prediction by using sRNA sequence-derived features. We investigate a variety of sRNA sequence-derived features, and evaluate the usefulness of features for the sRNA prediction. Then, we develop the sequence learning ensemble method, which uses the linear weighted sum of outputs from the individual feature-based predictors to predict sRNAs, and the genetic algorithm is adopted to optimize the parameters in the ensemble system. In the computational experiments, we compile a balanced dataset and four imbalanced datasets, and evaluate our method on these datasets by using 5-fold cross validation. The sequence learning ensemble method can achieve AUC scores greater than 0.9, and outperforms existing state-of-the-art sRNA prediction methods. In conclusion, the proposed method has a great potential for sRNA prediction. The source codes, datasets and supplementary are available in http://www.bioinfotech.cn/BIBM2017/SLEM.
Wen Zhang 0008, Jingwen Shi, Guifeng Tang, Wenjian Wu, Xiang Yue, Dingfang Li
BIBM1
2017 Predicting drug-disease associations based on the known association bipartite network
abstract
Recent studies show that drug-disease associations provide important information for drug discovery and drug repositioning. Wet experimental identification of drug-disease associations is time-consuming and labor-intensive. Therefore, the development of computational methods that predict drug-disease associations is an urgent task. In this paper, we propose a novel computational method named NTSIM, which only uses known drug-disease associations to predict unobserved associations. First of all, known drug-disease associations are represented as a drug-disease bipartite network, and a novel similarity measure named linear neighborhood similarity (LNS) is proposed to calculate drug-drug similarity and disease-disease similarity based on the bipartite network. Then, we predict unobserved drug-disease associations in the similarity-based graph by using label propagation process. In the computational experiments, this proposed method achieves high-accuracy performances, and outperforms representative state-of-the-art methods: PREDICT, TL-HGBI and LRSSL. Our studies reveal that known drug-disease associations can provide enough information to build the high-accuracy prediction models; linear neighbor similarity (LNS) can lead to better performances than other similarity measures such as Jaccard similarity, Gauss similarity and cosine similarity; the bipartite network-derived features outperform the drug biological features and disease semantic features.
Wen Zhang 0008, Xiang Yue, Yanlin Chen 0002, Weiran Lin, Bolin Li, Xiaohong Li 0003
BIBM1
2017 Predicting potential drug-drug interactions by integrating chemical, biological, phenotypic and network data
abstract
BACKGROUND: Drug-drug interactions (DDIs) are one of the major concerns in drug discovery. Accurate prediction of potential DDIs can help to reduce unexpected interactions in the entire lifecycle of drugs, and are important for the drug safety surveillance. RESULTS: Since many DDIs are not detected or observed in clinical trials, this work is aimed to predict unobserved or undetected DDIs. In this paper, we collect a variety of drug data that may influence drug-drug interactions, i.e., drug substructure data, drug target data, drug enzyme data, drug transporter data, drug pathway data, drug indication data, drug side effect data, drug off side effect data and known drug-drug interactions. We adopt three representative methods: the neighbor recommender method, the random walk method and the matrix perturbation method to build prediction models based on different data. Thus, we evaluate the usefulness of different information sources for the DDI prediction. Further, we present flexible frames of integrating different models with suitable ensemble rules, including weighted average ensemble rule and classifier ensemble rule, and develop ensemble models to achieve better performances. CONCLUSIONS: The experiments demonstrate that different data sources provide diverse information, and the DDI network based on known DDIs is one of most important information for DDI prediction. The ensemble methods can produce better performances than individual methods, and outperform existing state-of-the-art methods. The datasets and source codes are available at https://github.com/zw9977129/drug-drug-interaction/ .
Wen Zhang 0008, Yanlin Chen 0002, Fei Luo 0004, Gang Tian, Xiaohong Li 0003
BMC Bioinform.1
2017 Predicting human splicing branchpoints by combining sequence-derived features and multi-label learning methods
abstract
BACKGROUND: Alternative splicing is the critical process in a single gene coding, which removes introns and joins exons, and splicing branchpoints are indicators for the alternative splicing. Wet experiments have identified a great number of human splicing branchpoints, but many branchpoints are still unknown. In order to guide wet experiments, we develop computational methods to predict human splicing branchpoints. RESULTS: Considering the fact that an intron may have multiple branchpoints, we transform the branchpoint prediction as the multi-label learning problem, and attempt to predict branchpoint sites from intron sequences. First, we investigate a variety of intron sequence-derived features, such as sparse profile, dinucleotide profile, position weight matrix profile, Markov motif profile and polypyrimidine tract profile. Second, we consider several multi-label learning methods: partial least squares regression, canonical correlation analysis and regularized canonical correlation analysis, and use them as the basic classification engines. Third, we propose two ensemble learning schemes which integrate different features and different classifiers to build ensemble learning systems for the branchpoint prediction. One is the genetic algorithm-based weighted average ensemble method; the other is the logistic regression-based ensemble method. CONCLUSIONS: In the computational experiments, two ensemble learning methods outperform benchmark branchpoint prediction methods, and can produce high-accuracy results on the benchmark dataset.
Wen Zhang 0008, Junko Tsuji, Zhiping Weng
BMC Bioinform.1
2016 Drug side effect prediction through linear neighborhoods and multiple data source integration
abstract
Predicting drug side effects is a critical task in the drug discovery, which attracts great attentions in both academy and industry. Although lots of machine learning methods have been proposed, great challenges arise with boom of precision medicine. On one hand, many methods are based on the assumption that similar drugs may share same side effects, but measuring the drug-drug similarity appropriately is challenging. One the other hand, multi-source data provide diverse information for the analysis of side effects, and should be integrated for the high-accuracy prediction. In this paper, we tackle the side effect prediction problem through linear neighborhoods and multi-source data integration. In the feature space, linear neighborhoods are constructed to extract the drug-drug similarity, namely “linear neighborhood similarity”. By transferring the similarity into the side effect space, known side effect information is propagated through the similarity-based graph. Thus, we propose the linear neighborhood similarity method (LNSM), which utilizes single-source data for the side effect prediction. Further, we extend LNSM to deal with multi-source data, and propose two data integration methods: similarity matrix integration method (LNSM-SMI) and cost minimization integration method (LNSM-CMI), which integrate drug substructure data, drug target data, drug transporter data, drug enzyme data, drug pathway data and drug indication data to improve the prediction accuracy. The proposed methods are evaluated on the benchmark datasets. The linear neighborhood similarity method (LNSM) can produce satisfying results on the single-source data. Data integration methods (LNSM-SMI and LNSM-CMI) can effectively integrate multi-source data, and outperform other state-of-the-art side effect prediction methods in the cross validation and independent test. The proposed methods are promising for the drug side effect prediction.
Wen Zhang 0008, Yanlin Chen 0002, Shikui Tu, Qianlong Qu
BIBM1
2016 The prediction of human splicing branchpoints by multi-label learning
abstract
human splicing branchpoints are functional elements of the alternative splicing, and the study on branchpoints can help to understand the mechanism of human pre-mRNA transcript. There are a large number of human splicing branchpoints, but the wet methods that identify branchpoints are labor-intensive and time-consuming. In this paper, we utilize machine learning techniques to build models for the human branchpoint prediction. Since an intron may have multiple branchpoints, we formulate the original problem as a multi-label learning task, which predicts branchpoint sites of introns based on the characteristics of introns. First of all, we extract a diversity of intron sequence-derived features, including sparse profile, dinucleotide profile, position weight matrix profile, Markov motif profile, and polypyrimidine tract profile. Then, taking into account efficiency and effectiveness, we adopt three methods: partial least squares regression, canonical correlation analysis and regularized canonical correlation analysis, to build multi-label prediction models from different angles, by using intron sequence-derived features. Finally, we adopt the average scoring ensemble strategy to integrate different models, and develop the ensemble model for the branchpoint prediction. Computational experiments demonstrate that the proposed method can produce satisfying results on the experimentally verified dataset, and outperform other state-of-the-art methods. We develop a user-friendly web server for the human splicing branchpoint prediction, available at http://121.42.59.182:8080.
Wen Zhang 0008, Junko Tsuji, Zhiping Weng
BIBM1
2016 Multi-Domain Manifold Learning for Drug-Target Interaction Prediction
abstract
Drug-target interaction (DTI) provides novel insights about the genomic drug discovery, and is a critical technique to drug discovery. Recently, researchers try to incorporate different information about drugs and targets for prediction. However, the heterogeneous and high-dimensional data poses huge challenge to existing machine learning methods. In the last few years, extensive research efforts have been devoted to the utilization of manifold property on high dimensional data, e.g. dimension reduction methods preserving local structures of the manifolds. Motivated by the successes of these studies, we propose a general framework incorporating both manifold structures and known interaction/non-interaction information to predict the drug-target interactions. To overcome the challenges of domain scaling and information inconsistency, we formulate the problem with Semidefinite Programming (SDP), including new constraints to improve the robustness of the learning procedure. A variety of optimization techniques are also designed to enhance the scalability of the problem solver. Effectiveness of the method is evaluated by experiments on the benchmark dataset. Compared with state-of-the-art methods, the proposed methods generate much more accurate drug-target interaction prediction.
Ruichu Cai, Srinivasan Parthasarathy 0001, Anthony K. H. Tung, Wen Zhang 0008
SDM6
2016 A genetic algorithm-based weighted ensemble method for predicting transposon-derived piRNAs
abstract
BACKGROUND: Predicting piwi-interacting RNA (piRNA) is an important topic in the small non-coding RNAs, which provides clues for understanding the generation mechanism of gamete. To the best of our knowledge, several machine learning approaches have been proposed for the piRNA prediction, but there is still room for improvements. RESULTS: In this paper, we develop a genetic algorithm-based weighted ensemble method for predicting transposon-derived piRNAs. We construct datasets for three species: Human, Mouse and Drosophila. For each species, we compile the balanced dataset and imbalanced dataset, and thus obtain six datasets to build and evaluate prediction models. In the computational experiments, the genetic algorithm-based weighted ensemble method achieves 10-fold cross validation AUC of 0.932, 0.937 and 0.995 on the balanced Human dataset, Mouse dataset and Drosophila dataset, respectively, and achieves AUC of 0.935, 0.939 and 0.996 on the imbalanced datasets of three species. Further, we use the prediction models trained on the Mouse dataset to identify piRNAs of other species, and the models demonstrate the good performances in the cross-species prediction. CONCLUSIONS: Compared with other state-of-the-art methods, our method can lead to better performances. In conclusion, the proposed method is promising for the transposon-derived piRNA prediction. The source codes and datasets are available in https://github.com/zw9977129/piRNAPredictor .
Dingfang Li, Longqiang Luo, Wen Zhang 0008, Fei Luo 0004
BMC Bioinform.3
2016 Predicting potential side effects of drugs by recommender methods and ensemble learning
Wen Zhang 0008, Hua Zou 0002, Longqiang Luo, Qianchao Liu, Weijian Wu, Wenyi Xiao
Neurocomputing1
2015 Predicting drug side effects by multi-label learning and ensemble learning
abstract
BACKGROUND: Predicting drug side effects is an important topic in the drug discovery. Although several machine learning methods have been proposed to predict side effects, there is still space for improvements. Firstly, the side effect prediction is a multi-label learning task, and we can adopt the multi-label learning techniques for it. Secondly, drug-related features are associated with side effects, and feature dimensions have specific biological meanings. Recognizing critical dimensions and reducing irrelevant dimensions may help to reveal the causes of side effects. METHODS: In this paper, we propose a novel method 'feature selection-based multi-label k-nearest neighbor method' (FS-MLKNN), which can simultaneously determine critical feature dimensions and construct high-accuracy multi-label prediction models. RESULTS: Computational experiments demonstrate that FS-MLKNN leads to good performances as well as explainable results. To achieve better performances, we further develop the ensemble learning model by integrating individual feature-based FS-MLKNN models. When compared with other state-of-the-art methods, the ensemble method produces better performances on benchmark datasets. CONCLUSIONS: In conclusion, FS-MLKNN and the ensemble method are promising tools for the side effect prediction. The source code and datasets are available in the Additional file 1.
Wen Zhang 0008, Longqiang Luo, Jingxia Zhang
BMC Bioinform.1
2013 Predicting immunogenic T-cell epitopes by combining various sequence-derived features
abstract
The prediction of T-cell epitopes is of great help for facilitating vaccine design and understanding the immune system. In the bioinformatics, the MHC-binding peptides are defined as the T-cell epitopes, which will trigger the immune response to the antigens. However, binding peptides cannot necessarily activate the immune response, namely non-immunogenic. Until now, little attention has been paid to the immunogenic epitopes. Therefore, the recognition of immunogenic epitopes is a challenging task of the practical value. This paper systematically evaluates a wide variety of sequenced-derived features, which have been ever used for epitope prediction or similar tasks, and reveals their relationship with epitope immunogenicity. Then, we consider how to effectively exploit various features for the computational prediction of immunogenic epitopes. Subsequently, the random forest is adopted as the classification engine, and an ensemble model is developed by using the average scores of individual feature-based predictors. Compared with the previously published methods (POPI, POPISK and PAAQD), our models produce better performance on the benchmark datasets. Evaluated by t-test, the improvements of our method against existing methods are statistically significant (P<;0.01), showing the promise for the immunogenic epitope prediction. At present, only one MHC allele (HLA-A2) has sufficient data for the immunogenicity study. In the near future, with the increasing availability of immunogenic epitopes, we will carry out computational experiments on more MHC alleles. The source code for the ensemble model is available at: http://bcell.whu.edu.cn/sourcecode.html.
Wen Zhang 0008, Juan Liu 0007, Yi Xiong 0002, Meng Ke
BIBM1
2011 Prediction of Heme Binding Sites in Heme Proteins Using an Integrative Sequence Profile Coupling Evolutionary Information with Physicochemical Properties
abstract
Heme-protein interactions are essential for various biological processes such as electron transfer, catalysis, signal transduction and the control of gene expression. The knowledge of heme binding residues can provide crucial clues to understand the mechanism of heme- protein interactions and aid in functional annotation. In the present work, we propose a sequence-based approach for the accurate prediction of heme binding residues by a novel integrative sequence profile coupling position specific scoring matrices with heme specific physicochemical properties. Particularly, we design an intuitive feature selection scheme for informative physicochemical properties. As shown in the primary results, our integrative sequence profile approach for prediction of heme binding residues outperforms the conventional methods using amino acid and evolutionary information on the 5-fold cross validation and the independent test.
Yi Xiong 0002, Wen Zhang 0008, Tao Zeng 0003, Juan Liu 0007
BIBM2
2011 Prediction of conformational B-cell epitopes from 3D structures by random forest with a distance-based feature
abstract
BACKGROUND: Antigen-antibody interactions are key events in immune system, which provide important clues to the immune processes and responses. In Antigen-antibody interactions, the specific sites on the antigens that are directly bound by the B-cell produced antibodies are well known as B-cell epitopes. The identification of epitopes is a hot topic in bioinformatics because of their potential use in the epitope-based drug design. Although most B-cell epitopes are discontinuous (or conformational), insufficient effort has been put into the conformational epitope prediction, and the performance of existing methods is far from satisfaction. RESULTS: In order to develop the high-accuracy model, we focus on some possible aspects concerning the prediction performance, including the impact of interior residues, different contributions of adjacent residues, and the imbalanced data which contain much more non-epitope residues than epitope residues. In order to address above issues, we take following strategies. Firstly, a concept of 'thick surface patch' instead of 'surface patch' is introduced to describe the local spatial context of each surface residue, which considers the impact of interior residue. The comparison between the thick surface patch and the surface patch shows that interior residues contribute to the recognition of epitopes. Secondly, statistical significance of the distance distribution difference between non-epitope patches and epitope patches is observed, thus an adjacent residue distance feature is presented, which reflects the unequal contributions of adjacent residues to the location of binding sites. Thirdly, a bootstrapping and voting procedure is adopted to deal with the imbalanced dataset. Based on the above ideas, we propose a new method to identify the B-cell conformational epitopes from 3D structures by combining conventional features and the proposed feature, and the random forest (RF) algorithm is used as the classification engine. The experiments show that our method can predict conformational B-cell epitopes with high accuracy. Evaluated by leave-one-out cross validation (LOOCV), our method achieves the mean AUC value of 0.633 for the benchmark bound dataset, and the mean AUC value of 0.654 for the benchmark unbound dataset. When compared with the state-of-the-art prediction models in the independent test, our method demonstrates comparable or better performance. CONCLUSIONS: Our method is demonstrated to be effective for the prediction of conformational epitopes. Based on the study, we develop a tool to predict the conformational epitopes from 3D structures, available at http://code.google.com/p/my-project-bpredictor/downloads/list.
Wen Zhang 0008, Yi Xiong 0002, Hua Zou 0002, Xinghuo Ye, Juan Liu 0007
BMC Bioinform.1
2010 Quantitative prediction of MHC-II binding affinity using particle swarm optimization
Wen Zhang 0008, Juan Liu 0007, Yanqing Niu
Artif. Intell. Medicine1
2009 Quantitative prediction of MHC-II peptide binding affinity using relevance vector machine
Wen Zhang 0008, Juan Liu 0007, Yanqing Niu
Appl. Intell.1