EDBT 2026 Demo / reviewers in the wild / expert
Yun Tang 0001
dblp:67/764-1
· DBLP profile ↗
14ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0003-2340-1109ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BBANsh: a deep learning architecture based on BERT and bilinear attention networks to identify potent shRNAabstractRNA interference (RNAi) is a technique for precisely silencing the expression of specific genes by means of small RNA molecules and is essential in functional genomics. Among the commonly used RNAi molecules, short hairpin RNAs (shRNAs) exhibit advantages over small interfering RNAs, including longer half-life, comparable silencing efficiency, fewer off-target effects, and greater safety. However, traditional screening of potent shRNAs is costly and time-consuming. Advances in big data and artificial intelligence have enabled computational methods to significantly accelerate shRNA design and prediction. In this study, we propose BBANsh, a new shRNA prediction model based on bidirectional encoder representation from transformers (BERT) and bilinear attention network (BAN). We comprehensively evaluate the performance of BBANsh against traditional feature-based models, various feature fusion methods, and existing shRNA prediction models. The BBANsh has achieved an area under the precision-recall curve of 0.951 on five-cross validation and a prediction accuracy of 0.896 on a new external validation set, highlighting its superior predictive performance. Ablation experiments validate the significant contributions of BERT and BAN to model performance. The visualization of internal feature representations intuitively demonstrates the effectiveness of the feature fusion strategy of BBANsh. Furthermore, the attentional analysis reveals that nucleotides near the 5' end have the greatest impact on model predictions, highlighting sequence characteristics of potent shRNAs. Overall, BBANsh provides an efficient and reliable tool for shRNA prediction, which can offer valuable support for researchers in the precise selection and design of shRNA. Yuanting Chen, Weihua Li 0005, Yun Tang 0001, Guixia Liu |
Briefings Bioinform. | 5 |
| 2025 | miCGR: interpretable deep neural network for predicting both site-level and gene-level functional targets of microRNAabstractMicroRNAs (miRNAs) are critical regulators in various biological processes to cleave or repress translation of messenger RNAs (mRNAs). Accurately predicting miRNA targets is essential for developing miRNA-based therapies for diseases such as cancer and cardiovascular disease. Traditional miRNA target prediction methods often struggle due to incomplete knowledge of miRNA-target interactions and lack interpretability. To address these limitations, we propose miCGR, an end-to-end deep learning framework for predicting functional miRNA targets. MiCGR employs 2D convolutional neural networks alongside an enhanced Chaos Game Representation (CGR) of both miRNA sequences and their candidate target site (CTS) on mRNA. This advanced CGR transforms genetic sequences into informative 2D graphical representations based on sequence composition and subsequence frequencies, and explicitly incorporates important prior knowledge of seed regions and subsequence positions. Unlike one-dimensional methods based solely on sequence characters, this approach identifies functional motifs within sequences, even if they are distant in the original sequences. Our model outperforms existing methods in predicting functional targets at both the site and gene levels. To enhance interpretability, we incorporate Shapley value analysis for each subsequence within both miRNA sequences and their target sites, allowing miCGR to achieve improved accuracy, particularly with more lenient CTS selection criteria. Finally, two case studies demonstrate the practical applicability of miCGR, highlighting its potential to provide insights for optimizing artificial miRNA analogs that surpass endogenous counterparts. Lehan Zhang, Xiaochu Tong, Yitian Wang, Zimei Zhang, Xiangtai Kong, Shengkun Ni, Xiaomin Luo, Mingyue Zheng, Yun Tang 0001, Xutong Li |
Briefings Bioinform. | 10 |
| 2024 | Herb-CMap: a multimodal fusion framework for deciphering the mechanisms of action in traditional Chinese medicine using Suhuang antitussive capsule as a case studyabstractHerbal medicines, particularly traditional Chinese medicines (TCMs), are a rich source of natural products with significant therapeutic potential. However, understanding their mechanisms of action is challenging due to the complexity of their multi-ingredient compositions. We introduced Herb-CMap, a multimodal fusion framework leveraging protein-protein interactions and herb-perturbed gene expression signatures. Utilizing a network-based heat diffusion algorithm, Herb-CMap creates a connectivity map linking herb perturbations to their therapeutic targets, thereby facilitating the prioritization of active ingredients. As a case study, we applied Herb-CMap to Suhuang antitussive capsule (Suhuang), a TCM formula used for treating cough variant asthma (CVA). Using in vivo rat models, our analysis established the transcriptomic signatures of Suhuang and identified its key compounds, such as quercetin and luteolin, and their target genes, including IL17A, PIK3CB, PIK3CD, AKT1, and TNF. These drug-target interactions inhibit the IL-17 signaling pathway and deactivate PI3K, AKT, and NF-κB, effectively reducing lung inflammation and alleviating CVA. The study demonstrates the efficacy of Herb-CMap in elucidating the molecular mechanisms of herbal medicines, offering valuable insights for advancing drug discovery in TCM. Yinyin Wang, Yihang Sui, Qimeng Tian, Yun Tang 0001, Yongyu Ou, Jing Tang 0002, Ninghua Tan |
Briefings Bioinform. | 6 |
| 2024 | ToxGIN: an In silico prediction model for peptide toxicity via graph isomorphism networks integrating peptide sequence and structure informationabstractPeptide drugs have demonstrated enormous potential in treating a variety of diseases, yet toxicity prediction remains a significant challenge in drug development. Existing models for prediction of peptide toxicity largely rely on sequence information and often neglect the three-dimensional (3D) structures of peptides. This study introduced a novel model for short peptide toxicity prediction, named ToxGIN. The model utilizes Graph Isomorphism Network (GIN), integrating the underlying amino acid sequence composition and the 3D structures of peptides. ToxGIN comprises three primary modules: (i) Sequence processing module, converting peptide 3D structures and sequences into information of nodes and edges; (ii) Feature extraction module, utilizing GIN to learn discriminative features from nodes and edges; (iii) Classification module, employing a fully connected classifier for toxicity prediction. ToxGIN performed well on the independent test set with F1 score = 0.83, AUROC = 0.91, and Matthews correlation coefficient = 0.68, better than existing models for prediction of peptide toxicity. These results validated the effectiveness of integrating 3D structural information with sequence data using GIN for peptide toxicity prediction. The proposed ToxGIN and data can be freely accessible at https://github.com/cihebiyql/ToxGIN. Qiule Yu, Guixia Liu, Weihua Li 0005, Yun Tang 0001 |
Briefings Bioinform. | 5 |
| 2024 | MetaPredictor: in silico prediction of drug metabolites based on deep language models with prompt engineeringabstractMetabolic processes can transform a drug into metabolites with different properties that may affect its efficacy and safety. Therefore, investigation of the metabolic fate of a drug candidate is of great significance for drug discovery. Computational methods have been developed to predict drug metabolites, but most of them suffer from two main obstacles: the lack of model generalization due to restrictions on metabolic transformation rules or specific enzyme families, and high rate of false-positive predictions. Here, we presented MetaPredictor, a rule-free, end-to-end and prompt-based method to predict possible human metabolites of small molecules including drugs as a sequence translation problem. We innovatively introduced prompt engineering into deep language models to enrich domain knowledge and guide decision-making. The results showed that using prompts that specify the sites of metabolism (SoMs) can steer the model to propose more accurate metabolite predictions, achieving a 30.4% increase in recall and a 16.8% reduction in false positives over the baseline model. The transfer learning strategy was also utilized to tackle the limited availability of metabolic data. For the adaptation to automatic or non-expert prediction, MetaPredictor was designed as a two-stage schema consisting of automatic identification of SoMs followed by metabolite prediction. Compared to four available drug metabolite prediction tools, our method showed comparable performance on the major enzyme families and better generalization that could additionally identify metabolites catalyzed by less common enzymes. The results indicated that MetaPredictor could provide a more comprehensive and accurate prediction of drug metabolism through the effective combination of transfer learning and prompt-based learning strategies. Keyun Zhu, Mengting Huang, Yaxin Gu, Weihua Li 0005, Guixia Liu, Yun Tang 0001 |
Briefings Bioinform. | 7 |
| 2023 | Identification of vital chemical information via visualization of graph neural networksabstractQualitative or quantitative prediction models of structure-activity relationships based on graph neural networks (GNNs) are prevalent in drug discovery applications and commonly have excellently predictive power. However, the network information flows of GNNs are highly complex and accompanied by poor interpretability. Unfortunately, there are relatively less studies on GNN attributions, and their developments in drug research are still at the early stages. In this work, we adopted several advanced attribution techniques for different GNN frameworks and applied them to explain multiple drug molecule property prediction tasks, enabling the identification and visualization of vital chemical information in the networks. Additionally, we evaluated them quantitatively with attribution metrics such as accuracy, sparsity, fidelity and infidelity, stability and sensitivity; discussed their applicability and limitations; and provided an open-source benchmark platform for researchers. The results showed that all attribution techniques were effective, while those directly related to the predicted labels, such as integrated gradient, preferred to have better attribution performance. These attribution techniques we have implemented could be directly used for the vast majority of chemical GNN interpretation tasks. Mengting Huang, Weihua Li 0005, Zengrui Wu, Yun Tang 0001, Guixia Liu |
Briefings Bioinform. | 6 |
| 2022 | Profiling prediction of nuclear receptor modulators with multi-task deep learning methods: toward the virtual screeningabstractNuclear receptors (NRs) are ligand-activated transcription factors, which constitute one of the most important targets for drug discovery. Current computational strategies mainly focus on a single target, and the transfer of learned knowledge among NRs was not considered yet. Herein we proposed a novel computational framework named NR-Profiler for prediction of potential NR modulators with high affinity and specificity. First, we built a comprehensive NR data set including 42 684 interactions to connect 42 NRs and 31 033 compounds. Then, we used multi-task deep neural network and multi-task graph convolutional neural network architectures to construct multi-task multi-classification models. To improve the predictive capability and robustness, we built a consensus model with an area under the receiver operating characteristic curve (AUC) = 0.883. Compared with conventional machine learning and structure-based approaches, the consensus model showed better performance in external validation. Using this consensus model, we demonstrated the practical value of NR-Profiler in virtual screening for NRs. In addition, we designed a selectivity score to quantitatively measure the specificity of NR modulators. Finally, we developed a freely available standalone software for users to make profiling predictions for their compounds of interest. In summary, our NR-Profiler provides a useful tool for NR-profiling prediction and is expected to facilitate NR-based drug discovery. Jiye Wang, Chaofeng Lou, Guixia Liu, Weihua Li 0005, Zengrui Wu, Yun Tang 0001 |
Briefings Bioinform. | 6 |
| 2022 | ADENet: a novel network-based inference method for prediction of drug adverse eventsabstractIdentification of adverse drug events (ADEs) is crucial to reduce human health risks and improve drug safety assessment. With an increasing number of biological and medical data, computational methods such as network-based methods were proposed for ADE prediction with high efficiency and low cost. However, previous network-based methods rely on the topological information of known drug-ADE networks, and hence cannot make predictions for novel compounds without any known ADE. In this study, we introduced chemical substructures to bridge the gap between the drug-ADE network and novel compounds, and developed a novel network-based method named ADENet, which can predict potential ADEs for not only drugs within the drug-ADE network, but also novel compounds outside the network. To show the performance of ADENet, we collected drug-ADE associations from a comprehensive database named MetaADEDB and constructed a series of network-based prediction models. These models obtained high area under the receiver operating characteristic curve values ranging from 0.871 to 0.947 in 10-fold cross-validation. The best model further showed high performance in external validation, which outperformed a previous network-based and a recent deep learning-based method. Using several approved drugs as case studies, we found that 32-54% of the predicted ADEs can be validated by the literature, indicating the practical value of ADENet. Moreover, ADENet is freely available at our web server named NetInfer (http://lmmd.ecust.edu.cn/netinfer). In summary, our method would provide a promising tool for ADE prediction and drug safety assessment in drug discovery and development. Zhuohang Yu, Zengrui Wu, Weihua Li 0005, Guixia Liu, Yun Tang 0001 |
Briefings Bioinform. | 5 |
| 2021 | Drug repositioning by prediction of drug's anatomical therapeutic chemical code via network-based inference approachesabstractDrug discovery and development is a time-consuming and costly process. Therefore, drug repositioning has become an effective approach to address the issues by identifying new therapeutic or pharmacological actions for existing drugs. The drug's anatomical therapeutic chemical (ATC) code is a hierarchical classification system categorized as five levels according to the organs or systems that drugs act and the pharmacology, therapeutic and chemical properties of drugs. The 2nd-, 3rd- and 4th-level ATC codes reserved the therapeutic and pharmacological information of drugs. With the hypothesis that drugs with similar structures or targets would possess similar ATC codes, we exploited a network-based approach to predict the 2nd-, 3rd- and 4th-level ATC codes by constructing substructure drug-ATC (SD-ATC), target drug-ATC (TD-ATC) and Substructure&Target drug-ATC (STD-ATC) networks. After 10-fold cross validation and two external validations, the STD-ATC models outperformed the SD-ATC and TD-ATC ones. Furthermore, with KR as fingerprint, the STD-ATC model was identified as the optimal model with AUC values at 0.899 ± 0.015, 0.916 and 0.893 for 10-fold cross validation, external validation set 1 and external validation set 2, respectively. To illustrate the predictive capability of the STD-ATC model with KR fingerprint, as a case study, we predicted 25 FDA-approved drugs (22 drugs were actually purchased) to have potential activities on heart failure using that model. Experiments in vitro confirmed that 8 of the 22 old drugs have shown mild to potent cardioprotective activities on both hypoxia model and oxygen-glucose deprivation model, which demonstrated that our STD-ATC prediction model would be an effective tool for drug repositioning. Yayuan Peng, Manjiong Wang, Yixiang Xu, Zengrui Wu, Jiye Wang, Guixia Liu, Weihua Li 0005, Yun Tang 0001 |
Briefings Bioinform. | 10 |
| 2021 | MetaADEDB 2.0: a comprehensive database on adverse drug eventsabstractSUMMARY: MetaADEDB is an online database we developed to integrate comprehensive information on adverse drug events (ADEs). The first version of MetaADEDB was released in 2013 and has been widely used by researchers. However, it has not been updated for more than seven years. Here, we reported its second version by collecting more and newer data from the U.S. FDA Adverse Event Reporting System (FAERS) and Canada Vigilance Adverse Reaction Online Database, in addition to the original three sources. The new version consists of 744 709 drug-ADE associations between 8498 drugs and 13 193 ADEs, which has an over 40% increase in drug-ADE associations compared to the previous version. Meanwhile, we developed a new and user-friendly web interface for data search and analysis. We hope that MetaADEDB 2.0 could provide a useful tool for drug safety assessment and related studies in drug discovery and development. AVAILABILITY AND IMPLEMENTATION: The database is freely available at: http://lmmd.ecust.edu.cn/metaadedb/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhuohang Yu, Zengrui Wu, Weihua Li 0005, Guixia Liu, Yun Tang 0001 |
Bioinform. | 5 |
| 2019 | admetSAR 2.0: web-service for prediction and optimization of chemical ADMET propertiesabstractSUMMARY: admetSAR was developed as a comprehensive source and free tool for the prediction of chemical ADMET properties. Since its first release in 2012 containing 27 predictive models, admetSAR has been widely used in chemical and pharmaceutical fields. This update, admetSAR 2.0, focuses on extension and optimization of existing models with significant quantity and quality improvement on training data. Now 47 models are available for either drug discovery or environmental risk assessment. In addition, we added a new module named ADMETopt for lead optimization based on predicted ADMET properties. AVAILABILITY AND IMPLEMENTATION: Free available on the web at http://lmmd.ecust.edu.cn/admetsar2/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hongbin Yang 0002, Chaofeng Lou, Lixia Sun, Yingchun Cai, Weihua Li 0005, Guixia Liu, Yun Tang 0001 |
Bioinform. | 9 |
| 2019 | Predicting Meridian in Chinese traditional medicine using machine learning approachesabstractPlant-derived nature products, known as herb formulas, have been commonly used in Traditional Chinese Medicine (TCM) for disease prevention and treatment. The herbs have been traditionally classified into different categories according to the TCM Organ systems known as Meridians. Despite the increasing knowledge on the active components of the herbs, the rationale of Meridian classification remains poorly understood. In this study, we took a machine learning approach to explore the classification of Meridian. We determined the molecule features for 646 herbs and their active components including structure-based fingerprints and ADME properties (absorption, distribution, metabolism and excretion), and found that the Meridian can be predicted by machine learning approaches with a top accuracy of 0.83. We also identified the top compound features that were important for the Meridian prediction. To the best of our knowledge, this is the first time that molecular properties of the herb compounds are associated with the TCM Meridians. Taken together, the machine learning approach may provide novel insights for the understanding of molecular evidence of Meridians in TCM. Yinyin Wang, Mohieddin Jafari, Yun Tang 0001, Jing Tang 0002 |
PLoS Comput. Biol. | 3 |
| 2017 | SDTNBI: an integrated network and chemoinformatics tool for systematic prediction of drug-target interactions and drug repositioningabstractComputational prediction of drug-target interactions (DTIs) and drug repositioning provides a low-cost and high-efficiency approach for drug discovery and development. The traditional social network-derived methods based on the naïve DTI topology information cannot predict potential targets for new chemical entities or failed drugs in clinical trials. There are currently millions of commercially available molecules with biologically relevant representations in chemical databases. It is urgent to develop novel computational approaches to predict targets for new chemical entities and failed drugs on a large scale. In this study, we developed a useful tool, namely substructure-drug-target network-based inference (SDTNBI), to prioritize potential targets for old drugs, failed drugs and new chemical entities. SDTNBI incorporates network and chemoinformatics to bridge the gap between new chemical entities and known DTI network. High performance was yielded in 10-fold and leave-one-out cross validations using four benchmark data sets, covering G protein-coupled receptors, kinases, ion channels and nuclear receptors. Furthermore, the highest areas under the receiver operating characteristic curve were 0.797 and 0.863 for two external validation sets, respectively. Finally, we identified thousands of new potential DTIs via implementing SDTNBI on a global network. As a proof-of-principle, we showcased the use of SDTNBI to identify novel anticancer indications for nonsteroidal anti-inflammatory drugs by inhibiting AKR1C3, CA9 or CA12. In summary, SDTNBI is a powerful network-based approach that predicts potential targets for new chemical entities on a large scale and will provide a new tool for DTI prediction and drug repositioning. The program and predicted DTIs are available on request. Zengrui Wu, Feixiong Cheng, Weihua Li 0005, Guixia Liu, Yun Tang 0001 |
Briefings Bioinform. | 6 |
| 2012 | Prediction of Drug-Target Interactions and Drug Repositioning via Network-Based InferenceabstractDrug-target interaction (DTI) is the basis of drug discovery and design. It is time consuming and costly to determine DTI experimentally. Hence, it is necessary to develop computational methods for the prediction of potential DTI. Based on complex network theory, three supervised inference methods were developed here to predict DTI and used for drug repositioning, namely drug-based similarity inference (DBSI), target-based similarity inference (TBSI) and network-based inference (NBI). Among them, NBI performed best on four benchmark data sets. Then a drug-target network was created with NBI based on 12,483 FDA-approved and experimental drug-target binary links, and some new DTIs were further predicted. In vitro assays confirmed that five old drugs, namely montelukast, diclofenac, simvastatin, ketoconazole, and itraconazole, showed polypharmacological features on estrogen receptors or dipeptidyl peptidase-IV with half maximal inhibitory or effective concentration ranged from 0.2 to 10 µM. Moreover, simvastatin and ketoconazole showed potent antiproliferative activities on human MDA-MB-231 breast cancer cell line in MTT assays. The results indicated that these methods could be powerful tools in prediction of DTIs and drug repositioning. Feixiong Cheng, Weiqiang Lu, Weihua Li 0005, Guixia Liu, Wei-Xing Zhou, Yun Tang 0001 |
PLoS Comput. Biol. | 9 |