Mingyue Zheng

dblp:02/4808 · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-3323-3092ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MolPIF: a parameter interpolation flow model for molecule generation
abstract
MOTIVATION: Structure-based drug design (SBDD) has advanced with deep generative models, but bridging the gap between continuous atomic coordinates and discrete atom types remains a challenge. Current approaches, such as diffusion and flow matching models, often fail to unify these heterogeneous modalities, relying on separate strategies or ill-fitting Euclidean metrics for discrete variables. This lack of a consistent framework limits generative models' ability to capture the geometric and chemical structure of protein-ligand complexes. RESULTS: We present MolPIF, a parameter interpolation flow mechanism designed to unify the generation of continuous and discrete molecular variables. Unlike traditional flow models that operate in sample space, MolPIF interpolates between distributions in the parameter space, theoretically recovering Wasserstein-2 optimal transport for continuous coordinates and establishing Fisher-Rao geodesics for discrete atom types. We further incorporate a geometry-enhanced learning strategy to improve the capture of atomic contexts. Extensive evaluations on the CrossDocked2020 dataset demonstrate that MolPIF outperforms baselines in binding affinity, chemical validity, geometric fidelity, and chemical space coverage. Additionally, MolPIF exhibits versatility in lead optimization and offers flexible prior distribution selection (such as Laplace), establishing a robust paradigm for SBDD. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at https://github.com/BLEACH366/MolPIF.
Yaowei Jin, Yufan Tang, Wenkai Xiang, Duanhua Cao, Dan Teng, Zhehuan Fan, Jiacheng Xiong, Xia Sheng, Chuanlong Zeng, Duo An, Mingyue Zheng, Shuangjia Zheng, Qian Shi 0005
Bioinform.12
2025 Piloting Structure-Based Drug Design via Modality-Specific Optimal Schedule
abstract
Structure-Based Drug Design (SBDD) is crucial for identifying bioactive molecules. Recent deep generative models are faced with challenges in geometric structure modeling. A major bottleneck lies in the twisted probability path of multi-modalities—continuous 3D positions and discrete 2D topologies—which jointly determine molecular geometries. By establishing the fact that noise schedules decide the Variational Lower Bound (VLB) for the twisted probability path, we propose VLB-Optimal Scheduling (VOS) strategy in this under-explored area, which optimizes VLB as a path integral for SBDD. Our model effectively enhances molecular geometries and interaction modeling, achieving state-of-the-art PoseBusters passing rate of 95.9\% on CrossDock, more than 10\% improvement upon strong baselines, while maintaining high affinities and robust intramolecular validity evaluated on held-out test set. Code is available at https://github.com/AlgoMole/MolCRAFT.
Keyue Qiu, Yuxuan Song 0002, Zhehuan Fan, Mingyue Zheng, Hao Zhou 0012, Wei-Ying Ma
ICML6
2025 Empower Structure-Based Molecule Optimization with Gradient Guided Bayesian Flow Networks
abstract
Structure-based molecule optimization (SBMO) aims to optimize molecules with both continuous coordinates and discrete types against protein targets. A promising direction is to exert gradient guidance on generative models given its remarkable success in images, but it is challenging to guide discrete data and risks inconsistencies between modalities. To this end, we leverage a continuous and differentiable space derived through Bayesian inference, presenting Molecule Joint Optimization (MolJO), the gradient-based SBMO framework that facilitates joint guidance signals across different modalities while preserving SE(3)-equivariance. We introduce a novel backward correction strategy that optimizes within a sliding window of the past histories, allowing for a seamless trade-off between explore-and-exploit during optimization. MolJO achieves state-of-the-art performance on CrossDocked2020 benchmark (Success Rate 51.3\%, Vina Dock -9.05 and SA 0.78), more than 4x improvement in Success Rate compared to the gradient-based counterpart, and 2x ``Me-Better'' Ratio as much as 3D baselines. Furthermore, we extend MolJO to a wide range of optimization settings, including multi-objective optimization and challenging tasks in drug design such as R-group optimization and scaffold hopping, further underscoring its versatility. Code is available at https://github.com/AlgoMole/MolCRAFT.
Keyue Qiu, Yuxuan Song 0002, Hongbo Ma, Ziyao Cao, Yushuai Wu, Mingyue Zheng, Hao Zhou 0012, Wei-Ying Ma
ICML8
2025 miCGR: interpretable deep neural network for predicting both site-level and gene-level functional targets of microRNA
abstract
MicroRNAs (miRNAs) are critical regulators in various biological processes to cleave or repress translation of messenger RNAs (mRNAs). Accurately predicting miRNA targets is essential for developing miRNA-based therapies for diseases such as cancer and cardiovascular disease. Traditional miRNA target prediction methods often struggle due to incomplete knowledge of miRNA-target interactions and lack interpretability. To address these limitations, we propose miCGR, an end-to-end deep learning framework for predicting functional miRNA targets. MiCGR employs 2D convolutional neural networks alongside an enhanced Chaos Game Representation (CGR) of both miRNA sequences and their candidate target site (CTS) on mRNA. This advanced CGR transforms genetic sequences into informative 2D graphical representations based on sequence composition and subsequence frequencies, and explicitly incorporates important prior knowledge of seed regions and subsequence positions. Unlike one-dimensional methods based solely on sequence characters, this approach identifies functional motifs within sequences, even if they are distant in the original sequences. Our model outperforms existing methods in predicting functional targets at both the site and gene levels. To enhance interpretability, we incorporate Shapley value analysis for each subsequence within both miRNA sequences and their target sites, allowing miCGR to achieve improved accuracy, particularly with more lenient CTS selection criteria. Finally, two case studies demonstrate the practical applicability of miCGR, highlighting its potential to provide insights for optimizing artificial miRNA analogs that surpass endogenous counterparts.
Lehan Zhang, Xiaochu Tong, Yitian Wang, Zimei Zhang, Xiangtai Kong, Shengkun Ni, Xiaomin Luo, Mingyue Zheng, Yun Tang 0001, Xutong Li
Briefings Bioinform.9
2024 Unified Generative Modeling of 3D Molecules with Bayesian Flow Networks
abstract
Advanced generative model (\textit{e.g.}, diffusion model) derived from simplified continuity assumptions of data distribution, though showing promising progress, has been difficult to apply directly to geometry generation applications due to the \textit{multi-modality} and \textit{noise-sensitive} nature of molecule geometry. This work introduces Geometric Bayesian Flow Networks (GeoBFN), which naturally fits molecule geometry by modeling diverse modalities in the differentiable parameter space of distributions. GeoBFN maintains the SE-(3) invariant density modeling property by incorporating equivariant inter-dependency modeling on parameters of distributions and unifying the probabilistic modeling of different modalities. Through optimized training and sampling techniques, we demonstrate that GeoBFN achieves state-of-the-art performance on multiple 3D molecule generation benchmarks in terms of generation quality (90.87\% molecule stability in QM9 and 85.6\% atom stability in GEOM-DRUG\footnote{The scores are reported at 1k sampling steps for fair comparison, and our scores could be further improved if sampling sufficiently longer steps.}). GeoBFN can also conduct sampling with any number of steps to reach an optimal trade-off between efficiency and quality (\textit{e.g.}, 20$\times$ speedup without sacrificing performance).
Yuxuan Song 0002, Jingjing Gong, Hao Zhou 0012, Mingyue Zheng, Wei-Ying Ma
ICLR4
2024 MolCRAFT: Structure-Based Drug Design in Continuous Parameter Space
abstract
Generative models for structure-based drug design (SBDD) have shown promising results in recent years. Existing works mainly focus on how to generate molecules with higher binding affinity, ignoring the feasibility prerequisites for generated 3D poses and resulting in false positives. We conduct thorough studies on key factors of ill-conformational problems when applying autoregressive methods and diffusion to SBDD, including mode collapse and hybrid continuous-discrete space. In this paper, we introduce MolCRAFT, the first SBDD model that operates in the continuous parameter space, together with a novel noise reduced sampling strategy. Empirical results show that our model consistently achieves superior performance in binding affinity with more stable 3D structure, demonstrating our ability to accurately model interatomic interactions. To our best knowledge, MolCRAFT is the first to achieve reference-level Vina Scores (-6.59 kcal/mol) with comparable molecular size, outperforming other strong baselines by a wide margin (-0.84 kcal/mol). Code is available at https://github.com/AlgoMole/MolCRAFT.
Yanru Qu, Keyue Qiu, Yuxuan Song 0002, Jingjing Gong, Jiawei Han 0001, Mingyue Zheng, Hao Zhou 0012, Wei-Ying Ma
ICML6
2024 KinomeMETA: meta-learning enhanced kinome-wide polypharmacology profiling
abstract
Kinase inhibitors are crucial in cancer treatment, but drug resistance and side effects hinder the development of effective drugs. To address these challenges, it is essential to analyze the polypharmacology of kinase inhibitor and identify compound with high selectivity profile. This study presents KinomeMETA, a framework for profiling the activity of small molecule kinase inhibitors across a panel of 661 kinases. By training a meta-learner based on a graph neural network and fine-tuning it to create kinase-specific learners, KinomeMETA outperforms benchmark multi-task models and other kinase profiling models. It provides higher accuracy for understudied kinases with limited known data and broader coverage of kinase types, including important mutant kinases. Case studies on the discovery of new scaffold inhibitors for membrane-associated tyrosine- and threonine-specific cdc2-inhibitory kinase and selective inhibitors for fibroblast growth factor receptors demonstrate the role of KinomeMETA in virtual screening and kinome-wide activity profiling. Overall, KinomeMETA has the potential to accelerate kinase drug discovery by more effectively exploring the kinase polypharmacology landscape.
Qun Ren, Ning Qu, Lin Ni, Xiaochu Tong, Zimei Zhang, Xiangtai Kong, Yiming Wen, Yitian Wang, Dingyan Wang, Xiaomin Luo, Sulin Zhang, Mingyue Zheng, Xutong Li
Briefings Bioinform.15
2024 TG468: a text graph convolutional network for predicting clinical response to immune checkpoint inhibitor therapy
abstract
Enhancing cancer treatment efficacy remains a significant challenge in human health. Immunotherapy has witnessed considerable success in recent years as a treatment for tumors. However, due to the heterogeneity of diseases, only a fraction of patients exhibit a positive response to immune checkpoint inhibitor (ICI) therapy. Various single-gene-based biomarkers and tumor mutational burden (TMB) have been proposed for predicting clinical responses to ICI; however, their predictive ability is limited. We propose the utilization of the Text Graph Convolutional Network (GCN) method to comprehensively assess the impact of multiple genes, aiming to improve the predictive capability for ICI response. We developed TG468, a Text GCN model framing drug response prediction as a text classification task. By combining natural language processing (NLP) and graph neural network techniques, TG468 effectively handles sparse and high-dimensional exome sequencing data. As a result, TG468 can distinguish survival time for patients who received ICI therapy and outperforms single gene biomarkers, TMB and some classical machine learning models. Additionally, TG468's prediction results facilitate the identification of immune status differences among specific patient types in the Cancer Genome Atlas dataset, providing a rationale for the model's predictions. Our approach represents a pioneering use of a GCN model to analyze exome data in patients undergoing ICI therapy and offers inspiration for future research using NLP technology to analyze exome sequencing data.
Jiangshan Shi, Xiaochu Tong, Ning Qu, Xiangtai Kong, Shengkun Ni, Xutong Li, Mingyue Zheng
Briefings Bioinform.9
2024 Gram matrix: an efficient representation of molecular conformation and learning objective for molecular pretraining
abstract
Accurate prediction of molecular properties is fundamental in drug discovery and development, providing crucial guidance for effective drug design. A critical factor in achieving accurate molecular property prediction lies in the appropriate representation of molecular structures. Presently, prevalent deep learning-based molecular representations rely on 2D structure information as the primary molecular representation, often overlooking essential three-dimensional (3D) conformational information due to the inherent limitations of 2D structures in conveying atomic spatial relationships. In this study, we propose employing the Gram matrix as a condensed representation of 3D molecular structures and for efficient pretraining objectives. Subsequently, we leverage this matrix to construct a novel molecular representation model, Pre-GTM, which inherently encapsulates 3D information. The model accurately predicts the 3D structure of a molecule by estimating the Gram matrix. Our findings demonstrate that Pre-GTM model outperforms the baseline Graphormer model and other pretrained models in the QM9 and MoleculeNet quantitative property prediction task. The integration of the Gram matrix as a condensed representation of 3D molecular structure, incorporated into the Pre-GTM model, opens up promising avenues for its potential application across various domains of molecular research, including drug design, materials science, and chemical engineering.
Wenkai Xiang, Feisheng Zhong, Lin Ni, Mingyue Zheng, Xutong Li, Qian Shi 0005, Dingyan Wang
Briefings Bioinform.4
2024 FAPM: functional annotation of proteins using multimodal models beyond structural modeling
abstract
MOTIVATION: Assigning accurate property labels to proteins, like functional terms and catalytic activity, is challenging, especially for proteins without homologs and "tail labels" with few known examples. Previous methods mainly focused on protein sequence features, overlooking the semantic meaning of protein labels. RESULTS: We introduce functional annotation of proteins using multimodal models (FAPM), a contrastive multimodal model that links natural language with protein sequence language. This model combines a pretrained protein sequence model with a pretrained large language model to generate labels, such as Gene Ontology (GO) functional terms and catalytic activity predictions, in natural language. Our results show that FAPM excels in understanding protein properties, outperforming models based solely on protein sequences or structures. It achieves state-of-the-art performance on public benchmarks and in-house experimentally annotated phage proteins, which often have few known homologs. Additionally, FAPM's flexibility allows it to incorporate extra text prompts, like taxonomy information, enhancing both its predictive performance and explainability. This novel approach offers a promising alternative to current methods that rely on multiple sequence alignment for protein annotation. AVAILABILITY AND IMPLEMENTATION: The online demo is at: https://huggingface.co/spaces/wenkai/FAPM_demo.
Wenkai Xiang, Zhaoping Xiong, Jiacheng Xiong, Zunyun Fu, Mingyue Zheng, Qian Shi 0005
Bioinform.7
2024 GENNDTI: Drug-Target Interaction Prediction Using Graph Neural Network Enhanced by Router Nodes
abstract
Identifying drug-target interactions (DTI) is crucial in drug discovery and repurposing, and in silico techniques for DTI predictions are becoming increasingly important for reducing time and cost. Most interaction-based DTI models rely on the guilt-by-association principle that "similar drugs can interact with similar targets". However, such methods utilize precomputed similarity matrices and cannot dynamically discover intricate correlations. Meanwhile, some methods enrich DTI networks by incorporating additional networks like DDI and PPI networks, enriching biological signals to enhance DTI prediction. While these approaches have achieved promising performance in DTI prediction, such coarse-grained association data do not explain the specific biological mechanisms underlying DTIs. In this work, we propose GENNDTI, which constructs biologically meaningful routers to represent and integrate the salient properties of drugs and targets. Similar drugs or targets connect to more same router nodes, capturing property sharing. In addition, heterogeneous encoders are designed to distinguish different types of interactions, modeling both real and constructed interactions. This strategy enriches graph topology and enhances prediction efficiency as well. We evaluate the proposed method on benchmark datasets, demonstrating comparative performance over existing methods. We specifically analyze router nodes to validate their efficacy in improving predictions and providing biological explanations.
Beiyuan Yang, Yule Liu, Fang Bai, Mingyue Zheng, Jie Zheng 0002
IEEE J. Biomed. Health Informatics5
2023 An explainable molecular property prediction via multi-granularity
Haichao Sun, Guoyin Wang 0001, Qun Liu 0005, Jie Yang 0052, Mingyue Zheng
Inf. Sci.5
2022 Popular Music Production Trend Analysis and Prediction Research
Mingyue Zheng, Jiarui Jin
CHIRA1
2022 An inductive graph neural network model for compound-protein interaction prediction based on a homogeneous graph
abstract
Identifying the potential compound-protein interactions (CPIs) plays an essential role in drug development. The computational approaches for CPI prediction can reduce time and costs of experimental methods and have benefited from the continuously improved graph representation learning. However, most of the network-based methods use heterogeneous graphs, which is challenging due to their complex structures and heterogeneous attributes. Therefore, in this work, we transformed the compound-protein heterogeneous graph to a homogeneous graph by integrating the ligand-based protein representations and overall similarity associations. We then proposed an Inductive Graph AggrEgator-based framework, named CPI-IGAE, for CPI prediction. CPI-IGAE learns the low-dimensional representations of compounds and proteins from the homogeneous graph in an end-to-end manner. The results show that CPI-IGAE performs better than some state-of-the-art methods. Further ablation study and visualization of embeddings reveal the advantages of the model architecture and its role in feature extraction, and some of the top ranked CPIs by CPI-IGAE have been validated by a review of recent literature. The data and source codes are available at https://github.com/wanxiaozhe/CPI-IGAE.
Xiaozhe Wan, Dingyan Wang, Xiaoqin Tan, Xiaohong Liu 0002, Zunyun Fu, Hualiang Jiang, Mingyue Zheng, Xutong Li
Briefings Bioinform.8
2022 Multi-instance learning of graph neural networks for aqueous pKa prediction
abstract
MOTIVATION: The acid dissociation constant (pKa) is a critical parameter to reflect the ionization ability of chemical compounds and is widely applied in a variety of industries. However, the experimental determination of pKa is intricate and time-consuming, especially for the exact determination of micro-pKa information at the atomic level. Hence, a fast and accurate prediction of pKa values of chemical compounds is of broad interest. RESULTS: Here, we compiled a large-scale pKa dataset containing 16 595 compounds with 17 489 pKa values. Based on this dataset, a novel pKa prediction model, named Graph-pKa, was established using graph neural networks. Graph-pKa performed well on the prediction of macro-pKa values, with a mean absolute error around 0.55 and a coefficient of determination around 0.92 on the test dataset. Furthermore, combining multi-instance learning, Graph-pKa was also able to automatically deconvolute the predicted macro-pKa into discrete micro-pKa values. AVAILABILITY AND IMPLEMENTATION: The Graph-pKa model is now freely accessible via a web-based interface (https://pka.simm.ac.cn/). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiacheng Xiong, Guangchao Wang, Zunyun Fu, Feisheng Zhong, Tingyang Xu, Ziming Huang, Xiaohong Liu 0002, Kaixian Chen, Hualiang Jiang, Mingyue Zheng
Bioinform.12
2021 Drug repurposing against breast cancer by integrating drug-exposure expression profiles and drug-drug links based on graph neural network
abstract
MOTIVATION: Breast cancer is one of the leading causes of cancer deaths among women worldwide. It is necessary to develop new breast cancer drugs because of the shortcomings of existing therapies. The traditional discovery process is time-consuming and expensive. Repositioning of clinically approved drugs has emerged as a novel approach for breast cancer therapy. However, serendipitous or experiential repurposing cannot be used as a routine method. RESULTS: In this study, we proposed a graph neural network model GraphRepur based on GraphSAGE for drug repurposing against breast cancer. GraphRepur integrated two major classes of computational methods, drug network-based and drug signature-based. The differentially expressed genes of disease, drug-exposure gene expression data and the drug-drug links information were collected. By extracting the drug signatures and topological structure information contained in the drug relationships, GraphRepur can predict new drugs for breast cancer, outperforming previous state-of-the-art approaches and some classic machine learning methods. The high-ranked drugs have indeed been reported as new uses for breast cancer treatment recently. AVAILABILITYAND IMPLEMENTATION: The source code of our model and datasets are available at: https://github.com/cckamy/GraphRepur and https://figshare.com/articles/software/GraphRepur_Breast_Cancer_Drug_Repurposing/14220050. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dingyan Wang, Lifan Chen, Tingyang Xu, Mingyue Zheng, Xiaomin Luo, Hualiang Jiang, Kaixian Chen
Bioinform.7
2021 Crowdsourced identification of multi-target kinase inhibitors for RET- and TAU- based disease: The Multi-Targeting Drug DREAM Challenge
abstract
A continuing challenge in modern medicine is the identification of safer and more efficacious drugs. Precision therapeutics, which have one molecular target, have been long promised to be safer and more effective than traditional therapies. This approach has proven to be challenging for multiple reasons including lack of efficacy, rapidly acquired drug resistance, and narrow patient eligibility criteria. An alternative approach is the development of drugs that address the overall disease network by targeting multiple biological targets ('polypharmacology'). Rational development of these molecules will require improved methods for predicting single chemical structures that target multiple drug targets. To address this need, we developed the Multi-Targeting Drug DREAM Challenge, in which we challenged participants to predict single chemical entities that target pro-targets but avoid anti-targets for two unrelated diseases: RET-based tumors and a common form of inherited Tauopathy. Here, we report the results of this DREAM Challenge and the development of two neural network-based machine learning approaches that were applied to the challenge of rational polypharmacology. Together, these platforms provide a potentially useful first step towards developing lead therapeutic compounds that address disease complexity through rational polypharmacology.
Zhaoping Xiong, Minji Jeon, Robert J. Allaway, Jaewoo Kang, Donghyeon Park, Jinhyuk Lee, Hwisang Jeon, Miyoung Ko, Hualiang Jiang, Mingyue Zheng, Aik Choon Tan, Xindi Guo, Kristen K. Dang, Alexander Tropsha, Chana Hecht, Tirtha K. Das, Heather A. Carlson, Ruben Abagyan, Justin Guinney, Avner Schlessinger, Ross L. Cagan
PLoS Comput. Biol.10
2020 TransformerCPI: improving compound-protein interaction prediction by sequence-based deep learning with self-attention mechanism and label reversal experiments
abstract
MOTIVATION: Identifying compound-protein interaction (CPI) is a crucial task in drug discovery and chemogenomics studies, and proteins without three-dimensional structure account for a large part of potential biological targets, which requires developing methods using only protein sequence information to predict CPI. However, sequence-based CPI models may face some specific pitfalls, including using inappropriate datasets, hidden ligand bias and splitting datasets inappropriately, resulting in overestimation of their prediction performance. RESULTS: To address these issues, we here constructed new datasets specific for CPI prediction, proposed a novel transformer neural network named TransformerCPI, and introduced a more rigorous label reversal experiment to test whether a model learns true interaction features. TransformerCPI achieved much improved performance on the new experiments, and it can be deconvolved to highlight important interacting regions of protein sequences and compound atoms, which may contribute chemical biology studies with useful guidance for further ligand structural optimization. AVAILABILITY AND IMPLEMENTATION: https://github.com/lifanchen-simm/transformerCPI.
Lifan Chen, Xiaoqin Tan, Dingyan Wang, Feisheng Zhong, Xiaohong Liu 0002, Tianbiao Yang, Xiaomin Luo, Kaixian Chen, Hualiang Jiang, Mingyue Zheng, Arne Elofsson
Bioinform.10
2019 KinomeX: a web application for predicting kinome-wide polypharmacology effect of small molecules
abstract
MOTIVATION: The large-scale kinome-wide virtual profiling for small molecules is a daunting task by experimental and traditional in silico drug design approaches. Recent advances in deep learning algorithms have brought about new opportunities in promoting this process. RESULTS: KinomeX is an online platform to predict kinome-wide polypharmacology effect of small molecules based solely on their chemical structures. The prediction is made by a multi-task deep neural network model trained with over 140 000 bioactivity data points for 391 kinases. Extensive computational and experimental validations have been performed. Overall, KinomeX enables users to create a comprehensive kinome interaction network for designing novel chemical modulators, and is of practical value on exploring the previously less studied or untargeted kinases. AVAILABILITY AND IMPLEMENTATION: KinomeX is available at: https://kinome.dddc.ac.cn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xutong Li, Xiaohong Liu 0002, Zunyun Fu, Zhaoping Xiong, Xiaoqin Tan, Jihui Zhao, Feisheng Zhong, Xiaozhe Wan, Xiaomin Luo, Kaixian Chen, Hualiang Jiang, Mingyue Zheng
Bioinform.14
2015 TarPred: a web application for predicting therapeutic and side effect targets of chemical compounds
abstract
MOTIVATION: Discovering the relevant therapeutic targets for drug-like molecules, or their unintended 'off-targets' that predict adverse drug reactions, is a daunting task by experimental approaches alone. There is thus a high demand to develop computational methods capable of detecting these potential interacting targets efficiently. RESULTS: As biologically annotated chemical data are becoming increasingly available, it becomes feasible to explore such existing knowledge to identify potential ligand-target interactions. Here, we introduce an online implementation of a recently published computational model for target prediction, TarPred, based on a reference library containing 533 individual targets with 179 807 active ligands. TarPred accepts interactive graphical input or input in the chemical file format of SMILES. Given a query compound structure, it provides the top ranked 30 interacting targets. For each of them, TarPred not only shows the structures of three most similar ligands that are known to interact with the target but also highlights the disease indications associated with the target. This information is useful for understanding the mechanisms of action and toxicities of active compounds and can provide drug repositioning opportunities. AVAILABILITY AND IMPLEMENTATION: TarPred is available at: http://www.dddc.ac.cn/tarpred.
Jianlong Peng, Xiaomin Luo, Hualiang Jiang, Mingyue Zheng
Bioinform.10
2014 In silico site of metabolism prediction for human UGT-catalyzed reactions
abstract
MOTIVATION: The human uridine diphosphate-glucuronosyltransferase enzyme family catalyzes the glucuronidation of the glycosyl group of a nucleotide sugar to an acceptor compound (substrate), which is the most common conjugation pathway that serves to protect the organism from the potential toxicity of xenobiotics. Moreover, it could affect the pharmacological profile of a drug. Therefore, it is important to identify the metabolically labile sites for glucuronidation. RESULTS: In the present study, we developed four in silico models to predict sites of glucuronidation, for four major sites of metabolism functional groups, i.e. aliphatic hydroxyl, aromatic hydroxyl, carboxylic acid or amino nitrogen, respectively. According to the mechanism of glucuronidation, a series of 'local' and 'global' molecular descriptors characterizing the atomic reactivity, bonding strength and physical-chemical properties were calculated and selected with a genetic algorithm-based feature selection approach. The constructed support vector machine classification models show good prediction performance, with the balanced accuracy ranging from 0.88 to 0.96 on test set. For further validation, our models can successfully identify 84% of experimentally observed sites of metabolisms for an external test set containing 54 molecules. AVAILABILITY AND IMPLEMENTATION: The software somugt based on our models is available at www.dddc.ac.cn/adme/jlpeng/somugt_win32.zip.
Jianlong Peng, Qiancheng Shen, Mingyue Zheng, Xiaomin Luo, Weiliang Zhu, Hualiang Jiang, Kaixian Chen
Bioinform.4
2009 Site of metabolism prediction for six biotransformations mediated by cytochromes P450
abstract
MOTIVATION: One goal of metabolomics is to define and monitor the entire metabolite complement of a cell, while it is still far from reach since systematic and rapid approaches for determining the biotransformations of newly discovered metabolites are lacking. For drug development, such metabolic biotransformation of a new chemical entity (NCE) is of more interest because it may profoundly affect its bioavailability, activity and toxicity profile. The use of in silico methods to predict the site of metabolism (SOM) in phase I cytochromes P450-mediated reactions is usually a starting point of metabolic pathway studies, which may also assist in the process of drug/lead optimization. RESULTS: This article reports the Cytochromes P450 (CYP450)-mediated SOM prediction for the six most important metabolic reactions by incorporating the use of machine learning and semi-empirical quantum chemical calculations. Non-local models were developed on the basis of a large dataset comprising 1858 metabolic reactions extracted from 1034 heterogeneous chemicals. For validation, the overall accuracies of all six reaction types are higher than 0.81, four of which exceed 0.90. In further receiver operating characteristic (ROC) analyses, each of the SOM model gave a significant area under curve (AUC) value over 0.86, indicating a good predicting power. An external test was made on a previously published dataset, of which 80% of the experimentally observed SOMs can be correctly identified by applying the full set of our SOM models. AVAILABILITY: The program package SOME_v1.0 (Site Of Metabolism Estimator) developed based on our models is available at http://www.dddc.ac.cn/adme/myzheng/SOME_1_0.tar.gz.
Mingyue Zheng, Xiaomin Luo, Qiancheng Shen, Weiliang Zhu, Hualiang Jiang
Bioinform.1
2009 An effective docking strategy for virtual screening based on multi-objective optimization algorithm
abstract
BACKGROUND: Development of a fast and accurate scoring function in virtual screening remains a hot issue in current computer-aided drug research. Different scoring functions focus on diverse aspects of ligand binding, and no single scoring can satisfy the peculiarities of each target system. Therefore, the idea of a consensus score strategy was put forward. Integrating several scoring functions, consensus score re-assesses the docked conformations using a primary scoring function. However, it is not really robust and efficient from the perspective of optimization. Furthermore, to date, the majority of available methods are still based on single objective optimization design. RESULTS: In this paper, two multi-objective optimization methods, called MOSFOM, were developed for virtual screening, which simultaneously consider both the energy score and the contact score. Results suggest that MOSFOM can effectively enhance enrichment and performance compared with a single score. For three different kinds of binding sites, MOSFOM displays an excellent ability to differentiate active compounds through energy and shape complementarity. EFMOGA performed particularly well in the top 2% of database for all three cases, whereas MOEA_Nrg and MOEA_Cnt performed better than the corresponding individual scoring functions if the appropriate type of binding site was selected. CONCLUSION: The multi-objective optimization method was successfully applied in virtual screening with two different scoring functions that can yield reasonable binding poses and can furthermore, be ranked with the potentially compromised conformations of each compound, abandoning those conformations that can not satisfy overall objective functions.
Honglin Li 0003, Hailei Zhang, Mingyue Zheng, Ling Kang, Xiaofeng Liu 0005, Xicheng Wang, Hualiang Jiang
BMC Bioinform.3
2006 Mutagenic probability estimation of chemical compounds by a novel molecular electrophilicity vector and support vector machine
abstract
MOTIVATION: Mutagenicity is among the toxicological end points that pose the highest concern. The accelerated pace of drug discovery has heightened the need for efficient prediction methods. Currently, most available tools fall short of the desired degree of accuracy, and can only provide a binary classification. It is of significance to develop a discriminative and informative model for the mutagenicity prediction. RESULTS: Here we developed a mutagenic probability prediction model addressing the problem, based on datasets covering a large chemical space. A novel molecular electrophilicity vector (MEV) is first devised to represent the structure profile of chemical compounds. An extended support vector machine (SVM) method is then used to derive the posterior probabilistic estimation of mutagenicity from the MEVs of the training set. The results show that our model gives a better performance than TOPKAT (http://www.accelrys.com) and other previously published methods. In addition, a confidence level related to the prediction can be provided, which may help people make more flexible decisions on chemical ordering or synthesis. AVAILABILITY: The binary program (ZGTOX_1.1) based on our model and samples of input datasets on Windows PC are available at http://dddc.ac.cn/adme upon request from the authors.
Mingyue Zheng, Chunxia Xue, Weiliang Zhu, Kaixian Chen, Xiaomin Luo, Hualiang Jiang
Bioinform.1