Dong-Sheng Cao 0001

dblp:03/7930 · also Dongsheng Cao 0001 · DBLP profile ↗
← Back
50ranked-venue papers
3as first author
43since 2021 · last 2026
0000-0003-3604-3785ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 45 · 3 first-author · 38 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TRACE: Transformation-Aware Graph Refinement for Reaction Condition Prediction
abstract
Identifying suitable reaction conditions is critical for chemical synthesis, as they directly affect yield, selectivity, and transformation feasibility. While recent methods have shown promising results, most approaches either encode reactants and products independently or rely on rule-based reaction graphs, both of which constrain the ability of the model to capture condition-relevant structural transformations. In this work, we propose TRACE, a transformation-aware graph refinement framework for reaction condition prediction. TRACE constructs atom-level joint graphs that integrate both reactant and product structures to represent condition-relevant transformations. A structure-aware encoder enriches atom features with local chemical context, followed by a dynamic interaction refinement module that adaptively infers task-specific edges. To further guide the model toward condition-relevant patterns, a mechanism regularized graph encoder incorporates reaction center information, enabling more accurate modeling of transformation mechanisms. Experiments on benchmark datasets show that TRACE achieves state-of-the-art performance across multiple condition types. The integration of transformation-aware refinement leads to improvements in prediction accuracy and generalization, while maintaining robust performance in challenging and realistic synthesis planning scenarios.
Yujie Chen 0002, Tengfei Ma 0002, Yuansheng Liu, Leyi Wei, Dong-Sheng Cao 0001, Xiangxiang Zeng
AAAI6
2026 DynaTCR: dynamic hard-negative ensemble graph learning improves TCR-epitope binding prediction
abstract
MOTIVATION: T-cell receptors (TCRs) recognize antigenic peptides presented by major histocompatibility complex (MHC) molecules and are central to adaptive immunity. Computational prediction of TCR-epitope binding (TEB) can accelerate immunotherapy development, yet remains hampered by limited labeled data, false-negative noise in unobserved pairs, and over-smoothing in graph-based models. RESULTS: We present DynaTCR, a dynamic graph ensemble learning framework for TEB prediction. DynaTCR encodes TCR and epitope sequences with protein language model embeddings and organizes them into a bipartite interaction graph. A graph regularization-variance-preserving aggregation (GR-VPA) encoder stabilizes message propagation and alleviates over-smoothing, while a global attention layer captures long-range dependencies. Multiple base learners are trained with iteratively updated hard-negative samples to reduce false-negative predictions. Under the StrictTCR evaluation protocol on four public datasets, DynaTCR achieves AUC improvements of 4.0-8.2 percentage points over the strongest existing method and up to 15.8 percentage points in AUPR. On the most stringently curated dataset, DynaTCR attains an AUC of 95.1%. Furthermore, on an independent structure-derived test set, DynaTCR achieves the highest AUC (72.6%) among all compared methods, demonstrating its robustness and effectiveness for TEB prediction and candidate prioritization. AVAILABILITY: Source code and data can be downloaded from: https://github.com/2014402680/TEB/.
Xiangzheng Fu, Xinyu Zhang 0012, Linlin Zhuo, Dong-Sheng Cao 0001, Quan Zou 0001
Bioinform.5
2026 DrugKANs: A Paradigm to Enhance Drug-Target Interaction Prediction With KANs
abstract
Identifyingpotential drug-target interactions (DTIs) is crucial for understanding drug mechanisms, and recent computational methods have yielded promising results in this area. However, these methods face several challenges, including limited model generalization due to heavy reliance on multiple similarity datasets and complex feature extraction, as well as a lack of interpretability by ignoring intrinsic information about drugs and targets. To address these challenges, we propose DrugKANs, a novel DTI prediction model that enhances both the quality and interpretability of DTI representations by integrating a dual-tower architecture with Kolmogorov-Arnold Network (KAN) technology. Our model involves utilizing a pre-trained model to derive initial representations of drugs and targets, and employing a lightweight attention mechanism to capture key features, thereby improving representation quality. We leverage the dual-tower architecture and a lightweight feature interaction mechanism to extract high-level representations separately for drugs and targets, aiming to reduce complex feature interactions and mitigate overfitting. Additionally, we incorporate a contrastive learning strategy within the drug-target bipartite graph to address sparse neighborhood effects and enhance topological information. The inclusion of KAN technology further improves the interpretability of the DTI prediction model. Experimental results on public datasets demonstrate that our model predicts DTIs effectively, underscoring its potential as a valuable tool in drug discovery. This comprehensive methodology presents a balanced approach to overcoming the identified challenges in DTI prediction.
Xiangzheng Fu, Zhenya Du, Haiting Chen, Linlin Zhuo, Aiping Lu, Dong-Sheng Cao 0001
IEEE J. Biomed. Health Informatics7
2025 Hybrid approach for drug-target interaction predictions in ischemic stroke models
Jing-Jie Peng, Yi-Yue Zhang, Wen-Jun Zhu, Hong-Rui Liu, Hui-Yin Li, Dong-Sheng Cao 0001, Xiu-Ju Luo
Artif. Intell. Medicine8
2025 ET-PROTACs: modeling ternary complex interactions using cross-modal learning and ternary attention for accurate PROTAC-induced degradation prediction
abstract
MOTIVATION: Accurately predicting the degradation capabilities of proteolysis-targeting chimeras (PROTACs) for given target proteins and E3 ligases is important for PROTAC design. The distinctive ternary structure of PROTACs presents a challenge to traditional drug-target interaction prediction methods, necessitating more innovative approaches. While current state-of-the-art (SOTA) methods using graph neural networks (GNNs) can discern the molecular structure of PROTACs and proteins, thus enabling the efficient prediction of PROTACs' degradation capabilities, they rely heavily on limited crystal structure data of the POI-PROTAC-E3 ternary complex. This reliance underutilizes rich PROTAC experimental data and neglects intricate interaction relationships within ternary complexes. RESULTS: In this study, we propose a model based on cross-modal strategy and ternary attention technology, ET-PROTACs, to predict the targeted degradation capabilities of PROTACs. Our model capitalizes on the strengths of cross-modal methods by using equivariant GNN graph neural networks to process the graph structure and spatial coordinates of PROTAC molecules concurrently while utilizing sequence-based methods to learn the protein sequence information. This integration of cross-modal information is cohesively harnessed and channeled into a ternary attention mechanism, specially tailored for the unique structure of PROTACs, enabling the congruent modeling of both PROTAC and protein modalities. Experimental results demonstrate that the ET-PROTACs model outperforms existing SOTA methods. Moreover, visualizing attention scores illuminates crucial residues and atoms pivotal in specific POI-PROTAC-E3 interactions, thus offering invaluable insights and guidance for future pharmaceutical research. AVAILABILITY AND IMPLEMENTATION: The codes of our model are available at https://github.com/GuanyuYue/ET-PROTACs.
Guanyu Yue, Li Wang 0145, Quan Zou 0001, Xiangzheng Fu, Dong-Sheng Cao 0001
Briefings Bioinform.8
2025 GRAPE: graph-regularized protein language modeling unlocks TCR-epitope binding specificity
abstract
T-cell receptor (TCR)-epitope binding prediction is critical for immunotherapies but remains challenged by sparse interaction networks and severe class imbalance in training data. Current graph neural network (GNN) approaches for predicting TCR-epitope binding (TEB) fail to address two key limitations: over-smoothing during message propagation in sparse TCR-epitope graphs and biased predictions toward dominant epitope-TCR pairs. Here, we present GRAPE (Graph-Regularized Attentive Protein Embeddings), a framework unifying spectral graph regularization and imbalance-aware learning. GRAPE first leverages protein language models (ESM-2) to generate evolutionary-informed TCR/epitope embeddings, constructing a topology-aware interaction graph. To mitigate over-smoothing, we introduce spectral graph regularization, explicitly constraining node feature smoothness to preserve discriminative patterns in sparse neighborhoods. Simultaneously, a dynamic edge reweighting module prioritizes unobserved TCR-epitope edges during graph propagation, coupled with a differentiable area under the ROC curve-maximization objective that directly optimizes for imbalance resilience. Extensive benchmarking on public datasets demonstrates that GRAPE significantly outperforms state-of-the-art methods in TEB prediction. This work establishes GRAPE as a robust framework for elucidating TCR-epitope interactions, with broad applications in immunology research and therapeutic design.
Xiangzheng Fu, Mingqiang Rong, Dong-Sheng Cao 0001, Sisi Yuan, Aiping Lu
Briefings Bioinform.6
2025 Pushing the boundaries of few-shot learning for low-data drug discovery with a Bayesian meta-learning hypernetwork framework
abstract
Hunting for candidate compounds with favorable pharmacological, toxicological, and pharmacokinetic properties in drug discovery is essentially a low-data problem, as data acquisition is both challenging and costly. This inherent data limitation clashes with the requirements of many powerful deep learning models, which typically require large datasets. Here, we present Meta-Mol, a novel few-shot learning framework based on Bayesian Model-Agnostic Meta-Learning. Meta-Mol introduces a novel atom-bond graph isomorphism encoder that captures molecular structure information at the atomic and bond levels. This representation is further enhanced by a Bayesian meta-learning strategy, allowing for task-specific parameter adaptation and reducing overfitting risks. Additionally, a hypernetwork is employed to dynamically adjust weight updates across tasks, facilitating more complex posterior estimation. Our results demonstrate that Meta-Mol significantly outperforms existing models on several benchmarks, providing a robust solution to address data scarcity in drug discovery.
Jiacai Yi, Dejun Jiang 0002, Chengkun Wu, Xiao-Chen Zhang, Weixing He, Dong-Sheng Cao 0001
Briefings Bioinform.7
2025 CardiOT: Towards Interpretable Drug Cardiotoxicity Prediction Using Optimal Transport and Kolmogorov-Arnold Networks
abstract
Investigating the inhibitory effects of compounds on cardiac ion channels is essential for assessing cardiac drug safety. Consequently, researchers have developed computational models to evaluate combined cardiotoxicity (CCT) on cardiac ion channels. However, limitations in experimental data often cause issues like uneven data distribution and scarcity. Additionally, existing models primarily emphasize atomic information flow within graph neural networks (GNNs) while overlooking chemical bonds, leading to inadequate recognition of key structures. Therefore, this study integrates optimal transport (OT), structure remapping (SR), and Kolmogorov-Arnold networks (KANs) into a GNN-based CCT prediction model, CardiOT. First, the proposed CardiOT model employs OT pooling to optimize sample-feature joint distribution using expectation maximization, identifying "important" sample-feature pairs. Additionally, SR technology is used to emphasize the role of chemical bond information in message propagation. KAN technology is integrated to greatly enhance model interpretability. In summary, the model mitigates challenges related to uneven data distribution and scarcity. Multiple experiments on public datasets confirm the model's robust performance. We anticipate that this model will provide deeper insights into compound inhibition mechanisms on cardiac ion channels and reduce toxicity risks.
Xinyu Zhang 0012, Zhenya Du, Linlin Zhuo, Xiangzheng Fu, Dong-Sheng Cao 0001, Boqia Xie, Keqin Li 0001
IEEE J. Biomed. Health Informatics6
2024 Attribute-guided prototype network for few-shot molecular property prediction
abstract
The molecular property prediction (MPP) plays a crucial role in the drug discovery process, providing valuable insights for molecule evaluation and screening. Although deep learning has achieved numerous advances in this area, its success often depends on the availability of substantial labeled data. The few-shot MPP is a more challenging scenario, which aims to identify unseen property with only few available molecules. In this paper, we propose an attribute-guided prototype network (APN) to address the challenge. APN first introduces an molecular attribute extractor, which can not only extract three different types of fingerprint attributes (single fingerprint attributes, dual fingerprint attributes, triplet fingerprint attributes) by considering seven circular-based, five path-based, and two substructure-based fingerprints, but also automatically extract deep attributes from self-supervised learning methods. Furthermore, APN designs the Attribute-Guided Dual-channel Attention module to learn the relationship between the molecular graphs and attributes and refine the local and global representation of the molecules. Compared with existing works, APN leverages high-level human-defined attributes and helps the model to explicitly generalize knowledge in molecular graphs. Experiments on benchmark datasets show that APN can achieve state-of-the-art performance in most cases and demonstrate that the attributes are effective for improving few-shot MPP performance. In addition, the strong generalization ability of APN is verified by conducting experiments on data from different domains.
Linlin Hou, Hongxin Xiang, Xiangxiang Zeng, Dong-Sheng Cao 0001, Bosheng Song
Briefings Bioinform.4
2024 ChemMORT: an automatic ADMET optimization platform using deep learning and multi-objective particle swarm optimization
abstract
Drug discovery and development constitute a laborious and costly undertaking. The success of a drug hinges not only good efficacy but also acceptable absorption, distribution, metabolism, elimination, and toxicity (ADMET) properties. Overall, up to 50% of drug development failures have been contributed from undesirable ADMET profiles. As a multiple parameter objective, the optimization of the ADMET properties is extremely challenging owing to the vast chemical space and limited human expert knowledge. In this study, a freely available platform called Chemical Molecular Optimization, Representation and Translation (ChemMORT) is developed for the optimization of multiple ADMET endpoints without the loss of potency (https://cadd.nscc-tj.cn/deploy/chemmort/). ChemMORT contains three modules: Simplified Molecular Input Line Entry System (SMILES) Encoder, Descriptor Decoder and Molecular Optimizer. The SMILES Encoder can generate the molecular representation with a 512-dimensional vector, and the Descriptor Decoder is able to translate the above representation to the corresponding molecular structure with high accuracy. Based on reversible molecular representation and particle swarm optimization strategy, the Molecular Optimizer can be used to effectively optimize undesirable ADMET properties without the loss of bioactivity, which essentially accomplishes the design of inverse QSAR. The constrained multi-objective optimization of the poly (ADP-ribose) polymerase-1 inhibitor is provided as the case to explore the utility of ChemMORT.
Jiacai Yi, Wen-Tao Zhao, Zhi-Jiang Yang, Xiao-Chen Zhang, Chengkun Wu, Aiping Lu, Dong-Sheng Cao 0001
Briefings Bioinform.8
2024 Assembling spatial clustering framework for heterogeneous spatial transcriptomics data with GRAPHDeep
abstract
MOTIVATION: Spatial clustering is essential and challenging for spatial transcriptomics' data analysis to unravel tissue microenvironment and biological function. Graph neural networks are promising to address gene expression profiles and spatial location information in spatial transcriptomics to generate latent representations. However, choosing an appropriate graph deep learning module and graph neural network necessitates further exploration and investigation. RESULTS: In this article, we present GRAPHDeep to assemble a spatial clustering framework for heterogeneous spatial transcriptomics data. Through integrating 2 graph deep learning modules and 20 graph neural networks, the most appropriate combination is decided for each dataset. The constructed spatial clustering method is compared with state-of-the-art algorithms to demonstrate its effectiveness and superiority. The significant new findings include: (i) the number of genes or proteins of spatial omics data is quite crucial in spatial clustering algorithms; (ii) the variational graph autoencoder is more suitable for spatial clustering tasks than deep graph infomax module; (iii) UniMP, SAGE, SuperGAT, GATv2, GCN, and TAG are the recommended graph neural networks for spatial clustering tasks; and (iv) the used graph neural network in the existent spatial clustering frameworks is not the best candidate. This study could be regarded as desirable guidance for choosing an appropriate graph neural network for spatial clustering. AVAILABILITY AND IMPLEMENTATION: The source code of GRAPHDeep is available at https://github.com/narutoten520/GRAPHDeep. The studied spatial omics data are available at https://zenodo.org/record/8141084.
Zhaoyu Fang, Lining Zhang, Dong-Sheng Cao 0001, Min Li 0007, Mingzhu Yin
Bioinform.5
2024 Geometric deep learning for drug discovery
Mingquan Liu, Chunyan Li 0002, Ruizhe Chen, Dong-Sheng Cao 0001, Xiangxiang Zeng
Expert Syst. Appl.4
2024 Effective drug-target affinity prediction via generative active learning
Yuansheng Liu, Zhenran Zhou, Dong-Sheng Cao 0001, Xiangxiang Zeng
Inf. Sci.4
2023 GPMO: Gradient Perturbation-Based Contrastive Learning for Molecule Optimization
abstract
Optimizing molecules with desired properties is a crucial step in de novo drug design. While translation-based methods have achieved initial success, they continue to face the challenge of the “exposure bias” problem. The challenge of preventing the “exposure bias” problem of molecule optimization lies in the need for both positive and negative molecules of contrastive learning. That is because generating positive molecules through data augmentation requires domain-specific knowledge, and randomly sampled negative molecules are easily distinguished from the real molecules. Hence, in this work, we propose a molecule optimization method called GPMO, which leverages a gradient perturbation-based contrastive learning method to prevent the “exposure bias” problem in translation-based molecule optimization. With the assistance of positive and negative molecules, GPMO is able to effectively handle both real and artificial molecules. GPMO is a molecule optimization method that is conditioned on matched molecule pairs for drug discovery. Our empirical studies show that GPMO outperforms the state-of-the- art molecule optimization methods. Furthermore, the negative and positive perturbations improve the robustness of GPMO.
Xixi Yang, Yafeng Deng, Yuansheng Liu, Dong-Sheng Cao 0001, Xiangxiang Zeng
IJCAI5
2023 DKADE: a novel framework based on deep learning and knowledge graph for identifying adverse drug events and related medications
abstract
Adverse drug events (ADEs) are common in clinical practice and can cause significant harm to patients and increase resource use. Natural language processing (NLP) has been applied to automate ADE detection, but NLP systems become less adaptable when drug entities are missing or multiple medications are specified in clinical narratives. Additionally, no Chinese-language NLP system has been developed for ADE detection due to the complexity of Chinese semantics, despite ˃10 million cases of drug-related adverse events occurring annually in China. To address these challenges, we propose DKADE, a deep learning and knowledge graph-based framework for identifying ADEs. DKADE infers missing drug entities and evaluates their correlations with ADEs by combining medication orders and existing drug knowledge. Moreover, DKADE can automatically screen for new adverse drug reactions. Experimental results show that DKADE achieves an overall F1-score value of 91.13%. Furthermore, the adaptability of DKADE is validated using real-world external clinical data. In summary, DKADE is a powerful tool for studying drug safety and automating adverse event monitoring.
Ze-Ying Feng, Jun-Long Ma, Min Li 0007, Ge-Fei He, Dong-Sheng Cao 0001
Briefings Bioinform.6
2023 Comprehensive assessment of nine target prediction web services: which should we choose for target fishing?
abstract
Identification of potential targets for known bioactive compounds and novel synthetic analogs is of considerable significance. In silico target fishing (TF) has become an alternative strategy because of the expensive and laborious wet-lab experiments, explosive growth of bioactivity data and rapid development of high-throughput technologies. However, these TF methods are based on different algorithms, molecular representations and training datasets, which may lead to different results when predicting the same query molecules. This can be confusing for practitioners in practical applications. Therefore, this study systematically evaluated nine popular ligand-based TF methods based on target and ligand-target pair statistical strategies, which will help practitioners make choices among multiple TF methods. The evaluation results showed that SwissTargetPrediction was the best method to produce the most reliable predictions while enriching more targets. High-recall similarity ensemble approach (SEA) was able to find real targets for more compounds compared with other TF methods. Therefore, SwissTargetPrediction and SEA can be considered as primary selection methods in future studies. In addition, the results showed that k = 5 was the optimal number of experimental candidate targets. Finally, a novel ensemble TF method based on consensus voting is proposed to improve the prediction performance. The precision of the ensemble TF method outperforms the individual TF method, indicating that the ensemble TF method can more effectively identify real targets within a given top-k threshold. The results of this study can be used as a reference to guide practitioners in selecting the most effective methods in computational drug discovery.
Kai-Yue Ji, Zhao-Qian Liu, Yafeng Deng, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.6
2023 Comprehensive evaluation of deep and graph learning on drug-drug interactions prediction
abstract
Recent advances and achievements of artificial intelligence (AI) as well as deep and graph learning models have established their usefulness in biomedical applications, especially in drug-drug interactions (DDIs). DDIs refer to a change in the effect of one drug to the presence of another drug in the human body, which plays an essential role in drug discovery and clinical research. DDIs prediction through traditional clinical trials and experiments is an expensive and time-consuming process. To correctly apply the advanced AI and deep learning, the developer and user meet various challenges such as the availability and encoding of data resources, and the design of computational methods. This review summarizes chemical structure based, network based, natural language processing based and hybrid methods, providing an updated and accessible guide to the broad researchers and development community with different domain knowledge. We introduce widely used molecular representation and describe the theoretical frameworks of graph neural network models for representing molecular structures. We present the advantages and disadvantages of deep and graph learning methods by performing comparative experiments. We discuss the potential technical challenges and highlight future directions of deep and graph learning models for accelerating DDIs prediction.
Xuan Lin, Lichang Dai, Yafang Zhou, Jianyu Shi, Dong-Sheng Cao 0001, Bosheng Song, Philip S. Yu, Xiangxiang Zeng
Briefings Bioinform.7
2023 Graph deep learning enabled spatial domains identification for spatial transcriptomics
abstract
Advancing spatially resolved transcriptomics (ST) technologies help biologists comprehensively understand organ function and tissue microenvironment. Accurate spatial domain identification is the foundation for delineating genome heterogeneity and cellular interaction. Motivated by this perspective, a graph deep learning (GDL) based spatial clustering approach is constructed in this paper. First, the deep graph infomax module embedded with residual gated graph convolutional neural network is leveraged to address the gene expression profiles and spatial positions in ST. Then, the Bayesian Gaussian mixture model is applied to handle the latent embeddings to generate spatial domains. Designed experiments certify that the presented method is superior to other state-of-the-art GDL-enabled techniques on multiple ST datasets. The codes and dataset used in this manuscript are summarized at https://github.com/narutoten520/SCGDL.
Zhaoyu Fang, Li-Ning Zhang, Dong-Sheng Cao 0001, Mingzhu Yin
Briefings Bioinform.5
2023 Reducing false positive rate of docking-based virtual screening by active learning
abstract
Machine learning-based scoring functions (MLSFs) have become a very favorable alternative to classical scoring functions because of their potential superior screening performance. However, the information of negative data used to construct MLSFs was rarely reported in the literature, and meanwhile the putative inactive molecules recorded in existing databases usually have obvious bias from active molecules. Here we proposed an easy-to-use method named AMLSF that combines active learning using negative molecular selection strategies with MLSF, which can iteratively improve the quality of inactive sets and thus reduce the false positive rate of virtual screening. We chose energy auxiliary terms learning as the MLSF and validated our method on eight targets in the diverse subset of DUD-E. For each target, we screened the IterBioScreen database by AMLSF and compared the screening results with those of the four control models. The results illustrate that the number of active molecules in the top 1000 molecules identified by AMLSF was significantly higher than those identified by the control models. In addition, the free energy calculation results for the top 10 molecules screened out by the AMLSF, null model and control models based on DUD-E also proved that more active molecules can be identified, and the false positive rate can be reduced by AMLSF.
Shao-Hua Shi, Xiangxiang Zeng, Su-You Liu, Zhao-Qian Liu, Yafeng Deng, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.10
2022 Out-of-the-box deep learning prediction of quantum-mechanical partial charges by graph representation and transfer learning
abstract
Accurate prediction of atomic partial charges with high-level quantum mechanics (QM) methods suffers from high computational cost. Numerous feature-engineered machine learning (ML)-based predictors with favorable computability and reliability have been developed as alternatives. However, extensive expertise effort was needed for feature engineering of atom chemical environment, which may consequently introduce domain bias. In this study, SuperAtomicCharge, a data-driven deep graph learning framework, was proposed to predict three important types of partial charges (i.e. RESP, DDEC4 and DDEC78) derived from high-level QM calculations based on the structures of molecules. SuperAtomicCharge was designed to simultaneously exploit the 2D and 3D structural information of molecules, which was proved to be an effective way to improve the prediction accuracy of the model. Moreover, a simple transfer learning strategy and a multitask learning strategy based on self-supervised descriptors were also employed to further improve the prediction accuracy of the proposed model. Compared with the latest baselines, including one GNN-based predictor and two ML-based predictors, SuperAtomicCharge showed better performance on all the three external test sets and had better usability and portability. Furthermore, the QM partial charges of new molecules predicted by SuperAtomicCharge can be efficiently used in drug design applications such as structure-based virtual screening, where the predicted RESP and DDEC4 charges of new molecules showed more robust scoring and screening power than the commonly used partial charges. Finally, two tools including an online server (http://cadd.zju.edu.cn/deepchargepredictor) and the source code command lines (https://github.com/zjujdj/SuperAtomicCharge) were developed for the easy access of the SuperAtomicCharge services.
Dejun Jiang 0002, Huiyong Sun, Jike Wang, Chang-Yu Hsieh, Zhenxing Wu, Dong-Sheng Cao 0001, Jian Wu 0001, Tingjun Hou
Briefings Bioinform.7
2022 fastDRH: a webserver to predict and analyze protein-ligand complexes based on molecular docking and MM/PB(GB)SA computation
abstract
Predicting the native or near-native binding pose of a small molecule within a protein binding pocket is an extremely important task in structure-based drug design, especially in the hit-to-lead and lead optimization phases. In this study, fastDRH, a free and open accessed web server, was developed to predict and analyze protein-ligand complex structures. In fastDRH server, AutoDock Vina and AutoDock-GPU docking engines, structure-truncated MM/PB(GB)SA free energy calculation procedures and multiple poses based per-residue energy decomposition analysis were well integrated into a user-friendly and multifunctional online platform. Benefit from the modular architecture, users can flexibly use one or more of three features, including molecular docking, docking pose rescoring and hotspot residue prediction, to obtain the key information clearly based on a result analysis panel supported by 3Dmol.js and Apache ECharts. In terms of protein-ligand binding mode prediction, the integrated structure-truncated MM/PB(GB)SA rescoring procedures exhibit a success rate of >80% in benchmark, which is much better than the AutoDock Vina (~70%). For hotspot residue identification, our multiple poses based per-residue energy decomposition analysis strategy is a more reliable solution than the one using only a single pose, and the performance of our solution has been experimentally validated in several drug discovery projects. To summarize, the fastDRH server is a useful tool for predicting the ligand binding mode and the hotspot residue of protein for ligand binding. The fastDRH server is accessible free of charge at http://cadd.zju.edu.cn/fastdrh/.
Zhe Wang 0041, Huiyong Sun, Yu Kang 0002, Huanxiang Liu, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.6
2022 Comprehensive assessment of deep generative architectures for de novo drug design
abstract
Recently, deep learning (DL)-based de novo drug design represents a new trend in pharmaceutical research, and numerous DL-based methods have been developed for the generation of novel compounds with desired properties. However, a comprehensive understanding of the advantages and disadvantages of these methods is still lacking. In this study, the performances of different generative models were evaluated by analyzing the properties of the generated molecules in different scenarios, such as goal-directed (rediscovery, optimization and scaffold hopping of active compounds) and target-specific (generation of novel compounds for a given target) tasks. In overall, the DL-based models have significant advantages over the baseline models built by the traditional methods in learning the physicochemical property distributions of the training sets and may be more suitable for target-specific tasks. However, both the baselines and DL-based generative models cannot fully exploit the scaffolds of the training sets, and the molecules generated by the DL-based methods even have lower scaffold diversity than those generated by the traditional models. Moreover, our assessment illustrates that the DL-based methods do not exhibit obvious advantages over the genetic algorithm-based baselines in goal-directed tasks. We believe that our study provides valuable guidance for the effective use of generative models in de novo drug design.
Mingyang Wang 0004, Huiyong Sun, Jike Wang, Jinping Pang, Xin Chai, Lei Xu 0035, Honglin Li 0003, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.8
2022 Knowledge-based BERT: a method to extract molecular features like computational chemists
abstract
Molecular property prediction models based on machine learning algorithms have become important tools to triage unpromising lead molecules in the early stages of drug discovery. Compared with the mainstream descriptor- and graph-based methods for molecular property predictions, SMILES-based methods can directly extract molecular features from SMILES without human expert knowledge, but they require more powerful algorithms for feature extraction and a larger amount of data for training, which makes SMILES-based methods less popular. Here, we show the great potential of pre-training in promoting the predictions of important pharmaceutical properties. By utilizing three pre-training tasks based on atom feature prediction, molecular feature prediction and contrastive learning, a new pre-training method K-BERT, which can extract chemical information from SMILES like chemists, was developed. The calculation results on 15 pharmaceutical datasets show that K-BERT outperforms well-established descriptor-based (XGBoost) and graph-based (Attentive FP and HRGCN+) models. In addition, we found that the contrastive learning pre-training task enables K-BERT to 'understand' SMILES not limited to canonical SMILES. Moreover, the general fingerprints K-BERT-FP generated by K-BERT exhibit comparative predictive power to MACCS on 15 pharmaceutical datasets and can also capture molecular size and chirality information that traditional binary fingerprints cannot capture. Our results illustrate the great potential of K-BERT in the practical applications of molecular property predictions in drug discovery.
Zhenxing Wu, Dejun Jiang 0002, Jike Wang, Xujun Zhang, Hongyan Du, Lurong Pan, Chang-Yu Hsieh, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.8
2022 BioNet: a large-scale and heterogeneous biological network model for interaction prediction with graph convolution
abstract
MOTIVATION: Understanding chemical-gene interactions (CGIs) is crucial for screening drugs. Wet experiments are usually costly and laborious, which limits relevant studies to a small scale. On the contrary, computational studies enable efficient in-silico exploration. For the CGI prediction problem, a common method is to perform systematic analyses on a heterogeneous network involving various biomedical entities. Recently, graph neural networks become popular in the field of relation prediction. However, the inherent heterogeneous complexity of biological interaction networks and the massive amount of data pose enormous challenges. This paper aims to develop a data-driven model that is capable of learning latent information from the interaction network and making correct predictions. RESULTS: We developed BioNet, a deep biological networkmodel with a graph encoder-decoder architecture. The graph encoder utilizes graph convolution to learn latent information embedded in complex interactions among chemicals, genes, diseases and biological pathways. The learning process is featured by two consecutive steps. Then, embedded information learnt by the encoder is then employed to make multi-type interaction predictions between chemicals and genes with a tensor decomposition decoder based on the RESCAL algorithm. BioNet includes 79 325 entities as nodes, and 34 005 501 relations as edges. To train such a massive deep graph model, BioNet introduces a parallel training algorithm utilizing multiple Graphics Processing Unit (GPUs). The evaluation experiments indicated that BioNet exhibits outstanding prediction performance with a best area under Receiver Operating Characteristic (ROC) curve of 0.952, which significantly surpasses state-of-theart methods. For further validation, top predicted CGIs of cancer and COVID-19 by BioNet were verified by external curated data and published literature.
Xi Yang 0020, Jing-Lun Ma, Kai Lu 0001, Dong-Sheng Cao 0001, Chengkun Wu
Briefings Bioinform.6
2022 ABC-Net: a divide-and-conquer based deep learning architecture for SMILES recognition from molecular images
abstract
Structural information for chemical compounds is often described by pictorial images in most scientific documents, which cannot be easily understood and manipulated by computers. This dilemma makes optical chemical structure recognition (OCSR) an essential tool for automatically mining knowledge from an enormous amount of literature. However, existing OCSR methods fall far short of our expectations for realistic requirements due to their poor recovery accuracy. In this paper, we developed a deep neural network model named ABC-Net (Atom and Bond Center Network) to predict graph structures directly. Based on the divide-and-conquer principle, we propose to model an atom or a bond as a single point in the center. In this way, we can leverage a fully convolutional neural network (CNN) to generate a series of heat-maps to identify these points and predict relevant properties, such as atom types, atom charges, bond types and other properties. Thus, the molecular structure can be recovered by assembling the detected atoms and bonds. Our approach integrates all the detection and property prediction tasks into a single fully CNN, which is scalable and capable of processing molecular images quite efficiently. Experimental results demonstrate that our method could achieve a significant improvement in recognition performance compared with publicly available tools. The proposed method could be considered as a promising solution to OCSR problems and a starting point for the acquisition of molecular information in the literature.
Xiao-Chen Zhang, Jiacai Yi, Chengkun Wu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.6
2022 MICER: a pre-trained encoder-decoder architecture for molecular image captioning
abstract
MOTIVATION: Automatic recognition of chemical structures from molecular images provides an important avenue for the rediscovery of chemicals. Traditional rule-based approaches that rely on expert knowledge and fail to consider all the stylistic variations of molecular images usually suffer from cumbersome recognition processes and low generalization ability. Deep learning-based methods that integrate different image styles and automatically learn valuable features are flexible, but currently under-researched and have limitations, and are therefore not fully exploited. RESULTS: MICER, an encoder-decoder-based, reconstructed architecture for molecular image captioning, combines transfer learning, attention mechanisms and several strategies to strengthen effectiveness and plasticity in different datasets. The effects of stereochemical information, molecular complexity, data volume and pre-trained encoders on MICER performance were evaluated. Experimental results show that the intrinsic features of the molecular images and the sub-model match have a significant impact on the performance of this task. These findings inspire us to design the training dataset and the encoder for the final validation model, and the experimental results suggest that the MICER model consistently outperforms the state-of-the-art methods on four datasets. MICER was more reliable and scalable due to its interpretability and transfer capacity and provides a practical framework for developing comprehensive and accurate automated molecular structure identification tools to explore unknown chemical space. AVAILABILITY AND IMPLEMENTATION: https://github.com/Jiacai-Yi/MICER. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiacai Yi, Chengkun Wu, Xiao-Chen Zhang, Xinyi Xiao, Tingjun Hou, Dong-Sheng Cao 0001
Bioinform.8
2021 BioMedR: an R/CRAN package for integrated data analysis pipeline in biomedical study
abstract
BACKGROUND: With the increasing development of biotechnology and information technology, publicly available data in chemistry and biology are undergoing explosive growth. Such wealthy information in these resources needs to be extracted and then transformed to useful knowledge by various data mining methods. However, a main computational challenge is how to effectively represent or encode molecular objects under investigation such as chemicals, proteins, DNAs and even complicated interactions when data mining methods are employed. To further explore these complicated data, an integrated toolkit to represent different types of molecular objects and support various data mining algorithms is urgently needed. RESULTS: We developed a freely available R/CRAN package, called BioMedR, for molecular representations of chemicals, proteins, DNAs and pairwise samples of their interactions. The current version of BioMedR could calculate 293 molecular descriptors and 13 kinds of molecular fingerprints for small molecules, 9920 protein descriptors based on protein sequences and six types of generalized scale-based descriptors for proteochemometric modeling, more than 6000 DNA descriptors from nucleotide sequences and six types of interaction descriptors using three different combining strategies. Moreover, this package realized five similarity calculation methods and four powerful clustering algorithms as well as several useful auxiliary tools, which aims at building an integrated analysis pipeline for data acquisition, data checking, descriptor calculation and data modeling. CONCLUSION: BioMedR provides a comprehensive and uniform R package to link up different representations of molecular objects with each other and will benefit cheminformatics/bioinformatics and other biomedical users. It is available at: https://CRAN.R-project.org/package=BioMedR and https://github.com/wind22zhu/BioMedR/.
Yong-Huan Yun, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.6
2021 QSAR-assisted-MMPA to expand chemical transformation space for lead optimization
abstract
Matched molecular pairs analysis (MMPA) has become a powerful tool for automatically and systematically identifying medicinal chemistry transformations from compound/property datasets. However, accurate determination of matched molecular pair (MMP) transformations largely depend on the size and quality of existing experimental data. Lack of high-quality experimental data heavily hampers the extraction of more effective medicinal chemistry knowledge. Here, we developed a new strategy called quantitative structure-activity relationship (QSAR)-assisted-MMPA to expand the number of chemical transformations and took the logD7.4 property endpoint as an example to demonstrate the reliability of the new method. A reliable logD7.4 consensus prediction model was firstly established, and its applicability domain was strictly assessed. By applying the reliable logD7.4 prediction model to screen two chemical databases, we obtained more high-quality logD7.4 data by defining a strict applicability domain threshold. Then, MMPA was performed on the predicted data and experimental data to derive more chemical rules. To validate the reliability of the chemical rules, we compared the magnitude and directionality of the property changes of the predicted rules with those of the measured rules. Then, we compared the novel chemical rules generated by our proposed approach with the published chemical rules, and found that the magnitude and directionality of the property changes were consistent, indicating that the proposed QSAR-assisted-MMPA approach has the potential to enrich the collection of rule types or even identify completely novel rules. Finally, we found that the number of the MMP rules derived from the experimental data could be amplified by the predicted data, which is helpful for us to analyze the medicinal chemical rules in local chemical environment. In summary, the proposed QSAR-assisted-MMPA approach could be regarded as a very promising strategy to expand the chemical transformation space for lead optimization, especially when no enough experimental data can support MMPA.
Zhi-Jiang Yang, Mingzhu Yin, Aiping Lu, Shao Liu 0002, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.9
2021 Beware of the generic machine learning-based scoring functions in structure-based virtual screening
abstract
Machine learning-based scoring functions (MLSFs) have attracted extensive attention recently and are expected to be potential rescoring tools for structure-based virtual screening (SBVS). However, a major concern nowadays is whether MLSFs trained for generic uses rather than a given target can consistently be applicable for VS. In this study, a systematic assessment was carried out to re-evaluate the effectiveness of 14 reported MLSFs in VS. Overall, most of these MLSFs could hardly achieve satisfactory results for any dataset, and they could even not outperform the baseline of classical SFs such as Glide SP. An exception was observed for RFscore-VS trained on the Directory of Useful Decoys-Enhanced dataset, which showed its superiority for most targets. However, in most cases, it clearly illustrated rather limited performance on the targets that were dissimilar to the proteins in the corresponding training sets. We also used the top three docking poses rather than the top one for rescoring and retrained the models with the updated versions of the training set, but only minor improvements were observed. Taken together, generic MLSFs may have poor generalization capabilities to be applicable for the real VS campaigns. Therefore, it should be quite cautious to use this type of methods for VS.
Chao Shen 0008, Zhe Wang 0041, Xujun Zhang, Jinping Pang, Gaoang Wang, Haiyang Zhong, Lei Xu 0035, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.9
2021 Can machine learning consistently improve the scoring power of classical scoring functions? Insights into the role of machine learning in scoring functions
abstract
How to accurately estimate protein-ligand binding affinity remains a key challenge in computer-aided drug design (CADD). In many cases, it has been shown that the binding affinities predicted by classical scoring functions (SFs) cannot correlate well with experimentally measured biological activities. In the past few years, machine learning (ML)-based SFs have gradually emerged as potential alternatives and outperformed classical SFs in a series of studies. In this study, to better recognize the potential of classical SFs, we have conducted a comparative assessment of 25 commonly used SFs. Accordingly, the scoring power was systematically estimated by using the state-of-the-art ML methods that replaced the original multiple linear regression method to refit individual energy terms. The results show that the newly-developed ML-based SFs consistently performed better than classical ones. In particular, gradient boosting decision tree (GBDT) and random forest (RF) achieved the best predictions in most cases. The newly-developed ML-based SFs were also tested on another benchmark modified from PDBbind v2007, and the impacts of structural and sequence similarities were evaluated. The results indicated that the superiority of the ML-based SFs could be fully guaranteed when sufficient similar targets were contained in the training set. Moreover, the effect of the combinations of features from multiple SFs was explored, and the results indicated that combining NNscore2.0 with one to four other classical SFs could yield the best scoring power. However, it was not applicable to derive a generic target-specific SF or SF combination.
Chao Shen 0008, Zhe Wang 0041, Xujun Zhang, Haiyang Zhong, Gaoang Wang, Lei Xu 0035, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.9
2021 Accuracy or novelty: what can we gain from target-specific machine-learning-based scoring functions in virtual screening?
abstract
Machine-learning (ML)-based scoring functions (MLSFs) have gradually emerged as a promising alternative for protein-ligand binding affinity prediction and structure-based virtual screening. However, clouds of doubts have still been raised against the benefits of this novel type of scoring functions (SFs). In this study, to benchmark the performance of target-specific MLSFs on a relatively unbiased dataset, the MLSFs trained from three representative protein-ligand interaction representations were assessed on the LIT-PCBA dataset, and the classical Glide SP SF and three types of ligand-based quantitative structure-activity relationship (QSAR) models were also utilized for comparison. Two major aspects in virtual screening campaigns, including prediction accuracy and hit novelty, were systematically explored. The calculation results illustrate that the tested target-specific MLSFs yielded generally superior performance over the classical Glide SP SF, but they could hardly outperform the 2D fingerprint-based QSAR models. Although substantial improvements could be achieved by integrating multiple types of protein-ligand interaction features, the MLSFs were still not sufficient to exceed MACCS-based QSAR models. In terms of the correlations between the hit ranks or the structures of the top-ranked hits, the MLSFs developed by different featurization strategies would have the ability to identify quite different hits. Nevertheless, it seems that target-specific MLSFs do not have the intrinsic attributes of a traditional SF and may not be a substitute for classical SFs. In contrast, MLSFs can be regarded as a new derivative of ligand-based QSAR models. It is expected that our study may provide valuable guidance for the assessment and further development of target-specific MLSFs.
Chao Shen 0008, Gaoqi Weng, Xujun Zhang, Elaine Lai-Han Leung, Jinping Pang, Xin Chai, Dan Li 0013, Ercheng Wang, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.10
2021 DeepAtomicCharge: a new graph convolutional network-based architecture for accurate prediction of atomic charges
abstract
Atomic charges play a very important role in drug-target recognition. However, computation of atomic charges with high-level quantum mechanics (QM) calculations is very time-consuming. A number of machine learning (ML)-based atomic charge prediction methods have been proposed to speed up the calculation of high-accuracy atomic charges in recent years. However, most of them used a set of predefined molecular properties, such as molecular fingerprints, for model construction, which is knowledge-dependent and may lead to biased predictions due to the representation preference of different molecular properties used for training. To solve the problem, we present a new architecture based on graph convolutional network (GCN) and develop a high-accuracy atomic charge prediction model named DeepAtomicCharge. The new GCN architecture is designed with only the atomic properties and the connection information between the atoms in molecules and can dynamically learn and convert molecules into appropriate atomic features without any prior knowledge of the molecules. Using the designed GCN architecture, substantial improvement is achieved for the prediction accuracy of atomic charges. The average root-mean-square error (RMSE) of DeepAtomicCharge is 0.0121 e, which is obviously more accurate than that (0.0180 e) reported by the previous benchmark study on the same two external test sets. Moreover, the new GCN architecture needs much lower storage space compared with other methods, and the predicted DDEC atomic charges can be efficiently used in large-scale structure-based drug design, thus opening a new avenue for high-performance atomic charge prediction and application.
Jike Wang, Dong-Sheng Cao 0001, Cunchen Tang, Lei Xu 0035, Qiaojun He, Bo Yang 0023, Huiyong Sun, Tingjun Hou
Briefings Bioinform.2
2021 Hyperbolic relational graph convolution networks plus: a simple but highly efficient QSAR-modeling method
abstract
Accurate predictions of druggability and bioactivities of compounds are desirable to reduce the high cost and time of drug discovery. After more than five decades of continuing developments, quantitative structure-activity relationship (QSAR) methods have been established as indispensable tools that facilitate fast, reliable and affordable assessments of physicochemical and biological properties of compounds in drug-discovery programs. Currently, there are mainly two types of QSAR methods, descriptor-based methods and graph-based methods. The former is developed based on predefined molecular descriptors, whereas the latter is developed based on simple atomic and bond information. In this study, we presented a simple but highly efficient modeling method by combining molecular graphs and molecular descriptors as the input of a modified graph neural network, called hyperbolic relational graph convolution network plus (HRGCN+). The evaluation results show that HRGCN+ achieves state-of-the-art performance on 11 drug-discovery-related datasets. We also explored the impact of the addition of traditional molecular descriptors on the predictions of graph-based methods, and found that the addition of molecular descriptors can indeed boost the predictive power of graph-based methods. The results also highlight the strong anti-noise capability of our method. In addition, our method provides a way to interpret models at both the atom and descriptor levels, which can help medicinal chemists extract hidden information from complex datasets. We also offer an HRGCN+'s online prediction service at https://quantum.tencent.com/hrgcn/.
Zhenxing Wu, Dejun Jiang 0002, Chang-Yu Hsieh, Guangyong Chen, Ben Liao, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.6
2021 Do we need different machine learning algorithms for QSAR modeling? A comprehensive assessment of 16 machine learning algorithms on 14 QSAR data sets
abstract
Although a wide variety of machine learning (ML) algorithms have been utilized to learn quantitative structure-activity relationships (QSARs), there is no agreed single best algorithm for QSAR learning. Therefore, a comprehensive understanding of the performance characteristics of popular ML algorithms used in QSAR learning is highly desirable. In this study, five linear algorithms [linear function Gaussian process regression (linear-GPR), linear function support vector machine (linear-SVM), partial least squares regression (PLSR), multiple linear regression (MLR) and principal component regression (PCR)], three analogizers [radial basis function support vector machine (rbf-SVM), K-nearest neighbor (KNN) and radial basis function Gaussian process regression (rbf-GPR)], six symbolists [extreme gradient boosting (XGBoost), Cubist, random forest (RF), multiple adaptive regression splines (MARS), gradient boosting machine (GBM), and classification and regression tree (CART)] and two connectionists [principal component analysis artificial neural network (pca-ANN) and deep neural network (DNN)] were employed to learn the regression-based QSAR models for 14 public data sets comprising nine physicochemical properties and five toxicity endpoints. The results show that rbf-SVM, rbf-GPR, XGBoost and DNN generally illustrate better performances than the other algorithms. The overall performances of different algorithms can be ranked from the best to the worst as follows: rbf-SVM > XGBoost > rbf-GPR > Cubist > GBM > DNN > RF > pca-ANN > MARS > linear-GPR ≈ KNN > linear-SVM ≈ PLSR > CART ≈ PCR ≈ MLR. In terms of prediction accuracy and computational efficiency, SVM and XGBoost are recommended to the regression learning for small data sets, and XGBoost is an excellent choice for large data sets. We then investigated the performances of the ensemble models by integrating the predictions of multiple ML algorithms. The results illustrate that the ensembles of two or three algorithms in different categories can indeed improve the predictions of the best individual ML algorithms.
Zhenxing Wu, Yu Kang 0002, Elaine Lai-Han Leung, Tailong Lei, Chao Shen 0008, Dejun Jiang 0002, Zhe Wang 0041, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.9
2021 Learning to SMILES: BAN-based strategies to improve latent representation learning from molecules
abstract
Computational methods have become indispensable tools to accelerate the drug discovery process and alleviate the excessive dependence on time-consuming and labor-intensive experiments. Traditional feature-engineering approaches heavily rely on expert knowledge to devise useful features, which could be costly and sometimes biased. The emerging deep learning (DL) methods deliver a data-driven method to automatically learn expressive representations from complex raw data. Inspired by this, researchers have attempted to apply various deep neural network models to simplified molecular input line entry specification (SMILES) strings, which contain all the composition and structure information of molecules. However, current models usually suffer from the scarcity of labeled data. This results in a low generalization ability of SMILES-based DL models, which prevents them from competing with the state-of-the-art computational methods. In this study, we utilized the BiLSTM (bidirectional long short term merory) attention network (BAN) in which we employed a novel multi-step attention mechanism to facilitate the extracting of key features from the SMILES strings. Meanwhile, SMILES enumeration was utilized as a data augmentation method in the training phase to substantially increase the number of labeled data and enlarge the probability of mining more patterns from complex SMILES. We again took advantage of SMILES enumeration in the prediction phase to rectify model prediction bias and provide a more accurate prediction. Combined with the BAN model, our strategies can greatly improve the performance of latent features learned from SMILES strings. In 11 canonical absorption, distribution, metabolism, excretion and toxicity-related tasks, our method outperformed the state-of-the-art approaches.
Chengkun Wu, Xiao-Chen Zhang, Zhi-Jiang Yang, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.6
2021 Improving structure-based virtual screening performance via learning from scoring function components
abstract
Scoring functions (SFs) based on complex machine learning (ML) algorithms have gradually emerged as a promising alternative to overcome the weaknesses of classical SFs. However, extensive efforts have been devoted to the development of SFs based on new protein-ligand interaction representations and advanced alternative ML algorithms instead of the energy components obtained by the decomposition of existing SFs. Here, we propose a new method named energy auxiliary terms learning (EATL), in which the scoring components are extracted and used as the input for the development of three levels of ML SFs including EATL SFs, docking-EATL SFs and comprehensive SFs with ascending VS performance. The EATL approach not only outperforms classical SFs for the absolute performance (ROC) and initial enrichment (BEDROC) but also yields comparable performance compared with other advanced ML-based methods on the diverse subset of Directory of Useful Decoys: Enhanced (DUD-E). The test on the relatively unbiased actives as decoys (AD) dataset also proved the effectiveness of EATL. Furthermore, the idea of learning from SF components to yield improved screening power can also be extended to other docking programs and SFs available.
Guo-Li Xiong, Wenling Ye, Chao Shen 0008, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.6
2021 ChemFLuo: a web-server for structure analysis and identification of fluorescent compounds
abstract
BACKGROUND: Fluorescent detection methods are indispensable tools for chemical biology. However, the frequent appearance of potential fluorescent compound has greatly interfered with the recognition of compounds with genuine activity. Such fluorescence interference is especially difficult to identify as it is reproducible and possesses concentration-dependent characteristic. Therefore, the development of a credible screening tool to detect fluorescent compounds from chemical libraries is urgently needed in early stages of drug discovery. RESULTS: In this study, we developed a webserver ChemFLuo for fluorescent compound detection, based on two large and high-quality training datasets containing 4906 blue and 8632 green fluorescent compounds. These molecules were used to construct a group of prediction models based on the combination of three machine learning algorithms and seven types of molecular representations. The best blue fluorescence prediction model achieved with balanced accuracy (BA) = 0.858 and area under the receiver operating characteristic curve (AUC) = 0.931 for the validation set, and BA = 0.823 and AUC = 0.903 for the test set. The best green fluorescence prediction model achieved the prediction accuracy with BA = 0.810 and AUC = 0.887 for the validation set, and BA = 0.771 and AUC = 0.852 for the test set. Besides prediction model, 22 blue and 16 green representative fluorescent substructures were summarized for the screening of potential fluorescent compounds. The comparison with other fluorescence detection tools and theapplication to external validation sets and large molecule libraries have demonstrated the reliability of prediction model for fluorescent compound detection. CONCLUSION: ChemFLuo is a public webserver to filter out compounds with undesirable fluorescent properties, which will benefit the design of high-quality chemical libraries for drug discovery. It is freely available at http://admet.scbdd.com/chemfluo/index/.
Zhi-Jiang Yang, Mingzhu Yin, Hong-Li Jiang, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.9
2021 Scopy: an integrated negative design python library for desirable HTS/VS database design
abstract
BACKGROUND: High-throughput screening (HTS) and virtual screening (VS) have been widely used to identify potential hits from large chemical libraries. However, the frequent occurrence of 'noisy compounds' in the screened libraries, such as compounds with poor drug-likeness, poor selectivity or potential toxicity, has greatly weakened the enrichment capability of HTS and VS campaigns. Therefore, the development of comprehensive and credible tools to detect noisy compounds from chemical libraries is urgently needed in early stages of drug discovery. RESULTS: In this study, we developed a freely available integrated python library for negative design, called Scopy, which supports the functions of data preparation, calculation of descriptors, scaffolds and screening filters, and data visualization. The current version of Scopy can calculate 39 basic molecular properties, 3 comprehensive molecular evaluation scores, 2 types of molecular scaffolds, 6 types of substructure descriptors and 2 types of fingerprints. A number of important screening rules are also provided by Scopy, including 15 drug-likeness rules (13 drug-likeness rules and 2 building block rules), 8 frequent hitter rules (four assay interference substructure filters and four promiscuous compound substructure filters), and 11 toxicophore filters (five human-related toxicity substructure filters, three environment-related toxicity substructure filters and three comprehensive toxicity substructure filters). Moreover, this library supports four different visualization functions to help users to gain a better understanding of the screened data, including basic feature radar chart, feature-feature-related scatter diagram, functional group marker gram and cloud gram. CONCLUSION: Scopy provides a comprehensive Python package to filter out compounds with undesirable properties or substructures, which will benefit the design of high-quality chemical libraries for drug design and discovery. It is freely available at https://github.com/kotori-y/Scopy.
Zhi-Jiang Yang, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.5
2021 PySmash: Python package and individual executable program for representative substructure generation and application
abstract
BACKGROUND: Substructure screening is widely applied to evaluate the molecular potency and ADMET properties of compounds in drug discovery pipelines, and it can also be used to interpret QSAR models for the design of new compounds with desirable physicochemical and biological properties. With the continuous accumulation of more experimental data, data-driven computational systems which can derive representative substructures from large chemical libraries attract more attention. Therefore, the development of an integrated and convenient tool to generate and implement representative substructures is urgently needed. RESULTS: In this study, PySmash, a user-friendly and powerful tool to generate different types of representative substructures, was developed. The current version of PySmash provides both a Python package and an individual executable program, which achieves ease of operation and pipeline integration. Three types of substructure generation algorithms, including circular, path-based and functional group-based algorithms, are provided. Users can conveniently customize their own requirements for substructure size, accuracy and coverage, statistical significance and parallel computation during execution. Besides, PySmash provides the function for external data screening. CONCLUSION: PySmash, a user-friendly and integrated tool for the automatic generation and implementation of representative substructures, is presented. Three screening examples, including toxicophore derivation, privileged motif detection and the integration of substructures with machine learning (ML) models, are provided to illustrate the utility of PySmash in safety profile evaluation, therapeutic activity exploration and molecular optimization, respectively. Its executable program and Python package are available at https://github.com/kotori-y/pySmash.
Zhi-Jiang Yang, Mingzhu Yin, Aiping Lu, Shao Liu 0002, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.9
2021 Identification of active molecules against Mycobacterium tuberculosis through machine learning
abstract
Tuberculosis (TB) is an infectious disease caused by Mycobacterium tuberculosis (Mtb) and it has been one of the top 10 causes of death globally. Drug-resistant tuberculosis (XDR-TB), extensively resistant to the commonly used first-line drugs, has emerged as a major challenge to TB treatment. Hence, it is quite necessary to discover novel drug candidates for TB treatment. In this study, based on different types of molecular representations, four machine learning (ML) algorithms, including support vector machine, random forest (RF), extreme gradient boosting (XGBoost) and deep neural networks (DNN), were used to develop classification models to distinguish Mtb inhibitors from noninhibitors. The results demonstrate that the XGBoost model exhibits the best prediction performance. Then, two consensus strategies were employed to integrate the predictions from multiple models. The evaluation results illustrate that the consensus model by stacking the RF, XGBoost and DNN predictions offers the best predictions with area under the receiver operating characteristic curve of 0.842 and 0.942 for the 10-fold cross-validated training set and external test set, respectively. Besides, the association between the important descriptors and the bioactivities of molecules was interpreted by using the Shapley additive explanations method. Finally, an online webserver called ChemTB (http://cadd.zju.edu.cn/chemtb/) was developed, and it offers a freely available computational tool to detect potential Mtb inhibitors.
Xin Chai, Dejun Jiang 0002, Chao Shen 0008, Xujun Zhang, Dan Li 0013, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.8
2021 MG-BERT: leveraging unsupervised atomic representation learning for molecular property prediction
abstract
MOTIVATION: Accurate and efficient prediction of molecular properties is one of the fundamental issues in drug design and discovery pipelines. Traditional feature engineering-based approaches require extensive expertise in the feature design and selection process. With the development of artificial intelligence (AI) technologies, data-driven methods exhibit unparalleled advantages over the feature engineering-based methods in various domains. Nevertheless, when applied to molecular property prediction, AI models usually suffer from the scarcity of labeled data and show poor generalization ability. RESULTS: In this study, we proposed molecular graph BERT (MG-BERT), which integrates the local message passing mechanism of graph neural networks (GNNs) into the powerful BERT model to facilitate learning from molecular graphs. Furthermore, an effective self-supervised learning strategy named masked atoms prediction was proposed to pretrain the MG-BERT model on a large amount of unlabeled data to mine context information in molecules. We found the MG-BERT model can generate context-sensitive atomic representations after pretraining and transfer the learned knowledge to the prediction of a variety of molecular properties. The experimental results show that the pretrained MG-BERT model with a little extra fine-tuning can consistently outperform the state-of-the-art methods on all 11 ADMET datasets. Moreover, the MG-BERT model leverages attention mechanisms to focus on atomic features essential to the target property, providing excellent interpretability for the trained model. The MG-BERT model does not require any hand-crafted feature as input and is more reliable due to its excellent interpretability, providing a novel framework to develop state-of-the-art models for a wide range of drug discovery tasks.
Xiao-Chen Zhang, Chengkun Wu, Zhi-Jiang Yang, Zhen-Xing Wu, Jiacai Yi, Chang-Yu Hsieh, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.8
2021 Systematic comparison of ligand-based and structure-based virtual screening methods on poly (ADP-ribose) polymerase-1 inhibitors
abstract
The poly (ADP-ribose) polymerase-1 (PARP1) has been regarded as a vital target in recent years and PARP1 inhibitors can be used for ovarian and breast cancer therapies. However, it has been realized that most of PARP1 inhibitors have disadvantages of low solubility and permeability. Therefore, by discovering more molecules with novel frameworks, it would have greater opportunities to apply it into broader clinical fields and have a more profound significance. In the present study, multiple virtual screening (VS) methods had been employed to evaluate the screening efficiency of ligand-based, structure-based and data fusion methods on PARP1 target. The VS methods include 2D similarity screening, structure-activity relationship (SAR) models, docking and complex-based pharmacophore screening. Moreover, the sum rank, sum score and reciprocal rank were also adopted for data fusion methods. The evaluation results show that the similarity searching based on Torsion fingerprint, six SAR models, Glide docking and pharmacophore screening using Phase have excellent screening performance. The best data fusion method is the reciprocal rank, but the sum score also performs well in framework enrichment. In general, the ligand-based VS methods show better performance on PARP1 inhibitor screening. These findings confirmed that adding ligand-based methods to the early screening stage will greatly improve the screening efficiency, and be able to enrich more highly active PARP1 inhibitors with diverse structures.
Xiang-Gui Wang, Zhong-Ye Ma, Guo-Li Xiong, Zhi-Jiang Yang, Aiping Lu, Zhi-Jun Huang, Dong-Sheng Cao 0001
Briefings Bioinform.9
2021 DeepChargePredictor: a web server for predicting QM-based atomic charges via state-of-the-art machine-learning algorithms
abstract
SUMMARY: High-level quantum mechanics (QM) methods are no doubt the most reliable approaches for the prediction of atomic charges, but it usually needs very large computational resources, which apparently hinders the use of high-quality atomic charges in large-scale molecular modeling, such as high-throughput virtual screening. To solve this problem, several algorithms based on machine-learning (ML) have been developed to fit high-level QM atomic charges. Here, we proposed DeepChargePredictor, a web server that is able to generate the high-level QM atomic charges for small molecules based on two state-of-the-art ML algorithms developed in our group, namely AtomPathDescriptor and DeepAtomicCharge. These two algorithms were seamlessly integrated into the platform with the capability to predict three kinds of charges (i.e. RESP, AM1-BCC and DDEC) widely used in structure-based drug design. Moreover, we have comprehensively evaluated the performance of these charges generated by DeepChargePredictor for large-scale drug design applications, such as end-point binding free energy calculations and virtual screening, which all show reliable or even better performance compared with the baseline methods. AVAILABILITY AND IMPLEMENTATION: The data in the article can be obtained on the web page http://cadd.zju.edu.cn/deepchargepredictor/publication. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jike Wang, Huiyong Sun, Dejun Jiang 0002, Zhe Wang 0041, Zhenxing Wu, Dong-Sheng Cao 0001, Tingjun Hou
Bioinform.8
2020 Fast and accurate prediction of partial charges using Atom-Path-Descriptor-based machine learning
abstract
MOTIVATION: Partial atomic charges are usually used to calculate the electrostatic component of energy in many molecular modeling applications, such as molecular docking, molecular dynamics simulations, free energy calculations and so forth. High-level quantum mechanics calculations may provide the most accurate way to estimate the partial charges for small molecules, but they are too time-consuming to be used to process a large number of molecules for high throughput virtual screening. RESULTS: We proposed a new molecule descriptor named Atom-Path-Descriptor (APD) and developed a set of APD-based machine learning (ML) models to predict the partial charges for small molecules with high accuracy. In the APD algorithm, the 3D structures of molecules were assigned with atom centers and atom-pair path-based atom layers to characterize the local chemical environments of atoms. Then, based on the APDs, two representative ensemble ML algorithms, i.e. random forest (RF) and extreme gradient boosting (XGBoost), were employed to develop the regression models for partial charge assignment. The results illustrate that the RF models based on APDs give better predictions for all the atom types than those based on traditional molecular fingerprints reported in the previous study. More encouragingly, the models trained by XGBoost can improve the predictions of partial charges further, and they can achieve the average root-mean-square error 0.0116 e on the external test set, which is much lower than that (0.0195 e) reported in the previous study, suggesting that the proposed algorithm is quite promising to be used in partial charge assignment with high accuracy. AVAILABILITY AND IMPLEMENTATION: The software framework described in this paper is freely available at https://github.com/jkwang93/Atom-Path-Descriptor-based-machine-learning. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jike Wang, Dong-Sheng Cao 0001, Cunchen Tang, Huiyong Sun, Tingjun Hou
Bioinform.2
2015 Rcpi: R/Bioconductor package to generate various descriptors of proteins, compounds and their interactions
abstract
UNLABELLED: In chemoinformatics and bioinformatics fields, one of the main computational challenges in various predictive modeling is to find a suitable way to effectively represent the molecules under investigation, such as small molecules, proteins and even complex interactions. To solve this problem, we developed a freely available R/Bioconductor package, called Compound-Protein Interaction with R (Rcpi), for complex molecular representation from drugs, proteins and more complex interactions, including protein-protein and compound-protein interactions. Rcpi could calculate a large number of structural and physicochemical features of proteins and peptides from amino acid sequences, molecular descriptors of small molecules from their topology and protein-protein interaction and compound-protein interaction descriptors. In addition to main functionalities, Rcpi could also provide a number of useful auxiliary utilities to facilitate the user's need. With the descriptors calculated by this package, the users could conveniently apply various statistical machine learning methods in R to solve various biological and drug research questions in computational biology and drug discovery. AVAILABILITY AND IMPLEMENTATION: Rcpi is freely available from the Bioconductor site (http://bioconductor.org/packages/release/bioc/html/Rcpi.html).
Dong-Sheng Cao 0001, Nan Xiao 0002, Qingsong Xu 0003, Alex F. Chen
Bioinform.1
2015 protr/ProtrWeb: R package and web server for generating various numerical representation schemes of protein sequences
abstract
UNLABELLED: Amino acid sequence-derived structural and physiochemical descriptors are extensively utilized for the research of structural, functional, expression and interaction profiles of proteins and peptides. We developed protr, a comprehensive R package for generating various numerical representation schemes of proteins and peptides from amino acid sequence. The package calculates eight descriptor groups composed of 22 types of commonly used descriptors that include about 22 700 descriptor values. It allows users to select amino acid properties from the AAindex database, and use self-defined properties to construct customized descriptors. For proteochemometric modeling, it calculates six types of scales-based descriptors derived by various dimensionality reduction methods. The protr package also integrates the functionality of similarity score computation derived by protein sequence alignment and Gene Ontology semantic similarity measures within a list of proteins, and calculates profile-based protein features based on position-specific scoring matrix. We also developed ProtrWeb, a user-friendly web server for calculating descriptors presented in the protr package. AVAILABILITY AND IMPLEMENTATION: The protr package is freely available from CRAN: http://cran.r-project.org/package=protr, ProtrWeb, is freely available at http://protrweb.scbdd.com/.
Nan Xiao 0002, Dong-Sheng Cao 0001, Qingsong Xu 0003
Bioinform.2
2013 ChemoPy: freely available python package for computational biology and chemoinformatics
abstract
MOTIVATION: Molecular representation for small molecules has been routinely used in QSAR/SAR, virtual screening, database search, ranking, drug ADME/T prediction and other drug discovery processes. To facilitate extensive studies of drug molecules, we developed a freely available, open-source python package called chemoinformatics in python (ChemoPy) for calculating the commonly used structural and physicochemical features. It computes 16 drug feature groups composed of 19 descriptors that include 1135 descriptor values. In addition, it provides seven types of molecular fingerprint systems for drug molecules, including topological fingerprints, electro-topological state (E-state) fingerprints, MACCS keys, FP4 keys, atom pairs fingerprints, topological torsion fingerprints and Morgan/circular fingerprints. By applying a semi-empirical quantum chemistry program MOPAC, ChemoPy can also compute a large number of 3D molecular descriptors conveniently. AVAILABILITY: The python package, ChemoPy, is freely available via http://code.google.com/p/pychem/downloads/list, and it runs on Linux and MS-Windows. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dong-Sheng Cao 0001, Qingsong Xu 0003, Qian-Nan Hu, Yi-Zeng Liang
Bioinform.1
2013 propy: a tool to generate various modes of Chou's PseAAC
abstract
SUMMARY: Sequence-derived structural and physiochemical features have been frequently used for analysing and predicting structural, functional, expression and interaction profiles of proteins and peptides. To facilitate extensive studies of proteins and peptides, we developed a freely available, open source python package called protein in python (propy) for calculating the widely used structural and physicochemical features of proteins and peptides from amino acid sequence. It computes five feature groups composed of 13 features, including amino acid composition, dipeptide composition, tripeptide composition, normalized Moreau-Broto autocorrelation, Moran autocorrelation, Geary autocorrelation, sequence-order-coupling number, quasi-sequence-order descriptors, composition, transition and distribution of various structural and physicochemical properties and two types of pseudo amino acid composition (PseAAC) descriptors. These features could be generally regarded as different Chou's PseAAC modes. In addition, it can also easily compute the previous descriptors based on user-defined properties, which are automatically available from the AAindex database. AVAILABILITY: The python package, propy, is freely available via http://code.google.com/p/protpy/downloads/list, and it runs on Linux and MS-Windows. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dong-Sheng Cao 0001, Qingsong Xu 0003, Yi-Zeng Liang
Bioinform.1
2011 RxnFinder: biochemical reaction search engines using molecular structures, molecular fragments and reaction similarity
abstract
SUMMARY: Biochemical reactions play a key role to help sustain life and allow cells to grow. RxnFinder was developed to search biochemical reactions from KEGG reaction database using three search criteria: molecular structures, molecular fragments and reaction similarity. RxnFinder is helpful to get reference reactions for biosynthesis and xenobiotics metabolism. AVAILABILITY: RxnFinder is freely available via: http://sdd.whu.edu.cn/rxnfinder. CONTACT: [email protected].
Qian-Nan Hu, Huanan Hu, Dong-Sheng Cao 0001, Yi-Zeng Liang
Bioinform.4
2011 Recipe for Uncovering Predictive Genes Using Support Vector Machines Based on Model Population Analysis
abstract
Selecting a small number of informative genes for microarray-based tumor classification is central to cancer prediction and treatment. Based on model population analysis, here we present a new approach, called Margin Influence Analysis (MIA), designed to work with support vector machines (SVM) for selecting informative genes. The rationale for performing margin influence analysis lies in the fact that the margin of support vector machines is an important factor which underlies the generalization performance of SVM models. Briefly, MIA could reveal genes which have statistically significant influence on the margin by using Mann-Whitney U test. The reason for using the Mann-Whitney U test rather than two-sample t test is that Mann-Whitney U test is a nonparametric test method without any distribution-related assumptions and is also a robust method. Using two publicly available cancerous microarray data sets, it is demonstrated that MIA could typically select a small number of margin-influencing genes and further achieves comparable classification accuracy compared to those reported in the literature. The distinguished features and outstanding performance may make MIA a good alternative for gene selection of high dimensional microarray data. (The source code in MATLAB with GNU General Public License Version 2.0 is freely available at http://code.google.com/p/mia2009/).
Hong-Dong Li, Yi-Zeng Liang, Qingsong Xu 0003, Dong-Sheng Cao 0001, Bin-Bin Tan, Bai-Chuan Deng, Chen-Chen Lin
IEEE ACM Trans. Comput. Biol. Bioinform.4