Tingjun Hou

dblp:75/6017 · DBLP profile ↗
← Back
46ranked-venue papers
2as first author
37since 2021 · last 2025
0000-0001-7227-2580ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 43 · 2 first-author · 34 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer
abstract
Antibodies defend our health by binding to antigens with high specificity and potentiality, primarily relying on the Complementarity-Determining Region (CDR). Yet, current experimental methods of discovering new antibody CDRs are heavily time-consuming. Computational design could alleviate this burden; especially, protein language models have proven quite beneficial in many recent studies. However, most existing models solely focus on antibody potentiality and struggle to encapsulate the diverse range of plausible CDR candidates, limiting their effectiveness in real-world scenarios as binding is only one factor in the multitude of drug-forming criteria. In this paper, we introduce PG-AbD, a framework uniting Generative Flow Networks (GFlowNets) and pretrained Protein Language Models (PLMs) to successfully generate highly potent, diverse and novel antibody candidates. We innovatively construct a Products of Experts (PoE) composed by the global-distribution-modeling PLM and the local-distribution-modeling Potts Model to serve as the reward function of GFlowNet. The joint training paradigm is introduced, where PoE is trained by contrastive divergence with the negative samples generated by GFlowNet, and then guides GFlowNet to sample diverse antibody candidates. We evaluate PG-AbD on extensive antibody design benchmarks. It significantly outperforms existing methods in diversity (13.5% on RabDab, 31.1% on SabDab) while maintaining optimal potential and novelty. Generated antibodies are also found to form stable, regular 3D structures with their corresponding antigens, demonstrating the great potential of PG-AbD to accelerate real-world antibody discovery.
Mingze Yin, Hanjing Zhou, Yiheng Zhu 0002, Jialu Wu, Wei Wu 0045, Kun Fu 0002, Zheng Wang 0027, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001
AAAI10
2025 Improving the predictive performance of binding affinities and poses for protein-cyclic peptide complexes through fine-tuned MM/PBSA(GBSA)-based methods
abstract
Cyclic peptides represent a highly promising class of biopharmaceutical scaffolds. The screening of cyclic peptides against protein targets can be greatly facilitated using computational approaches, especially molecular docking. However, it remains a crucial challenge to accurately predict protein-cyclic peptide (P-cp) interactions employing scoring functions of molecular docking. End-point approaches, such as molecular mechanics generalized Born surface area (MM/GBSA) and molecular mechanics Poisson-Boltzmann surface area (MM/PBSA), provide theoretically more robust frameworks than conventional scoring functions, but their reliability in predicting binding affinities and discriminating native-like binding poses for P-cp complexes remains poorly quantified. Herein, we comprehensively assessed the predictive abilities of MM/PBSA(GBSA) in scoring binding affinities of P-cp complexes and re-ranking their binding poses. The binding affinity scoring ability of MM/PBSA(GBSA) was assessed on a carefully curated dataset consisting of 50 complexes involving P-cp binding affinities, and their re-ranking capability was evaluated on another dataset consisting of the decoys of 81 P-cp complexes. Based on these assessments, we proposed a two-step workflow for predicting P-cp binding affinities. First, we employed the assessed optimal re-ranking method to select the top-1 binding pose; second, we estimated the binding affinity based on the selected top-1 pose using the assessed optimal scoring method. Our proposed workflow, which requires only 3 s for each prediction, achieves binding affinity predictions with a Rp of -0.732 when compared to experimental values, which is twice as high as that of AutoDock CrankPep (Rp = -0.316). This study emphasizes the necessity of using fine-tuned MM/PBSA(GBSA) methods for predicting P-cp interactions.
Huifeng Zhao, Jianxiang Huang, Gaoqi Weng, Dejun Jiang 0002, Renling Hu, Yu Kang 0002, Tingjun Hou
Briefings Bioinform.7
2024 Multiscale topology in interactomic network: from transcriptome to antiaddiction drug repurposing
abstract
The escalating drug addiction crisis in the United States underscores the urgent need for innovative therapeutic strategies. This study embarked on an innovative and rigorous strategy to unearth potential drug repurposing candidates for opioid and cocaine addiction treatment, bridging the gap between transcriptomic data analysis and drug discovery. We initiated our approach by conducting differential gene expression analysis on addiction-related transcriptomic data to identify key genes. We propose a novel topological differentiation to identify key genes from a protein-protein interaction network derived from DEGs. This method utilizes persistent Laplacians to accurately single out pivotal nodes within the network, conducting this analysis in a multiscale manner to ensure high reliability. Through rigorous literature validation, pathway analysis and data-availability scrutiny, we identified three pivotal molecular targets, mTOR, mGluR5 and NMDAR, for drug repurposing from DrugBank. We crafted machine learning models employing two natural language processing (NLP)-based embeddings and a traditional 2D fingerprint, which demonstrated robust predictive ability in gauging binding affinities of DrugBank compounds to selected targets. Furthermore, we elucidated the interactions of promising drugs with the targets and evaluated their drug-likeness. This study delineates a multi-faceted and comprehensive analytical framework, amalgamating bioinformatics, topological data analysis and machine learning, for drug repurposing in addiction treatment, setting the stage for subsequent experimental validation. The versatility of the methods we developed allows for applications across a range of diseases and transcriptomic datasets.
Hongyan Du, Guo-Wei Wei 0001, Tingjun Hou
Briefings Bioinform.3
2024 AttABseq: an attention-based deep learning prediction method for antigen-antibody binding affinity changes based on protein sequences
abstract
The optimization of therapeutic antibodies through traditional techniques, such as candidate screening via hybridoma or phage display, is resource-intensive and time-consuming. In recent years, computational and artificial intelligence-based methods have been actively developed to accelerate and improve the development of therapeutic antibodies. In this study, we developed an end-to-end sequence-based deep learning model, termed AttABseq, for the predictions of the antigen-antibody binding affinity changes connected with antibody mutations. AttABseq is a highly efficient and generic attention-based model by utilizing diverse antigen-antibody complex sequences as the input to predict the binding affinity changes of residue mutations. The assessment on the three benchmark datasets illustrates that AttABseq is 120% more accurate than other sequence-based models in terms of the Pearson correlation coefficient between the predicted and experimental binding affinity changes. Moreover, AttABseq also either outperforms or competes favorably with the structure-based approaches. Furthermore, AttABseq consistently demonstrates robust predictive capabilities across a diverse array of conditions, underscoring its remarkable capacity for generalization across a wide spectrum of antigen-antibody complexes. It imposes no constraints on the quantity of altered residues, rendering it particularly applicable in scenarios where crystallographic structures remain unavailable. The attention-based interpretability analysis indicates that the causal effects of point mutations on antibody-antigen binding affinity changes can be visualized at the residue level, which might assist automated antibody sequence optimization. We believe that AttABseq provides a fiercely competitive answer to therapeutic antibody optimization.
Ruofan Jin, Jike Wang, Dejun Jiang 0002, Tianyue Wang, Yu Kang 0002, Wanting Xu, Chang-Yu Hsieh, Tingjun Hou
Briefings Bioinform.10
2024 Comprehensive assessment of protein loop modeling programs on large-scale datasets: prediction accuracy and efficiency
abstract
Protein loops play a critical role in the dynamics of proteins and are essential for numerous biological functions, and various computational approaches to loop modeling have been proposed over the past decades. However, a comprehensive understanding of the strengths and weaknesses of each method is lacking. In this work, we constructed two high-quality datasets (i.e. the General dataset and the CASP dataset) and systematically evaluated the accuracy and efficiency of 13 commonly used loop modeling approaches from the perspective of loop lengths, protein classes and residue types. The results indicate that the knowledge-based method FREAD generally outperforms the other tested programs in most cases, but encountered challenges when predicting loops longer than 15 and 30 residues on the CASP and General datasets, respectively. The ab initio method Rosetta NGK demonstrated exceptional modeling accuracy for short loops with four to eight residues and achieved the highest success rate on the CASP dataset. The well-known AlphaFold2 and RoseTTAFold require more resources for better performance, but they exhibit promise for predicting loops longer than 16 and 30 residues in the CASP and General datasets. These observations can provide valuable insights for selecting suitable methods for specific loop modeling tasks and contribute to future advancements in the field.
Tianyue Wang, Langcheng Wang, Xujun Zhang, Chao Shen 0008, Odin Zhang, Jike Wang, Jialu Wu, Ruofan Jin, Shicheng Chen, Chang-Yu Hsieh, Guangyong Chen, Peichen Pan, Yu Kang 0002, Tingjun Hou
Briefings Bioinform.17
2023 MolHF: A Hierarchical Normalizing Flow for Molecular Graph Generation
abstract
Molecular de novo design is a critical yet challenging task in scientific fields, aiming to design novel molecular structures with desired property profiles. Significant progress has been made by resorting to generative models for graphs. However, limited attention is paid to hierarchical generative models, which can exploit the inherent hierarchical structure (with rich semantic information) of the molecular graphs and generate complex molecules of larger size that we shall demonstrate to be difficult for most existing models. The primary challenge to hierarchical generation is the non-differentiable issue caused by the generation of intermediate discrete coarsened graph structures. To sidestep this issue, we cast the tricky hierarchical generation problem over discrete spaces as the reverse process of hierarchical representation learning and propose MolHF, a new hierarchical flow-based model that generates molecular graphs in a coarse-to-fine manner. Specifically, MolHF first generates bonds through a multi-scale architecture, then generates atoms based on the coarsened graph structure at each scale. We demonstrate that MolHF achieves state-of-the-art performance in random generation and property optimization, implying its high capacity to model data distribution. Furthermore, MolHF is the first flow-based model that can be applied to model larger molecules (polymer) with more than 100 heavy atoms. The code and models are available at https://github.com/violet-sto/MolHF.
Yiheng Zhu 0002, Zhenqiu Ouyang, Ben Liao, Jialu Wu, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001
IJCAI7
2023 Sample-efficient Multi-objective Molecular Optimization with GFlowNets
abstract
Many crucial scientific problems involve designing novel molecules with desired properties, which can be formulated as a black-box optimization problem over the *discrete* chemical space. In practice, multiple conflicting objectives and costly evaluations (e.g., wet-lab experiments) make the *diversity* of candidates paramount. Computational methods have achieved initial success but still struggle with considering diversity in both objective and search space. To fill this gap, we propose a multi-objective Bayesian optimization (MOBO) algorithm leveraging the hypernetwork-based GFlowNets (HN-GFN) as an acquisition function optimizer, with the purpose of sampling a diverse batch of candidate molecular graphs from an approximate Pareto front. Using a single preference-conditioned hypernetwork, HN-GFN learns to explore various trade-offs between objectives. We further propose a hindsight-like off-policy strategy to share high-performing molecules among different preferences in order to speed up learning for HN-GFN. We empirically illustrate that HN-GFN has adequate capacity to generalize over preferences. Moreover, experiments in various real-world MOBO settings demonstrate that our framework predominantly outperforms existing methods in terms of candidate quality and sample efficiency. The code is available at https://github.com/violet-sto/HN-GFN.
Yiheng Zhu 0002, Jialu Wu, Chaowen Hu, Jiahuan Yan, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001
NeurIPS6
2023 TransFoxMol: predicting molecular property with focused attention
abstract
Predicting the biological properties of molecules is crucial in computer-aided drug development, yet it's often impeded by data scarcity and imbalance in many practical applications. Existing approaches are based on self-supervised learning or 3D data and using an increasing number of parameters to improve performance. These approaches may not take full advantage of established chemical knowledge and could inadvertently introduce noise into the respective model. In this study, we introduce a more elegant transformer-based framework with focused attention for molecular representation (TransFoxMol) to improve the understanding of artificial intelligence (AI) of molecular structure property relationships. TransFoxMol incorporates a multi-scale 2D molecular environment into a graph neural network + Transformer module and uses prior chemical maps to obtain a more focused attention landscape compared to that obtained using existing approaches. Experimental results show that TransFoxMol achieves state-of-the-art performance on MoleculeNet benchmarks and surpasses the performance of baselines that use self-supervised learning or geometry-enhanced strategies on small-scale datasets. Subsequent analyses indicate that TransFoxMol's predictions are highly interpretable and the clever use of chemical knowledge enables AI to perceive molecules in a simple but rational way, enhancing performance.
Zheyuan Shen, Sikang Chen, Qingyu Bian, Yue Guo 0008, Liteng Shen, Jian Wu 0001, Binbin Zhou 0005, Tingjun Hou, Qiaojun He, Jinxin Che, Xiaowu Dong
Briefings Bioinform.12
2023 Can molecular dynamics simulations improve predictions of protein-ligand binding affinity with machine learning?
abstract
Binding affinity prediction largely determines the discovery efficiency of lead compounds in drug discovery. Recently, machine learning (ML)-based approaches have attracted much attention in hopes of enhancing the predictive performance of traditional physics-based approaches. In this study, we evaluated the impact of structural dynamic information on the binding affinity prediction by comparing the models trained on different dimensional descriptors, using three targets (i.e. JAK1, TAF1-BD2 and DDR1) and their corresponding ligands as the examples. Here, 2D descriptors are traditional ECFP4 fingerprints, 3D descriptors are the energy terms of the Smina and NNscore scoring functions and 4D descriptors contain the structural dynamic information derived from the trajectories based on molecular dynamics (MD) simulations. We systematically investigate the MD-refined binding affinity prediction performance of three classical ML algorithms (i.e. RF, SVR and XGB) as well as two common virtual screening methods, namely Glide docking and MM/PBSA. The outcomes of the ML models built using various dimensional descriptors and their combinations reveal that the MD refinement with the optimized protocol can improve the predictive performance on the TAF1-BD2 target with considerable structural flexibility, but not for the less flexible JAK1 and DDR1 targets, when taking docking poses as the initial structure instead of the crystal structures. The results highlight the importance of the initial structures to the final performance of the model through conformational analysis on the three targets with different flexibility.
Shukai Gu, Chao Shen 0008, Huanxiang Liu, Rong Sheng, Lei Xu 0035, Zhe Wang 0041, Tingjun Hou, Yu Kang 0002
Briefings Bioinform.10
2023 Comprehensive assessment of nine target prediction web services: which should we choose for target fishing?
abstract
Identification of potential targets for known bioactive compounds and novel synthetic analogs is of considerable significance. In silico target fishing (TF) has become an alternative strategy because of the expensive and laborious wet-lab experiments, explosive growth of bioactivity data and rapid development of high-throughput technologies. However, these TF methods are based on different algorithms, molecular representations and training datasets, which may lead to different results when predicting the same query molecules. This can be confusing for practitioners in practical applications. Therefore, this study systematically evaluated nine popular ligand-based TF methods based on target and ligand-target pair statistical strategies, which will help practitioners make choices among multiple TF methods. The evaluation results showed that SwissTargetPrediction was the best method to produce the most reliable predictions while enriching more targets. High-recall similarity ensemble approach (SEA) was able to find real targets for more compounds compared with other TF methods. Therefore, SwissTargetPrediction and SEA can be considered as primary selection methods in future studies. In addition, the results showed that k = 5 was the optimal number of experimental candidate targets. Finally, a novel ensemble TF method based on consensus voting is proposed to improve the prediction performance. The precision of the ensemble TF method outperforms the individual TF method, indicating that the ensemble TF method can more effectively identify real targets within a given top-k threshold. The results of this study can be used as a reference to guide practitioners in selecting the most effective methods in computational drug discovery.
Kai-Yue Ji, Zhao-Qian Liu, Yafeng Deng, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.5
2023 Reducing false positive rate of docking-based virtual screening by active learning
abstract
Machine learning-based scoring functions (MLSFs) have become a very favorable alternative to classical scoring functions because of their potential superior screening performance. However, the information of negative data used to construct MLSFs was rarely reported in the literature, and meanwhile the putative inactive molecules recorded in existing databases usually have obvious bias from active molecules. Here we proposed an easy-to-use method named AMLSF that combines active learning using negative molecular selection strategies with MLSF, which can iteratively improve the quality of inactive sets and thus reduce the false positive rate of virtual screening. We chose energy auxiliary terms learning as the MLSF and validated our method on eight targets in the diverse subset of DUD-E. For each target, we screened the IterBioScreen database by AMLSF and compared the screening results with those of the four control models. The results illustrate that the number of active molecules in the top 1000 molecules identified by AMLSF was significantly higher than those identified by the control models. In addition, the free energy calculation results for the top 10 molecules screened out by the AMLSF, null model and control models based on DUD-E also proved that more active molecules can be identified, and the false positive rate can be reduced by AMLSF.
Shao-Hua Shi, Xiangxiang Zeng, Su-You Liu, Zhao-Qian Liu, Yafeng Deng, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.9
2023 Learning with uncertainty to accelerate the discovery of histone lysine-specific demethylase 1A (KDM1A/LSD1) inhibitors
abstract
Machine learning including modern deep learning models has been extensively used in drug design and screening. However, reliable prediction of molecular properties is still challenging when exploring out-of-domain regimes, even for deep neural networks. Therefore, it is important to understand the uncertainty of model predictions, especially when the predictions are used to guide further experiments. In this study, we explored the utility and effectiveness of evidential uncertainty in compound screening. The evidential Graphormer model was proposed for uncertainty-guided discovery of KDM1A/LSD1 inhibitors. The benchmarking results illustrated that (i) Graphormer exhibited comparative predictive power to state-of-the-art models, and (ii) evidential regression enabled well-ranked uncertainty estimates and calibrated predictions. Subsequently, we leveraged time-splitting on the curated KDM1A/LSD1 dataset to simulate out-of-distribution predictions. The retrospective virtual screening showed that the evidential uncertainties helped reduce false positives among the top-acquired compounds and thus enabled higher experimental validation rates. The trained model was then used to virtually screen an independent in-house compound set. The top 50 compounds ranked by two different ranking strategies were experimentally validated, respectively. In general, our study highlighted the importance to understand the uncertainty in prediction, which can be recognized as an interpretable dimension to model predictions.
Zhenxing Wu, Chao Shen 0008, Lingjie Bao, Zhe Wang 0041, Hucheng Yao, De-Xin Kong, Tingjun Hou
Briefings Bioinform.10
2023 Cooperation of structural motifs controls drug selectivity in cyclin-dependent kinases: an advanced theoretical analysis
abstract
Understanding drug selectivity mechanism is a long-standing issue for helping design drugs with high specificity. Designing drugs targeting cyclin-dependent kinases (CDKs) with high selectivity is challenging because of their highly conserved binding pockets. To reveal the underlying general selectivity mechanism, we carried out comprehensive analyses from both the thermodynamics and kinetics points of view on a representative CDK12 inhibitor. To fully capture the binding features of the drug-target recognition process, we proposed to use kinetic residue energy analysis (KREA) in conjunction with the community network analysis (CNA) to reveal the underlying cooperation effect between individual residues/protein motifs to the binding/dissociating process of the ligand. The general mechanism of drug selectivity in CDKs can be summarized as that the difference of structural cooperation between the ligand and the protein motifs leads to the difference of the energetic contribution of the key residues to the ligand. The proposed mechanisms may be prevalent in drug selectivity issues, and the insights may help design new strategies to overcome/attenuate the drug selectivity associated problems.
Lei Xu 0035, Zhe Wang 0041, Tingjun Hou, Haiping Hao, Huiyong Sun
Briefings Bioinform.4
2023 ML-PLIC: a web platform for characterizing protein-ligand interactions and developing machine learning-based scoring functions
abstract
Cracking the entangling code of protein-ligand interaction (PLI) is of great importance to structure-based drug design and discovery. Different physical and biochemical representations can be used to describe PLI such as energy terms and interaction fingerprints, which can be analyzed by machine learning (ML) algorithms to create ML-based scoring functions (MLSFs). Here, we propose the ML-based PLI capturer (ML-PLIC), a web platform that automatically characterizes PLI and generates MLSFs to identify the potential binders of a specific protein target through virtual screening (VS). ML-PLIC comprises five modules, including Docking for ligand docking, Descriptors for PLI generation, Modeling for MLSF training, Screening for VS and Pipeline for the integration of the aforementioned functions. We validated the MLSFs constructed by ML-PLIC in three benchmark datasets (Directory of Useful Decoys-Enhanced, Active as Decoys and TocoDecoy), demonstrating accuracy outperforming traditional docking tools and competitive performance to the deep learning-based SF, and provided a case study of the Serine/threonine-protein kinase WEE1 in which MLSFs were developed by using the ML-based VS pipeline in ML-PLIC. Underpinning the latest version of ML-PLIC is a powerful platform that incorporates physical and biological knowledge about PLI, leveraging PLI characterization and MLSF generation into the design of structure-based VS pipeline. The ML-PLIC web platform is now freely available at http://cadd.zju.edu.cn/plic/.
Xujun Zhang, Chao Shen 0008, Tianyue Wang, Yafeng Deng, Yu Kang 0002, Dan Li 0013, Tingjun Hou, Peichen Pan
Briefings Bioinform.7
2022 Out-of-the-box deep learning prediction of quantum-mechanical partial charges by graph representation and transfer learning
abstract
Accurate prediction of atomic partial charges with high-level quantum mechanics (QM) methods suffers from high computational cost. Numerous feature-engineered machine learning (ML)-based predictors with favorable computability and reliability have been developed as alternatives. However, extensive expertise effort was needed for feature engineering of atom chemical environment, which may consequently introduce domain bias. In this study, SuperAtomicCharge, a data-driven deep graph learning framework, was proposed to predict three important types of partial charges (i.e. RESP, DDEC4 and DDEC78) derived from high-level QM calculations based on the structures of molecules. SuperAtomicCharge was designed to simultaneously exploit the 2D and 3D structural information of molecules, which was proved to be an effective way to improve the prediction accuracy of the model. Moreover, a simple transfer learning strategy and a multitask learning strategy based on self-supervised descriptors were also employed to further improve the prediction accuracy of the proposed model. Compared with the latest baselines, including one GNN-based predictor and two ML-based predictors, SuperAtomicCharge showed better performance on all the three external test sets and had better usability and portability. Furthermore, the QM partial charges of new molecules predicted by SuperAtomicCharge can be efficiently used in drug design applications such as structure-based virtual screening, where the predicted RESP and DDEC4 charges of new molecules showed more robust scoring and screening power than the commonly used partial charges. Finally, two tools including an online server (http://cadd.zju.edu.cn/deepchargepredictor) and the source code command lines (https://github.com/zjujdj/SuperAtomicCharge) were developed for the easy access of the SuperAtomicCharge services.
Dejun Jiang 0002, Huiyong Sun, Jike Wang, Chang-Yu Hsieh, Zhenxing Wu, Dong-Sheng Cao 0001, Jian Wu 0001, Tingjun Hou
Briefings Bioinform.9
2022 Characterizing the stabilization effects of stabilizers in protein-protein systems with end-point binding free energy calculations
abstract
Drug design targeting protein-protein interactions (PPIs) associated with the development of diseases has been one of the most important therapeutic strategies. Besides interrupting the PPIs with PPI inhibitors/blockers, increasing evidence shows that stabilizing the interaction between two interacting proteins may also benefit the therapy, such as the development of various types of molecular glues/stabilizers that mostly work by stabilizing the two interacting proteins to regulate the downstream biological effects. However, characterizing the stabilization effect of a stabilizer is usually hard or too complicated for traditional experiments since it involves ternary interactions [protein-protein-stabilizer (PPS) interaction]. Thus, developing reliable computational strategies will facilitate the discovery/design of molecular glues or PPI stabilizers. Here, by fully analyzing the energetic features of the binary interactions in the PPS ternary complex, we systematically investigated the performance of molecular mechanics Poisson-Boltzmann surface area (MM/PBSA) and molecular mechanics generalized Born surface area (MM/GBSA) methods on characterizing the stabilization effects of stabilizers in 14-3-3 systems. The results show that both MM/PBSA and MM/GBSA are powerful tools in distinguishing the stabilizers from the decoys (with area under the curves of 0.90-0.93 for all tested cases) and are reasonable for ranking protein-peptide interactions in the presence or absence of stabilizers as well (with the average Pearson correlation coefficient of ~0.6 at a relatively high dielectric constant for both methods). Moreover, to give a detailed picture of the stabilization effects, the stabilization mechanism is also analyzed from the structural and energetic points of view for individual systems containing strong or weak stabilizers. This study demonstrates a potential strategy to accelerate the discovery of PPI stabilizers.
Rongfan Tang, Zhe Wang 0041, Haiping Hao, Tingjun Hou, Huiyong Sun
Briefings Bioinform.6
2022 fastDRH: a webserver to predict and analyze protein-ligand complexes based on molecular docking and MM/PB(GB)SA computation
abstract
Predicting the native or near-native binding pose of a small molecule within a protein binding pocket is an extremely important task in structure-based drug design, especially in the hit-to-lead and lead optimization phases. In this study, fastDRH, a free and open accessed web server, was developed to predict and analyze protein-ligand complex structures. In fastDRH server, AutoDock Vina and AutoDock-GPU docking engines, structure-truncated MM/PB(GB)SA free energy calculation procedures and multiple poses based per-residue energy decomposition analysis were well integrated into a user-friendly and multifunctional online platform. Benefit from the modular architecture, users can flexibly use one or more of three features, including molecular docking, docking pose rescoring and hotspot residue prediction, to obtain the key information clearly based on a result analysis panel supported by 3Dmol.js and Apache ECharts. In terms of protein-ligand binding mode prediction, the integrated structure-truncated MM/PB(GB)SA rescoring procedures exhibit a success rate of >80% in benchmark, which is much better than the AutoDock Vina (~70%). For hotspot residue identification, our multiple poses based per-residue energy decomposition analysis strategy is a more reliable solution than the one using only a single pose, and the performance of our solution has been experimentally validated in several drug discovery projects. To summarize, the fastDRH server is a useful tool for predicting the ligand binding mode and the hotspot residue of protein for ligand binding. The fastDRH server is accessible free of charge at http://cadd.zju.edu.cn/fastdrh/.
Zhe Wang 0041, Huiyong Sun, Yu Kang 0002, Huanxiang Liu, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.7
2022 Comprehensive assessment of deep generative architectures for de novo drug design
abstract
Recently, deep learning (DL)-based de novo drug design represents a new trend in pharmaceutical research, and numerous DL-based methods have been developed for the generation of novel compounds with desired properties. However, a comprehensive understanding of the advantages and disadvantages of these methods is still lacking. In this study, the performances of different generative models were evaluated by analyzing the properties of the generated molecules in different scenarios, such as goal-directed (rediscovery, optimization and scaffold hopping of active compounds) and target-specific (generation of novel compounds for a given target) tasks. In overall, the DL-based models have significant advantages over the baseline models built by the traditional methods in learning the physicochemical property distributions of the training sets and may be more suitable for target-specific tasks. However, both the baselines and DL-based generative models cannot fully exploit the scaffolds of the training sets, and the molecules generated by the DL-based methods even have lower scaffold diversity than those generated by the traditional models. Moreover, our assessment illustrates that the DL-based methods do not exhibit obvious advantages over the genetic algorithm-based baselines in goal-directed tasks. We believe that our study provides valuable guidance for the effective use of generative models in de novo drug design.
Mingyang Wang 0004, Huiyong Sun, Jike Wang, Jinping Pang, Xin Chai, Lei Xu 0035, Honglin Li 0003, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.9
2022 Knowledge-based BERT: a method to extract molecular features like computational chemists
abstract
Molecular property prediction models based on machine learning algorithms have become important tools to triage unpromising lead molecules in the early stages of drug discovery. Compared with the mainstream descriptor- and graph-based methods for molecular property predictions, SMILES-based methods can directly extract molecular features from SMILES without human expert knowledge, but they require more powerful algorithms for feature extraction and a larger amount of data for training, which makes SMILES-based methods less popular. Here, we show the great potential of pre-training in promoting the predictions of important pharmaceutical properties. By utilizing three pre-training tasks based on atom feature prediction, molecular feature prediction and contrastive learning, a new pre-training method K-BERT, which can extract chemical information from SMILES like chemists, was developed. The calculation results on 15 pharmaceutical datasets show that K-BERT outperforms well-established descriptor-based (XGBoost) and graph-based (Attentive FP and HRGCN+) models. In addition, we found that the contrastive learning pre-training task enables K-BERT to 'understand' SMILES not limited to canonical SMILES. Moreover, the general fingerprints K-BERT-FP generated by K-BERT exhibit comparative predictive power to MACCS on 15 pharmaceutical datasets and can also capture molecular size and chirality information that traditional binary fingerprints cannot capture. Our results illustrate the great potential of K-BERT in the practical applications of molecular property predictions in drug discovery.
Zhenxing Wu, Dejun Jiang 0002, Jike Wang, Xujun Zhang, Hongyan Du, Lurong Pan, Chang-Yu Hsieh, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.9
2022 ABC-Net: a divide-and-conquer based deep learning architecture for SMILES recognition from molecular images
abstract
Structural information for chemical compounds is often described by pictorial images in most scientific documents, which cannot be easily understood and manipulated by computers. This dilemma makes optical chemical structure recognition (OCSR) an essential tool for automatically mining knowledge from an enormous amount of literature. However, existing OCSR methods fall far short of our expectations for realistic requirements due to their poor recovery accuracy. In this paper, we developed a deep neural network model named ABC-Net (Atom and Bond Center Network) to predict graph structures directly. Based on the divide-and-conquer principle, we propose to model an atom or a bond as a single point in the center. In this way, we can leverage a fully convolutional neural network (CNN) to generate a series of heat-maps to identify these points and predict relevant properties, such as atom types, atom charges, bond types and other properties. Thus, the molecular structure can be recovered by assembling the detected atoms and bonds. Our approach integrates all the detection and property prediction tasks into a single fully CNN, which is scalable and capable of processing molecular images quite efficiently. Experimental results demonstrate that our method could achieve a significant improvement in recognition performance compared with publicly available tools. The proposed method could be considered as a promising solution to OCSR problems and a starting point for the acquisition of molecular information in the literature.
Xiao-Chen Zhang, Jiacai Yi, Chengkun Wu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.5
2022 MICER: a pre-trained encoder-decoder architecture for molecular image captioning
abstract
MOTIVATION: Automatic recognition of chemical structures from molecular images provides an important avenue for the rediscovery of chemicals. Traditional rule-based approaches that rely on expert knowledge and fail to consider all the stylistic variations of molecular images usually suffer from cumbersome recognition processes and low generalization ability. Deep learning-based methods that integrate different image styles and automatically learn valuable features are flexible, but currently under-researched and have limitations, and are therefore not fully exploited. RESULTS: MICER, an encoder-decoder-based, reconstructed architecture for molecular image captioning, combines transfer learning, attention mechanisms and several strategies to strengthen effectiveness and plasticity in different datasets. The effects of stereochemical information, molecular complexity, data volume and pre-trained encoders on MICER performance were evaluated. Experimental results show that the intrinsic features of the molecular images and the sub-model match have a significant impact on the performance of this task. These findings inspire us to design the training dataset and the encoder for the final validation model, and the experimental results suggest that the MICER model consistently outperforms the state-of-the-art methods on four datasets. MICER was more reliable and scalable due to its interpretability and transfer capacity and provides a practical framework for developing comprehensive and accurate automated molecular structure identification tools to explore unknown chemical space. AVAILABILITY AND IMPLEMENTATION: https://github.com/Jiacai-Yi/MICER. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiacai Yi, Chengkun Wu, Xiao-Chen Zhang, Xinyi Xiao, Tingjun Hou, Dong-Sheng Cao 0001
Bioinform.7
2021 BioMedR: an R/CRAN package for integrated data analysis pipeline in biomedical study
abstract
BACKGROUND: With the increasing development of biotechnology and information technology, publicly available data in chemistry and biology are undergoing explosive growth. Such wealthy information in these resources needs to be extracted and then transformed to useful knowledge by various data mining methods. However, a main computational challenge is how to effectively represent or encode molecular objects under investigation such as chemicals, proteins, DNAs and even complicated interactions when data mining methods are employed. To further explore these complicated data, an integrated toolkit to represent different types of molecular objects and support various data mining algorithms is urgently needed. RESULTS: We developed a freely available R/CRAN package, called BioMedR, for molecular representations of chemicals, proteins, DNAs and pairwise samples of their interactions. The current version of BioMedR could calculate 293 molecular descriptors and 13 kinds of molecular fingerprints for small molecules, 9920 protein descriptors based on protein sequences and six types of generalized scale-based descriptors for proteochemometric modeling, more than 6000 DNA descriptors from nucleotide sequences and six types of interaction descriptors using three different combining strategies. Moreover, this package realized five similarity calculation methods and four powerful clustering algorithms as well as several useful auxiliary tools, which aims at building an integrated analysis pipeline for data acquisition, data checking, descriptor calculation and data modeling. CONCLUSION: BioMedR provides a comprehensive and uniform R package to link up different representations of molecular objects with each other and will benefit cheminformatics/bioinformatics and other biomedical users. It is available at: https://CRAN.R-project.org/package=BioMedR and https://github.com/wind22zhu/BioMedR/.
Yong-Huan Yun, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.5
2021 QSAR-assisted-MMPA to expand chemical transformation space for lead optimization
abstract
Matched molecular pairs analysis (MMPA) has become a powerful tool for automatically and systematically identifying medicinal chemistry transformations from compound/property datasets. However, accurate determination of matched molecular pair (MMP) transformations largely depend on the size and quality of existing experimental data. Lack of high-quality experimental data heavily hampers the extraction of more effective medicinal chemistry knowledge. Here, we developed a new strategy called quantitative structure-activity relationship (QSAR)-assisted-MMPA to expand the number of chemical transformations and took the logD7.4 property endpoint as an example to demonstrate the reliability of the new method. A reliable logD7.4 consensus prediction model was firstly established, and its applicability domain was strictly assessed. By applying the reliable logD7.4 prediction model to screen two chemical databases, we obtained more high-quality logD7.4 data by defining a strict applicability domain threshold. Then, MMPA was performed on the predicted data and experimental data to derive more chemical rules. To validate the reliability of the chemical rules, we compared the magnitude and directionality of the property changes of the predicted rules with those of the measured rules. Then, we compared the novel chemical rules generated by our proposed approach with the published chemical rules, and found that the magnitude and directionality of the property changes were consistent, indicating that the proposed QSAR-assisted-MMPA approach has the potential to enrich the collection of rule types or even identify completely novel rules. Finally, we found that the number of the MMP rules derived from the experimental data could be amplified by the predicted data, which is helpful for us to analyze the medicinal chemical rules in local chemical environment. In summary, the proposed QSAR-assisted-MMPA approach could be regarded as a very promising strategy to expand the chemical transformation space for lead optimization, especially when no enough experimental data can support MMPA.
Zhi-Jiang Yang, Mingzhu Yin, Aiping Lu, Shao Liu 0002, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.8
2021 Beware of the generic machine learning-based scoring functions in structure-based virtual screening
abstract
Machine learning-based scoring functions (MLSFs) have attracted extensive attention recently and are expected to be potential rescoring tools for structure-based virtual screening (SBVS). However, a major concern nowadays is whether MLSFs trained for generic uses rather than a given target can consistently be applicable for VS. In this study, a systematic assessment was carried out to re-evaluate the effectiveness of 14 reported MLSFs in VS. Overall, most of these MLSFs could hardly achieve satisfactory results for any dataset, and they could even not outperform the baseline of classical SFs such as Glide SP. An exception was observed for RFscore-VS trained on the Directory of Useful Decoys-Enhanced dataset, which showed its superiority for most targets. However, in most cases, it clearly illustrated rather limited performance on the targets that were dissimilar to the proteins in the corresponding training sets. We also used the top three docking poses rather than the top one for rescoring and retrained the models with the updated versions of the training set, but only minor improvements were observed. Taken together, generic MLSFs may have poor generalization capabilities to be applicable for the real VS campaigns. Therefore, it should be quite cautious to use this type of methods for VS.
Chao Shen 0008, Zhe Wang 0041, Xujun Zhang, Jinping Pang, Gaoang Wang, Haiyang Zhong, Lei Xu 0035, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.10
2021 Can machine learning consistently improve the scoring power of classical scoring functions? Insights into the role of machine learning in scoring functions
abstract
How to accurately estimate protein-ligand binding affinity remains a key challenge in computer-aided drug design (CADD). In many cases, it has been shown that the binding affinities predicted by classical scoring functions (SFs) cannot correlate well with experimentally measured biological activities. In the past few years, machine learning (ML)-based SFs have gradually emerged as potential alternatives and outperformed classical SFs in a series of studies. In this study, to better recognize the potential of classical SFs, we have conducted a comparative assessment of 25 commonly used SFs. Accordingly, the scoring power was systematically estimated by using the state-of-the-art ML methods that replaced the original multiple linear regression method to refit individual energy terms. The results show that the newly-developed ML-based SFs consistently performed better than classical ones. In particular, gradient boosting decision tree (GBDT) and random forest (RF) achieved the best predictions in most cases. The newly-developed ML-based SFs were also tested on another benchmark modified from PDBbind v2007, and the impacts of structural and sequence similarities were evaluated. The results indicated that the superiority of the ML-based SFs could be fully guaranteed when sufficient similar targets were contained in the training set. Moreover, the effect of the combinations of features from multiple SFs was explored, and the results indicated that combining NNscore2.0 with one to four other classical SFs could yield the best scoring power. However, it was not applicable to derive a generic target-specific SF or SF combination.
Chao Shen 0008, Zhe Wang 0041, Xujun Zhang, Haiyang Zhong, Gaoang Wang, Lei Xu 0035, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.10
2021 Accuracy or novelty: what can we gain from target-specific machine-learning-based scoring functions in virtual screening?
abstract
Machine-learning (ML)-based scoring functions (MLSFs) have gradually emerged as a promising alternative for protein-ligand binding affinity prediction and structure-based virtual screening. However, clouds of doubts have still been raised against the benefits of this novel type of scoring functions (SFs). In this study, to benchmark the performance of target-specific MLSFs on a relatively unbiased dataset, the MLSFs trained from three representative protein-ligand interaction representations were assessed on the LIT-PCBA dataset, and the classical Glide SP SF and three types of ligand-based quantitative structure-activity relationship (QSAR) models were also utilized for comparison. Two major aspects in virtual screening campaigns, including prediction accuracy and hit novelty, were systematically explored. The calculation results illustrate that the tested target-specific MLSFs yielded generally superior performance over the classical Glide SP SF, but they could hardly outperform the 2D fingerprint-based QSAR models. Although substantial improvements could be achieved by integrating multiple types of protein-ligand interaction features, the MLSFs were still not sufficient to exceed MACCS-based QSAR models. In terms of the correlations between the hit ranks or the structures of the top-ranked hits, the MLSFs developed by different featurization strategies would have the ability to identify quite different hits. Nevertheless, it seems that target-specific MLSFs do not have the intrinsic attributes of a traditional SF and may not be a substitute for classical SFs. In contrast, MLSFs can be regarded as a new derivative of ligand-based QSAR models. It is expected that our study may provide valuable guidance for the assessment and further development of target-specific MLSFs.
Chao Shen 0008, Gaoqi Weng, Xujun Zhang, Elaine Lai-Han Leung, Jinping Pang, Xin Chai, Dan Li 0013, Ercheng Wang, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.11
2021 DeepAtomicCharge: a new graph convolutional network-based architecture for accurate prediction of atomic charges
abstract
Atomic charges play a very important role in drug-target recognition. However, computation of atomic charges with high-level quantum mechanics (QM) calculations is very time-consuming. A number of machine learning (ML)-based atomic charge prediction methods have been proposed to speed up the calculation of high-accuracy atomic charges in recent years. However, most of them used a set of predefined molecular properties, such as molecular fingerprints, for model construction, which is knowledge-dependent and may lead to biased predictions due to the representation preference of different molecular properties used for training. To solve the problem, we present a new architecture based on graph convolutional network (GCN) and develop a high-accuracy atomic charge prediction model named DeepAtomicCharge. The new GCN architecture is designed with only the atomic properties and the connection information between the atoms in molecules and can dynamically learn and convert molecules into appropriate atomic features without any prior knowledge of the molecules. Using the designed GCN architecture, substantial improvement is achieved for the prediction accuracy of atomic charges. The average root-mean-square error (RMSE) of DeepAtomicCharge is 0.0121 e, which is obviously more accurate than that (0.0180 e) reported by the previous benchmark study on the same two external test sets. Moreover, the new GCN architecture needs much lower storage space compared with other methods, and the predicted DDEC atomic charges can be efficiently used in large-scale structure-based drug design, thus opening a new avenue for high-performance atomic charge prediction and application.
Jike Wang, Dong-Sheng Cao 0001, Cunchen Tang, Lei Xu 0035, Qiaojun He, Bo Yang 0023, Huiyong Sun, Tingjun Hou
Briefings Bioinform.9
2021 Hyperbolic relational graph convolution networks plus: a simple but highly efficient QSAR-modeling method
abstract
Accurate predictions of druggability and bioactivities of compounds are desirable to reduce the high cost and time of drug discovery. After more than five decades of continuing developments, quantitative structure-activity relationship (QSAR) methods have been established as indispensable tools that facilitate fast, reliable and affordable assessments of physicochemical and biological properties of compounds in drug-discovery programs. Currently, there are mainly two types of QSAR methods, descriptor-based methods and graph-based methods. The former is developed based on predefined molecular descriptors, whereas the latter is developed based on simple atomic and bond information. In this study, we presented a simple but highly efficient modeling method by combining molecular graphs and molecular descriptors as the input of a modified graph neural network, called hyperbolic relational graph convolution network plus (HRGCN+). The evaluation results show that HRGCN+ achieves state-of-the-art performance on 11 drug-discovery-related datasets. We also explored the impact of the addition of traditional molecular descriptors on the predictions of graph-based methods, and found that the addition of molecular descriptors can indeed boost the predictive power of graph-based methods. The results also highlight the strong anti-noise capability of our method. In addition, our method provides a way to interpret models at both the atom and descriptor levels, which can help medicinal chemists extract hidden information from complex datasets. We also offer an HRGCN+'s online prediction service at https://quantum.tencent.com/hrgcn/.
Zhenxing Wu, Dejun Jiang 0002, Chang-Yu Hsieh, Guangyong Chen, Ben Liao, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.7
2021 Do we need different machine learning algorithms for QSAR modeling? A comprehensive assessment of 16 machine learning algorithms on 14 QSAR data sets
abstract
Although a wide variety of machine learning (ML) algorithms have been utilized to learn quantitative structure-activity relationships (QSARs), there is no agreed single best algorithm for QSAR learning. Therefore, a comprehensive understanding of the performance characteristics of popular ML algorithms used in QSAR learning is highly desirable. In this study, five linear algorithms [linear function Gaussian process regression (linear-GPR), linear function support vector machine (linear-SVM), partial least squares regression (PLSR), multiple linear regression (MLR) and principal component regression (PCR)], three analogizers [radial basis function support vector machine (rbf-SVM), K-nearest neighbor (KNN) and radial basis function Gaussian process regression (rbf-GPR)], six symbolists [extreme gradient boosting (XGBoost), Cubist, random forest (RF), multiple adaptive regression splines (MARS), gradient boosting machine (GBM), and classification and regression tree (CART)] and two connectionists [principal component analysis artificial neural network (pca-ANN) and deep neural network (DNN)] were employed to learn the regression-based QSAR models for 14 public data sets comprising nine physicochemical properties and five toxicity endpoints. The results show that rbf-SVM, rbf-GPR, XGBoost and DNN generally illustrate better performances than the other algorithms. The overall performances of different algorithms can be ranked from the best to the worst as follows: rbf-SVM > XGBoost > rbf-GPR > Cubist > GBM > DNN > RF > pca-ANN > MARS > linear-GPR ≈ KNN > linear-SVM ≈ PLSR > CART ≈ PCR ≈ MLR. In terms of prediction accuracy and computational efficiency, SVM and XGBoost are recommended to the regression learning for small data sets, and XGBoost is an excellent choice for large data sets. We then investigated the performances of the ensemble models by integrating the predictions of multiple ML algorithms. The results illustrate that the ensembles of two or three algorithms in different categories can indeed improve the predictions of the best individual ML algorithms.
Zhenxing Wu, Yu Kang 0002, Elaine Lai-Han Leung, Tailong Lei, Chao Shen 0008, Dejun Jiang 0002, Zhe Wang 0041, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.10
2021 Learning to SMILES: BAN-based strategies to improve latent representation learning from molecules
abstract
Computational methods have become indispensable tools to accelerate the drug discovery process and alleviate the excessive dependence on time-consuming and labor-intensive experiments. Traditional feature-engineering approaches heavily rely on expert knowledge to devise useful features, which could be costly and sometimes biased. The emerging deep learning (DL) methods deliver a data-driven method to automatically learn expressive representations from complex raw data. Inspired by this, researchers have attempted to apply various deep neural network models to simplified molecular input line entry specification (SMILES) strings, which contain all the composition and structure information of molecules. However, current models usually suffer from the scarcity of labeled data. This results in a low generalization ability of SMILES-based DL models, which prevents them from competing with the state-of-the-art computational methods. In this study, we utilized the BiLSTM (bidirectional long short term merory) attention network (BAN) in which we employed a novel multi-step attention mechanism to facilitate the extracting of key features from the SMILES strings. Meanwhile, SMILES enumeration was utilized as a data augmentation method in the training phase to substantially increase the number of labeled data and enlarge the probability of mining more patterns from complex SMILES. We again took advantage of SMILES enumeration in the prediction phase to rectify model prediction bias and provide a more accurate prediction. Combined with the BAN model, our strategies can greatly improve the performance of latent features learned from SMILES strings. In 11 canonical absorption, distribution, metabolism, excretion and toxicity-related tasks, our method outperformed the state-of-the-art approaches.
Chengkun Wu, Xiao-Chen Zhang, Zhi-Jiang Yang, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.5
2021 Improving structure-based virtual screening performance via learning from scoring function components
abstract
Scoring functions (SFs) based on complex machine learning (ML) algorithms have gradually emerged as a promising alternative to overcome the weaknesses of classical SFs. However, extensive efforts have been devoted to the development of SFs based on new protein-ligand interaction representations and advanced alternative ML algorithms instead of the energy components obtained by the decomposition of existing SFs. Here, we propose a new method named energy auxiliary terms learning (EATL), in which the scoring components are extracted and used as the input for the development of three levels of ML SFs including EATL SFs, docking-EATL SFs and comprehensive SFs with ascending VS performance. The EATL approach not only outperforms classical SFs for the absolute performance (ROC) and initial enrichment (BEDROC) but also yields comparable performance compared with other advanced ML-based methods on the diverse subset of Directory of Useful Decoys: Enhanced (DUD-E). The test on the relatively unbiased actives as decoys (AD) dataset also proved the effectiveness of EATL. Furthermore, the idea of learning from SF components to yield improved screening power can also be extended to other docking programs and SFs available.
Guo-Li Xiong, Wenling Ye, Chao Shen 0008, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.5
2021 ChemFLuo: a web-server for structure analysis and identification of fluorescent compounds
abstract
BACKGROUND: Fluorescent detection methods are indispensable tools for chemical biology. However, the frequent appearance of potential fluorescent compound has greatly interfered with the recognition of compounds with genuine activity. Such fluorescence interference is especially difficult to identify as it is reproducible and possesses concentration-dependent characteristic. Therefore, the development of a credible screening tool to detect fluorescent compounds from chemical libraries is urgently needed in early stages of drug discovery. RESULTS: In this study, we developed a webserver ChemFLuo for fluorescent compound detection, based on two large and high-quality training datasets containing 4906 blue and 8632 green fluorescent compounds. These molecules were used to construct a group of prediction models based on the combination of three machine learning algorithms and seven types of molecular representations. The best blue fluorescence prediction model achieved with balanced accuracy (BA) = 0.858 and area under the receiver operating characteristic curve (AUC) = 0.931 for the validation set, and BA = 0.823 and AUC = 0.903 for the test set. The best green fluorescence prediction model achieved the prediction accuracy with BA = 0.810 and AUC = 0.887 for the validation set, and BA = 0.771 and AUC = 0.852 for the test set. Besides prediction model, 22 blue and 16 green representative fluorescent substructures were summarized for the screening of potential fluorescent compounds. The comparison with other fluorescence detection tools and theapplication to external validation sets and large molecule libraries have demonstrated the reliability of prediction model for fluorescent compound detection. CONCLUSION: ChemFLuo is a public webserver to filter out compounds with undesirable fluorescent properties, which will benefit the design of high-quality chemical libraries for drug discovery. It is freely available at http://admet.scbdd.com/chemfluo/index/.
Zhi-Jiang Yang, Mingzhu Yin, Hong-Li Jiang, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.8
2021 Scopy: an integrated negative design python library for desirable HTS/VS database design
abstract
BACKGROUND: High-throughput screening (HTS) and virtual screening (VS) have been widely used to identify potential hits from large chemical libraries. However, the frequent occurrence of 'noisy compounds' in the screened libraries, such as compounds with poor drug-likeness, poor selectivity or potential toxicity, has greatly weakened the enrichment capability of HTS and VS campaigns. Therefore, the development of comprehensive and credible tools to detect noisy compounds from chemical libraries is urgently needed in early stages of drug discovery. RESULTS: In this study, we developed a freely available integrated python library for negative design, called Scopy, which supports the functions of data preparation, calculation of descriptors, scaffolds and screening filters, and data visualization. The current version of Scopy can calculate 39 basic molecular properties, 3 comprehensive molecular evaluation scores, 2 types of molecular scaffolds, 6 types of substructure descriptors and 2 types of fingerprints. A number of important screening rules are also provided by Scopy, including 15 drug-likeness rules (13 drug-likeness rules and 2 building block rules), 8 frequent hitter rules (four assay interference substructure filters and four promiscuous compound substructure filters), and 11 toxicophore filters (five human-related toxicity substructure filters, three environment-related toxicity substructure filters and three comprehensive toxicity substructure filters). Moreover, this library supports four different visualization functions to help users to gain a better understanding of the screened data, including basic feature radar chart, feature-feature-related scatter diagram, functional group marker gram and cloud gram. CONCLUSION: Scopy provides a comprehensive Python package to filter out compounds with undesirable properties or substructures, which will benefit the design of high-quality chemical libraries for drug design and discovery. It is freely available at https://github.com/kotori-y/Scopy.
Zhi-Jiang Yang, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.4
2021 PySmash: Python package and individual executable program for representative substructure generation and application
abstract
BACKGROUND: Substructure screening is widely applied to evaluate the molecular potency and ADMET properties of compounds in drug discovery pipelines, and it can also be used to interpret QSAR models for the design of new compounds with desirable physicochemical and biological properties. With the continuous accumulation of more experimental data, data-driven computational systems which can derive representative substructures from large chemical libraries attract more attention. Therefore, the development of an integrated and convenient tool to generate and implement representative substructures is urgently needed. RESULTS: In this study, PySmash, a user-friendly and powerful tool to generate different types of representative substructures, was developed. The current version of PySmash provides both a Python package and an individual executable program, which achieves ease of operation and pipeline integration. Three types of substructure generation algorithms, including circular, path-based and functional group-based algorithms, are provided. Users can conveniently customize their own requirements for substructure size, accuracy and coverage, statistical significance and parallel computation during execution. Besides, PySmash provides the function for external data screening. CONCLUSION: PySmash, a user-friendly and integrated tool for the automatic generation and implementation of representative substructures, is presented. Three screening examples, including toxicophore derivation, privileged motif detection and the integration of substructures with machine learning (ML) models, are provided to illustrate the utility of PySmash in safety profile evaluation, therapeutic activity exploration and molecular optimization, respectively. Its executable program and Python package are available at https://github.com/kotori-y/pySmash.
Zhi-Jiang Yang, Mingzhu Yin, Aiping Lu, Shao Liu 0002, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.8
2021 Identification of active molecules against Mycobacterium tuberculosis through machine learning
abstract
Tuberculosis (TB) is an infectious disease caused by Mycobacterium tuberculosis (Mtb) and it has been one of the top 10 causes of death globally. Drug-resistant tuberculosis (XDR-TB), extensively resistant to the commonly used first-line drugs, has emerged as a major challenge to TB treatment. Hence, it is quite necessary to discover novel drug candidates for TB treatment. In this study, based on different types of molecular representations, four machine learning (ML) algorithms, including support vector machine, random forest (RF), extreme gradient boosting (XGBoost) and deep neural networks (DNN), were used to develop classification models to distinguish Mtb inhibitors from noninhibitors. The results demonstrate that the XGBoost model exhibits the best prediction performance. Then, two consensus strategies were employed to integrate the predictions from multiple models. The evaluation results illustrate that the consensus model by stacking the RF, XGBoost and DNN predictions offers the best predictions with area under the receiver operating characteristic curve of 0.842 and 0.942 for the 10-fold cross-validated training set and external test set, respectively. Besides, the association between the important descriptors and the bioactivities of molecules was interpreted by using the Shapley additive explanations method. Finally, an online webserver called ChemTB (http://cadd.zju.edu.cn/chemtb/) was developed, and it offers a freely available computational tool to detect potential Mtb inhibitors.
Xin Chai, Dejun Jiang 0002, Chao Shen 0008, Xujun Zhang, Dan Li 0013, Dong-Sheng Cao 0001, Tingjun Hou
Briefings Bioinform.9
2021 MG-BERT: leveraging unsupervised atomic representation learning for molecular property prediction
abstract
MOTIVATION: Accurate and efficient prediction of molecular properties is one of the fundamental issues in drug design and discovery pipelines. Traditional feature engineering-based approaches require extensive expertise in the feature design and selection process. With the development of artificial intelligence (AI) technologies, data-driven methods exhibit unparalleled advantages over the feature engineering-based methods in various domains. Nevertheless, when applied to molecular property prediction, AI models usually suffer from the scarcity of labeled data and show poor generalization ability. RESULTS: In this study, we proposed molecular graph BERT (MG-BERT), which integrates the local message passing mechanism of graph neural networks (GNNs) into the powerful BERT model to facilitate learning from molecular graphs. Furthermore, an effective self-supervised learning strategy named masked atoms prediction was proposed to pretrain the MG-BERT model on a large amount of unlabeled data to mine context information in molecules. We found the MG-BERT model can generate context-sensitive atomic representations after pretraining and transfer the learned knowledge to the prediction of a variety of molecular properties. The experimental results show that the pretrained MG-BERT model with a little extra fine-tuning can consistently outperform the state-of-the-art methods on all 11 ADMET datasets. Moreover, the MG-BERT model leverages attention mechanisms to focus on atomic features essential to the target property, providing excellent interpretability for the trained model. The MG-BERT model does not require any hand-crafted feature as input and is more reliable due to its excellent interpretability, providing a novel framework to develop state-of-the-art models for a wide range of drug discovery tasks.
Xiao-Chen Zhang, Chengkun Wu, Zhi-Jiang Yang, Zhen-Xing Wu, Jiacai Yi, Chang-Yu Hsieh, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.7
2021 DeepChargePredictor: a web server for predicting QM-based atomic charges via state-of-the-art machine-learning algorithms
abstract
SUMMARY: High-level quantum mechanics (QM) methods are no doubt the most reliable approaches for the prediction of atomic charges, but it usually needs very large computational resources, which apparently hinders the use of high-quality atomic charges in large-scale molecular modeling, such as high-throughput virtual screening. To solve this problem, several algorithms based on machine-learning (ML) have been developed to fit high-level QM atomic charges. Here, we proposed DeepChargePredictor, a web server that is able to generate the high-level QM atomic charges for small molecules based on two state-of-the-art ML algorithms developed in our group, namely AtomPathDescriptor and DeepAtomicCharge. These two algorithms were seamlessly integrated into the platform with the capability to predict three kinds of charges (i.e. RESP, AM1-BCC and DDEC) widely used in structure-based drug design. Moreover, we have comprehensively evaluated the performance of these charges generated by DeepChargePredictor for large-scale drug design applications, such as end-point binding free energy calculations and virtual screening, which all show reliable or even better performance compared with the baseline methods. AVAILABILITY AND IMPLEMENTATION: The data in the article can be obtained on the web page http://cadd.zju.edu.cn/deepchargepredictor/publication. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jike Wang, Huiyong Sun, Dejun Jiang 0002, Zhe Wang 0041, Zhenxing Wu, Dong-Sheng Cao 0001, Tingjun Hou
Bioinform.9
2020 Comprehensive assessment of nine docking programs on type II kinase inhibitors: prediction accuracy of sampling power, scoring power and screening power
abstract
Protein kinases have been regarded as important therapeutic targets for many diseases. Currently, a total of 41 kinase inhibitors have been approved by the Food and Drug Administration, along with a large number of kinase inhibitors being evaluated in clinical and preclinical trials. Among all, allosteric inhibitors, such as type II kinase inhibitors, have attracted extensive attention owing to their potential high selectivity. Nowadays, molecular docking has become a powerful tool to search for novel kinase inhibitors. However, as for type II kinase inhibitors, their allosteric characteristics may exert a deep influence on docking accuracy. In this study, a comprehensive assessment was conducted to evaluate the effectiveness of nine docking algorithms towards type II kinase inhibitors. The calculation results showed that most tested docking programs, especially Glide with XP scoring, LeDock and Surflex-Dock, succeeded in the accurate identification of near-native binding poses, with the success rates ranging from 0.80 to 0.90, and the scoring functions in GOLD and LeDock outperformed the others in the prediction of relative binding affinities. In terms of the P-values, areas under the curve and enrichment factors, Glide with XP scoring, Surflex-Dock, GOLD with Astex Statistical Potential scoring and LeDock had better screening power to discriminate between active compounds and decoys. However, the screening power is sensitive to different initial conformations of the same target. It is expected that our study can provide some guidance for docking-based virtual screening to discover novel type II kinase inhibitors, as well as other allosteric inhibitors.
Chao Shen 0008, Zhe Wang 0041, Youyong Li, Tailong Lei, Ercheng Wang, Lei Xu 0035, Feng Zhu 0004, Dan Li 0013, Tingjun Hou
Briefings Bioinform.10
2020 Fast and accurate prediction of partial charges using Atom-Path-Descriptor-based machine learning
abstract
MOTIVATION: Partial atomic charges are usually used to calculate the electrostatic component of energy in many molecular modeling applications, such as molecular docking, molecular dynamics simulations, free energy calculations and so forth. High-level quantum mechanics calculations may provide the most accurate way to estimate the partial charges for small molecules, but they are too time-consuming to be used to process a large number of molecules for high throughput virtual screening. RESULTS: We proposed a new molecule descriptor named Atom-Path-Descriptor (APD) and developed a set of APD-based machine learning (ML) models to predict the partial charges for small molecules with high accuracy. In the APD algorithm, the 3D structures of molecules were assigned with atom centers and atom-pair path-based atom layers to characterize the local chemical environments of atoms. Then, based on the APDs, two representative ensemble ML algorithms, i.e. random forest (RF) and extreme gradient boosting (XGBoost), were employed to develop the regression models for partial charge assignment. The results illustrate that the RF models based on APDs give better predictions for all the atom types than those based on traditional molecular fingerprints reported in the previous study. More encouragingly, the models trained by XGBoost can improve the predictions of partial charges further, and they can achieve the average root-mean-square error 0.0116 e on the external test set, which is much lower than that (0.0195 e) reported in the previous study, suggesting that the proposed algorithm is quite promising to be used in partial charge assignment with high accuracy. AVAILABILITY AND IMPLEMENTATION: The software framework described in this paper is freely available at https://github.com/jkwang93/Atom-Path-Descriptor-based-machine-learning. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jike Wang, Dong-Sheng Cao 0001, Cunchen Tang, Huiyong Sun, Tingjun Hou
Bioinform.6
2019 farPPI: a webserver for accurate prediction of protein-ligand binding structures for small-molecule PPI inhibitors by MM/PB(GB)SA methods
abstract
SUMMARY: Protein-protein interactions (PPIs) have been regarded as an attractive emerging class of therapeutic targets for the development of new treatments. Computational approaches, especially molecular docking, have been extensively employed to predict the binding structures of PPI-inhibitors or discover novel small molecule PPI inhibitors. However, due to the relatively 'undruggable' features of PPI interfaces, accurate predictions of the binding structures for ligands towards PPI targets are quite challenging for most docking algorithms. Here, we constructed a non-redundant pose ranking benchmark dataset for small-molecule PPI inhibitors, which contains 900 binding poses for 184 protein-ligand complexes. Then, we evaluated the performance of MM/PB(GB)SA approaches to identify the correct binding poses for PPI inhibitors, including two Prime MM/GBSA procedures from the Schrödinger suite and seven different MM/PB(GB)SA procedures from the Amber package. Our results showed that MM/PBSA outperformed the Glide SP scoring function (success rate of 58.6%) and MM/GBSA in most cases, especially the PB3 procedure which could achieve an overall success rate of ∼74%. Moreover, the GB6 procedure (success rate of 68.9%) performed much better than the other MM/GBSA procedures, highlighting the excellent potential of the GBNSR6 implicit solvation model for pose ranking. Finally, we developed the webserver of Fast Amber Rescoring for PPI Inhibitors (farPPI), which offers a freely available service to rescore the docking poses for PPI inhibitors by using the MM/PB(GB)SA methods. AVAILABILITY AND IMPLEMENTATION: farPPI web server is freely available at http://cadd.zju.edu.cn/farppi/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhe Wang 0041, Xuwen Wang, Youyong Li, Tailong Lei, Ercheng Wang, Dan Li 0013, Yu Kang 0002, Feng Zhu 0004, Tingjun Hou
Bioinform.9
2016 CaFE: a tool for binding affinity prediction using end-point free energy methods
abstract
UNLABELLED: Accurate prediction of binding free energy is of particular importance to computational biology and structure-based drug design. Among those methods for binding affinity predictions, the end-point approaches, such as MM/PBSA and LIE, have been widely used because they can achieve a good balance between prediction accuracy and computational cost. Here we present an easy-to-use pipeline tool named Calculation of Free Energy (CaFE) to conduct MM/PBSA and LIE calculations. Powered by the VMD and NAMD programs, CaFE is able to handle numerous static coordinate and molecular dynamics trajectory file formats generated by different molecular simulation packages and supports various force field parameters. AVAILABILITY AND IMPLEMENTATION: CaFE source code and documentation are freely available under the GNU General Public License via GitHub at https://github.com/huiliucode/cafe_plugin It is a VMD plugin written in Tcl and the usage is platform-independent. CONTACT: [email protected].
Hui Liu 0045, Tingjun Hou
Bioinform.2
2014 P-loop Conformation Governed Crizotinib Resistance in G2032R-Mutated ROS1 Tyrosine Kinase: Clues from Free Energy Landscape
abstract
Tyrosine kinases are regarded as excellent targets for chemical drug therapy of carcinomas. However, under strong purifying selection, drug resistance usually occurs in the cancer cells within a short term. Many cases of drug resistance have been found to be associated with secondary mutations in drug target, which lead to the attenuated drug-target interactions. For example, recently, an acquired secondary mutation, G2032R, has been detected in the drug target, ROS1 tyrosine kinase, from a crizotinib-resistant patient, who responded poorly to crizotinib within a very short therapeutic term. It was supposed that the mutation was located at the solvent front and might hinder the drug binding. However, a different fact could be uncovered by the simulations reported in this study. Here, free energy surfaces were characterized by the drug-target distance and the phosphate-binding loop (P-loop) conformational change of the crizotinib-ROS1 complex through advanced molecular dynamics techniques, and it was revealed that the more rigid P-loop region in the G2032R-mutated ROS1 was primarily responsible for the crizotinib resistance, which on one hand, impaired the binding of crizotinib directly, and on the other hand, shortened the residence time induced by the flattened free energy surface. Therefore, both of the binding affinity and the drug residence time should be emphasized in rational drug design to overcome the kinase resistance.
Huiyong Sun, Youyong Li, Junmei Wang, Tingjun Hou
PLoS Comput. Biol.5
2013 Allosite: a method for predicting allosteric sites
abstract
MOTIVATION: The use of allosteric modulators as preferred therapeutic agents against classic orthosteric ligands has colossal advantages, including higher specificity, fewer side effects and lower toxicity. Therefore, the computational prediction of allosteric sites in proteins is receiving increased attention in the field of drug discovery. Allosite is a newly developed automatic tool for the prediction of allosteric sites in proteins of interest and is now available through a web server. AVAILABILITY: The Allosite server and tutorials are freely available at http://mdl.shsmu.edu.cn/AST CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wenkang Huang, Shaoyong Lu, Zhimin Huang, Xinyi Liu 0003, Linkai Mou, Yaqin Liu, Zhongjie Chen, Tingjun Hou, Jian Zhang 0037
Bioinform.10
2011 Prediction of peptides binding to the PKA RIIα subunit using a hierarchical strategy
abstract
MOTIVATION: Favorable interaction between the regulatory subunit of the cAMP-dependent protein kinase (PKA) and a peptide in A-kinase anchoring proteins (AKAPs) is critical for translocating PKA to the subcellular sites where the enzyme phosphorylates its substrates. It is very hard to identify AKAPs peptides binding to PKA due to the high sequence diversity of AKAPs. RESULTS: We propose a hierarchical and efficient approach, which combines molecular dynamics (MD) simulations, free energy calculations, virtual mutagenesis (VM) and bioinformatics analyses, to predict peptides binding to the PKA RIIα regulatory subunit in the human proteome systematically. Our approach successfully retrieved 15 out of 18 documented RIIα-binding peptides. Literature curation supported that many newly predicted peptides might be true AKAPs. Here, we present the first systematic search for AKAP peptides in the human proteome, which is useful to further experimental identification of AKAPs and functional analysis of their biological roles.
Tingjun Hou, Youyong Li, Wei Wang 0051
Bioinform.1
2007 On the detection of functionally coherent groups of protein domains with an extension to protein annotation
abstract
BACKGROUND: Protein domains coordinate to perform multifaceted cellular functions, and domain combinations serve as the functional building blocks of the cell. The available methods to identify functional domain combinations are limited in their scope, e.g. to the identification of combinations falling within individual proteins or within specific regions in a translated genome. Further effort is needed to identify groups of domains that span across two or more proteins and are linked by a cooperative function. Such functional domain combinations can be useful for protein annotation. RESULTS: Using a new computational method, we have identified 114 groups of domains, referred to as domain assembly units (DASSEM units), in the proteome of budding yeast Saccharomyces cerevisiae. The units participate in many important cellular processes such as transcription regulation, translation initiation, and mRNA splicing. Within the units the domains were found to function in a cooperative manner; and each domain contributed to a different aspect of the unit's overall function. The member domains of DASSEM units were found to be significantly enriched among proteins contained in transcription modules, defined as genes sharing similar expression profiles and presumably similar functions. The observation further confirmed the functional coherence of DASSEM units. The functional linkages of units were found in both functionally characterized and uncharacterized proteins, which enabled the assessment of protein function based on domain composition. CONCLUSION: A new computational method was developed to identify groups of domains that are linked by a common function in the proteome of Saccharomyces cerevisiae. These groups can either lie within individual proteins or span across different proteins. We propose that the functional linkages among the domains within the DASSEM units can be used as a non-homology based tool to annotate uncharacterized proteins.
William A. McLaughlin, Ken Chen 0001, Tingjun Hou, Wei Wang 0051
BMC Bioinform.3
2006 Computational Analysis and Prediction of the Binding Motif and Protein Interacting Partners of the Abl SH3 Domain
abstract
Protein-protein interactions, particularly weak and transient ones, are often mediated by peptide recognition domains, such as Src Homology 2 and 3 (SH2 and SH3) domains, which bind to specific sequence and structural motifs. It is important but challenging to determine the binding specificity of these domains accurately and to predict their physiological interacting partners. In this study, the interactions between 35 peptide ligands (15 binders and 20 non-binders) and the Abl SH3 domain were analyzed using molecular dynamics simulation and the Molecular Mechanics/Poisson-Boltzmann Solvent Area method. The calculated binding free energies correlated well with the rank order of the binding peptides and clearly distinguished binders from non-binders. Free energy component analysis revealed that the van der Waals interactions dictate the binding strength of peptides, whereas the binding specificity is determined by the electrostatic interaction and the polar contribution of desolvation. The binding motif of the Abl SH3 domain was then determined by a virtual mutagenesis method, which mutates the residue at each position of the template peptide relative to all other 19 amino acids and calculates the binding free energy difference between the template and the mutated peptides using the Molecular Mechanics/Poisson-Boltzmann Solvent Area method. A single position mutation free energy profile was thus established and used as a scoring matrix to search peptides recognized by the Abl SH3 domain in the human genome. Our approach successfully picked ten out of 13 experimentally determined binding partners of the Abl SH3 domain among the top 600 candidates from the 218,540 decapeptides with the PXXP motif in the SWISS-PROT database. We expect that this physical-principle based method can be applied to other protein domains as well.
Tingjun Hou, Ken Chen 0001, William A. McLaughlin, Wei Wang 0051
PLoS Comput. Biol.1