EDBT 2026 Demo / reviewers in the wild / expert
Honglin Li 0003
dblp:67/5566-3
· DBLP profile ↗
22ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0003-2270-1900ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | COAT-GNN: Cooperative Attribute Learning and Topological Optimization for Protein-Protein Interaction Sites Prediction
Rongfan Tang, Chenglin Wang 0010, Danlin Liu, Jie Zhang 0012, Honglin Li 0003, Kai Zhang 0001 |
DASFAA (3) | 8 |
| 2025 | RPSubAlign: a novel sequence-based molecular representation method for retrosynthesis prediction with improved validity and robustnessabstractRetrosynthetic route planning is essential for designing efficient pathways to synthesize complex molecules, serving as a cornerstone in drug discovery and organic synthesis. Sequence-based models have become a predominant approach in retrosynthetic route planning, yet its validity and robustness remain limited by the challenges in molecular representation methods. Current methods typically treat reactants and products as independent molecules, overlooking structural relationships crucial for accurate synthesis predictions. Herein, we introduce RPSubAlign, a molecular sequence representation method specifically tailored for retrosynthetic tasks, which aligns common substructures between reactants and products to enhance the validity and robustness of sequence-based models. Compared with conventional random and root-alignment representations, RPSubAlign achieves better performance on the USPTO-50K and USPTO-MIT datasets, improving up to a 34.8% increase in Top-N accuracy (with Self-Referencing Embedded Strings representation) and demonstrating enhanced stability across various data augmentation scenarios. RPSubAlign significantly improves syntactic validity, reaching 86.64% on USPTO-50K and 96.45% on USPTO-MIT (with Simplified Molecular Input Line Entry System representation), outperforming baseline methods. These results highlight RPSubAlign as a robust, effective approach for molecular characterization method for retrosynthesis predictions. Code for RPSubAlign is available at https://github.com/Aminoacid1226/RPSubAlign. Hongwen Zhang 0008, Hongling Xu, Jixiang Gao, Wenshuai Deng, Zijing Tian, Qiaoyu Hu, Honglin Li 0003, Yanyan Diao |
Briefings Bioinform. | 9 |
| 2025 | Decoding Drug Response With Structurized Gridding Map-Based Cell RepresentationabstractA thorough understanding of cell-line drug response mechanisms is crucial for drug development, repurposing, and resistance reversal. While targeted anticancer therapies have shown promise, not all cancers have well-established biomarkers to stratify drug response. Single-gene associations only explain a small fraction of the observed drug sensitivity, so a more comprehensive method is needed. However, while deep learning models have shown promise in predicting drug response in cell lines, they still face significant challenges when it comes to their application in clinical applications. Therefore, this study proposed a new strategy called DD-Response for cell-line drug response prediction. First, a limitation of narrow modeling horizons was overcome to expand the model training domain by integrating multiple datasets through source-specific label binarization. Second, a modified representation based on a two-dimensional structurized gridding map (SGM) was developed for cell lines & drugs, avoiding feature correlation neglect and potential information loss. Third, a dual-branch, multi-channel convolutional neural network-based model for pairwise response prediction was constructed, enabling accurate outcomes and improved exploration of underlying mechanisms. As a result, the DD-Response demonstrated superior performance, captured cell-line characteristic variations, and provided insights into key factors impacting cell-line drug response. In addition, DD-Response exhibited scalability in predicting clinical patient responses to drug therapy. Overall, because of DD-response's excellent ability to predict drug response and capture key molecules behind them, DD-response is expected to greatly facilitate drug discovery, repurposing, resistance reversal, and therapeutic optimization. Jiayi Yin, Xiuna Sun, Nanxin You, Minjie Mou, Mingkun Lu, Feng Cheng Li, Honglin Li 0003, Su Zeng, Feng Zhu 0004 |
IEEE J. Biomed. Health Informatics | 9 |
| 2024 | Knowledge Graph Information Bottleneck for Drug-Drug Interaction PredictionabstractDrug-drug interaction (DDI) prediction is an important but challenging task in drug safety surveillance. With the accumulation of biological data, biomedical knowledge graphs (KGs) become available to model DDIs and related biological mechanisms. However, the presence of substantial noise in large-scale KGs hampers prediction performance and the identification of interpretable biological pathways. To fill the gaps, this paper proposes an information bottleneck-based (IB-based) framework that simultaneously denoises the KG and identifies key entities around drug pairs. Moreover, KG-based prediction methods rarely exploit the structural information of drug molecules. To this end, the proposed framework relates drug structures to IB objectives, together with a unique drug pair-centered readout to fuse molecular information into KG subgraph embeddings. Extensive experimental results and case studies demonstrate the effectiveness and interpretability of the framework. Gaoqi He, Kai Zhang 0001, Honglin Li 0003 |
IJCNN | 4 |
| 2024 | Prototype-based contrastive substructure identification for molecular property predictionabstractSubstructure-based representation learning has emerged as a powerful approach to featurize complex attributed graphs, with promising results in molecular property prediction (MPP). However, existing MPP methods mainly rely on manually defined rules to extract substructures. It remains an open challenge to adaptively identify meaningful substructures from numerous molecular graphs to accommodate MPP tasks. To this end, this paper proposes Prototype-based cOntrastive Substructure IdentificaTion (POSIT), a self-supervised framework to autonomously discover substructural prototypes across graphs so as to guide end-to-end molecular fragmentation. During pre-training, POSIT emphasizes two key aspects of substructure identification: firstly, it imposes a soft connectivity constraint to encourage the generation of topologically meaningful substructures; secondly, it aligns resultant substructures with derived prototypes through a prototype-substructure contrastive clustering objective, ensuring attribute-based similarity within clusters. In the fine-tuning stage, a cross-scale attention mechanism is designed to integrate substructure-level information to enhance molecular representations. The effectiveness of the POSIT framework is demonstrated by experimental results from diverse real-world datasets, covering both classification and regression tasks. Moreover, visualization analysis validates the consistency of chemical priors with identified substructures. The source code is publicly available at https://github.com/VRPharmer/POSIT. Gaoqi He, Changbo Wang, Kai Zhang 0001, Honglin Li 0003 |
Briefings Bioinform. | 6 |
| 2024 | MoDAFold: a strategy for predicting the structure of missense mutant protein based on AlphaFold2 and molecular dynamicsabstractProtein structure prediction is a longstanding issue crucial for identifying new drug targets and providing a mechanistic understanding of protein functions. To enhance the progress in this field, a spectrum of computational methodologies has been cultivated. AlphaFold2 has exhibited exceptional precision in predicting wild-type protein structures, with performance exceeding that of other methods. However, predicting the structures of missense mutant proteins using AlphaFold2 remains challenging due to the intricate and substantial structural alterations caused by minor sequence variations in the mutant proteins. Molecular dynamics (MD) has been validated for precisely capturing changes in amino acid interactions attributed to protein mutations. Therefore, for the first time, a strategy entitled 'MoDAFold' was proposed to improve the accuracy and reliability of missense mutant protein structure prediction by combining AlphaFold2 with MD. Multiple case studies have confirmed the superior performance of MoDAFold compared to other methods, particularly AlphaFold2. Lingyan Zheng, Shuiyang Shi, Xiuna Sun, Mingkun Lu, Yang Liao, Sisi Zhu, Hongning Zhang, Pan Fang, Zhenyu Zeng, Honglin Li 0003, Zhaorong Li, Weiwei Xue, Feng Zhu 0004 |
Briefings Bioinform. | 11 |
| 2024 | A Transformative Topological Representation for Link Modeling, Prediction and Cross-Domain Network AnalysisabstractMany complex social, biological, or physical systems are characterized as networks, and recovering the missing links of a network could shed important lights on its structure and dynamics. A good topological representation is crucial to accurate link modeling and prediction, yet how to account for the kaleidoscopic changes in link formation patterns remains a challenge, especially for analysis in cross-domain studies. We propose a new link representation scheme by projecting the local environment of a link into a "dipole plane", where neighboring nodes of the link are positioned via their relative proximity to the two anchors of the link, like a dipole. By doing this, complex and discrete topology arising from link formation is turned to differentiable point-cloud distribution, opening up new possibilities for topological feature-engineering with desired expressiveness, interpretability and generalization. Our approach has comparable or even superior results against state-of-the-art GNNs, meanwhile with a model up to hundreds of times smaller and running much faster. Furthermore, it provides a universal platform to systematically profile, study, and compare link-patterns from miscellaneous real-world networks. This allows building a global link-pattern atlas, based on which we have uncovered interesting common patterns of link formation, i.e., the bridge-style, the radiation-style, and the community-style across a wide collection of networks with highly different nature. Kai Zhang 0001, Junchen Shen, Gaoqi He, Yu Sun 0076, Haibin Ling, Hongyuan Zha, Honglin Li 0003, Jie Zhang 0012 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | MacFrag: segmenting large-scale molecules to obtain diverse fragments with high qualitiesabstractSUMMARY: Construction of high-quality fragment libraries by segmenting organic compounds is an important part of the drug discovery paradigm. This article presents a new method, MacFrag, for efficient molecule fragmentation. MacFrag utilized a modified version of BRICS rules to break chemical bonds and introduced an efficient subgraphs extraction algorithm for rapid enumeration of the fragment space. The evaluation results with ChEMBL dataset exhibited that MacFrag was overall faster than BRICS implemented in RDKit and modified molBLOCKS. Meanwhile, the fragments acquired through MacFrag were more compliant with the 'Rule of Three'. AVAILABILITY AND IMPLEMENTATION: https://github.com/yydiao1025/MacFrag. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yanyan Diao, Zihao Shen, Honglin Li 0003 |
Bioinform. | 4 |
| 2022 | e-TSN: an interactive visual exploration platform for target-disease knowledge mapping from literatureabstractTarget discovery and identification processes are driven by the increasing amount of biomedical data. The vast numbers of unstructured texts of biomedical publications provide a rich source of knowledge for drug target discovery research and demand the development of specific algorithms or tools to facilitate finding disease genes and proteins. Text mining is a method that can automatically mine helpful information related to drug target discovery from massive biomedical literature. However, there is a substantial lag between biomedical publications and the subsequent abstraction of information extracted by text mining to databases. The knowledge graph is introduced to integrate heterogeneous biomedical data. Here, we describe e-TSN (Target significance and novelty explorer, http://www.lilab-ecust.cn/etsn/), a knowledge visualization web server integrating the largest database of associations between targets and diseases from the full scientific literature by constructing significance and novelty scoring methods based on bibliometric statistics. The platform aims to visualize target-disease knowledge graphs to assist in prioritizing candidate disease-related proteins. Approved drugs and associated bioactivities for each interested target are also provided to facilitate the visualization of drug-target relationships. In summary, e-TSN is a fast and customizable visualization resource for investigating and analyzing the intricate target-disease networks, which could help researchers understand the mechanisms underlying complex disease phenotypes and improve the drug discovery and development efficiency, especially for the unexpected outbreak of infectious disease pandemics like COVID-19. Ziyan Feng, Zihao Shen, Honglin Li 0003, Shiliang Li |
Briefings Bioinform. | 3 |
| 2022 | Multi-modal chemical information reconstruction from images and texts for exploring the near-drug spaceabstractIdentification of new chemical compounds with desired structural diversity and biological properties plays an essential role in drug discovery, yet the construction of such a potential space with elements of 'near-drug' properties is still a challenging task. In this work, we proposed a multimodal chemical information reconstruction system to automatically process, extract and align heterogeneous information from the text descriptions and structural images of chemical patents. Our key innovation lies in a heterogeneous data generator that produces cross-modality training data in the form of text descriptions and Markush structure images, from which a two-branch model with image- and text-processing units can then learn to both recognize heterogeneous chemical entities and simultaneously capture their correspondence. In particular, we have collected chemical structures from ChEMBL database and chemical patents from the European Patent Office and the US Patent and Trademark Office using keywords 'A61P, compound, structure' in the years from 2010 to 2020, and generated heterogeneous chemical information datasets with 210K structural images and 7818 annotated text snippets. Based on the reconstructed results and substituent replacement rules, structural libraries of a huge number of near-drug compounds can be generated automatically. In quantitative evaluations, our model can correctly reconstruct 97% of the molecular images into structured format and achieve an F1-score around 97-98% in the recognition of chemical entities, which demonstrated the effectiveness of our model in automatic information extraction from chemical patents, and hopefully transforming them to a user-friendly, structured molecular database enriching the near-drug space to realize the intelligent retrieval technology of chemical knowledge. Jie Wang 0146, Zihao Shen, Yichen Liao, Shiliang Li, Gaoqi He, Man Lan, Xuhong Qian, Kai Zhang 0001, Honglin Li 0003 |
Briefings Bioinform. | 10 |
| 2022 | Comprehensive assessment of deep generative architectures for de novo drug designabstractRecently, deep learning (DL)-based de novo drug design represents a new trend in pharmaceutical research, and numerous DL-based methods have been developed for the generation of novel compounds with desired properties. However, a comprehensive understanding of the advantages and disadvantages of these methods is still lacking. In this study, the performances of different generative models were evaluated by analyzing the properties of the generated molecules in different scenarios, such as goal-directed (rediscovery, optimization and scaffold hopping of active compounds) and target-specific (generation of novel compounds for a given target) tasks. In overall, the DL-based models have significant advantages over the baseline models built by the traditional methods in learning the physicochemical property distributions of the training sets and may be more suitable for target-specific tasks. However, both the baselines and DL-based generative models cannot fully exploit the scaffolds of the training sets, and the molecules generated by the DL-based methods even have lower scaffold diversity than those generated by the traditional models. Moreover, our assessment illustrates that the DL-based methods do not exhibit obvious advantages over the genetic algorithm-based baselines in goal-directed tasks. We believe that our study provides valuable guidance for the effective use of generative models in de novo drug design. Mingyang Wang 0004, Huiyong Sun, Jike Wang, Jinping Pang, Xin Chai, Lei Xu 0035, Honglin Li 0003, Dong-Sheng Cao 0001, Tingjun Hou |
Briefings Bioinform. | 7 |
| 2022 | ncRNAInter: a novel strategy based on graph neural network to discover interactions between lncRNA and miRNAabstractIn recent years, many studies have illustrated the significant role that non-coding RNA (ncRNA) plays in biological activities, in which lncRNA, miRNA and especially their interactions have been proved to affect many biological processes. Some in silico methods have been proposed and applied to identify novel lncRNA-miRNA interactions (LMIs), but there are still imperfections in their RNA representation and information extraction approaches, which imply there is still room for further improving their performances. Meanwhile, only a few of them are accessible at present, which limits their practical applications. The construction of a new tool for LMI prediction is thus imperative for the better understanding of their relevant biological mechanisms. This study proposed a novel method, ncRNAInter, for LMI prediction. A comprehensive strategy for RNA representation and an optimized deep learning algorithm of graph neural network were utilized in this study. ncRNAInter was robust and showed better performance of 26.7% higher Matthews correlation coefficient than existing reputable methods for human LMI prediction. In addition, ncRNAInter proved its universal applicability in dealing with LMIs from various species and successfully identified novel LMIs associated with various diseases, which further verified its effectiveness and usability. All source code and datasets are freely available at https://github.com/idrblab/ncRNAInter. Xiuna Sun, Minjie Mou, Zhaorong Li, Honglin Li 0003, Feng Zhu 0004 |
Briefings Bioinform. | 8 |
| 2022 | VRPharmer: bringing virtual reality into pharmacophore-based virtual screening with interactive exploration and realistic visualizationabstractSUMMARY: Current pharmacophore-based virtual screening (VS) software has limited interactive capabilities and less intuitive screening processes. In this study, a novel tool named VRPharmer is proposed to perform the entire VS workflow in VR environments. VRPharmer enables users to interactively perceive computation processes and immersively observe molecular structures. Besides a typical screening mode (OPT mode), VRPharmer provides a unique interactive screening mode (SCORE mode) for freely exploring the optimal binding poses. Pharmacophore models are editable to study the impact of each feature and further refine the screening results. Moreover, molecular rendering algorithms are improved for precise representations. AVAILABILITY AND IMPLEMENTATION: VRPharmer is open-source software under the MIT license. The released version is available at https://github.com/VRPharmer/VRPharmer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianchao Zhou, Ziyan Feng, Zilong Jin, Chenfei Zhang, Shiliang Li, Gaoqi He, Honglin Li 0003 |
Bioinform. | 10 |
| 2016 | Accurate prediction of RNA-binding protein residues with two discriminative structural descriptorsabstractBACKGROUND: RNA-binding proteins participate in many important biological processes concerning RNA-mediated gene regulation, and several computational methods have been recently developed to predict the protein-RNA interactions of RNA-binding proteins. Newly developed discriminative descriptors will help to improve the prediction accuracy of these prediction methods and provide further meaningful information for researchers. RESULTS: In this work, we designed two structural features (residue electrostatic surface potential and triplet interface propensity) and according to the statistical and structural analysis of protein-RNA complexes, the two features were powerful for identifying RNA-binding protein residues. Using these two features and other excellent structure- and sequence-based features, a random forest classifier was constructed to predict RNA-binding residues. The area under the receiver operating characteristic curve (AUC) of five-fold cross-validation for our method on training set RBP195 was 0.900, and when applied to the test set RBP68, the prediction accuracy (ACC) was 0.868, and the F-score was 0.631. CONCLUSIONS: The good prediction performance of our method revealed that the two newly designed descriptors could be discriminative for inferring protein residues interacting with RNAs. To facilitate the use of our method, a web-server called RNAProSite, which implements the proposed method, was constructed and is freely available at http://lilab.ecust.edu.cn/NABind . Meijian Sun, Chuanxin Zou, Zenghui He, Honglin Li 0003 |
BMC Bioinform. | 6 |
| 2013 | ChemMapper: a versatile web server for exploring pharmacology and chemical structure association based on molecular 3D similarity methodabstractSUMMARY: ChemMapper is an online platform to predict polypharmacology effect and mode of action for small molecules based on 3D similarity computation. ChemMapper collects >350 000 chemical structures with bioactivities and associated target annotations (as well as >3 000 000 non-annotated compounds for virtual screening). Taking the user-provided chemical structure as the query, the top most similar compounds in terms of 3D similarity are returned with associated pharmacology annotations. ChemMapper is designed to provide versatile services in a variety of chemogenomics, drug repurposing, polypharmacology, novel bioactive compounds identification and scaffold hopping studies. AVAILABILITY: http://lilab.ecust.edu.cn/chemmapper/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiayu Gong, Chaoqian Cai, Xiaofeng Liu 0005, Xin Ku, Hualiang Jiang, Daqi Gao, Honglin Li 0003 |
Bioinform. | 7 |
| 2013 | PTID: an integrated web resource and computational tool for agrochemical discoveryabstractSUMMARY: Although in silico drug discovery approaches are crucial for the development of pharmaceuticals, their potential advantages in agrochemical industry have not been realized. The challenge for computer-aided methods in agrochemical arena is a lack of sufficient information for both pesticides and their targets. Therefore, it is important to establish such knowledge repertoire that contains comprehensive pesticides' profiles, which include physicochemical properties, environmental fates, toxicities and mode of actions. Here, we present an integrated platform called Pesticide-Target interaction database (PTID), which comprises a total of 1347 pesticides with rich annotation of ecotoxicological and toxicological data as well as 13 738 interactions of pesticide-target and 4245 protein terms via text mining. Additionally, through the integration of ChemMapper, an in-house computational approach to polypharmacology, PTID can be used as a computational platform to identify pesticides targets and design novel agrochemical products. AVAILABILITY: http://lilab.ecust.edu.cn/ptid/. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiayu Gong, Xiaofeng Liu 0005, Xianwen Cao, Yanyan Diao, Daqi Gao, Honglin Li 0003, Xuhong Qian |
Bioinform. | 6 |
| 2013 | An improved sequence based prediction protocol for DNA-binding proteins using SVM and comprehensive feature analysisabstractBACKGROUND: DNA-binding proteins (DNA-BPs) play a pivotal role in both eukaryotic and prokaryotic proteomes. There have been several computational methods proposed in the literature to deal with the DNA-BPs, many informative features and properties were used and proved to have significant impact on this problem. However the ultimate goal of Bioinformatics is to be able to predict the DNA-BPs directly from primary sequence. RESULTS: In this work, the focus is how to transform these informative features into uniform numeric representation appropriately and improve the prediction accuracy of our SVM-based classifier for DNA-BPs. A systematic representation of some selected features known to perform well is investigated here. Firstly, four kinds of protein properties are obtained and used to describe the protein sequence. Secondly, three different feature transformation methods (OCTD, AC and SAA) are adopted to obtain numeric feature vectors from three main levels: Global, Nonlocal and Local of protein sequence and their performances are exhaustively investigated. At last, the mRMR-IFS feature selection method and ensemble learning approach are utilized to determine the best prediction model. Besides, the optimal features selected by mRMR-IFS are illustrated based on the observed results which may provide useful insights for revealing the mechanisms of protein-DNA interactions. For five-fold cross-validation over the DNAdset and DNAaset, we obtained an overall accuracy of 0.940 and 0.811, MCC of 0.881 and 0.614 respectively. CONCLUSIONS: The good results suggest that it can efficiently develop an entirely sequence-based protocol that transforms and integrates informative features from different scales used by SVM to predict DNA-BPs accurately. Moreover, a novel systematic framework for sequence descriptor-based protein function prediction is proposed here. Chuanxin Zou, Jiayu Gong, Honglin Li 0003 |
BMC Bioinform. | 3 |
| 2010 | Bioactive Conformational Generation of Small Molecules: A Comparative Analysis between Force-Field and Multiple Empirical Criteria Based MethodsabstractBACKGROUND: Conformational sampling for small molecules plays an essential role in drug discovery research pipeline. Based on multi-objective evolution algorithm (MOEA), we have developed a conformational generation method called Cyndi in the previous study. In this work, in addition to Tripos force field in the previous version, Cyndi was updated by incorporation of MMFF94 force field to assess the conformational energy more rationally. With two force fields against a larger dataset of 742 bioactive conformations of small ligands extracted from PDB, a comparative analysis was performed between pure force field based method (FFBM) and multiple empirical criteria based method (MECBM) hybrided with different force fields. RESULTS: Our analysis reveals that incorporating multiple empirical rules can significantly improve the accuracy of conformational generation. MECBM, which takes both empirical and force field criteria as the objective functions, can reproduce about 54% (within 1Å RMSD) of the bioactive conformations in the 742-molecule testset, much higher than that of pure force field method (FFBM, about 37%). On the other hand, MECBM achieved a more complete and efficient sampling of the conformational space because the average size of unique conformations ensemble per molecule is about 6 times larger than that of FFBM, while the time scale for conformational generation is nearly the same as FFBM. Furthermore, as a complementary comparison study between the methods with and without empirical biases, we also tested the performance of the three conformational generation methods in MacroModel in combination with different force fields. Compared with the methods in MacroModel, MECBM is more competitive in retrieving the bioactive conformations in light of accuracy but has much lower computational cost. CONCLUSIONS: By incorporating different energy terms with several empirical criteria, the MECBM method can produce more reasonable conformational ensemble with high accuracy but approximately the same computational cost in comparison with FFBM method. Our analysis also reveals that the performance of conformational generation is irrelevant to the types of force field adopted in characterization of conformational accessibility. Moreover, post energy minimization is not necessary and may even undermine the diversity of conformational ensemble. All the results guide us to explore more empirical criteria like geometric restraints during the conformational process, which may improve the performance of conformational generation in combination with energetic accessibility, regardless of force field types adopted. Fang Bai, Xiaofeng Liu 0005, Haoyun Zhang, Hualiang Jiang, Xicheng Wang, Honglin Li 0003 |
BMC Bioinform. | 7 |
| 2009 | An effective docking strategy for virtual screening based on multi-objective optimization algorithmabstractBACKGROUND: Development of a fast and accurate scoring function in virtual screening remains a hot issue in current computer-aided drug research. Different scoring functions focus on diverse aspects of ligand binding, and no single scoring can satisfy the peculiarities of each target system. Therefore, the idea of a consensus score strategy was put forward. Integrating several scoring functions, consensus score re-assesses the docked conformations using a primary scoring function. However, it is not really robust and efficient from the perspective of optimization. Furthermore, to date, the majority of available methods are still based on single objective optimization design. RESULTS: In this paper, two multi-objective optimization methods, called MOSFOM, were developed for virtual screening, which simultaneously consider both the energy score and the contact score. Results suggest that MOSFOM can effectively enhance enrichment and performance compared with a single score. For three different kinds of binding sites, MOSFOM displays an excellent ability to differentiate active compounds through energy and shape complementarity. EFMOGA performed particularly well in the top 2% of database for all three cases, whereas MOEA_Nrg and MOEA_Cnt performed better than the corresponding individual scoring functions if the appropriate type of binding site was selected. CONCLUSION: The multi-objective optimization method was successfully applied in virtual screening with two different scoring functions that can yield reasonable binding poses and can furthermore, be ranked with the potentially compromised conformations of each compound, abandoning those conformations that can not satisfy overall objective functions. Honglin Li 0003, Hailei Zhang, Mingyue Zheng, Ling Kang, Xiaofeng Liu 0005, Xicheng Wang, Hualiang Jiang |
BMC Bioinform. | 1 |
| 2009 | Cyndi: a multi-objective evolution algorithm based method for bioactive molecular conformational generationabstractBACKGROUND: Conformation generation is a ubiquitous problem in molecule modelling. Many applications require sampling the broad molecular conformational space or perceiving the bioactive conformers to ensure success. Numerous in silico methods have been proposed in an attempt to resolve the problem, ranging from deterministic to non-deterministic and systemic to stochastic ones. In this work, we described an efficient conformation sampling method named Cyndi, which is based on multi-objective evolution algorithm. RESULTS: The conformational perturbation is subjected to evolutionary operation on the genome encoded with dihedral torsions. Various objectives are designated to render the generated Pareto optimal conformers to be energy-favoured as well as evenly scattered across the conformational space. An optional objective concerning the degree of molecular extension is added to achieve geometrically extended or compact conformations which have been observed to impact the molecular bioactivity (J Comput -Aided Mol Des 2002, 16: 105-112). Testing the performance of Cyndi against a test set consisting of 329 small molecules reveals an average minimum RMSD of 0.864 A to corresponding bioactive conformations, indicating Cyndi is highly competitive against other conformation generation methods. Meanwhile, the high-speed performance (0.49 +/- 0.18 seconds per molecule) renders Cyndi to be a practical toolkit for conformational database preparation and facilitates subsequent pharmacophore mapping or rigid docking. The copy of precompiled executable of Cyndi and the test set molecules in mol2 format are accessible in Additional file 1. CONCLUSION: On the basis of MOEA algorithm, we present a new, highly efficient conformation generation method, Cyndi, and report the results of validation and performance studies comparing with other four methods. The results reveal that Cyndi is capable of generating geometrically diverse conformers and outperforms other four multiple conformer generators in the case of reproducing the bioactive conformations against 329 structures. The speed advantage indicates Cyndi is a powerful alternative method for extensive conformational sampling and large-scale conformer database preparation. Xiaofeng Liu 0005, Fang Bai, Sisheng Ouyang, Xicheng Wang, Honglin Li 0003, Hualiang Jiang |
BMC Bioinform. | 5 |
| 2008 | PDTD: a web-accessible protein database for drug target identificationabstractBACKGROUND: Target identification is important for modern drug discovery. With the advances in the development of molecular docking, potential binding proteins may be discovered by docking a small molecule to a repository of proteins with three-dimensional (3D) structures. To complete this task, a reverse docking program and a drug target database with 3D structures are necessary. To this end, we have developed a web server tool, TarFisDock (Target Fishing Docking) http://www.dddc.ac.cn/tarfisdock, which has been used widely by others. Recently, we have constructed a protein target database, Potential Drug Target Database (PDTD), and have integrated PDTD with TarFisDock. This combination aims to assist target identification and validation. DESCRIPTION: PDTD is a web-accessible protein database for in silico target identification. It currently contains >1100 protein entries with 3D structures presented in the Protein Data Bank. The data are extracted from the literatures and several online databases such as TTD, DrugBank and Thomson Pharma. The database covers diverse information of >830 known or potential drug targets, including protein and active sites structures in both PDB and mol2 formats, related diseases, biological functions as well as associated regulating (signaling) pathways. Each target is categorized by both nosology and biochemical function. PDTD supports keyword search function, such as PDB ID, target name, and disease name. Data set generated by PDTD can be viewed with the plug-in of molecular visualization tools and also can be downloaded freely. Remarkably, PDTD is specially designed for target identification. In conjunction with TarFisDock, PDTD can be used to identify binding proteins for small molecules. The results can be downloaded in the form of mol2 file with the binding pose of the probe compound and a list of potential binding targets according to their ranking scores. CONCLUSION: PDTD serves as a comprehensive and unique repository of drug targets. Integrated with TarFisDock, PDTD is a useful resource to identify binding proteins for active compounds or existing drugs. Its potential applications include in silico drug target identification, virtual screening, and the discovery of the secondary effects of an old drug (i.e. new pharmacological usage) or an existing target (i.e. new pharmacological or toxic relevance), thus it may be a valuable platform for the pharmaceutical researchers. PDTD is available online at http://www.dddc.ac.cn/pdtd/. Zhenting Gao, Honglin Li 0003, Hailei Zhang, Xiaofeng Liu 0005, Ling Kang, Xiaomin Luo, Weiliang Zhu, Kaixian Chen, Xicheng Wang, Hualiang Jiang |
BMC Bioinform. | 2 |
| 2008 | Mechanics of Channel Gating of the Nicotinic Acetylcholine ReceptorabstractThe nicotinic acetylcholine receptor (nAChR) is a key molecule involved in the propagation of signals in the central nervous system and peripheral synapses. Although numerous computational and experimental studies have been performed on this receptor, the structural dynamics of the receptor underlying the gating mechanism is still unclear. To address the mechanical fundamentals of nAChR gating, both conventional molecular dynamics (CMD) and steered rotation molecular dynamics (SRMD) simulations have been conducted on the cryo-electron microscopy (cryo-EM) structure of nAChR embedded in a dipalmitoylphosphatidylcholine (DPPC) bilayer and water molecules. A 30-ns CMD simulation revealed a collective motion amongst C-loops, M1, and M2 helices. The inward movement of C-loops accompanying the shrinking of acetylcholine (ACh) binding pockets induced an inward and upward motion of the outer beta-sheet composed of beta9 and beta10 strands, which in turn causes M1 and M2 to undergo anticlockwise motions around the pore axis. Rotational motion of the entire receptor around the pore axis and twisting motions among extracellular (EC), transmembrane (TM), and intracellular MA domains were also detected by the CMD simulation. Moreover, M2 helices undergo a local twisting motion synthesized by their bending vibration and rotation. The hinge of either twisting motion or bending vibration is located at the middle of M2, possibly the gate of the receptor. A complementary twisting-to-open motion throughout the receptor was detected by a normal mode analysis (NMA). To mimic the pulsive action of ACh binding, nonequilibrium MD simulations were performed by using the SRMD method developed in one of our laboratories. The result confirmed all the motions derived from the CMD simulation and NMA. In addition, the SRMD simulation indicated that the channel may undergo an open-close (O <--> C) motion. The present MD simulations explore the structural dynamics of the receptor under its gating process and provide a new insight into the gating mechanism of nAChR at the atomic level. Xinli Liu, Yechun Xu, Honglin Li 0003, Xicheng Wang, Hualiang Jiang, Francisco J. Barrantes |
PLoS Comput. Biol. | 3 |