EDBT 2026 Demo / reviewers in the wild / expert
Yu Kang 0002
dblp:98/3089-2
· DBLP profile ↗
8ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-0999-8802ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving the predictive performance of binding affinities and poses for protein-cyclic peptide complexes through fine-tuned MM/PBSA(GBSA)-based methodsabstractCyclic peptides represent a highly promising class of biopharmaceutical scaffolds. The screening of cyclic peptides against protein targets can be greatly facilitated using computational approaches, especially molecular docking. However, it remains a crucial challenge to accurately predict protein-cyclic peptide (P-cp) interactions employing scoring functions of molecular docking. End-point approaches, such as molecular mechanics generalized Born surface area (MM/GBSA) and molecular mechanics Poisson-Boltzmann surface area (MM/PBSA), provide theoretically more robust frameworks than conventional scoring functions, but their reliability in predicting binding affinities and discriminating native-like binding poses for P-cp complexes remains poorly quantified. Herein, we comprehensively assessed the predictive abilities of MM/PBSA(GBSA) in scoring binding affinities of P-cp complexes and re-ranking their binding poses. The binding affinity scoring ability of MM/PBSA(GBSA) was assessed on a carefully curated dataset consisting of 50 complexes involving P-cp binding affinities, and their re-ranking capability was evaluated on another dataset consisting of the decoys of 81 P-cp complexes. Based on these assessments, we proposed a two-step workflow for predicting P-cp binding affinities. First, we employed the assessed optimal re-ranking method to select the top-1 binding pose; second, we estimated the binding affinity based on the selected top-1 pose using the assessed optimal scoring method. Our proposed workflow, which requires only 3 s for each prediction, achieves binding affinity predictions with a Rp of -0.732 when compared to experimental values, which is twice as high as that of AutoDock CrankPep (Rp = -0.316). This study emphasizes the necessity of using fine-tuned MM/PBSA(GBSA) methods for predicting P-cp interactions. Huifeng Zhao, Jianxiang Huang, Gaoqi Weng, Dejun Jiang 0002, Renling Hu, Yu Kang 0002, Tingjun Hou |
Briefings Bioinform. | 6 |
| 2024 | AttABseq: an attention-based deep learning prediction method for antigen-antibody binding affinity changes based on protein sequencesabstractThe optimization of therapeutic antibodies through traditional techniques, such as candidate screening via hybridoma or phage display, is resource-intensive and time-consuming. In recent years, computational and artificial intelligence-based methods have been actively developed to accelerate and improve the development of therapeutic antibodies. In this study, we developed an end-to-end sequence-based deep learning model, termed AttABseq, for the predictions of the antigen-antibody binding affinity changes connected with antibody mutations. AttABseq is a highly efficient and generic attention-based model by utilizing diverse antigen-antibody complex sequences as the input to predict the binding affinity changes of residue mutations. The assessment on the three benchmark datasets illustrates that AttABseq is 120% more accurate than other sequence-based models in terms of the Pearson correlation coefficient between the predicted and experimental binding affinity changes. Moreover, AttABseq also either outperforms or competes favorably with the structure-based approaches. Furthermore, AttABseq consistently demonstrates robust predictive capabilities across a diverse array of conditions, underscoring its remarkable capacity for generalization across a wide spectrum of antigen-antibody complexes. It imposes no constraints on the quantity of altered residues, rendering it particularly applicable in scenarios where crystallographic structures remain unavailable. The attention-based interpretability analysis indicates that the causal effects of point mutations on antibody-antigen binding affinity changes can be visualized at the residue level, which might assist automated antibody sequence optimization. We believe that AttABseq provides a fiercely competitive answer to therapeutic antibody optimization. Ruofan Jin, Jike Wang, Dejun Jiang 0002, Tianyue Wang, Yu Kang 0002, Wanting Xu, Chang-Yu Hsieh, Tingjun Hou |
Briefings Bioinform. | 7 |
| 2024 | Comprehensive assessment of protein loop modeling programs on large-scale datasets: prediction accuracy and efficiencyabstractProtein loops play a critical role in the dynamics of proteins and are essential for numerous biological functions, and various computational approaches to loop modeling have been proposed over the past decades. However, a comprehensive understanding of the strengths and weaknesses of each method is lacking. In this work, we constructed two high-quality datasets (i.e. the General dataset and the CASP dataset) and systematically evaluated the accuracy and efficiency of 13 commonly used loop modeling approaches from the perspective of loop lengths, protein classes and residue types. The results indicate that the knowledge-based method FREAD generally outperforms the other tested programs in most cases, but encountered challenges when predicting loops longer than 15 and 30 residues on the CASP and General datasets, respectively. The ab initio method Rosetta NGK demonstrated exceptional modeling accuracy for short loops with four to eight residues and achieved the highest success rate on the CASP dataset. The well-known AlphaFold2 and RoseTTAFold require more resources for better performance, but they exhibit promise for predicting loops longer than 16 and 30 residues in the CASP and General datasets. These observations can provide valuable insights for selecting suitable methods for specific loop modeling tasks and contribute to future advancements in the field. Tianyue Wang, Langcheng Wang, Xujun Zhang, Chao Shen 0008, Odin Zhang, Jike Wang, Jialu Wu, Ruofan Jin, Shicheng Chen, Chang-Yu Hsieh, Guangyong Chen, Peichen Pan, Yu Kang 0002, Tingjun Hou |
Briefings Bioinform. | 16 |
| 2023 | Can molecular dynamics simulations improve predictions of protein-ligand binding affinity with machine learning?abstractBinding affinity prediction largely determines the discovery efficiency of lead compounds in drug discovery. Recently, machine learning (ML)-based approaches have attracted much attention in hopes of enhancing the predictive performance of traditional physics-based approaches. In this study, we evaluated the impact of structural dynamic information on the binding affinity prediction by comparing the models trained on different dimensional descriptors, using three targets (i.e. JAK1, TAF1-BD2 and DDR1) and their corresponding ligands as the examples. Here, 2D descriptors are traditional ECFP4 fingerprints, 3D descriptors are the energy terms of the Smina and NNscore scoring functions and 4D descriptors contain the structural dynamic information derived from the trajectories based on molecular dynamics (MD) simulations. We systematically investigate the MD-refined binding affinity prediction performance of three classical ML algorithms (i.e. RF, SVR and XGB) as well as two common virtual screening methods, namely Glide docking and MM/PBSA. The outcomes of the ML models built using various dimensional descriptors and their combinations reveal that the MD refinement with the optimized protocol can improve the predictive performance on the TAF1-BD2 target with considerable structural flexibility, but not for the less flexible JAK1 and DDR1 targets, when taking docking poses as the initial structure instead of the crystal structures. The results highlight the importance of the initial structures to the final performance of the model through conformational analysis on the three targets with different flexibility. Shukai Gu, Chao Shen 0008, Huanxiang Liu, Rong Sheng, Lei Xu 0035, Zhe Wang 0041, Tingjun Hou, Yu Kang 0002 |
Briefings Bioinform. | 11 |
| 2023 | ML-PLIC: a web platform for characterizing protein-ligand interactions and developing machine learning-based scoring functionsabstractCracking the entangling code of protein-ligand interaction (PLI) is of great importance to structure-based drug design and discovery. Different physical and biochemical representations can be used to describe PLI such as energy terms and interaction fingerprints, which can be analyzed by machine learning (ML) algorithms to create ML-based scoring functions (MLSFs). Here, we propose the ML-based PLI capturer (ML-PLIC), a web platform that automatically characterizes PLI and generates MLSFs to identify the potential binders of a specific protein target through virtual screening (VS). ML-PLIC comprises five modules, including Docking for ligand docking, Descriptors for PLI generation, Modeling for MLSF training, Screening for VS and Pipeline for the integration of the aforementioned functions. We validated the MLSFs constructed by ML-PLIC in three benchmark datasets (Directory of Useful Decoys-Enhanced, Active as Decoys and TocoDecoy), demonstrating accuracy outperforming traditional docking tools and competitive performance to the deep learning-based SF, and provided a case study of the Serine/threonine-protein kinase WEE1 in which MLSFs were developed by using the ML-based VS pipeline in ML-PLIC. Underpinning the latest version of ML-PLIC is a powerful platform that incorporates physical and biological knowledge about PLI, leveraging PLI characterization and MLSF generation into the design of structure-based VS pipeline. The ML-PLIC web platform is now freely available at http://cadd.zju.edu.cn/plic/. Xujun Zhang, Chao Shen 0008, Tianyue Wang, Yafeng Deng, Yu Kang 0002, Dan Li 0013, Tingjun Hou, Peichen Pan |
Briefings Bioinform. | 5 |
| 2022 | fastDRH: a webserver to predict and analyze protein-ligand complexes based on molecular docking and MM/PB(GB)SA computationabstractPredicting the native or near-native binding pose of a small molecule within a protein binding pocket is an extremely important task in structure-based drug design, especially in the hit-to-lead and lead optimization phases. In this study, fastDRH, a free and open accessed web server, was developed to predict and analyze protein-ligand complex structures. In fastDRH server, AutoDock Vina and AutoDock-GPU docking engines, structure-truncated MM/PB(GB)SA free energy calculation procedures and multiple poses based per-residue energy decomposition analysis were well integrated into a user-friendly and multifunctional online platform. Benefit from the modular architecture, users can flexibly use one or more of three features, including molecular docking, docking pose rescoring and hotspot residue prediction, to obtain the key information clearly based on a result analysis panel supported by 3Dmol.js and Apache ECharts. In terms of protein-ligand binding mode prediction, the integrated structure-truncated MM/PB(GB)SA rescoring procedures exhibit a success rate of >80% in benchmark, which is much better than the AutoDock Vina (~70%). For hotspot residue identification, our multiple poses based per-residue energy decomposition analysis strategy is a more reliable solution than the one using only a single pose, and the performance of our solution has been experimentally validated in several drug discovery projects. To summarize, the fastDRH server is a useful tool for predicting the ligand binding mode and the hotspot residue of protein for ligand binding. The fastDRH server is accessible free of charge at http://cadd.zju.edu.cn/fastdrh/. Zhe Wang 0041, Huiyong Sun, Yu Kang 0002, Huanxiang Liu, Dong-Sheng Cao 0001, Tingjun Hou |
Briefings Bioinform. | 4 |
| 2021 | Do we need different machine learning algorithms for QSAR modeling? A comprehensive assessment of 16 machine learning algorithms on 14 QSAR data setsabstractAlthough a wide variety of machine learning (ML) algorithms have been utilized to learn quantitative structure-activity relationships (QSARs), there is no agreed single best algorithm for QSAR learning. Therefore, a comprehensive understanding of the performance characteristics of popular ML algorithms used in QSAR learning is highly desirable. In this study, five linear algorithms [linear function Gaussian process regression (linear-GPR), linear function support vector machine (linear-SVM), partial least squares regression (PLSR), multiple linear regression (MLR) and principal component regression (PCR)], three analogizers [radial basis function support vector machine (rbf-SVM), K-nearest neighbor (KNN) and radial basis function Gaussian process regression (rbf-GPR)], six symbolists [extreme gradient boosting (XGBoost), Cubist, random forest (RF), multiple adaptive regression splines (MARS), gradient boosting machine (GBM), and classification and regression tree (CART)] and two connectionists [principal component analysis artificial neural network (pca-ANN) and deep neural network (DNN)] were employed to learn the regression-based QSAR models for 14 public data sets comprising nine physicochemical properties and five toxicity endpoints. The results show that rbf-SVM, rbf-GPR, XGBoost and DNN generally illustrate better performances than the other algorithms. The overall performances of different algorithms can be ranked from the best to the worst as follows: rbf-SVM > XGBoost > rbf-GPR > Cubist > GBM > DNN > RF > pca-ANN > MARS > linear-GPR ≈ KNN > linear-SVM ≈ PLSR > CART ≈ PCR ≈ MLR. In terms of prediction accuracy and computational efficiency, SVM and XGBoost are recommended to the regression learning for small data sets, and XGBoost is an excellent choice for large data sets. We then investigated the performances of the ensemble models by integrating the predictions of multiple ML algorithms. The results illustrate that the ensembles of two or three algorithms in different categories can indeed improve the predictions of the best individual ML algorithms. Zhenxing Wu, Yu Kang 0002, Elaine Lai-Han Leung, Tailong Lei, Chao Shen 0008, Dejun Jiang 0002, Zhe Wang 0041, Dong-Sheng Cao 0001, Tingjun Hou |
Briefings Bioinform. | 3 |
| 2019 | farPPI: a webserver for accurate prediction of protein-ligand binding structures for small-molecule PPI inhibitors by MM/PB(GB)SA methodsabstractSUMMARY: Protein-protein interactions (PPIs) have been regarded as an attractive emerging class of therapeutic targets for the development of new treatments. Computational approaches, especially molecular docking, have been extensively employed to predict the binding structures of PPI-inhibitors or discover novel small molecule PPI inhibitors. However, due to the relatively 'undruggable' features of PPI interfaces, accurate predictions of the binding structures for ligands towards PPI targets are quite challenging for most docking algorithms. Here, we constructed a non-redundant pose ranking benchmark dataset for small-molecule PPI inhibitors, which contains 900 binding poses for 184 protein-ligand complexes. Then, we evaluated the performance of MM/PB(GB)SA approaches to identify the correct binding poses for PPI inhibitors, including two Prime MM/GBSA procedures from the Schrödinger suite and seven different MM/PB(GB)SA procedures from the Amber package. Our results showed that MM/PBSA outperformed the Glide SP scoring function (success rate of 58.6%) and MM/GBSA in most cases, especially the PB3 procedure which could achieve an overall success rate of ∼74%. Moreover, the GB6 procedure (success rate of 68.9%) performed much better than the other MM/GBSA procedures, highlighting the excellent potential of the GBNSR6 implicit solvation model for pose ranking. Finally, we developed the webserver of Fast Amber Rescoring for PPI Inhibitors (farPPI), which offers a freely available service to rescore the docking poses for PPI inhibitors by using the MM/PB(GB)SA methods. AVAILABILITY AND IMPLEMENTATION: farPPI web server is freely available at http://cadd.zju.edu.cn/farppi/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhe Wang 0041, Xuwen Wang, Youyong Li, Tailong Lei, Ercheng Wang, Dan Li 0013, Yu Kang 0002, Feng Zhu 0004, Tingjun Hou |
Bioinform. | 7 |