EDBT 2026 Demo / reviewers in the wild / expert
Yuguang Mu
dblp:42/10998
· DBLP profile ↗
15ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-2499-026XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Could statistical potential models achieve comparable or better performance than deep learning models?abstractAccurately predicting protein-ligand interactions is vital for structure-based drug discovery. Although deep learning (DL) models have shown strong performance, the potential of traditional statistical potentials under data-limited conditions remains underexplored. Here, we systematically assess several statistical potential models in docking and virtual screening. We find that docking benefits from distance-dependent pairwise atom-atom potentials with clear physical meanings, while screening relies more on orientation-dependent atom-residue potentials that capture local chemical environments. Based on these findings, we propose HybridSP, a hybrid potential combining distance-dependent atom-atom, atom-residue, and orientation-dependent atom-residue terms. An affinity-weighted scheme is applied to correct biases in statistical distributions. On the CASF-2016 benchmark, HybridSP achieves a 91.6% docking success rate and an enrichment factor of 29.35 at the top 1%, rivaling and even surpassing state-of-the-art DL models. Its strong screening ability is further validated on directory of useful decoys-enhanced and directory of useful decoys-adjusted. These results demonstrate that well-designed statistical potentials can achieve high performance and interpretability without complex DL architectures, offering an efficient alternative for scoring function design. The models are available at: https://github.com/zelixirSH/HybridSP.git. Sheng Wang 0001, Yuguang Mu, Liangzhen Zheng |
Briefings Bioinform. | 4 |
| 2025 | Enhancing Bioactivity Prediction via Spatial Emptiness Representation of Protein-ligand Complex and Union of Multiple PocketsabstractPredicting the bioactivity of candidate ligands remains a central challenge in drug discovery. Ligands and endogenous substrates often compete for the same binding sites on target proteins, and the extent to which a ligand can modulate protein function depends not only on its binding but also on how effectively it occupies the relevant pocket. However, most existing methods focus narrowly on local interactions within protein–ligand complexes and neglect spatial emptiness—the unoccupied regions within the binding site that may permit endogenous molecules to engage or interfere. Such unfilled space can diminish the ligand’s functional impact, regardless of binding affinity. To overcome this key limitation in protein–ligand modeling, we propose LigoSpace, a novel method integrating three core components. LigoSpace introduces GeoREC (Geometric Representation of Spatial Emptiness in Complexes) to quantify atomic-level empty space and Union-Pocket to unify multiple protein pockets, providing a global view of binding sites. Additionally, LigoSpace employs a pairwise loss instead of commonly used MSE loss, to better capture relative relationships critical for drug discovery. Extensive experiments on multiple datasets with diverse bioactivity types demonstrate that LigoSpace significantly improves performance when integrated into state-of-the-art models, highlighting the effectiveness of its novel components. Zhiyuan Zhou 0001, Yueming Yin, Yuguang Mu, Hoi-Yeung Li, Adams Wai-Kin Kong |
NeurIPS | 4 |
| 2025 | BridgeNet: a high-efficiency framework integrating sequence and structure for protein and enzyme function predictionabstractUnderstanding the relationship between protein sequences and structures is essential for accurate protein property prediction. We propose BridgeNet, a pre-trained deep learning framework that integrates sequence and structural information through a novel latent environment matrix, enabling seamless alignment of these two modalities. The model's modular architecture-comprising sequence encoding, structural encoding, and a bridge module-effectively captures complementary features without requiring explicit structural inputs during inference. Extensive evaluations on tasks such as enzyme classification, Gene Ontology annotation, coenzyme specificity prediction, and peptide toxicity prediction demonstrate its superior performance over state-of-the-art models. BridgeNet provides a scalable and robust solution, advancing protein representation learning and enabling applications in computational biology and structural bioinformatics. Hongliang Duan, Yuguang Mu |
Briefings Bioinform. | 3 |
| 2024 | Protein language models are performant in structure-free virtual screeningabstractHitherto virtual screening (VS) has been typically performed using a structure-based drug design paradigm. Such methods typically require the use of molecular docking on high-resolution three-dimensional structures of a target protein-a computationally-intensive and time-consuming exercise. This work demonstrates that by employing protein language models and molecular graphs as inputs to a novel graph-to-transformer cross-attention mechanism, a screening power comparable to state-of-the-art structure-based models can be achieved. The implications thereof include highly expedited VS due to the greatly reduced compute required to run this model, and the ability to perform early stages of computer-aided drug design in the complete absence of 3D protein structures. Hilbert Yuen In Lam, Jia Sheng Guan, Xing Er Ong, Robbe Pincket, Yuguang Mu |
Briefings Bioinform. | 5 |
| 2024 | RmsdXNA: RMSD prediction of nucleic acid-ligand docking poses using machine-learning methodabstractSmall molecule drugs can be used to target nucleic acids (NA) to regulate biological processes. Computational modeling methods, such as molecular docking or scoring functions, are commonly employed to facilitate drug design. However, the accuracy of the scoring function in predicting the closest-to-native docking pose is often suboptimal. To overcome this problem, a machine learning model, RmsdXNA, was developed to predict the root-mean-square-deviation (RMSD) of ligand docking poses in NA complexes. The versatility of RmsdXNA has been demonstrated by its successful application to various complexes involving different types of NA receptors and ligands, including metal complexes and short peptides. The predicted RMSD by RmsdXNA was strongly correlated with the actual RMSD of the docked poses. RmsdXNA also outperformed the rDock scoring function in ranking and identifying closest-to-native docking poses across different structural groups and on the testing dataset. Using experimental validated results conducted on polyadenylated nuclear element for nuclear expression triplex, RmsdXNA demonstrated better screening power for the RNA-small molecule complex compared to rDock. Molecular dynamics simulations were subsequently employed to validate the binding of top-scoring ligand candidates selected by RmsdXNA and rDock on MALAT1. The results showed that RmsdXNA has a higher success rate in identifying promising ligands that can bind well to the receptor. The development of an accurate docking score for a NA-ligand complex can aid in drug discovery and development advancements. The code to use RmsdXNA is available at the GitHub repository https://github.com/laiheng001/RmsdXNA. Lai Heng Tan, Chee Keong Kwoh 0001, Yuguang Mu |
Briefings Bioinform. | 3 |
| 2024 | A new paradigm for applying deep learning to protein-ligand interaction predictionabstractProtein-ligand interaction prediction presents a significant challenge in drug design. Numerous machine learning and deep learning (DL) models have been developed to accurately identify docking poses of ligands and active compounds against specific targets. However, current models often suffer from inadequate accuracy or lack practical physical significance in their scoring systems. In this research paper, we introduce IGModel, a novel approach that utilizes the geometric information of protein-ligand complexes as input for predicting the root mean square deviation of docking poses and the binding strength (pKd, the negative value of the logarithm of binding affinity) within the same prediction framework. This ensures that the output scores carry intuitive meaning. We extensively evaluate the performance of IGModel on various docking power test sets, including the CASF-2016 benchmark, PDBbind-CrossDocked-Core and DISCO set, consistently achieving state-of-the-art accuracies. Furthermore, we assess IGModel's generalizability and robustness by evaluating it on unbiased test sets and sets containing target structures generated by AlphaFold2. The exceptional performance of IGModel on these sets demonstrates its efficacy. Additionally, we visualize the latent space of protein-ligand interactions encoded by IGModel and conduct interpretability analysis, providing valuable insights. This study presents a novel framework for DL-based prediction of protein-ligand interactions, contributing to the advancement of this field. The IGModel is available at GitHub repository https://github.com/zchwang/IGModel. Zechen Wang, Sheng Wang 0001, Yanjie Wei, Yuguang Mu, Liangzhen Zheng |
Briefings Bioinform. | 6 |
| 2024 | OpenDock: a pytorch-based open-source framework for protein-ligand docking and modellingabstractMOTIVATION: Molecular docking is an invaluable computational tool with broad applications in computer-aided drug design and enzyme engineering. However, current molecular docking tools are typically implemented in languages such as C++ for calculation speed, which lack flexibility and user-friendliness for further development. Moreover, validating the effectiveness of external scoring functions for molecular docking and screening within these frameworks is challenging, and implementing more efficient sampling strategies is not straightforward. RESULTS: To address these limitations, we have developed an open-source molecular docking framework, OpenDock, based on Python and PyTorch. This framework supports the integration of multiple scoring functions; some can be utilized during molecular docking and pose optimization, while others can be used for post-processing scoring. In terms of sampling, the current version of this framework supports simulated annealing and Monte Carlo optimization. Additionally, it can be extended to include methods such as genetic algorithms and particle swarm optimization for sampling docking poses and protein side chain orientations. Distance constraints are also implemented to enable covalent docking, restricted docking or distance map constraints guided pose sampling. Overall, this framework serves as a valuable tool in drug design and enzyme engineering, offering significant flexibility for most protein-ligand modelling tasks. AVAILABILITY AND IMPLEMENTATION: OpenDock is publicly available at: https://github.com/guyuehuo/opendock. Qiuyue Hu, Zechen Wang, Jintao Meng 0001, Yuguang Mu, Sheng Wang 0001, Liangzhen Zheng, Yanjie Wei |
Bioinform. | 6 |
| 2024 | Systematic benchmarking of deep-learning methods for tertiary RNA structure predictionabstractThe 3D structure of RNA critically influences its functionality, and understanding this structure is vital for deciphering RNA biology. Experimental methods for determining RNA structures are labour-intensive, expensive, and time-consuming. Computational approaches have emerged as valuable tools, leveraging physics-based-principles and machine learning to predict RNA structures rapidly. Despite advancements, the accuracy of computational methods remains modest, especially when compared to protein structure prediction. Deep learning methods, while successful in protein structure prediction, have shown some promise for RNA structure prediction as well, but face unique challenges. This study systematically benchmarks state-of-the-art deep learning methods for RNA structure prediction across diverse datasets. Our aim is to identify factors influencing performance variation, such as RNA family diversity, sequence length, RNA type, multiple sequence alignment (MSA) quality, and deep learning model architecture. We show that generally ML-based methods perform much better than non-ML methods on most RNA targets, although the performance difference isn't substantial when working with unseen novel or synthetic RNAs. The quality of the MSA and secondary structure prediction both play an important role and most methods aren't able to predict non-Watson-Crick pairs in the RNAs. Overall among the automated 3D RNA structure prediction methods, DeepFoldRNA has the best prediction results followed by DRFold as the second best method. Finally, we also suggest possible mitigations to improve the quality of the prediction for future method development. Akash Bahai, Chee Keong Kwoh 0001, Yuguang Mu |
PLoS Comput. Biol. | 3 |
| 2024 | Advancing Bioactivity Prediction Through Molecular Docking and Self-AttentionabstractBioactivity refers to the ability of a substance to induce biological effects within living systems, often describing the influence of molecules, drugs, or chemicals on organisms. In drug discovery, predicting bioactivity streamlines early-stage candidate screening by swiftly identifying potential active molecules. The popular deep learning methods in bioactivity prediction primarily model the ligand structure-bioactivity relationship under the premise of Quantitative Structure-Activity Relationship (QSAR). However, bioactivity is determined by multiple factors, including not only the ligand structure but also drug-target interactions, signaling pathways, reaction environments, pharmacokinetic properties, and species differences. Our study first integrates drug-target interactions into bioactivity prediction using protein-ligand complex data from molecular docking. We devise a Drug-Target Interaction Graph Neural Network (DTIGN), infusing interatomic forces into intermolecular graphs. DTIGN employs multi-head self-attention to identify native-like binding pockets and poses within molecular docking results. To validate the fidelity of the self-attention mechanism, we gather ground truth data from crystal structure databases. Subsequently, we employ these limited native structures to refine bioactivity prediction via semi-supervised learning. For this study, we establish a unique benchmark dataset for evaluating bioactivity prediction models in the context of protein-ligand complexes, showcasing the superior performance of our method (with an average improvement of 27.03%) through comparison with 9 leading deep learning-based bioactivity prediction methods. Yueming Yin, Hilbert Yuen In Lam, Yuguang Mu, Hoi-Yeung Li, Adams Wai-Kin Kong |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | A fully differentiable ligand pose optimization framework guided by deep learning and a traditional scoring functionabstractThe recently reported machine learning- or deep learning-based scoring functions (SFs) have shown exciting performance in predicting protein-ligand binding affinities with fruitful application prospects. However, the differentiation between highly similar ligand conformations, including the native binding pose (the global energy minimum state), remains challenging that could greatly enhance the docking. In this work, we propose a fully differentiable, end-to-end framework for ligand pose optimization based on a hybrid SF called DeepRMSD+Vina combined with a multi-layer perceptron (DeepRMSD) and the traditional AutoDock Vina SF. The DeepRMSD+Vina, which combines (1) the root mean square deviation (RMSD) of the docking pose with respect to the native pose and (2) the AutoDock Vina score, is fully differentiable; thus is capable of optimizing the ligand binding pose to the energy-lowest conformation. Evaluated by the CASF-2016 docking power dataset, the DeepRMSD+Vina reaches a success rate of 94.4%, which outperforms most reported SFs to date. We evaluated the ligand conformation optimization framework in practical molecular docking scenarios (redocking and cross-docking tasks), revealing the high potentialities of this framework in drug design and discovery. Structural analysis shows that this framework has the ability to identify key physical interactions in protein-ligand binding, such as hydrogen-bonding. Our work provides a paradigm for optimizing ligand conformations based on deep learning algorithms. The DeepRMSD+Vina model and the optimization framework are available at GitHub repository https://github.com/zchwang/DeepRMSD-Vina_Optimization. Zechen Wang, Liangzhen Zheng, Sheng Wang 0001, Mingzhi Lin, Adams Wai-Kin Kong, Yuguang Mu, Yanjie Wei |
Briefings Bioinform. | 7 |
| 2022 | Improving protein-ligand docking and screening accuracies by incorporating a scoring function correction termabstractScoring functions are important components in molecular docking for structure-based drug discovery. Traditional scoring functions, generally empirical- or force field-based, are robust and have proven to be useful for identifying hits and lead optimizations. Although multiple highly accurate deep learning- or machine learning-based scoring functions have been developed, their direct applications for docking and screening are limited. We describe a novel strategy to develop a reliable protein-ligand scoring function by augmenting the traditional scoring function Vina score using a correction term (OnionNet-SFCT). The correction term is developed based on an AdaBoost random forest model, utilizing multiple layers of contacts formed between protein residues and ligand atoms. In addition to the Vina score, the model considerably enhances the AutoDock Vina prediction abilities for docking and screening tasks based on different benchmarks (such as cross-docking dataset, CASF-2016, DUD-E and DUD-AD). Furthermore, our model could be combined with multiple docking applications to increase pose selection accuracies and screening abilities, indicating its wide usage for structure-based drug discoveries. Furthermore, in a reverse practice, the combined scoring strategy successfully identified multiple known receptors of a plant hormone. To summarize, the results show that the combination of data-driven model (OnionNet-SFCT) and empirical scoring function (Vina score) is a good scoring strategy that could be useful for structure-based drug discoveries and potentially target fishing in future. Liangzhen Zheng, Jintao Meng 0001, Haidong Lan, Zechen Wang, Mingzhi Lin, Yanjie Wei, Yuguang Mu |
Briefings Bioinform. | 10 |
| 2015 | Fast, accurate, and reliable molecular docking with QuickVina 2abstractMOTIVATION: The need for efficient molecular docking tools for high-throughput screening is growing alongside the rapid growth of drug-fragment databases. AutoDock Vina ('Vina') is a widely used docking tool with parallelization for speed. QuickVina ('QVina 1') then further enhanced the speed via a heuristics, requiring high exhaustiveness. With low exhaustiveness, its accuracy was compromised. We present in this article the latest version of QuickVina ('QVina 2') that inherits both the speed of QVina 1 and the reliability of the original Vina. RESULTS: We tested the efficacy of QVina 2 on the core set of PDBbind 2014. With the default exhaustiveness level of Vina (i.e. 8), a maximum of 20.49-fold and an average of 2.30-fold acceleration with a correlation coefficient of 0.967 for the first mode and 0.911 for the sum of all modes were attained over the original Vina. A tendency for higher acceleration with increased number of rotatable bonds as the design variables was observed. On the accuracy, Vina wins over QVina 2 on 30% of the data with average energy difference of only 0.58 kcal/mol. On the same dataset, GOLD produced RMSD smaller than 2 Å on 56.9% of the data while QVina 2 attained 63.1%. AVAILABILITY AND IMPLEMENTATION: The C++ source code of QVina 2 is available at (www.qvina.org). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Amr Alhossary, Stephanus Daniel Handoko, Yuguang Mu, Chee Keong Kwoh 0001 |
Bioinform. | 3 |
| 2012 | Accurate detection of SNPs using base-specific cleavage and mass spectrometryabstractAccurate detection of single-nucleotide polymorphisms (SNPs) is crucial for the success of many downstream analyses such as clinical diagnosis, virus identification, genetic mapping and association studies. Among many others, one valuable approach for SNP detection is based on the base-specific cleavage of single-stranded nucleic acids followed by matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) analysis. In this paper, we present a new SNP detection algorithm, which in particular permits an efficient and effective integration of the information in four complementary base-specific mass spectra. The new algorithm was implemented in a program called SnpMs. Comparative evaluation has been carried out on both simulated and real biological datasets, where experimental results clearly demonstrated the high ability of SnpMs as a tool to accurately detect SNPs. Ruimin Sun, Xiang Gao 0008, Nanyu Han, Yuguang Mu |
BIBM | 5 |
| 2011 | Dynamically-Driven Inactivation of the Catalytic Machinery of the SARS 3C-Like Protease by the N214A Mutation on the Extra DomainabstractDespite utilizing the same chymotrypsin fold to host the catalytic machinery, coronavirus 3C-like proteases (3CLpro) noticeably differ from picornavirus 3C proteases in acquiring an extra helical domain in evolution. Previously, the extra domain was demonstrated to regulate the catalysis of the SARS-CoV 3CLpro by controlling its dimerization. Here, we studied N214A, another mutant with only a doubled dissociation constant but significantly abolished activity. Unexpectedly, N214A still adopts the dimeric structure almost identical to that of the wild-type (WT) enzyme. Thus, we conducted 30-ns molecular dynamics (MD) simulations for N214A, WT, and R298A which we previously characterized to be a monomer with the collapsed catalytic machinery. Remarkably, three proteases display distinctive dynamical behaviors. While in WT, the catalytic machinery stably retains in the activated state; in R298A it remains largely collapsed in the inactivated state, thus implying that two states are not only structurally very distinguishable but also dynamically well separated. Surprisingly, in N214A the catalytic dyad becomes dynamically unstable and many residues constituting the catalytic machinery jump to sample the conformations highly resembling those of R298A. Therefore, the N214A mutation appears to trigger the dramatic change of the enzyme dynamics in the context of the dimeric form which ultimately inactivates the catalytic machinery. The present MD simulations represent the longest reported so far for the SARS-CoV 3CLpro, unveiling that its catalysis is critically dependent on the dynamics, which can be amazingly modulated by the extra domain. Consequently, mediating the dynamics may offer a potential avenue to inhibit the SARS-CoV 3CLpro. Jiahai Shi, Nanyu Han, Liangzhong Lim, Shixiong Lua, J. Sivaraman, Lushan Wang, Yuguang Mu, Jianxing Song |
PLoS Comput. Biol. | 7 |
| 2009 | Amyloidogenesis Abolished by Proline Substitutions but Enhanced by Lipid BindingabstractThe influence of lipid molecules on the aggregation of a highly amyloidogenic segment of human islet amyloid polypeptide, hIAPP20-29, and the corresponding sequence from rat has been studied by all-atom replica exchange molecular dynamics (REMD) simulations with explicit solvent model. hIAPP20-29 fragments aggregate into partially ordered beta-sheet oligomers and then undergo large conformational reorganization and convert into parallel/antiparallel beta-sheet oligomers in mixed in-register and out-of-register patterns. The hydrophobic interaction between lipid tails and residues at positions 23-25 is found to stabilize the ordered beta-sheet structure, indicating a catalysis role of lipid molecules in hIAPP20-29 self-assembly. The rat IAPP variants with three proline residues maintain unstructured micelle-like oligomers, which is consistent with non-amyloidogenic behavior observed in experimental studies. Our study provides the atomic resolution descriptions of the catalytic function of lipid molecules on the aggregation of IAPP peptides. Yuguang Mu |
PLoS Comput. Biol. | 3 |