EDBT 2026 Demo / reviewers in the wild / expert
Lei Xu 0035
dblp:19/360-35
· DBLP profile ↗
9ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-4095-6539ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Utilization of Machine Learning Approaches in Predicting the Prominent Physicochemical Properties of Polychlorinated BiphenylabstractGiven the difficulties in experimental measurement, machine learning offers a viable alternative method to reduce and minimize costs and time. Machine learning enhances the effectiveness and efficiency of environmental pollutant monitoring, supporting better management and protection of ecosystems and public health. Random forest is indeed a significant machine learning method, widely used for various applications due to its robustness and versatility. Polychlorinated biphenyls (PCBs), a vital class of persistent organic pollutants, have garnered significant attention from the scientific community due to their detrimental impacts. In this study, a computational prediction model for octanol–water partition coefficient (logP) and bioconcentration factor (logBCF) of PCBs was developed by using random forest (RF), sparrow search algorithm random forest (SSA-RF), particle swarm optimization random forest (PSO-RF), and gray wolf optimization random forest (GWO-RF) methods. We performed a comprehensive validation, evaluation, and mechanistic explanation of the model for ensuring its reliability and applicability. Overall, the internal and external validation statistical parameters of the eight models have good robustness and predictive power. The Williams plots further show that all models are built in a wide range of application domains, and therefore they can be applied to unrecognized PCBs already in the environment to fill in the gaps in relevant experimental data. SSA-RF was superior to other methods, suggesting that it is more appropriate for computational studies of PCBs. Hongqin Zhang, Lei Xu 0035 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2023 | Can molecular dynamics simulations improve predictions of protein-ligand binding affinity with machine learning?abstractBinding affinity prediction largely determines the discovery efficiency of lead compounds in drug discovery. Recently, machine learning (ML)-based approaches have attracted much attention in hopes of enhancing the predictive performance of traditional physics-based approaches. In this study, we evaluated the impact of structural dynamic information on the binding affinity prediction by comparing the models trained on different dimensional descriptors, using three targets (i.e. JAK1, TAF1-BD2 and DDR1) and their corresponding ligands as the examples. Here, 2D descriptors are traditional ECFP4 fingerprints, 3D descriptors are the energy terms of the Smina and NNscore scoring functions and 4D descriptors contain the structural dynamic information derived from the trajectories based on molecular dynamics (MD) simulations. We systematically investigate the MD-refined binding affinity prediction performance of three classical ML algorithms (i.e. RF, SVR and XGB) as well as two common virtual screening methods, namely Glide docking and MM/PBSA. The outcomes of the ML models built using various dimensional descriptors and their combinations reveal that the MD refinement with the optimized protocol can improve the predictive performance on the TAF1-BD2 target with considerable structural flexibility, but not for the less flexible JAK1 and DDR1 targets, when taking docking poses as the initial structure instead of the crystal structures. The results highlight the importance of the initial structures to the final performance of the model through conformational analysis on the three targets with different flexibility. Shukai Gu, Chao Shen 0008, Huanxiang Liu, Rong Sheng, Lei Xu 0035, Zhe Wang 0041, Tingjun Hou, Yu Kang 0002 |
Briefings Bioinform. | 8 |
| 2023 | Cooperation of structural motifs controls drug selectivity in cyclin-dependent kinases: an advanced theoretical analysisabstractUnderstanding drug selectivity mechanism is a long-standing issue for helping design drugs with high specificity. Designing drugs targeting cyclin-dependent kinases (CDKs) with high selectivity is challenging because of their highly conserved binding pockets. To reveal the underlying general selectivity mechanism, we carried out comprehensive analyses from both the thermodynamics and kinetics points of view on a representative CDK12 inhibitor. To fully capture the binding features of the drug-target recognition process, we proposed to use kinetic residue energy analysis (KREA) in conjunction with the community network analysis (CNA) to reveal the underlying cooperation effect between individual residues/protein motifs to the binding/dissociating process of the ligand. The general mechanism of drug selectivity in CDKs can be summarized as that the difference of structural cooperation between the ligand and the protein motifs leads to the difference of the energetic contribution of the key residues to the ligand. The proposed mechanisms may be prevalent in drug selectivity issues, and the insights may help design new strategies to overcome/attenuate the drug selectivity associated problems. Lei Xu 0035, Zhe Wang 0041, Tingjun Hou, Haiping Hao, Huiyong Sun |
Briefings Bioinform. | 2 |
| 2022 | Improving the Performance of Lattice Boltzmann Method with Pipelined Algorithm on A Heterogeneous Multi-zone Processor
Qingyang Zhang 0009, Lei Xu 0035, Rongliang Chen, Lin Chen 0028, Xinhai Chen 0001, Jie Liu 0002, Bo Yang 0023 |
PDCAT | 2 |
| 2022 | Comprehensive assessment of deep generative architectures for de novo drug designabstractRecently, deep learning (DL)-based de novo drug design represents a new trend in pharmaceutical research, and numerous DL-based methods have been developed for the generation of novel compounds with desired properties. However, a comprehensive understanding of the advantages and disadvantages of these methods is still lacking. In this study, the performances of different generative models were evaluated by analyzing the properties of the generated molecules in different scenarios, such as goal-directed (rediscovery, optimization and scaffold hopping of active compounds) and target-specific (generation of novel compounds for a given target) tasks. In overall, the DL-based models have significant advantages over the baseline models built by the traditional methods in learning the physicochemical property distributions of the training sets and may be more suitable for target-specific tasks. However, both the baselines and DL-based generative models cannot fully exploit the scaffolds of the training sets, and the molecules generated by the DL-based methods even have lower scaffold diversity than those generated by the traditional models. Moreover, our assessment illustrates that the DL-based methods do not exhibit obvious advantages over the genetic algorithm-based baselines in goal-directed tasks. We believe that our study provides valuable guidance for the effective use of generative models in de novo drug design. Mingyang Wang 0004, Huiyong Sun, Jike Wang, Jinping Pang, Xin Chai, Lei Xu 0035, Honglin Li 0003, Dong-Sheng Cao 0001, Tingjun Hou |
Briefings Bioinform. | 6 |
| 2021 | Beware of the generic machine learning-based scoring functions in structure-based virtual screeningabstractMachine learning-based scoring functions (MLSFs) have attracted extensive attention recently and are expected to be potential rescoring tools for structure-based virtual screening (SBVS). However, a major concern nowadays is whether MLSFs trained for generic uses rather than a given target can consistently be applicable for VS. In this study, a systematic assessment was carried out to re-evaluate the effectiveness of 14 reported MLSFs in VS. Overall, most of these MLSFs could hardly achieve satisfactory results for any dataset, and they could even not outperform the baseline of classical SFs such as Glide SP. An exception was observed for RFscore-VS trained on the Directory of Useful Decoys-Enhanced dataset, which showed its superiority for most targets. However, in most cases, it clearly illustrated rather limited performance on the targets that were dissimilar to the proteins in the corresponding training sets. We also used the top three docking poses rather than the top one for rescoring and retrained the models with the updated versions of the training set, but only minor improvements were observed. Taken together, generic MLSFs may have poor generalization capabilities to be applicable for the real VS campaigns. Therefore, it should be quite cautious to use this type of methods for VS. Chao Shen 0008, Zhe Wang 0041, Xujun Zhang, Jinping Pang, Gaoang Wang, Haiyang Zhong, Lei Xu 0035, Dong-Sheng Cao 0001, Tingjun Hou |
Briefings Bioinform. | 8 |
| 2021 | Can machine learning consistently improve the scoring power of classical scoring functions? Insights into the role of machine learning in scoring functionsabstractHow to accurately estimate protein-ligand binding affinity remains a key challenge in computer-aided drug design (CADD). In many cases, it has been shown that the binding affinities predicted by classical scoring functions (SFs) cannot correlate well with experimentally measured biological activities. In the past few years, machine learning (ML)-based SFs have gradually emerged as potential alternatives and outperformed classical SFs in a series of studies. In this study, to better recognize the potential of classical SFs, we have conducted a comparative assessment of 25 commonly used SFs. Accordingly, the scoring power was systematically estimated by using the state-of-the-art ML methods that replaced the original multiple linear regression method to refit individual energy terms. The results show that the newly-developed ML-based SFs consistently performed better than classical ones. In particular, gradient boosting decision tree (GBDT) and random forest (RF) achieved the best predictions in most cases. The newly-developed ML-based SFs were also tested on another benchmark modified from PDBbind v2007, and the impacts of structural and sequence similarities were evaluated. The results indicated that the superiority of the ML-based SFs could be fully guaranteed when sufficient similar targets were contained in the training set. Moreover, the effect of the combinations of features from multiple SFs was explored, and the results indicated that combining NNscore2.0 with one to four other classical SFs could yield the best scoring power. However, it was not applicable to derive a generic target-specific SF or SF combination. Chao Shen 0008, Zhe Wang 0041, Xujun Zhang, Haiyang Zhong, Gaoang Wang, Lei Xu 0035, Dong-Sheng Cao 0001, Tingjun Hou |
Briefings Bioinform. | 8 |
| 2021 | DeepAtomicCharge: a new graph convolutional network-based architecture for accurate prediction of atomic chargesabstractAtomic charges play a very important role in drug-target recognition. However, computation of atomic charges with high-level quantum mechanics (QM) calculations is very time-consuming. A number of machine learning (ML)-based atomic charge prediction methods have been proposed to speed up the calculation of high-accuracy atomic charges in recent years. However, most of them used a set of predefined molecular properties, such as molecular fingerprints, for model construction, which is knowledge-dependent and may lead to biased predictions due to the representation preference of different molecular properties used for training. To solve the problem, we present a new architecture based on graph convolutional network (GCN) and develop a high-accuracy atomic charge prediction model named DeepAtomicCharge. The new GCN architecture is designed with only the atomic properties and the connection information between the atoms in molecules and can dynamically learn and convert molecules into appropriate atomic features without any prior knowledge of the molecules. Using the designed GCN architecture, substantial improvement is achieved for the prediction accuracy of atomic charges. The average root-mean-square error (RMSE) of DeepAtomicCharge is 0.0121 e, which is obviously more accurate than that (0.0180 e) reported by the previous benchmark study on the same two external test sets. Moreover, the new GCN architecture needs much lower storage space compared with other methods, and the predicted DDEC atomic charges can be efficiently used in large-scale structure-based drug design, thus opening a new avenue for high-performance atomic charge prediction and application. Jike Wang, Dong-Sheng Cao 0001, Cunchen Tang, Lei Xu 0035, Qiaojun He, Bo Yang 0023, Huiyong Sun, Tingjun Hou |
Briefings Bioinform. | 4 |
| 2020 | Comprehensive assessment of nine docking programs on type II kinase inhibitors: prediction accuracy of sampling power, scoring power and screening powerabstractProtein kinases have been regarded as important therapeutic targets for many diseases. Currently, a total of 41 kinase inhibitors have been approved by the Food and Drug Administration, along with a large number of kinase inhibitors being evaluated in clinical and preclinical trials. Among all, allosteric inhibitors, such as type II kinase inhibitors, have attracted extensive attention owing to their potential high selectivity. Nowadays, molecular docking has become a powerful tool to search for novel kinase inhibitors. However, as for type II kinase inhibitors, their allosteric characteristics may exert a deep influence on docking accuracy. In this study, a comprehensive assessment was conducted to evaluate the effectiveness of nine docking algorithms towards type II kinase inhibitors. The calculation results showed that most tested docking programs, especially Glide with XP scoring, LeDock and Surflex-Dock, succeeded in the accurate identification of near-native binding poses, with the success rates ranging from 0.80 to 0.90, and the scoring functions in GOLD and LeDock outperformed the others in the prediction of relative binding affinities. In terms of the P-values, areas under the curve and enrichment factors, Glide with XP scoring, Surflex-Dock, GOLD with Astex Statistical Potential scoring and LeDock had better screening power to discriminate between active compounds and decoys. However, the screening power is sensitive to different initial conformations of the same target. It is expected that our study can provide some guidance for docking-based virtual screening to discover novel type II kinase inhibitors, as well as other allosteric inhibitors. Chao Shen 0008, Zhe Wang 0041, Youyong Li, Tailong Lei, Ercheng Wang, Lei Xu 0035, Feng Zhu 0004, Dan Li 0013, Tingjun Hou |
Briefings Bioinform. | 7 |