VLDB 2026 Research / reviewers in the wild / expert
Kai-Long Zhao
dblp:310/7616
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | De Novo Protein Structure Prediction by Model Quality Assessment Dynamic Feedback Mechanism Using Deep LearningabstractAccurate de novo protein structure prediction remains a fundamental challenge, particularly in cases where homologous templates are unavailable or evolutionary information is weak. While end-to-end methods such as AlphaFold2 have achieved unprecedented accuracy, their closed box nature provides limited insight into the folding process and offers little flexibility for incorporating external evaluation. Here, we investigate whether model quality assessment (MQA) can be integrated into the structure prediction pipeline as a closed-loop feedback mechanism to iteratively improve prediction accuracy. In this study, we propose DGMFold, a de novo protein structure prediction method that establishes a feedback loop among three components: the geometric constraint prediction network (GeomNet), the structural simulation module, and the model quality evaluation network (EmaNet). In GeomNet, co-evolutionary features extracted from multiple sequence alignments (MSAs) are fed into an improved residual neural network to predict inter-residue geometric constraints, which are then used to guide structure folding. EmaNet then extracts 1D and 2D features from the folded structure model and employs a deep residual neural network to estimate the inter-residue distance deviation and per-residue lDDT. These evaluations are subsequently fed back into GeomNet as dynamic features, enabling iterative refinement of the predicted geometries and overall model accuracy. DGMFold was tested on 437 benchmark proteins and 20 FM targets of CASP14. Experimental results demonstrate that the closed-loop feedback mechanism significantly contributes to the performance of DGMFold, and the prediction accuracy of DGMFold outperforms that of the state-of-the-art de novo methods trRosetta and RaptorX at the time. When evaluated on the 124 human proteins for which AlphaFold2 yields TM-scores below 0.9, DGMFold achieves higher prediction accuracy than AlphaFold2 and RoseTTAFold on 71 and 72 targets, respectively, and outperforms both on 58 proteins. Jun Liu 0078, Guang-Xing He, Kai-Long Zhao, Guijun Zhang |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Improved model quality assessment using sequence and structural information by enhanced deep neural networksabstractProtein model quality assessment plays an important role in protein structure prediction, protein design and drug discovery. In this work, DeepUMQA2, a substantially improved version of DeepUMQA for protein model quality assessment, is proposed. First, sequence features containing protein co-evolution information and structural features reflecting family information are extracted to complement model-dependent features. Second, a novel backbone network based on triangular multiplication update and axial attention mechanism is designed to enhance information exchange between inter-residue pairs. On CASP13 and CASP14 datasets, the performance of DeepUMQA2 increases by 20.5 and 20.4% compared with DeepUMQA, respectively (measured by top 1 loss). Moreover, on the three-month CAMEO dataset (11 March to 04 June 2022), DeepUMQA2 outperforms DeepUMQA by 15.5% (measured by local AUC0,0.2) and ranks first among all competing server methods in CAMEO blind test. Experimental results show that DeepUMQA2 outperforms state-of-the-art model quality assessment methods, such as ProQ3D-LDDT, ModFOLD8, and DeepAccNet and DeepUMQA2 can select more suitable best models than state-of-the-art protein structure methods, such as AlphaFold2, RoseTTAFold and I-TASSER, provided themselves. Jun Liu 0078, Kai-Long Zhao, Guijun Zhang |
Briefings Bioinform. | 2 |
| 2022 | Construct a variable-length fragment library for de novo protein structure predictionabstractAlthough remarkable achievements, such as AlphaFold2, have been made in end-to-end structure prediction, fragment libraries remain essential for de novo protein structure prediction, which can help explore and understand the protein-folding mechanism. In this work, we developed a variable-length fragment library (VFlib). In VFlib, a master structure database was first constructed from the Protein Data Bank through sequence clustering. The hidden Markov model (HMM) profile of each protein in the master structure database was generated by HHsuite, and the secondary structure of each protein was calculated by DSSP. For the query sequence, the HMM-profile was first constructed. Then, variable-length fragments were retrieved from the master structure database through dynamically variable-length profile-profile comparison. A complete method for chopping the query HMM-profile during this process was proposed to obtain fragments with increased diversity. Finally, secondary structure information was used to further screen the retrieved fragments to generate the final fragment library of specific query sequence. The experimental results obtained with a set of 120 nonredundant proteins show that the global precision and coverage of the fragment library generated by VFlib were 55.04% and 94.95% at the RMSD cutoff of 1.5 Å, respectively. Compared with the benchmark method of NNMake, the global precision of our fragment library had increased by 62.89% with equivalent coverage. Furthermore, the fragments generated by VFlib and NNMake were used to predict structure models through fragment assembly. Controlled experimental results demonstrate that the average TM-score of VFlib was 16.00% higher than that of NNMake. Qiongqiong Feng, Minghua Hou, Jun Liu 0078, Kai-Long Zhao, Guijun Zhang |
Briefings Bioinform. | 4 |
| 2021 | A de novo protein structure prediction by iterative partition sampling, topology adjustment and residue-level distance deviation optimizationabstractMOTIVATION: With the great progress of deep learning-based inter-residue contact/distance prediction, the discrete space formed by fragment assembly cannot satisfy the distance constraint well. Thus, the optimal solution of the continuous space may not be achieved. Designing an effective closed-loop continuous dihedral angle optimization strategy that complements the discrete fragment assembly is crucial to improve the performance of the distance-assisted fragment assembly method. RESULTS: In this article, we proposed a de novo protein structure prediction method called IPTDFold based on closed-loop iterative partition sampling, topology adjustment and residue-level distance deviation optimization. First, local dihedral angle crossover and mutation operators are designed to explore the conformational space extensively and achieve information exchange between the conformations in the population. Then, the dihedral angle rotation model of loop region with partial inter-residue distance constraints is constructed, and the rotation angle satisfying the constraints is obtained by differential evolution algorithm, so as to adjust the spatial position relationship between the secondary structures. Finally, the residue distance deviation is evaluated according to the difference between the conformation and the predicted distance, and the dihedral angle of the residue is optimized with biased probability. The final model is generated by iterating the above three steps. IPTDFold is tested on 462 benchmark proteins, 24 FM targets of CASP13 and 20 FM targets of CASP14. Results show that IPTDFold is significantly superior to the distance-assisted fragment assembly method Rosetta_D (Rosetta with distance). In particular, the prediction accuracy of IPTDFold does not decrease as the length of the protein increases. When using the same FastRelax protocol, the prediction accuracy of IPTDFold is significantly superior to that of trRosetta without orientation constraints, and is equivalent to that of the full version of trRosetta. AVAILABILITYAND IMPLEMENTATION: The source code and executable are freely available at https://github.com/iobio-zjut/IPTDFold. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun Liu 0078, Kai-Long Zhao, Guang-Xing He, Liu-Jing Wang, Guijun Zhang |
Bioinform. | 2 |
| 2021 | MMpred: a distance-assisted multimodal conformation sampling for de novo protein structure predictionabstractMOTIVATION: The mathematically optimal solution in computational protein folding simulations does not always correspond to the native structure, due to the imperfection of the energy force fields. There is therefore a need to search for more diverse suboptimal solutions in order to identify the states close to the native. We propose a novel multimodal optimization protocol to improve the conformation sampling efficiency and modeling accuracy of de novo protein structure folding simulations. RESULTS: A distance-assisted multimodal optimization sampling algorithm, MMpred, is proposed for de novo protein structure prediction. The protocol consists of three stages: The first is a modal exploration stage, in which a structural similarity evaluation model DMscore is designed to control the diversity of conformations, generating a population of diverse structures in different low-energy basins. The second is a modal maintaining stage, where an adaptive clustering algorithm MNDcluster is proposed to divide the populations and merge the modal by adjusting the annealing temperature to locate the promising basins. In the last stage of modal exploitation, a greedy search strategy is used to accelerate the convergence of the modal. Distance constraint information is used to construct the conformation scoring model to guide sampling. MMpred is tested on a large set of 320 non-redundant proteins, where MMpred obtains models with TM-score≥0.5 on 291 cases, which is 28% higher than that of Rosetta guided with the same set of distance constraints. In addition, on 320 benchmark proteins, the enhanced version of MMpred (E-MMpred) has 167 targets better than trRosetta when the best of five models are evaluated. The average TM-score of the best model of E-MMpred is 0.732, which is comparable to trRosetta (0.730). AVAILABILITY AND IMPLEMENTATION: The source code and executable are freely available at https://github.com/iobio-zjut/MMpred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kai-Long Zhao, Jun Liu 0078, Jianzhong Su, Yang Zhang 0040, Guijun Zhang |
Bioinform. | 1 |