Jinjin Li 0003

dblp:05/7569-3 · also Jin-Jin Li 0003 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2023
0000-0003-4661-4051ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
YearPublicationVenuePosition
2023 A deep transfer learning-based protocol accelerates full quantum mechanics calculation of protein
abstract
Effective full quantum mechanics (FQM) calculation of protein remains a grand challenge and of great interest in computational biology with substantial applications in drug discovery, protein dynamic simulation and protein folding. However, the huge computational complexity of the existing QM methods impends their applications in large systems. Here, we design a transfer-learning-based deep learning (TDL) protocol for effective FQM calculations (TDL-FQM) on proteins. By incorporating a transfer-learning algorithm into deep neural network (DNN), the TDL-FQM protocol is capable of performing calculations at any given accuracy using models trained from small datasets with high-precision and knowledge learned from large amount of low-level calculations. The high-level double-hybrid DFT functional and high-level quality of basis set is used in this work as a case study to evaluate the performance of TDL-FQM, where the selected 15 proteins are predicted to have a mean absolute error of 0.01 kcal/mol/atom for potential energy and an average root mean square error of 1.47 kcal/mol/$ {\rm A^{^{ \!\!\!o}}} $ for atomic forces. The proposed TDL-FQM approach accelerates the FQM calculation more than thirty thousand times faster in average and presents more significant benefits in efficiency as the size of protein increases. The ability to learn knowledge from one task to solve related problems demonstrates that the proposed TDL-FQM overcomes the limitation of standard DNN and has a strong power to predict proteins with high precision, which solves the challenge of high precision prediction in large chemical and biological systems.
Yanqiang Han, Zhilong Wang 0003, An Chen 0003, Junfei Cai, Simin Ye, Zhiyun Wei, Jinjin Li 0003
Briefings Bioinform.8
2022 An inductive transfer learning force field (ITLFF) protocol builds protein force fields in seconds
abstract
Accurate simulation of protein folding is a unique challenge in understanding the physical process of protein folding, with important implications for protein design and drug discovery. Molecular dynamics simulation strongly requires advanced force fields with high accuracy to achieve correct folding. However, the current force fields are inaccurate, inapplicable and inefficient. We propose a machine learning protocol, the inductive transfer learning force field (ITLFF), to construct protein force fields in seconds with any level of accuracy from a small dataset. This process is achieved by incorporating an inductive transfer learning algorithm into deep neural networks, which learn knowledge of any high-level calculations from a large dataset of low-level method. Here, we use a double-hybrid density functional theory (DFT) as a case functional, but ITLFF is suitable for any high-precision functional. The performance of the selected 18 proteins indicates that compared with the fragment-based double-hybrid DFT algorithm, the force field constructed by ITLFF achieves considerable accuracy with a mean absolute error of 0.0039 kcal/mol/atom for energy and a root mean square error of 2.57 $\mathrm{kcal}/\mathrm{mol}/{\AA}$ for force, and it is more than 30 000 times faster and obtains more significant efficiency benefits as the system increases. The outstanding performance of ITLFF provides promising prospects for accurate and efficient protein dynamic simulations and makes an important step toward protein folding simulation. Due to the ability of ITLFF to utilize the knowledge acquired in one task to solve related problems, it is also applicable for various problems in biology, chemistry and material science.
Yanqiang Han, Zhilong Wang 0003, An Chen 0003, Junfei Cai, Simin Ye, Jinjin Li 0003
Briefings Bioinform.7
2022 Clustered tree regression to learn protein energy change with mutated amino acid
abstract
Accurate and effective prediction of mutation-induced protein energy change remains a great challenge and of great interest in computational biology. However, high resource consumption and insufficient structural information of proteins severely limit the experimental techniques and structure-based prediction methods. Here, we design a structure-independent protocol to accurately and effectively predict the mutation-induced protein folding free energy change with only sequence, physicochemical and evolutionary features. The proposed clustered tree regression protocol is capable of effectively exploiting the inherent data patterns by integrating unsupervised feature clustering by K-means and supervised tree regression using XGBoost, and thus enabling fast and accurate protein predictions with different mutations, with an average Pearson correlation coefficient of 0.83 and an average root-mean-square error of 0.94kcal/mol. The proposed sequence-based method not only eliminates the dependence on protein structures, but also has potential applications in protein predictions with rare structural information.
Hongwei Tu, Yanqiang Han, Zhilong Wang 0003, Jinjin Li 0003
Briefings Bioinform.4
2021 Potential inhibitors for the novel coronavirus (SARS-CoV-2)
abstract
The lack of a vaccine or any effective treatment for the aggressive novel coronavirus disease (COVID-19) has created a sense of urgency for the discovery of effective drugs. Several repurposing pharmaceutical candidates have been reported or envisaged to inhibit the emerging infections of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), but their binding sites, binding affinities and inhibitory mechanisms are still unavailable. In this study, we use the ligand-protein docking program and molecular dynamic simulation to ab initio investigate the binding mechanism and inhibitory ability of seven clinically approved drugs (Chloroquine, Hydroxychloroquine, Remdesivir, Ritonavir, Beclabuvir, Indinavir and Favipiravir) and a recently designed α-ketoamide inhibitor (13b) at the molecular level. The results suggest that Chloroquine has the strongest binding affinity with 3CL hydrolase (Mpro) among clinically approved drugs, indicating its effective inhibitory ability for SARS-CoV-2. However, the newly designed inhibitor 13b shows potentially improved inhibition efficiency with larger binding energy compared with Chloroquine. We further calculate the important binding site residues at the active site and demonstrate that the MET 165 and HIE 163 contribute the most for 13b, while the MET 165 and GLN 189 for Chloroquine, based on residual energy decomposition analysis. The proposed work offers a higher research priority for 13b to treat the infection of SARS-CoV-2 and provides theoretical basis for further design of effective drug molecules with stronger inhibition.
Yanqiang Han, Zhilong Wang 0003, Jiahao Ren, Zhiyun Wei, Jinjin Li 0003
Briefings Bioinform.5
2021 Machine learning builds full-QM precision protein force fields in seconds
abstract
Full-quantum mechanics (QM) calculations are extraordinarily precise but difficult to apply to large systems, such as biomolecules. Motivated by the massive demand for efficient calculations for large systems at the full-QM level and by the significant advances in machine learning, we have designed a neural network-based two-body molecular fractionation with conjugate caps (NN-TMFCC) approach to accelerate the energy and atomic force calculations of proteins. The results show very high precision for the proposed NN potential energy surface models of residue-based fragments, with energy root-mean-squared errors (RMSEs) less than 1.0 kcal/mol and force RMSEs less than 1.3 kcal/mol/Å for both training and testing sets. The proposed NN-TMFCC method calculates the energies and atomic forces of 15 representative proteins with full-QM precision in 10-100 s, which is thousands of times faster than the full-QM calculations. The computational complexity of the NN-TMFCC method is independent of the protein size and only depends on the number of residue species, which makes this method particularly suitable for rapid prediction of large systems with tens of thousands or even hundreds of thousands of times acceleration. This highly precise and efficient NN-TMFCC approach exhibits considerable potential for performing energy and force calculations, structure predictions and molecular dynamics simulations of proteins with full-QM precision.
Yanqiang Han, Zhilong Wang 0003, Zhiyun Wei, Jinyun Liu, Jinjin Li 0003
Briefings Bioinform.5