Luhua Lai

dblp:61/6993 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0002-8343-7587ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 68% Computational science and engineering · 32%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
contrastive learning
0.712023
Hierarchical graph transformer with contrastive learning for protein function prediction · Bioinform. 2023
Computational science and engineering › graph learning
graph neural network
0.712023
Hierarchical graph transformer with contrastive learning for protein function prediction · Bioinform. 2023
Bioinformatics and computational biology
protein function prediction
0.712023
Hierarchical graph transformer with contrastive learning for protein function prediction · Bioinform. 2023
Bioinformatics and computational biology
protein structure analysis
0.112007
Identification of amyloid fibril-forming segments based on structure and residue-based statistical potential · Bioinform. 2007

Methods — techniques the papers use, named apart from their topics

graph transformer · 0.7contrastive learning · 0.7class activation mapping · 0.7residue-based statistical potential · 0.1microcrystal structure · 0.1
YearPublicationVenuePosition
2023 Hierarchical graph transformer with contrastive learning for protein function prediction
abstract
MOTIVATION: In recent years, high-throughput sequencing technologies have made large-scale protein sequences accessible. However, their functional annotations usually rely on low-throughput and pricey experimental studies. Computational prediction models offer a promising alternative to accelerate this process. Graph neural networks have shown significant progress in protein research, but capturing long-distance structural correlations and identifying key residues in protein graphs remains challenging. RESULTS: In the present study, we propose a novel deep learning model named Hierarchical graph transformEr with contrAstive Learning (HEAL) for protein function prediction. The core feature of HEAL is its ability to capture structural semantics using a hierarchical graph Transformer, which introduces a range of super-nodes mimicking functional motifs to interact with nodes in the protein graph. These semantic-aware super-node embeddings are then aggregated with varying emphasis to produce a graph representation. To optimize the network, we utilized graph contrastive learning as a regularization technique to maximize the similarity between different views of the graph representation. Evaluation of the PDBch test set shows that HEAL-PDB, trained on fewer data, achieves comparable performance to the recent state-of-the-art methods, such as DeepFRI. Moreover, HEAL, with the added benefit of unresolved protein structures predicted by AlphaFold2, outperforms DeepFRI by a significant margin on Fmax, AUPR, and Smin metrics on PDBch test set. Additionally, when there are no experimentally resolved structures available for the proteins of interest, HEAL can still achieve better performance on AFch test set than DeepFRI and DeepGOPlus by taking advantage of AlphaFold2 predicted structures. Finally, HEAL is capable of finding functional sites through class activation mapping. AVAILABILITY AND IMPLEMENTATION: Implementations of our HEAL can be found at https://github.com/ZhonghuiGu/HEAL.
Zhonghui Gu, Xiao Luo 0001, Jiaxiao Chen, Minghua Deng, Luhua Lai
Bioinform.5
2022 Prediction of liquid-liquid phase separating proteins using machine learning
abstract
BACKGROUND: The liquid-liquid phase separation (LLPS) of biomolecules in cell underpins the formation of membraneless organelles, which are the condensates of protein, nucleic acid, or both, and play critical roles in cellular function. Dysregulation of LLPS is implicated in a number of diseases. Although the LLPS of biomolecules has been investigated intensively in recent years, the knowledge of the prevalence and distribution of phase separation proteins (PSPs) is still lag behind. Development of computational methods to predict PSPs is therefore of great importance for comprehensive understanding of the biological function of LLPS. RESULTS: Based on the PSPs collected in LLPSDB, we developed a sequence-based prediction tool for LLPS proteins (PSPredictor), which is an attempt at general purpose of PSP prediction that does not depend on specific protein types. Our method combines the componential and sequential information during the protein embedding stage, and, adopts the machine learning algorithm for final predicting. The proposed method achieves a tenfold cross-validation accuracy of 94.71%, and outperforms previously reported PSPs prediction tools. For further applications, we built a user-friendly PSPredictor web server ( http://www.pkumdl.cn/PSPredictor ), which is accessible for prediction of potential PSPs. CONCLUSIONS: PSPredictor could identifie novel scaffold proteins for stress granules and predict PSPs candidates in the human genome for further study. For further applications, we built a user-friendly PSPredictor web server ( http://www.pkumdl.cn/PSPredictor ), which provides valuable information for potential PSPs recognition.
Xiaoquan Chu, Tanlin Sun, Youjun Xu, Zhuqing Zhang, Luhua Lai, Jianfeng Pei
BMC Bioinform.6
2021 A transferable deep learning approach to fast screen potential antiviral drugs against SARS-CoV-2
abstract
The COVID-19 pandemic calls for rapid development of effective treatments. Although various drug repurpose approaches have been used to screen the FDA-approved drugs and drug candidates in clinical phases against SARS-CoV-2, the coronavirus that causes this disease, no magic bullets have been found until now. In this study, we used directed message passing neural network to first build a broad-spectrum anti-beta-coronavirus compound prediction model, which gave satisfactory predictions on newly reported active compounds against SARS-CoV-2. Then, we applied transfer learning to fine-tune the model with the recently reported anti-SARS-CoV-2 compounds and derived a SARS-CoV-2 specific prediction model COVIDVS-3. We used COVIDVS-3 to screen a large compound library with 4.9 million drug-like molecules from ZINC15 database and recommended a list of potential anti-SARS-CoV-2 compounds for further experimental testing. As a proof-of-concept, we experimentally tested seven high-scored compounds that also demonstrated good binding strength in docking studies against the 3C-like protease of SARS-CoV-2 and found one novel compound that can inhibit the enzyme. Our model is highly efficient and can be used to screen large compound databases with millions or more compounds to accelerate the drug discovery process for the treatment of COVID-19.
Youjun Xu, Jianfeng Pei, Luhua Lai
Briefings Bioinform.5
2017 Sequence-based prediction of protein protein interaction using a deep-learning algorithm
abstract
BACKGROUND: Protein-protein interactions (PPIs) are critical for many biological processes. It is therefore important to develop accurate high-throughput methods for identifying PPI to better understand protein function, disease occurrence, and therapy design. Though various computational methods for predicting PPI have been developed, their robustness for prediction with external datasets is unknown. Deep-learning algorithms have achieved successful results in diverse areas, but their effectiveness for PPI prediction has not been tested. RESULTS: We used a stacked autoencoder, a type of deep-learning algorithm, to study the sequence-based PPI prediction. The best model achieved an average accuracy of 97.19% with 10-fold cross-validation. The prediction accuracies for various external datasets ranged from 87.99% to 99.21%, which are superior to those achieved with previous methods. CONCLUSIONS: To our knowledge, this research is the first to apply a deep-learning algorithm to sequence-based PPI prediction, and the results demonstrate its potential in this field.
Tanlin Sun, Luhua Lai, Jianfeng Pei
BMC Bioinform.3
2013 Ligand Clouds around Protein Clouds: A Scenario of Ligand Binding with Intrinsically Disordered Proteins
abstract
Intrinsically disordered proteins (IDPs) were found to be widely associated with human diseases and may serve as potential drug design targets. However, drug design targeting IDPs is still in the very early stages. Progress in drug design is usually achieved using experimental screening; however, the structural disorder of IDPs makes it difficult to characterize their interaction with ligands using experiments alone. To better understand the structure of IDPs and their interactions with small molecule ligands, we performed extensive simulations on the c-Myc₃₇₀₋₄₀₉ peptide and its binding to a reported small molecule inhibitor, ligand 10074-A4. We found that the conformational space of the apo c-Myc₃₇₀₋₄₀₉ peptide was rather dispersed and that the conformations of the peptide were stabilized mainly by charge interactions and hydrogen bonds. Under the binding of the ligand, c-Myc₃₇₀₋₄₀₉ remained disordered. The ligand was found to bind to c-Myc₃₇₀₋₄₀₉ at different sites along the chain and behaved like a 'ligand cloud'. In contrast to ligand binding to more rigid target proteins that usually results in a dominant bound structure, ligand binding to IDPs may better be described as ligand clouds around protein clouds. Nevertheless, the binding of the ligand and a non-ligand to the c-Myc₃₇₀₋₄₀₉ target could be clearly distinguished. The present study provides insights that will help improve rational drug design that targets IDPs.
Luhua Lai
PLoS Comput. Biol.3
2007 Identification of amyloid fibril-forming segments based on structure and residue-based statistical potential
abstract
MOTIVATION: Experimental evidence suggests that certain short protein segments have stronger amyloidogenic propensities than others. Identification of the fibril-forming segments of proteins is crucial for understanding diseases associated with protein misfolding and for finding favorable targets for therapeutic strategies. RESULT: In this study, we used the microcrystal structure of the NNQQNY peptide from yeast prion protein and residue-based statistical potentials to establish an algorithm to identify the amyloid fibril-forming segment of proteins. Using the same sets of sequences, a comparable prediction performance was obtained from this study to that from 3D profile method based on the physical atomic-level potential ROSETTADESIGN. The predicted results are consistent with experiments for several representative proteins associated with amyloidosis, and also agree with the idea that peptides that can form fibrils may have strong sequence signatures. Application of the residue-based statistical potentials is computationally more efficient than using atomic-level potentials and can be applied in whole proteome analysis to investigate the evolutionary pressure effect or forecast other latent diseases related to amyloid deposits. AVAILABILITY: The fibril prediction program is available at ftp://mdl.ipc.pku.edu.cn/pub/software/pre-amyl/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhuqing Zhang, Luhua Lai
Bioinform.3
2007 Prediction of potential drug targets based on simple sequence properties
abstract
BACKGROUND: During the past decades, research and development in drug discovery have attracted much attention and efforts. However, only 324 drug targets are known for clinical drugs up to now. Identifying potential drug targets is the first step in the process of modern drug discovery for developing novel therapeutic agents. Therefore, the identification and validation of new and effective drug targets are of great value for drug discovery in both academia and pharmaceutical industry. If a protein can be predicted in advance for its potential application as a drug target, the drug discovery process targeting this protein will be greatly speeded up. In the current study, based on the properties of known drug targets, we have developed a sequence-based drug target prediction method for fast identification of novel drug targets. RESULTS: Based on simple physicochemical properties extracted from protein sequences of known drug targets, several support vector machine models have been constructed in this study. The best model can distinguish currently known drug targets from non drug targets at an accuracy of 84%. Using this model, potential protein drug targets of human origin from Swiss-Prot were predicted, some of which have already attracted much attention as potential drug targets in pharmaceutical research. CONCLUSION: We have developed a drug target prediction method based solely on protein sequence information without the knowledge of family/domain annotation, or the protein 3D structure. This method can be applied in novel drug target identification and validation, as well as genome scale drug target predictions.
Qingliang Li 0002, Luhua Lai
BMC Bioinform.2
2007 Dynamic Simulations on the Arachidonic Acid Metabolic Network
abstract
Drug molecules not only interact with specific targets, but also alter the state and function of the associated biological network. How to design drugs and evaluate their functions at the systems level becomes a key issue in highly efficient and low-side-effect drug design. The arachidonic acid metabolic network is the network that produces inflammatory mediators, in which several enzymes, including cyclooxygenase-2 (COX-2), have been used as targets for anti-inflammatory drugs. However, neither the century-old nonsteriodal anti-inflammatory drugs nor the recently revocatory Vioxx have provided completely successful anti-inflammatory treatment. To gain more insights into the anti-inflammatory drug design, the authors have studied the dynamic properties of arachidonic acid (AA) metabolic network in human polymorphous leukocytes. Metabolic flux, exogenous AA effects, and drug efficacy have been analyzed using ordinary differential equations. The flux balance in the AA network was found to be important for efficient and safe drug design. When only the 5-lipoxygenase (5-LOX) inhibitor was used, the flux of the COX-2 pathway was increased significantly, showing that a single functional inhibitor cannot effectively control the production of inflammatory mediators. When both COX-2 and 5-LOX were blocked, the production of inflammatory mediators could be completely shut off. The authors have also investigated the differences between a dual-functional COX-2 and 5-LOX inhibitor and a mixture of these two types of inhibitors. Their work provides an example for the integration of systems biology and drug discovery.
Kun Yang 0003, Wenzhe Ma, Huanhuan Liang, Qi Ouyang, Chao Tang 0003, Luhua Lai
PLoS Comput. Biol.6