Kai Wang 0017

dblp:78/2022-17 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-6309-8370ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Hypergraph learning with multi-dimensional metabolite feature extractions and static-dynamic attention mechanisms to fill missing reactions in metabolic networks
abstract
Genome-scale metabolic models (GEMs) can effectively facilitate many fields in synthetic biology, biomanufacturing, and biomedicine. Reconstructing high-quality GEMs is crucial for accurate phenotype predictions of organisms. However, draft GEMs generated by automated reconstruction tools contain many knowledge gaps, especially missing reactions. The existing machine learning-based gap-filling approaches need to be further developed. In this article, we propose a novel HyperGraph Learning approach with Multi-dimensional metabolite feature extractions and static-dynamic Attention mechanisms (HGLMA) for predicting and teasing out missing reactions in GEM gap-fillings. HGLMA simultaneously uses two pretrained language models to proceed multi-dimensional metabolite feature extractions, which are further fused and regarded as node embeddings for graph learning. The directed and high-order associations between metabolites in reactions of GEMs are deeply mined by successively employing a directional graph network and a hypergraph neural network. Before outputting the predicted confidence score for candidate reactions, the static-dynamic multi-head attention mechanism is utilized to automatically learn attention weights and to identify key metabolites within any candidate reaction. The five-fold cross-validation results on 108 BiGG GEMs show that HGLMA significantly outperforms other state-of-the-art machine learning-based approaches both in prediction performances and in the ability of discovering missing reactions from metabolic reaction pools. The ablation study shows the contributions of multi-dimensional feature extractions and static-dynamic attention mechanisms. In addition, the phenotype prediction results of 24 bacterial organisms demonstrate the effectiveness and superiority of gap-fillings by HGLMA.
Kai Wang 0017, Jiajun Qu, Fei Liu 0001, Xiaoli Luan
Briefings Bioinform.1
2025 GRLGRN: graph representation-based learning to infer gene regulatory networks from single-cell RNA-seq data
abstract
BACKGROUND: A gene regulatory network (GRN) is a graph-level representation that describes the regulatory relationships between transcription factors and target genes in cells. The reconstruction of GRNs can help investigate cellular dynamics, drug design, and metabolic systems, and the rapid development of single-cell RNA sequencing (scRNA-seq) technology provides important opportunities while posing significant challenges for reconstructing GRNs. A number of methods for inferring GRNs have been proposed in recent years based on traditional machine learning and deep learning algorithms. However, inferring the GRN from scRNA-seq data remains challenging owing to cellular heterogeneity, measurement noise, and data dropout. RESULTS: In this study, we propose a deep learning model called graph representational learning GRN (GRLGRN) to infer the latent regulatory dependencies between genes based on a prior GRN and data on the profiles of single-cell gene expressions. GRLGRN uses a graph transformer network to extract implicit links from the prior GRN, and encodes the features of genes by using both an adjacency matrix of implicit links and a matrix of the profile of gene expression. Moreover, it uses attention mechanisms to improve feature extraction, and feeds the refined gene embeddings into an output module to infer gene regulatory relationships. To evaluate the performance of GRLGRN, we compared it with prevalent models and performed ablation experiments on seven cell-line datasets with three ground-truth networks. The results showed that GRLGRN achieved the best predictions in AUROC and AUPRC on 78.6% and 80.9% of the datasets, and achieved an average improvement of 7.3% in AUROC and 30.7% in AUPRC. The interpretation discussion and the network visualization were conducted. CONCLUSIONS: The experimental results and case studies illustrate the considerable performance of GRLGRN in predicting gene interactions and provide interpretability for the prediction tasks, such as identifying hub genes in the network and uncovering implicit links.
Kai Wang 0017, Fei Liu 0001, Xiaoli Luan, Xinglong Wang
BMC Bioinform.1
2025 PLMAM-PLA: A Method Using Pretrained Language Models and Attention Mechanisms for Protein-Ligand Binding Affinity Prediction
abstract
Protein-ligand binding affinity measures the strength of interactions between proteins and ligands. Accurately predicting this value is crucial for drug discovery and estimating enzyme kinetic parameters. In recent years, various computational models based on deep learning algorithms have been developed for predicting protein-ligand binding affinity. Most of these require data on protein structure or pockets in addition to protein sequences and ligand SMILES strings. Although integrating structural or pocket information can enhance prediction performances, sequence-based affinity prediction methods using only protein sequences and ligand SMILES strings are more convenient and efficient in practice. We have developed a novel sequence-based deep learning model, called PLMAM-PLA, to predict protein-ligand binding affinity. This model simultaneously extracts global and local features from both protein sequences and ligand SMILES by leveraging pretrained language models (ESM-2 and MolFormer) and dilated convolutional neural networks. The features are enhanced by SKNets and SENets and are further fused by successively using cross-attention and self-attention mechanisms. The output module provides the final affinity prediction value. Ablation studies emphasize the important contributions of the different modules, while visualization experiments demonstrate the efficacy of PLMAM-PLA in capturing meaningful feature representations. Additionally, case studies highlight the powerful generalization capabilities of the model, while comparisons with state-of-the-art models confirm its superior performance in predicting protein-ligand binding affinities.
Kai Wang 0017, Aijie Song, Fei Liu 0001, Xiaoli Luan, Xinglong Wang
IEEE Trans. Comput. Biol. Bioinform.1
2024 Machine learning-assisted substrate binding pocket engineering based on structural information
abstract
Engineering enzyme-substrate binding pockets is the most efficient approach for modifying catalytic activity, but is limited if the substrate binding sites are indistinct. Here, we developed a 3D convolutional neural network for predicting protein-ligand binding sites. The network was integrated by DenseNet, UNet, and self-attention for extracting features and recovering sample size. We attempted to enlarge the dataset by data augmentation, and the model achieved success rates of 48.4%, 35.5%, and 43.6% at a precision of ≥50% and 52%, 47.6%, and 58.1%. The distance of predicted and real center is ≤4 Å, which is based on SC6K, COACH420, and BU48 validation datasets. The substrate binding sites of Klebsiella variicola acid phosphatase (KvAP) and Bacillus anthracis proline 4-hydroxylase (BaP4H) were predicted using DUnet, showing high competitive performance of 53.8% and 56% of the predicted binding sites that critically affected the catalysis of KvAP and BaP4H. Virtual saturation mutagenesis was applied based on the predicted binding sites of KvAP, and the top-ranked 10 single mutations contributed to stronger enzyme-substrate binding varied while the predicted sites were different. The advantage of DUnet for predicting key residues responsible for enzyme activity further promoted the success rate of virtual mutagenesis. This study highlighted the significance of correctly predicting key binding sites for enzyme engineering.
Xinglong Wang, Kangjie Xu, Kai Linghu, Beichen Zhao, Shangyang Yu, Shuyao Yu, Weizhu Zeng, Kai Wang 0017
Briefings Bioinform.11
2024 BERT-TFBS: a novel BERT-based model for predicting transcription factor binding sites by transfer learning
abstract
Transcription factors (TFs) are proteins essential for regulating genetic transcriptions by binding to transcription factor binding sites (TFBSs) in DNA sequences. Accurate predictions of TFBSs can contribute to the design and construction of metabolic regulatory systems based on TFs. Although various deep-learning algorithms have been developed for predicting TFBSs, the prediction performance needs to be improved. This paper proposes a bidirectional encoder representations from transformers (BERT)-based model, called BERT-TFBS, to predict TFBSs solely based on DNA sequences. The model consists of a pre-trained BERT module (DNABERT-2), a convolutional neural network (CNN) module, a convolutional block attention module (CBAM) and an output module. The BERT-TFBS model utilizes the pre-trained DNABERT-2 module to acquire the complex long-term dependencies in DNA sequences through a transfer learning approach, and applies the CNN module and the CBAM to extract high-order local features. The proposed model is trained and tested based on 165 ENCODE ChIP-seq datasets. We conducted experiments with model variants, cross-cell-line validations and comparisons with other models. The experimental results demonstrate the effectiveness and generalization capability of BERT-TFBS in predicting TFBSs, and they show that the proposed model outperforms other deep-learning models. The source code for BERT-TFBS is available at https://github.com/ZX1998-12/BERT-TFBS.
Kai Wang 0017, Fei Liu 0001, Xiaoli Luan, Xinglong Wang
Briefings Bioinform.1