EDBT 2026 Demo / reviewers in the wild / expert
Guanglei Yu
dblp:312/3767
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OmiImp: A Cross-Omics Imputation Framework Based on Improved Generative Adversarial NetworkabstractThe integration of multi-omics data has emerged as a powerful approach to elucidate interactions across different biological levels. However, throughput limitations and high costs of sequencing technologies often result in sparse multi-omics datasets, where only a subset of samples contains complete omics profiles - a challenge known as the “block missing”. To address this challenge, we propose OmiImp, a novel computational framework based on an improved generative adversarial network (GAN) for cross-omics data imputation. Through performance evaluation on independent datasets, we demonstrate that OmiImp outperforms existing state-of-the-art imputation methods while maintaining stable performance across different missing rates. In addition, we perform enrichment analysis and the results demonstrate that differential expressed features are uniformly distributed across pathways, and the synthetic data retains utility in diverse prognostic analyses. Collectively, this methodological advancement facilitates more reliable multi-omics integration studies, particularly when handling incomplete datasets. Kai Zhao 0010, Xuehua Bi, Guanglei Yu, Linlin Zhang 0005 |
BIBM | 4 |
| 2025 | MOGATFF: An Explainable Multi-Omics Prediction Model with Feature Enhancement for Genotype-Phenotype Association Analysis
Guanglei Yu, Xuehua Bi |
ISBRA (2) | 3 |
| 2025 | CATH-ddG: towards robust mutation effect prediction on protein-protein interactions out of CATH homologous superfamilyabstractMOTIVATION: Protein-protein interactions (PPIs) are fundamental aspects in understanding biological processes. Accurately predicting the effects of mutations on PPIs remains a critical requirement for drug design and disease mechanistic studies. Recently, deep learning models using protein 3D structures have become predominant for predicting mutation effects. However, significant challenges remain in practical applications, in part due to the considerable disparity in generalization capabilities between easy and hard mutations. Specifically, a hard mutation is defined as one with its maximum TM-score <0.6 when compared to the training set. Additionally, compared to physics-based approaches, deep learning models may overestimate performance due to potential data leakage. RESULTS: We propose new training/test splits that mitigate data leakage according to the CATH homologous superfamily. Under the constraints of physical energy, protein 3D structures, and CATH domain objectives, we employ a hybrid noise strategy as data augmentation and present a geometric encoder scenario, named CATH-ddG, to represent the mutational microenvironment differences between wild-type and mutated protein complexes. Additionally, we fine-tune ESM2 representations by incorporating a lightweight nonlinear module to achieve the transferability of sequence co-evolutionary information. Finally, our study demonstrates that CATH-ddG framework provides enhanced generalization by outperforming other baselines on non-superfamily leakage splits, which plays a crucial role in exploring robust mutation effect regression prediction. Independent case studies demonstrate successful enhancement of binding affinity on 419 antibody variants to human epidermal growth factor receptor 2 (HER2) and 285 variants in the receptor-binding domain (RBD) of SARS-CoV-2 to angiotensin-converting enzyme 2 (ACE2) receptor. AVAILABILITY AND IMPLEMENTATION: CATH-ddG is available at https://github.com/ak422/CATH-ddG. Guanglei Yu, Xuehua Bi, Yaohang Li, Jianxin Wang 0001 |
Bioinform. | 1 |
| 2025 | HPOseq: a deep ensemble model for predicting the protein-phenotype relationships based on protein sequencesabstractBACKGROUND: Understanding the relationships between proteins and specific disease phenotypes contributes to the early detection of diseases and advances the development of personalized medicine. The acquisition of a large amount of proteomics data has facilitated this process. To improve discovery efficiency and reduce the time and financial costs associated with biological experiments, various computational methods have yielded promising results. However, the lack of rich and reliable protein-related information still presents challenges in this process. RESULTS: In this paper, we propose an ensemble prediction model, named HPOseq, which predicts human protein-phenotype relationships based only on sequence information. HPOseq establishes two base models to achieve objectives. One directly extracts internal information from amino acid sequences as protein features to predict the associated phenotypes. The other builds a protein-protein network based on sequence similarity, extracting information between proteins for phenotype prediction. Ultimately, an ensemble module is employed to integrate the predictions from both base models, resulting in the final prediction. CONCLUSION: The results of 5-fold cross-validation reveal that HPOseq outperforms seven baseline methods for predicting protein-phenotype relationships. Moreover, we conduct case studies from the points of phenotype annotation and protein analysis to verify the practical significance of HPOseq. Zhuocheng Ji, Na Quan, Guanglei Yu, Xuehua Bi |
BMC Bioinform. | 6 |
| 2024 | TKMBR: Temporal Knowledge Graph-based Multi-Behavior RecommendationabstractStriving to enhance predictive performance by leveraging auxiliary behaviors, multi-behavior recommendation models have emerged in in many different fields. These models aim to address the diversity and effectiveness of interactive behaviors. While some methods have shown promising effects, they still exhibit certain limitations, such as overlooking dynamic nature of user interactions. In this paper, we present TKMBR, a temporal knowledge graph-based framework for multi-behavior recommendation. TKMBR incorporates a temporal knowledge graph to capture the temporal dynamics of user behaviors, which allows for the identification of underlying temporal patterns and the capturing of evolving user preferences over time. To augment the understanding of user preferences, heterogeneous signals are integrated and an item-side information knowledge graph is constructed based on various user-item interactions. Moreover, contrastive learning tasks are employed to alleviate the issue of data sparsity. Evaluation on three datasets using HR and NDCG shows TKMBR’s effectiveness in improving recommendation quality. Xiaoman Zhang, Xuehua Bi, Guanglei Yu, Ruyi Cao |
IJCNN | 4 |
| 2024 | DDAffinity: predicting the changes in binding affinity of multiple point mutations using protein 3D structureabstractMOTIVATION: Mutations are the crucial driving force for biological evolution as they can disrupt protein stability and protein-protein interactions which have notable impacts on protein structure, function, and expression. However, existing computational methods for protein mutation effects prediction are generally limited to single point mutations with global dependencies, and do not systematically take into account the local and global synergistic epistasis inherent in multiple point mutations. RESULTS: To this end, we propose a novel spatial and sequential message passing neural network, named DDAffinity, to predict the changes in binding affinity caused by multiple point mutations based on protein 3D structures. Specifically, instead of being on the whole protein, we perform message passing on the k-nearest neighbor residue graphs to extract pocket features of the protein 3D structures. Furthermore, to learn global topological features, a two-step additive Gaussian noising strategy during training is applied to blur out local details of protein geometry. We evaluate DDAffinity on benchmark datasets and external validation datasets. Overall, the predictive performance of DDAffinity is significantly improved compared with state-of-the-art baselines on multiple point mutations, including end-to-end and pre-training based methods. The ablation studies indicate the reasonable design of all components of DDAffinity. In addition, applications in nonredundant blind testing, predicting mutation effects of SARS-CoV-2 RBD variants, and optimizing human antibody against SARS-CoV-2 illustrate the effectiveness of DDAffinity. AVAILABILITY AND IMPLEMENTATION: DDAffinity is available at https://github.com/ak422/DDAffinity. Guanglei Yu, Qichang Zhao, Xuehua Bi, Jianxin Wang 0001 |
Bioinform. | 1 |
| 2023 | Enhancing Protein Subcellular Localization Prediction Through Multi-Feature FusionabstractAccurately determining the subcellular location of proteins is essential for comprehending their functions, as it provides crucial insights into biochemical pathways and regulatory mechanisms. Although some methods have achieved promising effects, there are still some negative aspects, such as inappropriate feature engineering. In this paper, we propose a method for predicting the subcellular location of proteins that combines multiple features taken from several data sources. Firstly, we obtain three features, Di-peptide Composition, Moran correlation and Conjoint-Triad, from the amino acid sequence. We also employ node2vec to extract features from protein-protein interaction networks and combine them with gene ontology. To eliminate redundant information between features, we then fuse the multiple features from different data source with an auto-encoder. Finally, we employ a supervised learning model, Wide and Deep, to predict the subcellular location of proteins. The experimental results demonstrate that our approach achieves higher accuracy than state-of-the-art methods. This approach provides a promising solution for accurately predicting the subcellular location of proteins. Weiyang Liang, Xuehua Bi, Guanglei Yu, Na Quan |
SMC | 4 |
| 2023 | Predicting disease genes based on multi-head attention fusionabstractBACKGROUND: The identification of disease-related genes is of great significance for the diagnosis and treatment of human disease. Most studies have focused on developing efficient and accurate computational methods to predict disease-causing genes. Due to the sparsity and complexity of biomedical data, it is still a challenge to develop an effective multi-feature fusion model to identify disease genes. RESULTS: This paper proposes an approach to predict the pathogenic gene based on multi-head attention fusion (MHAGP). Firstly, the heterogeneous biological information networks of disease genes are constructed by integrating multiple biomedical knowledge databases. Secondly, two graph representation learning algorithms are used to capture the feature vectors of gene-disease pairs from the network, and the features are fused by introducing multi-head attention. Finally, multi-layer perceptron model is used to predict the gene-disease association. CONCLUSIONS: The MHAGP model outperforms all of other methods in comparative experiments. Case studies also show that MHAGP is able to predict genes potentially associated with diseases. In the future, more biological entity association data, such as gene-drug, disease phenotype-gene ontology and so on, can be added to expand the information in heterogeneous biological networks and achieve more accurate predictions. In addition, MHAGP with strong expansibility can be used for potential tasks such as gene-drug association and drug-disease association prediction. Dianrong Lu, Xuehua Bi, Guanglei Yu, Na Quan |
BMC Bioinform. | 5 |