VLDB 2026 Research / reviewers in the wild / expert
Ruimeng Li
dblp:262/7135
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal learning on heterogeneous subgraphs and LLMs representation for MHC-peptide binding affinity predictionabstractAccurate prediction of MHC-peptide binding affinity remains a challenge for immunotherapeutic development. Existing methods struggle to jointly model functional semantics of polymorphic residues, evolutionary conservation constraints, and structural dynamic. We propose the Contrast learning-based Multi-feature Heterogeneous Subgraph model (CMHS) with sequence and structural representation. For sequence representation, we introduce LoRA fine-tuning to obtain the MHC-exclusive sequence representation from ESM2, then jointly BLOSUM50 to capture long-range functional dependencies and evolutionarily conserved residues. For structural representation, we use the biophysics-guided heterogeneous graph network. Constructing an MHC-peptide graph with a novel trainable Gaussian noise layer guided by crystallographic B-factors to dynamically simulate electron density uncertainty, coupled with a three-stage message-passing framework with subgraph aggregation, subgraph extraction and heterogeneous. Finally, to align sequence and graph representation spaces, we use contrastive learning to obtain a more comprehensive representation and to enhance the ability of model prediction. Evaluations on 16 HLA allele benchmarks show average SRCC improvements of 8.7%, with improvements of average AUC of 7.6%. This work establishes a new paradigm for predicting hypervariable immune interactions. The corresponding code can be founded in github. Ruimeng Li, Haozhou Li, Biyi Zhou, Qinke Peng |
BMC Bioinform. | 1 |
| 2026 | SEHLP: A summary-enhanced large language model for financial report sentiment analysis via hybrid LoRA and dynamic prefix tuning
Haozhou Li, Qinke Peng, Xu Mou, Zeyuan Zeng, Ruimeng Li, Jinzhi Wang, Wentong Sun |
Inf. Process. Manag. | 5 |
| 2025 | ESM2_AMP: an interpretable framework for protein-protein interactions prediction and biological mechanism discoveryabstractThe prediction of binary protein-protein interactions (PPIs) is essential for protein engineering, but a major challenge in deep learning-based methods is the unknown decision-making process of the model. To address this challenge, we propose the ESM2_AMP framework, which utilizes the ESM2 protein language model for extracting segment features from actual amino acid sequences and integrates the Transformer model for feature fusion in binary PPIs prediction. Further, the two distinct models, ESM2_AMPS and ESM2_AMP_CSE are developed to systematically explore the contributions of segment features and combine with special tokens features in the decision-making process. The experimental results reveal that the model relying on segment features demonstrates strong correlations between segments with high attention weights and known functional regions of amino acid sequences. This insight suggests that attention to these segments helps capture biologically relevant functional and interaction-related information. By analyzing the coverage relationship between high-attention sequence fragments and functional regions, we validated the model's ability to capture key segment features of PPIs and revealed the critical role of functional domains in PPIs. This finding not only enhances the interpretability methods for sequence-based prediction models but also provides biological evidence supporting the important regulatory role of functional sequences in protein-protein interactions. It offers cross-disciplinary insights for algorithm optimization and experimental validation research in the field of computational biology. Yawen Sun, Zeyu Luo, Lejia Tan, Ruimeng Li, Yu-Juan Zhang |
Briefings Bioinform. | 6 |
| 2025 | CSTE: A Context-enhanced Speaker-aware Triple Encoding model for intelligent triage and diagnosis in medical dialogue
Haozhou Li, Qinke Peng, Xinyuan Wang 0011, Wentong Sun, Defu Li, Ruimeng Li |
Inf. Process. Manag. | 6 |
| 2025 | Invariant representation learning via decoupling style and spurious features
Ruimeng Li, Yuanhao Pu, Chenwang Wu, Hong Xie 0004, Defu Lian |
Mach. Learn. | 1 |
| 2025 | Compound Interaction Presentation Learning for MHC-Peptide Binding Affinity PredictionabstractThe interaction between peptides and Major Histocompatibility Complex Class I (MHC-I) molecules plays a critical role in adaptive immune recognition. Although computational prediction algorithms have advanced over traditional experimental methods, challenges still remain. There is a scarcity of standardized datasets that provide comprehensive profiles of MHC-peptide structure. The polymorphism of MHC molecules introduces diverse binding patterns, complicating the characterization of specific amino acid interaction pairs. To address these issues, we introduce GSM, a novel deep learning model that combines a Graph Attention Neural Network with a Self-Attention Convolutional Neural Network to predict MHC-peptide binding affinities. By integrating self-attention mechanisms to capture global peptide-MHC interactions and graph-based modeling to represent local amino acid pairwise interactions, GSM provides a comprehensive understanding of binding mode. Compared to existing algorithms, GSM shows superior performance and greater stability across diverse allele datasets, as demonstrated on the benchmarks. Furthermore, by leveraging real 3D structural data and attention visualization, GSM is capacity of selecting interaction sites, offering valuable insights for vaccine design and advancing immunological research. Ruimeng Li, Qinke Peng, Haozhou Li, Zeyuan Zeng |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | Dynamic Graph Attention Meets Pretrained Language Models: Adaptive K-Mer Decomposition for LncRNA-Protein Interaction PredictionabstractProtein-RNA complexes, particularly those involving RNA-binding proteins and long non-coding RNAs (lncRNA), are commonly found to influence gene expression and mediate fundamental cellular processes. Despite significant advances in representations for these biological sequences, sequence decomposition based on k-mer generally results in fix-length substrings, failing to detect the information of variable-length biological functional regions. In this paper, we develop a concept of expressiveness for k-mer decompositions as a theoretical underpinning for traversing all k-mer decompositions. Based on this concept, we propose an advanced approach, BERTDGA-LPI, to detect the information of variable-length biological functional regions utilizing dynamic graph attention and to capture the influence of RNA and protein context leveraging pretrained language models. The experimental results demonstrate the outperformance of BERTDGA-LPI over state-of-the-art methods across two homo sapiens datasets, one plant species dataset, and two species-unspecific datasets. Furthermore, BERTDGA-LPI is validated as effective in predicting unknown RNA-protein interactions (RPI) with 100% prediction accuracy in six independent validation sets from different species. This study lays a theoretical underpinning for traversing all k-mer decompositions and innovatively offers a broadly applicable and efficient tool for LPI prediction and RPI prediction based only on sequences. Zeyuan Zeng, Jingxian Zeng, Defu Li, Qinke Peng, Haozhou Li, Ruimeng Li, Wentong Sun, Jinzhi Wang |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2025 | A Finetuning Deep Learning Framework for Pan-Species Promoters Identification With Pseudo Time Series Analysis on Time and Frequency SpaceabstractPromoters are genomic sequences harbouring specific motifs, such as the TATA- box for eukaryotes and the Pribnow box for prokaryotes, which are known as regulatory elements. Accurate identification of these regulatory elements is essential for deciphering transcriptional regulation mechanisms. However, the heterogeneity of promoters across different species poses a significant challenge in this task. In our study, we introduce two deep learning methods, ProTriCNN and TransPro, designed for promoter identification. Based on promoter representation, ProTriCNN treats promoters as pseudo-time series, utilizing this approach to capture the intricate heterogeneity of promoter elements. TransPro is a ProTriCNN-based Fine-tuning framework to improve identification performance across different species. TransPro lies in utilizes elements and species evolutionary trees to represent the locality difference between source and target species across various levels and time-frequency space, respectively. With systematic experiments using real datasets, we demonstrate that ProTriCNN outperfroms state-of-the-art methods across all species, achieving an average accuracy improvement of 2.1% and a 20% enhancement in the Matthews coefficient. TransPro further attains accuracy improvement of the highest 8% and a 25% enhancement in the Matthews coefficient compared to ProTriCNN. Ruimeng Li, Qinke Peng, Haozhou Li, Wentong Sun |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Graph Neural Networks With High-Order Polynomial Spectral FiltersabstractGeneral graph neural networks (GNNs) implement convolution operations on graphs based on polynomial spectral filters. Existing filters with high-order polynomial approximations can detect more structural information when reaching high-order neighborhoods but produce indistinguishable representations of nodes, which indicates their inefficiency of processing information in high-order neighborhoods, resulting in performance degradation. In this article, we theoretically identify the feasibility of avoiding this problem and attribute it to overfitting polynomial coefficients. To cope with it, the coefficients are restricted in two steps, dimensionality reduction of the coefficients' domain and sequential assignment of the forgetting factor. We transform the optimization of coefficients to the tuning of a hyperparameter and propose a flexible spectral-domain graph filter, which significantly reduces the memory demand and the adverse impacts on message transmission under large receptive fields. Utilizing our filter, the performance of GNNs is improved significantly in large receptive fields and the receptive fields of GNNs are multiplied as well. Meanwhile, the superiority of applying a high-order approximation is verified across various datasets, notably in strongly hyperbolic datasets. Codes are publicly available at: https://github.com/cengzeyuan/TNNLS-FFKSF. Zeyuan Zeng, Qinke Peng, Xu Mou, Ying Wang 0065, Ruimeng Li |
IEEE Trans. Neural Networks Learn. Syst. | 5 |