EDBT 2026 Demo / reviewers in the wild / expert
Xinrui Li 0004
dblp:151/4579-4
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0002-4730-5277ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving the ability of pre-trained language model by imparting large language model's experience
Chao Ni 0001, Xinrui Li 0004, Xiaohu Yang 0001 |
J. Syst. Softw. | 3 |
| 2026 | Abundant Modalities Offer More Nutrients: Multi-Modal-Based Function-Level Vulnerability DetectionabstractSoftware vulnerabilities are weaknesses in software systems that can lead to significant cybersecurity risks. Recently, several deep learning (DL)-based approaches have been proposed to detect vulnerabilities at the function level. These approaches typically utilize one or a few different modalities (e.g., text representation and graph-based representation) of the function, and have shown promising performance. However, existing studies have not fully leveraged diverse modalities, particularly those that use images to represent functions for vulnerability detection. These approaches often fail to make sufficient use of the important graph structure underlying the images. In this article, we propose MVulD+, a multi-modal-based function-level vulnerability detection approach, which fuses multi-modal features of the function (i.e., text representation, graph representation, and image representation) to detect vulnerabilities. Specifically, MVulD+ leverages a pre-trained model (i.e., UniXcoder) to capture the semantic information of the textual source code, uses a graph neural network to extract graph representations, and employs computer vision techniques to obtain image representations while preserving the graph structure of the function. To investigate the effectiveness of MVulD+, we conduct a large-scale experiment by comparing our approach with nine state-of-the-art baselines. Experimental results demonstrate that MVulD+ improves the DL-based baselines by 24.3–125.7%, 5.2–31.4%, 40.6–192.2%, and 22.3–186.9% in terms of F1-score, Accuracy, Precision, and PR-AUC, respectively. Chao Ni 0001, Xinrui Li 0004, Xiaodan Xu |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | Reliable Code Generation with Test Case Prioritization and Cognitive ValidationabstractLarge Language Models (LLMs) have shown impressive capabilities in code generation. However, they often struggle in complex programming scenarios due to incomplete semantic understanding and limited ability to correct misunderstandinginduced errors. While recent efforts have incorporated test cases to guide task comprehension, they typically overlook the quality and relevance of the test cases, reducing their effectiveness in steering accurate code generation. To overcome these limitations, we present PriGen, a multiagent collaborative framework for prioritized and test case driven code generation. PriGen introduces a novel test case prioritization mechanism that selects a high-value subset based on semantic coverage, boundary sensitivity, and error-triggering potential. These curated test cases assist in refining the LLM’s task understanding. Additionally, PriGen integrates a Cognitive Validation Loop, which iteratively verifies and improves the model’s comprehension through interactive evaluation and dynamic test injection, ensuring semantic alignment before code synthesis. We evaluate PriGen on two enhanced benchmarks, HumanEvalET and MBPP-ET, using three representative open-source LLMs: DeepSeek-Coder, Qwen2.5-Coder, and Llama-3.1. Experimental results show that PriGen consistently outperforms state-of-the-art baselines in both correctness and efficiency, demonstrating its effectiveness and generalizability in enhancing LLM-based code generation. Lingyun Huang, Xinrui Li 0004, Chao Ni 0001 |
APSEC | 3 |
| 2025 | Enhancing Commit Classification for Software Maintenance with Adversarial LearningabstractAccurately classifying developer contributions is essential for improving open-source software development workflows and enabling effective contributor incentive mechanisms. However, existing commit message classification methods primarily rely on traditional machine learning or standard deep learning models, which often fail to capture the rich semantics embedded in commit messages, leading to suboptimal performance. This paper introduces CoMAL, a novel framework that combines adversarial training with pre-trained BERT models to enhance the robustness and accuracy of commit message classification. To support evaluation, we construct the GitHub Commit Dataset (GCD)-a large-scale, manually labeled dataset comprising 123,325 commit messages from six widely-used opensource projects across three programming languages (C, Python, and Java), categorized into six contribution types: Fix, Feature Addition, Test, Refactoring, Docs, and Environment. We conduct comprehensive empirical studies comparing CoMAL with three SOTA baselines across five evaluation metrics. Experimental results show that CoMAL consistently outperforms baselines, achieving an accuracy of 0.90 and a macro-average F1-score of 0.87, representing improvements of 7% to 43% in accuracy and 5% to 38% in F1-score over baselines. Xinrui Li 0004, Chao Ni 0001 |
APSEC | 1 |
| 2025 | Sembug: Detecting Logic Bugs in Dbms Through Generating Semantic-Aware Non-Optimizing QueryabstractLogic bugs, which cause Database Management Systems (DBMSs) to return incorrect results, are challenging to detect due to the absence of explicit signs such as system crashes. The majority of these bugs originate from the query optimizer and are commonly referred to as optimization bugs. Many approaches have been proposed for detecting logic bugs, which can be divided into two groups. The first group aims to detect the optimization bugs but only focuses on those with incorrect results cardinality, neglecting to check semantic correctness and consequently limiting the detection of bugs in advanced DBMS features. For the second group, though it can verify the correctness of the results for both their cardinality and semantics, it is ineffective in handling optimization bugs, which restricts its practical usage effectiveness. In this paper, we propose Semantic-aware Non-Optimizing Query (SemBug), a novel approach for logic bug detection in DBMSs. SemBug focuses on optimization bugs by transforming the queries that can be highly optimized by DBMS into equivalent but less optimized ones. Additionally, SemBug integrates semantic analysis technology, enabling it to identify semantic logic bugs and support testing advanced DBMS features. Any discrepancy in cardinality or content between the original and transformed queries indicates a logic bug. To investigate the effectiveness of SemBug, we conduct a large-scale experiment on five widelyused DBMS systems (i.e., MySQL, TiDB, MariaDB, SQLite, and PostgreSQL) and compare it with three state-of-the-art (SOTA) approaches (i.e., Pinolo, TLP, and NoREC). The experimental results indicate that SemBug outperforms three SOTAs. Over 24 hours, SemBug found 34 unique logic bugs, which are 19, 14, and 13 more bugs than each of the three SOTAs, marking an improvement of$126 \%, 70 \%$, and 61 % respectively. As of the time of paper submission, SemBug has uncovered 37 unique logic bugs, of which 29 have been verified by developers, and 11 have been fixed. SemBug helps developers identify these bugs, providing insights into such inconsistencies and assisting in resolving them. Shiyang Ye, Chao Ni 0001, Qianqian Pang, Xinrui Li 0004, Xiaodan Xu |
ICPC | 5 |
| 2025 | Enhancing LLM's Ability to Generate More Repository-Aware Unit Tests Through Precise Context InjectionabstractRecently, Large Language Models (LLMs) have gained attention for their ability to handle a broad range of tasks, including unit test generation. Despite their success, LLMs may exhibit hallucinations when generating unit tests for focal methods or functions due to their lack of awareness regarding the project’s global context. While many studies have explored the role of context, they often extract fixed patterns of context for different models and focal methods, which may not be suitable for all generation processes (e.g., excessive irrelevant context could lead to redundancy, preventing the model from focusing on essential information).To overcome this limitation, we propose RATester, which integrates language servers to provide dynamic definition lookup to assist the LLM. When RATester encounters an unfamiliar identifier, it first leverages language servers (e.g., Gopls) to fetch relevant definitions and documentation comments, and then uses this global knowledge to guide the LLM. We evaluate the effectiveness and efficiency of RATester by constructing a new Golang dataset from real-world projects. On our Golang dataset, RATester achieves an average line coverage of 26.25%, representing an improvement of 9.10% to 165.69% over the baselines. In mutation testing, RATester shows superior performance by successfully killing 18 to 147 more mutants than the baselines. Additionally, our model-agnostic and generalizability analysis confirms RATester’s effectiveness across different models, programming languages, and model scales, validating its broad applicability. Chao Ni 0001, Xinrui Li 0004, Liushan Chen, Guojun Ma, Xiaohu Yang 0001 |
ASE | 3 |