VLDB 2026 Research / reviewers in the wild / expert
Zhenlei Ye
dblp:345/9135
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0008-4000-9639ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aspect-centric vulnerability understanding via semantics-aware commit representation learning
Xiaobing Sun 0001, Sicong Cao, Zhenlei Ye |
Empir. Softw. Eng. | 4 |
| 2025 | Evaluating the Test Adequacy of Benchmarks for LLMs on Code GenerationabstractABSTRACT Code generation for users' intent has become increasingly prevalent with the large language models (LLMs). To automatically evaluate the effectiveness of these models, multiple execution‐based benchmarks are proposed, including specially crafted tasks, accompanied by some test cases and a ground truth solution. LLMs are regarded as well‐performed in code generation tasks if they can pass the test cases corresponding to most tasks in these benchmarks. However, it is unknown whether the test cases have sufficient test adequacy and whether the test adequacy can affect the evaluation. In this paper, we conducted an empirical study to evaluate the test adequacy of the execution‐based benchmarks and to explore their effects during evaluation for LLMs. Based on the evaluation of the widely used benchmarks, HumanEval, MBPP, and two enhanced benchmarks HumanEval+ and MBPP+, we obtained the following results: (1) All the evaluated benchmarks have high statement coverage (above 99.16%), low branch coverage (74.39%) and low mutation score (87.69%). Especially for the tasks with higher cyclomatic complexities in the HumanEval and MBPP, the mutation score of test cases is lower. (2) No significant correlation exists between test adequacy (statement coverage, branch coverage and mutation score) of benchmarks and evaluating results on LLMs at the individual task level. (3) There is a significant positive correlation between mutation score‐based evaluation and another execution‐based evaluation metric () on LLMs at the individual task level. (4) The existing test case augmentation techniques have limited improvement in the coverage of test cases in the benchmark, while significantly improving the mutation score by approximately 34.60% and also can bring a more rigorous evaluation to LLMs on code generation. (5) The LLM‐based test case generation technique (EvalPlus) performs better than the traditional search‐based technique (Pynguin) in improving the benchmarks' test quality and evaluation ability of code generation. Xiangyue Liu 0002, Xiaobing Sun 0001, Lili Bo, Yufei Hu, Zhenlei Ye |
J. Softw. Evol. Process. | 6 |
| 2025 | KG4VA: Constructing Vulnerability Knowledge Graph for Software Vulnerability AssessmentabstractSoftware vulnerabilities pose serious threats to software security. When faced with multiple software vulnerabilities at the same time, it is urgent to determine whether the vulnerabilities are high-risk. Existing vulnerability assessment approaches only learn the mapping relationships between vulnerability descriptions and severity levels, while ignoring the sharing of the same or similar elements between vulnerabilities. Furthermore, solely focusing on vulnerability descriptions fails to accurately characterize the vulnerability behavior. In this paper, we propose a novel vulnerability knowledge graph (KG) to capture the relationships between vulnerabilities. To construct the vulnerability KG automatically, we propose to leverage vulnerability elements extracted from vulnerability descriptions to link different vulnerabilities. Based on the constructed KG, we further propose a novel KG-based vulnerability assessment (VA) approach KG4VA, which precisely finds the similar vulnerability for an encountered vulnerability description by analyzing and matching the elements entities based on the vulnerability KG. The experiment results show that KG4VA outperforms the baselines in almost all metrics (e.g., 3.27%-10.83% accuracy improvements). Moreover, our ablation experiments demonstrate that the vulnerability knowledge graph can indeed offer valuable information for vulnerability assessment. Zhenlei Ye, Xiaobing Sun 0001, Lili Bo, Sicong Cao, Xiaoxue Ren, Lianyong Qi, Jiale Zhang 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Automatic software vulnerability assessment by extracting vulnerability elements
Xiaobing Sun 0001, Zhenlei Ye, Lili Bo, Xiaoxue Wu 0001, Ying Wei 0012, Tao Zhang 0001, Bin Li 0006 |
J. Syst. Softw. | 2 |