VLDB 2026 Research / reviewers in the wild / expert
Xiaoheng Xie
dblp:359/0480
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0006-9700-0542ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Datalog-Based Language-Agnostic Change Impact Analysis for MicroservicesabstractThe shift-left principle in the industry requires us to test a software application as early as possible. In particular, when code changes in a microservice application are committed to the code repository, we have to efficiently identify all public microservice interfaces affected by the changes so that the impacted interfaces can be tested as soon as possible. However, developing an efficient change impact analysis is extremely challenging in microservices due to the multilingual problem: microservice applications are often implemented using varying programming languages and involve diverse frameworks and configuration files. To address this issue, this paper presents MICROSCOPE, a language-agnostic change impact analysis that uniformly represents code, configuration files, frameworks, and code changes by relational Datalog rules. MICROSCOPE then benefits from an efficient Datalog solver to identify impacted interfaces. Experiments based on the use of MICROSCOPE in Ant Group, a leading software vendor, demonstrate that MICROSCOPE is both effective and fast, as it successfully identifies interfaces affected by 112 code commits, with moderate time overhead, and could reduce 97% of interfaces to test and save 73% of testing time after code changes. Qingkai Shi, Xiaoheng Xie, Xianjin Fu, Peng Di, Ang Zhou, Gang Fan |
ICSE | 2 |
| 2024 | RepoGenix: Dual Context-Aided Repository-Level Code Completion with Language ModelsabstractThe success of language models in code assistance has spurred the proposal of repository-level code completion as a means to enhance prediction accuracy, utilizing the context from the entire codebase. However, this comprehensive context comes at a cost: while it enhances model performance, it also increases inference latency. This balance between improved accuracy and computational efficiency poses a significant challenge in real-world applications. We present RepoGenix, a solution that enhances repository-level code completion without increased latency. RepoGenix combines analogous context and relevant context, using Context-Aware Selection technology to efficiently compress these contexts into limited-size prompts. Our experiments on CrossCodeEval demonstrate that RepoGenix not only achieves a substantial 48.41% reduction in inference time, but also yields improvement in performance compared to baseline methods. We have successfully implemented and tested RepoGenix within AntGroup's development environments. This approach is being extended to multiple programming languages and will be open-sourced, aiming to enhance code completion efficiency for the broader developer community. Xiaoheng Xie, Gehao Zhang, Xunjin Zheng, Peng Di, Wei Jiang 0041, Chengpeng Wang 0001, Gang Fan |
ASE | 2 |
| 2024 | GrayDuck: The Sword of Damocles for Duck Typing in Dynamic Language DeserializationabstractDuck typing is a flexible programming style in dynamic languages, enabling the achievement of complex behaviors using less code. The use of duck typing is currently widespread; however, the question is whether its use in code is truly safe. In fact, improper use of duck typing may introduce unexpected security threats. In this paper, we reveal another side of duck typing, showing how it can exacerbate the impact of deserialization vulnerabilities and expand the range of attack options for attackers. We present three cases of duck typing misuse and theoretically demonstrate how such misuse can expand the attack surface of deserialization vulnerabilities. Additionally, we design a static analysis tool, GrayDuck, to construct a Class Relation Graph (CRG) that clearly delineates the range of classes accessible through each deserialization operation and identify instances of duck typing misuse along with the associated attack surfaces so that to assess the potential harm. We utilized this tool to scan 5 Python programs known to have real deserialization vulnerabilities, detecting 7 issues of deserialized object duck typing misuse and calculating the corresponding expansions of the attack surfaces. Xunjin Zheng, Cai Fu, Xiaoheng Xie, Peng Di |
ASE | 4 |
| 2024 | LLMDFA: Analyzing Dataflow in Code with Large Language ModelsabstractDataflow analysis is a fundamental code analysis technique that identifies dependencies between program values. Traditional approaches typically necessitate successful compilation and expert customization, hindering their applicability and usability for analyzing uncompilable programs with evolving analysis needs in real-world scenarios. This paper presents LLMDFA, an LLM-powered compilation-free and customizable dataflow analysis framework. To address hallucinations for reliable results, we decompose the problem into several subtasks and introduce a series of novel strategies. Specifically, we leverage LLMs to synthesize code that outsources delicate reasoning to external expert tools, such as using a parsing library to extract program values of interest and invoking an automated theorem prover to validate path feasibility. Additionally, we adopt a few-shot chain-of-thought prompting to summarize dataflow facts in individual functions, aligning the LLMs with the program semantics of small code snippets to mitigate hallucinations. We evaluate LLMDFA on synthetic programs to detect three representative types of bugs and on real-world Android applications for customized bug detection. On average, LLMDFA achieves 87.10% precision and 80.77% recall, surpassing existing techniques with F1 score improvements of up to 0.35. We have open-sourced LLMDFA at https://github.com/chengpeng-wang/LLMDFA. Chengpeng Wang 0001, Wuqi Zhang, Zian Su, Xiangzhe Xu, Xiaoheng Xie, Xiangyu Zhang 0001 |
NeurIPS | 5 |