VLDB 2026 Research / reviewers in the wild / expert
Xinyu She
dblp:359/6063
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0001-2988-7042ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pitfalls in Language Models for Code Intelligence: A Taxonomy and SurveyabstractModern Language Models (LMs) have been successfully employed in source code generation and understanding, leading to a significant increase in research focused on learning-based code intelligence, such as automated bug repair and test case generation. Despite their great potential, LMs for code intelligence (LM4Code) are susceptible to potential pitfalls , which hinder realistic performance and further impact their reliability and applicability in real-world deployment . Such challenges drive the need for a comprehensive understanding—not just identifying these issues but delving into their possible implications and existing solutions to build more reliable LMs tailored to code intelligence. Based on a well-defined systematic research approach, we conducted an extensive literature review to uncover the pitfalls inherent in LM4Code. Finally, 121 primary studies from top-tier venues have been identified. After carefully examining these studies, we designed a taxonomy of pitfalls in LM4Code research and conducted a systematic study to summarize the issues, current solutions, implications, and challenges of different pitfalls for LM4Code systems. We developed a comprehensive classification scheme that dissects pitfalls across four crucial aspects: data collection and labeling, system design and learning, performance evaluation, and deployment and maintenance. Through this study, we aim to provide a roadmap for researchers and practitioners, facilitating their understanding and utilization of LM4Code in reliable and trustworthy ways. Xinyu She, Yue Liu 0011, Yanjie Zhao 0001, Yiling He, Li Li 0029, Chakkrit Tantithamthavorn, Zhan Qin, Haoyu Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | WASMaker: Differential Testing of WebAssembly Runtimes via Semantic-Aware Binary GenerationabstractA fundamental component of the Wasm ecosystem is the Wasm runtime, as it directly impacts whether Wasm applications can be executed as expected. Bugs in Wasm runtimes are frequently reported, so the research community has made a few attempts to design automated testing frameworks to detect bugs in Wasm runtimes. However, existing testing frameworks are limited by the quality of test cases, i.e., they face challenges in generating Wasm binaries that are both semantically rich and syntactically correct. As a result, complicated bugs cannot be triggered effectively. In this work, we present WASMaker, a novel differential testing framework that can generate complicated Wasm test cases by disassembling and assembling real-world Wasm binaries, which can trigger hidden inconsistencies among Wasm runtimes. To further pinpoint the root causes of unexpected behaviors, we design a runtime-agnostic root cause location method to locate bugs accurately. Extensive evaluation suggests that WASMaker outperforms state-of-the-art techniques in terms of both efficiency and effectiveness. We have uncovered 33 unique bugs in popular Wasm runtimes, among which 25 have been confirmed. Shangtong Cao, Ningyu He, Xinyu She, Mu Zhang 0001, Haoyu Wang 0001 |
ISSTA | 3 |
| 2024 | WaDec: Decompiling WebAssembly Using Large Language ModelabstractWebAssembly (abbreviated Wasm) has emerged as a cornerstone of web development, offering a compact binary format that allows high-performance applications to run at near-native speeds in web browsers. Despite its advantages, Wasm's binary nature presents significant challenges for developers and researchers, particularly regarding readability when debugging or analyzing web applications. Therefore, effective decompilation becomes crucial. Unfortunately, traditional decompilers often struggle with producing readable outputs. While some large language model (LLM)-based decompilers have shown good compatibility with general binary files, they still face specific challenges when dealing with Wasm. Xinyu She, Yanjie Zhao 0001, Haoyu Wang 0001 |
ASE | 1 |