VLDB 2026 Research / reviewers in the wild / expert
Qi Xin 0001
dblp:77/2901-1
· DBLP profile ↗
17ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-0543-4935ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 4 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PreMulBVD: A pretraining-based multi-modal binary vulnerability detection framework
Chenliang Xing, Xiaoyuan Xie, Qi Xin 0001, Gong Chen 0007 |
J. Syst. Softw. | 3 |
| 2026 | Studying and Improving the Soundness of Input-Based Feature-Oriented DebloatingabstractThe paper consists of two parts: a study of the soundness of feature-oriented debloating techniques that use inputs as the feature specification and a new blocking method BLOCKAUGwe proposed for soundness improvement. Feature-oriented debloating techniques aim to eliminate code bloat related to unneeded program features. Many of these techniques rely on a usage profile, typically provided as a set of inputs, for specification. Such input-based techniques tend to produce debloated programs that are overfitted to the inputs provided, introducing soundness issues often as bugs and vulnerabilities that pose severe threats to the program correctness and security. No prior work has systematically investigated the soundness of current debloating techniques and analyzed the types and causes of the soundness issues they introduce. To fill this gap, we conducted a study in which we applied 7 input-based techniques to 18 programs from two existing benchmarks for debloating and used three fuzzers with various sanitizers to detect soundness issues introduced by these techniques. Our results show that current techniques are highly unsound, as they can introduce a number of issues that lead to program crashes. A key reason for the issue introduction is the inappropriate deletion of soundness-related code such as conditional statements checking invalid cases, which, if missing, can result in unexpected program state and unconditioned execution.To improve the soundness of input-based debloating, we explored a blocking method that can be applied to coverage-based debloating. The core idea is to identify every deleted branch resulted from coverage-based code pruning and, instead of leaving the branch as empty, augment it to prevent any execution from passing through the branch and causing problems. To assess the effectiveness of the method, we used it to augment the debloated programs generated by four coverage-based techniques and evaluated the soundness and generality of the augmented programs. We found that the blocking method can significantly improve soundness, at the cost of only slightly increasing the program size. Although it can change program semantics, it does not significantly affect the generality by weakening the program’s ability in handling other feature-related inputs not seen while debugging. Moreover, the blocking method can forbid unexpected execution of any inputs the program should not have processed, thereby improving the program trustworthiness. Jiahao Yuan 0001, Weinuo Leng, Xuan Wei 0002, Qi Xin 0001, Xiaoyuan Xie, Jifeng Xuan |
IEEE Trans. Software Eng. | 4 |
| 2025 | Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code SearchabstractCode search is a vital activity in software engineering, focused on identifying and retrieving the correct code snippets based on a query provided in natural language. Approaches based on deep learning techniques have been increasingly adopted for this task, enhancing the initial representations of both code and its natural language descriptions. Despite this progress, there remains an unexplored gap in ensuring consistency between the representation spaces of code and its descriptions. Furthermore, existing methods have not fully leveraged the potential relevance between code snippets and their descriptions, presenting a challenge in discerning fine-grained semantic distinctions among similar code snippets. To address these challenges, we introduce a multi-task hedging contrastive Learning framework for Code Search, referred to as HedgeCode. HedgeCode is structured around two primary training phases. The first phase, known as the representation alignment stage, proposes a hedging contrastive learning approach. This method aims to detect subtle differences between code and natural language text, thereby aligning their representation spaces by identifying relevance. The subsequent phase involves multi-task joint learning, wherein the previously trained model serves as the encoder. This stage optimizes the model through a combination of supervised and self-supervised contrastive learning tasks. Our framework's effectiveness is demonstrated through its performance on the CodeSearchNet benchmark, showcasing HedgeCode's ability to address the mentioned limitations in code search tasks. Gong Chen 0007, Xiaoyuan Xie, Daniel Tang, Qi Xin 0001 |
ICSE | 4 |
| 2025 | Revisit the Intuition of Mutation-Based Fault Localization in Real-world ProgramsabstractMutation-based fault localization (MBFL) is an automated fault localization method that has been extensively studied in recent years.The intuition behind MBFL is based on the assumption that mutation operations can correct faults in a program.However, this assumption has only been experimented and validated on simulated datasets, and whether it truly holds in the real world has never been investigated.Fault types in simulated datasets are simple and differ significantly from the complex and diverse faults found in realworld programs.Therefore, to investigate whether MBFL works in the real world, it is necessary to validate its intuition in the real world.The goal of this study is to analyze whether the intuition of MBFL still holds in the real world.We quantified the MBFL intuition by establishing an algorithm, which eliminated the interference of factors unrelated to MBFL itself, allowing us to directly validate the intuition of MBFL.Based on this algorithm, we conducted extensive experiments on both real-world programs and programs in simulated datasets.The results revealed an interesting trend: due to the complexity of faults in real-world programs compared to those in simulated datasets, MBFL's intuition probably cannot hold in the real world.This indicates that MBFL's intuition is difficult to hold in the real world.Consequently, we focused on analyzing the real-world faulty versions and summarized a set of mutation operators that perform better in the real world by studying the types and effects of each mutant, providing guidance for the application of MBFL. Chenliang Xing, Gong Chen 0007, Qi Xin 0001, Xiaoyuan Xie |
Internetware | 3 |
| 2025 | Kotsuite: Unit Test Generation for Kotlin Programs in Android ApplicationsabstractUnit testing plays a pivotal role in safeguarding functional requirements and supporting the maintenance during the development of Android applications. The Kotlin programming language emerges in developing Android applications due to its simplicity, safety, and interoperability with Java. It is timeconsuming to manually write unit test cases for Kotlin programs. To mitigate labor costs, automated unit test generation techniques are developed. However, existing tools of unit test generation, such as EvoSuite and Randoop, are primarily optimized for traditional Java projects. This makes these tools incapable of generating test cases for Kotlin projects in Android. In this paper, we introduce KotSuite, an automated tool of unit test generation for Kotlin applications in Android. KotSuite employs static analysis techniques to extract the syntactic structure of the target methods and transforms the syntactic structure into the control flow representation. Then, KotSuite automatically generates a suite of test cases using a genetic algorithm and test reuse. We evaluate KotSuite on eight modules from four widely-used and opensource Kotlin projects in Android. Experimental results show that KotSuite can effectively generate high-coverage test cases with average line coverage of 66.0 % and branch coverage of 60.4%. Qi Xin 0001, Zhilei Ren, Jifeng Xuan |
ICPC | 2 |
| 2025 | ROSE: An IDE-Based Interactive Repair Framework for DebuggingabstractDebugging is costly. Automated program repair (APR) holds the promise of reducing its cost by automatically fixing errors. However, current techniques are not easily applicable in a realistic debugging scenario because they assume a high-quality test suite and frequent program re-execution, have low repair efficiency, and only handle a limited set of errors. To improve the practicality of APR for debugging, we propose ROSE, an interactive repair framework that is able to suggest quick and effective repairs of semantic errors while debugging in an Integrated Development Environment (IDE). ROSE allows an easy integration of existing APR patch generators and can do program repair without assuming the existence of a test suite and without requiring program re-execution. It works in conjunction with an IDE debugger and assumes a debugger stopping point where a problem symptom is observed. ROSE asks the developer to quickly describe the symptom. Then it uses the stopping point, the identified symptom, and the current environment to identify potentially faulty lines, uses a variety of APR techniques to suggest repairs at those lines, and validates those repairs without re-executing the program. Finally, it presents the results so the developer can examine, select, and make the appropriate repair. ROSE uses novel approaches to achieve effective fault localization and patch validation without a test suite or program re-execution. For fault localization, ROSE builds on a fast abstract interpretation-based flow analysis to compute a static backward slice approximating the real dynamic slice while taking into account the symptom and the current execution. For patch validation without re-running the program, ROSE generates simulated traces based on a live-programming system for both the original and repaired executions and compares the traces with respect to the problem symptoms to infer patch correctness. We implemented a prototype of ROSE that works in an Eclipse-based IDE and evaluated its potency and utility with an effectiveness study and a user study. We found that ROSE’s fault localization and validation are highly effective and a ROSE-based tool using existing APR patch generators generated correct repair suggestions for many errors in only seconds. Moreover, the user study demonstrated that ROSE was helpful for debugging and developers liked to use it. Steven P. Reiss, Xuan Wei 0002, Jiahao Yuan 0001, Qi Xin 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | Do not neglect what's on your hands: localizing software faults with exception trigger streamabstractExisting fault localization techniques typically analyze static information and run-time profiles of faulty software programs, and subsequently calculate suspiciousness values for each program entity. Such strategies typically have overbroad information to be analyzed and lead to unsatisfactory results. Exception is a widely-used programming language feature. It is closely related to the execution status during the execution of programs, and thus can be incorporated into automatic fault localization techniques for better effectiveness. Based on this intuition, we propose EXPECT, a novel fault localization technique that makes use of exception information, a valuable source of data for fault localization while being often ignored in previous research. Specifically, EXPECT first constructs exception trigger streams (including exception trigger information and execution traces), and then localizes faults by tracing bifurcation points between different exception trigger streams. Moreover, the tie-breaking problem can be also benefited from the use of exception trigger streams. Experimental results demonstrate the advantages of EXPECT: it achieves as high as 38.26% improvements in localizing faults regarding the Exam metric in comparison to the state-of-the-art fault localization technique, and it reduces the scales of ties in existing FL methods by up to 99.08%. Xihao Zhang, Xiaoyuan Xie, Qi Xin 0001, Chenliang Xing |
ASE | 4 |
| 2023 | Potential Solutions to Challenges in C Program Repair: A Practical PerspectiveabstractAutomated program repair is to reduce the manual work for bug fixing by human developers. In recent 15 years, the research community of program repair has created many novel techniques. However, these techniques share several assumptions that cannot always be satisfied in daily software development. This badly hurts the application of program repair in practice. For example, many repair techniques assume that test cases are well written before patch generation; many techniques assume that specific language features can be ignored (or already-processed). In this paper, we propose a framework of C program repair, which mainly addresses two challenges: test-independent repair and preprocessor directive processing. Our solution to test-independent repair is to automatically construct patch conditions for C programs via parsing the syntax structures; our solution to preprocessor directive processing is to generate code symbols to replace preprocessor directives. We plan to implement these potential solutions with program analysis techniques. The goal of this paper is to present practical solutions for developers to automate C program repair. Jifeng Xuan, Qi Xin 0001, Liqian Chen, Xiaoguang Mao |
ASE | 2 |
| 2022 | Automated test generation for REST APIs: no time to rest yetabstractModern web services routinely provide REST APIs for clients to access their functionality. These APIs present unique challenges and opportunities for automated testing, driving the recent development of many techniques and tools that generate test cases for API endpoints using various strategies. Understanding how these techniques compare to one another is difficult, as they have been evaluated on different benchmarks and using different metrics. To fill this gap, we performed an empirical study aimed to understand the landscape in automated testing of REST APIs and guide future research in this area. We first identified, through a systematic selection process, a set of 10 state-of-the-art REST API testing tools that included tools developed by both researchers and practitioners. We then applied these tools to a benchmark of 20 real-world open-source RESTful services and analyzed their performance in terms of code coverage achieved and unique failures triggered. This analysis allowed us to identify strengths, weaknesses, and limitations of the tools considered and of their underlying strategies, as well as implications of our findings for future research in this area. Myeongsoo Kim, Qi Xin 0001, Saurabh Sinha 0003, Alessandro Orso |
ISSTA | 2 |
| 2022 | Studying and Understanding the Tradeoffs Between Generality and Reduction in Software DebloatingabstractExisting approaches for program debloating often use a usage profile, typically provided as a set of inputs, for identifying the features of a program to be preserved. Specifically, given a program and a set of inputs, these techniques produce a reduced program that behaves correctly for these inputs. Focusing only on reduction, however, would typically result in programs that are overfitted to the inputs used for debloating. For this reason, another important factor to consider in the context of debloating is generality, which measures the extent to which a debloated program behaves correctly also for inputs that were not in the initial usage profile. Unfortunately, most evaluations of existing debloating approaches only consider reduction, thus providing partial information on the effectiveness of these approaches. To address this limitation, we perform an empirical evaluation of the reduction and generality of 4 debloating techniques, 3 state-of-the-art ones, and a baseline, on a set of 25 programs and different sets of inputs for these programs. Our results show that these approaches can indeed produce programs that are overfitted to the inputs used and have low generality. Based on these results, we also propose two new augmentation approaches and evaluate their effectiveness. The results of this additional evaluation show that these two approaches can help improve program generality without significantly affecting size reduction. Finally, because different approaches have different strengths and weaknesses, we also provide guidelines to help users choose the most suitable approach based on their specific needs and context. Qi Xin 0001, Qirun Zhang, Alessandro Orso |
ASE | 1 |
| 2020 | Subdomain-Based Generality-Aware DebloatingabstractPrograms are becoming increasingly complex and typically contain an abundance of unneeded features, which can degrade the performance and security of the software. Recently, we have witnessed a surge of debloating techniques that aim to create a reduced version of a program by eliminating the unneeded features therein. To debloat a program, most existing techniques require a usage profile of the program, typically provided as a set of inputs I. Unfortunately, these techniques tend to generate a reduced program that is over-fitted to I and thus fails to behave correctly for other inputs. To address this limitation, we propose DomGad, which has two main advantages over existing debloating approaches. First, it produces a reduced program that is guaranteed to work for subdomains, rather than for specific inputs. Second, it uses stochastic optimization to generate reduced programs that achieve a close-to-optimal tradeoff between reduction and generality (i.e., the extent to which the reduced program is able to correctly handle inputs in its whole domain). To assess the effectiveness of DomGad, we applied our approach to a benchmark of ten Unix utility programs. Our results are promising, as they show that DomGad could produce debloated programs that achieve, on average, 50% code reduction and 95% generality. Our results also show that DomGad performs well when compared with two state-of-the-art debloating approaches. Qi Xin 0001, Myeongsoo Kim, Qirun Zhang, Alessandro Orso |
ASE | 1 |
| 2019 | Automated API-usage update for Android appsabstractMobile apps rely heavily on the application programming interface (API) provided by their underlying operating system (OS). Because OS and API can change frequently, developers must quickly update their apps to ensure that the apps behave as intended with new API and OS versions. To help developers with this tedious, error prone, and time consuming task, we developed a technique that can automatically perform app updates for API changes based on examples of how other developers evolved their apps for the same changes. Given a target app to be updated and information about the changes in the API, our technique performs four main steps. First, it analyzes the target app to identify code affected by API changes. Second, it searches existing code bases for examples of updates to the new version of the API. Third, it analyzes, ranks, and transforms into generic patches the update examples found in the previous step. Finally, it applies the generated patches to the target app in order of ranking, while performing differential testing to validate the update. We implemented our technique and performed an empirical evaluation on 15 real-world apps with promising results. Overall, our technique was able to update 85% of the API changes considered and automatically validate 68% of the updates performed. Mattia Fazzini, Qi Xin 0001, Alessandro Orso |
ISSTA | 2 |
| 2018 | SEEDE: simultaneous execution and editing in a development environmentabstractWe introduce a tool within the Code Bubbles development environment that allows for continuous execution as the programmer edits. The tool, SEEDE, shows both the intermediate and final results of execution in terms of variables, control and data flow, output, and graphics. These results are updated as the user edits. The tool can be used to help the user write new code or to find and fix bugs. The tool is explicitly designed to let the user quickly explore the execution of a method along with all the code it invokes, possibly while writing or modifying the code. The user can start continuous execution either at a breakpoint or for a test case. This paper describes the tool, its implementation, and its user interface. It presents an initial user study of the tool demonstrating its potential utility. Steven P. Reiss, Qi Xin 0001, Jeff Huang 0002 |
ASE | 2 |
| 2018 | Seeking the user interface
Steven P. Reiss, Yun Miao, Qi Xin 0001 |
Autom. Softw. Eng. | 3 |
| 2017 | Identifying test-suite-overfitted patches through test case generationabstractA typical automatic program repair technique that uses a test suite as the correct criterion can produce a patched program that is test-suite-overfitted, or overfitting, which passes the test suite but does not actually repair the bug. In this paper, we propose DiffTGen which identifies a patched program to be overfitting by first generating new test inputs that uncover semantic differences between the original faulty program and the patched program, then testing the patched program based on the semantic differences, and finally generating test cases. Such a test case could be added to the original test suite to make it stronger and could prevent the repair technique from generating a similar overfitting patch again. We evaluated DiffTGen on 89 patches generated by four automatic repair techniques for Java with 79 of them being likely to be overfitting and incorrect. DiffTGen identifies in total 39 (49.4%) overfitting patches and yields the corresponding test cases. We further show that an automatic repair technique, if configured with DiffTGen, could avoid yielding overfitting patches and potentially produce correct ones. Qi Xin 0001, Steven P. Reiss |
ISSTA | 1 |
| 2017 | A demonstration of simultaneous execution and editing in a development environment
Steven P. Reiss, Qi Xin 0001 |
ASE | 2 |
| 2017 | Leveraging syntax-related code for automated program repairabstractWe present our automated program repair technique ssFix which leverages existing code (from a code database) that is syntax-related to the context of a bug to produce patches for its repair. Given a faulty program and a fault-exposing test suite, ssFix does fault localization to identify suspicious statements that are likely to be faulty. For each such statement, ssFix identifies a code chunk (or target chunk) including the statement and its local context. ssFix works on the target chunk to produce patches. To do so, it first performs syntactic code search to find candidate code chunks that are syntax-related, i.e., structurally similar and conceptually related, to the target chunk from a code database (or codebase) consisting of the local faulty program and an external code repository. ssFix assumes the correct fix to be contained in the candidate chunks, and it leverages each candidate chunk to produce patches for the target chunk. To do so, ssFix translates the candidate chunk by unifying the names used in the candidate chunk with those in the target chunk; matches the chunk components (expressions and statements) between the translated candidate chunk and the target chunk; and produces patches for the target chunk based on the syntactic differences that exist between the matched components and in the unmatched components. ssFix finally validates the patched programs generated against the test suite and reports the first one that passes the test suite. We evaluated ssFix on 357 bugs in the Defects4J bug dataset. Our results show that ssFix successfully repaired 20 bugs with valid patches generated and that it outperformed five other repair techniques for Java. Qi Xin 0001, Steven P. Reiss |
ASE | 1 |