Jiangchang Wu

dblp:348/0991 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-9932-7315ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Boosting Compiler Fault Localization: Getting the Best of Both Worlds by Fusing Dynamic and Historical Data
abstract
Compilers are prone to bugs that can have severe consequences for downstream applications. Accurately identifying and localizing compiler faults poses unique challenges due to the inherent complexity and large scale of modern compiler infrastructures. Existing studies have proposed various techniques to construct passing and failing executions by generating witness test programs from bug-inducing test cases or by producing adversarial compilation configurations for the same test program. These executions are then leveraged to apply spectrum-based fault localization (SBFL) techniques for isolating compiler faults, yielding promising results. Recently, Yang et al. revisited SBFL-based techniques and showed that a simple yet widely adopted debugging practice—treating files modified in bug-inducing commits (BICs) as potential fault candidates—can surprisingly outperform SBFL-based techniques on the most critical localization metrics. Moreover, they further demonstrated that BIC-based and SBFL-based techniques are highly complementary, as they tend to localize different subsets of compiler faults. Consequently, effectively integrating these two sources of information to improve compiler fault localization remains an open and largely unexplored challenge. To address this problem, we propose DUALTRACK, a hybrid approach that integrates dynamic execution information from SBFL with historical information derived from BICs. DUALTRACKemploys a two-layer framework that first prioritizes files modified in bug-inducing commits and then refines their rankings using suspiciousness scores computed by SBFL formulas. An evaluation on 120 real-world compiler bugs from GCC and LLVM shows that DUALTRACK successfully identifies 52% of faulty files at the Top-1 rank, demonstrating a substantial improvement over existing state-of-the-art compiler fault localization techniques.
Qingyang Li 0006, Yibiao Yang, Jiangchang Wu, Qingkai Shi, Yuming Zhou, Baowen Xu
IEEE Trans. Software Eng.4
2025 Debugger Toolchain Validation via Cross-Level Debugging
abstract
Ensuring the correctness of debugger toolchains is of paramount importance, as they play a vital role in understanding and resolving programming errors during software development. Bugs hidden within these toolchains can significantly mislead developers. Unfortunately, comprehensive testing of debugger toolchains is lacking due to the absence of effective test oracles. Existing studies on debugger toolchain validation have primarily focused on validating the debug information within optimized executables by comparing the traces between debugging optimized and unoptimized executables (i.e., different executables) in the debugger, under the assumption that the traces obtained from debugging unoptimized executables serve as a reliable oracle. However, these techniques suffer from inherent limitations, as compiler optimizations can drastically alter source code elements, variable representations, and instruction order, rendering the traces obtained from debugging different executables incomparable and failing to uncover bugs in debugger toolchains when debugging unoptimized executables. To address these limitations, we propose a novel concept called Cross-Level Debugging (CLD) for validating the debugger toolchain. CLD compares the traces obtained from debugging the same executable using source-level and instruction-level strategies within the same debugger. The core insight of CLD is that the execution traces obtained from different debugging levels for the same executable should adhere to specific relationships, regardless of whether the executable is generated with or without optimization. We formulate three key relations in CLD: reachability preservation of program locations, order preservation for reachable program locations, and value consistency at program locations, which apply to traces at different debugging levels. We implement Devil, a practical framework that employs these relations for debugger toolchain validation. We evaluate the effectiveness of Devil using two widely used production debugger toolchains, GDB and LLDB. Ultimately, Devil successfully identified 27 new bug reports, of which 18 have been confirmed and 12 have been fixed by developers.
Yibiao Yang, Jiangchang Wu, Qingyang Li 0006, Yuming Zhou
ASPLOS (1)3
2025 Clozemaster: Fuzzing Rust Compiler by Harnessing Llms for Infilling Masked Real Programs
abstract
Ensuring the reliability of the Rust compiler is of paramount importance, given increasing adoption of Rust for critical systems development, due to its emphasis on memory and thread safety. However, generating valid test programs for the Rust compiler poses significant challenges, given Rust's complex syntax and strict requirements. With the growing popularity of large language models (LLMs), much research in software testing has explored using LLMs to generate test cases. Still, directly using LLMs to generate Rust programs often results in a large number of invalid test cases. Existing studies have indicated that test cases triggering historical compiler bugs can assist in software testing. Our investigation into Rust compiler bug issues supports this observation. Inspired by existing work and our empirical research, we introduce a bracket-based masking and filling strategy called clozeMask. The clozeMask strategy involves extracting test code from historical issue reports, identifying and masking code snippets with specific structures, and using an LLM to fill in the masked portions for synthesizing new test programs. This approach harnesses the generative capabilities of LLMs while retaining the ability to trigger Rust compiler bugs. It enables comprehensive testing of the compiler's behavior, particularly exploring edge cases. We implemented our approach as a prototype ClozeMaster. ClozeMaster has identified 27 confirmed bugs for rustc and mrustc, of which 10 have been fixed by developers. Furthermore, our experimental results indicate that ClozeMaster outperforms existing fuzzers in terms of code coverage and effectiveness.
Hongyan Gao, Yibiao Yang, Jiangchang Wu, Yuming Zhou, Baowen Xu
ICSE4
2025 Unveiling Compiler Faults via Attribute-Guided Compilation Space Exploration
Jiangchang Wu, Yibiao Yang, Yuming Zhou
USENIX ATC1
2025 Validating SMT Rewriters via Rewrite Space Exploration Supported by Generative Equality Saturation
abstract
Satisfiability Modulo Theories (SMT) solvers are widely used for program analysis and other applications that require automated reasoning. Rewrite systems, as crucial integral components of SMT solvers, are responsible for simplifying and transforming formulas to optimize the solving process. The effectiveness of an SMT solver heavily depends on the robustness of its rewrite system, making its validation crucial. Despite ongoing advancements in SMT solver testing, rewrite system validation remains largely unexplored. Our empirical analysis reveals that developers invest significant effort in ensuring the correctness and reliability of rewrite systems. However, existing testing techniques do not adequately address this aspect. In this paper, we introduce Aries , a novel technique designed to validate SMT solver rewrite systems. First, Aries employs mimetic mutation , a targeted strategy that actively reshapes input formulas to provoke and diversify rewrite opportunities. By aligning mutated terms with known rewrite patterns, Aries can conduct a thorough exploration of the rewrite space in the following phase. Second, Aries utilizes deductive rewriting , leveraging generative equality saturation to effectively explore rewrite space and produce semantically equivalent mutants for the purpose of validation. We implemented Aries as a practical validation tool and evaluated it on leading SMT solvers, including Z3 and cvc5. Our experiments demonstrate that Aries effectively identifies bugs, with 27 new issues detected, of which 22 have been confirmed or fixed by developers. Most of these issues involve the rewrite systems, highlighting Aries ’s strength in exploring the rewrite space.
Yibiao Yang, Jiangchang Wu, Yuming Zhou
Proc. ACM Program. Lang.3
2025 Isolating Compiler Faults Through Differentiated Compilation Configurations
abstract
Compilation optimization bugs are prevalent and can significantly affect the correctness of software products, posing serious challenges to software development. Identifying and localizing these bugs are critical tasks for compiler developers. However, the intricate nature and extensive scale of modern compilers make it difficult to pinpointing the root causes of such bugs. Previous research has introduced innovative techniques that generatewitness test programs–tests that pass–by mutating bug-triggering test cases, highlighting the importance of this problem and demonstrating the effectiveness of such approaches. Nevertheless, existing techniques based on witness test programs generation suffer from inherent limitations. Specifically, they do not guarantee the successful creation of witness test programs via mutation and are often time-consuming, typically requiring extensive iterations to produce a valid witness test program. In this study, we present Odfl, a simple yet effective approach for automatically isolating compiler optimization faults by introducing the concept ofdifferentiated compilation configurations. The core insight behind Odfl is that modifying compilation settings such as disabling fine-grained compilation flags in GCC or reducing the number of fine-grained compilation passes in LLVM, can suppress the manifestation of compiler bugs triggered by the same test program. Through adjusting these settings, Odfl creates differentiated compilation configuration that produce multiple compiler executions with distinct pass/-fail outcomes. We utilize these differentiated configurations to collect both passing and failing compiler coverage, and then applySpectrum-Based Fault Localization (SBFL)techniques to rank compiler source files based on their suspiciousness. Our evaluation of 60 GCC and 50 LLVM compiler bugs demonstrates that Odfl substantially outperforms state-of-the-art compiler fault localization techniques in terms of both effectiveness and efficiency. Notably, Odfl achieves over 90% improvement in accurately ranking the top-1 faulty source files compared to three existing techniques–DiWi, RecBi, and LLM4CBI–and reduces fault localization time by more than 99% on average.
Yibiao Yang, Qingyang Li 0006, Jing Yang 0051, Jiangchang Wu, Yuming Zhou
IEEE Trans. Software Eng.5
2023 Boosting Compiler Testing via Eliminating Test Programs with Long-Execution-Time
abstract
Compiler testing is crucially important as compiler is the fundamental infrastructure in software development. One common compiler testing practice leverages a random program generator such as Csmith to generate a huge number of test programs to stress-test compilers. Each of the test programs will be compiled to different executables at different optimization levels and then their outputs will be compared against each other to differentially test compilers. However, the execution time of different test programs varies a lot. Therefore, in practice, developers often set a time limit, such as 60 or 300 seconds, to control the execution of different executables. If the execution exceeds the time limit, it will be terminated. Nevertheless, it is still unclear which time limit is more suitable for compiler testing in this context. We therefore perform the first empirical analysis to investigate how different time limits afffect the efficiency of compiler testing. We found that a time limit of 0.1 seconds can achieve the maximum benefits for compiler testing with the randomly generated test programs. At the same time, we found that 12% test programs requires more than 300 seconds for the execution and these test programs with long-execution-time (LET) consume more than 90% of the entire testing resources which makes compiler testing not cost-effectiveness. We thus propose a framework named ELECT to automatically identify and exclude LET test programs for boosting compiler testing. Our extensive experiments on two popular compilers GCC and LLVM have shown that ELECT can significantly improve the cost-effectiveness of compiler testing as it can respectively detect about 19% and 10% more bugs than the two baseline approaches under the same testing time budget. Besides, ELECT can respectively save about 12% and 38% time than the two other approaches for detecting the same number of bugs.
Jiangchang Wu, Yibiao Yang, Yuming Zhou
SANER1