Bo Wang 0050

dblp:72/6811-50 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0001-7944-9182ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 17 · 9 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2026 Assessing the effectiveness of recent closed-source large language models in fault localization and automated program repair
Bo Wang 0050, Mingda Chen, Youfang Lin, Jie Zhang 0050
Autom. Softw. Eng.1
2025 CAMUS: Context-Aware Neural Mutation Selection
abstract
Mutation analysis is a fundamental technique in software engineering, playing a vital role in software testing and debugging. It introduces small artificial faults (mutants) into a program to simulate realistic defects. These mutants are systematically injected and then leveraged in downstream tasks such as mutation testing, mutation-based fault localization (MBFL), and mutation-based test case prioritization (MBTCP). However, all of these applications require generating and executing a large number of mutants, leading to substantial computational overhead. This scalability issue significantly limits the applicability of mutation analysis to large-scale software systems. To address this challenge, we propose CAMUS, a context-aware neural mutation selection approach. CAMUS models each mutation and its surrounding context as a graph enriched with AST hierarchy and type information. It then employs an encoder with self-attention mechanisms to produce embeddings, from which the selection probability of each mutant is predicted. We evaluate CAMUS on three downstream tasks, i.e., mutation testing, MBFL, and MBTCP, by comparing it with state-of-the-art mutation selection techniques, including the recent large language model DeepSeek-v3. Experiments are conducted on the Defects4J 2.0 and ConDefects benchmarks. Our results show that CAMUS consistently outperforms existing methods across a wide range of mutation retention rates. In mutation testing, CAMUS achieves over 95% accuracy using only 50% of the mutants. In MBTCP, it reaches near-optimal prioritization performance with as little as 5% of the mutants, demonstrating its strong efficiency and effectiveness.
Mingda Chen, Bo Wang 0050, Youfang Lin, Jie Zhang 0050
APSEC2
2025 RustMap: Towards Project-Scale C-to-Rust Migration via Program Analysis and LLM
Xuemeng Cai, Xiping Huang, Yijun Yu 0001, Chunmiao Li, Bo Wang 0050, Imam Nur Bani Yusuf, Lingxiao Jiang
ICECCS7
2025 Boosting Redundancy-Based Automated Program Repair by Fine-Grained Pattern Mining
abstract
Redundancy-based automated program repair (APR), which generates patches by referencing existing source code, has gained much attention since they are effective in repairing real-world bugs with good interpretability. However, since existing approaches either demand the existence of multiline similar code or randomly reference existing code, they can only repair a small number of bugs with many incorrect patches, hindering their wide application in practice. In this work, we aim to improve the effectiveness of redundancy-based APRs by exploring more effective source code reuse methods for improving the number of correct patches and reducing incorrect patches. Specifically, we have proposed a new repair technique named REPATT, which incorporates a two-level pattern mining process for guiding effective patch generation (i.e., token and expression levels). We have conducted an extensive experiment on the widely-used Defects4J benchmark and compared Repatt with ten state-of-the-art APR approaches. The results show that it complements existing approaches by repairing 9 unique bugs compared with the latest Large Language Model (LLM)-based and deep learning-based methods and 19 unique bugs compared with traditional repair methods when providing the perfect fault localization. In addition, when the perfect fault localization is unknown in real practice, REPATT significantly outperforms the baseline approaches by achieving much higher patch precision, i.e., 83.8 %, although it repairs fewer bugs. Moreover, we further proposed an effective patch ranking strategy for combining the strength of REPATT and the baseline methods. The result shows that it repairs 124 bugs when only considering the Top-1 patches and improves the best-performing repair method by repairing 39 more bugs. The results demonstrate the effectiveness of our approach for practical use.
Jiajun Jiang, Zhirui Ye, Mengjiao Liu, Bo Wang 0050, Hongyu Zhang 0002, Junjie Chen 0003
ICSME6
2025 Deep learning-based software engineering: progress, challenges, and opportunities
abstract
Abstract Researchers have recently achieved significant advances in deep learning techniques, which in turn has substantially advanced other research disciplines, such as natural language processing, image processing, speech recognition, and software engineering. Various deep learning techniques have been successfully employed to facilitate software engineering tasks, including code generation, software refactoring, and fault localization. Many studies have also been presented in top conferences and journals, demonstrating the applications of deep learning techniques in resolving various software engineering tasks. However, although several surveys have provided overall pictures of the application of deep learning techniques in software engineering, they focus more on learning techniques, that is, what kind of deep learning techniques are employed and how deep models are trained or fine-tuned for software engineering tasks. We still lack surveys explaining the advances of subareas in software engineering driven by deep learning techniques, as well as challenges and opportunities in each subarea. To this end, in this study, we present the first task-oriented survey on deep learning-based software engineering. It covers twelve major software engineering subareas significantly impacted by deep learning techniques. Such subareas spread out through the whole lifecycle of software development and maintenance, including requirements engineering, software development, testing, maintenance, and developer collaboration. As we believe that deep learning may provide an opportunity to revolutionize the whole discipline of software engineering, providing one survey covering as many subareas as possible in software engineering can help future research push forward the frontier of deep learning-based software engineering more systematically. For each of the selected subareas, we highlight the major advances achieved by applying deep learning techniques with pointers to the available datasets in such a subarea. We also discuss the challenges and opportunities concerning each of the surveyed software engineering subareas.
Xiangping Chen, Xing Hu 0008, Yuan Huang 0002, He Jiang 0001, Weixing Ji, Yanjie Jiang, Yanyan Jiang 0001, Bo Liu 0094, Hui Liu 0003, Xiaoli Lian, Guozhu Meng, Xin Peng 0001, Hailong Sun 0001, Lin Shi 0006, Bo Wang 0050, Chong Wang 0013, Jifeng Xuan, Xin Xia 0001, Yibiao Yang, Yixin Yang 0006, Li Zhang 0029, Yuming Zhou, Lu Zhang 0023
Sci. China Inf. Sci.16
2025 Fuzzing C++ Compilers via Type-Driven Mutation
abstract
C++ is a system-level programming language for modern software development, which supports multiple programming paradigms, including object-oriented, generic, and functional programming. The intrinsic complexity of these paradigms and their interactions grants C++ powerful expressiveness while posing significant challenges for compilers in correctly implementing its type system. A type system encompasses various aspects such as type inference, type checking, subtyping, type conversions, generics, scoping, and binding. However, systematic testing of the type systems of C++ compilers remains largely underexplored in existing studies. In this work, we present TyMut, the first approach specifically designed to test the C++ type system. TyMut is a mutation-based compiler fuzzer equipped with advanced type-driven mutation operators, carefully crafted to target intricate type-related features such as template generics, type conversions, and inheritance. Beyond differential testing, TyMut introduces enhanced test oracles through a must analysis that partially confirms the validity of generated programs. Specifically, mutation operators are classified into well-formed and not-well-formed : Programs generated by well-formed mutation operators are valid and must be accepted by compilers. Programs generated by not-well-formed operators are validated against a set of well-formedness rules . Any violation indicates the program is invalid and must be rejected. For programs that pass the rules but lack a definitive oracle, TyMut applies differential testing to identify behavioral inconsistencies across compilers. The testing campaign took about 32 hours to generate and test 250584 programs. The must analysis provides definite test oracles for nearly 80% of all generated programs. TyMut uncovered 102 bugs in the recent versions of GCC and Clang, with 56 confirmed as new bugs by compiler developers. Among the confirmed bugs, 26 of them cause compiler crashes, and more than 50% cause miscompilation. Additionally, 7 of them had remained hidden for over 20 years, 22 for over 10 years, and 39 for over 5 years. One long-standing bug discovered by TyMut was later confirmed as the root cause of a real-world issue in TensorFlow. Before submitting this paper, 13 bugs were fixed, most of which were fixed within 60 days. Notably, some unconfirmed bugs have led to in-depth discussions among developers. For instance, one bug led a compiler developer to submit a new issue to the C++ language standard, showing that we uncovered ambiguities in the language specification.
Bo Wang 0050, Chong Chen 0002, Junjie Chen 0003, Youfang Lin, Dan Hao 0001, Jun Sun 0001
Proc. ACM Program. Lang.1
2025 A Systematic Exploration of Mutation-Based Fault Localization Formulae
abstract
ABSTRACT Fault localization (FL) aims to automatically find the location of bugs in a software program. In the family of FL approaches, spectrum‐based fault localization (SBFL) is the most widely used and has been extensively studied, which computes the suspicious scores of a code element to be buggy via test coverage information. Mutation‐based fault localization (MBFL) further analyses the relationship between mutants and test results. Although MBFL involves more information, it also faces the following limitations. (1) To precisely evaluate suspicious sources, SBFL approaches proposed dozens of formulae based on coverage, while only a few of them have been adopted by MBFL. (2) The current MBFL approaches are based on the assumption that the capability in localizing bugs of each mutant is the same, so they assign the same weight to the mutants that change the test outputs and consider the mutant from the same code element equivalent. In this paper, we intend to enrich the MBFL family by the following two approaches. First, we collect 25 typical SBFL formulae and transform them into MBFL versions. Second, we propose new assumptions that the mutants should have different weights in computing suspicious sources and propose BLMu, a novel MBFL approach that treats mutants differently. We also propose novel metrics for MBFL by considering the high cost of mutation analysis. We perform large‐scale experiments by evaluating all the MBFL approaches against 395 real‐world Java bugs of the Defects4J benchmark. Our evaluation results reveal that BLMu demonstrates a substantial improvement over both MUSE and Metallaxis, the most popular MBFL approaches, at both the statement level and the method level. Specifically, in terms of Top‐1 , BLMu improves by 106% at the statement level and 69% at the method level. When compared with other types of FL approaches, MBFL outperforms typical SBFL approaches while still far behind the state‐of‐the‐art learning‐based FL approaches.
Bo Wang 0050, Jinkang Wei, Mingda Chen, Chong Chen 0002, Youfang Lin, Jie Zhang 0050
Softw. Test. Verification Reliab.1
2025 A Comprehensive Study of OOP-Related Bugs in C++ Compilers
abstract
Modern C++, a programming language characterized by its extensive use of object-oriented programming (OOP) features, is widely used for system programming. However, C++ compilers often struggle to correctly handle these sophisticated OOP features, resulting in numerous high-profile compiler bugs that can lead to crashes or miscompilation. Despite the significance of OOP-related bugs, existing studies largely overlook OOP features, hindering their ability to discover such bugs. To assist both compiler fuzzer designers and compiler developers, we conduct a comprehensive study of the compiler bugs caused by incorrectly handling C++ OOP-related features. First, we systematically extract 788 OOP-related C++ compiler bugs from GCC and LLVM. Second, derived from the core concepts of OOP and C++, we manually identified a two-level taxonomy of the OOP-related features leading to compiler bugs, which consists of 6 primary categories (e.g.,Abstraction & Encapsulation,Inheritance, andRuntime Polymorphism), along with 17 secondary categories (e.g.,Constructors & DestructorsandMultiple Inheritance). Third, we systematically analyze the root causes, symptoms, fixes, options, and C++ standard versions of these bugs. Our analysis yields 13 key findings, highlighting that features related to the construction and destruction of objects lead to the highest number of bugs, crashes are the most frequent symptom, and while the average time from bug introduction to discovery is 1856 days, fixing the bug once discovered takes only 174 days on average. Additionally, more than half of the bugs can be triggered without any compiler options. These findings offer valuable insights not only for developing new compiler testing approaches but also for improving language design and compiler engineering. Inspired by these findings, we developed a proof-of-concept compiler fuzzer OOPFuzz, specifically targeting OOP-related bugs in C++ compilers. We applied it against the newest release versions of GCC and LLVM. In about 3 hours, it detected 9 bugs, of which 3 have been confirmed by the developers, including a bug of LLVM that had persisted for 13 years. The results indicate our taxonomy and analysis provide valuable insights for future research in compiler testing.
Bo Wang 0050, Chong Chen 0002, Junjie Chen 0003, Youfang Lin, Guoliang Dong, Jun Sun 0001
IEEE Trans. Software Eng.1
2024 On the Evaluation of Large Language Models in Unit Test Generation
abstract
Unit testing is an essential activity in software development for verifying the correctness of software components. However, manually writing unit tests is challenging and time-consuming. The emergence of Large Language Models (LLMs) offers a new direction for automating unit test generation. Existing research primarily focuses on closed-source LLMs (e.g., ChatGPT and CodeX) with fixed prompting strategies, leaving the capabilities of advanced open-source LLMs with various prompting settings unexplored. Particularly, open-source LLMs offer advantages in data privacy protection and have demonstrated superior performance in some tasks. Moreover, effective prompting is crucial for maximizing LLMs' capabilities. In this paper, we conduct the first empirical study to fill this gap, based on 17 Java projects, five widely-used open-source LLMs with different structures and parameter sizes, and comprehensive evaluation metrics. Our findings highlight the significant influence of various prompt factors, show the performance of open-source LLMs compared to the commercial GPT-4 and the traditional Evosuite, and identify limitations in LLM-based unit test generation. We then derive a series of implications from our study to guide future research and practical use of LLM-based unit test generation.
Lin Yang 0030, Shutao Gao, Weijing Wang, Bo Wang 0050, Qihao Zhu, Xiao Chu, Guangtai Liang, Qianxiang Wang, Junjie Chen 0003
ASE5
2024 Enhanced evolutionary automated program repair by finer-granularity ingredients and better search algorithms
abstract
Summary Bug repair is time consuming and tedious, which hampers software maintenance. To alleviate the burden, automated program repair (APR) is proposed and has been fruitful in the last decade. Evolutionary repair is the seminal work of this field and proliferated a family of approaches. The performance of evolutionary repair approaches is affected by two main factors: (1) search space, which defines all possible patches, and (2) search algorithms, which navigate the space. Although recent approaches have achieved remarkable progress, the main challenges of the two factors still remain. On one hand, the different kinds of search space are very coarse for containing correct patches. On the other hand, the search process guided by genetic algorithms is inefficient in finding the correct patches in an appropriate time budget. In this paper, we propose MicroRepair, a new evolutionary repair approach to address the two challenges. Rather than finding statement‐level patches like existing genetic repair approaches, MicroRepair enlarges the search space by breaking the statements into finer‐granularity ingredients that consist of AST leaves. As the search space grows exponentially, the former search algorithms may become inefficient in navigating the larger space. We utilize the best multiobjective search algorithm selected from our empirical comparison of a set of search algorithms. At last, we find redundancies search in the existing genetic process, and we further design a history‐aware search strategy to accelerate the process. We evaluated MicroRepair on 224 bugs of real‐world from the benchmark Defects4J and compared it with several state‐of‐the‐art repair approaches. The evaluation results show that MicroRepair correctly repaired 26 bugs with a precision of 62%, which significantly outperforms the state‐of‐the‐art evolutionary APR approaches in terms of precision. Moreover, the history‐aware search boosts the repair execution speed by 4% on average.
Bo Wang 0050, Guizhuang Liu, Youfang Lin, Shuang Ren, Dalin Zhang 0003
J. Softw. Evol. Process.1
2024 Accelerating Patch Validation for Program Repair With Interception-Based Execution Scheduling
abstract
Long patch validation time is a limiting factor for automated program repair (APR). Though the duality between patch validation and mutation testing is recognized, so far there exists no study of systematically adapting mutation testing techniques to general-purpose patch validation. To address this gap, we investigate existing mutation testing techniques and identify five classes of acceleration techniques that are suitable for general-purpose patch validation. Among them, mutant schemata and mutant deduplication have not been adapted to general-purpose patch validation due to the arbitrary changes that third-party APR approaches may introduce. This presents two problems for adaption: 1) the difficulty of implementing the static equivalence analysis required by the state-of-the-art mutant deduplication approach; 2) the difficulty of capturing the changes of patches to the system state at runtime. To overcome these problems, we propose two novel approaches: 1) execution scheduling, which detects the equivalence between patches online, avoiding the static equivalence analysis and its imprecision; 2) interception-based instrumentation, which intercepts the changes of patches to the system state, avoiding a full interpreter and its overhead. Based on the contributions above, we implement ExpressAPR, a general-purpose patch validator for Java that integrates all recognized classes of techniques suitable for patch validation. Our large-scale evaluation with four APR approaches shows that ExpressAPR accelerates patch validation by 137.1x over plain validation or 8.8x over the state-of-the-art approach, making patch validation no longer the time bottleneck of APR. Patch validation time for a single bug can be reduced to within a few minutes on mainstream CPUs.
Yuan-an Xiao, Chenyang Yang 0002, Bo Wang 0050, Yingfei Xiong 0001
IEEE Trans. Software Eng.3
2023 ExpressAPR: Efficient Patch Validation for Java Automated Program Repair Systems
abstract
Automated program repair (APR) approaches suffer from long patch validation time, which limits their practical application and receives relatively low attention. The patch validation process repeatedly executes tests to filter patches, and has been recognized as the dual of mutation analysis. We systematically investigate existing mutation testing techniques and recognize five families of acceleration techniques that are suitable for patch validation, two of which are never adapted to a general-purpose patch validator. We implement and demonstrate ExpressAPR, the first framework that combines five families of acceleration techniques for patch validation as the complete set. In our evaluation on 30 random Defects4J bugs and four APR systems, ExpressAPR accelerates patch validation for two order-of-magnitudes over plain validation or one order-of-magnitude over the state-of-the-art approach, benefiting APR researchers and users with a much shorter patch validation time. Demo video available at https://youtu.be/7AB-4VvBuuM Tool repo (source code + Docker image + evaluation dataset) available at https://github.com/ExpressAPR/ExpressAPR
Yuan-an Xiao, Chenyang Yang 0002, Bo Wang 0050, Yingfei Xiong 0001
ASE3
2022 Enhanced Evolutionary Automated Program Repair by Finer-Granularity Ingredients and Better Search Algorithms
abstract
Bug repair is time-consuming and tedious, which hampers software maintenance. To alleviate the burden, automated program repair (APR) is proposed and has been fruitful in the last decade. Evolutionary repair is the seminal work of this field and proliferated a family of approaches. The performance of evolutionary repair approaches is affected by two main factors: (1) search space, which defines all possible patches, and (2) search algorithms, which navigates the space. Although recent approaches have achieved remarkable progress, the main challenges of the two factors still remain. On one hand, the different kinds of search space are very coarse for containing correct patches. On the other hand, the search process guided by genetic algorithms is inefficient to find the correct patches in an appropriate time budget.
Bo Wang 0050, Guizhuang Liu, Youfang Lin, Shuang Ren, Dalin Zhang 0003
Internetware1
2022 L2S: A Framework for Synthesizing the Most Probable Program under a Specification
abstract
In many scenarios, we need to find the most likely program that meets a specification under a local context, where the local context can be an incomplete program, a partial specification, natural language description, and so on. We call such a problem program estimation . In this article, we propose a framework, LingLong Synthesis Framework (L2S) , to address this problem. Compared with existing work, our work is novel in the following aspects. (1) We propose a theory of expansion rules to describe how to decompose a program into choices. (2) We propose an approach based on abstract interpretation to efficiently prune off the program sub-space that does not satisfy the specification. (3) We prove that the probability of a program is the product of the probabilities of choosing expansion rules, regardless of the choosing order. (4) We reduce the program estimation problem to a pathfinding problem, enabling existing pathfinding algorithms to solve this problem. L2S has been applied to program generation and program repair. In this article, we report our instantiation of this framework for synthesizing conditional expressions (L2S-Cond) and repairing conditional statements (L2S-Hanabi). The experiments on L2S-Cond show that each option enabled by L2S, including the expansion rules, the pruning technique, and the use of different pathfinding algorithms, plays a major role in the performance of the approach. The default configuration of L2S-Cond correctly predicts nearly 60% of the conditional expressions in the top 5 candidates. Moreover, we evaluate L2S-Hanabi on 272 bugs from two real-world Java defects benchmarks, namely Defects4J and Bugs.jar. L2S-Hanabi correctly fixes 32 bugs with a high precision of 84%. In terms of repairing conditional statement bugs, L2S-Hanabi significantly outperforms all existing approaches in both precision and recall.
Yingfei Xiong 0001, Bo Wang 0050
ACM Trans. Softw. Eng. Methodol.2
2021 Faster Mutation Analysis with Fewer Processes and Smaller Overheads
abstract
Mutation analysis is a powerful dynamic approach that has many applications, such as measuring the quality of test suites or automatically locating faults. However, the inherent low scalability hampers its practical use. To accelerate mutation analysis, researchers propose approaches to reduce redundant executions. A family of fork-based approaches tries to share identical executions among mutants. Fork-based approaches carry all mutants in one process and decide whether to fork new child processes when reaching a mutated statement. The mutants carried by the parent process are split into groups and distributed to different processes to finish the remaining executions. However, existing fork-based approaches have two limitations: (1) the limited analysis scope on a single statement to compare and cluster mutants prevents their systems from detecting more equivalent mutants, and (2) the interpretation of the mutants and the runtime equivalence analysis introduce significant overhead.In this paper, we present a novel fork-based mutation analysis approach WinMut, which (1) groups mutants in a scope of mutated statements and, (2) removes redundant computations inside interpreters. WinMut not only reduces the number of invoked processes but also has a lower cost for executing a single process. Our experiments show that our approach can further accelerate mutation analysis with an average speedup of 5.57x on top of the state-of-the-art fork-based approach, AccMut.
Bo Wang 0050, Sirui Lu, Yingfei Xiong 0001, Feng Liu 0061
ASE1
2021 Beyond Tests: Program Vulnerability Repair via Crash Constraint Extraction
abstract
Automated program repair is an emerging technology that seeks to automatically rectify program errors and vulnerabilities. Repair techniques are driven by a correctness criterion that is often in the form of a test suite. Such test-based repair may produce overfitting patches, where the patches produced fail on tests outside the test suite driving the repair. In this work, we present a repair method that fixes program vulnerabilities without the need for a voluminous test suite. Given a vulnerability as evidenced by an exploit, the technique extracts a constraint representing the vulnerability with the help of sanitizers. The extracted constraint serves as a proof obligation that our synthesized patch should satisfy. The proof obligation is met by propagating the extracted constraint to locations that are deemed to be “suitable” fix locations. An implementation of our approach (E xtract F ix ) on top of the KLEE symbolic execution engine shows its efficacy in fixing a wide range of vulnerabilities taken from the ManyBugs benchmark, real-world CVEs and Google’s OSS-Fuzz framework. We believe that our work presents a way forward for the overfitting problem in program repair by generalizing observable hazards/vulnerabilities (as constraint) from a single failing test or exploit.
Xiang Gao 0012, Bo Wang 0050, Gregory J. Duck, Ruyi Ji, Yingfei Xiong 0001, Abhik Roychoudhury
ACM Trans. Softw. Eng. Methodol.2
2017 Faster mutation analysis via equivalence modulo states
abstract
Mutation analysis has many applications, such as asserting the quality of test suites and localizing faults. One important bottleneck of mutation analysis is scalability. The latest work explores the possibility of reducing the redundant execution via split-stream execution. However, split-stream execution is only able to remove redundant execution before the first mutated statement.
Bo Wang 0050, Yingfei Xiong 0001, Yangqingwei Shi, Lu Zhang 0023, Dan Hao 0001
ISSTA1
2016 Dynamic analysis of shared execution in software product line testing
abstract
Software product line (SPL), a family-based software development process, has proven to be a more effective technology than single software systems. Testing SPL products individually is redundant for product lines testing. Meanwhile, the complexity of systematically testing SPL programs is combinatorial, which limits the scalability of testing SPL.
Bo Wang 0050
SPLC1