Hengchen Yuan

dblp:351/6812 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0005-3695-103XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Detecting Flaky Tests by Controlling Nondeterministic API Behavior
abstract
Regression testing is an essential part of software development to ensure high-quality software; but the presence of flaky tests makes the testing outcomes unreliable. It is essential to proactively detect flaky tests, so developers are aware of them early on and can react appropriately in case they fail. While there has been prior work in detecting flaky tests, they either are developed to focus on specific flaky test types or can be imprecise in their analyses, such as by relying on AI models for prediction. We present ChaosAPI, a framework to support detecting a variety of different types of flaky tests. Our insight is that the most effective approach to detecting flaky tests is to target the nondeterministic components that lead tests to have the flaky behavior in the first place. In particular, we target specific APIs within the Java Standard Library that all Java code relies on and that are known to exhibit nondeterministic behavior, such as those related to system time, concurrency, and environmental factors. During test execution, ChaosAPI modifies the behavior of these API calls, perturbing inputs and return values of these APIs in a systematic manner while still remaining compliant with the API specification. We can detect flaky tests by observing whether tests that previously passed would now fail when run through ChaosAPI. Our evaluation on a prior dataset of known flaky tests, as well as running on test suites of other popular open-source projects, demonstrates that ChaosAPI not only detects more flaky tests than simple rerunning across a wide range of projects but also detects them more efficiently, making ChaosAPI a practical addition to the toolbox of flaky test detection techniques.
Hengchen Yuan, Jiefang Lin, August Shi
Proc. ACM Program. Lang.1
2024 Test Scheduling Across Heterogeneous Machines While Balancing Running Time, Price, and Flakiness
abstract
Scheduling tests to run in parallel across different machines is an effective way to reduce overall test running time. Prior work has focused on scheduling tests across homogeneous machines, namely, machines all of the same configuration. However, using all the same configuration may not be the most cost-effective way to reduce test running time. We propose scheduling tests across machines with different configurations, namely heterogeneous machines. Doing so allows us to balance various factors, e.g., price, as tests may have similar running times on different machine configurations but result in drastically different monetary prices. Furthermore, there can be flaky tests that fail more often on different machine configurations, so scheduling them across heterogeneous machines gives better control over their flaky-failure rates. Our approach, GASearch, leverages genetic algorithms and a fitness function to balance running time and price to efficiently generate a heterogeneous machine configuration on which to run tests. We also model flaky-failure rate of tests on different machines within the fitness function as a factor of running time, where a failing flaky test would be rerun until it passes (or to a maximum number of runs) to confirm if it is a flaky failure, so we can balance all factors at once. We evaluate our approach on test suites from 24 modules in open-source Maven projects. Compared against baselines that schedule across homogeneous machines, we find that scheduling across heterogeneous ones can achieve a lower running time and price.
Hengchen Yuan, Jiefang Lin, Wing Lam, August Shi
ICSME1
2023 Evaluating and Improving Hybrid Fuzzing
abstract
To date, various hybrid fuzzers have been proposed for maximal program vulnerability exposure by integrating the power of fuzzing strategies and concolic executors. While the existing hybrid fuzzers have shown their superiority over conventional coverage-guided fuzzers, they seldom follow equivalent evaluation setups, e.g., benchmarks and seed corpora. Thus, there is a pressing need for a comprehensive study on the existing hybrid fuzzers to provide implications and guidance for future research in this area. To this end, in this paper, we conduct the first extensive study on state-of-the-art hybrid fuzzers. Surprisingly, our study shows that the performance of existing hybrid fuzzers may not well generalize to other experimental settings. Meanwhile, their performance advantages over conventional coverage-guided fuzzers are overall limited. In addition, instead of simply updating the fuzzing strategies or concolic executors, updating their coordination modes potentially poses crucial performance impact of hybrid fuzzers. Accordingly, we propose CoFuzz to improve the effectiveness of hybrid fuzzers by upgrading their coordination modes. Specifically, based on the baseline hybrid fuzzer QSYM, CoFuzz adopts edge-oriented scheduling to schedule edges for applying concolic execution via an online linear regression model with Stochastic Gradient Descent. It also adopts sampling-augmenting synchronization to derive seeds for applying fuzzing strategies via the interval path abstraction and John walk as well as incrementally updating the model. Our evaluation results indicate that CoFuzz can significantly increase the edge coverage (e.g., 16.31% higher than the best existing hybrid fuzzer in our study) and expose around 2X more unique crashes than all studied hybrid fuzzers. Moreover, CoFuzz successfully detects 37 previously unknown bugs where 30 are confirmed with 8 new CVEs and 20 are fixed.
Hengchen Yuan, Mingyuan Wu, Lingming Zhang 0001, Yuqun Zhang
ICSE2
2023 Third-Party Library Dependency for Large-Scale SCA in the C/C++ Ecosystem: How Far Are We?
abstract
Existing software composition analysis (SCA) techniques for the C/C++ ecosystem tend to identify the reused components through feature matching between target software project and collected third-party libraries (TPLs). However, feature duplication caused by internal code clone can cause inaccurate SCA results. To mitigate this issue, Centris, a state-of-the-art SCA technique for the C/C++ ecosystem, was proposed to adopt function-level code clone detection to derive the TPL dependencies for eliminating the redundant features before performing SCA tasks. Although Centris has been shown effective in the original paper, the accuracy of the derived TPL dependencies is not evaluated. Additionally, the dataset to evaluate the impact of TPL dependency on SCA is limited. To further investigate the efficacy and limitations of Centris, we first construct two large-scale ground-truth datasets for evaluating the accuracy of deriving TPL dependency and SCA results respectively. Then we extensively evaluate Centris where the evaluation results suggest that the accuracy of TPL dependencies derived by Centris may not well generalize to our evaluation dataset. We further infer the key factors that degrade the performance can be the inaccurate function birth time and the threshold-based recall. In addition, the impact on SCA from the TPL dependencies derived by Centris can be somewhat limited. Inspired by our findings, we propose TPLite with function-level origin TPL detection and graph-based dependency recall to enhance the accuracy of TPL reuse detection in the C/C++ ecosystem. Our evaluation results indicate that TPLite effectively increases the precision from 35.71% to 88.33% and the recall from 49.44% to 62.65% of deriving TPL dependencies compared with Centris. Moreover, TPLite increases the precision from 21.08% to 75.90% and the recall from 57.62% to 64.17% compared with the SOTA academic SCA tool B2SFinder and even outperforms the well-adopted commercial SCA tool BDBA, i.e., increasing the precision from 72.46% to 75.90% and the recall from 58.55% to 64.17%.
Hengchen Yuan, Qiyi Tang 0003, Sen Nie, Shi Wu, Yuqun Zhang
ISSTA2