EDBT 2026 Demo / reviewers in the wild / expert
Huan Xie 0002
dblp:13/9949-2
· DBLP profile ↗
24ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0002-5265-0051ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 6 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeepFlaky: Deep hybrid representation learning for flaky test prediction
Yan Lei 0005, Huan Xie 0002 |
Inf. Softw. Technol. | 5 |
| 2026 | Test-free fault localization using large language models and a Transformer-Mamba hybrid architecture
Yujian Huang, Yan Lei 0005, Huan Xie 0002 |
Inf. Softw. Technol. | 3 |
| 2026 | HiFlaky: Hierarchy-aware flakiness classification
Yan Lei 0005, Huan Xie 0002, Maojin Li |
J. Syst. Softw. | 4 |
| 2026 | Enhanced Feature Representation via Hybrid Feature Fusion for Coincidental Correctness DetectionabstractCoincidental Correctness (CC) arises when a test case executes faulty entity in a program without causing a failure. This phenomenon injects noise into coverage information, as CC tests weaken the connection between faulty entities and test failures. Since many fault localization (FL) approaches relies on analyzing test execution traces to locate faulty entities, the compromised reliability of test results directly undermines FL accuracy. Furthermore, the detrimental effects of CC extend beyond fault localization to subsequent software maintenance tasks like automatic program repair. Therefore, identifying and mitigating CC tests becomes critical not only for enhancing FL but also for ensuring robust software quality assurance. Thus, we propose FusionCC: an approach that applies multiscale coverage features and handcrafted features to fuse complementary feature representations for CC test case detection. Specifically, FusionCC first refines original coverage data by filtering out noisy irrelevant elements, then extracts multiscale features from the refined matrix, and finally fuses the coverage and handcrafted features to generate highly informative feature representations for CC detection. FusionCC realizes a comprehensive fusion of complementary features across different scales and from diverse sources, which significantly enhances the accuracy of CC detection. To evaluate the effectiveness of FusionCC, we conduct large-scale experiments on 277 faulty versions of six representative benchmarks. The experimental results show that FusionCC significantly improves CC detection (e.g., average improvements of 50.93% precision and 82.03% in$F_{1}$value compared to state-of-the-art CC detection approaches) and fault localization effectiveness (e.g., 10.33, 19.33, 25.67 average faults can be found in terms of Top-1, Top-3, Top-5 metrics at relabel strategy compared with state-of-the-art FL approaches). Tao Zhang 0195, Yan Lei 0005, Huan Xie 0002 |
IEEE Trans. Reliab. | 3 |
| 2026 | Unleashing the Potential of Coverage Representation in Deep Learning-Based Fault Localization
Yan Lei 0005, Huan Xie 0002 |
IEEE Trans. Software Eng. | 3 |
| 2026 | Mutants Will Tell: Statistical Mutation-Based Multiple Fault Localization for Deep Learning ProgramsabstractAs deep learning (DL) systems are increasingly deployed in safety-critical domains, e.g., intelligent planning and autonomous driving, localizing faults that occur in such systems becomes indispensable. Inevitably, DL systems also suffer from faults like traditional software. Although single fault localization for DL programs has been studied, the multiple-fault localization for DL programs remains underexplored. We notice that mutation analysis is a powerful technique for locating multiple faults since it can simulate the faulty behaviors of a DL program by generating multiple mutants simultaneously. Thus, we propose MuMuFL: StatisticalMutation-basedMultipleFaultLocalization approach to locate the multiple faulty statements residing in a faulty DL program. The insight of MuMuFL is that the different behaviors of mutants provide valuable information for pinpointing the faulty statements of a DL fault. MuMuFL defines and leverages DL mutation operators on a DL program to simulate the faulty DL behavior. Then, MuMuFL evaluates the difference in the accuracy between the original DL model and the mutated DL model to quantify the suspiciousness of each statement being faulty. Finally, the large-scale experiments show that MuMuFL effectively localizes DL faults, e.g., localizing 36% of multiple-fault DL programs, whereas the best-performing baseline can only localize 14% of them. Huan Xie 0002, Zhengxiong Deng, Yan Lei 0005, Maojin Li, Meng Yan 0001, David Lo 0001 |
IEEE Trans. Software Eng. | 1 |
| 2025 | AgentTCP: A Collaborative Multi-Agent Framework for Change-Aware Test Case PrioritizationabstractTest Case Prioritization (TCP) is a critical technique for improving efficiency in CI/CD pipelines. While applying Large Language Models (LLMs) to this task is a promising direction due to their advanced code comprehension, naively using them as monolithic tools fails to address key engineering challenges of scale, tool-integration, and structured reasoning. To address these shortcomings, we propose AgentTCP, a novel collaborative multi-agent framework for change-aware test case prioritization. Our framework decomposes the TCP task into a structured, three-stage workflow managed by specialized, LLM-driven agents: 1) a Code Change Analyst assesses the intent and risk of new commits; 2) a Test Coverage Strategist correlates changes with test cases by interacting with coverage data via tool-integration; 3) a Risk-aware Prioritizer synthesizes all information to generate a final, ranked list with reasoning. By delegating distinct responsibilities, AgentTCP mitigates the context and reasoning issues of monolithic models and produces verifiable intermediate results, enhancing overall trustworthiness. Experimental results on the widely used Defects4J benchmark demonstrate that AgentTCP surpasses the monolithic-LLM baseline by 11.75 points in terms of the APFD metric, highlighting its superior effectiveness in prioritizing fault-revealing test cases. Huan Xie 0002, Yan Lei 0005 |
APSEC | 4 |
| 2025 | Sifting Truth from Coincidences: A Two-Stage Positive and Unlabeled Learning Model for Coincidental Correctness DetectionabstractFault localization (FL) can identify the fault's location by analyzing the execution information from test cases in the program. This execution information serves as the foundation for FL to infer latent causal relationships between fault entities and failed results. However, this execution information contains coincidental correctness (CC), which reduces the accuracy of FL. CC arises when a test case executes faulty program entities but still produces the correct output, leading to misleading FL inferences. In widely used datasets, the presence of CC compromises the reliability of passed test cases (i.e., negative labels). In contrast, failed test cases (i.e., positive labels) remain definitive. In FL scenarios, unlabeled data is typically abundant and primarily consists of passed test cases. Therefore, systematically leveraging positive and unlabeled data for accurate CC detection is essential, which is beneficial to FL. To tackle the problem, we propose a two-stagE positiVe and unlAbeled learning model for coiNcidental correctneSs detection, EVANS. EVANS defines failed test cases as positive samples and treats the remaining ones as unlabeled data. It comprises two core modules: (1) A module for selecting high-quality pseudo-negative samples. This module leverages vector distance metrics to identify high-quality pseudo-negative test cases, using inter-class distances computed via a pre-trained model. (2) A weakly supervised contrastive learning module. This module utilizes the labeled samples from Stage (1) to train a contrastive learning model, which then detects CC in unlabeled test cases. Experimental results demonstrate that EVANS significantly outperforms current CC detection methods. Huan Xie 0002, Yan Lei 0005 |
ASE | 2 |
| 2025 | From Sparse to Structured: A Diffusion-Enhanced and Feature-Aligned Framework for Coincidental Correctness DetectionabstractCoincidental correctness (CC) refers to test cases that execute faulty code but still produce excepted outputs. This phenomenon introduces noise into the data of software testing-related tasks. As demonstrated in the literature, CC has negative impact on test suite reduction, test case prioritization, fault localization, and automated program repair. Thus, it is essential to detect and mitigate the impact of CC. Although CC is commonly observed across a large number of programs, CC test cases are typically sparse within each program’s test suite. In other words, CC test cases generally make up merely a small portion of the passing test cases. The proportions vary from 3.27% to 31.74% within Defects4J V1.4. This results in a highly imbalanced distribution of CC versus non-CC test cases, posing challenges for accurate detection.To address this issue, we propose a Diffusion-Enhanced and Feature-Aligned Framework for Coincidental Correctness detection, named DEFACC, to obtain more structured representations of test cases. Specifically, DEFACC first introduces a diffusion-based generation module. This module generates new CC samples from original samples to alleviate class imbalance issue and enhance the diversity of CC samples. However, generated feature samples may deviate from the distribution of real CC samples. Such shifts can hurt model reliability and generalization. To resolve this, DEFACC integrates a feature alignment module that is founded on the Maximum Mean Discrepancy (MMD) loss. This module enforces distributional consistency between generated and original CC samples during training. Together, these components ensure that the augmented samples are from sparse to structured, which is not only quantitatively balanced but also semantically faithful. Experimental results show that the DEFACC significantly improves the performance of existing CC detection methods and provides a stronger representation foundation for accurate fault localization. Huan Xie 0002, Yan Lei 0005 |
ASE | 1 |
| 2024 | Combining Coverage and Expert Features with Semantic Representation for Coincidental Correctness DetectionabstractCoincidental correctness (CC) can be misleading for developers because it gives the impression that the code is functioning correctly when there are hidden faults. To mitigate the negative impacts of CC test cases, extensive research has been conducted on their detection, employing either coverage-based or expert-based features. These studies have yielded promising results. Coverage and expert features each provide unique insights into program execution, yet the literature has not fully explored the combined potential of these two feature sets to enhance the detection of CC. Additionally, the rich semantics of the test code and focal method have not been fully utilized. Therefore, we propose to build a unified model, CORE, that integrates coverage and expert features with semantic representations of test and focal methods to improve the detection of CC test cases. We make a comprehensive evaluation with six state-of-the-art baselines on the widely-used Defects4J benchmark. The experimental results show that CORE outperforms the baselines in terms of CC detection accuracy, with a substantial improvement (i.e., 40% improvement on average in terms of F1 score). Then, we conduct the ablation experiment to show that the coverage, expert, and semantics contribute to CORE. CORE can also improve the effectiveness of spectrum-based and mutation-based fault localization performance (e.g., 50% improvements for spectrum-based formula Dstar and 44% improvements for mutation-based method MUSE under relabeling strategy). Huan Xie 0002, Yan Lei 0005, Maojin Li, Meng Yan 0001 |
ASE | 1 |
| 2024 | Flakyrank: Predicting Flaky Tests Using Augmented Learning to RankabstractThe ideal principle of software testing is that test results ought to be deterministic: a test failure indicates the presence of a software bug, while a test success suggests the absence of a bug. Nevertheless, flaky tests break the principle. Flaky tests yield inconsistent results when executed repeatedly under the same conditions. The most straightforward approach runs the tests multiple times to predict flaky tests whereas it is highly time-consuming. Many researchers have proposed efficient approaches to reduce the cost, e.g., recent approaches leverage machine learning techniques for the prediction of flaky tests. However, traditional machine learning primarily focuses on predicting specific instances, which is not conducive to identify flaky tests across an entire project. Therefore, we propose Flakyrank, a ranking framework based on augmented learning to rank to predict flaky tests. The insight is that learning to rank, as compared with traditional machine learning, not only concentrates on individual samples but also optimizes the overall ranking. Since flaky tests constitute a small proportion of the dataset (i.e., approximately 3.6% of the total tests), we utilize generative adversarial networks to generate some synthetic flaky tests to augment the dataset. Based on the augmented dataset, FLAKYRANK treats predicting flaky tests as an information retrieval task, where newly detected flaky tests and test cases serve as queries and documents, respectively. For each newly detected flaky test (i.e., query), FLAKYRANK combines multiple relevant features into a learning to rank model to predict flaky tests candidate tests. We conduct large-scale experiments on different learning to rank models, and the results show that FLAKYRANK with the LambdaMART algorithm yields the best performance. In addition, the experimental results on 23 benchmark projects show that FLAKYRANK outperforms the state-of-the-art predictors. Jiaguo Wang, Yan Lei 0005, Maojin Li, Guanyu Ren, Huan Xie 0002, Shifeng Jin |
SANER | 5 |
| 2024 | Labelrepair: Sequence Labelling for Compilation Errors RepairabstractManual fixing of compilation errors could be a tedious and time-consuming task for novice programmers, and even for experienced ones. In recent years, an increasing number of automated repair techniques have been proposed to guide novice programmers and improve the efficiency of software development. Among them, learning-based automated repair techniques have achieved promising results in terms of repair accuracy. However, existing approaches neglect the time efficiency of patch generation, and often treat the compilation errors repair as a neural machine translation task. The end-to-end repair model decoding cannot be parallelized during the inference stage and suffers from redundant decoding search space. Furthermore, the large search space brought by the model poses a potential risk of semantic tampering. To this end, we propose Labelrepair, a novel repair technique that treats the repair of compilation errors as a sequence labelling task. Labelrepair discards the decoding model and converts the search for patches to a search for the error mapping actions between broken code and patch code pairs. In this way, the search space for patch tokens is plum-meted from the entire vocabulary to the size of edit action labels, and the logical semantics of the original program are preserved to some extent. The time complexity of inference is reduced from$O(n)$to$O$(1) owing to the parallel generation of edit action labels. Through a comprehensive evaluation of Labelrepair on two datasets, we demonstrate that Labelrepair is able to generate patches instantly (0.71ms on average), which is 28 times faster than existing end-to-end repair models. Compared with existing edit-based repair models, Labelrepair achieves the state-of-art repair accuracy. Deheng Yang, Yan Lei 0005, Huan Xie 0002, Minghua Tang, Maojin Li |
SANER | 4 |
| 2024 | Demystifying API misuses in deep learning applications
Deheng Yang, Kui Liu 0001, Yan Lei 0005, Li Li 0029, Huan Xie 0002, Xiaoguang Mao, Tegawendé F. Bissyandé |
Empir. Softw. Eng. | 5 |
| 2024 | Towards More Precise Coincidental Correctness Detection With Deep Semantic LearningabstractCoincidental correctness (CC) is a situation during the execution of a test case, the buggy entity is executed, but the program behaves correctly as expected. Many automated fault localization (FL) techniques use runtime information to discover the underlying connection between the executed buggy entity and the failing test result. The existence of CC will weaken such connection, mislead the FL algorithms to build inaccurate models, and consequently, decrease the localization accuracy. To alleviate the adverse effect of CC on FL, CC detection techniques have been proposed to identify the possible CC tests via heuristic or machine learning algorithms. However, their performance on precision is not satisfactory since they overestimate the possible CC tests and are insufficient in learning the deep semantic features. In this work, we propose a novelTriplet network-basedCoincidentalCorrectness detection technique (i.e.,TriCoCo) to overcome the limitations of the prior works.TriCoConarrows the possible CC tests by designing three features to identify genuine passing tests. Instead of using all tests as inputs by existing techniques,TriCoCotakes the identified genuine passing tests and failing ones to train a triplet model that can evaluate their relative distance. Finally,TriCoCoinfers the probability of being a CC test of the test in the rest of the passing tests by using the trained triplet model. We conduct large-scale experiments to evaluateTriCoCobased on the widely-used Defects4J benchmark. The results demonstrate thatTriCoCocan improve not only the precision of CC detection but also the effectiveness of FL techniques,e.g.,the precision ofTriCoCois 80.33$\%$on average, andTriCoCoboosts the efficacy of DStar by 18$\%$–74$\%$in terms of MFR metric when compared to seven state-of-the-art CC detection baselines. Huan Xie 0002, Yan Lei 0005, Meng Yan 0001, Shanshan Li 0001, Xiaoguang Mao, Yue Yu 0001, David Lo 0001 |
IEEE Trans. Software Eng. | 1 |
| 2023 | On the Reliability of Coverage Data for Fault LocalizationabstractThe high quality of input data serves as the foundation for various tasks. Inaccurate data may decrease the effectiveness of elaborate algorithms and significantly impact the output. This also applies to fault localization, as accurate and reliable data is crucial for effective fault localization techniques. Many fault localization techniques analyze the coverage information for detecting bug positions. However, the source coverage data suffers from various problems, such as the imbalanced data and the coincidental correctness. These problems make the source coverage data unreliable for fault localization. To mitigate the potential adverse effect of these unreliable factors, we propose Orlando, a cOveRage-based decoupLing And recoNstructingData apprOach for fault localization. Or-landooptimizes the coverage data by synthesizing passing coverage with less coincidental correctness and failing coverage with more balanced data. The reconstructed data can provide more reliable source data for fault localization. We evaluate Orlando using the widely used Defects4J benchmark and demonstrate its effectiveness in improving two spectrum-based and two deep learning-based methods. Furthermore, Orlando outperforms state-of-the-art data optimization approaches in fault localization. Huan Xie 0002, Maojin Li, Yan Lei 0005, Shanshan Li 0001, Xiaoguang Mao, Yue Yu 0001 |
APSEC | 1 |
| 2023 | Contrastive Coincidental Correctness Representation LearningabstractA test suite is indispensable for fault localization by providing useful execution information of its test cases for locating suspicious statements of being faulty. There exists a type of test cases known as coincidental correctness (CC) test cases, which executes the faulty statement whereas produces the anticipated output. The existing studies have shown CC test cases harmfully impact fault localization effectiveness. Therefore, it is crucial to detect CC test cases to mitigate the adverse impact of CC test cases on fault localization.To address this issue, we propose ContraCC: a CC test cases detection method using contrastive learning. The insight of ContraCC is that the internal structural information of source test case execution data should be beneficial for CC detection whereas there is a lack of suitable representation methods. Inspired by the insight, ContraCC uses contrastive learning to learn new differentiated representations as test case vectors, which differentiate between similar and dissimilar pairs of test cases by maximizing their similarity within the same class and minimizing it between different classes. Based on the contrastive learning representations (i.e., test case vectors), ContraCC adopts multi-layer perceptron for binary classification to detect CC in downstream tasks. To evaluate the effectiveness of ContraCC, we conduct large-scale experiments on widely-used benchmarks by comparing ContraCC with five state-of-the-art CC test cases detection methods and applying ContraCC for fault localization. The experimental results show that ContraCC outperforms four state-of-the-art methods (e.g., from 10% to 84% improvement in Top-N on the best-performing baseline NeuralCCD) and significantly improves fault localization effectiveness (e.g., 24% improvement on the best-performing baseline Dstar). Maojin Li, Yan Lei 0005, Huan Xie 0002, Jiaguo Wang, Zhengxiong Deng |
ISSRE | 3 |
| 2023 | Mitigating the Effect of Class Imbalance in Fault Localization Using Context-aware Generative Adversarial NetworkabstractFault localization (FL) analyzes the execution information of a test suite to pinpoint the root cause of a failure. The class imbalance of a test suite, i.e., the imbalanced class proportion between passing test cases (i.e., majority class) and failing ones (i.e., minority class), adversely affects FL effectiveness.To mitigate the effect of class imbalance in FL, we propose CGAN4FL: a data augmentation approach using Context-aware Generative Adversarial Network for Fault Localization. Specifically, CGAN4FL uses program dependencies to construct a failure-inducing context showing how a failure is caused. Then, CGAN4FL leverages a generative adversarial network to analyze the failure-inducing context and synthesize the minority class of test cases (i.e., failing test cases). Finally, CGAN4FL augments the synthesized data into original test cases to acquire a class-balanced dataset for FL. Our experiments show that CGAN4FL significantly improves FL effectiveness, e.g., promoting MLP-FL by 200.00%, 25.49%, and 17.81% under the Top-1, Top-5, and Top-10 respectively. Yan Lei 0005, Tiantian Wen, Huan Xie 0002, Lingfeng Fu |
ICPC | 3 |
| 2023 | NeuralCCD: Integrating Multiple Features for Neural Coincidental Correctness DetectionabstractFault localization seeks to locate the suspicious statements possible for causing a program failure. Experimental evidence shows that fault localization effectiveness is affected adversely by the existence of coincidental correctness (CC) test cases, where a CC test case denotes the test case which executes a fault but no failure occurs. Even worse, CC test cases are prevailing in realistic testing and debugging, leading to a severe issue on fault localization effectiveness. Thus, it is indispensable to accurately detect CC test cases and alleviate their harmful effect on fault localization effectiveness.To address this problem, we propose NeuralCCD: a neural coincidental correctness detection approach by integrating multiple features. Specifically, NeuralCCD first leverages suspiciousness score, coverage ratio and similarity to define three CC detection features. Based on these CC detection features and CC labels, NeuralCCD utilizes multi-layer perceptron to learn a different feature-based model for a program, and finally combine the trained models of different programs as an ensemble system to detect CC test cases. To evaluate the effectiveness of NeuralCCD, we conduct large-scale experiments on 247 faulty version of five representative benchmarks and compare NeuralCCD with four state-of-the-art CC detection approaches. The experimental results show that NeuralCCD significantly improves the effectiveness of CC detection, e.g., NeuralCCD yields by at most 109.5%, 93% and 81.3% improvement of Top-1, Top-3 and Top-5 over Tech-I when utilized in Dstar formular. Zhou Tao, Yan Lei 0005, Huan Xie 0002 |
SANER | 3 |
| 2023 | A light-weight data augmentation method for fault localization
Huan Xie 0002, Yan Lei 0005 |
Inf. Softw. Technol. | 2 |
| 2022 | A Universal Data Augmentation Approach for Fault LocalizationabstractData is the fuel to models, and it is still applicable in fault localization (FL). Many existing elaborate FL techniques take the code coverage matrix and failure vector as inputs, expecting the techniques could find the correlation between program entities and failures. However, the input data is high-dimensional and extremely unbalanced since the real-world programs are large in size and the number of failing test cases is much less than that of passing test cases, which are posing severe threats to the effectiveness of FL techniques. Huan Xie 0002, Yan Lei 0005, Meng Yan 0001, Yue Yu 0001, Xin Xia 0001, Xiaoguang Mao |
ICSE | 1 |
| 2022 | Context-based cluster fault localizationabstractAutomated fault localization techniques collect runtime information as input data to identify suspicious statement potentially responsible for program failures. To discover the statistical coincidences between test results (i.e., failing or passing) and the executions of the different statements of a program (i.e., executed or not executed), researchers developed a suspiciousness methodology (e.g., spectrum-based formulas and deep neural network models). However, the occurrences of coincidental correctness (CC) which means the faulty statements were executed but the output of the program was right affect the effectiveness of fault localization. Many researchers seek to identify CC tests using cluster analysis. However, the high-dimensional data containing too much noise reduce the effectiveness of cluster analysis. Junji Yu, Yan Lei 0005, Huan Xie 0002, Lingfeng Fu |
ICPC | 3 |
| 2022 | BCL-FL: A Data Augmentation Approach with Between-Class Learning for Fault LocalizationabstractAutomated fault localization (FL) techniques collect runtime information as input data and then analyze input data to identify the relationship between program statements and failures. They usually take advantages of the statistics of the input data to develop a suspiciousness evaluation methodology (e.g., spectrum-based formulas and deep neural network models) by exploring the underlying correlation rooted in the input data. Thus, the quality of input data is critical for FL. In the actual process of development, developers seek to generate adequate test cases for testing the function or the robustness of a subject program. However, regarding a fault, most test cases are passed test cases and a very few ones are failed test cases since a very small portion of inputs in input domain will lead to a program failure. It means that FL usually faces a problem of imbalanced data, and this problem has been proven to pose an adverse effect on FL effectiveness. To address this problem, we propose BCL-FL: a data augmentation approach based on between-class learning, which produces new synthesized failed test samples by mixing two classes of real test cases (i.e., a passed test case and a failed one) with a random ratio. Specifically, BCL-FL uses the characteristics of real failed test cases to design a data synthesis formula suitable for failed test samples, which can make the synthesized failed test samples closer to real test cases. Since the synthesized data is different from real data, we ingeniously assign a continuous value between 0 and 1 to label the synthesized sample according to the mixing ratio of original labels. We take the synthesized failed test samples and the original test cases as the balanced input data for FL techniques to address the imbalanced data problem. To evaluate the effectiveness of BCL-FL, we conduct large-scale experiments on 287 faulty versions of eight large-sized programs (from ManyBugs and Defects4J) using six state-of-the-art FL approaches. The experimental results show that BCL-FL significantly improves the effectiveness of existing FL techniques, e.g., BCL-FL improves the CNN-FL approach in Top-1, Top-5, and Top-10 by 150%, 136.36%, and 193.1%, respectively. Yan Lei 0005, Huan Xie 0002, Sheng Huang 0001, Meng Yan 0001, Zhou Xu 0003 |
SANER | 3 |
| 2022 | Feature-FL: Feature-Based Fault LocalizationabstractFault localization aims at developing an effective methodology identifying suspicious statements potentially responsible for program failures. The spectrum-based fault localization is the widely used methodology by analyzing the statistical coincidences viewed from the spectrum to evaluate the suspiciousness of each statement of being faulty. However, just analyzing statistical coincidences in the coverage information perspective and without combining diverse amount of information may restrict fault localization effectiveness. Thus, this article proposes feature-based fault localization (Feature-FL): A family fault localization methodology of feature-based metrics by combining the feature diversity from the view of program features into suspiciousness evaluation. Specifically,Feature-FLdefines a concept of branching execution probability to abstract program behaviors as the values of features. Then,Feature-FLuses feature selection (i.e., a family of feature-based metrics) to evaluate the relevance of each feature with program failures. Finally,Feature-FLassociates each feature with its corresponding statement, and uses the relevance as the suspiciousness to locate suspicious statements. We present six feature-based metrics forFeature-FL, and conduct an extensive study to evaluate the effectiveness ofFeature-FLand its potential over the state-of-the-art spectrum-based formulas. Our results provide insight into the potential among different feature-based metrics and also showFeature-FLsignificantly outperforms the state-of-the-art spectrum-based formulas, e.g., an averagesavingof at least 30% over spectrum-based formulas in case of real faults. Yan Lei 0005, Huan Xie 0002, Tao Zhang 0001, Meng Yan 0001, Zhou Xu 0003, Chengnian Sun |
IEEE Trans. Reliab. | 2 |
| 2021 | Is the Ground Truth Really Accurate? Dataset Purification for Automated Program RepairabstractDatasets of real-world bugs shipped with human-written patches are intensively used in the evaluation of existing automated program repair (APR) techniques, wherein the human-written patches always serve as the ground truth, for manual or automated assessment approaches, to evaluate the correctness of test-suite adequate patches. An inaccurate human-written patch tangled with other code changes will pose threats to the reliability of the assessment results. Therefore, the construction of such datasets always requires much manual effort on isolating real bug fixes from bug fixing commits. However, the manual work is time-consuming and prone to mistakes, and little has been known on whether the ground truth in such datasets is really accurate.In this paper, we propose DEPTEST, an automated DatasEt Purification technique from the perspective of triggering Tests. Leveraging coverage analysis and delta debugging, DEPTEST can automatically identify and filter out the code changes irrelevant to the bug exposed by triggering tests. To measure the strength of DEPTEST, we run it on the most extensively used dataset (i.e., Defects4J) that claims to already exclude all irrelevant code changes for each bug fix via manual purification. Our experiment indicates that even in a dataset where the bug fix is claimed to be well isolated, 41.01% of human-written patches can be further reduced by 4.3 lines on average, with the largest reduction reaching up to 53 lines. This indicates its great potential in assisting in the construction of datasets of accurate bug fixes. Furthermore, based on the purified patches, we re-dissect Defects4J and systematically revisit the APR of multi-chunk bugs to provide insights for future research targeting such bugs. Deheng Yang, Yan Lei 0005, Xiaoguang Mao, David Lo 0001, Huan Xie 0002, Meng Yan 0001 |
SANER | 5 |