VLDB 2026 Research / reviewers in the wild / expert
Rahul Gopinath
dblp:93/11109
· DBLP profile ↗
30ranked-venue papers
12as first author
10since 2021 · last 2025
0000-0001-9953-0930ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 27 · 11 first-author · 8 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automatic Data Repair without Format SpecificationsabstractIn data processing, datasets are expected to adhere to specific formats. However, inconsistencies due to human error, data corruption, or partial transmission can render these datasets nonconforming, hindering automated processing. This necessitates manual data repair, a time-consuming and errorprone task, especially when formal specifications are unavailable. To address this challenge, we introduce $\epsilon$REPAIR, a novel format-free approach to automating data repair. $\epsilon$REPAIR leverages parser feedback to detect and correct data inconsistencies, making it a versatile solution for data cleansing. In evaluation, $\epsilon$REPAIR achieves $2.6 \times$ higher-quality repairs than its closest competitor, DDMax, in terms of the number of edits required to restore corrupted data, while reducing data loss by $2.8 \times$ compared to DDMax, with only $1.4 \times$ runtime overhead. This work presents a practical, robust, and flexible formatfree data repair alternative to DDMax. Its applications extend to domains such as data science, software development, and other human-centric systems, where handling diverse and inconsistent datasets is critical. Zijian Luo, Lukas Kirschner, Ezekiel O. Soremekun, Rahul Gopinath |
ISSRE | 4 |
| 2024 | Empirical Evaluation of Frequency Based Statistical Models for Estimating Killable MutantsabstractBackground. Mutation analysis is the premier technique for evaluating test suite quality estimating residual software defects. However, the reliability of mutation analysis is hampered by equivalent mutants which are undetectable by test cases. Reliably detecting and eliminating killable mutants is difficult as it is highly program and location dependent. Statistical estimation of killable mutants seems to be a promising approach to tackle this problem. Aims. Frequency-based species estimation methods have been proposed as a solution for several related problems in software testing. This paper investigates whether such frequency-based estimation methods can accurately estimate the number of killable mutants. Method. We conducted a large-scale empirical study on the ability of twelve widely known frequency-based estimators to predict the number of killable mutants in ten mature software projects. Result. Our investigation finds limited or no evidence that any of the statistical estimators are able to consistently predict the number of killable mutants in projects evaluated. Conclusion. We found that the investigated estimators lack sufficient predictive power and cannot produce reliable and useful estimates of killable mutants. Konstantin Kuznetsov 0001, Alessio Gambi, Saikrishna Dhiddi, Julia Hess, Rahul Gopinath |
ESEM | 5 |
| 2024 | FormatFuzzer: Effective Fuzzing of Binary File FormatsabstractEffective fuzzing of programs that process structured binary inputs, such as multimedia files, is a challenging task, since those programs expect a very specific input format. Existing fuzzers, however, are mostly format-agnostic, which makes them versatile, but also ineffective when a specific format is required. We present FormatFuzzer , a generator for format-specific fuzzers . FormatFuzzer takes as input a binary template (a format specification used by the 010 Editor) and compiles it into C++ code that acts as parser, mutator, and highly efficient generator of inputs conforming to the rules of the language. The resulting format-specific fuzzer can be used as a standalone producer or mutator in black-box settings, where no guidance from the program is available. In addition, by providing mutable decision seeds, it can be easily integrated with arbitrary format-agnostic fuzzers such as AFL to make them format-aware. In our evaluation on complex formats such as MP4 or ZIP, FormatFuzzer showed to be a highly effective producer of valid inputs that also detected previously unknown memory errors in ffmpeg and timidity . Rafael Dutra, Rahul Gopinath, Andreas Zeller |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | SEAL: Capability-Based Access Control for Data-Analytic ScenariosabstractData science is the basis for various disciplines in the Big-Data era. Due to the high volume, velocity, and variety of big data, data owners often store their data in data servers. Past few years, many computation techniques have emerged to protect the security and privacy of such shared data while enabling analysis thereon. Hence, access-control systems must provide a fine-grained, multi-layer mechanism to protect data. However, the existing systems and frameworks fail to satisfy all these requirements and resolve the trust issue between data owners and analysts. Hamed Rasifard, Rahul Gopinath, Michael Backes 0001, Hamed Nemati |
SACMAT | 2 |
| 2023 | Systematic Assessment of Fuzzers using Mutation Analysis
Philipp Görz, Björn Mathis, Keno Hassler, Emre Güler, Thorsten Holz, Andreas Zeller, Rahul Gopinath |
USENIX Security Symposium | 7 |
| 2022 | "Synthesizing input grammars": a replication studyabstractWhen producing test inputs for a program, test generators ("fuzzers") can greatly profit from grammars that formally describe the language of expected inputs. In recent years, researchers thus have studied means to recover input grammars from programs and their executions. The GLADE algorithm by Bastani et al., published at PLDI 2017, was the first black-box approach to claim context-free approximation of input specification for non-trivial languages such as XML, Lisp, URLs, and more. Bachir Bendrissou, Rahul Gopinath, Andreas Zeller |
PLDI | 2 |
| 2022 | CLIFuzzer: mining grammars for command-line invocationsabstractThe behavior of command-line utilities can be very much influenced by passing command-line options and arguments—configuration settings that enable, disable, or otherwise influence parts of the code to be executed. Hence, systematic testing of command-line utilities requires testing them with diverse configurations of supported command-line options. Abhilash Gupta, Rahul Gopinath, Andreas Zeller |
ESEC/SIGSOFT FSE | 2 |
| 2022 | Mutation analysis and its industrial applicationsabstractWe are pleased to introduce the first part of the special issue on Mutation Testing of the Journal of Software Testing, Verification & Reliability. The issue features three papers selected by a review board of 40 mutation testing experts. Our board was selected from an extensive pool of experts comprising of reviewers from the Mutation workshop series, authors of published manuscripts in established journals related to mutation testing as well as the specific subdomain where mutation was applied. Mutation Testing was introduced in 1970s as a rigorous and automated alternative to fault injection. With over four decades of research and development, it is now accepted as a premier measure of the fault reveling potential of test suites. Mutation Testing is a fault-based testing technique that introduces candidate faults by making syntactic changes. The fault revealing ability of a test suite is then measured as the ratio of mutants it can detect against the total number of detectable mutants. A test suite is mutation testing adequate if the test suite detects all detectable mutants. While mutation testing is usually limited to simple faults, the twin axioms of mutation testing—finite neighborhood and coupling effect—provide confidence that a test suite that is mutation adequate can also reveal complex faults composed of simple faults. While the field of mutation testing has seen active research, numerous important problems remain open. These include (1) how to adapt mutation testing to non-traditional fields such as machine learning, (2) how to apply mutation testing for automatic test generation such as fuzzers, and (3) how to effectively solve the question of equivalent mutants. We thank the editors, authors, and reviewers for their efforts to successfully complete this special issue. Rahul Gopinath, Jie Zhang 0050, Marinos Kintis, Mike Papadakis |
Softw. Test. Verification Reliab. | 1 |
| 2022 | Mutation analysis and its industrial applicationsabstractWe are pleased to introduce the second part of the special issue on Mutation Testing of the Journal of Software Testing, Verification & Reliability. We thank the editors, authors, and reviewers for their efforts to successfully complete this special issue. Rahul Gopinath, Jie Zhang 0050, Marinos Kintis, Mike Papadakis |
Softw. Test. Verification Reliab. | 1 |
| 2021 | Input AlgebrasabstractGrammar-based test generators are highly efficient in producing syntactically valid test inputs, and give their user precise control over which test inputs should be generated. Adapting a grammar or a test generator towards a particular testing goal can be tedious, though. We introduce the concept of a grammar transformer, specializing a grammar towards inclusion or exclusion of specific patterns: "The phone number must not start with 011 or +1". To the best of our knowledge, ours is the first approach to allow for arbitrary Boolean combinations of patterns, giving testers unprecedented flexibility in creating targeted software tests. The resulting specialized grammars can be used with any grammar-based fuzzer for targeted test generation, but also as validators to check whether the given specialization is met or not, opening up additional usage scenarios. In our evaluation on real-world bugs, we show that specialized grammars are accurate both in producing and validating targeted inputs. Rahul Gopinath, Hamed Nemati, Andreas Zeller |
ICSE | 1 |
| 2020 | Abstracting failure-inducing inputsabstractA program fails. Under which circumstances does the failure occur? Starting with a single failure-inducing input ("The input ((4)) fails") and an input grammar, the DDSET algorithm uses systematic tests to automatically generalize the input to an abstract failure-inducing input that contains both (concrete) terminal symbols and (abstract) nonterminal symbols from the grammar—for instance, "(( ))", which represents any expression in double parentheses. Such an abstract failure-inducing input can be used (1) as a debugging diagnostic, characterizing the circumstances under which a failure occurs ("The error occurs whenever an expression is enclosed in double parentheses"); (2) as a producer of additional failure-inducing tests to help design and validate fixes and repair candidates ("The inputs ((1)), ((3 * 4)), and many more also fail"). In its evaluation on real-world bugs in JavaScript, Clojure, Lua, and UNIX command line utilities, DDSET’s abstract failure-inducing inputs provided to-the-point diagnostics, and precise producers for further failure inducing inputs. Rahul Gopinath, Alexander Kampmann, Nikolas Havrikov, Ezekiel O. Soremekun, Andreas Zeller |
ISSTA | 1 |
| 2020 | Learning input tokens for effective fuzzingabstractModern fuzzing tools like AFL operate at a lexical level: They explore the input space of tested programs one byte after another. For inputs with complex syntactical properties, this is very inefficient, as keywords and other tokens have to be composed one character at a time. Fuzzers thus allow to specify dictionaries listing possible tokens the input can be composed from; such dictionaries speed up fuzzers dramatically. Also, fuzzers make use of dynamic tainting to track input tokens and infer values that are expected in the input validation phase. Unfortunately, such tokens are usually implicitly converted to program specific values which causes a loss of the taints attached to the input data in the lexical phase. In this paper, we present a technique to extend dynamic tainting to not only track explicit data flows but also taint implicitly converted data without suffering from taint explosion. This extension makes it possible to augment existing techniques and automatically infer a set of tokens and seed inputs for the input language of a program given nothing but the source code. Specifically targeting the lexical analysis of an input processor, our lFuzzer test generator systematically explores branches of the lexical analysis, producing a set of tokens that fully cover all decisions seen. The resulting set of tokens can be directly used as a dictionary for fuzzing. Along with the token extraction seed inputs are generated which give further fuzzing processes a head start. In our experiments, the lFuzzer-AFL combination achieves up to 17% more coverage on complex input formats like json, lisp, tinyC, and JavaScript compared to AFL. Björn Mathis, Rahul Gopinath, Andreas Zeller |
ISSTA | 2 |
| 2020 | Revisiting the Relationship Between Fault Detection, Test Adequacy Criteria, and Test Set SizeabstractThe research community has long recognized a complex interrelationship between fault detection, test adequacy criteria, and test set size. However, there is substantial confusion about whether and how to experimentally control for test set size when assessing how well an adequacy criterion is correlated with fault detection and when comparing test adequacy criteria. Resolving the confusion, this paper makes the following contributions: (1) A review of contradictory analyses of the relationships between fault detection, test adequacy criteria, and test set size. Specifically, this paper addresses the supposed contradiction of prior work and explains why test set size is neither a confounding variable, as previously suggested, nor an independent variable that should be experimentally manipulated. (2) An explication and discussion of the experimental designs of prior work, together with a discussion of conceptual and statistical problems, as well as specific guidelines for future work. (3) A methodology for comparing test adequacy criteria on an equal basis, which accounts for test set size without directly manipulating it through unrealistic stratification. (4) An empirical evaluation that compares the effectiveness of coverage-based testing, mutation-based testing, and random testing. Additionally, this paper proposes probabilistic coupling, a methodology for assessing the representativeness of a set of test goals for a given fault and for approximating the fault-detection probability of adequate test sets. Yiqun Chen 0001, Rahul Gopinath, Anita Tadakamalla, Michael D. Ernst, Reid Holmes, Gordon Fraser 0001, Paul Ammann, René Just |
ASE | 2 |
| 2020 | Mining input grammars from dynamic control flowabstractOne of the key properties of a program is its input specification. Having a formal input specification can be critical in fields such as vulnerability analysis, reverse engineering, software testing, clone detection, or refactoring. Unfortunately, accurate input specifications for typical programs are often unavailable or out of date. Rahul Gopinath, Björn Mathis, Andreas Zeller |
ESEC/SIGSOFT FSE | 1 |
| 2020 | Using Relative Lines of Code to Guide Automated Test Generation for PythonabstractRaw lines of code (LOC) is a metric that does not, at first glance, seem extremely useful for automated test generation. It is both highly language-dependent and not extremely meaningful, semantically, within a language: one coder can produce the same effect with many fewer lines than another. However, relative LOC , between components of the same project, turns out to be a highly useful metric for automated testing. In this article, we make use of a heuristic based on LOC counts for tested functions to dramatically improve the effectiveness of automated test generation. This approach is particularly valuable in languages where collecting code coverage data to guide testing has a very high overhead. We apply the heuristic to property-based Python testing using the TSTL (Template Scripting Testing Language) tool. In our experiments, the simple LOC heuristic can improve branch and statement coverage by large margins (often more than 20%, up to 40% or more) and improve fault detection by an even larger margin (usually more than 75% and up to 400% or more). The LOC heuristic is also easy to combine with other approaches and is comparable to, and possibly more effective than, two well-established approaches for guiding random testing. Josie Holmes, Iftekhar Ahmed 0001, Caius Brindescu, Rahul Gopinath, He Zhang 0025, Alex Groce |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2019 | Parser-directed fuzzingabstractTo be effective, software test generation needs to well cover the space of possible inputs. Traditional fuzzing generates large numbers of random inputs, which however are unlikely to contain keywords and other specific inputs of non-trivial input languages. Constraint-based test generation solves conditions of paths leading to uncovered code, but fails on programs with complex input conditions because of path explosion. In this paper, we present a test generation technique specifically directed at input parsers. We systematically produce inputs for the parser and track comparisons made; after every rejection, we satisfy the comparisons leading to rejection. This approach effectively covers the input space: Evaluated on five subjects, from CSV files to JavaScript, our pFuzzer prototype covers more tokens than both random-based and constraint-based approaches, while requiring no symbolic analysis and far fewer tests than random fuzzers. Björn Mathis, Rahul Gopinath, Michaël Mera, Alexander Kampmann, Matthias Höschele, Andreas Zeller |
PLDI | 2 |
| 2019 | Evaluating Fault Localization for Resource Adaptation via Test-Based Software ModificationabstractThe ability to dynamically adapt to resource variations is critical for modern-day mission-critical systems that operate in ever-changing resource environments. Test-based Software Modification (TBSM) is a recently proposed technique to build Resource Adaptive Software (RAS) that relies on existing test infrastructure, test labeling, and program modifications. TBSM is simple and applicable, but an inefficient technique; the primary reason for inefficiency is the sheer size of the search space. In this paper, we propose AdFL, a repurposing of Fault Localization (FL) that can shrink (and prioritize) the search space for TBSM more effectively than previously proposed heuristics. We present complete case studies and an empirical analysis of a set of open source projects as evidence that AdFL can significantly reduce the search space in TBSM. We show how to combine AdFL with previous heuristics for TBSM, and propose an incremental, best-effort variant of TBSM that uses AdFL to prioritize the search. Arpit Christi, Alex Groce, Rahul Gopinath |
QRS | 3 |
| 2017 | The Theory of Composite FaultsabstractFault masking happens when the effect of one fault serves to mask that of another fault for particular test inputs. The coupling effect is relied upon by testing practitioners to ensure that fault masking is rare. It states that complex faults are coupled to simple faults in such a way that a test data set that detects all simple faults in a program will detect a high percentage of the complex faults. While this effect has been empirically evaluated, our theoretical understanding of the coupling effect is as yet incomplete. Wah proposed a theory of the coupling effect on finite bijective (or near bijective) functions with the same domain and co-domain and assuming a uniform distribution for candidate functions. This model, however, was criticized as being too simple to model real systems, as it did not account for differing domain and co-domain in real programs, or for the syntactic neighborhood. We propose a new theory of fault coupling for general functions (with certain constraints). We show that there are two kinds of fault interactions, of which only the weak interaction can be modeled by the theory of the coupling effect. The strong interaction can produce faults that are semantically different from the original faults. These faults should hence be considered as independent atomic faults. Our analysis shows that the theory holds even when the effect of the syntactic neighborhood of the program is considered. We analyze numerous real-world programs with real faults to validate our hypothesis. Rahul Gopinath, Carlos Jensen, Alex Groce |
ICST | 1 |
| 2017 | Does choice of mutation tool matter?
Rahul Gopinath, Iftekhar Ahmed 0001, Mohammad Amin Alipour, Carlos Jensen, Alex Groce |
Softw. Qual. J. | 1 |
| 2017 | Mutation Reduction Strategies Considered HarmfulabstractMutation analysis is a well known yet unfortunately costly method for measuring test suite quality. Researchers have proposed numerous mutation reduction strategies in order to reduce the high cost of mutation analysis, while preserving the representativeness of the original set of mutants. As mutation reduction is an area of active research, it is important to understand the limits of possible improvements. We theoretically and empirically investigate the limits of improvement in effectiveness from using mutation reduction strategies compared to random sampling. Using real-world open source programs as subjects, we find an absolute limit in improvement of effectiveness over random sampling- 13.078%. Given our findings with respect to absolute limits, one may ask: How effective are the extant mutation reduction strategies? We evaluate the effectiveness of multiple mutation reduction strategies in comparison to random sampling. We find that none of the mutation reduction strategies evaluated-many forms of operator selection, and stratified sampling (on operators or program elements)-produced an effectiveness advantage larger than 5% in comparison with random sampling. Given the poor performance of mutation selection strategies-they may have a negligible advantage at best, and often perform worse than random sampling- we caution practicing testers against applying mutation reduction strategies without adequate justification. Rahul Gopinath, Iftekhar Ahmed 0001, Mohammad Amin Alipour, Carlos Jensen, Alex Groce |
IEEE Trans. Reliab. | 1 |
| 2016 | On the limits of mutation reduction strategiesabstractAlthough mutation analysis is considered the best way to evaluate the effectiveness of a test suite, hefty computational cost often limits its use. To address this problem, various mutation reduction strategies have been proposed, all seeking to reduce the number of mutants while maintaining the representativeness of an exhaustive mutation analysis. While research has focused on the reduction achieved, the effectiveness of these strategies in selecting representative mutants, and the limits in doing so have not been investigated, either theoretically or empirically. Rahul Gopinath, Mohammad Amin Alipour, Iftekhar Ahmed 0001, Carlos Jensen, Alex Groce |
ICSE | 1 |
| 2016 | Generating focused random tests using directed swarm testingabstractRandom testing can be a powerful and scalable method for finding faults in software. However, sophisticated random testers usually test a whole program, not individual components. Writing random testers for individual components of complex programs may require unreasonable effort. In this paper we present a novel method, directed swarm testing, that uses statistics and a variation of random testing to produce random tests that focus on only part of a program, increasing the frequency with which tests cover the targeted code. We demonstrate the effectiveness of this technique using real-world programs and test systems (the YAFFS2 file system, GCC, and Mozilla's SpiderMonkey JavaScript engine), and discuss various strategies for directed swarm testing. The best strategies can improve coverage frequency for targeted code by a factor ranging from 1.1-4.5x on average, and from nearly 3x to nearly 9x in the best case. For YAFFS2, directed swarm testing never decreased coverage, and for GCC and SpiderMonkey coverage increased for over 99% and 73% of targets, respectively, using the best strategies. Directed swarm testing improves detection rates for real SpiderMonkey faults, when the code in the introducing commit is targeted. This lightweight technique is applicable to existing industrial-strength random testers. Mohammad Amin Alipour, Alex Groce, Rahul Gopinath, Arpit Christi |
ISSTA | 3 |
| 2016 | Evaluating non-adequate test-case reductionabstractGiven two test cases, one larger and one smaller, the smaller test case is preferred for many purposes. A smaller test case usually runs faster, is easier to understand, and is more convenient for debugging. However, smaller test cases also tend to cover less code and detect fewer faults than larger test cases. Whereas traditional research focused on reducing test suites while preserving code coverage, recent work has introduced the idea of reducing individual test cases, rather than test suites, while still preserving code coverage. Other recent work has proposed non-adequately reducing test suites by not even preserving all the code coverage. This paper empirically evaluates a new combination of these two ideas, non-adequate reduction of test cases, which allows for a wide range of trade-offs between test case size and fault detection. Our study introduces and evaluates C%-coverage reduction (where a test case is reduced to retain at least C% of its original coverage) and N -mutant reduction (where a test case is reduced to kill at least N of the mutants it originally killed). We evaluate the reduction trade-offs with varying values of C% and N for four real-world C projects: Mozilla’s SpiderMonkey JavaScript engine, the YAFFS2 flash file system, Grep, and Gzip. The results show that it is possible to greatly reduce the size of many test cases while still preserving much of their fault-detection capability. Mohammad Amin Alipour, August Shi, Rahul Gopinath, Darko Marinov, Alex Groce |
ASE | 3 |
| 2016 | Can testedness be effectively measured?abstractAmong the major questions that a practicing tester faces are deciding where to focus additional testing effort, and deciding when to stop testing. Test the least-tested code, and stop when all code is well-tested, is a reasonable answer. Many measures of "testedness" have been proposed; unfortunately, we do not know whether these are truly effective. In this paper we propose a novel evaluation of two of the most important and widely-used measures of test suite quality. The first measure is statement coverage, the simplest and best-known code coverage measure. The second measure is mutation score, a supposedly more powerful, though expensive, measure. Iftekhar Ahmed 0001, Rahul Gopinath, Caius Brindescu, Alex Groce, Carlos Jensen |
SIGSOFT FSE | 2 |
| 2015 | An Empirical Study of Design Degradation: How Software Projects Get Worse over TimeabstractContext: Software decay is a key concern for large, long-lived software projects. Systems degrade over time as design and implementation compromises and exceptions pile up. Goal: Quantify design decay and understand how software projects deal with this issue. Method: We conducted an empirical study on the presence and evolution of code smells, used as an indicator of design degradation in 220 open source projects. Results: The best approach to maintain the quality of a project is to spend time reducing both software defects (bugs) and design issues (refactoring). We found that design issues are frequently ignored in favor of fixing defects. We also found that design issues have a higher chance of being fixed in the early stages of a project, and that efforts to correct these stall as projects mature and the code base grows, leading to a build-up of problems. Conclusions: From studying a large set of open source projects, our research suggests that while core contributors tend to fix design issues more often than non-core contributors, there is no difference once the relative quantity of commits is accounted for. We also show that design issues tend to build up over time. Iftekhar Ahmed 0001, Umme Ayda Mannan, Rahul Gopinath, Carlos Jensen |
ESEM | 3 |
| 2015 | How hard does mutation analysis have to be, anyway?abstractMutation analysis is considered the best method for measuring the adequacy of test suites. However, the number of test runs required for a full mutation analysis grows faster than project size, which is not feasible for real-world software projects, which often have more than a million lines of code. It is for projects of this size, however, that developers most need a method for evaluating the efficacy of a test suite. Various strategies have been proposed to deal with the explosion of mutants. However, these strategies at best reduce the number of mutants required to a fraction of overall mutants, which still grows with program size. Running, e.g., 5% of all mutants of a 2MLOC program usually requires analyzing over 100,000 mutants. Similarly, while various approaches have been proposed to tackle equivalent mutants, none completely eliminate the problem, and the fraction of equivalent mutants remaining is hard to estimate, often requiring manual analysis of equivalence. In this paper, we provide both theoretical analysis and empirical evidence that a small constant sample of mutants yields statistically similar results to running a full mutation analysis, regardless of the size of the program or similarity between mutants. We show that a similar approach, using a constant sample of inputs can estimate the degree of stubbornness in mutants remaining to a high degree of statistical confidence, and provide a mutation analysis framework for Python that incorporates the analysis of stubbornness of mutants. Rahul Gopinath, Mohammad Amin Alipour, Iftekhar Ahmed 0001, Carlos Jensen, Alex Groce |
ISSRE | 1 |
| 2014 | Code coverage for suite evaluation by developersabstractOne of the key challenges of developers testing code is determining a test suite's quality -- its ability to find faults. The most common approach is to use code coverage as a measure for test suite quality, and diminishing returns in coverage or high absolute coverage as a stopping rule. In testing research, suite quality is often evaluated by a suite's ability to kill mutants (artificially seeded potential faults). Determining which criteria best predict mutation kills is critical to practical estimation of test suite quality. Previous work has only used small sets of programs, and usually compares multiple suites for a single program. Practitioners, however, seldom compare suites --- they evaluate one suite. Using suites (both manual and automatically generated) from a large set of real-world open-source projects shows that evaluation results differ from those for suite-comparison: statement (not block, branch, or path) coverage predicts mutation kills best. Rahul Gopinath, Carlos Jensen, Alex Groce |
ICSE | 1 |
| 2014 | Mutations: How Close are they to Real Faults?abstractMutation analysis is often used to compare the effectiveness of different test suites or testing techniques. One of the main assumptions underlying this technique is the Competent Programmer Hypothesis, which proposes that programs are very close to a correct version, or that the difference between current and correct code for each fault is very small. Researchers have assumed on the basis of the Competent Programmer Hypothesis that the faults produced by mutation analysis are similar to real faults. While there exists some evidence that supports this assumption, these studies are based on analysis of a limited and potentially non-representative set of programs and are hence not conclusive. In this paper, we separately investigate the characteristics of bug-fixes and other changes in a very large set of randomly selected projects using four different programming languages. Our analysis suggests that a typical fault involves about three to four tokens, and is seldom equivalent to any traditional mutation operator. We also find the most frequently occurring syntactical patterns, and identify the factors that affect the real bug-fix change distribution. Our analysis suggests that different languages have different distributions, which in turn suggests that operators optimal in one language may not be optimal for others. Moreover, our results suggest that mutation analysis stands in need of better empirical support of the connection between mutant detection and detection of actual program faults in a larger body of real programs. Rahul Gopinath, Carlos Jensen, Alex Groce |
ISSRE | 1 |
| 2014 | MuCheck: an extensible tool for mutation testing of haskell programsabstractThis paper presents MuCheck, a mutation testing tool for Haskell programs. MuCheck is a counterpart to the widely used QuickCheck random testing tool for functional programs, and can be used to evaluate the efficacy of QuickCheck property definitions. The tool implements mutation operators that are specifically designed for functional programs, and makes use of the type system of Haskell to achieve a more relevant set of mutants than otherwise possible. Mutation coverage is particularly valuable for functional programs due to highly compact code, referential transparency, and clean semantics; these make augmenting a test suite or specification based on surviving mutants a practical method for improved testing. Mohammad Amin Alipour, Rahul Gopinath, Alex Groce |
ISSTA | 3 |
| 2012 | Explanations for Regular Expressions
Martin Erwig, Rahul Gopinath |
FASE | 2 |