Josie Holmes

dblp:196/4995 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
3 papers
Software testing · 82% Debugging and program repair · 18%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software testing › test optimization
test case reduction
0.622017
A suite of tools for making effective use of automatically generated tests · ISSTA 2017
One test to rule them all · ISSTA 2017
Software testing › test generation
automated test generation
0.412020
Using Relative Lines of Code to Guide Automated Test Generation for Python · ACM Trans. Softw. Eng. Methodol. 2020
Software testing › test generation
coverage-based test generation
0.412020
Using Relative Lines of Code to Guide Automated Test Generation for Python · ACM Trans. Softw. Eng. Methodol. 2020
Software testing
fault detection
0.412020
Using Relative Lines of Code to Guide Automated Test Generation for Python · ACM Trans. Softw. Eng. Methodol. 2020
Software testing › random testing
property-based testing
0.412020
Using Relative Lines of Code to Guide Automated Test Generation for Python · ACM Trans. Softw. Eng. Methodol. 2020
Debugging and program repair › automated debugging
delta debugging
0.312017
A suite of tools for making effective use of automatically generated tests · ISSTA 2017
Debugging and program repair
fault localization
0.312017
One test to rule them all · ISSTA 2017
Software testing
test generation
0.312017
A suite of tools for making effective use of automatically generated tests · ISSTA 2017

Methods — techniques the papers use, named apart from their topics

delta debugging · 0.6random testing · 0.4lines-of-code heuristic · 0.4
YearPublicationVenuePosition
2020 Practical Automatic Lightweight Nondeterminism and Flaky Test Detection and Debugging for Python
abstract
A critically important, but surprisingly neglected, aspect of system reliability is system predictability. Many soft-ware systems are implemented using mechanisms (unsafe languages, concurrency, caching, stochastic algorithms, environmental dependencies) that can introduce unexpected and unwanted behavioral nondeterminism. Such nondeterministic behavior can result in software bugs and flaky tests as well as causing problems for test reduction, differential testing, and automated regression test generation. We show that lightweight techniques, requiring little effort on the part of developers, can extend an existing testing system to allow detection and debugging of nondeterminism. We show how to make delta-debugging effective for probabilistic faults in general, and that our methods can improve mutation score by 6% for a strong, full differential test harness for a widely used mock file system.
Alex Groce, Josie Holmes
QRS2
2020 Using mutants to help developers distinguish and debug (compiler) faults
abstract
Summary Measuring the distance between two program executions is a fundamental problem in dynamic analysis of software and useful in many test generation and debugging algorithms. This paper proposes a metric for measuring distance between executions and specializes it to an important application: determining similarity of failing test cases for the purpose of automated fault identification and localization in debugging based on automatically generated compiler tests. The metric is based on a causal concept of distance where executions are similar to the degree that changes in the program itself, introduced by mutation, cause similar changes in the correctness of the executions. Specifically, if two failing test cases for a compiler can be made to pass by applying the same mutant, those two tests are more likely to be due to the same fault. We evaluate our metric using more than 50 faults and 2,800 test cases for two widely used real‐world compilers and demonstrate improvements over state‐of‐the‐art methods for fault identification and localization. A simple operator selection approach to reducing the number of mutants can reduce the cost of our approach by 70%, while producing a gain in fault identification accuracy. We additionally show that our approach, although devised for compilers, is applicable as a conservative fault localization algorithm for other types of programs and can help triage certain types of crashes found in fuzzing non‐compiler programs more effectively than a state‐of‐the‐art technique.
Josie Holmes, Alex Groce
Softw. Test. Verification Reliab.1
2020 Using Relative Lines of Code to Guide Automated Test Generation for Python
abstract
Raw lines of code (LOC) is a metric that does not, at first glance, seem extremely useful for automated test generation. It is both highly language-dependent and not extremely meaningful, semantically, within a language: one coder can produce the same effect with many fewer lines than another. However, relative LOC , between components of the same project, turns out to be a highly useful metric for automated testing. In this article, we make use of a heuristic based on LOC counts for tested functions to dramatically improve the effectiveness of automated test generation. This approach is particularly valuable in languages where collecting code coverage data to guide testing has a very high overhead. We apply the heuristic to property-based Python testing using the TSTL (Template Scripting Testing Language) tool. In our experiments, the simple LOC heuristic can improve branch and statement coverage by large margins (often more than 20%, up to 40% or more) and improve fault detection by an even larger margin (usually more than 75% and up to 400% or more). The LOC heuristic is also easy to combine with other approaches and is comparable to, and possibly more effective than, two well-established approaches for guiding random testing.
Josie Holmes, Iftekhar Ahmed 0001, Caius Brindescu, Rahul Gopinath, He Zhang 0025, Alex Groce
ACM Trans. Softw. Eng. Methodol.1
2018 Causal Distance-Metric-Based Assistance for Debugging after Compiler Fuzzing
abstract
Measuring the distance between two program executions is a fundamental problem in dynamic analysis of software, and useful in many test generation and debugging algorithms. This paper proposes a metric for measuring distance between executions, and specializes it to an important application: determining similarity of failing test cases for the purpose of automated fault identification and localization in debugging based on automatically generated compiler tests. The metric is based on a causal concept of distance where executions are similar to the degree that changes in the program itself, introduced by mutation, cause similar changes in the correctness of the executions. Specifically, if two failing test cases (for the original compiler) become successful due to the same mutant, they are more likely to be due to the same fault. We evaluate our metric using more than 50 faults and 2,800 test cases for two widely-used real-world compilers, and demonstrate improvements over state-of-the-art methods for fault identification and localization.
Josie Holmes, Alex Groce
ISSRE1
2018 How verified (or tested) is my code? Falsification-driven verification and testing
Alex Groce, Iftekhar Ahmed 0001, Carlos Jensen, Paul E. McKenney, Josie Holmes
Autom. Softw. Eng.5
2018 TSTL: the template scripting testing language
Josie Holmes, Alex Groce, Jervis Pinto, Pranjal Mittal, Pooria Azimi, Kevin Kellar, James O'Brien
Int. J. Softw. Tools Technol. Transf.1
2017 One test to rule them all
abstract
Test reduction has long been seen as critical for automated testing. However, traditional test reduction simply reduces the length of a test, but does not attempt to reduce semantic complexity. This paper extends previous efforts with algorithms for normalizing and generalizing tests. Rewriting tests into a normal form can reduce semantic complexity and even remove steps from an already delta-debugged test. Moreover, normalization dramatically reduces the number of tests that a reader must examine, partially addressing the ``fuzzer taming'' problem of discovering distinct faults in a set of failing tests. Generalization, in contrast, takes a test and reports what aspects of the test could have been changed while preserving the property that the test fails. Normalization plus generalization aids understanding of tests, including tests for complex and widely used APIs such as the NumPy numeric computation library and the ArcPy GIS scripting package. Normalization frequently reduces the number of tests to be examined by well over an order of magnitude, and often to just one test per fault. Together, ideally, normalization and generalization allow a user to replace reading a large set of tests that vary in unimportant ways with reading one annotated summary test.
Alex Groce, Josie Holmes, Kevin Kellar
ISSTA2
2017 A suite of tools for making effective use of automatically generated tests
abstract
Automated test generation tools (we hope) produce failing tests from time to time. In a world of fault-free code this would not be true, but in such a world we would not need automated test generation tools. Failing tests are generally speaking the most valuable products of the testing process, and users need tools that extract their full value. This paper describes the tools provided by the TSTL testing language for making use of tests (which are not limited to failing tests). In addition to the usual tools for simple delta-debugging and executing tests as regressions, TSTL provides tools for 1) minimizing tests by criteria other than failure, such as code coverage, 2) normalizing tests to achieve further reduction and canonicalization than provided by delta-debugging, 3) generalizing tests to describe the neighborhood of similar tests that fail in the same fashion, and 4) avoiding slippage, where delta-debugging causes a failing test to change underlying fault. These tools can be accessed both by easy-to-use command-line tools and via a powerful API that supports more complex custom test manipulations.
Josie Holmes, Alex Groce
ISSTA1