EDBT 2026 Demo / reviewers in the wild / expert
Matthew Staats
dblp:88/6925 · also Matt Staats
· DBLP profile ↗
26ranked-venue papers
10as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 25 · 10 first-authorSecurity and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
18 papers |
Software testing · 76% Empirical software engineering · 10% Program analysis · 7% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 30 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software testing
test oracle |
0.9 | 6 | 2015 | Automated Oracle Data Selection Support · IEEE Trans. Software Eng. 2015 Dodona: automated oracle data set selection · ISSTA 2014 Automated oracle creation support, or: How I learned to stop worrying about fault propagation and love mutation testing · ICSE 2012 |
Software testing › test oracle
oracle data selection |
0.6 | 3 | 2015 | Automated Oracle Data Selection Support · IEEE Trans. Software Eng. 2015 Dodona: automated oracle data set selection · ISSTA 2014 Automated oracle creation support, or: How I learned to stop worrying about fault propagation and love mutation testing · ICSE 2012 |
Software testing
test generation |
0.5 | 3 | 2015 | The Risks of Coverage-Directed Test Case Generation · IEEE Trans. Software Eng. 2015 Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study · ACM Trans. Softw. Eng. Methodol. 2015 Parallel symbolic execution for structural test generation · ISSTA 2010 |
Software testing › test adequacy › coverage criteria
structural coverage criteria |
0.5 | 4 | 2015 | A Flexible and Non-intrusive Approach for Computing Complex Structural Coverage Metrics · ICSE (1) 2015 Observable modified Condition/Decision coverage · ICSE 2013 ReqsCov: A Tool for Measuring Test-Adequacy over Requirements · ASE 2008 |
Software testing › test adequacy › coverage criteria › structural coverage criteria
MC/DC coverage |
0.4 | 2 | 2016 | The Effect of Program and Model Structure on the Effectiveness of MC/DC Test Adequacy Coverage · ACM Trans. Softw. Eng. Methodol. 2016 Observable modified Condition/Decision coverage · ICSE 2013 |
Empirical software engineering
developer studies |
0.4 | 3 | 2015 | Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study · ACM Trans. Softw. Eng. Methodol. 2015 Understanding user understanding: determining correctness of generated program invariants · ISSTA 2012 Does automated white-box test generation really help software testers? · ISSTA 2013 |
Software testing
test adequacy |
0.3 | 2 | 2016 | The Effect of Program and Model Structure on the Effectiveness of MC/DC Test Adequacy Coverage · ACM Trans. Softw. Eng. Methodol. 2016 ReqsCov: A Tool for Measuring Test-Adequacy over Requirements · ASE 2008 |
Software testing › test adequacy
coverage criteria |
0.2 | 1 | 2016 | The Effect of Program and Model Structure on the Effectiveness of MC/DC Test Adequacy Coverage · ACM Trans. Softw. Eng. Methodol. 2016 |
Software testing › test generation
coverage-based test generation |
0.2 | 1 | 2015 | The Risks of Coverage-Directed Test Case Generation · IEEE Trans. Software Eng. 2015 |
Empirical software engineering
mining software repositories |
0.2 | 1 | 2015 | The Impact of View Histories on Edit Recommendations · IEEE Trans. Software Eng. 2015 |
Software testing › test generation
unit test generation |
0.2 | 1 | 2015 | Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study · ACM Trans. Softw. Eng. Methodol. 2015 |
Software testing › test generation
automated test generation |
0.2 | 1 | 2013 | Does automated white-box test generation really help software testers? · ISSTA 2013 |
Software maintenance and evolution › program comprehension
code exploration |
0.2 | 1 | 2013 | NavClus: a graphical recommender for assisting code exploration · ICSE 2013 |
Software testing › test generation
white-box test generation |
0.2 | 1 | 2013 | Does automated white-box test generation really help software testers? · ISSTA 2013 |
Program analysis › specification mining
dynamic invariant detection |
0.1 | 1 | 2012 | Understanding user understanding: determining correctness of generated program invariants · ISSTA 2012 |
Software testing
mutation testing |
0.1 | 1 | 2012 | Automated oracle creation support, or: How I learned to stop worrying about fault propagation and love mutation testing · ICSE 2012 |
Software testing › test oracle
test oracle generation |
0.1 | 1 | 2012 | Understanding user understanding: determining correctness of generated program invariants · ISSTA 2012 |
Software testing › test optimization
test case selection |
0.1 | 1 | 2011 | Better testing through oracle selection · ICSE 2011 |
Software testing
testing theory |
0.1 | 1 | 2011 | Programs, tests, and oracles: the foundations of testing revisited · ICSE 2011 |
Empirical software engineering
controlled experiment |
0.1 | 1 | 2010 | The influence of multiple artifacts on the effectiveness of software testing · ASE 2010 |
Program analysis › symbolic execution
parallel symbolic execution |
0.1 | 1 | 2010 | Parallel symbolic execution for structural test generation · ISSTA 2010 |
Software testing › structural testing
structural test generation |
0.1 | 1 | 2010 | Parallel symbolic execution for structural test generation · ISSTA 2010 |
Program analysis
symbolic execution |
0.1 | 1 | 2010 | Parallel symbolic execution for structural test generation · ISSTA 2010 |
Software testing › test coverage
code coverage |
0.1 | 1 | 2015 | Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study · ACM Trans. Softw. Eng. Methodol. 2015 |
Compilers and program optimization
partial evaluation |
0.1 | 1 | 2015 | A Flexible and Non-intrusive Approach for Computing Complex Structural Coverage Metrics · ICSE (1) 2015 |
Program analysis
static analysis |
0.1 | 1 | 2015 | A Flexible and Non-intrusive Approach for Computing Complex Structural Coverage Metrics · ICSE (1) 2015 |
Software testing
test coverage |
0.1 | 1 | 2015 | Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study · ACM Trans. Softw. Eng. Methodol. 2015 |
Software testing › test process
test design |
0.1 | 1 | 2015 | 2nd International Workshop on Requirements Engineering and Testing (RET 2015) · ICSE (2) 2015 |
Program analysis
dynamic analysis |
0.1 | 1 | 2014 | Dodona: automated oracle data set selection · ISSTA 2014 |
Recommender systems › domain-specific recommendation
software engineering recommender systems |
0.0 | 1 | 2013 | NavClus: a graphical recommender for assisting code exploration · ICSE 2013 |
Methods — techniques the papers use, named apart from their topics
mutation analysis · 0.6symbolic execution · 0.4controlled experiment · 0.4empirical study · 0.4simulation · 0.2random test generation · 0.2fault-finding effectiveness ranking · 0.2counterexample-based test generation · 0.2association rule mining · 0.2program execution analysis · 0.2interaction trace mining · 0.2static partitioning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | The Effect of Program and Model Structure on the Effectiveness of MC/DC Test Adequacy CoverageabstractTest adequacy metrics defined over the structure of a program, such as Modified Condition and Decision Coverage (MC/DC), are used to assess testing efforts. However, MC/DC can be “cheated” by restructuring a program to make it easier to achieve the desired coverage. This is concerning, given the importance of MC/DC in assessing the adequacy of test suites for critical systems domains. In this work, we have explored the impact of implementation structure on the efficacy of test suites satisfying the MC/DC criterion using four real-world avionics systems. Our results demonstrate that test suites achieving MC/DC over implementations with structurally complex Boolean expressions are generally larger and more effective than test suites achieving MC/DC over functionally equivalent, but structurally simpler, implementations. Additionally, we found that test suites generated over simpler implementations achieve significantly lower MC/DC and fault-finding effectiveness when applied to complex implementations, whereas test suites generated over the complex implementation still achieve high MC/DC and attain high fault finding over the simpler implementation. By measuring MC/DC over simple implementations, we can significantly reduce the cost of testing, but in doing so, we also reduce the effectiveness of the testing process. Thus, developers have an economic incentive to “cheat” the MC/DC criterion, but this cheating leads to negative consequences. Accordingly, we recommend that organizations require MC/DC over a structurally complex implementation for testing purposes to avoid these consequences. Gregory Gay 0002, Ajitha Rajan, Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2015 | 2nd International Workshop on Requirements Engineering and Testing (RET 2015)abstractThe RET (Requirements Engineering and Testing) workshop provides a meeting point for researchers and practitioners from the two separate fields of Requirements Engineering (RE) and Testing. The goal is to improve the connection and alignment of these two areas through an exchange of ideas, challenges, practices, experiences and results. The long term aim is to build a community and a body of knowledge within the intersection of RE and Testing. One of the main outputs of the 1st workshop was a collaboratively constructed map of the area of RET showing the topics relevant to RET for these. The 2nd workshop will continue in the same interactive vein and include a keynote, paper presentations with ample time for discussions, and a group exercise. For true impact and relevance this cross-cutting area requires contribution from both RE and Testing, and from both researchers and practitioners. For that reason we welcome a range of paper contributions from short experience papers to full research papers that both clearly cover connections between the two fields. Elizabeth Bjarnason, Mirko Morandini, Markus Borg, Michael Unterkalmsteiner, Michael Felderer, Matthew Staats |
ICSE (2) | 6 |
| 2015 | A Flexible and Non-intrusive Approach for Computing Complex Structural Coverage MetricsabstractSoftware analysis tools and techniques often leverage structural code coverage information to reason about the dynamic behavior of software. Existing techniques instrument the code with the required structural obligations and then monitor the execution of the compiled code to report coverage. Instrumentation based approaches often incur considerable runtime overhead for complex structural coverage metrics such as Modified Condition/Decision (MC/DC). Code instrumentation, in general, has to be approached with great care to ensure it does not modify the behavior of the original code. Furthermore, instrumented code cannot be used in conjunction with other analyses that reason about the structure and semantics of the code under test. In this work, we introduce a non-intrusive preprocessing approach for computing structural coverage information. It uses a static partial evaluation of the decisions in the source code and a source-to-bytecode mapping to generate the information necessary to efficiently track structural coverage metrics during execution. Our technique is flexible; the results of the preprocessing can be used by a variety of coverage-driven software analysis tasks, including automated analyses that are not possible for instrumented code. Experimental results in the context of symbolic execution show the efficiency and flexibility of our non- intrusive approach for computing code coverage information. Michael W. Whalen, Suzette Person, Neha Rungta, Matthew Staats, Daniela Grijincu |
ICSE (1) | 4 |
| 2015 | Are concurrency coverage metrics effective for testing: a comprehensive empirical investigationabstractSummary Testing multithreaded programs is inherently challenging, as programs can exhibit numerous thread interactions. To help engineers test these programs cost‐effectively, researchers have proposed concurrency coverage metrics. These metrics are intended to be used as predictors for testing effectiveness and provide targets for test generation. The effectiveness of these metrics, however, remains largely unexamined. In this work, we explore the impact of concurrency coverage metrics on testing effectiveness and examine the relationship between coverage, fault detection, and test suite size. We study eight existing concurrency coverage metrics and six new metrics formed by combining complementary metrics. Our results indicate that the metrics are moderate to strong predictors of testing effectiveness and effective at providing test generation targets. Nevertheless, metric effectiveness varies across programs, and even combinations of complementary metrics do not consistently provide effective testing. These results highlight the need for additional work on concurrency coverage metrics. Copyright © 2014 John Wiley & Sons, Ltd. Shin Hong, Matthew Staats, Moonzoo Kim, Gregg Rothermel |
Softw. Test. Verification Reliab. | 2 |
| 2015 | Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical StudyabstractWork on automated test generation has produced several tools capable of generating test data which achieves high structural coverage over a program. In the absence of a specification, developers are expected to manually construct or verify the test oracle for each test input. Nevertheless, it is assumed that these generated tests ease the task of testing for the developer, as testing is reduced to checking the results of tests. While this assumption has persisted for decades, there has been no conclusive evidence to date confirming it. However, the limited adoption in industry indicates this assumption may not be correct, and calls into question the practical value of test generation tools. To investigate this issue, we performed two controlled experiments comparing a total of 97 subjects split between writing tests manually and writing tests with the aid of an automated unit test generation tool, E vo S uite . We found that, on one hand, tool support leads to clear improvements in commonly applied quality metrics such as code coverage (up to 300% increase). However, on the other hand, there was no measurable improvement in the number of bugs actually found by developers. Our results not only cast some doubt on how the research community evaluates test generation tools, but also point to improvements and future work necessary before automated test generation tools will be widely adopted by practitioners. Gordon Fraser 0001, Matthew Staats, Phil McMinn, Andrea Arcuri, Frank Padberg |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2015 | The Risks of Coverage-Directed Test Case GenerationabstractA number of structural coverage criteria have been proposed to measure the adequacy of testing efforts. In the avionics and other critical systems domains, test suites satisfying structural coverage criteria are mandated by standards. With the advent of powerful automated test generation tools, it is tempting to simply generate test inputs to satisfy these structural coverage criteria. However, while techniques to produce coverage-providing tests are well established, the effectiveness of such approaches in terms of fault detection ability has not been adequately studied. In this work, we evaluate the effectiveness of test suites generated to satisfy four coverage criteria through counterexample-based test generation and a random generation approach-where tests are randomly generated until coverage is achieved-contrasted against purely random test suites of equal size. Our results yield three key conclusions. First, coverage criteria satisfaction alone can be a poor indication of fault finding effectiveness, with inconsistent results between the seven case examples (and random test suites of equal size often providing similar-or even higher-levels of fault finding). Second, the use of structural coverage as a supplement-rather than a target-for test generation can have a positive impact, with random test suites reduced to a coverage-providing subset detecting up to 13.5 percent more faults than test suites generated specifically to achieve coverage. Finally, Observable MC/DC, a criterion designed to account for program structure and the selection of the test oracle, can-in part-address the failings of traditional structural coverage criteria, allowing for the generation of test suites achieving higher levels of fault detection than random test suites of equal size. These observations point to risks inherent in the increase in test automation in critical systems, and the need for more research in how coverage criteria, test generation approaches, the test oracle used, and system structure jointly influence test effectiveness. Gregory Gay 0002, Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl |
IEEE Trans. Software Eng. | 2 |
| 2015 | Automated Oracle Data Selection SupportabstractThe choice of test oracle-the artifact that determines whether an application under test executes correctly-can significantly impact the effectiveness of the testing process. However, despite the prevalence of tools that support test input selection, little work exists for supporting oracle creation. We propose a method of supporting test oracle creation that automatically selects the oracle data-the set of variables monitored during testing-for expected value test oracles. This approach is based on the use of mutation analysis to rank variables in terms of fault-finding effectiveness, thus automating the selection of the oracle data. Experimental results obtained by employing our method over six industrial systems (while varying test input types and the number of generated mutants) indicate that our method-when paired with test inputs generated either at random or to satisfy specific structural coverage criteria-may be a cost-effective approach for producing small, effective oracle data sets, with fault finding improvements over current industrial best practice of up to 1,435 percent observed (with typical improvements of up to 50 percent). Gregory Gay 0002, Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl |
IEEE Trans. Software Eng. | 2 |
| 2015 | The Impact of View Histories on Edit RecommendationsabstractRecommendation systems are intended to increase developer productivity by recommending files to edit. These systems mine association rules in software revision histories. However, mining coarse-grained rules using only edit histories produces recommendations with low accuracy, and can only produce recommendations after a developer edits a file. In this work, we explore the use of finer-grained association rules, based on the insight that view histories help characterize the contexts of files to edit. To leverage this additional context and fine-grained association rules, we have developed MI, a recommendation system extending ROSE, an existing edit-based recommendation system. We then conducted a comparative simulation of ROSE and MI using the interaction histories stored in the Eclipse Bugzilla system. The simulation demonstrates that MI predicts the files to edit with significantly higher recommendation accuracy than ROSE (about 63 over 35 percent), and makes recommendations earlier, often before developers begin editing. Our results clearly demonstrate the value of considering both views and edits in systems to recommend files to edit, and results in more accurate, earlier, and more flexible recommendations. Seonah Lee 0001, Sungwon Kang, Sunghun Kim 0001, Matthew Staats |
IEEE Trans. Software Eng. | 4 |
| 2014 | Test Case Prioritization Based on Information Retrieval ConceptsabstractIn regression testing, running all a system's test cases can require a great deal of time and resources. Test case prioritization (TCP) attempts to schedule test cases to achieve goals such as higher coverage or faster fault detection. While code coverage-based approaches are typical in TCP, recent work has explored the use of additional information to improve effectiveness. In this work, we explore the use of Information Retrieval (IR) techniques to improve the effectiveness of TCP, particularly for testing infrequently tested code. Our approach considers the frequency at which elements have been tested, in additional to traditional coverage information, balancing these factors using linear regression modeling. Our empirical study demonstrates that our approach is generally more effective than both random and traditional code coverage-based approaches, with improvements in rate of fault detection of up to 4.7%. Jung-Hyun Kwon, In-Young Ko, Gregg Rothermel, Matthew Staats |
APSEC (1) | 4 |
| 2014 | Dodona: automated oracle data set selectionabstractSoftware complexity has increased the need for automated software testing. Most research on automating testing, however, has focused on creating test input data. While careful selection of input data is necessary to reach faulty states in a system under test, test oracles are needed to actually detect failures. In this work, we describe Dodona, a system that supports the generation of test oracles. Dodona ranks program variables based on the interactions and dependencies observed between them during program execution. Using this ranking, Dodona proposes a set of variables to be monitored, that can be used by engineers to construct assertion-based oracles. Our empirical study of Dodona reveals that it is more effective and efficient than the current state-of-the-art approach for generating oracle data sets, and can often yield oracles that are almost as effective as oracles hand-crafted by engineers without support. Pablo S. Loyola, Matthew Staats, In-Young Ko, Gregg Rothermel |
ISSTA | 2 |
| 2013 | NavClus: a graphical recommender for assisting code explorationabstractRecently, several graphical tools have been proposed to help developers avoid becoming disoriented when working with large software projects. These tools visualize the locations that developers have visited, allowing them to quickly recall where they have already visited. However, developers also spend a significant amount of time exploring source locations to visit, which is a task that is not currently supported by existing tools. In this work, we propose a graphical code recommender NavClus, which helps developers find relevant, unexplored source locations to visit. NavClus operates by mining a developer's daily interaction traces, comparing the developer's current working context with previously seen contexts, and then predicting relevant source locations to visit. These locations are displayed graphically along with the already explored locations in a class diagram. As a result, with NavClus developers can quickly find, reach, and focus on source locations relevant to their working contexts. http://www.youtube.com/watch?v=rbrc5ERyWjQ. Seonah Lee 0001, Sungwon Kang, Matthew Staats |
ICSE | 3 |
| 2013 | Observable modified Condition/Decision coverageabstractIn many critical systems domains, test suite adequacy is currently measured using structural coverage metrics over the source code. Of particular interest is the modified condition/decision coverage (MC/DC) criterion required for, e.g., critical avionics systems. In previous investigations we have found that the efficacy of such test suites is highly dependent on the structure of the program under test and the choice of variables monitored by the oracle. MC/DC adequate tests would frequently exercise faulty code, but the effects of the faults would not propagate to the monitored oracle variables. In this report, we combine the MC/DC coverage metric with a notion of observability that helps ensure that the result of a fault encountered when covering a structural obligation propagates to a monitored variable; we term this new coverage criterion Observable MC/DC (OMC/DC). We hypothesize this path requirement will make structural coverage metrics 1.) more effective at revealing faults, 2.) more robust to changes in program structure, and 3.) more robust to the choice of variables monitored. We assess the efficacy and sensitivity to program structure of OMC/DC as compared to masking MC/DC using four subsystems from the civil avionics domain and the control logic of a microwave. We have found that test suites satisfying OMC/DC are significantly more effective than test suites satisfying MC/DC, revealing up to 88% more faults, and are less sensitive to program structure and the choice of monitored variables. Michael W. Whalen, Gregory Gay 0002, Dongjiang You, Mats P. E. Heimdahl, Matthew Staats |
ICSE | 5 |
| 2013 | The Impact of Concurrent Coverage Metrics on Testing EffectivenessabstractWhen testing multithreaded programs, the number of possible thread interactions makes exploring all interactions infeasible in practice. In response, researchers have developed concurrent coverage metrics for multithreaded programs. These metrics allow them to estimate how well they have exercised concurrent program behavior, just as branch and statement coverage metrics do for sequential program testing. However, unlike sequential coverage metrics, the effectiveness of concurrent coverage metrics in testing remains largely unexamined. In this paper, we explore the relationship between concurrent coverage and fault detection effectiveness by studying the application of eight concurrent coverage metrics in testing nine concurrent programs. Our results show that existing concurrent coverage metrics are often moderate to strong predictors of concurrent testing effectiveness, and are generally reasonable targets for test suite generation. Nevertheless, their relative effectiveness as predictors and test generation targets varies across programs, and thus additional work is needed in this area. Shin Hong, Matthew Staats, Moonzoo Kim, Gregg Rothermel |
ICST | 2 |
| 2013 | Does automated white-box test generation really help software testers?abstractAutomated test generation techniques can efficiently produce test data that systematically cover structural aspects of a program. In the absence of a specification, a common assumption is that these tests relieve a developer of most of the work, as the act of testing is reduced to checking the results of the tests. Although this assumption has persisted for decades, there has been no conclusive evidence to date confirming it. However, the fact that the approach has only seen a limited uptake in industry suggests the contrary, and calls into question its practical usefulness. To investigate this issue, we performed a controlled experiment comparing a total of 49 subjects split between writing tests manually and writing tests with the aid of an automated unit test generation tool, EvoSuite. We found that, on one hand, tool support leads to clear improvements in commonly applied quality metrics such as code coverage (up to 300% increase). However, on the other hand, there was no measurable improvement in the number of bugs actually found by developers. Our results not only cast some doubt on how the research community evaluates test generation tools, but also point to improvements and future work necessary before automated test generation tools will be widely adopted by practitioners. Gordon Fraser 0001, Matthew Staats, Phil McMinn, Andrea Arcuri, Frank Padberg |
ISSTA | 2 |
| 2012 | On the Danger of Coverage Directed Test Case Generation
Matthew Staats, Gregory Gay 0002, Michael W. Whalen, Mats P. E. Heimdahl |
FASE | 1 |
| 2012 | Automated oracle creation support, or: How I learned to stop worrying about fault propagation and love mutation testingabstractIn testing, the test oracle is the artifact that determines whether an application under test executes correctly. The choice of test oracle can significantly impact the effectiveness of the testing process. However, despite the prevalence of tools that support the selection of test inputs, little work exists for supporting oracle creation. In this work, we propose a method of supporting test oracle creation. This method automatically selects the oracle data — the set of variables monitored during testing — for expected value test oracles. This approach is based on the use of mutation analysis to rank variables in terms of fault-finding effectiveness, thus automating the selection of the oracle data. Experiments over four industrial examples demonstrate that our method may be a cost-effective approach for producing small, effective oracle data, with fault finding improvements over current industrial best practice of up to 145.8% observed. Matthew Staats, Gregory Gay 0002, Mats P. E. Heimdahl |
ICSE | 1 |
| 2012 | Oracle-Centric Test Case PrioritizationabstractRecent work in testing has demonstrated the benefits of considering test oracles in the testing process. Unfortunately, this work has focused primarily on developing techniques for generating test oracles, in particular techniques based on mutation testing. While effective for test case generation, existing research has not considered the impact of test oracles in the context of regression testing tasks. Of interest here is the problem of test case prioritization, in which a set of test cases are ordered to attempt to detect faults earlier and to improve the effectiveness of testing when the entire set cannot be executed. In this work, we propose a technique for prioritizing test cases that explicitly takes into account the impact of test oracles on the effectiveness of testing. Our technique operates by first capturing the flow of information from variable assignments to test oracles for each test case, and then prioritizing to ``cover'' variables using the shortest paths possible to a test oracle. As a result, we favor test orderings in which many variables impact the test oracle's result early in test execution. Our results demonstrate improvements in rate of fault detection relative to both random and structural coverage based prioritization techniques when applied to faulty versions of three synchronous reactive systems. Matthew Staats, Pablo S. Loyola, Gregg Rothermel |
ISSRE | 1 |
| 2012 | Understanding user understanding: determining correctness of generated program invariantsabstractRecently, work has begun on automating the generation of test oracles, which are necessary to fully automate the testing process. One approach to such automation involves dynamic invariant generation which extracts invariants from program executions. To use such invariants as test oracles, however, it is necessary to distinguish correct from incorrect invariants, a process that currently requires human intervention. In this work we examine this process. In particular, we examine the ability of 30 users, across two empirical studies, to classify invariants generated from three Java programs. Our results indicate that users struggle to classify generated invariants: on average, they misclassify 9.1% to 31.7% of correct invariants and 26.1%-58.6% of incorrect invariants. These results contradict prior studies that suggest that classification by users is easy, and indicate that further work needs to be done to bridge the gap between the effectiveness of dynamic invariant generation in theory, and the ability of users to apply it in practice. Along these lines, we suggest several areas for future work. Matthew Staats, Shin Hong, Moonzoo Kim, Gregg Rothermel |
ISSTA | 1 |
| 2011 | Programs, tests, and oracles: the foundations of testing revisitedabstractIn previous decades, researchers have explored the formal foundations of program testing. By exploring the foundations of testing largely separate from any specific method of testing, these researchers provided a general discussion of the testing process, including the goals, the underlying problems, and the limitations of testing. Unfortunately, a common, rigorous foundation has not been widely adopted in empirical software testing research, making it difficult to generalize and compare empirical research. Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl |
ICSE | 1 |
| 2011 | Better testing through oracle selectionabstractIn software testing, the test oracle determines if the application under test has performed an execution correctly. In current testing practice and research, significant effort and thought is placed on selecting test inputs, with the selection of test oracles largely neglected. Here, we argue that improvements to the testing process can be made by considering the problem of oracle selection. In particular, we argue that selecting the test oracle and test inputs together to complement one another may yield improvements testing effectiveness. We illustrate this using an example and present selected results from an ongoing study demonstrating the relationship between test suite selection, oracle selection, and fault finding. Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl |
ICSE | 1 |
| 2010 | Parallel symbolic execution for structural test generationabstractSymbolic execution is a popular technique for automatically generating test cases achieving high structural coverage. Symbolic execution suffers from scalability issues since the number of symbolic paths that need to be explored is very large (or even infinite) for most realistic programs. To address this problem, we propose a technique, Simple Static Partitioning, for parallelizing symbolic execution. The technique uses a set of pre-conditions to partition the symbolic execution tree, allowing us to effectively distribute symbolic execution and decrease the time needed to explore the symbolic execution tree. The proposed technique requires little communication between parallel instances and is designed to work with a variety of architectures, ranging from fast multi-core machines to cloud or grid computing environments. We implement our technique in the Java PathFinder verification tool-set and evaluate it on six case studies with respect to the performance improvement when exploring a finite symbolic execution tree and performing automatic test generation. Matthew Staats, Corina Pasareanu |
ISSTA | 1 |
| 2010 | The influence of multiple artifacts on the effectiveness of software testingabstractThe effectiveness of the software testing process is determined by artifacts used in testing, including the program, the set of tests, and the test oracle. However, in evaluating software testing techniques, including automated software testing techniques, the influence of these testing artifacts is often overlooked. In my upcoming dissertation, we intend to explore the interrelationship between these three testing artifacts, with the goal of establishing a solid scientific foundation for understanding how they interact. We plan to provide two contributions towards this goal. First, we propose a theoretical framework for discussing testing based on previous work in the theory of testing. Second, we intend to perform a rigorous empirical study controlling for program structure, test coverage criteria, and oracle selection in the domain of safety critical avionics software. Matthew Staats |
ASE | 1 |
| 2008 | Requirements Coverage as an Adequacy Measure for Conformance Testing
Ajitha Rajan, Michael W. Whalen, Matthew Staats, Mats P. E. Heimdahl |
ICFEM | 3 |
| 2008 | Partial Translation Verification for Untrusted Code-Generators
Matthew Staats, Mats P. E. Heimdahl |
ICFEM | 1 |
| 2008 | ReqsCov: A Tool for Measuring Test-Adequacy over RequirementsabstractWhen creating test cases for software, a common approach is to create tests that exercise requirements. Determining the adequacy of test cases, however, is generally done through inspection or indirectly by measuring structural coverage of an executable artifact (such as source code or a software model). We present ReqsCov, a tool to directly measure requirements coverage provided by test cases. ReqsCov allows users to measure Linear Temporal Logic requirements coverage using three increasingly rigorous requirements coverage metrics: naive coverage, antecedent coverage, and Unique First Cause coverage. By measuring requirements coverage, users are given insight into the quality of test suites beyond what is available when solely using structural coverage metrics over an implementation. Matthew Staats, Weijia Deng, Ajitha Rajan, Mats P. E. Heimdahl, Kurt Woodham |
ASE | 1 |
| 2008 | Breaking and Provably Fixing Minx
Erik Shimshock, Matthew Staats, Nicholas Hopper |
Privacy Enhancing Technologies | 2 |