Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Matthew Staats

dblp:88/6925 · also Matt Staats · DBLP profile ↗
← Back
26ranked-venue papers
10as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 25 · 10 first-authorSecurity and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
18 papers
Software testing · 76% Empirical software engineering · 10% Program analysis · 7%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software testing
test oracle
0.962015
Automated Oracle Data Selection Support · IEEE Trans. Software Eng. 2015
Dodona: automated oracle data set selection · ISSTA 2014
Automated oracle creation support, or: How I learned to stop worrying about fault propagation and love mutation testing · ICSE 2012
Software testing › test oracle
oracle data selection
0.632015
Automated Oracle Data Selection Support · IEEE Trans. Software Eng. 2015
Dodona: automated oracle data set selection · ISSTA 2014
Automated oracle creation support, or: How I learned to stop worrying about fault propagation and love mutation testing · ICSE 2012
Software testing
test generation
0.532015
The Risks of Coverage-Directed Test Case Generation · IEEE Trans. Software Eng. 2015
Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study · ACM Trans. Softw. Eng. Methodol. 2015
Parallel symbolic execution for structural test generation · ISSTA 2010
Software testing › test adequacy › coverage criteria
structural coverage criteria
0.542015
A Flexible and Non-intrusive Approach for Computing Complex Structural Coverage Metrics · ICSE (1) 2015
Observable modified Condition/Decision coverage · ICSE 2013
ReqsCov: A Tool for Measuring Test-Adequacy over Requirements · ASE 2008
Software testing › test adequacy › coverage criteria › structural coverage criteria
MC/DC coverage
0.422016
The Effect of Program and Model Structure on the Effectiveness of MC/DC Test Adequacy Coverage · ACM Trans. Softw. Eng. Methodol. 2016
Observable modified Condition/Decision coverage · ICSE 2013
Empirical software engineering
developer studies
0.432015
Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study · ACM Trans. Softw. Eng. Methodol. 2015
Understanding user understanding: determining correctness of generated program invariants · ISSTA 2012
Does automated white-box test generation really help software testers? · ISSTA 2013
Software testing
test adequacy
0.322016
The Effect of Program and Model Structure on the Effectiveness of MC/DC Test Adequacy Coverage · ACM Trans. Softw. Eng. Methodol. 2016
ReqsCov: A Tool for Measuring Test-Adequacy over Requirements · ASE 2008
Software testing › test adequacy
coverage criteria
0.212016
The Effect of Program and Model Structure on the Effectiveness of MC/DC Test Adequacy Coverage · ACM Trans. Softw. Eng. Methodol. 2016
Software testing › test generation
coverage-based test generation
0.212015
The Risks of Coverage-Directed Test Case Generation · IEEE Trans. Software Eng. 2015
Empirical software engineering
mining software repositories
0.212015
The Impact of View Histories on Edit Recommendations · IEEE Trans. Software Eng. 2015
Software testing › test generation
unit test generation
0.212015
Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study · ACM Trans. Softw. Eng. Methodol. 2015
Software testing › test generation
automated test generation
0.212013
Does automated white-box test generation really help software testers? · ISSTA 2013
Software maintenance and evolution › program comprehension
code exploration
0.212013
NavClus: a graphical recommender for assisting code exploration · ICSE 2013
Software testing › test generation
white-box test generation
0.212013
Does automated white-box test generation really help software testers? · ISSTA 2013
Program analysis › specification mining
dynamic invariant detection
0.112012
Understanding user understanding: determining correctness of generated program invariants · ISSTA 2012
Software testing
mutation testing
0.112012
Automated oracle creation support, or: How I learned to stop worrying about fault propagation and love mutation testing · ICSE 2012
Software testing › test oracle
test oracle generation
0.112012
Understanding user understanding: determining correctness of generated program invariants · ISSTA 2012
Software testing › test optimization
test case selection
0.112011
Better testing through oracle selection · ICSE 2011
Software testing
testing theory
0.112011
Programs, tests, and oracles: the foundations of testing revisited · ICSE 2011
Empirical software engineering
controlled experiment
0.112010
The influence of multiple artifacts on the effectiveness of software testing · ASE 2010
Program analysis › symbolic execution
parallel symbolic execution
0.112010
Parallel symbolic execution for structural test generation · ISSTA 2010
Software testing › structural testing
structural test generation
0.112010
Parallel symbolic execution for structural test generation · ISSTA 2010
Program analysis
symbolic execution
0.112010
Parallel symbolic execution for structural test generation · ISSTA 2010
Software testing › test coverage
code coverage
0.112015
Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study · ACM Trans. Softw. Eng. Methodol. 2015
Compilers and program optimization
partial evaluation
0.112015
A Flexible and Non-intrusive Approach for Computing Complex Structural Coverage Metrics · ICSE (1) 2015
Program analysis
static analysis
0.112015
A Flexible and Non-intrusive Approach for Computing Complex Structural Coverage Metrics · ICSE (1) 2015
Software testing
test coverage
0.112015
Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study · ACM Trans. Softw. Eng. Methodol. 2015
Software testing › test process
test design
0.112015
2nd International Workshop on Requirements Engineering and Testing (RET 2015) · ICSE (2) 2015
Program analysis
dynamic analysis
0.112014
Dodona: automated oracle data set selection · ISSTA 2014
Recommender systems › domain-specific recommendation
software engineering recommender systems
0.012013
NavClus: a graphical recommender for assisting code exploration · ICSE 2013

Methods — techniques the papers use, named apart from their topics

mutation analysis · 0.6symbolic execution · 0.4controlled experiment · 0.4empirical study · 0.4simulation · 0.2random test generation · 0.2fault-finding effectiveness ranking · 0.2counterexample-based test generation · 0.2association rule mining · 0.2program execution analysis · 0.2interaction trace mining · 0.2static partitioning · 0.1
YearPublicationVenuePosition
2016 The Effect of Program and Model Structure on the Effectiveness of MC/DC Test Adequacy Coverage
abstract
Test adequacy metrics defined over the structure of a program, such as Modified Condition and Decision Coverage (MC/DC), are used to assess testing efforts. However, MC/DC can be “cheated” by restructuring a program to make it easier to achieve the desired coverage. This is concerning, given the importance of MC/DC in assessing the adequacy of test suites for critical systems domains. In this work, we have explored the impact of implementation structure on the efficacy of test suites satisfying the MC/DC criterion using four real-world avionics systems. Our results demonstrate that test suites achieving MC/DC over implementations with structurally complex Boolean expressions are generally larger and more effective than test suites achieving MC/DC over functionally equivalent, but structurally simpler, implementations. Additionally, we found that test suites generated over simpler implementations achieve significantly lower MC/DC and fault-finding effectiveness when applied to complex implementations, whereas test suites generated over the complex implementation still achieve high MC/DC and attain high fault finding over the simpler implementation. By measuring MC/DC over simple implementations, we can significantly reduce the cost of testing, but in doing so, we also reduce the effectiveness of the testing process. Thus, developers have an economic incentive to “cheat” the MC/DC criterion, but this cheating leads to negative consequences. Accordingly, we recommend that organizations require MC/DC over a structurally complex implementation for testing purposes to avoid these consequences.
Gregory Gay 0002, Ajitha Rajan, Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl
ACM Trans. Softw. Eng. Methodol.3
2015 2nd International Workshop on Requirements Engineering and Testing (RET 2015)
abstract
The RET (Requirements Engineering and Testing) workshop provides a meeting point for researchers and practitioners from the two separate fields of Requirements Engineering (RE) and Testing. The goal is to improve the connection and alignment of these two areas through an exchange of ideas, challenges, practices, experiences and results. The long term aim is to build a community and a body of knowledge within the intersection of RE and Testing. One of the main outputs of the 1st workshop was a collaboratively constructed map of the area of RET showing the topics relevant to RET for these. The 2nd workshop will continue in the same interactive vein and include a keynote, paper presentations with ample time for discussions, and a group exercise. For true impact and relevance this cross-cutting area requires contribution from both RE and Testing, and from both researchers and practitioners. For that reason we welcome a range of paper contributions from short experience papers to full research papers that both clearly cover connections between the two fields.
Elizabeth Bjarnason, Mirko Morandini, Markus Borg, Michael Unterkalmsteiner, Michael Felderer, Matthew Staats
ICSE (2)6
2015 A Flexible and Non-intrusive Approach for Computing Complex Structural Coverage Metrics
abstract
Software analysis tools and techniques often leverage structural code coverage information to reason about the dynamic behavior of software. Existing techniques instrument the code with the required structural obligations and then monitor the execution of the compiled code to report coverage. Instrumentation based approaches often incur considerable runtime overhead for complex structural coverage metrics such as Modified Condition/Decision (MC/DC). Code instrumentation, in general, has to be approached with great care to ensure it does not modify the behavior of the original code. Furthermore, instrumented code cannot be used in conjunction with other analyses that reason about the structure and semantics of the code under test. In this work, we introduce a non-intrusive preprocessing approach for computing structural coverage information. It uses a static partial evaluation of the decisions in the source code and a source-to-bytecode mapping to generate the information necessary to efficiently track structural coverage metrics during execution. Our technique is flexible; the results of the preprocessing can be used by a variety of coverage-driven software analysis tasks, including automated analyses that are not possible for instrumented code. Experimental results in the context of symbolic execution show the efficiency and flexibility of our non- intrusive approach for computing code coverage information.
Michael W. Whalen, Suzette Person, Neha Rungta, Matthew Staats, Daniela Grijincu
ICSE (1)4
2015 Are concurrency coverage metrics effective for testing: a comprehensive empirical investigation
abstract
Summary Testing multithreaded programs is inherently challenging, as programs can exhibit numerous thread interactions. To help engineers test these programs cost‐effectively, researchers have proposed concurrency coverage metrics. These metrics are intended to be used as predictors for testing effectiveness and provide targets for test generation. The effectiveness of these metrics, however, remains largely unexamined. In this work, we explore the impact of concurrency coverage metrics on testing effectiveness and examine the relationship between coverage, fault detection, and test suite size. We study eight existing concurrency coverage metrics and six new metrics formed by combining complementary metrics. Our results indicate that the metrics are moderate to strong predictors of testing effectiveness and effective at providing test generation targets. Nevertheless, metric effectiveness varies across programs, and even combinations of complementary metrics do not consistently provide effective testing. These results highlight the need for additional work on concurrency coverage metrics. Copyright © 2014 John Wiley & Sons, Ltd.
Shin Hong, Matthew Staats, Moonzoo Kim, Gregg Rothermel
Softw. Test. Verification Reliab.2
2015 Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study
abstract
Work on automated test generation has produced several tools capable of generating test data which achieves high structural coverage over a program. In the absence of a specification, developers are expected to manually construct or verify the test oracle for each test input. Nevertheless, it is assumed that these generated tests ease the task of testing for the developer, as testing is reduced to checking the results of tests. While this assumption has persisted for decades, there has been no conclusive evidence to date confirming it. However, the limited adoption in industry indicates this assumption may not be correct, and calls into question the practical value of test generation tools. To investigate this issue, we performed two controlled experiments comparing a total of 97 subjects split between writing tests manually and writing tests with the aid of an automated unit test generation tool, E vo S uite . We found that, on one hand, tool support leads to clear improvements in commonly applied quality metrics such as code coverage (up to 300% increase). However, on the other hand, there was no measurable improvement in the number of bugs actually found by developers. Our results not only cast some doubt on how the research community evaluates test generation tools, but also point to improvements and future work necessary before automated test generation tools will be widely adopted by practitioners.
Gordon Fraser 0001, Matthew Staats, Phil McMinn, Andrea Arcuri, Frank Padberg
ACM Trans. Softw. Eng. Methodol.2
2015 The Risks of Coverage-Directed Test Case Generation
abstract
A number of structural coverage criteria have been proposed to measure the adequacy of testing efforts. In the avionics and other critical systems domains, test suites satisfying structural coverage criteria are mandated by standards. With the advent of powerful automated test generation tools, it is tempting to simply generate test inputs to satisfy these structural coverage criteria. However, while techniques to produce coverage-providing tests are well established, the effectiveness of such approaches in terms of fault detection ability has not been adequately studied. In this work, we evaluate the effectiveness of test suites generated to satisfy four coverage criteria through counterexample-based test generation and a random generation approach-where tests are randomly generated until coverage is achieved-contrasted against purely random test suites of equal size. Our results yield three key conclusions. First, coverage criteria satisfaction alone can be a poor indication of fault finding effectiveness, with inconsistent results between the seven case examples (and random test suites of equal size often providing similar-or even higher-levels of fault finding). Second, the use of structural coverage as a supplement-rather than a target-for test generation can have a positive impact, with random test suites reduced to a coverage-providing subset detecting up to 13.5 percent more faults than test suites generated specifically to achieve coverage. Finally, Observable MC/DC, a criterion designed to account for program structure and the selection of the test oracle, can-in part-address the failings of traditional structural coverage criteria, allowing for the generation of test suites achieving higher levels of fault detection than random test suites of equal size. These observations point to risks inherent in the increase in test automation in critical systems, and the need for more research in how coverage criteria, test generation approaches, the test oracle used, and system structure jointly influence test effectiveness.
Gregory Gay 0002, Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl
IEEE Trans. Software Eng.2
2015 Automated Oracle Data Selection Support
abstract
The choice of test oracle-the artifact that determines whether an application under test executes correctly-can significantly impact the effectiveness of the testing process. However, despite the prevalence of tools that support test input selection, little work exists for supporting oracle creation. We propose a method of supporting test oracle creation that automatically selects the oracle data-the set of variables monitored during testing-for expected value test oracles. This approach is based on the use of mutation analysis to rank variables in terms of fault-finding effectiveness, thus automating the selection of the oracle data. Experimental results obtained by employing our method over six industrial systems (while varying test input types and the number of generated mutants) indicate that our method-when paired with test inputs generated either at random or to satisfy specific structural coverage criteria-may be a cost-effective approach for producing small, effective oracle data sets, with fault finding improvements over current industrial best practice of up to 1,435 percent observed (with typical improvements of up to 50 percent).
Gregory Gay 0002, Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl
IEEE Trans. Software Eng.2
2015 The Impact of View Histories on Edit Recommendations
abstract
Recommendation systems are intended to increase developer productivity by recommending files to edit. These systems mine association rules in software revision histories. However, mining coarse-grained rules using only edit histories produces recommendations with low accuracy, and can only produce recommendations after a developer edits a file. In this work, we explore the use of finer-grained association rules, based on the insight that view histories help characterize the contexts of files to edit. To leverage this additional context and fine-grained association rules, we have developed MI, a recommendation system extending ROSE, an existing edit-based recommendation system. We then conducted a comparative simulation of ROSE and MI using the interaction histories stored in the Eclipse Bugzilla system. The simulation demonstrates that MI predicts the files to edit with significantly higher recommendation accuracy than ROSE (about 63 over 35 percent), and makes recommendations earlier, often before developers begin editing. Our results clearly demonstrate the value of considering both views and edits in systems to recommend files to edit, and results in more accurate, earlier, and more flexible recommendations.
Seonah Lee 0001, Sungwon Kang, Sunghun Kim 0001, Matthew Staats
IEEE Trans. Software Eng.4
2014 Test Case Prioritization Based on Information Retrieval Concepts
abstract
In regression testing, running all a system's test cases can require a great deal of time and resources. Test case prioritization (TCP) attempts to schedule test cases to achieve goals such as higher coverage or faster fault detection. While code coverage-based approaches are typical in TCP, recent work has explored the use of additional information to improve effectiveness. In this work, we explore the use of Information Retrieval (IR) techniques to improve the effectiveness of TCP, particularly for testing infrequently tested code. Our approach considers the frequency at which elements have been tested, in additional to traditional coverage information, balancing these factors using linear regression modeling. Our empirical study demonstrates that our approach is generally more effective than both random and traditional code coverage-based approaches, with improvements in rate of fault detection of up to 4.7%.
Jung-Hyun Kwon, In-Young Ko, Gregg Rothermel, Matthew Staats
APSEC (1)4
2014 Dodona: automated oracle data set selection
abstract
Software complexity has increased the need for automated software testing. Most research on automating testing, however, has focused on creating test input data. While careful selection of input data is necessary to reach faulty states in a system under test, test oracles are needed to actually detect failures. In this work, we describe Dodona, a system that supports the generation of test oracles. Dodona ranks program variables based on the interactions and dependencies observed between them during program execution. Using this ranking, Dodona proposes a set of variables to be monitored, that can be used by engineers to construct assertion-based oracles. Our empirical study of Dodona reveals that it is more effective and efficient than the current state-of-the-art approach for generating oracle data sets, and can often yield oracles that are almost as effective as oracles hand-crafted by engineers without support.
Pablo S. Loyola, Matthew Staats, In-Young Ko, Gregg Rothermel
ISSTA2
2013 NavClus: a graphical recommender for assisting code exploration
abstract
Recently, several graphical tools have been proposed to help developers avoid becoming disoriented when working with large software projects. These tools visualize the locations that developers have visited, allowing them to quickly recall where they have already visited. However, developers also spend a significant amount of time exploring source locations to visit, which is a task that is not currently supported by existing tools. In this work, we propose a graphical code recommender NavClus, which helps developers find relevant, unexplored source locations to visit. NavClus operates by mining a developer's daily interaction traces, comparing the developer's current working context with previously seen contexts, and then predicting relevant source locations to visit. These locations are displayed graphically along with the already explored locations in a class diagram. As a result, with NavClus developers can quickly find, reach, and focus on source locations relevant to their working contexts. http://www.youtube.com/watch?v=rbrc5ERyWjQ.
Seonah Lee 0001, Sungwon Kang, Matthew Staats
ICSE3
2013 Observable modified Condition/Decision coverage
abstract
In many critical systems domains, test suite adequacy is currently measured using structural coverage metrics over the source code. Of particular interest is the modified condition/decision coverage (MC/DC) criterion required for, e.g., critical avionics systems. In previous investigations we have found that the efficacy of such test suites is highly dependent on the structure of the program under test and the choice of variables monitored by the oracle. MC/DC adequate tests would frequently exercise faulty code, but the effects of the faults would not propagate to the monitored oracle variables. In this report, we combine the MC/DC coverage metric with a notion of observability that helps ensure that the result of a fault encountered when covering a structural obligation propagates to a monitored variable; we term this new coverage criterion Observable MC/DC (OMC/DC). We hypothesize this path requirement will make structural coverage metrics 1.) more effective at revealing faults, 2.) more robust to changes in program structure, and 3.) more robust to the choice of variables monitored. We assess the efficacy and sensitivity to program structure of OMC/DC as compared to masking MC/DC using four subsystems from the civil avionics domain and the control logic of a microwave. We have found that test suites satisfying OMC/DC are significantly more effective than test suites satisfying MC/DC, revealing up to 88% more faults, and are less sensitive to program structure and the choice of monitored variables.
Michael W. Whalen, Gregory Gay 0002, Dongjiang You, Mats P. E. Heimdahl, Matthew Staats
ICSE5
2013 The Impact of Concurrent Coverage Metrics on Testing Effectiveness
abstract
When testing multithreaded programs, the number of possible thread interactions makes exploring all interactions infeasible in practice. In response, researchers have developed concurrent coverage metrics for multithreaded programs. These metrics allow them to estimate how well they have exercised concurrent program behavior, just as branch and statement coverage metrics do for sequential program testing. However, unlike sequential coverage metrics, the effectiveness of concurrent coverage metrics in testing remains largely unexamined. In this paper, we explore the relationship between concurrent coverage and fault detection effectiveness by studying the application of eight concurrent coverage metrics in testing nine concurrent programs. Our results show that existing concurrent coverage metrics are often moderate to strong predictors of concurrent testing effectiveness, and are generally reasonable targets for test suite generation. Nevertheless, their relative effectiveness as predictors and test generation targets varies across programs, and thus additional work is needed in this area.
Shin Hong, Matthew Staats, Moonzoo Kim, Gregg Rothermel
ICST2
2013 Does automated white-box test generation really help software testers?
abstract
Automated test generation techniques can efficiently produce test data that systematically cover structural aspects of a program. In the absence of a specification, a common assumption is that these tests relieve a developer of most of the work, as the act of testing is reduced to checking the results of the tests. Although this assumption has persisted for decades, there has been no conclusive evidence to date confirming it. However, the fact that the approach has only seen a limited uptake in industry suggests the contrary, and calls into question its practical usefulness. To investigate this issue, we performed a controlled experiment comparing a total of 49 subjects split between writing tests manually and writing tests with the aid of an automated unit test generation tool, EvoSuite. We found that, on one hand, tool support leads to clear improvements in commonly applied quality metrics such as code coverage (up to 300% increase). However, on the other hand, there was no measurable improvement in the number of bugs actually found by developers. Our results not only cast some doubt on how the research community evaluates test generation tools, but also point to improvements and future work necessary before automated test generation tools will be widely adopted by practitioners.
Gordon Fraser 0001, Matthew Staats, Phil McMinn, Andrea Arcuri, Frank Padberg
ISSTA2
2012 On the Danger of Coverage Directed Test Case Generation
Matthew Staats, Gregory Gay 0002, Michael W. Whalen, Mats P. E. Heimdahl
FASE1
2012 Automated oracle creation support, or: How I learned to stop worrying about fault propagation and love mutation testing
abstract
In testing, the test oracle is the artifact that determines whether an application under test executes correctly. The choice of test oracle can significantly impact the effectiveness of the testing process. However, despite the prevalence of tools that support the selection of test inputs, little work exists for supporting oracle creation. In this work, we propose a method of supporting test oracle creation. This method automatically selects the oracle data — the set of variables monitored during testing — for expected value test oracles. This approach is based on the use of mutation analysis to rank variables in terms of fault-finding effectiveness, thus automating the selection of the oracle data. Experiments over four industrial examples demonstrate that our method may be a cost-effective approach for producing small, effective oracle data, with fault finding improvements over current industrial best practice of up to 145.8% observed.
Matthew Staats, Gregory Gay 0002, Mats P. E. Heimdahl
ICSE1
2012 Oracle-Centric Test Case Prioritization
abstract
Recent work in testing has demonstrated the benefits of considering test oracles in the testing process. Unfortunately, this work has focused primarily on developing techniques for generating test oracles, in particular techniques based on mutation testing. While effective for test case generation, existing research has not considered the impact of test oracles in the context of regression testing tasks. Of interest here is the problem of test case prioritization, in which a set of test cases are ordered to attempt to detect faults earlier and to improve the effectiveness of testing when the entire set cannot be executed. In this work, we propose a technique for prioritizing test cases that explicitly takes into account the impact of test oracles on the effectiveness of testing. Our technique operates by first capturing the flow of information from variable assignments to test oracles for each test case, and then prioritizing to ``cover'' variables using the shortest paths possible to a test oracle. As a result, we favor test orderings in which many variables impact the test oracle's result early in test execution. Our results demonstrate improvements in rate of fault detection relative to both random and structural coverage based prioritization techniques when applied to faulty versions of three synchronous reactive systems.
Matthew Staats, Pablo S. Loyola, Gregg Rothermel
ISSRE1
2012 Understanding user understanding: determining correctness of generated program invariants
abstract
Recently, work has begun on automating the generation of test oracles, which are necessary to fully automate the testing process. One approach to such automation involves dynamic invariant generation which extracts invariants from program executions. To use such invariants as test oracles, however, it is necessary to distinguish correct from incorrect invariants, a process that currently requires human intervention. In this work we examine this process. In particular, we examine the ability of 30 users, across two empirical studies, to classify invariants generated from three Java programs. Our results indicate that users struggle to classify generated invariants: on average, they misclassify 9.1% to 31.7% of correct invariants and 26.1%-58.6% of incorrect invariants. These results contradict prior studies that suggest that classification by users is easy, and indicate that further work needs to be done to bridge the gap between the effectiveness of dynamic invariant generation in theory, and the ability of users to apply it in practice. Along these lines, we suggest several areas for future work.
Matthew Staats, Shin Hong, Moonzoo Kim, Gregg Rothermel
ISSTA1
2011 Programs, tests, and oracles: the foundations of testing revisited
abstract
In previous decades, researchers have explored the formal foundations of program testing. By exploring the foundations of testing largely separate from any specific method of testing, these researchers provided a general discussion of the testing process, including the goals, the underlying problems, and the limitations of testing. Unfortunately, a common, rigorous foundation has not been widely adopted in empirical software testing research, making it difficult to generalize and compare empirical research.
Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl
ICSE1
2011 Better testing through oracle selection
abstract
In software testing, the test oracle determines if the application under test has performed an execution correctly. In current testing practice and research, significant effort and thought is placed on selecting test inputs, with the selection of test oracles largely neglected. Here, we argue that improvements to the testing process can be made by considering the problem of oracle selection. In particular, we argue that selecting the test oracle and test inputs together to complement one another may yield improvements testing effectiveness. We illustrate this using an example and present selected results from an ongoing study demonstrating the relationship between test suite selection, oracle selection, and fault finding.
Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl
ICSE1
2010 Parallel symbolic execution for structural test generation
abstract
Symbolic execution is a popular technique for automatically generating test cases achieving high structural coverage. Symbolic execution suffers from scalability issues since the number of symbolic paths that need to be explored is very large (or even infinite) for most realistic programs. To address this problem, we propose a technique, Simple Static Partitioning, for parallelizing symbolic execution. The technique uses a set of pre-conditions to partition the symbolic execution tree, allowing us to effectively distribute symbolic execution and decrease the time needed to explore the symbolic execution tree. The proposed technique requires little communication between parallel instances and is designed to work with a variety of architectures, ranging from fast multi-core machines to cloud or grid computing environments. We implement our technique in the Java PathFinder verification tool-set and evaluate it on six case studies with respect to the performance improvement when exploring a finite symbolic execution tree and performing automatic test generation.
Matthew Staats, Corina Pasareanu
ISSTA1
2010 The influence of multiple artifacts on the effectiveness of software testing
abstract
The effectiveness of the software testing process is determined by artifacts used in testing, including the program, the set of tests, and the test oracle. However, in evaluating software testing techniques, including automated software testing techniques, the influence of these testing artifacts is often overlooked. In my upcoming dissertation, we intend to explore the interrelationship between these three testing artifacts, with the goal of establishing a solid scientific foundation for understanding how they interact. We plan to provide two contributions towards this goal. First, we propose a theoretical framework for discussing testing based on previous work in the theory of testing. Second, we intend to perform a rigorous empirical study controlling for program structure, test coverage criteria, and oracle selection in the domain of safety critical avionics software.
Matthew Staats
ASE1
2008 Requirements Coverage as an Adequacy Measure for Conformance Testing
Ajitha Rajan, Michael W. Whalen, Matthew Staats, Mats P. E. Heimdahl
ICFEM3
2008 Partial Translation Verification for Untrusted Code-Generators
Matthew Staats, Mats P. E. Heimdahl
ICFEM1
2008 ReqsCov: A Tool for Measuring Test-Adequacy over Requirements
abstract
When creating test cases for software, a common approach is to create tests that exercise requirements. Determining the adequacy of test cases, however, is generally done through inspection or indirectly by measuring structural coverage of an executable artifact (such as source code or a software model). We present ReqsCov, a tool to directly measure requirements coverage provided by test cases. ReqsCov allows users to measure Linear Temporal Logic requirements coverage using three increasingly rigorous requirements coverage metrics: naive coverage, antecedent coverage, and Unique First Cause coverage. By measuring requirements coverage, users are given insight into the quality of test suites beyond what is available when solely using structural coverage metrics over an implementation.
Matthew Staats, Weijia Deng, Ajitha Rajan, Mats P. E. Heimdahl, Kurt Woodham
ASE1
2008 Breaking and Provably Fixing Minx
Erik Shimshock, Matthew Staats, Nicholas Hopper
Privacy Enhancing Technologies2