VLDB 2026 Research / reviewers in the wild / expert
Gregory Gay 0002
dblp:39/7539 · also Greg Gay 0002
· DBLP profile ↗
51ranked-venue papers
18as first author
21since 2021 · last 2026
0000-0001-6794-9585ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 50 · 18 first-author · 20 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Logs to Lessons: An Exploration of LLM-based Log Summarization for Debugging Automotive SoftwareabstractIdentifying where faults occur is an essential part of debugging, yet examining extensive system logs can be slow and mentally demanding, especially in complex software environments. One emerging strategy to enhance log analysis is to employ large language models (LLMs) to distill log information into more manageable summaries that can guide human reasoning during diagnosis. We report on a case study carried out in an automotive setting, where engineers investigated actual failures with and without support from an LLM-based summarization tool. During fault localization sessions where participants analyzed real failure logs, we collected cognitive load measurements, observed their reasoning processes, and gathered feedback on both the LLM-based summarization and the workflow through post-session interviews. Our results indicate that although the use of summaries raised certain cognitive demands, particularly related to mental effort and time pressure, participants experienced less frustration overall and considered the support helpful in focusing their attention. They also expressed a clear interest in being able to shape and refine summaries as their understanding evolved. These findings offer insights into how LLM-generated summaries influence practitioners’ diagnostic work and point toward the need for more adaptive, interactive, and workflow-aware support. Anton Ekström, Hampus Rhedin Stam, Francisco Gomes de Oliveira Neto, Gregory Gay 0002, Sabina Edenlund |
AST | 4 |
| 2025 | An intelligent test management system for optimizing decision making during software testing
Albin Lönnfält, Viktor Tu, Gregory Gay 0002, Sahar Tahvili |
J. Syst. Softw. | 3 |
| 2024 | Exploring the Role of Automation in Duplicate Bug Report Detection: An Industrial Case StudyabstractDuplicate bug reports can increase technical debt and tester work-load in long-running software projects. Many automated techniques have been proposed to detect potential duplicate reports. However, such techniques have not seen widespread industrial adoption. Our objective in this study is to better understand how automated techniques could effectively be employed within a tester's duplicate detection workflow. We are particularly interested in exploring the potential of a human-in-the-loop scenario where tools and humans work together to make duplicate determinations. Malte Götharsson, Karl Stahre, Gregory Gay 0002, Francisco Gomes de Oliveira Neto |
AST | 3 |
| 2024 | Message from the Program Co-Chairs; ICST 2024abstractIt is our distinct pleasure to welcome you to the proceedings of the 17thIEEE International Conference on Software Testing, Verification, and Validation (ICST 2024) - the premier forum for scientific research on topics ranging from automated test generation, to software reliability, to formal verification and model checking, to studies on the human aspects of the quality assurance process. Gregory Gay 0002, Shiva Nejati 0001 |
ICST | 1 |
| 2024 | Integrating Mutation Testing Into Developer Workflow: An Industrial Case StudyabstractMutation testing is a potentially effective method to assess test suite adequacy. Researchers have made mutation testing more computationally efficient, and new frameworks are regularly emerging. However, there is still limited adoption of mutation testing in industry. We hypothesize that such adoption is hindered by a lack of guidance on how to effectively and efficiently utilize mutation testing in a development workflow. To that end, we have conducted an industrial case study exploring the technical challenges of implementing mutation testing in continuous integration, what information from mutation testing is of use to developers, and how that information should be presented (in textual and visual form). Our results reveal five technical challenges of integrating mutation testing and nine key findings regarding how the results of mutation testing are used and presented. We also offer a dashboard to visualize mutation testing results, as well as 16 recommendations for making effective use of mutation testing in practice1. Stefan Alexander van Heijningen, Theo Wiik, Francisco Gomes de Oliveira Neto, Gregory Gay 0002, Kim Viggedal, David Friberg |
ASE | 4 |
| 2024 | Scoping of Non-Functional Requirements for Machine Learning SystemsabstractMachine Learning (ML) systems increasingly perform complex decision-making and prediction tasks—e.g., in autonomous driving—based on patterns inferred from large quantities of data. The inclusion of ML increases the capabilities of software systems, but also introduces or exacerbates challenges. ML systems can be more complex, time-consuming and expensive to specify, develop, and test than traditional systems, and can suffer from issues related to safety, lack of explainability, limited maintainability, and bias [1], [2]. As in other domains, ML systems must satisfy certain quality requirements—known as non-functional requirements (NFRs)—to be considered fit for purpose [1]. Khan Mohammad Habibullah, Juan García-Bellido, Gregory Gay 0002, Jennifer Horkoff |
RE | 3 |
| 2024 | Requirements and software engineering for automotive perception systems: an interview studyabstractAbstract Driving automation systems, including autonomous driving and advanced driver assistance, are an important safety-critical domain. Such systems often incorporate perception systems that use machine learning to analyze the vehicle environment. We explore new or differing topics and challenges experienced by practitioners in this domain, which relate to requirements engineering (RE), quality, and systems and software engineering. We have conducted a semi-structured interview study with 19 participants across five companies and performed thematic analysis of the transcriptions. Practitioners have difficulty specifying upfront requirements and often rely on scenarios and operational design domains (ODDs) as RE artifacts. RE challenges relate to ODD detection and ODD exit detection, realistic scenarios, edge case specification, breaking down requirements, traceability, creating specifications for data and annotations, and quantifying quality requirements. Practitioners consider performance, reliability, robustness, user comfort, and—most importantly—safety as important quality attributes. Quality is assessed using statistical analysis of key metrics, and quality assurance is complicated by the addition of ML, simulation realism, and evolving standards. Systems are developed using a mix of methods, but these methods may not be sufficient for the needs of ML. Data quality methods must be a part of development methods. ML also requires a data-intensive verification and validation process, introducing data, analysis, and simulation challenges. Our findings contribute to understanding RE, safety engineering, and development methodologies for perception systems. This understanding and the collected challenges can drive future research for driving automation and other ML systems. Khan Mohammad Habibullah, Hans-Martin Heyn, Gregory Gay 0002, Jennifer Horkoff, Eric Knauss, Markus Borg, Alessia Knauss, Håkan Sivencrona, Polly Jing Li |
Requir. Eng. | 3 |
| 2023 | Search-Based Test Generation Targeting Non-Functional Quality Attributes of Android AppsabstractMobile apps form a major proportion of the software marketplace and it is crucial to ensure that they meet both functional and nonfunctional quality thresholds. Automated test input generation can reduce the cost of the testing process. However, existing Android test generation approaches are focused on code coverage and cannot be customized to a tester's diverse goals---in particular, quality attributes such as resource use. Teklit Gereziher, Selam Gebrekrstos, Gregory Gay 0002 |
GECCO | 3 |
| 2023 | Test Maintenance for Machine Learning Systems: A Case Study in the Automotive IndustryabstractMachine Learning (ML) systems have seen widespread use for automated decision making. Testing is essential to ensure the quality of these systems, especially safety-critical autonomous systems in the automotive domain. ML systems introduce new challenges with the potential to affect test maintenance, the process of updating test cases to match the evolving system. We conducted an exploratory case study in the automotive domain to identify factors that affect test maintenance for ML systems, as well as to make recommendations to improve the maintenance process. Based on interview and artifact analysis, we identified 14 factors affecting maintenance, including five especially relevant for ML systems—with the most important relating to non-determinism and large input spaces. We also proposed ten recommendations for improving test maintenance, including four targeting ML systems—in particular, emphasizing the use of test oracles tolerant to acceptable non-determinism. The study’s findings expand our knowledge of test maintenance for an emerging class of systems, benefiting the practitioners testing these systems. Lukas Berglund, Tim Grube, Gregory Gay 0002, Francisco Gomes de Oliveira Neto, Dimitrios Platis |
ICST | 3 |
| 2023 | How Closely are Common Mutation Operators Coupled to Real Faults?abstractIn mutation testing, faulty versions of a program are generated through automated modifications of source code. These mutants are used to assess and improve test suite quality, under the assumption that detection of mutants is indicative of a test suite’s ability to detect real faults—i.e., that mutants and faults have a semantic relationship. Improving the effectiveness—in both cost and quality—of mutation testing may lie in better understanding this relationship, in particular with regard to how individual mutation operators (types) couple to real faults.In this study, we examine coupling between 32,002 mutants produced by 31 mutation operators and 144 real faults, using a scale based on number of failing tests and reasons for failure. Ultimately, we observed that 9.92% of the mutants are strongly coupled to real faults, and 51.03% of the faults have at least one strongly coupled mutant. We identify and examine mutation operators with the highest median coupling, as well as the operators that tend to produce non-compiling mutants, undetected mutants, and mutants that cause tests other than those that detect the actual fault to fail. We also examine how coupling could be used to filter the set of operators employed, leading to potentially significant cost savings during mutation testing. Our findings could lead to improvements in how mutation testing is applied, improved implementation of specific mutation operators, and inspiration for new mutation operators. Gregory Gay 0002, Alireza Salahirad |
ICST | 1 |
| 2023 | Understanding Problem Solving in Software Testing: An Exploration of Tester Routines and Behavior
Eduard Paul Enoiu, Gregory Gay 0002, Jameel Esber, Robert Feldt |
ICTSS | 2 |
| 2023 | How Do Different Types of Testing Goals Affect Test Case Design?
Dia Istanbuly, Max Zimmer, Gregory Gay 0002 |
ICTSS | 3 |
| 2023 | Requirements Engineering for Automotive Perception Systems: An Interview Study
Khan Mohammad Habibullah, Hans-Martin Heyn, Gregory Gay 0002, Jennifer Horkoff, Eric Knauss, Markus Borg, Alessia Knauss, Håkan Sivencrona, Polly Jing Li |
REFSQ | 3 |
| 2023 | Improving the Readability of Generated Tests Using GPT-4 and ChatGPT Code Interpreter
Gregory Gay 0002 |
SSBSE | 1 |
| 2023 | Developer Views on Software Carbon Footprint and Its Potential for Automated Reduction
Haozhou Lyu, Gregory Gay 0002, Maiko Sakamoto |
SSBSE | 2 |
| 2023 | Exploring Genetic Improvement of the Carbon Footprint of Web Pages
Haozhou Lyu, Gregory Gay 0002, Maiko Sakamoto |
SSBSE | 2 |
| 2023 | Mapping the structure and evolution of software testing research over the past three decadesabstractThe field of software testing is growing and rapidly-evolving. Based on keywords assigned to publications, we seek to identify predominant research topics and understand how they are connected and have evolved. We apply co-word analysis to map the topology of testing research as a network where author-assigned keywords are connected by edges indicating co-occurrence in publications. Keywords are clustered based on edge density and frequency of connection. We examine the most popular keywords, summarize clusters into high-level research topics examine how topics connect, and examine how the field is changing. Testing research can be divided into 16 high-level topics and 18 subtopics. Creation guidance, automated test generation, evolution and maintenance, and test oracles have particularly strong connections to other topics, highlighting their multidisciplinary nature. Emerging keywords relate to web and mobile apps, machine learning, energy consumption, automated program repair and test generation, while emerging connections have formed between web apps, test oracles, and machine learning with many topics. Random and requirements-based testing show potential decline. Our observations, advice, and map data offer a deeper understanding of the field and inspiration regarding challenges and connections to explore. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Alireza Salahirad, Gregory Gay 0002, Ehsan Mohammadi |
J. Syst. Softw. | 2 |
| 2023 | Non-functional requirements for machine learning: understanding current use and challenges among practitionersabstractAbstract Systems that rely on Machine Learning (ML systems) have differing demands on quality—known as non-functional requirements (NFRs)—from traditional systems. NFRs for ML systems may differ in their definition, measurement, scope, and comparative importance. Despite the importance of NFRs in ensuring the quality ML systems, our understanding of all of these aspects is lacking compared to our understanding of NFRs in traditional domains. We have conducted interviews and a survey to understand how NFRs for ML systems are perceived among practitioners from both industry and academia. We have identified the degree of importance that practitioners place on different NFRs, including cases where practitioners are in agreement or have differences of opinion. We explore how NFRs are defined and measured over different aspects of a ML system (i.e., model, data, or whole system). We also identify challenges associated with NFR definition and measurement. Finally, we explore differences in perspective between practitioners in industry, academia, or a blended context. This knowledge illustrates how NFRs for ML systems are treated in current practice, and helps to guide future RE for ML efforts. Khan Mohammad Habibullah, Gregory Gay 0002, Jennifer Horkoff |
Requir. Eng. | 2 |
| 2023 | The integration of machine learning into automated test generation: A systematic mapping studyabstractAbstract Machine learning (ML) may enable effective automated test generation. We characterize emerging research, examining testing practices, researcher goals, ML techniques applied, evaluation, and challenges in this intersection by performing. We perform a systematic mapping study on a sample of 124 publications. ML generates input for system, GUI, unit, performance, and combinatorial testing or improves the performance of existing generation methods. ML is also used to generate test verdicts, property‐based, and expected output oracles. Supervised learning—often based on neural networks—and reinforcement learning—often based on Q‐learning—are common, and some publications also employ unsupervised or semi‐supervised learning. (Semi‐/Un‐)Supervised approaches are evaluated using both traditional testing metrics and ML‐related metrics (e.g., accuracy), while reinforcement learning is often evaluated using testing metrics tied to the reward function. The work‐to‐date shows great promise, but there are open challenges regarding training data, retraining, scalability, evaluation complexity, ML algorithms employed—and how they are applied—benchmarks, and replicability. Our findings can serve as a roadmap and inspiration for researchers in this field. Afonso Fontes, Gregory Gay 0002 |
Softw. Test. Verification Reliab. | 2 |
| 2022 | Learning how to search: generating effective test cases through adaptive fitness function selectionabstractAbstract Search-based test generation is guided by feedback from one or more fitness functions—scoring functions that judge solution optimality. Choosing informative fitness functions is crucial to meeting the goals of a tester. Unfortunately, many goals—such as forcing the class-under-test to throw exceptions, increasing test suite diversity, and attaining Strong Mutation Coverage—do not have effective fitness function formulations. We propose that meeting such goals requires treating fitness function identification as a secondary optimization step. An adaptive algorithm that can vary the selection of fitness functions could adjust its selection throughout the generation process to maximize goal attainment, based on the current population of test suites. To test this hypothesis, we have implemented two reinforcement learning algorithms in the EvoSuite unit test generation framework, and used these algorithms to dynamically set the fitness functions used during generation for the three goals identified above. We have evaluated our framework, EvoSuiteFIT, on a set of Java case examples. EvoSuiteFIT techniques attain significant improvements for two of the three goals, and show limited improvements on the third when the number of generations of evolution is fixed. Additionally, for two of the three goals, EvoSuiteFIT detects faults missed by the other techniques. The ability to adjust fitness functions allows strategic choices that efficiently produce more effective test suites, and examining these choices offers insight into how to attain our testing goals. We find that adaptive fitness function selection is a powerful technique to apply when an effective fitness function does not already exist for achieving a testing goal. Hussein K. Almulla, Gregory Gay 0002 |
Empir. Softw. Eng. | 2 |
| 2021 | Special Section on the 2019 Symposium on Search-Based Software Engineering
Gregory Gay 0002, Shiva Nejati 0001 |
Inf. Softw. Technol. | 1 |
| 2020 | Understanding The Impact of Solver Choice in Model-Based Test GenerationabstractBackground: In model-based test generation, SMT solvers explore the state-space of the model in search of violations of specified properties. If the solver finds that a predicate can be violated, it produces a partial test specification demonstrating the violation. Gregory Gay 0002 |
ESEM | 2 |
| 2020 | Learning How to Search: Generating Exception-Triggering Tests Through Adaptive Fitness Function SelectionabstractSearch-based test generation is guided by feedback from one or more fitness functions-scoring functions that judge solution optimality. Choosing informative fitness functions is crucial to meeting the goals of a tester. Unfortunately, many goals-such as forcing the class-under-test to throw exceptions- do not have a known fitness function formulation. We propose that meeting such goals requires treating fitness function identification as a secondary optimization step. An adaptive algorithm that can vary the selection of fitness functions could adjust its selection throughout the generation process to maximize goal attainment, based on the current population of test suites. To test this hypothesis, we have implemented two reinforcement learning algorithms in the EvoSuite framework, and used these algorithms to dynamically set the fitness functions used during generation.We have evaluated our framework, EvoSuiteFIT, on a set of 386 real faults. EvoSuiteFIT discovers and retains more exception-triggering input and produces suites that detect a variety of faults missed by the other techniques. The ability to adjust fitness functions allows EvoSuiteFIT to make strategic choices that efficiently produce more effective test suites. Hussein K. Almulla, Gregory Gay 0002 |
ICST | 2 |
| 2020 | Generating Diverse Test Suites for Gson Through Adaptive Fitness Function Selection
Hussein K. Almulla, Gregory Gay 0002 |
SSBSE | 2 |
| 2020 | Bytecode-Based Multiple Condition Coverage: An Initial Investigation
Srujana Bollina, Gregory Gay 0002 |
SSBSE | 2 |
| 2020 | Defects4J as a Challenge Case for the Search-Based Software Engineering Community
Gregory Gay 0002, René Just |
SSBSE | 1 |
| 2020 | Choosing the fitness function for the job: Automated generation of test suites that detect real faultsabstractThe article from this special issue was previously published in Software Testing, Verification and Reliability, Volume 29, Issue 4–5, 2019. For completeness we are including the title page of the article below. The full text of the article can be read in Issue 29:4–5 on Wiley Online Library: https://onlinelibrary.wiley.com/doi/10.1002/stvr.1701 Alireza Salahirad, Hussein K. Almulla, Gregory Gay 0002 |
Softw. Test. Verification Reliab. | 3 |
| 2020 | Ensuring the Observability of Structural Test ObligationsabstractTest adequacy criteria are widely used to guide test creation. However, many of these criteria are sensitive to statement structure or the choice of test oracle. This is because such criteria ensure that execution reaches the element of interest, but impose no constraints on the execution path after this point. We are not guaranteed to observe a failure just because a fault is triggered. To address this issue, we have proposed the concept of observability-an extension to coverage criteria based on Boolean expressions that combines the obligations of a host criterion with an additional path condition that increases the likelihood that a fault encountered will propagate to a monitored variable. Our study, conducted over five industrial systems and an additional forty open-source systems, has revealed that adding observability tends to improve efficacy over satisfaction of the traditional criteria, with average improvements of 125.98 percent in mutation detection with the common output-only test oracle and per-model improvements of up to 1760.52 percent. Ultimately, there is merit to our hypothesis-observability reduces sensitivity to the choice of oracle and to the program structure. Gregory Gay 0002, Michael W. Whalen |
IEEE Trans. Software Eng. | 2 |
| 2019 | Choosing the fitness function for the job: Automated generation of test suites that detect real faultsabstractSummary Search‐based unit test generation, if effective at fault detection, can lower the cost of testing. Such techniques rely on fitness functions to guide the search. Ultimately, such functions represent test goals that approximate—but do not ensure—fault detection. The need to rely on approximations leads to two questions—can fitness functions produce effective tests and, if so, which should be used to generate tests? To answer these questions, we have assessed the fault‐detection capabilities of unit test suites generated to satisfy eight white‐box fitness functions on 597 real faults from the Defects4J database. Our analysis has found that the strongest indicators of effectiveness are a high level of code coverage over the targeted class and high satisfaction of a criterion's obligations. Consequently, the branch coverage fitness function is the most effective. Our findings indicate that fitness functions that thoroughly explore system structure should be used as primary generation objectives—supported by secondary fitness functions that explore orthogonal, supporting scenarios. Our results also provide further evidence that future approaches to test generation should focus on attaining higher coverage of private code and better initialization and manipulation of class dependencies. Alireza Salahirad, Hussein K. Almulla, Gregory Gay 0002 |
Softw. Test. Verification Reliab. | 3 |
| 2018 | Detecting Real Faults in the Gson Library Through Search-Based Unit Test GenerationabstractAn important benchmark for test generation tools is their ability to detect real faults . We have identified 16 real faults in Gson—a Java library for manipulating JSON data—and added them to the Defects4J fault database. Tests generated using the EvoSuite framework are able to detect seven faults. Analysis of the remaining faults offers lessons in how to improve generation. We offer these faults to the community to assist future research. Gregory Gay 0002 |
SSBSE | 1 |
| 2018 | Mapping Class Dependencies for Fun and ProfitabstractClasses depend on other classes to perform certain tasks. By mapping these dependencies, we may be able to improve software quality. We have developed a prototype framework for generating optimized groupings of classes coupled to targets of interest. From a pilot study investigating the value of coupling information in test generation, we have seen that coupled classes generally have minimal impact on results. However, we found 23 cases where the inclusion of coupled classes improves test suite efficacy, with an average improvement of 120.26% in the likelihood of fault detection. Seven faults were detected only through the inclusion of coupled classes. These results offer lessons on how coupling information could improve automated test generation. Allen Kanapala, Gregory Gay 0002 |
SSBSE | 2 |
| 2018 | Investigating faults missed by test suites achieving high code coverage
Amanda Schwartz, Daniel Puckett 0002, Gregory Gay 0002 |
J. Syst. Softw. | 4 |
| 2017 | The Fitness Function for the Job: Search-Based Generation of Test Suites That Detect Real FaultsabstractSearch-based test generation, if effective at fault detection, can lower the cost of testing. Such techniques rely on fitness functions to guide the search. Ultimately, such functions represent test goals that approximate - but do not ensure - fault detection. The need to rely on approximations leads to two questions - can fitness functions produce effective tests and, if so, which should be used to generate tests? To answer these questions, we have assessed the fault-detection capabilities of the EvoSuite framework and eight of its fitness functions on 353 real faults from the Defects4J database. Our analysis has found that the strongest indicator of effectiveness is a high level of code coverage. Consequently, the branch coverage fitness function is the most effective. Our findings indicate that fitness functions that thoroughly explore system structure should be used as primary generation objectives - supported by secondary fitness functions that vary the scenarios explored. Gregory Gay 0002 |
ICST | 1 |
| 2017 | Using Search-Based Test Generation to Discover Real Faults in Guava
Hussein K. Almulla, Alireza Salahirad, Gregory Gay 0002 |
SSBSE | 3 |
| 2017 | Generating Effective Test Suites by Combining Coverage Criteria
Gregory Gay 0002 |
SSBSE | 1 |
| 2017 | Automated Steering of Model-Based Test Oracles to Admit Real Program BehaviorsabstractThe test oracle-a judge of the correctness of the system under test (SUT)-is a major component of the testing process. Specifying test oracles is challenging for some domains, such as real-time embedded systems, where small changes in timing or sensory input may cause large behavioral differences. Models of such systems, often built for analysis and simulation, are appealing for reuse as test oracles. These models, however, typically represent an idealized system, abstracting away certain issues such as non-deterministic timing behavior and sensor noise. Thus, even with the same inputs, the model's behavior may fail to match an acceptable behavior of the SUT, leading to many false positives reported by the test oracle. We propose an automated steering framework that can adjust the behavior of the model to better match the behavior of the SUT to reduce the rate of false positives. This model steering is limited by a set of constraints (defining the differences in behavior that are acceptable) and is based on a search process attempting to minimize a dissimilarity metric. This framework allows non-deterministic, but bounded, behavioral differences, while preventing future mismatches by guiding the oracle-within limits-to match the execution of the SUT. Results show that steering significantly increases SUT-oracle conformance with minimal masking of real faults and, thus, has significant potential for reducing false positives and, consequently, testing and debugging costs while improving the quality of the testing process. Gregory Gay 0002, Sanjai Rayadurgam, Mats P. E. Heimdahl |
IEEE Trans. Software Eng. | 1 |
| 2016 | Challenges in Using Search-Based Test Generation to Identify Real Faults in Mockito
Gregory Gay 0002 |
SSBSE | 1 |
| 2016 | The Effect of Program and Model Structure on the Effectiveness of MC/DC Test Adequacy CoverageabstractTest adequacy metrics defined over the structure of a program, such as Modified Condition and Decision Coverage (MC/DC), are used to assess testing efforts. However, MC/DC can be “cheated” by restructuring a program to make it easier to achieve the desired coverage. This is concerning, given the importance of MC/DC in assessing the adequacy of test suites for critical systems domains. In this work, we have explored the impact of implementation structure on the efficacy of test suites satisfying the MC/DC criterion using four real-world avionics systems. Our results demonstrate that test suites achieving MC/DC over implementations with structurally complex Boolean expressions are generally larger and more effective than test suites achieving MC/DC over functionally equivalent, but structurally simpler, implementations. Additionally, we found that test suites generated over simpler implementations achieve significantly lower MC/DC and fault-finding effectiveness when applied to complex implementations, whereas test suites generated over the complex implementation still achieve high MC/DC and attain high fault finding over the simpler implementation. By measuring MC/DC over simple implementations, we can significantly reduce the cost of testing, but in doing so, we also reduce the effectiveness of the testing process. Thus, developers have an economic incentive to “cheat” the MC/DC criterion, but this cheating leads to negative consequences. Accordingly, we recommend that organizations require MC/DC over a structurally complex implementation for testing purposes to avoid these consequences. Gregory Gay 0002, Ajitha Rajan, Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2015 | 8th International Workshop on Search-Based Software Testing (SBST 2015)abstractThis paper is a report on the 8th International Workshop on Search-Based Software Testing at the 37th International Conference on Sofrware Engineering (ICSE). Search-Based Software Testing (SBST) is a form of Search-Based Software Engineering (SBSE) that optimizes testing through the use of computational search. SBST is used to generate test data, prioritize test cases, minimize test suites, reduce human oracle cost, verify software models, test service-orientated architectures, construct test suites for interaction testing, and validate real time properties. The objectives of this workshop are to bring together researchers and industrial practitioners from SBST and the wider software engineering community to share experience and provide directions for future research, and to encourage the use of search techniques to combine aspects of testing with other aspects of the software engineering lifecycle.Three full research papers, three short papers, and threeposition papers will be presented in the two-day workshop. Additionally, six development groups have pitted their test generation tools against a common set of programs and benchmarks, and will present their techniques and results. This report will give the background of the workshop and detail the provisional program. Gregory Gay 0002, Giuliano Antoniol |
ICSE (2) | 1 |
| 2015 | Efficient observability-based test generation by dynamic symbolic executionabstractStructural coverage metrics have been widely used to measure test suite adequacy as well as to generate test cases. In previous investigations, we have found that the fault-finding effectiveness of tests satisfying structural coverage criteria is highly dependent on program syntax - even if the faulty code is exercised, its effect may not be observable at the output. To address these problems, observability-based coverage metrics have been defined. Specifically, Observable MC/DC (OMC/DC) is a criterion that appears to be both more effective at detecting faults and more robust to program restructuring than MC/DC. Traditional counterexample-based test generation for OMC/DC, however, can be infeasible on large systems. In this study, we propose an incremental test generation approach that combines the notion of observability with dynamic symbolic execution. We evaluated the efficiency and effectiveness of our approach using seven systems from the avionics and medical device domains. Our results show that the incremental approach requires much lower generation time, while achieving even higher fault finding effectiveness compared with regular OMC/DC generation. Dongjiang You, Sanjai Rayadurgam, Michael W. Whalen, Mats P. E. Heimdahl, Gregory Gay 0002 |
ISSRE | 5 |
| 2015 | The Risks of Coverage-Directed Test Case GenerationabstractA number of structural coverage criteria have been proposed to measure the adequacy of testing efforts. In the avionics and other critical systems domains, test suites satisfying structural coverage criteria are mandated by standards. With the advent of powerful automated test generation tools, it is tempting to simply generate test inputs to satisfy these structural coverage criteria. However, while techniques to produce coverage-providing tests are well established, the effectiveness of such approaches in terms of fault detection ability has not been adequately studied. In this work, we evaluate the effectiveness of test suites generated to satisfy four coverage criteria through counterexample-based test generation and a random generation approach-where tests are randomly generated until coverage is achieved-contrasted against purely random test suites of equal size. Our results yield three key conclusions. First, coverage criteria satisfaction alone can be a poor indication of fault finding effectiveness, with inconsistent results between the seven case examples (and random test suites of equal size often providing similar-or even higher-levels of fault finding). Second, the use of structural coverage as a supplement-rather than a target-for test generation can have a positive impact, with random test suites reduced to a coverage-providing subset detecting up to 13.5 percent more faults than test suites generated specifically to achieve coverage. Finally, Observable MC/DC, a criterion designed to account for program structure and the selection of the test oracle, can-in part-address the failings of traditional structural coverage criteria, allowing for the generation of test suites achieving higher levels of fault detection than random test suites of equal size. These observations point to risks inherent in the increase in test automation in critical systems, and the need for more research in how coverage criteria, test generation approaches, the test oracle used, and system structure jointly influence test effectiveness. Gregory Gay 0002, Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl |
IEEE Trans. Software Eng. | 1 |
| 2015 | Automated Oracle Data Selection SupportabstractThe choice of test oracle-the artifact that determines whether an application under test executes correctly-can significantly impact the effectiveness of the testing process. However, despite the prevalence of tools that support test input selection, little work exists for supporting oracle creation. We propose a method of supporting test oracle creation that automatically selects the oracle data-the set of variables monitored during testing-for expected value test oracles. This approach is based on the use of mutation analysis to rank variables in terms of fault-finding effectiveness, thus automating the selection of the oracle data. Experimental results obtained by employing our method over six industrial systems (while varying test input types and the number of generated mutants) indicate that our method-when paired with test inputs generated either at random or to satisfy specific structural coverage criteria-may be a cost-effective approach for producing small, effective oracle data sets, with fault finding improvements over current industrial best practice of up to 1,435 percent observed (with typical improvements of up to 50 percent). Gregory Gay 0002, Matthew Staats, Michael W. Whalen, Mats P. E. Heimdahl |
IEEE Trans. Software Eng. | 1 |
| 2014 | Improving the accuracy of oracle verdicts through automated model steeringabstractThe oracle - a judge of the correctness of the system under test (SUT) - is a major component of the testing process. Specifying test oracles is challenging for some domains, such as real-time embedded systems, where small changes in timing or sensory input may cause large behavioral differences. Models of such systems, often built for analysis and simulation, are appealing for reuse as oracles. These models, however, typically represent an idealized system, abstracting away certain issues such as non-deterministic timing behavior and sensor noise. Thus, even with the same inputs, the model's behavior may fail to match an acceptable behavior of the SUT, leading to many false positives reported by the oracle. Gregory Gay 0002, Sanjai Rayadurgam, Mats P. E. Heimdahl |
ASE | 1 |
| 2013 | Observable modified Condition/Decision coverageabstractIn many critical systems domains, test suite adequacy is currently measured using structural coverage metrics over the source code. Of particular interest is the modified condition/decision coverage (MC/DC) criterion required for, e.g., critical avionics systems. In previous investigations we have found that the efficacy of such test suites is highly dependent on the structure of the program under test and the choice of variables monitored by the oracle. MC/DC adequate tests would frequently exercise faulty code, but the effects of the faults would not propagate to the monitored oracle variables. In this report, we combine the MC/DC coverage metric with a notion of observability that helps ensure that the result of a fault encountered when covering a structural obligation propagates to a monitored variable; we term this new coverage criterion Observable MC/DC (OMC/DC). We hypothesize this path requirement will make structural coverage metrics 1.) more effective at revealing faults, 2.) more robust to changes in program structure, and 3.) more robust to the choice of variables monitored. We assess the efficacy and sensitivity to program structure of OMC/DC as compared to masking MC/DC using four subsystems from the civil avionics domain and the control logic of a microwave. We have found that test suites satisfying OMC/DC are significantly more effective than test suites satisfying MC/DC, revealing up to 88% more faults, and are less sensitive to program structure and the choice of monitored variables. Michael W. Whalen, Gregory Gay 0002, Dongjiang You, Mats P. E. Heimdahl, Matthew Staats |
ICSE | 2 |
| 2012 | On the Danger of Coverage Directed Test Case Generation
Matthew Staats, Gregory Gay 0002, Michael W. Whalen, Mats P. E. Heimdahl |
FASE | 2 |
| 2012 | Automated oracle creation support, or: How I learned to stop worrying about fault propagation and love mutation testingabstractIn testing, the test oracle is the artifact that determines whether an application under test executes correctly. The choice of test oracle can significantly impact the effectiveness of the testing process. However, despite the prevalence of tools that support the selection of test inputs, little work exists for supporting oracle creation. In this work, we propose a method of supporting test oracle creation. This method automatically selects the oracle data — the set of variables monitored during testing — for expected value test oracles. This approach is based on the use of mutation analysis to rank variables in terms of fault-finding effectiveness, thus automating the selection of the oracle data. Experiments over four industrial examples demonstrate that our method may be a cost-effective approach for producing small, effective oracle data, with fault finding improvements over current industrial best practice of up to 145.8% observed. Matthew Staats, Gregory Gay 0002, Mats P. E. Heimdahl |
ICSE | 2 |
| 2011 | Sharing experiments using open-source softwareabstractAbstract When researchers want to repeat, improve or refute prior conclusions, it is useful to have a complete and operational description of prior experiments. If those descriptions are overly long or complex, then sharing their details may not be informative. OURMINE is a scripting environment for the development and deployment of data mining experiments. Using OURMINE, data mining novices can specify and execute intricate experiments, while researchers can publish their complete experimental rig alongside their conclusions. This is achievable because of OURMINE's succinctness. For example, this paper presents two experiments documented in the OURMINE syntax. Thus, the brevity and simplicity of OURMINE recommends it as a better tool for documenting, executing, and sharing data mining experiments. Copyright © 2010 John Wiley & Sons, Ltd. Adam Nelson, Tim Menzies, Gregory Gay 0002 |
Softw. Pract. Exp. | 3 |
| 2010 | When to use data from other projects for effort estimationabstractCollecting the data required for quality prediction within a development team is time-consuming and expensive. An alternative to make predictions using data that crosses from other projects or even other companies. We show that with/without relevancy filtering, imported data performs the same/worse (respectively) than using local data. Therefore, we recommend the use of relevancy filtering whenever generating estimates using data from another project. Ekrem Kocaguneli, Gregory Gay 0002, Tim Menzies, Jacky W. Keung |
ASE | 2 |
| 2010 | Automatically finding the control variables for complex system behavior
Gregory Gay 0002, Tim Menzies, Misty D. Davies, Karen Gundy-Burlet |
Autom. Softw. Eng. | 1 |
| 2010 | Finding robust solutions in requirements models
Gregory Gay 0002, Tim Menzies, Omid Jalali, Gregory E. Mundy, Beau Gilkerson, Martin Feather, James D. Kiper |
Autom. Softw. Eng. | 1 |
| 2009 | On the use of relevance feedback in IR-based concept locationabstractConcept location is a critical activity during software evolution as it produces the location where a change is to start in response to a modification request, such as, a bug report or a new feature request. Lexical-based concept location techniques rely on matching the text embedded in the source code to queries formulated by the developers. The efficiency of such techniques is strongly dependent on the ability of the developer to write good queries. We propose an approach to augment information retrieval (IR) based concept location via an explicit relevance feedback (RF) mechanism. RF is a two-part process in which the developer judges existing results returned by a search and the IR system uses this information to perform a new search, returning more relevant information to the user. A set of case studies performed on open source software systems reveals the impact of RF on IR based concept location. Gregory Gay 0002, Sonia Haiduc, Andrian Marcus, Tim Menzies |
ICSM | 1 |