EDBT 2026 Demo / reviewers in the wild / expert
Phil McMinn
dblp:33/13 · also Philip McMinn
· DBLP profile ↗
90ranked-venue papers
16as first author
25since 2021 · last 2026
0000-0001-9137-7433ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 73 · 11 first-author · 23 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Effective are Coverage- and Diversity-Based Test Selection at Killing Stubborn Mutants?
Islam T. Elgendy, Robert M. Hierons, Phil McMinn |
ICST | 3 |
| 2026 | Dynamic Mutation Scheduling: Highly Parallel, Efficient Evaluation of Mutations for Rust Programs through Program Splitting
Zalán Lévai, Donghwan Shin 0001, Phil McMinn |
ICST | 3 |
| 2026 | mutest-rs: Flexible, Efficient Mutation Analysis Tool for Rust Programs, using Extensive Static AnalysisabstractDetermining the adequacy of software tests, and where testing gaps might lie, is crucial for improving and maintaining the strength of test suites. Mutation analysis facilitates this by evaluating tests against generated program faults; however, no mature mutation analysis tooling exists for the safety-focused Rust systems programming language, to date. This paper introduces mutest-rs, a mature, end-to-end mutation analysis tool for Rust programs that is based on extensive static program analysis, and integrates directly with the rustc Rust compiler. Our tool overcomes the numerous challenges of generating valid Rust code mutants, and does so efficiently through a Rustspecific meta-mutant approach. Our open-source tool, mutest-rs, is available online at https://mutest.rs. A video demonstrating mutest-rs is available at https://youtu.be/8yEYAU6P63I. Zalán Lévai, Donghwan Shin 0001, Phil McMinn |
ICST | 3 |
| 2026 | Multi-Fidelity Bayesian Optimization for Simulation Based Autonomous Driving Systems Testing
Olek Osikowicz, Phil McMinn, Donghwan Shin 0001 |
IV | 2 |
| 2026 | Fuzz3 : Entropy as a Third Oracle
Karine Even-Mendoza, Janine Obiri, Aidan Dakhama, Phil McMinn, William B. Langdon |
SSBSE | 4 |
| 2025 | Empirically Evaluating the Use of Bytecode for Diversity-Based Test Case PrioritisationabstractRegression testing assures software correctness after changes but is resource-intensive. Test Case Prioritisation (TCP) mitigates this by ordering tests to maximise early fault detection. Diversity-based TCP prioritises dissimilar tests, assuming they exercise different system parts and uncover more faults. Traditional static diversity-based TCP approaches (i.e., methods that utilise the dissimilarity of tests), like the state-of-the-art FAST approach, rely on textual diversity from test source code, which is effective but inefficient due to its relative verbosity and redundancies affecting similarity calculations. This paper is the first to study bytecode as the basis of diversity in TCP, leveraging its compactness for improved efficiency and accuracy. An empirical study on seven Defects4J projects shows that bytecode diversity improves fault detection by 2.3–7.8% over text-based TCP. It is also 2–3 orders of magnitude faster in one TCP approach and 2.5–6 times faster in FAST-based TCP. Filtering specific bytecode instructions improves efficiency up to fourfold while maintaining effectiveness, making bytecode diversity a superior static approach. Islam T. Elgendy, Robert M. Hierons, Phil McMinn |
EASE | 3 |
| 2025 | Systemic Flakiness: An Empirical Analysis of Co-Occurring Flaky Test FailuresabstractFlaky tests produce inconsistent outcomes without code changes, creating major challenges for software developers. An industrial case study reported that developers spend 1.28% of their time repairing flaky tests at a monthly cost of $2,250. This paper reveals that flaky tests often exist in clusters, with co-occurring failures that share the same root causes, which we call systemic flakiness. This result suggests that developers can reduce test repair costs by addressing shared root causes, enabling them to fix multiple flaky tests at once rather than tackling them individually. This study represents an inflection point by challenging the deep-seated assumption that flaky test failures are isolated occurrences. We used an established dataset of 10,000 test suite runs from 24 Java projects on GitHub, spanning domains from data orchestration to job scheduling. Using a data set that contains 810 flaky tests, we performed a mixed-method empirical analysis of co-occurring flaky test failures, revealing that systemic flakiness is significant and widespread. Owain Parry, Gregory M. Kapfhammer, Michael Hilton 0001, Phil McMinn |
EASE | 4 |
| 2025 | Where Tests Fall Short: Empirically Analyzing Oracle Gaps in Covered CodeabstractBackground: Developers often rely on statement coverage to assess test suite quality. However, statement coverage alone may only lead to 10 % fault detection, necessitating more rigorous approaches. While mutation testing is effective, its execution and human analysis costs remain high. Identifying covered statements that are not checked by oracles (e.g., assertions) offers a cost-effective alternative; however, the lack of empirical evidence for selecting the appropriate Oracle Gap Calculation Approach (OGCA) prevents developers from making informed choices. Aims: This knowledge-seeking study compares oracle gap characteristics determined by different OGCAs to assist developers in choosing the most valuable approach for their use cases. Method: Using mixed-method empirical analysis, we conduct an in-depth evaluation of the oracle gaps produced using three OGCAs: Checked Coverage using a Dynamic Slicer ($\mathbf{C C}_{D S}$), Checked Coverage using an Observational Slicer$\left(\mathbf{C C}_{O S}\right)$, and Pseudo-Tested Statement Identification (PTSI). Across 30 Java classes from six open-source projects, we report on a quantitative evaluation of gap prominence, distribution, fault detection correlation and execution times, as well as results from a qualitative manual inspection of the statement types found in the oracle gaps. Results: The qualitative analysis showed data-loading statements, iteration statements and output updates to be most prominent in the oracle gaps. PTSI identified the oracle gaps with the lowest median mutation score (0.32), highlighting areas requiring more fault detection improvement compared to$\mathbf{C C}_{D S} \boldsymbol{(} \mathbf{0. 7 6) }$and$\mathbf{C C}_{O S} \boldsymbol{(} \mathbf{0. 5 0} \boldsymbol{)}$. PTSI also had the shortest median execution time (19.9 seconds), far quicker than both$\text{CC}_{D S}$(273.2 seconds) and$\text{CC}_{O S}$(5957.1 seconds). Conclusions: PTSI quickly reveals the priority testing areas for improved fault detection, making it an effective OGCA for developers to identify where tests fall short. Megan Maton, Gregory M. Kapfhammer, Phil McMinn |
ESEM | 3 |
| 2025 | A Systematic Mapping Study of the Metrics, Uses and Subjects of Diversity-Based Testing TechniquesabstractABSTRACT There has been a significant amount of interest regarding the use of DBTtsfull in software testing over the past two decades. Diversity‐based testing (DBT) technique uses similarity metrics to leverage the dissimilarity between software artefacts—such as requirements, abstract models, programme structures or inputs—in order to address a software testing problem. DBT techniques have been used to assist in finding solutions to several different types of problems including generating test cases, prioritizing them and reducing very large test suites. This paper is a systematic mapping study of DBT techniques that summarizes the key aspects and trends of 167 papers that report the use of 79 different similarity metrics with 22 different types of software artefacts, which have been used by researchers to tackle 11 different types of software testing problems. We further present an analysis of the recent trends in DBT techniques and review the different application domains to which the techniques have been applied, giving an overview of the tools developed by researchers in order to do so. Finally, the paper identifies some DBT challenges that are potential topics for future work, such as exploring other diversity artefacts and measuring diversity for complex data input. Islam T. Elgendy, Robert M. Hierons, Phil McMinn |
Softw. Test. Verification Reliab. | 3 |
| 2024 | Evaluating String Distance Metrics for Reducing Automatically Generated Test SuitesabstractRegression test suites can have a large number of test cases, especially automatically generated ones, and tend to grow in size, making it costly to run the entire test suite. Test suite reduction aims to eliminate some test cases to reduce the test suite size and therefore reduce the cost of running it. In this paper, string distances on the text of the test cases are used as measures of similarity for reduction. A practical benefit of using string distance is that there is no need to run the test cases: the test suite source code is the only requirement, making the approach fast. We reduce test suites generated from Randoop and EvoSuite; two well-known test generation tools of Java programs. We implemented a string-based similarity reduction and compared it against random reduction. In the experiments, mutation scores using reduced test suites based on maximising string dissimilarity of test cases were higher than those for random reduction in over 70% of the test suites generated. Also, the results showed that test suites generated by Randoop can be drastically reduced in one case by 99% using the string-based similarity reduction approach while maintaining the fault-finding capabilities of the original test suite. Finally, on average, the normalised compression distance was found to be the best similarity metric choice in terms of fault-detection. Islam T. Elgendy, Robert M. Hierons, Phil McMinn |
AST | 3 |
| 2024 | Do Automatic Test Generation Tools Generate Flaky Tests?abstractNon-deterministic test behavior, or flakiness, is common and dreaded among developers. Researchers have studied the issue and proposed approaches to mitigate it. However, the vast majority of previous work has only considered developer-written tests. The prevalence and nature of flaky tests produced by test generation tools remain largely unknown. We ask whether such tools also produce flaky tests and how these differ from developer-written ones. Furthermore, we evaluate mechanisms that suppress flaky test generation. We sample 6 356 projects written in Java or Python. For each project, we generate tests using EvoSuite (Java) and Pynguin (Python), and execute each test 200 times, looking for inconsistent outcomes. Our results show that flakiness is at least as common in generated tests as in developer-written tests. Nevertheless, existing flakiness suppression mechanisms implemented in EvoSuite are effective in alleviating this issue (71.7 % fewer flaky tests). Compared to developer-written flaky tests, the causes of generated flaky tests are distributed differently. Their non-deterministic behavior is more frequently caused by randomness, rather than by networking and concurrency. Using flakiness suppression, the remaining flaky tests differ significantly from any flakiness previously reported, where most are attributable to runtime optimizations and EvoSuite-internal resource thresholds. These insights, with the accompanying dataset, can help maintainers to improve test generation tools, give recommendations for developers using these tools, and serve as a foundation for future research in test flakiness or test generation. Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck, Phil McMinn, Gordon Fraser 0001 |
ICSE | 5 |
| 2024 | Exploring Pseudo-Testedness: Empirically Evaluating Extreme Mutation Testing at the Statement LevelabstractExtreme mutation testing (XMT) detects undesirable pseudo-testedness in a program by deleting the method bodies of covered code and observing whether the test suite can detect their absence. Even though XMT may identify test limitations, its coarse granularity means that it may overlook testing inadequacies, particularly at the statement level, that developers may want to address before committing the resources demanded by traditional mutation testing. This paper proposes the use of the statement deletion mutation operator (SDL) to uncover pseudo-tested statements in addition to complete methods. In an experimental evaluation involving four frequently-studied, large, Apache Commons Java projects and 23 projects randomly selected from the Maven Central Repository, we found 722 different cases of pseudo-tested statements. Critically, we discovered that 48% of these statements exist outside of pseudo-tested methods, meaning that the detection of testing deficiencies related to these statements would normally be left to traditional, resource-intensive, mutation testing. Also, we found that a popular Java mutation testing tool would not have mutated some of the statement types involved in the first place, effectively rendering these issues, hitherto, hard to discover. This paper therefore demonstrates that XMT alone is insufficient and should be combined with pseudo-tested statement evaluation to pinpoint subtle, yet important, testing oversights that a developer should tackle before applying traditional mutation testing. Megan Maton, Gregory M. Kapfhammer, Phil McMinn |
ICSME | 3 |
| 2024 | PseudoSweep: A Pseudo-Tested Code IdentifierabstractSoftware testing remains a crucial practice for en-suring and maintaining code quality. Yet, a critical issue remains: the existence of pseudo-tested statements. Tests cover these statements, but removing them does not trigger test failures. Since no established tools address this challenge, this paper introduces Pseudosweep, a novel tool that automatically identifies pseudo-tested methods and statements in Java projects. PSEUDOSWEEP combines method and statement deletion techniques to reveal these maintenance problems. In addition to explaining the approach used by Pseudo Sweep,this paper details use-cases and overviews results from experiments with Pseudosweep. The tool is available (including set-up instructions and examples) at https://github.com/PseudoTested/PseudoSweep and there is a video demonstration at https://youtu.be/5QCsu7MbiXI. Megan Maton, Gregory M. Kapfhammer, Phil McMinn |
ICSME | 3 |
| 2024 | Private-Keep Out? Understanding How Developers Account for Code Visibility in Unit TestingabstractRegression test maintenance costs can be reduced by striving to write tests that will require as few changes as possible in the future. Writing unit tests against behavior, as opposed to implementation, is one way to try to achieve this, because as long as the public API remains constant, units can be safely refactored without the need to also change the tests. However, in a study on 4,801 open-source Java projects reported in this paper, we found that 28% of projects contradict this advice, with tests that side-step the public API by directly calling non-public methods. We investigated why developers do not solely test public APIs-potentially increasing future test maintenance costs-by surveying 73 developers and conducting a systematic review of 60 StackOverflow posts dating from 2008–2023. Through numerical and thematic analyses, we uncover several findings, including (1) developers are disunited on whether to test only through public APIs or not; (2) those in favor of only testing through the public API tend to be more experienced and believe the need or desire to break with this is borne out of poor software design; while (3) those that test non-public methods directly are concerned about untested code complexity and overly intricate tests. Our findings provide multiple implications for future work, including automated developer support in the form of automated non-public method sequence replacement, and automated refactoring of production code using problematic public API-avoiding tests. Muhammad Firhard Roslan, José Miguel Rojas, Phil McMinn |
ICSME | 3 |
| 2024 | Viscount: A Direct Method Call Coverage Tool for JavaabstractWriting unit tests against implementation detail in production code, often embodied in non-public methods, is considered bad practice in formal and gray literature. This is because it leads to fragile tests that break easily when underlying implementation details change. For this reason, tests that focus on behavior are encouraged. One way to achieve this is to test units exclusively through their public API. However, our recent developer survey shows that this advice is not always followed in practice. Moreover, code coverage tools do not provide a way to determine which methods were called directly from tests, meaning there is no easy way to identify whether units make calls to non-public methods, other than through manual examination. To address this problem, we developed Viscount, a tool that can determine direct method call coverage for Java tests written in JUnit. Viscount reports the percentage of methods invoked directly from tests, according to their visibility - i.e., public or non-public (protected, package-private, or private). This can help developers and researchers identify tests that potentially need to be refactored or rewritten. In this paper, we describe Viscount's overall architecture, its core features, and how to use it. Viscount is also publicly available on GitHub: https://github.com/unittesting-nonpublic/viscount. A demo video of Viscount is available at: https://youtu.be/ZUyRtiUnbsU. Muhammad Firhard Roslan, José Miguel Rojas, Phil McMinn |
ICSME | 3 |
| 2023 | Batching Non-Conflicting Mutations for Efficient, Safe, Parallel Mutation Analysis in RustabstractRust is a relatively young, memory safe systems programming language which is increasingly being adopted by projects requiring both performance, and safety. While automated testing is built into the language, tool support for mutation analysis is almost non-existent, having not been the subject of past research. This leaves Rust developers without a way to determine test thoroughness. To address this problem, we design a mutation analysis process for Rust that overcomes challenges related to generating viable mutations due to the strictness of the language in terms of its type system, memory restrictions, and the potential to introduce undefined behavior in unsafe code blocks. Our technique efficiently evaluates mutations simultaneously through a process we refer to as "batching" — the use of static analysis to determine mutations that are non-conflicting, and therefore are able to be evaluated together. Batching enables our technique to maximize thread usage, executing more tests in parallel, and further reducing the time required to evaluate mutations. We implemented these techniques into a tool, mutest-rs, which we empirically evaluated on a diverse set of common subject libraries and Rust programs, and found that our batching method for increasing parallelism is able to reduce the overall runtime of mutation analysis by up to 66.4%, compared to not applying batching. Zalán Lévai, Phil McMinn |
ICST | 2 |
| 2023 | Empirically evaluating flaky test detection techniques combining test case rerunning and machine learning modelsabstractAbstract A flaky test is a test case whose outcome changes without modification to the code of the test case or the program under test. These tests disrupt continuous integration, cause a loss of developer productivity, and limit the efficiency of testing. Many flaky test detection techniques are rerunning-based, meaning they require repeated test case executions at a considerable time cost, or are machine learning-based, and thus they are fast but offer only an approximate solution with variable detection performance. These two extremes leave developers with a stark choice. This paper introduces CANNIER, an approach for reducing the time cost of rerunning-based detection techniques by combining them with machine learning models. The empirical evaluation involving 89,668 test cases from 30 Python projects demonstrates that CANNIER can reduce the time cost of existing rerunning-based techniques by an order of magnitude while maintaining a detection performance that is significantly better than machine learning models alone. Furthermore, the comprehensive study extends existing work on machine learning-based detection and reveals a number of additional findings, including (1) the performance of machine learning models for detecting polluter test cases; (2) using the mean values of dynamic test case features from repeated measurements can slightly improve the detection performance of machine learning models; and (3) correlations between various test case features and the probability of the test case being flaky. Owain Parry, Gregory M. Kapfhammer, Michael Hilton 0001, Phil McMinn |
Empir. Softw. Eng. | 4 |
| 2022 | What Do Developer-Repaired Flaky Tests Tell Us About the Effectiveness of Automated Flaky Test Detection?abstractBecause they pass or fail without code changes, flaky tests cause serious problems such as spuriously failing builds and the eroding of developers' trust in tests. Many previous evaluations of automated flaky test detection techniques do not accurately assess their usefulness for the developers who identify the flaky tests to repair. This is because researchers evaluate detection techniques against baselines that are not derived from past developer behavior or against no baselines at all. To study the effectiveness of an automated test rerunning technique, a common baseline for other approaches to detection, this paper uses 75 commits --- authored by human software developers --- that repair test flakiness in 31 real-world Python projects. Surprisingly, automated rerunning detects the developer-repaired flaky tests in only 40% of the studied commits. This result suggests that automated rerunning does not often find those flaky tests that developers fix, implying that it makes an unsuitable baseline for assessing a detection technique's usefulness for developers. Owain Parry, Michael Hilton 0001, Gregory M. Kapfhammer, Phil McMinn |
AST | 4 |
| 2022 | Automated Repair of Responsive Web Page LayoutsabstractResponsive Web Design (RWD) is a strategy that allows developers to create webpages that adjust their layout according to available screen size. Since modern web applications must format correctly on the small displays of mobile devices up to the large displays on desktop computers, and given this dramatic difference in screen space, Responsive Layout Failures (RLFs) - visual discrepancies that are only apparent at certain screen sizes - can easily creep into live production webpages. These can include, for example, HTML elements protruding off the edge of the page or into one another as layout space becomes scarce. This leaves webpages looking unprofessional at best and non-functional at worst. This paper presents a technique for repairing RLFs, implemented into a tool called Layout Dr. After detecting an RLF, Layout Dr harvests layouts from the page's responsive design that are closest to the point of failure, but where the RLF does not occur. It then transforms these layouts so that they can be transplanted over the failure, effectively “hiding” the original RLF from the end user. We evaluated Layout Dr on 19 subjects, containing 55 RLFs in total. Layout Dr could find a suitable fix for each of them. When we conducted a human study of the repairs, 92% of the participants preferred the repaired version of the page compared to the original containing the RLF. Ibrahim Althomali, Gregory M. Kapfhammer, Phil McMinn |
ICST | 3 |
| 2022 | Evaluating Features for Machine Learning Detection of Order- and Non-Order-Dependent Flaky TestsabstractFlaky tests are test cases that can pass or fail without code changes. They often waste the time of software developers and obstruct the use of continuous integration. Previous work has presented several automated techniques for detecting flaky tests, though many involve repeated test executions and a lot of source code instrumentation and thus may be both intrusive and expensive. While this motivates researchers to evaluate machine learning models for detecting flaky tests, prior work on the features used to encode a test case is limited. Without further study of this topic, machine learning models cannot perform to their full potential in this domain. Previous studies also exclude a specific, yet prevalent and problematic, category of flaky tests: order-dependent (OD) flaky tests. This means that prior research only addresses part of the challenge of detecting flaky tests with machine learning. Closing this knowledge gap, this paper presents a new feature set for encoding tests, called Flake16. Using 54 distinct pipelines of data preprocessing, data balancing, and machine learning models for detecting both non-order-dependent (NOD) and OD flaky tests, this paper compares Flake16 to another well-established feature set. To assess the new feature set's effectiveness, this paper's experiments use the test suites of 26 Python projects, consisting of over 67,000 tests. Along with identifying the most impactful metrics for using machine learning to detect both types of flaky test, the empirical study shows how Flake16 is better than prior work, including (1) a 13% increase in overall F1 score when detecting NOD flaky tests and (2) a 17% increase in overall F1 score when detecting OD flaky tests. Owain Parry, Gregory M. Kapfhammer, Michael Hilton 0001, Phil McMinn |
ICST | 4 |
| 2022 | An Empirical Comparison of EvoSuite and DSpot for Improving Developer-Written Test Suites with Respect to Mutation Score
Muhammad Firhard Roslan, José Miguel Rojas, Phil McMinn |
SSBSE | 3 |
| 2022 | A Survey of Flaky TestsabstractTests that fail inconsistently, without changes to the code under test, are described as flaky . Flaky tests do not give a clear indication of the presence of software bugs and thus limit the reliability of the test suites that contain them. A recent survey of software developers found that 59% claimed to deal with flaky tests on a monthly, weekly, or daily basis. As well as being detrimental to developers, flaky tests have also been shown to limit the applicability of useful techniques in software testing research. In general, one can think of flaky tests as being a threat to the validity of any methodology that assumes the outcome of a test only depends on the source code it covers. In this article, we systematically survey the body of literature relevant to flaky test research, amounting to 76 papers. We split our analysis into four parts: addressing the causes of flaky tests, their costs and consequences, detection strategies, and approaches for their mitigation and repair. Our findings and their implications have consequences for how the software-testing community deals with test flakiness, pertinent to practitioners and of interest to those wanting to familiarize themselves with the research area. Owain Parry, Gregory M. Kapfhammer, Michael Hilton 0001, Phil McMinn |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2021 | An Empirical Study to Determine if Mutants Can Effectively Simulate Students' Programming Mistakes to Increase Tutors' Confidence in AutogradingabstractAutomated grading often requires automated test suites to identify students' faults. However, tests may not detect some faults, limiting feedback, and providing inaccurate grades. This issue can be mitigated by first ensuring that tests can detect faults. Mutation analysis is a technique that generates artificial faulty variants of a program for this purpose, called mutants. Mutants that are not detected by tests reveal their inadequacies, providing knowledge on how they can be improved. By using mutants to improve test suites, tutors can gain the confidence that: a) generated grades will not be biased by unidentified faults, and b) students will receive appropriate feedback for their mistakes. Existing work has shown that mutants are suitable substitutes for faults in real world software, but no work has shown that this holds for students' faults. In this paper, we investigate whether mutants are capable of replicating mistakes made by students. We conducted a quantitative study on 197 Java classes written by students across three introductory programming assignments, and mutants generated from the assignments' model solutions. We found that generated mutants capture the observed faulty behaviour of students' solutions. We also found that mutants better assess test adequacy than code coverage in some cases. Our results indicate that tutors can use mutants to identify and remedy deficiencies in grading test suites. Ben Clegg 0002, Phil McMinn, Gordon Fraser 0001 |
SIGCSE | 2 |
| 2021 | Automated visual classification of DOM-based presentation failure reports for responsive web pagesabstractSummary Since it is common for the users of a web page to access it through a wide variety of devices—including desktops, laptops, tablets and phones—web developers rely on responsive web design (RWD) principles and frameworks to create sites that are useful on all devices. A correctly implemented responsive web page adjusts its layout according to the viewport width of the device in use, thereby ensuring that its design suitably features the content. Since the use of complex RWD frameworks often leads to web pages with hard‐to‐detect responsive layout failures (RLFs), developers employ testing tools that generate reports of potential RLFs. Since testing tools for responsive web pages, like ReDeCheck, analyse a web page representation called the Document Object Model (DOM), they may inadvertently flag concerns that are not human visible, thereby requiring developers to manually confirm and classify each potential RLF as a true positive (TP), false positive (FP), or non‐observable issue (NOI)—a process that is time consuming and error prone. The conference version of this paper presented Viser, a tool that automatically classified three types of RLFs reported by ReDeCheck. Since Viser was not designed to automatically confirm and classify two types of RLFs that ReDeCheck's DOM‐based analysis could surface, this paper introduces Verve, a tool that automatically classifies all RLF types reported by ReDeCheck. Along with manipulating the opacity of HTML elements in a web page, as does Viser, the Verve tool also uses histogram‐based image comparison to classify RLFs in web pages. Incorporating both the 25 web pages used in prior experiments and 20 new pages not previously considered, this paper's empirical study reveals that Verve's classification of all five types of RLFs frequently agrees with classifications produced manually by humans. The experiments also reveal that Verve took on average about 4 s to classify any of the RLFs among the 469 reported by ReDeCheck. Since this paper demonstrates that classifying an RLF as a TP, FP, or NOI with Verve, a publicly available tool, is less subjective and error prone than the same manual process done by a human web developer, we argue that it is well‐suited for supporting the testing of complex responsive web pages. Ibrahim Althomali, Gregory M. Kapfhammer, Phil McMinn |
Softw. Test. Verification Reliab. | 3 |
| 2021 | Effective automated repair of internationalization presentation failures in web applications using style similarity clustering and search-based techniquesabstractSummary Companies often employ (i18n) frameworks to provide translated text and localized media content on their websites in order to effectively communicate with a global audience. However, the varying lengths of text from different languages can cause undesired distortions in the layout of a web page. Such distortions, called Internationalization Presentation Failures (IPFs), can negatively affect the aesthetics or usability of the website. Most of the existing automated techniques developed for assisting repair of IPFs either produce fixes that are likely to significantly reduce the legibility and attractiveness of the pages or are limited to only detecting IPFs, with the actual repair itself remaining a labour intensive manual task. To address this problem, we propose a search‐based technique for automatically repairing IPFs in web applications, while ensuring a legible and attractive page. The empirical evaluation of our approach reported that our approach was able to successfully resolve 94% of the detected IPFs for 46 real‐world web pages. In a user study, participants rated the visual quality of our fixes significantly higher than the unfixed versions and also considered the repairs generated by our approach to be notably more legible and visually appealing than the repairs generated by existing techniques. Sonal Mahajan, Abdulmajeed Alameer, Phil McMinn, William G. J. Halfond |
Softw. Test. Verification Reliab. | 3 |
| 2020 | The Influence of Test Suite Properties on Automated Grading of Programming ExercisesabstractAutomated grading allows for the scalable assessment of large programming courses, often using test cases to determine the correctness of students' programs. However, test suites can vary in multiple ways, such as quality, size, and coverage. In this paper, we investigate how much test suites with varying properties can impact generated grades, and how these properties cause this impact. We conduct a study on artificial faulty programs that simulate students' programming mistakes and test suites generated from manually written tests. We find that these test suites generate greatly varying grades, with the standard deviation of grades for each fault typically representing ~84% of the grades not apportioned to the fault. We show that different properties of test suites can influence the grades that they produce, with coverage typically making the greatest effect, and mutation score and the potentially redundant repeated coverage of lines also having a significant impact. We offer suggestions based on our findings to assist tutors with building grading test suites that assess students' code in a fair and consistent manner. These suggestions include ensuring that test suites have 100% coverage, avoiding unnecessarily recovering lines, and checking test suites using real or artificial faults. Ben Clegg 0002, Phil McMinn, Gordon Fraser 0001 |
CSEE&T | 2 |
| 2020 | STICCER: Fast and Effective Database Test Suite Reduction Through Merging of Similar Test CasesabstractSince relational databases support many software applications, industry professionals recommend testing both database queries and the underlying database schema that contains complex integrity constraints. These constraints, which include primary and foreign keys, NOT NULL, and arbitrary CHECK constraints, are important because they protect the consistency and coherency of data in the relational database. Since testing integrity constraints is potentially an arduous task, human testers can use new tools to automatically generate test suites that effectively find schema faults. However, these tool-generated test suites often contain many lengthy tests that may both increase the time overhead of regression testing and limit the ability of human testers to understand them. Aiming to reduce the size of automatically generated test suites for database schemas, this paper introduces STICCER, a technique that finds overlaps between test cases, merging database interactions from similar tests and removing others. By systematically discarding and merging redundant tests, STICCER creates a reduced test suite that is guaranteed to have the same coverage as the original one. Using thirty-four relational database schemas, we experimentally compared STICCER to two greedy test suite reduction techniques and a random method. The results show that, compared to the greedy and random methods, STICCER is the most effective at reducing the number of test cases and database interactions while maintaining test effectiveness as measured by the mutation score. Abdullah Alsharif, Gregory M. Kapfhammer, Phil McMinn |
ICST | 3 |
| 2020 | An Investigation into the Effect of Control and Data Dependence Paths on Predicate TestabilityabstractThe following topics are dealt with: program diagnostics; program verification; Java; software maintenance; program debugging; public domain software; Internet; program testing; learning (artificial intelligence); and object-oriented programming. Dave W. Binkley, James Glenn, Abdullah Alsharif, Phil McMinn |
SCAM | 4 |
| 2020 | Automatically identifying potential regressions in the layout of responsive web pagesabstractSummary Providing a good user experience on the ever‐increasing number and variety of devices being used to browse the web is a difficult, yet critical, task. With responsive web design, front‐end web developers design web pages so that they dynamically resize and rearrange content to best fit the dimensions of a device's screen. However, when making code modifications to a responsive page, developers can easily introduce regressions from the correct layout that have detrimental effects at unpredictable screen sizes. For instance, the source code change that a developer makes to improve the layout at one screen size may obscure a page's content at other sizes. Current approaches to testing are often insufficient because they rely on limited tools and error‐prone manual inspections of web pages. As such, many unintended regressions in web page layout often go undetected and ultimately manifest in production websites. To address the challenge of detecting regressions in responsive web pages, this paper presents an automated approach that extracts the responsive layout of two versions of a page and compares them, alerting developers to the differences in layout that they may wish to investigate further. We implemented the approach and empirically evaluated it on 15 real‐world responsive web pages. Leveraging code mutations that a tool automatically injected into the pages as a systematic simulation of developer changes, the experiments show that the approach was highly effective. When compared with manual and automated baseline testing techniques, it detected 12.5% and 18.75% more injected changes, respectively. Along with identifying the best parameters for the method that extracts the responsive layout, the experiments show that the approach surpasses the baselines across changes that vary in their impact, but works particularly well for subtle, hard‐to‐detect mutants, showing the benefits of automatically identifying regressions in web page layout. Thomas A. Walsh 0001, Gregory M. Kapfhammer, Phil McMinn |
Softw. Test. Verification Reliab. | 3 |
| 2019 | What Factors Make SQL Test Cases Understandable for Testers? A Human Study of Automated Test Data Generation TechniquesabstractSince relational databases are a key component of software systems ranging from small mobile to large enterprise applications, there are well-studied methods that automatically generate test cases for database-related functionality. Yet, there has been no research to analyze how well testers - who must often serve as an "oracle" - both understand tests involving SQL and decide if they reveal flaws. This paper reports on a human study of test comprehension in the context of automatically generated tests that assess the correct specification of the integrity constraints in a relational database schema. In this domain, a tool generates INSERT statements with data values designed to either satisfy (i.e., be accepted into the database) or violate the schema (i.e., be rejected from the database). The study reveals two key findings. First, the choice of data values in INSERTs influences human understandability: the use of default values for elements not involved in the test (but necessary for adhering to SQL's syntax rules) aided participants, allowing them to easily identify and understand the important test values. Yet, negative numbers and "garbage" strings hindered this process. The second finding is more far reaching: humans found the outcome of test cases very difficult to predict when NULL was used in conjunction with foreign keys and CHECK constraints. This suggests that, while including NULLs can surface the confusing semantics of database schemas, their use makes tests less understandable for humans. Abdullah Alsharif, Gregory M. Kapfhammer, Phil McMinn |
ICSME | 3 |
| 2019 | Automatic Visual Verification of Layout Failures in Responsively Designed Web PagesabstractResponsively designed web pages adjust their layout according to the viewport width of the device in use. Although tools exist to help developers test the layout of a responsive web page, they often rely on humans to flag problems. Yet, the considerable number of web-enabled devices with unique viewport widths makes this manual process both time-consuming and error-prone. Capable of detecting some common responsive layout failures, the ReDeCheck tool partially automates this process. Since ReDeCheck focuses on a web page's document object model (DOM), some of the issues it finds are not observable by humans. This paper presents a tool, called Viser, that renders a ReDeCheck-reported layout issue in a browser, adjusting the opacity of certain elements and checking for a visible difference. Unless Viser classifies an issue as a human-observable layout failure, a web developer can ignore it. This paper's experiments reveal the benefit of using Viser to support automated visual verification of layout failures in responsively designed web pages. Viser automatically classified all of the 117 layout failures that ReDeCheck reported for 20 web pages, each of which had to be manually analyzed in a prior study. Viser's automated manipulation of element opacity also highlighted manual classification's subjectivity: it categorized 28 issues differently to manual analysis, including three correctly reclassified as false positives. Ibrahim Althomali, Gregory M. Kapfhammer, Phil McMinn |
ICST | 3 |
| 2019 | An Empirical Study on the Use of Defect Prediction for Test Case PrioritizationabstractTest case prioritization has been extensively re-searched as a means for reducing the time taken to discover regressions in software. While many different strategies have been developed and evaluated, prior experiments have shown them to not be effective at prioritizing test suites to find real faults. This paper presents a test case prioritization strategy based on defect prediction, a technique that analyzes code features - such as the number of revisions and authors - to estimate the likelihood that any given Java class will contain a bug. Intuitively, if defect prediction can accurately predict the class that is most likely to be buggy, a tool can prioritize tests to rapidly detect the defects in that class. We investigated how to configure a defect prediction tool, called Schwa, to maximize the likelihood of an accurate prediction, surfacing the link between perfect defect prediction and test case prioritization effectiveness. Using 6 real-world Java programs containing 395 real faults, we conducted an empirical evaluation comparing this paper's strategy, called G-clef, against eight existing test case prioritization strategies. The experiments reveal that using defect prediction to prioritize test cases reduces the number of test cases required to find a fault by on average 9.48% when compared with existing coverage-based strategies, and 10.4% when compared with existing history-based strategies. David Paterson, José Campos 0001, Rui Abreu 0001, Gregory M. Kapfhammer, Gordon Fraser 0001, Phil McMinn |
ICST | 6 |
| 2019 | Automatic Detection and Removal of Ineffective Mutants for the Mutation Analysis of Relational Database SchemasabstractData is one of an organization's most valuable and strategic assets. Testing the relational database schema, which protects the integrity of this data, is of paramount importance. Mutation analysis is a means of estimating the fault-finding “strength” of a test suite. As with program mutation, however, relational database schema mutation results in many “ineffective” mutants that both degrade test suite quality estimates and make mutation analysis more time consuming. This paper presents a taxonomy of ineffective mutants for relational database schemas, summarizing the root causes of ineffectiveness with a series of key patterns evident in database schemas. On the basis of these, we introduce algorithms that automatically detect and remove ineffective mutants. In an experimental study involving the mutation analysis of 34 schemas used with three popular relational database management systems-HyperSQL, PostgreSQL, and SQLite-the results show that our algorithms can identify and discard large numbers of ineffective mutants that can account for up to 24 percent of mutants, leading to a change in mutation score for 33 out of 34 schemas. The tests for seven schemas were found to achieve 100 percent scores, indicating that they were capable of detecting and killing all non-equivalent mutants. The results also reveal that the execution cost of mutation analysis may be significantly reduced, especially with “heavyweight” DBMSs like PostgreSQL. Phil McMinn, Chris J. Wright, Colton J. McCurdy, Gregory M. Kapfhammer |
IEEE Trans. Software Eng. | 1 |
| 2018 | Automated repair of mobile friendly problems in web pagesabstractMobile devices have become a primary means of accessing the Internet. Unfortunately, many websites are not designed to be mobile friendly. This results in problems such as unreadable text, cluttered navigation, and content overflowing a device's viewport; all of which can lead to a frustrating and poor user experience. Existing techniques are limited in helping developers repair these mobile friendly problems. To address this limitation of prior work, we designed a novel automated approach for repairing mobile friendly problems in web pages. Our empirical evaluation showed that our approach was able to successfully resolve mobile friendly problems in 95% of the evaluation subjects. In a user study, participants preferred our repaired versions of the subjects and also considered the repaired pages to be more readable than the originals. Sonal Mahajan, Negarsadat Abolhassani, Phil McMinn, William G. J. Halfond |
ICSE | 3 |
| 2018 | DOMINO: Fast and Effective Test Data Generation for Relational Database Schemas
Abdullah Alsharif, Gregory M. Kapfhammer, Phil McMinn |
ICST | 3 |
| 2018 | Automated Repair of Internationalization Presentation Failures in Web Pages Using Style Similarity Clustering and Search-Based Techniques
Sonal Mahajan, Abdulmajeed Alameer, Phil McMinn, William G. J. Halfond |
ICST | 3 |
| 2018 | Search-based detection of deviation failures in the migration of legacy spreadsheet applicationsabstractMany legacy financial applications exist as a collection of formulas implemented in spreadsheets. Migration of these spreadsheets to a full-fledged system, written in a language such as Java, is an error- prone process. While small differences in the outputs of numerical calculations from the two systems are inevitable and tolerable, large discrepancies can have serious financial implications. Such discrepancies are likely due to faults in the migrated implementation, and are referred to as deviation failures. In this paper, we present a search-based technique that seeks to reveal deviation failures automatically. We evaluate different variants of this approach on two financial applications involving 40 formulas. These applications were produced by SEB Life & Pension Holding AB, who migrated their Microsoft Excel spreadsheets to a Java application. While traditional random and branch coverage-based test generation techniques were only able to detect approximately 25% and 32% of known faults in the migrated code respectively, our search-based approach detected up to 70% of faults with the same test generation budget. Without restriction of the search budget, up to 90% of known deviation failures were detected. In addition, three previously unknown faults were detected by this method that were confirmed by SEB experts. Mohammad Moein Almasi, Hadi Hemmati, Gordon Fraser 0001, Phil McMinn, Janis Benefelds |
ISSTA | 4 |
| 2018 | Random or evolutionary search for object-oriented test suite generation?abstractSummary An important aim in software testing is constructing a test suite with high structural code coverage, that is, ensuring that most if not all of the code under test have been executed by the test cases comprising the test suite. Several search‐based techniques have proved successful at automatically generating tests that achieve high coverage. However, despite the well‐established arguments behind using evolutionary search algorithms (eg, genetic algorithms) in preference to random search, it remains an open question whether the benefits can actually be observed in practice when generating unit test suites for object‐oriented classes. In this paper, we report an empirical study on the effects of using evolutionary algorithms (including a genetic algorithm and chemical reaction optimization) to generate test suites, compared with generating test suites incrementally with random search. We apply the EVOSUITEunit test suite generator to 1000 classes randomly selected from the SF110 corpus of open‐source projects. Surprisingly, the results show that the difference is much smaller than one might expect: While evolutionary search covers more branches of the type where standard fitness functions provide guidance, we observed that, in practice, the vast majority of branches do not provide any guidance to the search. These results suggest that, although evolutionary algorithms are more effective at covering complex branches, a random search may suffice to achieve high coverage of most object‐oriented classes. Sina Shamshiri, José Miguel Rojas, Luca Gazzola, Gordon Fraser 0001, Phil McMinn, Leonardo Mariani, Andrea Arcuri |
Softw. Test. Verification Reliab. | 5 |
| 2018 | Effectively Incorporating Expert Knowledge in Automated Software RemodularisationabstractRemodularising the components of a software system is challenging: sound design principles (e.g., coupling and cohesion) need to be balanced against developer intuition of which entities conceptually belong together. Despite this, automated approaches to remodularisation tend to ignore domain knowledge, leading to results that can be nonsensical to developers. Nevertheless, suppling such knowledge is a potentially burdensome task to perform manually. A lot information may need to be specified, particularly for large systems. Addressing these concerns, we propose the SUpervised reMOdularisation (SUMO) approach. SUMO is a technique that aims to leverage a small subset of domain knowledge about a system to produce a remodularisation that will be acceptable to a developer. With SUMO, developers refine a modularisation by iteratively supplying corrections. These corrections constrain the type of remodularisation eventually required, enabling SUMO to dramatically reduce the solution space. This in turn reduces the amount of feedback the developer needs to supply. We perform a comprehensive systematic evaluation using 100 real world subject systems. Our results show that SUMO guarantees convergence on a target remodularisation with a tractable amount of user interaction. Mathew Hall, Neil Walkinshaw, Phil McMinn |
IEEE Trans. Software Eng. | 3 |
| 2017 | Automated repair of layout cross browser issues using search-based techniquesabstractA consistent cross-browser user experience is crucial for the success of a website. Layout Cross Browser Issues (XBIs) can severely undermine a website’s success by causing web pages to render incorrectly in certain browsers, thereby negatively impacting users’ impression of the quality and services that the web page delivers. Existing Cross Browser Testing (XBT) techniques can only detect XBIs in websites. Repairing them is, hitherto, a manual task that is labor intensive and requires significant expertise. Addressing this concern, our paper proposes a technique for automatically repairing layout XBIs in websites using guided search-based techniques. Our empirical evaluation showed that our approach was able to successfully fix 86% of layout XBIs reported for 15 different web pages studied, thereby improving their cross-browser consistency. Sonal Mahajan, Abdulmajeed Alameer, Phil McMinn, William G. J. Halfond |
ISSTA | 3 |
| 2017 | XFix: an automated tool for the repair of layout cross browser issuesabstractDifferences in the rendering of a website across different browsers can cause inconsistencies in its appearance and usability, resulting in Layout Cross Browser Issues (XBIs). Such XBIs can negatively impact the functionality of a website as well as users’ impressions of its trustworthiness and reliability. Existing techniques can only detect XBIs, and therefore require developers to manually perform the labor intensive task of repair. In this demo paper we introduce our tool, XFix, that automatically repairs layout XBIs in web applications. To the best of our knowledge, XFix is the first automated technique for generating XBI repairs. Sonal Mahajan, Abdulmajeed Alameer, Phil McMinn, William G. J. Halfond |
ISSTA | 3 |
| 2017 | Automated layout failure detection for responsive web pages without an explicit oracleabstractAs the number and variety of devices being used to access the World Wide Web grows exponentially, ensuring the correct presentation of a web page, regardless of the device used to browse it, is an important and challenging task. When developers adopt responsive web design (RWD) techniques, web pages modify their appearance to accommodate a device’s display constraints. However, a current lack of automated support means that presentation failures may go undetected in a page’s layout when rendered for different viewport sizes. A central problem is the difficulty in providing an automated “oracle” to validate RWD layouts against, meaning that checking for failures is largely a manual process in practice, which results in layout failures in many live responsive web sites. This paper presents an automated failure detection technique that checks the consistency of a responsive page’s layout across a range of viewport widths, obviating the need for an explicit oracle. In an empirical study, this method found failures in 16 of 26 real-world production pages studied, detecting 33 distinct failures in total. Thomas A. Walsh 0001, Gregory M. Kapfhammer, Phil McMinn |
ISSTA | 3 |
| 2017 | ReDeCheck: an automatic layout failure checking tool for responsively designed web pagesabstractSince people frequently access websites with a wide variety of devices (e.g., mobile phones, laptops, and desktops), developers need frameworks and tools for creating layouts that are useful at many viewport widths. While responsive web design (RWD) principles and frameworks facilitate the development of such sites, there is a lack of tools supporting the detection of failures in their layout. Since the quality assurance process for responsively designed websites is often manual, time-consuming, and error-prone, this paper presents ReDeCheck, an automated layout checking tool that alerts developers to both potential unintended regressions in responsive layout and common types of layout failure. In addition to summarizing ReDeCheck’s benefits, this paper explores two different usage scenarios for this tool that is publicly available on GitHub. Thomas A. Walsh 0001, Gregory M. Kapfhammer, Phil McMinn |
ISSTA | 3 |
| 2017 | Evaluating CAVM: A New Search-Based Test Data Generation Tool for C
Junhwi Kim, Byeonghyeon You, Minhyuk Kwon, Phil McMinn, Shin Yoo |
SSBSE | 4 |
| 2016 | mrstudyr: Retrospectively Studying the Effectiveness of Mutant Reduction TechniquesabstractMutation testing is a well-known method for measuring a test suite's quality. However, due to its computational expense and intrinsic difficulties (e.g., detecting equivalent mutants and potentially checking a mutant's status for each test), mutation testing is often challenging to practically use. To control the computational cost of mutation testing, many reduction strategies have been proposed (e.g., uniform random sampling over mutants). Yet, a stand-alone tool to compare the efficiency and effectiveness of these methods is heretofore unavailable. Since existing mutation testing tools are often complex and language-dependent, this paper presents a tool, called mrstudyr, that enables the "retrospective" study of mutant reduction methods using the data collected from a prior analysis of all mutants. Focusing on the mutation operators and the mutants that they produce, the presented tool allows developers to prototype and evaluate mutant reducers without being burdened by the implementation details of mutation testing tools. Along with describing mrstudyr's design and overviewing the experimental results from using it, this paper inaugurates the public release of this open-source tool. Colton J. McCurdy, Phil McMinn, Gregory M. Kapfhammer |
ICSME | 2 |
| 2016 | SchemaAnalyst: Search-Based Test Data Generation for Relational Database SchemasabstractData stored in relational databases plays a vital role in many aspects of society. When this data is incorrect, the services that depend on it may be compromised. The database schema is the artefact responsible for maintaining the integrity of stored data. Because of its critical function, the proper testing of the database schema is a task of great importance. Employing a search-based approach to generate high-quality test data for database schemas, SchemaAnalyst is a tool that supports testing this key software component. This presented tool is extensible and includes both an evaluation framework for assessing the quality of the generated tests and full-featured documentation. In addition to describing the design and implementation of SchemaAnalyst and overviewing its efficiency and effectiveness, this paper coincides with the tool's public release, thereby enhancing practitioners' ability to test relational database schemas. Phil McMinn, Chris J. Wright, Cody Kinneer, Colton J. McCurdy, Michael Camara, Gregory M. Kapfhammer |
ICSME | 1 |
| 2016 | AVMf: An Open-Source Framework and Implementation of the Alternating Variable Method
Phil McMinn, Gregory M. Kapfhammer |
SSBSE | 1 |
| 2015 | Random or Genetic Algorithm Search for Object-Oriented Test Suite Generation?abstractAchieving high structural coverage is an important aim in software testing. Several search-based techniques have proved successful at automatically generating tests that achieve high coverage. However, despite the well- established arguments behind using evolutionary search algorithms (e.g., genetic algorithms) in preference to random search, it remains an open question whether the benefits can actually be observed in practice when generating unit test suites for object-oriented classes. In this paper, we report an empirical study on the effects of using a genetic algorithm (GA) to generate test suites over generating test suites incrementally with random search, by applying the EvoSuite unit test suite generator to 1,000 classes randomly selected from the SF110 corpus of open source projects. Surprisingly, the results show little difference between the coverage achieved by test suites generated with evolutionary search compared to those generated using random search. A detailed analysis reveals that the genetic algorithm covers more branches of the type where standard fitness functions provide guidance. In practice, however, we observed that the vast majority of branches in the analyzed projects provide no such guidance. Sina Shamshiri, José Miguel Rojas, Gordon Fraser 0001, Phil McMinn |
GECCO | 4 |
| 2015 | Do Automatically Generated Unit Tests Find Real Faults? An Empirical Study of Effectiveness and Challenges (T)abstractRather than tediously writing unit tests manually, tools can be used to generate them automatically - sometimes even resulting in higher code coverage than manual testing. But how good are these tests at actually finding faults? To answer this question, we applied three state-of-the-art unit test generation tools for Java (Randoop, EvoSuite, and Agitar) to the 357 real faults in the Defects4J dataset and investigated how well the generated test suites perform at detecting these faults. Although the automatically generated test suites detected 55.7% of the faults overall, only 19.9% of all the individual test suites detected a fault. By studying the effectiveness and problems of the individual tools and the tests they generate, we derive insights to support the development of automated unit test generators that achieve a higher fault detection rate. These insights include 1) improving the obtained code coverage so that faulty statements are executed in the first instance, 2) improving the propagation of faulty program states to an observable output, coupled with the generation of more sensitive assertions, and 3) improving the simulation of the execution environment to detect faults that are dependent on external factors such as date and time. Sina Shamshiri, René Just, José Miguel Rojas, Gordon Fraser 0001, Phil McMinn, Andrea Arcuri |
ASE | 5 |
| 2015 | Automatic Detection of Potential Layout Faults Following Changes to Responsive Web Pages (N)abstractDue to the exponential increase in the number ofmobile devices being used to access the World Wide Web, it iscrucial that Web sites are functional and user-friendly across awide range of Web-enabled devices. This necessity has resulted in the introduction of responsive Web design (RWD), which usescomplex cascading style sheets (CSS) to fluidly modify a Web site's appearance depending on the viewport width of the device in use. Although existing tools may support the testing of responsive Web sites, they are time consuming and error-prone to use because theyrequire manual screenshot inspection at specified viewport widths. Addressing these concerns, this paper presents a method thatcan automatically detect potential layout faults in responsively designed Web sites. To experimentally evaluate this approach, weimplemented it as a tool, called ReDeCheck, and applied itto 5 real-world web sites that vary in both their approach toresponsive design and their complexity. The experiments revealthat ReDeCheck finds 91% of the inserted layout faults. Thomas A. Walsh 0001, Phil McMinn, Gregory M. Kapfhammer |
ASE | 2 |
| 2015 | Automatically Evaluating the Efficiency of Search-Based Test Data Generation for Relational Database SchemasabstractThe characterization of an algorithm's worst-case time complexity is useful because it succinctly captures how its runtime will grow as the input size becomes arbitrarily large.However, for certain algorithms-such as those performing search-based test data generation-a theoretical analysis to determine worst-case time complexity is difficult to generalize and thus not often reported in the literature.This paper introduces a framework that empirically determines an algorithm's worst-case time complexity by doubling the size of the input and observing the change in runtime.Since the relational database is a centerpiece of modern software and the database's schema is frequently untested, we apply the doubling technique to the domain of data generation for relational database schemas, a field where worst-case time complexities are often unknown.In addition to demonstrating the feasibility of suggesting the worst-case runtimes of the chosen algorithms and configurations, the results of our study reveal performance tradeoffs in testing strategies for relational database schemas. Cody Kinneer, Gregory M. Kapfhammer, Phil McMinn, Chris J. Wright |
SEKE | 3 |
| 2015 | ExpOse: Inferring Worst-case Time Complexity by Automatic Empirical Study
Cody Kinneer, Gregory M. Kapfhammer, Chris J. Wright, Phil McMinn |
SEKE | 4 |
| 2015 | A Memetic Algorithm for whole test suite generationabstractThe generation of unit-level test cases for structural code coverage is a task well-suited to Genetic Algorithms. Method call sequences must be created that construct objects, put them into the right state and then execute uncovered code. However, the generation of primitive values, such as integers and doubles, characters that appear in strings, and arrays of primitive values, are not so straightforward. Often, small local changes are required to drive the value toward the one needed to execute some target structure. However, global searches like Genetic Algorithms tend to make larger changes that are not concentrated on any particular aspect of a test case. In this paper, we extend the Genetic Algorithm behind the EvoSuite test generation tool into a Memetic Algorithm, by equipping it with several local search operators. These operators are designed to efficiently optimize primitive values and other aspects of a test suite that allow the search for test cases to function more effectively. We evaluate our operators using a rigorous experimental methodology on over 12,000 Java classes, comprising open source classes of various different kinds, including numerical applications and text processors. Our study shows that increases in branch coverage of up to 53% are possible for an individual class in practice. Gordon Fraser 0001, Andrea Arcuri, Phil McMinn |
J. Syst. Softw. | 3 |
| 2015 | Automatic generation of valid and invalid test data for string validation routines using web searches and regular expressions
Muzammil Shahbaz, Phil McMinn, Mark Stevenson 0001 |
Sci. Comput. Program. | 2 |
| 2015 | Design and analysis of different alternating variable searches for search-based software testingabstractManual software testing is a notoriously expensive part of the software development process, and its automation is of high concern. One aspect of the testing process is the automatic generation of test inputs. This paper studies the Alternating Variable Method (AVM) approach to search-based test input generation. The AVM has been shown to be an effective and efficient means of generating branch-covering inputs for procedural programs. However, there has been little work that has sought to analyse the technique and further improve its performance. This paper proposes two different local searches that may be used in conjunction with the AVM, Geometric and Lattice Search. A theoretical runtime analysis proves that under certain conditions, the use of these searches results in better performance compared to the original AVM. These theoretical results are confirmed by an empirical study with five programs, which shows that increases of speed of over 50% are possible in practice. Joseph Kempka, Phil McMinn, Dirk Sudholt |
Theor. Comput. Sci. | 2 |
| 2015 | Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical StudyabstractWork on automated test generation has produced several tools capable of generating test data which achieves high structural coverage over a program. In the absence of a specification, developers are expected to manually construct or verify the test oracle for each test input. Nevertheless, it is assumed that these generated tests ease the task of testing for the developer, as testing is reduced to checking the results of tests. While this assumption has persisted for decades, there has been no conclusive evidence to date confirming it. However, the limited adoption in industry indicates this assumption may not be correct, and calls into question the practical value of test generation tools. To investigate this issue, we performed two controlled experiments comparing a total of 97 subjects split between writing tests manually and writing tests with the aid of an automated unit test generation tool, E vo S uite . We found that, on one hand, tool support leads to clear improvements in commonly applied quality metrics such as code coverage (up to 300% increase). However, on the other hand, there was no measurable improvement in the number of bugs actually found by developers. Our results not only cast some doubt on how the research community evaluates test generation tools, but also point to improvements and future work necessary before automated test generation tools will be widely adopted by practitioners. Gordon Fraser 0001, Matthew Staats, Phil McMinn, Andrea Arcuri, Frank Padberg |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2015 | The Effectiveness of Test Coverage Criteria for Relational Database Schema Integrity ConstraintsabstractDespite industry advice to the contrary, there has been little work that has sought to test that a relational database's schema has correctly specified integrity constraints. These critically important constraints ensure the coherence of data in a database, defending it from manipulations that could violate requirements such as “usernames must be unique” or “the host name cannot be missing or unknown.” This article is the first to propose coverage criteria, derived from logic coverage criteria, that establish different levels of testing for the formulation of integrity constraints in a database schema. These range from simple criteria that mandate the testing of successful and unsuccessful INSERT statements into tables to more advanced criteria that test the formulation of complex integrity constraints such as multi-column PRIMARY KEYs and arbitrary CHECK constraints. Due to different vendor interpretations of the structured query language (SQL) specification with regard to how integrity constraints should actually function in practice, our criteria crucially account for the underlying semantics of the database management system (DBMS). After formally defining these coverage criteria and relating them in a subsumption hierarchy, we present two approaches for automatically generating tests that satisfy the criteria. We then describe the results of an empirical study that uses mutation analysis to investigate the fault-finding capability of data generated when our coverage criteria are applied to a wide variety of relational schemas hosted by three well-known and representative DBMSs—HyperSQL, PostgreSQL, and SQLite. In addition to revealing the complementary fault-finding capabilities of the presented criteria, the results show that mutation scores range from as low as just 12% of mutants being killed with the simplest of criteria to 96% with the most advanced. Phil McMinn, Chris J. Wright, Gregory M. Kapfhammer |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2015 | The Oracle Problem in Software Testing: A SurveyabstractTesting involves examining the behaviour of a system in order to discover potential faults. Given an input for a system, the challenge of distinguishing the corresponding desired, correct behaviour from potentially incorrect behavior is called the “test oracle problem”. Test oracle automation is important to remove a current bottleneck that inhibits greater overall test automation. Without test oracle automation, the human has to determine whether observed behaviour is correct. The literature on test oracles has introduced techniques for oracle automation, including modelling, specifications, contract-driven development and metamorphic testing. When none of these is completely adequate, the final source of test oracle information remains the human, who may be aware of informal specifications, expectations, norms and domain specific information that provide informal oracle guidance. All forms of test oracles, even the humble human, involve challenges of reducing cost and increasing benefit. This paper provides a comprehensive survey of current approaches to the test oracle problem and an analysis of trends in this important area of software testing research and practice. Earl T. Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, Shin Yoo |
IEEE Trans. Software Eng. | 3 |
| 2014 | Establishing the Source Code Disruption Caused by Automated Remodularisation ToolsabstractCurrent software remodularisation tools only operate on abstractions of a software system. In this paper, we investigate the actual impact of automated remodularisation on source code using a tool that automatically applies remodularisations as refactorings. This shows us that a typical remodularisation (as computed by the Bunch tool) will require changes to thousands of lines of code, spread throughout the system (typically no code files remain untouched). In a typical multi-developer project this presents a serious integration challenge, and could contribute to the low uptake of such tools in an industrial context. We relate these findings with our ongoing research into techniques that produce iterative commit friendly" code changes to address this problem. Mathew Hall, Muhammad Ali Khojaye, Neil Walkinshaw, Phil McMinn |
ICSME | 4 |
| 2013 | Test suite generation with memetic algorithmsabstractGenetic Algorithms have been successfully applied to the generation of unit tests for classes, and are well suited to create complex objects through sequences of method calls. However, because the neighborhood in the search space for method sequences is huge, even supposedly simple optimizations on primitive variables (e.g., numbers and strings) can be ineffective or unsuccessful. To overcome this problem, we extend the global search applied in the EVOSUITE test generation tool with local search on the individual statements of method sequences. In contrast to previous work on local search, we also consider complex datatypes including strings and arrays. A rigorous experimental methodology has been applied to properly evaluate these new local search operators. In our experiments on a set of open source classes of different kinds (e.g., numerical applications and text processing), the resulting test data generation technique increased branch coverage by up to 32% on average over the normal Genetic Algorithm. Gordon Fraser 0001, Andrea Arcuri, Phil McMinn |
GECCO | 3 |
| 2013 | A theoretical runtime and empirical analysis of different alternating variable searches for search-based testingabstractThe Alternating Variable Method (AVM) has been shown to be a surprisingly effective and efficient means of generating branch-covering inputs for procedural programs. However, there has been little work that has sought to analyse the technique and further improve its performance. This paper proposes two new local searches that may be used in conjunction with the AVM, Geometric and Lattice Search. A theoretical runtime analysis shows that under certain conditions, the use of these searches is proven to outperform the original AVM. These theoretical results are confirmed by an empirical study with four programs, which shows that increases of speed of over 50% are possible in practice. Joseph Kempka, Phil McMinn, Dirk Sudholt |
GECCO | 2 |
| 2013 | Evolving Readable String Test Inputs Using a Natural Language Model to Reduce Human Oracle CostabstractThe frequent non-availability of an automated oracle means that, in practice, checking software behaviour is frequently a painstakingly manual task. Despite the high cost of human oracle involvement, there has been little research investigating how to make the role easier and less time-consuming. One source of human oracle cost is the inherent unreadability of machine-generated test inputs. In particular, automatically generated string inputs tend to be arbitrary sequences of characters that are awkward to read. This makes test cases hard to comprehend and time-consuming to check. In this paper we present an approach in which a natural language model is incorporated into a search-based input data generation process with the aim of improving the human readability of generated strings. We further present a human study of test inputs generated using the technique on 17 open source Java case studies. For 10 of the case studies, the participants recorded significantly faster times when evaluating inputs produced using the language model, with medium to large effect sizes 60% of the time. In addition, the study found that accuracy of test input evaluation was also significantly improved for 3 of the case studies. Sheeva Afshan, Phil McMinn, Mark Stevenson 0001 |
ICST | 2 |
| 2013 | Search-Based Testing of Relational Schema Integrity Constraints Across Multiple Database Management SystemsabstractThere has been much attention to testing applications that interact with database management systems, and the testing of individual database management systems themselves. However, there has been very little work devoted to testing arguably the most important artefact involving an application supported by a relational database - the underlying schema. This paper introduces a search-based technique for generating database table data with the intention of exercising the integrity constraints placed on table columns. The development of a schema is a process open to flaws like any stage of application development. Its cornerstone nature to an application means that defects need to be found early in order to prevent knock-on effects to other parts of a project and the spiralling bug-fixing costs that may be incurred. Examples of such flaws include incomplete primary keys, incorrect foreign keys, and omissions of NOT NULL declarations. Using mutation analysis, this paper presents an empirical study evaluating the effectiveness of our proposed technique and comparing it against a popular tool for generating table data, DBMonster. With competitive or faster data generation times, our method outperforms DBMonster in terms of both constraint coverage and mutation score. Gregory M. Kapfhammer, Phil McMinn, Chris J. Wright |
ICST | 2 |
| 2013 | Does automated white-box test generation really help software testers?abstractAutomated test generation techniques can efficiently produce test data that systematically cover structural aspects of a program. In the absence of a specification, a common assumption is that these tests relieve a developer of most of the work, as the act of testing is reduced to checking the results of the tests. Although this assumption has persisted for decades, there has been no conclusive evidence to date confirming it. However, the fact that the approach has only seen a limited uptake in industry suggests the contrary, and calls into question its practical usefulness. To investigate this issue, we performed a controlled experiment comparing a total of 49 subjects split between writing tests manually and writing tests with the aid of an automated unit test generation tool, EvoSuite. We found that, on one hand, tool support leads to clear improvements in commonly applied quality metrics such as code coverage (up to 300% increase). However, on the other hand, there was no measurable improvement in the number of bugs actually found by developers. Our results not only cast some doubt on how the research community evaluates test generation tools, but also point to improvements and future work necessary before automated test generation tools will be widely adopted by practitioners. Gordon Fraser 0001, Matthew Staats, Phil McMinn, Andrea Arcuri, Frank Padberg |
ISSTA | 3 |
| 2013 | An identification of program factors that impact crossover performance in evolutionary test input generation for the branch coverage of C programs
Phil McMinn |
Inf. Softw. Technol. | 1 |
| 2013 | An orchestrated survey of methodologies for automated software test case generation
Saswat Anand, Edmund K. Burke, Tsong Yueh Chen, John A. Clark, Myra B. Cohen, Wolfgang Grieskamp, Mark Harman, Mary Jean Harrold, Phil McMinn |
J. Syst. Softw. | 9 |
| 2012 | Supervised software modularisationabstractThis paper is concerned with the challenge of reorganising a software system into modules that both obey sound design principles and are sensible to domain experts. The problem has given rise to several unsupervised automated approaches that use techniques such as clustering and Formal Concept Analysis. Although results are often partially correct, they usually require refinement to enable the developer to integrate domain knowledge. This paper presents the SUMO algorithm, an approach that is complementary to existing techniques and enables the maintainer to refine their results. The algorithm is guaranteed to eventually yield a result that is satisfactory to the maintainer, and the evaluation on a diverse range of systems shows that this occurs with a reasonably low amount of effort. Mathew Hall, Neil Walkinshaw, Phil McMinn |
ICSM | 3 |
| 2012 | Search-Based Test Input Generation for String Data Types Using the Results of Web QueriesabstractGenerating realistic, branch-covering string inputs is a challenging problem, due to the diverse and complex types of real-world data that are naturally encodable as strings, for example resource locators, dates of different localised formats, international banking codes, and national identity numbers. This paper presents an approach in which examples of inputs are sought from the Internet by reformulating program identifiers into web queries. The resultant URLs are downloaded, split into tokens, and used to augment and seed a search-based test data generation technique. The use of the Internet as part of test input generation has two key advantages. Firstly, web pages are a rich source of valid inputs for various types of string data that may be used to improve test coverage. Secondly, the web pages tend to contain realistic, human-readable values, which are invaluable when test cases need manual confirmation due to the lack of an automated oracle. An empirical evaluation of the approach is presented, involving string input validation code from 10 open source projects. Well-formed, valid string inputs were retrieved from the web for 96% of the different string types analysed. Using the approach, coverage was improved for 75% of the Java classes studied by an average increase of 14%. Phil McMinn, Muzammil Shahbaz, Mark Stevenson 0001 |
ICST | 1 |
| 2012 | Towards the Automatic Identification of Faulty Multi-Agent Based Simulation Runs Using MASTER
Chris J. Wright, Phil McMinn, Julio Gallardo |
MABS | 2 |
| 2012 | Input Domain Reduction through Irrelevant Variable Removal and Its Effect on Local, Global, and Hybrid Search-Based Structural Test Data GenerationabstractSearch-Based Test Data Generation reformulates testing goals as fitness functions so that test input generation can be automated by some chosen search-based optimization algorithm. The optimization algorithm searches the space of potential inputs, seeking those that are “fit for purpose,” guided by the fitness function. The search space of potential inputs can be very large, even for very small systems under test. Its size is, of course, a key determining factor affecting the performance of any search-based approach. However, despite the large volume of work on Search-Based Software Testing, the literature contains little that concerns the performance impact of search space reduction. This paper proposes a static dependence analysis derived from program slicing that can be used to support search space reduction. The paper presents both a theoretical and empirical analysis of the application of this approach to open source and industrial production code. The results provide evidence to support the claim that input domain reduction has a significant effect on the performance of local, global, and hybrid search, while a purely random search is unaffected. Phil McMinn, Mark Harman, Kiran Lakhotia, Youssef Hassoun, Joachim Wegener |
IEEE Trans. Software Eng. | 1 |
| 2011 | A multiobjective optimisation approach for the dynamic inference and refinement of agent-based model specificationsabstractDespite their increasing popularity, agent-based models are hard to test, and so far no established testing technique has been devised for this kind of software applications. Reverse engineering an agent-based model specification from model simulations can help establish a confidence level about the implemented model and in some cases reveal discrepancies between observed and normal or expected behaviour. In this study, a multiobjective optimisation technique based on a simple random search algorithm is deployed to dynamically infer and refine the specification of three agent-based models from their simulations. The multiobjective optimisation technique also incorporates a dynamic invariant detection technique which serves to guide the search towards uncovering new model behaviour that better captures the model specification. The Non-dominated Sorting Genetic Algorithm (NSGA-II) was also deployed to replace the random search algorithm, and the results from both approaches were compared. While both algorithms revealed good potential in capturing the model specifications, the pure exploratory nature of random search was found more suitable for the application at hand, compared to the balanced exploitation/exploration nature of genetic algorithms in general. Salem Fawaz Adra, Mariam Kiran, Phil McMinn, Neil Walkinshaw |
IEEE Congress on Evolutionary Computation | 3 |
| 2011 | Symbolic search-based testingabstractWe present an algorithm for constructing fitness functions that improve the efficiency of search-based testing when trying to generate branch adequate test data. The algorithm combines symbolic information with dynamic analysis and has two key advantages: It does not require any change in the underlying test data generation technique and it avoids many problems traditionally associated with symbolic execution, in particular the presence of loops. We have evaluated the algorithm on industrial closed source and open source systems using both local and global search-based testing techniques, demonstrating that both are statistically significantly more efficient using our approach. The test for significance was done using a one-sided, paired Wilcoxon signed rank test. On average, the local search requires 23.41% and the global search 7.78% fewer fitness evaluations when using a symbolic execution based fitness function generated by the algorithm. Arthur I. Baars, Mark Harman, Youssef Hassoun, Kiran Lakhotia, Phil McMinn, Paolo Tonella, Tanja E. J. Vos |
ASE | 5 |
| 2010 | Superstate identification for state machines using search-based clusteringabstractState machines are a popular method of representing a system at a high level of abstraction that enables developers to gain an overview of the system they represent and quickly understand it. Mathew Hall, Phil McMinn, Neil Walkinshaw |
GECCO | 2 |
| 2010 | An empirical investigation into branch coverage for C programs using CUTE and AUSTIN
Kiran Lakhotia, Phil McMinn, Mark Harman |
J. Syst. Softw. | 2 |
| 2010 | A Theoretical and Empirical Study of Search-Based Testing: Local, Global, and Hybrid SearchabstractSearch-based optimization techniques have been applied to structural software test data generation since 1992, with a recent upsurge in interest and activity within this area. However, despite the large number of recent studies on the applicability of different search-based optimization approaches, there has been very little theoretical analysis of the types of testing problem for which these techniques are well suited. There are also few empirical studies that present results for larger programs. This paper presents a theoretical exploration of the most widely studied approach, the global search technique embodied by Genetic Algorithms. It also presents results from a large empirical study that compares the behavior of both global and local search-based optimization on real-world programs. The results of this study reveal that cases exist of test data generation problem that suit each algorithm, thereby suggesting that a hybrid global-local search (a Memetic Algorithm) may be appropriate. The paper presents a Memetic Algorithm along with further empirical results studying its performance. Mark Harman, Phil McMinn |
IEEE Trans. Software Eng. | 2 |
| 2009 | Search-based failure discovery using testability transformations to generate pseudo-oraclesabstractTestability transformations are source-to-source program transformations that are designed to improve the testability of a program. This paper introduces a novel approach in which transformations are used to improve testability of a program by generating a pseudo-oracle. A pseudo-oracle is an alternative version of a program under test whose output can be compared with the original. Differences in output between the two programs may indicate a fault in the original program. Two transformations are presented. The first can highlight numerical inaccuracies in programs and cumulative roundoff errors, whilst the second may detect the presence of race conditions in multi-threaded code. Once a pseudo-oracle is generated, techniques are applied from the field of search-based testing to automatically find differences in output between the two versions of the program. The results of an experimental study presented in the paper show that both random testing and genetic algorithms are capable of utilizing the pseudo-oracles to automatically find program failures. Using genetic algorithms it is possible to explicitly maximize the discrepancies between the original programs and their pseudo-oracles. This allows for the production of test cases where the observable failure is highly pronounced, enabling the tester to establish the seriousness of the underlying fault. Phil McMinn |
GECCO | 1 |
| 2009 | TAIC PART 2007 and Mutation 2007 special issue editorial
Mark Harman, Zheng Li 0002, Phil McMinn, A. Jefferson Offutt, John A. Clark |
J. Syst. Softw. | 3 |
| 2009 | Empirical evaluation of a nesting testability transformation for evolutionary testingabstractEvolutionary testing is an approach to automating test data generation that uses an evolutionary algorithm to search a test object's input domain for test data. Nested predicates can cause problems for evolutionary testing, because information needed for guiding the search only becomes available as each nested conditional is satisfied. This means that the search process can overfit to early information, making it harder, and sometimes near impossible, to satisfy constraints that only become apparent later in the search. The article presents a testability transformation that allows the evaluation of all nested conditionals at once. Two empirical studies are presented. The first study shows that the form of nesting handled is prevalent in practice. The second study shows how the approach improves evolutionary test data generation. Phil McMinn, Dave W. Binkley, Mark Harman |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2008 | Handling dynamic data structures in search based testingabstractThere has been little attention to search based test data generation in the presence of pointer inputs and dynamic data structures, an area in which recent concolic methods have excelled. This paper introduces a search based testing approach which is able to handle pointers and dynamic data structures. It combines an alternating variable hill climb with a set of constraint solving rules for pointer inputs. The result is a lightweight and efficient method, as shown in the results from a case study, which compares the method to CUTE, a concolic unit testing tool. Kiran Lakhotia, Mark Harman, Phil McMinn |
GECCO | 3 |
| 2008 | Editorial: Testing practice and researchabstractThe first ‘Testing: Academic & Industrial Conference—Practice and Research Techniques’ (TAIC PART 2006) was held at Cumberland Lodge Windsor during 29–31 August 2006. The general chair was Mark Harman (King's College London), the programme chair was Phil McMinn (University of Sheffield), and the local arrangements chair was Zheng Li (King's College London). TAIC PART is firmly grounded in fostering collaboration between industry and academia. It aims to bring together industrial software developers and users together with academic researchers working on the theory and practice of software testing. TAIC PART 2006 was a unique, not-for-profit conference that combined what we believe to be the best aspects of three kinds of event: a formal academic conference, a research workshop, and a retreat. The aim was to act not only as a forum for the exchange of ideas, but also as a vehicle to stimulate, deepen, and widen partnership between the academia and industry in software testing internationally. In all, 54 delegates from 11 different countries, comprising 32 academics and 22 industrialists, attended the conference. The event featured two keynotes, regular paper sessions, a PhD symposium, and a ‘speed dating’ session devoted to stimulating collaboration between attendees. The first keynote was given by Bill Woodworth, Corporate Director of IBM Quality Software Engineering. Bill spoke on test management at IBM, of which he has over 25 years of experience. John Hatcliff, who delivered the second keynote, is Professor in the Computing and Information Sciences Department at Kansas State University. John spoke on his internationally leading work on testing and software model checking. TAIC PART 2006 received a total of 50 full-paper submissions. After a rigorous reviewing process, 24 papers were accepted, with 8 of those papers from industry, 10 papers containing academic research, and a further 6 short papers accepted for the special PhD programme. Accepted papers covered a wide spectrum of state-of-the-art testing practice and research, including fault prediction, model-based testing, test specifications, the testing life cycle, search-based testing, database testing, web service testing, test requirements analysis, integration testing, empirical and case studies, and industrial challenges. The conference proceedings were published by IEEE and are available online. Two papers in this special issue are extended versions of some of the best papers originally presented at the conference. They have also been through a further reviewing process. The first paper is the result of an academic–industrial collaboration. In their paper ‘Quality Assurance for TTCN-3 Test Specifications’, Helmut Neukirchen, Benjamin Zeiss, and Jens Grabowski, of the University of Göttingen, and Paul Baker and Dominic Evans, of Motorola Labs, propose a technique to (1) assess the quality of existing test suites through metrics; (2) improve it through refactoring; and (3) detect refactoring opportunities by means of rules that are based on quality metrics. They focus on test suites expressed in the Testing and Test Control Notation TTCN-3, a language designed to support the specification of test suites in the telecommunication domain. The quality attribute considered in this work is maintainability, decomposed into analysability and changeability. Size, complexity, and coupling metrics are defined to characterize such quality attributes. The refactoring catalogue includes 23 TTCN-3-specific refactorings and 28 Java refactorings that are applicable to TTCN-3 as well. Eight rules are defined to check the applicability of refactorings automatically. These are implemented in a tool called ‘TRex’. The second paper, entitled ‘Automated Discovery of State Transitions and their Functions in Source Code’, by Neil Walkinshaw, Shaukat Ali, Kirill Bogdanov, and Mike Holcombe, presents a technique to reverse engineer source code into a state machine. It allows a developer to identify the states at a given point and statements that are responsible for state transitions. The technique also combines several ingredients, including symbolic execution and state abstraction, and is demonstrated with examples. Finally, we are grateful to our sponsors, whose financial contributions made it possible for TAIC PART 2006 to happen. Funding was received from the EPSRC and also from industry, including support from DaimlerChrysler, Ericsson, IPL Ltd., LDRA Ltd., Motorola, and Vizuri. The TAIC PART website (http://www2006.taicpart.org) serves as lasting resource to the event, containing the programme, photographs, and downloadable presentations of all the talks. Mark Harman, Zheng Li 0002, Phil McMinn |
Softw. Test. Verification Reliab. | 3 |
| 2007 | A multi-objective approach to search-based test data generationabstractThere has been a considerable body of work on search-based test data generation for branch coverage. However, hitherto, there has been no work on multi-objective branch coverage. In many scenarios a single-objective formulation is unrealistic; testers will want to find test sets that meet several objectives simultaneously in order to maximize the value obtained from the inherently expensive process of running the test cases and examining the output they produce. This paper introduces multi-objective branch coverage.The paper presents results from a case study of the twin objectives of branch coverage and dynamic memory consumption for both real and synthetic programs. Several multi-objective evolutionary algorithms are applied. The results show that multi-objective evolutionary algorithms are suitable for this problem, and illustrates the way in which a Pareto optimal search can yield insights into the trade-offs between the two simultaneous objectives. Kiran Lakhotia, Mark Harman, Phil McMinn |
GECCO | 3 |
| 2007 | A theoretical & empirical znalysis of evolutionary testing and hill climbing for structural test data generationabstractEvolutionary testing has been widely studied as a technique for automating the process of test case generation. However, to date, there has been no theoretical examination of when and why it works. Furthermore, the empirical evidence for the effectiveness of evolutionary testing consists largely of small scale laboratory studies. This paper presents a first theoretical analysis of the scenarios in which evolutionary algorithms are suitable for structural test case generation. The theory is backed up by an empirical study that considers real world programs, the search spaces of which are several orders of magnitude larger than those previously considered. Mark Harman, Phil McMinn |
ISSTA | 2 |
| 2007 | The impact of input domain reduction on search-based test data generationabstractThere has recently been a great deal of interest in search-based test data generation, with many local and global search algorithms being proposed. However, to date, there has been no investigation ofthe relationship between the size of the input domain (the search space) and performance of search-based algorithms. Static analysis can be used to remove irrelevant variables for a given test data generation problem, thereby reducing the search space size. This paper studies the effect of this domain reduction, presenting results from the application of local and global search algorithms to real world examples. This provides evidence to support the claimthat domain reduction has implications for practical search-based test data generation. Mark Harman, Youssef Hassoun, Kiran Lakhotia, Phil McMinn, Joachim Wegener |
ESEC/SIGSOFT FSE | 4 |
| 2006 | The species per path approach to SearchBased test data generationabstractThis paper introduces the Species per Path approach to search-based software test data generation. The approach transforms the program under test into a version in which multiple paths to the search target are factored out. Test data are then sought for each individual path by dedicated 'species' operating in parallel. The factoring out of paths results in several individual search landscapes, with feasible paths giving rise to landscapes that are potentially more conducive to test data discovery than the original overall landscape.The paper presents the results of two empirical studies that validate and verify the approach. The validation study supports the claim that the approach is widely applicable and practical. The verification study shows that it is possible to generate test data for targets with the approach that are troublesome for the standard evolutionary method. Phil McMinn, Mark Harman, Dave W. Binkley, Paolo Tonella |
ISSTA | 1 |
| 2006 | Evolutionary Testing Using an Extended Chaining ApproachabstractFitness functions derived from certain types of white-box test goals can be inadequate for evolutionary software test data generation (Evolutionary Testing), due to a lack of search guidance to the required test data. Often this is because the fitness function does not take into account data dependencies within the program under test, and the fact that certain program statements may need to have been executed prior to the target structure in order for it to be feasible. This paper proposes a solution to this problem by hybridizing Evolutionary Testing with an extended Chaining Approach. The Chaining Approach is a method which identifies statements on which the target structure is data dependent, and incrementally develops chains of dependencies in an event sequence. By incorporating this facility into Evolutionary Testing, and by performing a test data search for each generated event sequence, the search can be directed into potentially promising, unexplored areas of the test object's input domain. Results presented in the paper show that test data can be found for a number of test goals with this hybrid approach that could not be found by using the original Evolutionary Testing approach alone. One such test goal is drawn from code found in the publicly available libpng library. Phil McMinn, Mike Holcombe |
Evol. Comput. | 1 |
| 2006 | Editorial: Addressing industrial challenges - UKTest 2005 and beyond
Phil McMinn, Robert M. Hierons |
Softw. Test. Verification Reliab. | 1 |
| 2005 | Evolutionary testing of state-based programsabstractThe application of Evolutionary Algorithms to structural test data generation, known as Evolutionary Testing, has to date largely focused on programs with input-output behavior. However, the existence of state behavior in test objects presents additional challenges for Evolutionary Testing, not least because certain test goals may require a search for a sequence of inputs to the test object. Furthermore, state-based test objects often make use of internal variables such as boolean flags, enumerations and counters for managing or querying their internal state. These types of variables can lead to a loss of information in computing fitness values, producing coarse or flat fitness landscapes. This results in the search receiving less guidance, and the chances of finding required test data are decreased.This paper proposes an extended approach based on previous works. Input sequences are generated, and internal variable problems are addressed through hybridization with an extended Chaining Approach. The basic idea of the Chaining Approach is to find a sequence of statements, involving internal variables, which need to be executed prior to the test goal. By requiring these statements are executed, information previously unavailable to the search can be made use of, possibly guiding it into potentially promising and unexplored areas of the test object's input domain. A number of experiments demonstrate the value of the approach. Phil McMinn, Mike Holcombe |
GECCO | 1 |
| 2004 | Hybridizing Evolutionary Testing with the Chaining Approach
Phil McMinn, Mike Holcombe |
GECCO (2) | 1 |
| 2004 | Search-based software test data generation: a surveyabstractAbstract The use of metaheuristic search techniques for the automatic generation of test data has been a burgeoning interest for many researchers in recent years. Previous attempts to automate the test generation process have been limited, having been constrained by the size and complexity of software, and the basic fact that, in general, test data generation is an undecidable problem. Metaheuristic search techniques offer much promise in regard to these problems. Metaheuristic search techniques are high‐level frameworks, which utilize heuristics to seek solutions for combinatorial problems at a reasonable computational cost. To date, metaheuristic search techniques have been applied to automate test data generation for structural and functional testing; the testing of grey‐box properties, for example safety constraints; and also non‐functional properties, such as worst‐case execution time. This paper surveys some of the work undertaken in this field, discussing possible new future directions of research for each of its different individual areas. Copyright © 2004 John Wiley & Sons, Ltd. Phil McMinn |
Softw. Test. Verification Reliab. | 1 |
| 2003 | The State Problem for Evolutionary Testing
Phil McMinn, Mike Holcombe |
GECCO | 1 |