VLDB 2026 Research / reviewers in the wild / expert
Neil Walkinshaw
dblp:36/3976
· DBLP profile ↗
46ranked-venue papers
15as first author
13since 2021 · last 2025
0000-0003-2134-6548ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 42 · 15 first-author · 13 since 2021Artificial intelligence and machine learning · 6Theory of computation · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Using Causal Inference to Test Systems with Hidden and Interacting Variables: An Evaluative Case StudyabstractSoftware systems with large parameter spaces, nondeterminism and high computational cost are challenging to test. Recently, software testing techniques based on causal inference have been successfully applied to systems that exhibit such characteristics, including scientific models and autonomous driving systems. One significant limitation is that these are restricted to test properties where all of the variables involved can be observed and where there are no interactions between variables. In practice, this is rarely guaranteed; the logging infrastructure may not be available to record all of the necessary runtime variable values, and it can often be the case that an output of the system can be affected by complex interactions between variables. To address this, we leverage two additional concepts from causal inference, namely effect modification and instrumental variable methods. We build these concepts into an existing causal testing tool and conduct an evaluative case study which uses the concepts to test three system-level requirements of CARLA, a high-fidelity driving simulator widely used in autonomous vehicle development and testing. The results show that we can obtain reliable test outcomes without requiring large amounts of highly controlled test data or instrumentation of the code, even when variables interact with each other and are not recorded in the test data. Michael Foster 0001, Robert M. Hierons, Donghwan Shin 0001, Neil Walkinshaw, Christopher Wild |
EASE | 4 |
| 2025 | Exploratory Software Testing in Scrum: A Qualitative StudyabstractAbstract Exploratory Testing (ET) is a dynamic software testing approach that emphasises creativity, real-time learning, and defect discovery. The integration of ET into structured frameworks like Scrum remains insufficiently explored and presents distinct challenges. This qualitative study investigates how ET is implemented in Scrum workflows and identifies key factors enabling its effective application. Interviews with 20 industry professionals highlight ET’s role in enhancing test coverage, uncovering usability issues, and addressing edge cases often missed by automated or scripted tests. The results demonstrated that the critical enablers of effective ET are the tester’s eagerness to learn about the system under test and the ability to adopt a user-centric perspective. Other key factors include testers’ curiosity, creativity, domain knowledge, and organisational support. Participants noted that ET complements Scrum’s iterative cycles, enabling teams to identify defects dynamically and improve software quality. Despite its advantages, ET faces challenges within Scrum, including time constraints and the need for traceability. Lightweight documentation practices, such as annotated mind maps and screen recordings, emerged as effective strategies to bridge these gaps. This study underscores ET’s potential to enhance Scrum workflows, providing actionable insights for optimising testing strategies in Agile environments. Giulia Neri, Rob Marchand, Neil Walkinshaw |
XP | 3 |
| 2025 | Configuration Testing of an Artificial Pancreas System Using a Digital Twin: An Evaluative Case StudyabstractABSTRACT The recent growth in popularity of wearable medical devices has improved the quality of life of people with medical conditions. Testing such devices may require users to configure these systems using physical trials, putting themselves in potentially dangerous scenarios. Misconfiguration of such devices has caused disease misdiagnoses and incorrect drug prescriptions. Digital twins have been proposed as an opportunity to reduce such risks of testing system configurations in simulated environments, decoupling the user from the system under test. In this paper, we perform an evaluative case study to assess the use of a digital twin for configuration testing of an artificial pancreas system (APS) control algorithm. These systems regulate the blood glucose levels in people with type 1 diabetes mellitus, and so misconfigurations can cause severe hypoglycaemia or hyperglycaemia, which can be life‐threatening. We tested the OpenAPS control algorithm against 156 people's clinical data. We found that our digital twin provided an accurate simulation environment to perform configuration testing and accurately predict blood glucose–insulin behaviour. We evaluated different APS configurations, identifying a potentially unsafe configuration without the risks associated with a physical trial. We identified the challenges associated with modelling clinical data, which could lead to misinterpretations in configuration testing and the reduction of test reliability when modelling stochastic body dynamics. Richard J. Somers, Neil Walkinshaw, Robert M. Hierons, Jackie Elliott, Ahmed Iqbal, Emma Walkinshaw |
Softw. Test. Verification Reliab. | 2 |
| 2024 | Causal Test AdequacyabstractCausal reasoning is becoming an increasingly popular technique for testing software. In this setting, the tester starts from a simple directed graph that captures their underlying understanding of causal relationships between relevant variables in the program, and this knowledge is then used to reason about causal input-output relationships that are observed during testing. One question that has not yet been addressed in this context is how to measure test adequacy: How do we know whether a causal relationship (or set of relationships) has been properly established by a test set? In this paper we present a metric inspired by Weyuker's notion of inference adequacy. For a given causal relationship, we estimate the causal effect from the test data. The basis of our adequacy metric is then an estimate of the convergence of this estimate, which we calculate using statistical bootstrapping. We evaluate our metric on tests for three diverse computational models. The results show a statistically significant correlation between our metric and a test suite's ability to detect mutants, and also that it is a good indicator of whether a sufficient number of system executions have been observed to trust the outcome of the test. Michael Foster 0001, Christopher Wild, Robert M. Hierons, Neil Walkinshaw |
ICST | 4 |
| 2024 | Autonomous Driving System Testing: Traffic Density Does Matter
Guannan Lou, Donghwan Shin 0001, Neil Walkinshaw, Robert M. Hierons |
ICTSS | 3 |
| 2024 | Testing Causality in Scientific Modelling SoftwareabstractFrom simulating galaxy formation to viral transmission in a pandemic, scientific models play a pivotal role in developing scientific theories and supporting government policy decisions that affect us all. Given these critical applications, a poor modelling assumption or bug could have far-reaching consequences. However, scientific models possess several properties that make them notoriously difficult to test, including a complex input space, long execution times, and non-determinism, rendering existing testing techniques impractical. In fields such as epidemiology, where researchers seek answers to challenging causal questions, a statistical methodology known as Causal inference has addressed similar problems, enabling the inference of causal conclusions from noisy, biased, and sparse data instead of costly experiments. This article introduces the causal testing framework: a framework that uses causal inference techniques to establish causal effects from existing data, enabling users to conduct software testing activities concerning the effect of a change, such as metamorphic testing, a posteriori . We present three case studies covering real-world scientific models, demonstrating how the causal testing framework can infer metamorphic test outcomes from reused, confounded test data to provide an efficient solution for testing scientific modelling software. Andrew G. Clark, Michael Foster 0001, Benedikt Prifling, Neil Walkinshaw, Robert M. Hierons, Volker Schmidt, Robert D. Turner |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | Active Inference of EFSMs Without Reset
Michael Foster 0001, Roland Groz, Catherine Oriat, Adenilso da Silva Simão, Germán Vega, Neil Walkinshaw |
ICFEM | 6 |
| 2023 | Metamorphic Testing with Causal GraphsabstractMetamorphic testing provides a means by which to generate succinct test oracles that can apply to large input spaces. For this it depends on the formulation of metamorphic relations, which generally require extensive domain expertise and human input. To address this problem, we present a model-based testing approach that can automatically generate metamorphic relations and associated tests. Our approach is motivated by the observation that metamorphic testing is a fundamentally causal task. We show how it is possible to leverage lightweight graph-based modelling techniques from the field of causal inference to specify causal properties of the system-under-test. Through a series of controlled experiments, we find that the proposed approach is robust to misspecification and can test evasive causal relationships (i.e. those that are difficult to exercise and observe) when combined with an appropriate test generation strategy. We also apply the approach to two case studies from the Defects4J framework with known bugs that affect causal behaviour. The results of these case studies suggest that the approach is not only useful for catching bugs affecting causal structure, but also alerting the user to inaccuracies in the specification. Andrew G. Clark, Michael Foster 0001, Neil Walkinshaw, Robert M. Hierons |
ICST | 3 |
| 2023 | Digital-twin-based testing for cyber-physical systems: A systematic literature reviewabstractCyber–physical systems present a challenge to testers, bringing complexity and scale to safety-critical and collaborative environments. Digital twins enhance these systems through data-driven and simulation based models coupled to physical systems to provide visualisation, predict future states and communication. Due to the coupling between digital and physical worlds, digital twins provide a new perspective into cyber–physical system testing. The objectives of this study are to summarise the existing literature on digital-twin-based testing. We aim to uncover emerging areas of adoptions, the testing techniques used in these areas and identify future research areas. We conducted a systematic literature review which answered the following research questions: What cyber–physical systems are digital twins currently being used to test? How are test oracles defined for cyber–physical systems? What is the distribution of white-box, black-box and grey-box modelling techniques used for digital twins in the context of testing? How are test cases defined and how does this affect test inputs? We uncovered 26 relevant studies from 480 produced by searching with a curated search query. These studies showed an adoption of digital-twin-based testing following the introduction of digital twins in industry as well as the increasing accessibility of the technology. The oracles used in testing are the digital twin themselves and therefore rely on both system specification and data derivation. Cyber–physical systems are tested through passive testing techniques, as opposed to either active testing through test cases or predictive testing using digital twin prediction. This review uncovers the existing areas in which digital twins are used to test cyber–physical systems as well as outlining future research areas in the field. We outline how the infancy of digital twins has affected their wide variety of definitions, emerging specialised testing and modelling techniques as well as the current lack of predictive ability. Richard J. Somers, James A. Douthwaite, David James Wagg, Neil Walkinshaw, Robert M. Hierons |
Inf. Softw. Technol. | 4 |
| 2023 | Modelling Second-Order Uncertainty in State MachinesabstractModelling the behaviour of state-based systems can be challenging, especially when the modeller is not entirely certain about its intended interactions with the user or the environment. Currently, it is possible to associate a stated level of uncertainty with a given event by attaching probabilities to transitions (producing ‘Probabilistic State Machines’). This captures the ‘First-order uncertainty’ - the (un-)certainty that a given event will occur. However, this does not permit the modeller to capture their own uncertainty (or lack thereof) about that stated probability - also known as ‘Second-order uncertainty’. In this article we introduce a generalisation of probabilistic finite state machines that makes it possible to incorporate this important additional dimension of uncertainty. For this we adopt a formalism for reasoning about uncertainty called Subjective Logic. We present an algorithm to create these enhanced state machines automatically from a conventional state machine and a set of observed sequences. We show how this approach can be used for reverse-engineering predictive state machines from traces. Neil Walkinshaw, Robert M. Hierons |
IEEE Trans. Software Eng. | 1 |
| 2022 | Deep State Inference: Toward Behavioral Model Inference of Black-Box Software Systems
Foozhan Ataiefard, Mohammad Jafar Mashhadi, Hadi Hemmati, Neil Walkinshaw |
IEEE Trans. Software Eng. | 4 |
| 2021 | Reverse-Engineering EFSMs with Data Dependencies
Michael Foster 0001, John Derrick, Neil Walkinshaw |
ICTSS | 3 |
| 2021 | Test case generation for agent-based models: A systematic literature review
Andrew G. Clark, Neil Walkinshaw, Robert M. Hierons |
Inf. Softw. Technol. | 2 |
| 2020 | Reasoning about Uncertainty in Empirical ResultsabstractConclusions that are drawn from experiments are subject to varying degrees of uncertainty. For example, they might rely on small data sets, employ statistical techniques that make assumptions that are hard to verify, or there may be unknown confounding factors. In this paper we propose an alternative but complementary mechanism to explicitly incorporate these various sources of uncertainty into reasoning about empirical findings, by applying Subjective Logic. To do this we show how typical traditional results can be encoded as "subjective opinions" -- the building blocks of Subjective Logic. We demonstrate the value of the approach by using Subjective Logic to aggregate empirical results from two large published studies that explore the relationship between programming languages and defects or failures. Neil Walkinshaw, Martin J. Shepperd |
EASE | 1 |
| 2018 | Are 20% of files responsible for 80% of defects?abstractBackground: Over the past two decades a mixture of anecdote from the industry and empirical studies from academia have suggested that the 80:20 rule (otherwise known as the Pareto Principle) applies to the relationship between source code files and the number of defects in the system: a small minority of files (roughly 20%) are responsible for a majority of defects (roughly 80%). Neil Walkinshaw, Leandro L. Minku |
ESEM | 1 |
| 2018 | How Do Automatically Generated Unit Tests Influence Software Maintenance?abstractGenerating unit tests automatically saves time over writing tests manually and can lead to higher code coverage. However, automatically generated tests are usually not based on realistic scenarios, and are therefore generally considered to be less readable. This places a question mark over their practical value: Every time a test fails, a developer has to decide whether this failure has revealed a regression fault in the program under test, or whether the test itself needs to be updated. Does the fact that automatically generated tests are harder to read outweigh the time-savings gained by their automated generation, and render them more of a hindrance than a help for software maintenance? In order to answer this question, we performed an empirical study in which participants were presented with an automatically generated or manually written failing test, and were asked to identify and fix the cause of the failure. Our experiment and two replications resulted in a total of 150 data points based on 75 participants. Whilst maintenance activities take longer when working with automatically generated tests, we found developers to be equally effective with manually written and automatically generated tests. This has implications on how automated test generation is best used in practice, and it indicates a need for research into the generation of more realistic tests. Sina Shamshiri, José Miguel Rojas, Juan P. Galeotti, Neil Walkinshaw, Gordon Fraser 0001 |
ICST | 4 |
| 2018 | Comparison of Search-Based Algorithms for Stress-Testing Integrated CircuitsabstractThis paper is concerned with the task of ‘stress testing’an integrated circuit in its operational environment with the goal of identifying any circumstances under which the circuit might suffer from performance issues. Previous attempts to use simple hill-climbing algorithms to automate the generation of tests have faltered because the behaviour of the circuits can be subject to non-determinism, with a search space that can give rise to local maxima. In this paper we seek to work around these problems by experimenting with different search algorithms which ought to be better at handling such search-space properties (random-restart hill-climbing and simulated annealing). We evaluate these enhancements by applying the approach to test the Arm Cache Coherent Interconnect Unit (CCI) on a new 64-bit development platform, and show that both simulated annealing and random-restart hill-climbing outperforms simple hill-climbing algorithm. Basil Eljuse, Neil Walkinshaw |
SSBSE | 2 |
| 2018 | Effectively Incorporating Expert Knowledge in Automated Software RemodularisationabstractRemodularising the components of a software system is challenging: sound design principles (e.g., coupling and cohesion) need to be balanced against developer intuition of which entities conceptually belong together. Despite this, automated approaches to remodularisation tend to ignore domain knowledge, leading to results that can be nonsensical to developers. Nevertheless, suppling such knowledge is a potentially burdensome task to perform manually. A lot information may need to be specified, particularly for large systems. Addressing these concerns, we propose the SUpervised reMOdularisation (SUMO) approach. SUMO is a technique that aims to leverage a small subset of domain knowledge about a system to produce a remodularisation that will be acceptable to a developer. With SUMO, developers refine a modularisation by iteratively supplying corrections. These corrections constrain the type of remodularisation eventually required, enabling SUMO to dramatically reduce the solution space. This in turn reduces the amount of feedback the developer needs to supply. We perform a comprehensive systematic evaluation using 100 real world subject systems. Our results show that SUMO guarantees convergence on a target remodularisation with a tractable amount of user interaction. Mathew Hall, Neil Walkinshaw, Phil McMinn |
IEEE Trans. Software Eng. | 2 |
| 2017 | Uncertainty-Driven Black-Box Test Data GenerationabstractWe can never be certain that a software system is correct simply by testing it, but with every additional successful test we become less uncertain about its correctness. In absence of source code or elaborate specifications and models, tests are usually generated or chosen randomly. However, rather than randomly choosing tests, it would be preferable to choose those tests that decrease our uncertainty about correctness the most. In order to guide test generation, we apply what is referred to in Machine Learning as "Query Strategy Framework": We infer a behavioural model of the system under test and select those tests which the inferred model is "least certain" about. Running these tests on the system under test thus directly targets those parts about which tests so far have failed to inform the model. We provide an implementation that uses a genetic programming engine for model inference in order to enable an uncertainty sampling technique known as "query by committee", and evaluate it on eight subject systems from the Apache Commons Math framework and JodaTime. The results indicate that test generation using uncertainty sampling outperforms conventional and Adaptive Random Testing. Neil Walkinshaw, Gordon Fraser 0001 |
ICST | 1 |
| 2017 | Using Segment-Based Alignment to Extract Packet Structures from Network TracesabstractMany applications in security, from understanding unfamiliar protocols to fuzz-testing and guarding against potential attacks, rely on analysing network protocols. In many situations we cannot rely on access to a specification or even an implementation of the protocol, and must instead rely on raw network data "sniffed" from the network. When this is the case, one of the key challenges is to discern from the raw data the underlying packet structures - a task that is commonly carried out by using alignment algorithms to identify commonalities (e.g. field delimiters) between packets. For this, most approaches have used variants of the Needleman Wunsch algorthm to perform byte-wise alignment. However, they can suffer when messages are heterogeneous, or in cases where protocol fields are separated by long variable fields. In this paper, we present an alternative alignment algorithm known as segment-based alignment. We show how this technique can produce accurate results on traces from several common protocols, and how the results tend to be more intuitive than those produced by state-of-the-art techniques. Othman Esoul, Neil Walkinshaw |
QRS | 2 |
| 2016 | Data and Analysis Code for GP EFSM Inference
Mathew Hall, Neil Walkinshaw |
ICSME | 2 |
| 2016 | Inferring Computational State Machine Models from Program ExecutionsabstractThe challenge of inferring state machines from log data or execution traces is well-established, and has led to the development of several powerful techniques. Current approaches tend to focus on the inference of conventional finite state machines or, in few cases, state machines with guards. However, these machines are ultimately only partial, because they fail to model how any underlying variables are computed during the course of an execution, they are not computational. In this paper we introduce a technique based upon Genetic Programming to infer these data transformation functions, which in turn render inferred automata fully computational. Instead of merely determining whether or not a sequence is possible, they can be simulated, and be used to compute the variable values throughout the course of an execution. We demonstrate the approach by using a Cross-Validation study to reverse-engineer complete (computational) EFSMs from traces of established implementations. Neil Walkinshaw, Mathew Hall |
ICSME | 1 |
| 2016 | Choreography-Based Analysis of Distributed Message Passing ProgramsabstractWe report on the analysis of gen_server, a popular Erlang library to build client-server applications. Our analysis uses a tool based on choreographic models. We discuss how, once the library has been modelled in terms of communicating finite state machines, an automated analysis can be used to detect potential communication errors. The results of our analysis suggest how to properly use gen_server in order to guarantee the absence of communication errors. Ramsay Taylor, Emilio Tuosto, Neil Walkinshaw, John Derrick |
PDP | 3 |
| 2016 | A Search Based Approach for Stress-Testing Integrated Circuits
Basil Eljuse, Neil Walkinshaw |
SSBSE | 2 |
| 2016 | Inferring extended finite state machine models from software executions
Neil Walkinshaw, Ramsay Taylor, John Derrick |
Empir. Softw. Eng. | 1 |
| 2015 | SEPIA: Search for Proofs Using Inferred Automata
Thomas Gransden, Neil Walkinshaw, Rajeev Raman |
CADE | 2 |
| 2015 | An evidential reasoning approach for assessing confidence in safety evidenceabstractSafety cases present the arguments and evidence that can be used to justify the acceptable safety of a system. Many secondary factors such as the tools used, the techniques applied, and the experience of the people who created the evidence, can affect an assessor's confidence in the evidence cited by a safety case. One means of reasoning about this confidence and its inherent uncertainties is to present a `confidence argument' that explicitly justifies the provenance of the evidence used. In this paper, we propose a novel approach to automatically construct these confidence arguments by enabling assessors to provide individual judgements concerning the trustworthiness and the appropriateness of the evidence. The approach is based on Evidential Reasoning and enables the derivation of a quantified aggregate of the overall confidence. The proposed approach is supported by a prototype tool (EviCA) and has been evaluated using the Technology Acceptance Model. Sunil Nair, Neil Walkinshaw, Tim Kelly, Jose Luis de la Vara |
ISSRE | 2 |
| 2015 | Visualising software as a particle systemabstractCurrent metrics-based approaches to visualise unfamiliar software systems face two key limitations: (1) They are limited in terms of the number of dimensions that can be projected, and (2) they use fixed layout algorithms where the resulting positions of entities can be vulnerable to mis-interpretation. In this paper we show how computer games technology can be used to address these problems. We present the PhysVis software exploration system, where software metrics can be variably mapped to parameters of a physical model and displayed via a particle system. Entities can be imbued with attributes such as mass, gravity, and (for relationships) strength or springiness, alongside traditional attributes such as position, colour and size. The resulting visualisation is a dynamic scene; the relative positions of entities are not determined by a fixed layout algorithm, but by intuitive physical notions such as gravity, mass, and drag. The implementation is openly available, and we evaluate it on a selection of visualisation tasks for two openly-available software systems. Simon Scarle, Neil Walkinshaw |
VISSOFT | 2 |
| 2015 | Assessing and generating test sets in terms of behavioural adequacyabstractSummary Identifying a finite test set that adequately captures the essential behaviour of a program such that all faults are identified is a well‐established problem. This is traditionally addressed with syntactic adequacy metrics (e.g. branch coverage), but these can be impractical and may be misleading even if they are satisfied. One intuitive notion of adequacy, which has been discussed in theoretical terms over the past three decades, is the idea ofbehavioural coverage: If it is possible to infer an accurate model of a system from its test executions, then the test set can be deemed to be adequate. Despite its intuitive basis, it has remained almost entirely in the theoretical domain because inferred models have been expected to be exact (generally an infeasible task) and have not allowed for any pragmatic interim measures of adequacy to guide test set generation. This paper presents a practical approach to incorporate behavioural coverage. OurBESTESTapproach (1) enables the use of machine learning algorithms to augment standard syntactic testing approaches and (2) shows how search‐based testing techniques can be applied to generate test sets with respect to this criterion. An empirical study on a selection of Java units demonstrates that test sets with higher behavioural coverage significantly outperform current baseline test criteria in terms of detected faults. © 2015 The Authors.Software Testing, Verification and Reliabilitypublished by John Wiley & Sons, Ltd. Gordon Fraser 0001, Neil Walkinshaw |
Softw. Test. Verification Reliab. | 2 |
| 2014 | Establishing the Source Code Disruption Caused by Automated Remodularisation ToolsabstractCurrent software remodularisation tools only operate on abstractions of a software system. In this paper, we investigate the actual impact of automated remodularisation on source code using a tool that automatically applies remodularisations as refactorings. This shows us that a typical remodularisation (as computed by the Bunch tool) will require changes to thousands of lines of code, spread throughout the system (typically no code files remain untouched). In a typical multi-developer project this presents a serious integration challenge, and could contribute to the low uptake of such tools in an industrial context. We relate these findings with our ongoing research into techniques that produce iterative commit friendly" code changes to address this problem. Mathew Hall, Muhammad Ali Khojaye, Neil Walkinshaw, Phil McMinn |
ICSME | 3 |
| 2014 | Mining State-Based Models from Proof Corpora
Thomas Gransden, Neil Walkinshaw, Rajeev Raman |
CICM | 2 |
| 2013 | STAMINA: a competition to encourage the development and assessment of software model inference techniquesabstractModels play a crucial role in the development and maintenance of software systems, but are often neglected during the development process due to the considerable manual effort required to produce them. In response to this problem, numerous techniques have been developed that seek to automate the model generation task with the aid of increasingly accurate algorithms from the domain of Machine Learning. From an empirical perspective, these are extremely challenging to compare; there are many factors that are difficult to control (e.g. the richness of the input and the complexity of subject systems), and numerous practical issues that are just as troublesome (e.g. tool availability). This paper describes the StaMinA ( Sta te M achine In ference A pproaches) competiton, that was designed to address these problems. The competition attracted numerous submissions, many of which were improved or adapted versions of techniques that had not been subjected to extensive empirical evaluations, and had not been evaluated with respect to their ability to infer models of software systems. This paper shows how many of these techniques substantially improve on the state of the art, providing insights into some of the factors that could underpin the success of the best techniques. In a more general sense it demonstrates the potential for competitions to act as a useful basis for empirical software engineering by (a) spurring the development of new techniques and (b) facilitating their comparative evaluation to an extent that would usually be prohibitively challenging without the active participation of the developers. Neil Walkinshaw, Bernard Lambeau, Christophe Damas, Kirill Bogdanov 0002, Pierre Dupont |
Empir. Softw. Eng. | 1 |
| 2013 | Automated Comparison of State-Based Software Models in Terms of Their Language and StructureabstractState machines capture the sequential behavior of software systems. Their intuitive visual notation, along with a range of powerful verification and testing techniques render them an important part of the model-driven software engineering process. There are several situations that require the ability to identify and quantify the differences between two state machines (e.g. to evaluate the accuracy of state machine inference techniques is measured by the similarity of a reverse-engineered model to its reference model). State machines can be compared from two complementary perspectives: (1) In terms of theirlanguage-- the externally observable sequences of events that are permitted or not, and (2) in terms of theirstructure-- the actual states and transitions that govern the behavior. This article describes two techniques to compare models in terms of these two perspectives. It shows how the difference can be quantified and measured by adapting existing binary classification performance measures for the purpose. The approaches have been implemented by the authors, and the implementation is openly available. Feasibility is demonstrated via a case study to compare two real state machine inference approaches. Scalability and accuracy are assessed experimentally with respect to a large collection of randomly synthesized models. Neil Walkinshaw, Kirill Bogdanov 0002 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2012 | Supervised software modularisationabstractThis paper is concerned with the challenge of reorganising a software system into modules that both obey sound design principles and are sensible to domain experts. The problem has given rise to several unsupervised automated approaches that use techniques such as clustering and Formal Concept Analysis. Although results are often partially correct, they usually require refinement to enable the developer to integrate domain knowledge. This paper presents the SUMO algorithm, an approach that is complementary to existing techniques and enables the maintainer to refine their results. The algorithm is guaranteed to eventually yield a result that is satisfactory to the maintainer, and the evaluation on a diverse range of systems shows that this occurs with a reasonably low amount of effort. Mathew Hall, Neil Walkinshaw, Phil McMinn |
ICSM | 2 |
| 2012 | Behaviourally Adequate Software TestingabstractIdentifying a finite test set that adequately captures the essential behaviour of a program such that all faults are identified is a well-established problem. Traditional adequacy metrics can be impractical, and may be misleading even if they are satisfied. One intuitive notion of adequacy, which has been discussed in theoretical terms over the past three decades, is the idea of behavioural coverage, if it is possible to infer an accurate model of a system from its test executions, then the test set must be adequate. Despite its intuitive basis, it has remained almost entirely in the theoretical domain because inferred models have been expected to be exact (generally an infeasible task), and have not allowed for any pragmatic interim measures of adequacy to guide test set generation. In this work we present a new test generation technique that is founded on behavioural adequacy, which combines a model evaluation framework from the domain of statistical learning theory with search-based white-box test generation strategies. Experiments with our BESTEST prototype indicate that such test sets not only come with a statistically valid measurement of adequacy, but also detect significantly more defects. Gordon Fraser 0001, Neil Walkinshaw |
ICST | 2 |
| 2012 | Model-Based Testing and Model Inference
Karl Meinke, Neil Walkinshaw |
ISoLA (1) | 2 |
| 2011 | A multiobjective optimisation approach for the dynamic inference and refinement of agent-based model specificationsabstractDespite their increasing popularity, agent-based models are hard to test, and so far no established testing technique has been devised for this kind of software applications. Reverse engineering an agent-based model specification from model simulations can help establish a confidence level about the implemented model and in some cases reveal discrepancies between observed and normal or expected behaviour. In this study, a multiobjective optimisation technique based on a simple random search algorithm is deployed to dynamically infer and refine the specification of three agent-based models from their simulations. The multiobjective optimisation technique also incorporates a dynamic invariant detection technique which serves to guide the search towards uncovering new model behaviour that better captures the model specification. The Non-dominated Sorting Genetic Algorithm (NSGA-II) was also deployed to replace the random search algorithm, and the results from both approaches were compared. While both algorithms revealed good potential in capturing the model specifications, the pure exploratory nature of random search was found more suitable for the application at hand, compared to the balanced exploitation/exploration nature of genetic algorithms in general. Salem Fawaz Adra, Mariam Kiran, Phil McMinn, Neil Walkinshaw |
IEEE Congress on Evolutionary Computation | 4 |
| 2011 | Assessing Test Adequacy for Black-Box Systems without Specifications
Neil Walkinshaw |
ICTSS | 1 |
| 2010 | Superstate identification for state machines using search-based clusteringabstractState machines are a popular method of representing a system at a high level of abstraction that enables developers to gain an overview of the system they represent and quickly understand it. Mathew Hall, Phil McMinn, Neil Walkinshaw |
GECCO | 3 |
| 2010 | Increasing Functional Coverage by Inductive Testing: A Case Study
Neil Walkinshaw, Kirill Bogdanov 0002, John Derrick, Javier París |
ICTSS | 1 |
| 2010 | TAIC-PART 2009 - Testing: Academic & Industrial Conference - Practice And Research Techniques: Special Section Editorial
Leonardo Bottaci, Gregory M. Kapfhammer, Neil Walkinshaw |
J. Syst. Softw. | 3 |
| 2009 | Iterative Refinement of Reverse-Engineered Models by Model-Based Testing
Neil Walkinshaw, John Derrick, Qiang Guo 0001 |
FM | 1 |
| 2008 | Inferring Finite-State Models with Temporal ConstraintsabstractFinite state machine-based abstractions of software behaviour are popular because they can be used as the basis for a wide range of (semi-) automated verification and validation techniques. These can however rarely be applied in practice, because the specifications are rarely kept up- to-date or even generated in the first place. Several techniques to reverse-engineer these specifications have been proposed, but they are rarely used in practice because their input requirements (i.e. the number of execution traces) are often very high if they are to produce an accurate result. An insufficient set of traces usually results in a state machine that is either too general, or incomplete. Temporal logic formulae can often be used to concisely express constraints on system behaviour that might otherwise require thousands of execution traces to identify. This paper describes an extension of an existing state machine inference technique that accounts for temporal logic formulae, and encourages the addition of new formulae as the inference process converges on a solution. The implementation of this process is openly available, and some preliminary results are provided. Neil Walkinshaw, Kirill Bogdanov 0002 |
ASE | 1 |
| 2008 | Improving dynamic software analysis by applying grammar inference principlesabstractAbstract Grammar inference is a family of machine learning techniques that aim to infer grammars from a sample of sentences in some (unknown) language. Dynamic analysis is a family of techniques in the domain of software engineering that attempts to infer rules that govern the behaviour of software systems from a sample of executions. Despite their disparate domains, both fields have broadly similar aims; they try to infer rules that govern the behaviour of some unknown system from a sample of observations. Deriving general rules about program behaviour from dynamic analysis is difficult because it is virtually impossible to identify and supply a complete sample of necessary program executions. The problems that arise with incomplete input samples have been extensively investigated in the grammar inference community. This has resulted in a number of advances that have produced increasingly sophisticated solutions that are more successful at accurately inferring grammars from (potentially) sparse information about the underlying system. This paper investigates the similarities and shows how many of these advances can be applied with similar effect to dynamic analysis problems by a series of small experiments on random state machines. Copyright © 2008 John Wiley & Sons, Ltd. Neil Walkinshaw, Kirill Bogdanov 0002, Mike Holcombe, Sarah Salahuddin |
J. Softw. Maintenance Res. Pract. | 1 |
| 2008 | Automated discovery of state transitions and their functions in source codeabstractAbstract Finite‐state machine specifications form the basis for a number of rigorous state‐based testing techniques and can help to understand program behaviour. Unfortunately they are rarely maintained during software development, which means that these benefits can rarely be fully exploited. This paper describes a technique that, given a set of states that are of interest to a developer, uses symbolic execution to reverse‐engineer state transitions from source code. A particularly novel aspect of our approach is that, besides determining whether or not a state transition can take place, it also identifies the paths through the source code that govern a transition. The technique has been implemented as a prototype, enabling its preliminary evaluation with respect to real software systems. Copyright © 2007 John Wiley & Sons, Ltd. Neil Walkinshaw, Kirill Bogdanov 0002, Shaukat Ali 0001, Mike Holcombe |
Softw. Test. Verification Reliab. | 1 |
| 2007 | Feature Location and Extraction using Landmarks and BarriersabstractIdentifying and isolating the source code associated with a particular feature is a problem that frequently arises in many maintenance tasks. The delocalised nature of object-oriented systems, where the code associated with a feature is distributed across many interrelated objects, makes this problem particularly challenging. This paper presents an approach that combines 'landmark' methods that have a key role in the execution of a particular feature with slicing to create a call graph of related code. The size of this call graph is constrained by the identification of 'barrier' methods which exclude parts of the graph that are not of interest. The approach is supported by a tool, and the evaluation on three open-source systems yields encouraging results and demonstrates the practical applicability of the technique. Neil Walkinshaw, Marc Roper, Murray Wood |
ICSM | 1 |