Matthew Leeke

dblp:73/7674 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 2 first-author · 3 since 2021Security and privacy · 7 · 2 first-author · 3 since 2021Systems, architecture and hardware · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2024 SARAD: Spatial Association-Aware Anomaly Detection and Diagnosis for Multivariate Time Series
abstract
Anomaly detection in time series data is fundamental to the design, deployment, and evaluation of industrial control systems. Temporal modeling has been the natural focus of anomaly detection approaches for time series data. However, the focus on temporal modeling can obscure or dilute the spatial information that can be used to capture complex interactions in multivariate time series. In this paper, we propose SARAD, an approach that leverages spatial information beyond data autoencoding errors to improve the detection and diagnosis of anomalies. SARAD trains a Transformer to learn the spatial associations, the pairwise inter-feature relationships which ubiquitously characterize such feedback-controlled systems. As new associations form and old ones dissolve, SARAD applies subseries division to capture their changes over time. Anomalies exhibit association descending patterns, a key phenomenon we exclusively observe and attribute to the disruptive nature of anomalies detaching anomalous features from others. To exploit the phenomenon and yet dismiss non-anomalous descent, SARAD performs anomaly detection via autoencoding in the association space. We present experimental results to demonstrate that SARAD achieves state-of-the-art performance, providing robust anomaly detection and a nuanced understanding of anomalous events.
Zhihao Dai, Ligang He, Shuang-Hua Yang, Matthew Leeke
NeurIPS4
2022 A Heterogeneous Redundant Architecture for Industrial Control System Security
abstract
Component-level heterogeneous redundancy is gaining popularity as an approach for preventing single-point security breaches in Industrial Control Systems (ICSs), especially with regard to core components such as Programmable Logic Controllers (PLCs). To take control of a system with component-level heterogeneous redundancy, an adversary must uncover and concurrently exploit vulnerabilities across multiple versions of hardened components. As such, attackers incur increased costs and delays when seeking to launch a successful attack. Existing approaches advocate attack resilience via pairwise comparison among outputs from multiple PLCs. These approaches incur increased resource costs due to them having a high degree of redundancy and do not address concurrent attacks. In this paper we address both issues, demonstrating a data-driven component selection approach that achieves a trade-off between resources cost and security. In particular, we propose (i) a novel dual-PLC ICS architecture with native pairwise comparison which can offer limited yet comparable defence against single-point breaches, (ii) a machine-learning based selection mechanisms which can deliver resilience against non-concurrent attacks under resource constraints, (iii) a scaled up variant of the proposed architecture to counteract concurrent attacks with modest resource implications.
Zhihao Dai, Matthew Leeke, Shuang-Hua Yang
PRDC2
2022 Towards a Dependable Energy Market: Proof of Authority in a Blockchain-based Peer-to-Peer Microgrid
abstract
For reasons of energy security, affordability and environment, the foundations of our energy markets are in flux. A promising approach for addressing some of the challenges faced comes in the form of blockchain-based peer-to-peer (P2P) microgrids, which we argue are a potential solution to many of the pitfalls of existing grid architectures. More specifically, this paper analyses consensus mechanisms that could be employed in blockchain-based microgrids, demonstrating that proof of authority (PoA) is a promising direction when seeking to improve the dependability of the energy market. We go on to specify a viable architecture for a PoA-based microgrid and provide experimental results to demonstrate that PoA is superior to proof of work (PoW) in this context.
Joe Hewett, Mark Etman, Robbie Marseglia, Tomas Mella Pickersgill, Matthew Leeke
PRDC5
2022 Developing a GPT-3-Based Automated Victim for Advance Fee Fraud Disruption
abstract
Advance Fee Fraud (AFF) is amongst the most prevalent and destructive forms of cybercrime. Scammers typically commit AFF by tricking victims into making upfront payments for goods or services that are never provided. These payments are small compared to the alleged gains, and can thus be attractive for victims, particularly if they are vulnerable or in a heightened emotional state. Given that approximately three billion fraudulent emails are sent every day, the scale and impact of AFF demands innovative approaches. In this paper we document the development of an automated victim for AFF. The system leverages GPT-3, a large language model, in conjunction with deliberately engineered prompts to generate plausible responses to AFF emails, allowing fraud to be disrupted and actionable information relating to perpetrators to be obtained.
Joe Hewett, Matthew Leeke
PRDC2
2020 Simultaneous Fault Models for the Generation and Location of Efficient Error Detection Mechanisms
abstract
Abstract The application of machine learning to software fault injection data has been shown to be an effective approach for the generation of efficient error detection mechanisms (EDMs). However, such approaches to the design of EDMs have invariably adopted a fault model with a single-fault assumption, limiting the relevance of the detectors and their evaluation. Software containing more than a single fault is commonplace, with safety standards recognizing that critical failures are often the result of unlikely or unforeseen combinations of faults. This paper addresses this shortcoming, demonstrating that it is possible to generate efficient EDMs under simultaneous fault models. In particular, it is shown that (i) efficient EDMs can be designed using fault injection data collected under models accounting for the occurrence of simultaneous faults, (ii) exhaustive fault injection under a simultaneous bit flip model can yield improved EDM efficiency, (iii) exhaustive fault injection under a simultaneous bit flip model can be made non-exhaustive and (iv) EDMs can be relocated within a software system using program slicing, reducing the resource costs of experimentation to practicable levels without sacrificing EDM efficiency.
Matthew Leeke
Comput. J.1
2020 WolfGraph: The edge-centric graph processing on GPU
Huanzhou Zhu, Ligang He, Matthew Leeke, Rui Mao 0001
Future Gener. Comput. Syst.3
2018 Hybrid online protocols for source location privacy in wireless sensor networks
abstract
Wireless sensor networks (WSNs) will form the building blocks of many novel applications such as asset monitoring. These applications will have to guarantee that the location of the occurrence of specific events is kept private from attackers, in what is called the source location privacy (SLP) problem. Fake sources have been used in numerous techniques, however, the solution’s efficiency is typically achieved by fine-tuning parameters at compile time. This is undesirable as WSN conditions may change. In this paper, we first present an SLP algorithm – Dynamic – that estimates the relevant parameters at runtime and show that it provides a high level of SLP, albeit at the expense of a high number of messages. To address this, we provide a hybrid online algorithm – DynamicSPR – that uses directed random walks for the fake sources allocation strategy to reduce energy usage. We perform simulations of the various protocols we present and our results show that DynamicSPR provides a similar level of SLP as when parameters are optimised at compile-time, with a lower number of messages sent.
Matthew Bradbury, Arshad Jhumka, Matthew Leeke
J. Parallel Distributed Comput.3
2017 Simultaneous Fault Models for the Generation of Efficient Error Detection Mechanisms
abstract
The application of machine learning to software fault injection data has been shown to be an effective approach for the generation of efficient error detection mechanisms (EDMs). However, such approaches to the design of EDMs have invariably adopted a fault model with a single-fault assumption, limiting the practical relevance of the detectors and their evaluation. Software containing more than a single fault is commonplace, with prominent safety standards recognising that critical failures are often the result of unlikely or unforeseen combinations of faults. This paper addresses this shortcoming, demonstrating that it is possible to generate similarly efficient EDMs under more realistic fault models. In particular, it is shown that (i) efficient EDMs can be designed using fault data collected under models accounting for the occurrence of simultaneous faults, (ii) exhaustive fault injection under a simultaneous bit flip model can yield improvements to EDM efficiency, and (iii) exhaustive fault injection under a simultaneous bit flip model can made non-exhaustive, reducing the resource costs of experimentation to practicable levels, without sacrificing resultant EDM efficiency.
Matthew Leeke
ISSRE1
2015 Assessing the Performance of Phantom Routing on Source Location Privacy in Wireless Sensor Networks
abstract
As wireless sensor networks (WSNs) have been applied across a spectrum of application domains, the problem of source location privacy (SLP) has emerged as a significant issue, particularly in safety-critical situations. In seminal work on SLP, phantom routing was proposed as an approach to addressing the issue. However, results presented in support of phantom routing have not included considerations for practical network configurations, omitting simulations and analyses with larger network sizes. This paper addresses this shortcoming by conducting an in-depth investigation of phantom routing under various network configurations. The results presented demonstrate that previous work in phantom routing does not generalise well to different network configurations. Specifically, under certain configurations, it is shown that the afforded SLP is reduced by a factor of up to 75.
Chen Gu, Matthew Bradbury, Arshad Jhumka, Matthew Leeke
PRDC4
2015 Fake source-based source location privacy in wireless sensor networks
abstract
Summary The development of novel wireless sensor network (WSN) applications, such as asset monitoring, has led to novel reliability requirements. One such property is source location privacy (SLP). The original SLP problem is to protect the location of a source node in a WSN from a singledistributed eavesdropperattacker. Several techniques have been proposed to address the SLP problem, and most of them use some form of traffic analysis and engineering to provide enhanced SLP. The use of fake sources is considered to be promising for providing SLP, and several works have investigated the effectiveness of the fake sources approach under various attacker models. However, very little work has been done to understand the theoretical underpinnings of the fake source technique. In this paper, we (i) provide a novel formalisation of the fake sources selection problem; (ii) prove the fake sources selection problem to be NP‐complete; (iii) provide parametric heuristics for three different network configurations; and (iv) show that these heuristics provide (near) optimal levels of SLP under appropriate parameterisation. Our results show that fake sources can provide a high level of SLP. Our work is the first to investigate the theoretical underpinnings of the fake source technique. Copyright © 2014 John Wiley & Sons, Ltd.
Arshad Jhumka, Matthew Bradbury, Matthew Leeke
Concurr. Comput. Pract. Exp.3
2014 Towards unified secure on- and off-line analytics at scale
abstract
Data scientists have applied various analytic models and techniques to address the oft-cited problems of large volume, high velocity data rates and diversity in semantics. Such approaches have traditionally employed analytic techniques in a streaming or batch processing paradigm. This paper presents CRUCIBLE, a first-in-class framework for the analysis of large-scale datasets that exploits both streaming and batch paradigms in a unified manner. The CRUCIBLE framework includes a domain specific language for describing analyses as a set of communicating sequential processes, a common runtime model for analytic execution in multiple streamed and batch environments, and an approach to automating the management of cell-level security labelling that is applied uniformly across runtimes. This paper shows the applicability of CRUCIBLE to a variety of state-of-the-art analytic environments, and compares a range of runtime models for their scalability and performance against a series of native implementations. The work demonstrates the significant impact of runtime model selection, including improvements of between 2.3× and 480× between runtime models, with an average performance gap of just 14× between CRUCIBLE and a suite of equivalent native implementations.
Peter Coetzee, Matthew Leeke, Stephen A. Jarvis
Parallel Comput.2
2013 Towards the Design of Efficient Error Detection Mechanisms for Transient Data Errors
abstract
A dependable software system must contain two dependability components: (i) error detection mechanisms (EDMs) and (ii) error recovery mechanisms. Currently, EDMs are generally designed based on some system specification or based on the experience of software engineers, with their efficiency typically being measured using fault injection and software measures such as coverage and latency. In contrast to finite-state programs, for which efficient EDMs can be obtained by design, no systematic design approach exists for real-world software systems. In this paper, we bridge this gap by developing an approach for the design of highly efficient error detection predicates for EDMs for such software systems. Our approach is based on the use of data mining techniques to classify states as safe or failure-inducing. The results presented, under a transient data value fault model, demonstrate the viability of the approach for the development of efficient EDMs, as the EDMs generated yield a true positive rate of nearly 100% and a false positive rate close to 0% for the detection of failure-inducing states.
Matthew Leeke, Arshad Jhumka, Sarabjot S. Anand
Comput. J.1
2012 Towards Understanding Source Location Privacy in Wireless Sensor Networks through Fake Sources
abstract
Source location privacy is becoming an increasingly important property in wireless sensor network applications, such as asset monitoring. The original source location problem is to protect the location of a source in a wireless sensor network from a single distributed eavesdropper attack. Several techniques have been proposed to address the source location problem, where most of these apply some form of traffic analysis and engineering to provide enhanced privacy. One such technique, namely fake sources, has proved to be promising for providing source location privacy. Recent research has concentrated on investigating the efficiency of fake source approaches under various attacker models. In this paper, we (i) provide a novel formalisation of the source location privacy problem, (ii) prove the source location privacy problem to be NP-complete, and (iii) provide a heuristic that yields an optimal level of privacy under appropriate parameterisation. Crucially, the results presented show that fake sources can provide a high, sometimes optimal, level of privacy.
Arshad Jhumka, Matthew Bradbury, Matthew Leeke
TrustCom3
2011 A methodology for the generation of efficient error detection mechanisms
abstract
A dependable software system must contain error detection mechanisms and error recovery mechanisms. Software components for the detection of errors are typically designed based on a system specification or the experience of software engineers, with their efficiency typically being measured using fault injection and metrics such as coverage and latency. In this paper, we introduce a methodology for the design of highly efficient error detection mechanisms. The proposed methodology combines fault injection analysis and data mining techniques in order to generate predicates for efficient error detection mechanisms. The results presented demonstrate the viability of the methodology as an approach for the development of efficient error detection mechanisms, as the predicates generated yield a true positive rate of almost 100% and a false positive rate very close to 0% for the detection of failure-inducing states. The main advantage of the proposed methodology over current state-of-the-art approaches is that efficient detectors are obtained by design, rather than by using specification-based detector design or the experience of software engineers.
Matthew Leeke, Saima Arif, Arshad Jhumka, Sarabjot S. Anand
DSN1
2011 The Early Identification of Detector Locations in Dependable Software
abstract
The dependability properties of a software system are usually assessed and refined towards the end of the software development lifecycle. Problems pertaining to software dependability may necessitate costly system redesign. Hence, early insights into the potential for error propagation within a software system would be beneficial. Further, the refinement of the dependability properties of software involves the design and location of dependability components called detectors and correctors. Recently, a metric, called spatial impact, has been proposed to capture the extent of error propagation in a software system, providing insights into the location of detectors and correctors. However, the metric only provides insights towards the end of the software development life cycle. In this paper, our objective is to investigate whether spatial impact can enable the early identification of locations for detectors. To achieve this we first hypothesise that spatial impact is correlated with module coupling, a metric that can be evaluated early in the software development life cycle, and show this relationship to hold. We then evaluate module coupling for the modules of a complex software system, identifying modules with high coupling values as potential locations for detectors. We then enhanced these modules with detectors and perform fault-injection analysis to determine the suitability of these locations. The results presented demonstrate that our approach can permit the early identification of possible detector locations.
Arshad Jhumka, Matthew Leeke
ISSRE2
2011 On the Use of Fake Sources for Source Location Privacy: Trade-Offs Between Energy and Privacy
abstract
Wireless sensor networks have enabled novel applications such as monitoring, where security is invariably a requirement. One aspect of security, namely source location privacy, is becoming an increasingly important property of some wireless sensor network applications. The fake source technique has been proposed as an efficient technique to handle the source location privacy problem. However, there are several factors that limit the usefulness of current results: (i) the selection of fake sources is dependent on sophisticated nodes, (ii) fake sources are known a priori and (iii) the selection of fake sources is based on a prohibitively expensive pre-configuration phase. In this paper, we investigate the privacy enhancement and energy efficiency of different implementations of the fake source technique that circumvents these limitations. Our results show that the fake source technique is indeed effective in enhancing privacy. Specifically, one implementation achieves near-perfect privacy when there is at least one fake source in the network, at the expense of increased energy consumption. In the presence of multiple attackers, the same implementation yields only a 30% decrease in capture ratio with respect to flooding. To address this problem, we propose a hybrid technique which achieves a corresponding 50% reduction in the capture ratio and a near-perfect privacy whenever at least one fake source exists in the network.
Arshad Jhumka, Matthew Leeke, Sambid Shrestha
Comput. J.2
2009 Issues on the Design of Efficient Fail-Safe Fault Tolerance
abstract
The design of a fault-tolerant program is known to be an inherently difficult task. Decisions taken during the design process will invariably have an impact on the efficiency of the resulting fault-tolerant program. In this paper, we focus on two such decisions, namely (i) the class of faults the program is to tolerate, and (ii) the variables that can be read and written. The impact these design issues have on the overall fault tolerance of the system needs to be well-understood, failure of which can lead to costly redesigns. For the case of understanding the impact of fault classes on the efficiency of fail-safe fault tolerance, we show that, under the assumption of a general fault model, it is impossible to preserve the original behavior of the fault-intolerant program. For the second problem of read and write constraints of variables, we again show that it is impossible to preserve the original behavior of the fault-intolerant program. We analyze the reasons that lead to these impossibility results, and suggest possible ways of circumventing them.
Arshad Jhumka, Matthew Leeke
ISSRE2
2009 Evaluating the Use of Reference Run Models in Fault Injection Analysis
abstract
Fault injection (FI) has been shown to be an effective approach to assessing the dependability of software systems. To determine the impact of faults injected during FI, a given oracle is needed. Oracles can take a variety of forms, including (i) specifications, (ii) error detection mechanisms and (iii) golden runs. Focusing on golden runs, in this paper we show that there are classes of software which a golden run based approach can not be used to analyse. Specifically, we demonstrate that a golden run based approach can not be used in the analysis of systems which employ a main control loop with an irregular period. Further, we show how a simple model, which has been refined using FI experiments, can be employed as an oracle in the analysis of such a system.
Matthew Leeke, Arshad Jhumka
PRDC1