VLDB 2026 Research / reviewers in the wild / expert
Philip A. Legg
dblp:25/7426 · also Phil Legg
· DBLP profile ↗
17ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0003-3460-5609ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-authorSecurity and privacy · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LOTL-hunter: Detecting multi-stage living-off-the-land attacks in cyber-physical systems using decision fusion techniques with digital twinsabstract• Digital Twin-based testbed enables safe and repeatable cyber-physical threat simulation. • Multimodal anomaly fusion improves detection of stealthy, multi-stage APT attacks • Confidence-aware decision logic resolves missing data and prediction conflicts • Per-minute aggregation and alignment of anomalies support IT/OT threat hunting • Cyber-physical threat summary integrates IT/OT anomalies for explainability The integration of smart sensors and actuators in industrial environments has expanded the cyber-physical attack surface, making it increasingly difficult to distinguish anomalies caused by cyberattacks from those due to mechanical or electrical faults. This challenge is exacerbated by stealthy, multi-stage attacks leveraging Living off the Land (LOTL) techniques, which often evade conventional anomaly detection or intrusion detection systems (IDS). This study presents a Digital Twin-based testbed for safe, repeatable simulation of multi-stage cyber-physical attacks targeting Cyber-Physical Systems (CPS) and Industrial Control Systems (ICS). We propose a two-level decision fusion method that aggregates and aligns anomalies across network, process, and host domains in synchronized 1-minute intervals. The first-level fusion improves OT-layer detection by applying confidence-aware decision logic to outputs combined from (a) a supervised deep learning model (LSTM-FCN) for process anomalies, (b) an unsupervised model (Isolation Forest) for OPC UA network anomalies, and (c) process alarm signals. The second-level fusion integrates these results with host-based anomalies, computed through point-based scoring of Wazuh alerts, to provide comprehensive IT/OT situational awareness. Experimental results demonstrate improved detection of stealthy, multi-stage APT attack behaviours. Additionally, Large Language Models (LLM) provide summarization of the integrated IT/OT anomaly logs into human-readable insights, enhancing interpretability and supporting cyber threat hunting. Carol Lo, Thu Yein Win, Zeinab Rezaeifar, Zaheer Abbas Khan, Philip A. Legg |
Future Gener. Comput. Syst. | 5 |
| 2026 | Advancing fuzzing with unbiased random generator and Feistel network-based mutationsabstractThis research tackles challenges in traditional fuzzing, such as limited coverage, instability, and inefficiency in bug discovery. We propose two novel models and their combination to enhance mutation processes and improve its reliability through unbiased randomisation, building on cryptographic techniques from our prior work. To our knowledge, we are the first to apply this approach to AFL++ , extending Feistel-inspired mutation and high-performance randomisation to generate high-quality test cases, with potential to attract attention in the fuzzing community. • Integrate and assess Feistel-inspired mutations’ impact on AFL++ performance, focusing on code coverage and stability. • Integrate the Permuted Congruential Generator ( PCG ) into AFL++ and evaluate its performance compared to traditional random number generators ( RNGs ). • Evaluate a hybrid model combining Feistel and PCG randomness for better stability and coverage. We enhance AFL++ with algorithmic improvements and RNGs modifications. Our models include: • CAFL++ ( Cryptographic - AFL++ ): Integrates Feistel-inspired transformations for improved coverage. • PCGAFL++ : Refines the AFL ’s RNG with PCG to reduce bias. • CPCGAFL++ : Combines Feistel-inspired swaps and PCG -based RNG for a robust fuzzing approach. Performance was analysed using metrics like Code Coverage and the Vargha-Delaney_A12 statistic across 20 Fuzzbench targets, and bug discovery on three targets. Our models showed significant improvements over AFL++ . CAFL++ outperformed AFL++ in 75% of test targets, offering better code coverage and stability. PCGAFL++ surpassed AFL++ in 60% of targets by enhancing randomness, resulting in more efficient fuzzing. CPCGAFL++ demonstrated improved stability and enhanced bug discovery performance, while achieving code coverage comparable to AFL++ . These results highlight the key improvements introduced by our two models for fuzz testing. Our models advance fuzzing by improving code coverage and stability. Integrating Feistel-inspired swaps and PCG -based RNG overcomes traditional fuzzing limitations, offering a more efficient and reliable method. These models represent a step forward in fuzzing techniques, influencing both academic research and industrial practices. Sadegh Bamohabbat Chafjiri, Philip A. Legg, Jun Hong 0001, Michail-Antisthenis I. Tsompanas |
Inf. Softw. Technol. | 2 |
| 2024 | Cyber Funfair: Creating Immersive and Educational Experiences for Teaching Cyber Physical Systems SecurityabstractDelivering meaningful and inspiring cyber security education for younger audiences can often be a challenge due to limited expertise and resources. Key to any outreach activity is that it both develops a learner's curiosity, as well as providing educational objectives. To address this need, we developed a novel learning and awareness activity that addresses the Cyber Physical Systems (CPS) Security knowledge area as mapped by the Cyber Security Body of Knowledge (CyBOK). At the core of our activity is the integration of the Raspberry Pi device with LEGO SPIKE kits. LEGO SPIKE is part of the LEGO Education system that combines colourful LEGO building blocks with motors and sensors, creating an adaptable and engaging learning environment. This hands-on approach allows participants to witness the tangible consequences of cyber and network actions in a physical and engaging format. To evaluate the effectiveness of the activity, we used the activity as part of an outreach activity day attended by approximately 300 students aged between 12-14 from schools across the West of England. Participants of the activity were surveyed and the results showed an increase in understanding of CPS specific and wider cyber security for over 90% of respondents. Activity engagement was also well received with no negative feedback. We report on our survey findings and discuss best practices to support other practitioners in developing hands-on interactive experiences for engaging and educational cyber security activities. Alan Mills, Philip A. Legg |
SIGCSE (1) | 3 |
| 2024 | Vulnerability detection through machine learning-based fuzzing: A systematic reviewabstractModern software and networks underpin our digital society, yet the rapid growth of vulnerabilities that are uncovered within these threaten our cyber security posture. Addressing these issues at scale requires automated proactive approaches that can identify and mitigate these vulnerabilities in a suitable time frame. Fuzzing techniques have emerged as crucial methods to preemptively tackle these risks. However, traditional fuzzing methods encounter various challenges, such as a lack of strategy for deep bug identification, time-intensive bug analysis, quality of inputs, seed scheduling and others. To overcome these challenges, diverse Machine Learning (ML) models and optimisation techniques have been employed, including advanced feature engineering, optimised seed selection, refined predictive/fitness models, and Gradient-based optimisation. Furthermore, the use of ML architectures such as Long Short-Term Memory (LSTM), Generative Adversarial Network (GAN), Sequence-to-Sequence (Seq2Seq), and Generative Randomised Unit (GRU), have demonstrated greater effectiveness within ML-based fuzzing. In this paper, we delve into this paradigm shift, aiming to address fundamental challenges across different ML categories. We survey popular ML categories such as Traditional Machine Learning (TML), Deep Learning (DL), Reinforcement Learning (RL), and Deep Reinforcement Learning (DRL), to investigate their potential for enhancing traditional fuzzing approaches. We explore the respective advantages in each category of ML-based fuzzing, while also analysing the challenges unique to each category. Our work provides a comprehensive survey across the fuzzing domain and how machine learning techniques have been utilised, that we believe will be of use to future researchers in this domain. Sadegh Bamohabbat Chafjiri, Philip A. Legg, Jun Hong 0001, Michail-Antisthenis I. Tsompanas |
Comput. Secur. | 2 |
| 2023 | Longitudinal risk-based security assessment of docker software container imagesabstractAs the use of software containerisation has increased, so too has the need for security research on their usage, with various surveys and studies conducted to assess the overall security posture of software container images. To date, there has been very little work that has taken a longitudinal view of container security to observe whether vulnerabilities are being resolved over time, as well as understanding the real-world implications of reported vulnerabilities, to assess the evolving security posture. In this work, we study the evolution of 380 software container images across 3 analysis periods between July 2022 and January 2023 to analyse maintenance and vulnerabilities factors over time. We sample across the 3 DockerHub categories: Official, Verified and OSS (Sponsored) Open Source Software. We found that the number of vulnerabilities present increased over time despite many containers receiving regular updates by providers. We also found that the choice of container OS can dramatically impact the number of reported vulnerabilities present over time, with Debian-based images typically having many more vulnerabilities that other Linux distributions, and with some containers still reporting vulnerabilities that date back as far as 1999. However, when taking into account additional reported attributes such as the attack vector required and the existence of a public exploit rated higher than negligible, we found that for each analysis period, less than 1% of all vulnerabilities present what we would consider as high risk real-world impact. Through our investigation, we aim to improve the understanding of the threat landscape posed by software containerisation that is further complicated by the discrepancies between different vulnerability reporting tools. Alan Mills, Philip A. Legg |
Comput. Secur. | 3 |
| 2023 | Defending against adversarial machine learning attacks using hierarchical learning: A case study on network traffic attack classificationabstractMachine learning is key for automated detection of malicious network activity to ensure that computer networks and organizations are protected against cyber security attacks. Recently, there has been growing interest in the domain of adversarial machine learning, which explores how a machine learning model can be compromised by an adversary, resulting in misclassified output. Whilst to date, most focus has been given to visual domains, the challenge is present in all applications of machine learning where a malicious attacker would want to cause unintended functionality, including cyber security and network traffic analysis. We first present a study on conducting adversarial attacks against a well-trained network traffic classification model. We show how well-crafted adversarial examples can be constructed so that known attack types are misclassified by the model as benign activity. To combat this, we present a novel defensive strategy based on hierarchical learning to help reduce the attack surface that an adversarial example can exploit within the constraints of the parameter space of the intended attack. Our results show that our defensive learning model can withstand crafted adversarial attacks and can achieve classification accuracy in line with our original model when not under attack. Andrew McCarthy, Essam Ghadafi, Panagiotis Andriotis, Philip A. Legg |
J. Inf. Secur. Appl. | 4 |
| 2021 | Deep Learning-Based Security Behaviour Analysis in IoT Environments: A SurveyabstractInternet of Things (IoT) applications have been used in a wide variety of domains ranging from smart home, healthcare, smart energy, and Industrial 4.0. While IoT brings a number of benefits including convenience and efficiency, it also introduces a number of emerging threats. The number of IoT devices that may be connected, along with the ad hoc nature of such systems, often exacerbates the situation. Security and privacy have emerged as significant challenges for managing IoT. Recent work has demonstrated that deep learning algorithms are very efficient for conducting security analysis of IoT systems and have many advantages compared with the other methods. This paper aims to provide a thorough survey related to deep learning applications in IoT for security and privacy concerns. Our primary focus is on deep learning enhanced IoT security. First, from the view of system architecture and the methodologies used, we investigate applications of deep learning in IoT security. Second, from the security perspective of IoT systems, we analyse the suitability of deep learning to improve security. Finally, we evaluate the performance of deep learning in IoT system security. Yawei Yue, Shancang Li, Philip A. Legg, Fuzhong Li |
Secur. Commun. Networks | 3 |
| 2018 | Predicting User Confidence During Visual Decision MakingabstractPeople are not infallible consistent “oracles”: their confidence in decision-making may vary significantly between tasks and over time. We have previously reported the benefits of using an interface and algorithms that explicitly captured and exploited users’ confidence: error rates were reduced by up to 50% for an industrial multi-class learning problem; and the number of interactions required in a design-optimisation context was reduced by 33%. Having access to users’ confidence judgements could significantly benefit intelligent interactive systems in industry, in areas such as intelligent tutoring systems and in health care. There are many reasons for wanting to capture information about confidence implicitly . Some are ergonomic, but others are more “social”—such as wishing to understand (and possibly take account of) users’ cognitive state without interrupting them. We investigate the hypothesis that users’ confidence can be accurately predicted from measurements of their behaviour. Eye-tracking systems were used to capture users’ gaze patterns as they undertook a series of visual decision tasks, after each of which they reported their confidence on a 5-point Likert scale. Subsequently, predictive models were built using “conventional” machine learning approaches for numerical summary features derived from users’ behaviour. We also investigate the extent to which the deep learning paradigm can reduce the need to design features specific to each application by creating “gaze maps”—visual representations of the trajectories and durations of users’ gaze fixations—and then training deep convolutional networks on these images. Treating the prediction of user confidence as a two-class problem (confident/not confident), we attained classification accuracy of 88% for the scenario of new users on known tasks, and 87% for known users on new tasks. Considering the confidence as an ordinal variable, we produced regression models with a mean absolute error of ≈0.7 in both cases. Capturing just a simple subset of non-task-specific numerical features gave slightly worse, but still quite high accuracy (e.g., MAE ≈ 1.0). Results obtained with gaze maps and convolutional networks are competitive, despite not having access to longer-term information about users and tasks, which was vital for the “summary” feature sets. This suggests that the gaze-map-based approach forms a viable, transferable alternative to handcrafting features for each different application. These results provide significant evidence to confirm our hypothesis, and offer a way of substantially improving many interactive artificial intelligence applications via the addition of cheap non-intrusive hardware and computationally cheap prediction algorithms. Jim E. Smith, Philip A. Legg, Milos Matovic, Kristofer Kinsey |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2015 | Visualizing the insider threat: challenges and tools for identifying malicious user activityabstractOne of the greatest challenges for managing organisational cyber security is the threat that comes from those who operate within the organisation. With entitled access and knowledge of organisational processes, insiders who choose to attack have the potential to cause serious impact, such as financial loss, reputational damage, and in severe cases, could even threaten the existence of the organisation. Security analysts therefore require sophisticated tools that allow them to explore and identify user activity that could be indicative of an imminent threat to the organisation. In this work, we discuss the challenges associated with identifying insider threat activity, along with the tools that can help to combat this problem. We present a visual analytics approach that incorporates multiple views, including a user selection tool that indicates anomalous behaviour, an interactive Principal Component Analysis (iPCA) tool that aids the analyst to assess the reasoning behind the anomaly detection results, and an activity plot that visualizes user and role activity over time. We demonstrate our approach using the Carnegie Mellon University CERT Insider Threat Dataset to show how the visual analytics workflow supports the Information-Seeking mantra. Philip A. Legg |
VizSEC | 1 |
| 2015 | Feature Neighbourhood Mutual Information for multi-modal image registration: An application to eye fundus imaging
Philip A. Legg, Paul L. Rosin, David Marshall 0001, James E. Morgan |
Pattern Recognit. | 1 |
| 2013 | Force-Directed Parallel CoordinatesabstractParallel coordinates are a well-known and valuable technique for the analysis and visualization of high dimen-sional data sets. However, while Inselberg emphasizes that the strength of parallel coordinates as a methodology is rooted in exploration and interactivity, the set of interac-tion techniques is currently limited. Axes can be re-ordered and brushing (simple, angular or multi-dimensional) can be performed. In this paper, we propose a force-directed algorithm and related interaction techniques to support the exploration of parallel coordinate plots through a physi-cal metaphor. Our parallel-coordinates visualization of-fers novel user interaction beyond the standard techniques by allowing the user to rotate the axis according to force-directed polylines. The new interaction provides the user with a more immersive experience for data exploration that results in greater intuition of the data, especially in cases where many polylines overlap. We demonstrate our ap-proach, then present the results of a qualitative evaluation of the system. 1 Rick Walker, Philip A. Legg, Serban R. Pop, Zhao Geng, Robert S. Laramee, Jonathan Roberts 0002 |
IV | 2 |
| 2013 | Automated 3-D Animation From Snooker Videos With Information-Theoretical OptimizationabstractAutomated 3-D modeling from real sports videos can provide useful resources for visual design in sports-related computer games, saving a lot of effort in manual design of visual contents. However, image-based 3-D reconstruction usually suffers from inaccuracy caused by statistic image analysis. In this paper, we propose an information-theoretical scheme to minimize errors of automated 3-D modeling from monocular sports videos. In the proposed scheme, mutual information (MI) was exploited to compute the fitting scores of a 3-D model against the observed single-view scene, and the optimization of model fitting was carried out subsequently. With this optimization scheme, errors in model fitting were minimized without human intervention, allowing automated reconstruction of 3-D animation from consecutive monocular video frames at high accuracy. In our work, the Snooker videos were taken as our case study, balls were positioned in 3-D space from single-view frames, and 3-D animation was reproduced from real Snooker videos. Our experimental results validated that the proposed information-theoretical scheme can help attain better accuracy in the automated reconstruction of 3-D animation, and demonstrated that information-theoretical evaluation can be an effective approach for model-based reconstruction from single-view videos. Richard Jiang 0001, Matthew L. Parry, Philip A. Legg, David H. S. Chung, Iwan W. Griffiths |
IEEE Trans. Comput. Intell. AI Games | 3 |
| 2013 | Transformation of an Uncertain Video Search Pipeline to a Sketch-Based Visual Analytics LoopabstractTraditional sketch-based image or video search systems rely on machine learning concepts as their core technology. However, in many applications, machine learning alone is impractical since videos may not be semantically annotated sufficiently, there may be a lack of suitable training data, and the search requirements of the user may frequently change for different tasks. In this work, we develop a visual analytics systems that overcomes the shortcomings of the traditional approach. We make use of a sketch-based interface to enable users to specify search requirement in a flexible manner without depending on semantic annotation. We employ active machine learning to train different analytical models for different types of search requirements. We use visualization to facilitate knowledge discovery at the different stages of visual analytics. This includes visualizing the parameter space of the trained model, visualizing the search space to support interactive browsing, visualizing candidature search results to support rapid interaction for active learning while minimizing watching videos, and visualizing aggregated information of the search results. We demonstrate the system for searching spatiotemporal attributes from sports video to identify key instances of the team and player performance. Philip A. Legg, David H. S. Chung, Matthew L. Parry, Rhodri Bown, Mark W. Jones 0001, Iwan W. Griffiths, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2012 | MatchPad: Interactive Glyph-Based Visualization for Real-Time Sports Performance AnalysisabstractAbstract Today real‐time sports performance analysis is a crucial aspect of matches in many major sports. For example, in soccer and rugby, team analysts may annotate videos during the matches by tagging specific actions and events, which typically result in some summary statistics and a large spreadsheet of recorded actions and events. To a coach, the summary statistics (e.g., the percentage of ball possession) lacks sufficient details, while reading the spreadsheet is time‐consuming and making decisions based on the spreadsheet in real‐time is thereby impossible. In this paper, we present a visualization solution to the current problem in real‐time sports performance analysis. We adopt a glyph‐based visual design to enable coaching staff and analysts to visualize actions and events “at a glance”. We discuss the relative merits of metaphoric glyphs in comparison with other types of glyph designs in this particular application. We describe an algorithm for managing the glyph layout at different spatial scales in interactive visualization. We demonstrate the use of this technical approach through its application in rugby, for which we delivered the visualization software,MatchPad, on a tablet computer. The MatchPad was used by the Welsh Rugby Union during the Rugby World Cup 2011. It successfully helped coaching staff and team analysts to examine actions and events in detail whilst maintaining a clear overview of the match, and assisted in their decision making during the matches. It also allows coaches to convey crucial information back to the players in a visually‐engaging manner to help improve their performance. Philip A. Legg, David H. S. Chung, Matthew L. Parry, Mark W. Jones 0001, R. Long, Iwan W. Griffiths, Min Chen 0001 |
Comput. Graph. Forum | 1 |
| 2011 | Intelligent filtering by semantic importance for single-view 3D reconstruction from Snooker videoabstractIn this paper we investigate the challenge of 3D reconstruction from Snooker video data. We propose a system pipeline for intelligent filtering based on semantic importance in Snooker. The system can be divided into table detection and correction, followed by ball detection, classification and tracking. It is apparent from previous work that there are several challenges presented here. Firstly, previous methods tend to use a fixed top-down camera mounted above the table. To capture a full table view from this is challenging due to space limitations above the table. Instead, we capture video data from a tripod and correct the viewpoint through processing. Secondly, previous methods tend to simply detect the balls without considering other interfering objects such as player and cue. This becomes even more apparent when the player strikes the cue ball. Our intelligent filtering avoids such issues to give accurate 3D table reconstruction. Philip A. Legg, Matthew L. Parry, David H. S. Chung, Richard Jiang 0001, Adrian Morris, Iwan W. Griffiths, David Marshall 0001, Min Chen 0001 |
ICIP | 1 |
| 2011 | Hierarchical Event Selection for Video Storyboards with a Case Study on Snooker Video VisualizationabstractVideo storyboard, which is a form of video visualization, summarizes the major events in a video using illustrative visualization. There are three main technical challenges in creating a video storyboard, (a) event classification, (b) event selection and (c) event illustration. Among these challenges, (a) is highly application-dependent and requires a significant amount of application specific semantics to be encoded in a system or manually specified by users. This paper focuses on challenges (b) and (c). In particular, we present a framework for hierarchical event representation, and an importance-based selection algorithm for supporting the creation of a video storyboard from a video. We consider the storyboard to be an event summarization for the whole video, whilst each individual illustration on the board is also an event summarization but for a smaller time window. We utilized a 3D visualization template for depicting and annotating events in illustrations. To demonstrate the concepts and algorithms developed, we use Snooker video visualization as a case study, because it has a concrete and agreeable set of semantic definitions for events and can make use of existing techniques of event detection and 3D reconstruction in a reliable manner. Nevertheless, most of our concepts and algorithms developed for challenges (b) and (c) can be applied to other application areas. Matthew L. Parry, Philip A. Legg, David H. S. Chung, Iwan W. Griffiths, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2009 | A Robust Solution to Multi-modal Image Registration by Combining Mutual Information with Multi-scale Derivatives
Philip A. Legg, Paul L. Rosin, David Marshall 0001, James E. Morgan |
MICCAI (1) | 1 |