Chris Hicks

dblp:262/6509 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DRMD: Deep Reinforcement Learning for Malware Detection Under Concept Drift
abstract
Malware detection in real-world settings must deal with evolving threats, limited labeling budgets, and uncertain predictions. Traditional classifiers, without additional mechanisms, struggle to maintain performance under concept drift in malware domains, as their supervised learning formulation cannot optimize when to defer decisions to manual labeling and adaptation. Modern malware detection pipelines combine classifiers with monthly active learning (AL) and rejection mechanisms to mitigate the impact of concept drift. In this work, we develop a novel formulation of malware detection as a one-step Markov Decision Process and train a deep reinforcement learning (DRL) agent, simultaneously optimizing sample classification performance and rejecting high-risk samples for manual labeling. We evaluated the joint detection and drift mitigation policy learned by the DRL-based Malware Detection (DRMD) agent through time-aware evaluations on Android malware datasets subject to realistic drift requiring multi-year performance stability. The policies learned under these conditions achieve a higher Area Under Time (AUT) performance compared to standard classification approaches used in the domain, showing improved resilience to concept drift. Specifically, the DRMD agent achieved an average AUT improvement of 8.66 and 10.90 for the classification-only and classification-rejection policies, respectively. Our results demonstrate for the first time that DRL can facilitate effective malware detection and improved resiliency to concept drift in the dynamic setting of Android malware detection.
Shae McFadden, Myles Foley, Mario D'Onghia, Chris Hicks, Vasilios Mavroudis, Nicola Paoletti, Fabio Pierazzi
AAAI4
2026 Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
abstract
Deep Reinforcement Learning (DRL) systems are increasingly used in safety-critical applications, yet their security remains severely underexplored. This work investigates backdoor attacks, which implant hidden triggers that cause malicious actions only when specific inputs appear in the observation space. Existing DRL backdoor research focuses solely on training-time attacks requiring full adversarial access to the training pipeline. In contrast, we reveal critical vulnerabilities across the DRL supply chain where backdoors can be embedded with significantly reduced adversarial privileges. We introduce two novel attacks: (1) TrojanentRL, which exploits component-level flaws to implant a persistent backdoor that survives full model retraining; and (2) InfrectroRL, a post-training backdoor attack which requires no access to training, validation, or test data. Empirical and analytical evaluations across six Atari environments show our attacks rival state-of-the-art training-time backdoor attacks while operating under much stricter adversarial constraints. We also demonstrate that InfrectroRL further evades two leading DRL backdoor defenses. These findings challenge the current research focus and highlight the urgent need for robust defenses.
Sanyam Vyas, Alberto Caron, Chris Hicks, Pete Burnap, Vasilios Mavroudis
AAAI3
2026 CoinJoin ecosystem insights for Wasabi 1.x, Wasabi 2.x and Whirlpool coordinator-based privacy mixers
abstract
Coinjoin is a collaborative Bitcoin transaction designed to enhance users' privacy by joining coins of multiple parties. Coinjoin implementations with a centralized trust-minimized coordinator have mixed more than 391,000 bitcoins since 2018. In June 2024, law enforcement actions led to the shutdown of coordinators of all three leading coinjoin designs-Wasabi 1.x, Wasabi 2.x, and Whirlpool-prompting a proliferation of new independent coordinators. Prior studies of this ecosystem offer little information beyond coinjoin detection and aggregation of related statistics, primarily due to the unavailability of information (intentionally) obscured by coinjoins. To overcome this limitation, we combine on-chain transaction analysis, coordinator monitoring, and active coinjoin participation, complemented by tailored analysis and visualizations, to: 1) provide a model for estimation of the number of active users at a time, 2) detect previously unknown coordinators, 3) retrospectively observe activity patterns of prominent parties, and 4) deliver continuously updated insights into the ecosystem. We show that while Wasabi 1.x remains inactive and recently resumed Whirlpool exhibits little activity, the Wasabi 2.x mixing has exceeded pre-shutdown levels under the new coordinator kruw.io, which now accounts for 99% of mixed inputs (over 46,000 bitcoins in the past 18 months) with raising concurrently active users now estimated to 70+ in average coinjoin, providing the first data-based estimate of the participants’ anonymity set. We revealed a previously unknown but significant coordinator active from early May to November 2024, with a method capable of identifying emerging coordinators. Although WW2 conceals large-input wallet activity more effectively than WW1, it is still detected by the proposed transaction-centric liquidity analysis.
Petr Svenda, Jiri Gavenda, Vasilios Mavroudis, Chris Hicks
Proc. Priv. Enhancing Technol.4
2025 Towards Autonomous Cyber Defence: Applying Systems Theoretic Process Analysis to Human-Machine Teaming
abstract
With cyberspace becoming ever more fiercely contested, blue team operations face a growing demand for more efficient and sophisticated capabilities. One way to meet this demand is by integrating Human-Machine Teaming (HMT) and autonomous cyber defence (ACD), enhancing blue team collaboration with AI-driven machines. A rigorous analysis of their interaction dynamics is essential to ensure these technologies are safely deployed in high-stakes environments.This study employs Systems Theoretic Process Analysis (STPA) to examine the control actions and feedback loops between machines and controlled processors in the context of ACD, identifying potential hazards and losses arising from unsafe control actions. Additionally, it establishes system constraints to mitigate risks and enhance the safety and reliability of AI-driven HMT operations.
Shu-Jui Chang, Chris Hicks, Vasilios Mavroudis, Iain Phillips 0002, Tim Watson
IJCNN2
2025 Statistical Parity with Exponential Weights
abstract
Statistical parity is one of the most foundational constraints in algorithmic fairness and privacy. In this paper, we show that statistical parity can be enforced efficiently in the adversarial contextual bandit setting while retaining strong performance guarantees. Specifically, we present a meta-algorithm that transforms any efficient implementation of Hedge (or, equivalently, any discrete Bayesian inference algorithm) into an efficient contextual bandit algorithm that guarantees exact statistical parity on every trial. Compared to any comparator that satisfies the same statistical parity constraint, the algorithm achieves the same asymptotic regret bound as running the equivalent instance of Exp4 for each group. We also address the scenario where the target parity distribution is unknown and must be estimated online. Finally, using online-to-batch conversion, we extend our approach to the batch classification setting.
Stephen Pasteris, Chris Hicks, Vasilios Mavroudis
NeurIPS2
2024 Online Convex Optimisation: The Optimal Switching Regret for all Segmentations Simultaneously
abstract
We consider the classic problem of online convex optimisation. Whereas the notion of static regret is relevant for stationary problems, the notion of switching regret is more appropriate for non-stationary problems. A switching regret is defined relative to any segmentation of the trial sequence, and is equal to the sum of the static regrets of each segment. In this paper we show that, perhaps surprisingly, we can achieve the asymptotically optimal switching regret on every possible segmentation simultaneously. Our algorithm for doing so is very efficient: having a space and per-trial time complexity that is logarithmic in the time-horizon. Our algorithm also obtains novel bounds on its dynamic regret: being adaptive to variations in the rate of change of the comparator sequence.
Stephen Pasteris, Chris Hicks, Vasilios Mavroudis, Mark Herbster
NeurIPS2
2023 Nearest Neighbour with Bandit Feedback
abstract
In this paper we adapt the nearest neighbour rule to the contextual bandit problem. Our algorithm handles the fully adversarial setting in which no assumptions at all are made about the data-generation process. When combined with a sufficiently fast data-structure for (perhaps approximate) adaptive nearest neighbour search, such as a navigating net, our algorithm is extremely efficient - having a per trial running time polylogarithmic in both the number of trials and actions, and taking only quasi-linear space. We give generic regret bounds for our algorithm and further analyse them when applied to the stochastic bandit problem in euclidean space. A side result of this paper is that, when applied to the online classification problem with stochastic labels, our algorithm can, under certain conditions, have sublinear regret whilst only finding a single nearest neighbour per trial - in stark contrast to the k-nearest neighbours algorithm.
Stephen Pasteris, Chris Hicks, Vasilios Mavroudis
NeurIPS2
2022 Autonomous Network Defence using Reinforcement Learning
abstract
In the network security arms race, the defender is significantly disadvantaged as they need to successfully detect and counter every malicious attack. In contrast, the attacker needs to succeed only once. To level the playing field, we investigate the effectiveness of autonomous agents in a realistic network defence scenario. We first outline the problem, provide the background on reinforcement learning and detail our proposed agent design. Using a network environment simulation, with 13 hosts spanning 3 subnets, we train a novel reinforcement learning agent and show that it can reliably defend continual attacks by two advanced persistent threat (APT) red agents: one with complete knowledge of the network layout and another which must discover resources through exploration but is more general.
Myles Foley, Chris Hicks, Kate Highnam, Vasilios Mavroudis
AsiaCCS2