Sam Devlin

dblp:64/7502 · DBLP profile ↗
← Back
39ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0002-7769-3090ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 since 2021Human-computer interaction and ubiquitous computing · 8 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2025 Scaling Laws for Pre-training Agents and World Models
abstract
The performance of embodied agents has been shown to improve by increasing model parameters, dataset size, and compute. This has been demonstrated in domains from robotics to video games, when generative learning objectives on offline datasets (pre-training) are used to model an agent’s behavior (imitation learning) or their environment (world modeling). This paper characterizes the role of scale in these tasks more precisely. Going beyond the simple intuition that ‘bigger is better’, we show that the same types of power laws found in language modeling also arise in world modeling and imitation learning (e.g. between loss and optimal model size). However, the coefficients of these laws are heavily influenced by the tokenizer, task & architecture – this has important implications on the optimal sizing of models and data.
Tim Pearce, Tabish Rashid, David Bignell, Raluca Georgescu, Sam Devlin, Katja Hofmann
ICML5
2025 Difference rewards policy gradients
abstract
Abstract Policy gradient methods have become one of the most popular classes of algorithms for multi-agent reinforcement learning. A key challenge, however, that is not addressed by many of these methods is multi-agent credit assignment: assessing an agent’s contribution to the overall performance, which is crucial for learning good policies. We propose a novel algorithm called Dr.Reinforce that explicitly tackles this by combining difference rewards with policy gradients to allow for learning decentralized policies when the reward function is known. By differencing the reward function directly, Dr.Reinforce avoids difficulties associated with learning the Q-function as done by counterfactual multi-agent policy gradients (COMA), a state-of-the-art difference rewards method. For applications where the reward function is unknown, we show the effectiveness of a version of Dr.Reinforce that learns an additional reward network that is used to estimate the difference rewards.
Jacopo Castellini, Sam Devlin, Frans A. Oliehoek, Rahul Savani
Neural Comput. Appl.2
2023 Navigates Like Me: Understanding How People Evaluate Human-Like AI in Video Games
abstract
We aim to understand how people assess human likeness in navigation produced by people and artificially intelligent (AI) agents in a video game. To this end, we propose a novel AI agent with the goal of generating more human-like behavior. We collect hundreds of crowd-sourced assessments comparing the human-likeness of navigation behavior generated by our agent and baseline AI agents with human-generated behavior. Our proposed agent passes a Turing Test, while the baseline agents do not. By passing a Turing Test, we mean that human judges could not quantitatively distinguish between videos of a person and an AI agent navigating. To understand what people believe constitutes human-like navigation, we extensively analyze the justifications of these assessments. This work provides insights into the characteristics that people consider human-like in the context of goal-directed video game navigation, which is a key step for further improving human interactions with AI agents.
Stephanie Milani, Arthur Juliani, Ida Momennejad, Raluca Georgescu, Jaroslaw Rzepecki, Ali Shaw, Gavin Costello, Fei Fang 0001, Sam Devlin, Katja Hofmann
CHI9
2023 Contrastive Meta-Learning for Partially Observable Few-Shot Learning
Adam Jelley, Amos J. Storkey, Antreas Antoniou, Sam Devlin
ICLR4
2023 Imitating Human Behaviour with Diffusion Models
Tim Pearce, Tabish Rashid, Anssi Kanervisto, David Bignell, Mingfei Sun 0001, Raluca Georgescu, Sergio Valcarcel Macua, Shan Zheng Tan, Ida Momennejad, Katja Hofmann, Sam Devlin
ICLR11
2022 Deterministic and Discriminative Imitation (D2-Imitation): Revisiting Adversarial Imitation for Sample Efficiency
abstract
Sample efficiency is crucial for imitation learning methods to be applicable in real-world applications. Many studies improve sample efficiency by extending adversarial imitation to be off-policy regardless of the fact that these off-policy extensions could either change the original objective or involve complicated optimization. We revisit the foundation of adversarial imitation and propose an off-policy sample efficient approach that requires no adversarial training or min-max optimization. Our formulation capitalizes on two key insights: (1) the similarity between the Bellman equation and the stationary state-action distribution equation allows us to derive a novel temporal difference (TD) learning approach; and (2) the use of a deterministic policy simplifies the TD learning. Combined, these insights yield a practical algorithm, Deterministic and Discriminative Imitation (D2-Imitation), which oper- ates by first partitioning samples into two replay buffers and then learning a deterministic policy via off-policy reinforcement learning. Our empirical results show that D2-Imitation is effective in achieving good sample efficiency, outperforming several off-policy extension approaches of adversarial imitation on many control tasks.
Mingfei Sun 0001, Sam Devlin, Katja Hofmann, Shimon Whiteson
AAAI2
2022 Adaptive Scaffolding in Block-Based Programming via Synthesizing New Tasks as Pop Quizzes
Ahana Ghosh, Sebastian Tschiatschek, Sam Devlin, Adish Singla
AIED (1)3
2022 Turning Zeroes into Non-Zeroes: Sample Efficient Exploration with Monte Carlo Graph Search
abstract
Monte Carlo Tree Search (MCTS) has proven to be a staple method in Game Artificial Intelligence for creating agents that can perform well in complex environments without requiring domain-specific knowledge. The main downside of this planning based algorithm is the high computational budget needed to recommend an action. The fundamental cause of this is a vast search space caused by a high branching factor, and the difficulty to create a good heuristic function to guide the search without leveraging domain-specific knowledge. Recent advances in the field proposed a new planning based method called Monte Carlo Graph Search (MCGS), which uses a graph instead of a tree to plan its next action, reducing the branching factor and consequently increasing the performance of the search. In this paper, we propose several modifications that optimize the performance by increasing the sample efficiency of MCGS. The use of frontier for node selection, improving the rollout phase by doing stored rollouts, and a generalized approach to guide the search by incorporating a domain-independent online novelty detection method. Together these enhancements enable MCGS to solve sparse reward environments while using a significantly lower computational budget than MCTS.
Marko Tot, Michelangelo Conserva, Diego Perez Liebana, Sam Devlin
CoG4
2022 Imitating Playstyle with Dynamic Time Warping Imitation
abstract
Imitation learning has been demonstrated as a useful technique in automatic game testing and the development of believable Non-Player Characters (NPCs). However, imitation learning methods typically focus on learning a policy to complete a task without consideration about the playstyle used. In this work we consider the case where the task is to imitate a given playstyle. We defined a player’s playstyle based on the strategies they use in order to complete the overall task. This has been achieved by rewarding a learning agent based on the similarity of the agent and demonstration trajectories, within a learnt representation space. This allows the playstyle to be learnt in levels that differ to the one the demonstrations were collected in.
Mark Ferguson, Sam Devlin, Daniel Kudenko, James Alfred Walker
FDG2
2022 Uni[MASK]: Unified Inference in Sequential Decision Problems
abstract
Randomly masking and predicting word tokens has been a successful approach in pre-training language models for a variety of downstream tasks. In this work, we observe that the same idea also applies naturally to sequential decision making, where many well-studied tasks like behavior cloning, offline RL, inverse dynamics, and waypoint conditioning correspond to different sequence maskings over a sequence of states, actions, and returns. We introduce the UniMASK framework, which provides a unified way to specify models which can be trained on many different sequential decision making tasks. We show that a single UniMASK model is often capable of carrying out many tasks with performance similar to or better than single-task models. Additionally, after fine-tuning, our UniMASK models consistently outperform comparable single-task models.
Micah Carroll, Orr Paradise, Jessy Lin, Raluca Georgescu, Mingfei Sun 0001, David Bignell, Stephanie Milani, Katja Hofmann, Matthew J. Hausknecht, Anca D. Dragan, Sam Devlin
NeurIPS11
2022 Rolling Horizon Evolutionary Algorithms for General Video Game Playing
abstract
Game-playing evolutionary algorithms, specifically rolling horizon evolutionary algorithms (RHEA), have recently managed to beat the state of the art in win rate across many video games. However, the best results in a game are highly dependent on the specific configuration of modifications introduced over several papers, each adding additional parameters to the core algorithm. Furthermore, the best previously published parameters have been found from only a few human-picked combinations, as the possibility space has grown beyond exhaustive search. This article presents the state of the art in RHEA, combining all modifications described in the literature, as well as new ones. We then use a parameter optimizer, the$N$-tuple bandit evolutionary algorithm, to find the best combination of parameters in 20 games from the general video game Artificial Intelligence (AI) framework. Furthermore, we analyze the algorithm’s parameters and some interesting combinations revealed through the optimization process. Finally, we find new state of the art solutions on several games by automatically exploring the large parameter space of RHEA.
Raluca D. Gaina, Sam Devlin, Simon M. Lucas, Diego Perez Liebana
IEEE Trans. Games2
2022 A Comparison of Self-Play Algorithms Under a Generalized Framework
abstract
The notion of self-play, albeit often cited in multiagent reinforcement learning as a process by which to train agent policies from scratch, has received little efforts to be taxonomized within a formal model. We present a formalized framework, with clearly defined assumptions, which encapsulates the meaning of self-play as abstracted from various existing self-play algorithms. This framework is framed as an approximation to a theoretical solution concept for multiagent training. Through a novel qualitative visualization metric, on a simple environment, we show that different self-play algorithms generate different distributions of episode trajectories, leading to different explorations of the policy space by the learning agents. Quantitatively, on two environments, we analyze the learning dynamics of policies trained under different self-play algorithms captured under our framework and perform cross self-play performance comparisons. Our results indicate that, throughout training, various widely used self-play algorithms exhibit cyclic policy evolutions and that the choice of self-play algorithm significantly affects the final performance of trained agents.
Daniel Hernández 0008, Kevin Denamganaï, Sam Devlin, Spyridon Samothrakis, James Alfred Walker
IEEE Trans. Games3
2021 Meta-Learning Divergences for Variational Inference
abstract
Variational inference (VI) plays an essential role in approximate Bayesian inference due to its computational efficiency and broad applicability. Crucial to the performance of VI is the selection of the associated divergence measure, as VI approximates the intractable distribution by minimizing this divergence. In this paper we propose a meta-learning algorithm to learn the divergence metric suited for the task of interest, automating the design of VI methods. In addition, we learn the initialization of the variational parameters without additional cost when our method is deployed in the few-shot learning scenarios. We demonstrate our approach outperforms standard VI on Gaussian mixture distribution approximation, Bayesian neural network regression, image generation with variational autoencoders and recommender systems with a partial variational autoencoder.
Ruqi Zhang, Yingzhen Li, Christopher De Sa, Sam Devlin, Cheng Zhang 0005
AISTATS4
2021 Navigation Turing Test (NTT): Learning to Evaluate Human-Like Navigation
abstract
A key challenge on the path to developing agents that learn complex human-like behavior is the need to quickly and accurately quantify human-likeness. While human assessments of such behavior can be highly accurate, speed and scalability are limited. We address these limitations through a novel automated Navigation Turing Test (ANTT) that learns to predict human judgments of human-likeness. We demonstrate the effectiveness of our automated NTT on a navigation task in a complex 3D environment. We investigate six classification models to shed light on the types of architectures best suited to this task, and validate them against data collected through a human NTT. Our best models achieve high accuracy when distinguishing true human and agent behavior. At the same time, we show that predicting finer-grained human assessment of agents’ progress towards human-like behavior remains unsolved. Our work takes an important step towards agents that more effectively learn complex human-like behavior.
Sam Devlin, Raluca Georgescu, Ida Momennejad, Jaroslaw Rzepecki, Evelyn Zuniga, Gavin Costello, Guy Leroy, Ali Shaw, Katja Hofmann
ICML1
2021 Strategically efficient exploration in competitive multi-agent reinforcement learning
abstract
High sample complexity remains a barrier to the application of reinforcement learning (RL), particularly in multi-agent systems. A large body of work has demonstrated that exploration mechanisms based on the principle of optimism under uncertainty can significantly improve the sample efficiency of RL in single agent tasks. This work seeks to understand the role of optimistic exploration in non-cooperative multi-agent settings. We will show that, in zero-sum games, optimistic exploration can cause the learner to waste time sampling parts of the state space that are irrelevant to strategic play, as they can only be reached through cooperation between both players. To address this issue, we introduce a formal notion of strategically efficient exploration in Markov games, and use this to develop two strategically efficient learning algorithms for finite Markov games. We demonstrate that these methods can be significantly more sample efficient than their optimistic counterparts.
Robert Tyler Loftin, Aadirupa Saha, Sam Devlin, Katja Hofmann
UAI3
2021 Win Prediction in Multiplayer Esports: Live Professional Match Prediction
abstract
Esports are competitive videogames watched by audiences. Most esports generate detailed data for each match that are publicly available. Esports analytics research is focused on predicting match outcomes. Previous research has emphasized prematch prediction and used data from amateur games, which are more easily available than those from professional level. However, the commercial value of win prediction exists at the professional level. Furthermore, predicting real-time data is unexplored, as is its potential for informing audiences. Here, we present the first comprehensive case study on live win prediction in a professional esport. We provide a literature review for win prediction in a multiplayer online battle arena (MOBA) esport. This article evaluates the first professional-level prediction models for live DotA 2 matches, one of the most popular MOBA games, and trials it at a major international esports tournament. Using standard machine learning models, feature engineering and optimization, our model is up to 85% accurate after 5 min of gameplay. Our analyses highlight the need for algorithm evaluation and optimization. Finally, we present implications for the esports/game analytics domains, describe commercial opportunities and practical challenges, and propose a set of evaluation criteria for research on esports win prediction.
Victoria J. Hodge, Sam Devlin, Nick Sephton, Florian Block, Peter I. Cowling, Anders Drachen
IEEE Trans. Games2
2020 Player Style Clustering without Game Variables
abstract
Player clustering when applied to the field of video games has several potential applications. For example, the evaluation of the composition of a player base or the generation of AI agents with identified playing styles. These agents can then be used for either the testing of new game content or used directly to enhance a player’s gaming experience. Most current player clustering techniques focus on the use of internal game variables. This raises two main issues: (1) the availability of game variables, as source code access is required to log them and hence limits the data sources that can be used, and (2) the choice of game variables can introduce unintended bias in the types of play style extracted. In this work, a hybrid unsupervised frame encoder and a ‘reference-based’ clustering algorithm are both proposed and combined to allow clustering from raw game play videos. It is shown that the proposed methods are most beneficial when the types of play styles are unknown.
Mark Ferguson, Sam Devlin, Daniel Kudenko, James Alfred Walker
FDG2
2020 Automatic Similarity Detection in LEGO Ducks
Mark Ferguson, Sebastian Deterding, Andreas Lieberoth, Marc Malmdorf Andersen, Sam Devlin, Daniel Kudenko, James Alfred Walker
ICCC5
2020 AMRL: Aggregated Memory For Reinforcement Learning
Jacob Beck, Kamil Ciosek, Sam Devlin, Sebastian Tschiatschek, Cheng Zhang 0005, Katja Hofmann
ICLR3
2019 MazeExplorer: A Customisable 3D Benchmark for Assessing Generalisation in Reinforcement Learning
abstract
This paper presents a customisable 3D benchmark for assessing generalisability of reinforcement learning agents based on the 3D first-person game Doom and open source environment VizDoom. As a sample use-case we show that different domain randomisation techniques during training in a key-collection navigation task can help to improve agent performance on unseen evaluation maps.
Luke Harries, Sebastian Lee, Jaroslaw Rzepecki, Katja Hofmann, Sam Devlin
CoG5
2019 A Generalized Framework for Self-Play Training
abstract
Throughout scientific history, overarching theoretical frameworks have allowed researchers to grow beyond personal intuitions and culturally biased theories. They allow to verify and replicate existing findings, and to link disconnected results. The notion of self-play, albeit often cited in multiagent Reinforcement Learning, has never been grounded in a formal model. We present a formalized framework, with clearly defined assumptions, which encapsulates the meaning of self-play as abstracted from various existing self-play algorithms. This framework is framed as an approximation to a theoretical solution concept for multiagent training. On a simple environment, we qualitatively measure how well a subset of the captured self-play methods approximate this solution when paired with the famous PPO algorithm. The results indicate that throughout training the trained policies exhibit cyclic evolutions, showing that self-play research is still at an early stage.
Daniel Hernández 0008, Kevin Denamganaï, Alex Yuan Gao, Peter York, Sam Devlin, Spyridon Samothrakis, James Alfred Walker
CoG5
2019 Win or Learn Fast Proximal Policy Optimisation
abstract
AI agents within video games are often required to compete within an environment shared by many other agents. This problem can be tackled by multi-agent reinforcement learning (MARL). One solution to MARL is to learn a Nash Equilibrium Strategy (NES) that guarantees a known minimum payoff when playing against other rational agents. We focus on one approach for learning a NES, Win or Learn Fast (WoLF), WoLF has been shown to converge towards a NES in a variety of matrix-games and grid based games. Research into Deep MARL has focused on performance against opponent agents and with limited quantitative results regarding learning a NES. We present a systematic empirical investigation into the ability of Proximal Policy Optimisation (PPO) to learn a NES, showing instability in certain matrix games. We then present an extension, WoLF-PPO, that is able to learn a policy that is closer to the NES.
Dino Stephen Ratcliffe, Katja Hofmann, Sam Devlin
CoG3
2019 Generalization in Reinforcement Learning with Selective Noise Injection and Information Bottleneck
abstract
The ability for policies to generalize to new environments is key to the broad application of RL agents. A promising approach to prevent an agent’s policy from overfitting to a limited set of training environments is to apply regularization techniques originally developed for supervised learning. However, there are stark differences between supervised learning and RL. We discuss those differences and propose modifications to existing regularization techniques in order to better adapt them to RL. In particular, we focus on regularization techniques relying on the injection of noise into the learned function, a family that includes some of the most widely used approaches such as Dropout and Batch Normalization. To adapt them to RL, we propose Selective Noise Injection (SNI), which maintains the regularizing effect the injected noise has, while mitigating the adverse effects it has on the gradient quality. Furthermore, we demonstrate that the Information Bottleneck (IB) is a particularly well suited regularization technique for RL as it is effective in the low-data regime encountered early on in training RL agents. Combining the IB with SNI, we significantly outperform current state of the art results, including on the recently proposed generalization benchmark Coinrun.
Maximilian Igl, Kamil Ciosek, Yingzhen Li, Sebastian Tschiatschek, Cheng Zhang 0005, Sam Devlin, Katja Hofmann
NeurIPS6
2019 Type inference in flexible model-driven engineering using classification algorithms
abstract
Flexible or bottom-up model-driven engineering (MDE) is an emerging approach to domain and systems modelling. Domain experts, who have detailed domain knowledge, typically lack the technical expertise to transfer this knowledge using traditional MDE tools. Flexible MDE approaches tackle this challenge by promoting the use of simple drawing tools to increase the involvement of domain experts in the language definition process. In such approaches, no metamodel is created upfront, but instead the process starts with the definition of example models that will be used to infer the metamodel. Pre-defined metamodels created by MDE experts may miss important concepts of the domain and thus restrict their expressiveness. However, the lack of a metamodel, that encodes the semantics of conforming models has some drawbacks, among others that of having models with elements that are unintentionally left untyped. In this paper, we propose the use of classification algorithms to help with the inference of such untyped elements. We evaluate the proposed approach in a number of random generated example models from various domains. The correct type prediction varies from 23 to 100% depending on the domain, the proportion of elements that were left untyped and the prediction algorithm used.
Athanasios Zolotas, Nicholas Drivalos Matragkas, Sam Devlin, Dimitrios S. Kolovos, Richard F. Paige
Softw. Syst. Model.3
2019 The Text-Based Adventure AI Competition
abstract
In 2016-2018 at the IEEE Conference on Computational Intelligence in Games, the authors of this paper ran a competition for agents that can play classic text-based adventure games. This competition fills a gap in existing game artificial intelligence (AI) competitions that have typically focused on traditional card/board games or modern video games with graphical interfaces. By providing a platform for evaluating agents in textbased adventures, the competition provides a novel benchmark for game AI with unique challenges for natural language understanding and generation. This paper summarizes the three competitions ran in 2016-2018 (including details of open-source implementations of both the competition framework and our competitors) and presents the results of an improved evaluation of these competitors across 20 games.
Timothy Atkinson 0001, Hendrik Baier, Tara Copplestone, Sam Devlin, Jerry Swan
IEEE Trans. Games4
2019 Emulating Human Play in a Leading Mobile Card Game
abstract
Monte Carlo tree search (MCTS) has become a popular solution for game artificial intelligence (AI), capable of creating strong game playing opponents. However, the emergent playstyle of agents using MCTS is not necessarily human-like, believable or enjoyable. AI Factory Spades, currently the top rated Spades game in the Google Play store, uses a variant of MCTS to control AI allies and opponents. In collaboration with the developers, we showed in a previous study that the playstyle of human players significantly differed from that of the AI players. This paper presents a method for player modeling using gameplay data and neural networks that does not require domain knowledge, and a method of biasing MCTS with such a player model to create Spades playing agents that emulate human play whilst maintaining strong, competitive performance. The methods of player modeling and biasing MCTS presented in this study are applied to the commercial codebase of AI Factory Spades, and are transferable to MCTS implementations for discrete-action games where relevant gameplay data are available.
Hendrik Baier, Adam Sattaur, Edward J. Powley, Sam Devlin, Jeff Rollason, Peter I. Cowling
IEEE Trans. Games4
2019 How the Business Model of Customizable Card Games Influences Player Engagement
abstract
In this paper, we analyze the gameplay data of three popular customizable card games where players build decks prior to gameplay. We analyze the data from a player engagement perspective, how the business model affects players, how players influence the business model and provide strategic insights for players themselves. Sifa et al. found a lack of cross-game analytics, whereas Marchand and Hennig-Thurau identified a lack of understanding of how a game's business model and strategies affect players. We address both issues. The three games have similar business models but differ in one aspect: the distribution model for the cards used in the game. Our longitudinal analysis highlights this variation's impact. A uniform distribution creates a spread of decks with slowly emerging trends while a random distribution creates stripes of deck building activity that switch suddenly each update. Our method is simple, easily understandable, independent of the specific game's structure, and able to compare multiple games. It is applicable to games that release updates and enables comparison across games. Optimizing a game's updates strategy is the key, as it affects player engagement and retention, which directly influence businesses' revenues and profitability in the $95 billion global games market.
Victoria J. Hodge, Feng Li 0021, Nick Sephton, Sam Devlin, Peter I. Cowling, Nikolaos Goumagias, Kieran Purvis, Ignazio Cabras, Kiran Jude Fernandes 0001
IEEE Trans. Games4
2017 Exploration and Skill Acquisition in a Major Online Game
Tom Stafford 0002, Sam Devlin, Rafet Sifa, Anders Drachen
CogSci2
2017 Policy invariance under reward transformations for multi-objective reinforcement learning
Patrick Mannion, Sam Devlin, Karl Mason, Jim Duggan, Enda Howley
Neurocomputing2
2016 Potential-based reward shaping for finite horizon online POMDP planning
Adam Eck, Leen-Kiat Soh, Sam Devlin, Daniel Kudenko
Auton. Agents Multi Agent Syst.3
2015 Expressing Arbitrary Reward Functions as Potential-Based Advice
abstract
Effectively incorporating external advice is an important problem in reinforcement learning, especially as it moves into the real world. Potential-based reward shaping is a way to provide the agent with a specific form of additional reward, with the guarantee of policy invariance. In this work we give a novel way to incorporate an arbitrary reward function with the same guarantee, by implicitly translating it into the specific form of dynamic advice potentials, which are maintained as an auxiliary value function learnt at the same time. We show that advice provided in this way captures the input reward function in expectation, and demonstrate its efficacy empirically.
Anna Harutyunyan, Sam Devlin, Peter Vrancx, Ann Nowé
AAAI2
2015 Type Inference in Flexible Model-Driven Engineering
Athanasios Zolotas, Nicholas Drivalos Matragkas, Sam Devlin, Dimitrios S. Kolovos, Richard F. Paige
ECMFA3
2015 Preface to the special issue: Adaptive Learning Agents Part 3
abstract
1. An adaptive learning agent is capable of adapting its behaviour in order to react to changes in its environment and using previous experience to improve its performance with respect to some eval...
Sam Devlin, Daniel Hennes, Samuel Barrett
Connect. Sci.1
2015 Distributed reinforcement learning for adaptive and robust network intrusion response
abstract
Distributed denial of service (DDoS) attacks constitute a rapidly evolving threat in the current Internet. Multiagent Router Throttling is a novel approach to defend against DDoS attacks where multiple reinforcement learning agents are installed on a set of routers and learn to rate-limit or throttle traffic towards a victim server. The focus of this paper is on online learning and scalability. We propose an approach that incorporates task decomposition, team rewards and a form of reward shaping called difference rewards. One of the novel characteristics of the proposed system is that it provides a decentralised coordinated response to the DDoS problem, thus being resilient to DDoS attacks themselves. The proposed system learns remarkably fast, thus being suitable for online learning. Furthermore, its scalability is successfully demonstrated in experiments involving 1000 learning agents. We compare our approach against a baseline and a popular state-of-the-art throttling technique from the network security literature and show that the proposed approach is more effective, adaptive to sophisticated attack rate dynamics and robust to agent failures.
Kleanthis Malialis, Sam Devlin, Daniel Kudenko
Connect. Sci.2
2015 Player Preference and Style in a Leading Mobile Card Game
abstract
Tuning game difficulty prior to release requires careful consideration. Players can quickly lose interest in a game if it is too hard or too easy. Assessing how players will cope prior to release is often inaccurate. However, modern games can now collect sufficient data to perform large scale analysis post deployment and update the product based on these insights. AI Factory Spades is currently the top rated Spades game in the Google Play store. In collaboration with the developers, we have collected gameplay data from 27 592 games and statistics regarding wins/losses for 99 866 games using Google Analytics. Using the data collected, this study analyses the difficulty and behavior of an Information Set Monte Carlo Tree Search player we developed and deployed in the game previously. The methods of data collection and analysis presented in this study are generally applicable. The same workflow could be used to analyze the difficulty and typical player or opponent behavior in any game. Furthermore, addressing issues of difficulty or nonhuman-like opponents postdeployment can positively affect player retention.
Peter I. Cowling, Sam Devlin, Edward J. Powley, Daniel Whitehouse, Jeff Rollason
IEEE Trans. Comput. Intell. AI Games2
2014 Coordinated Team Learning and Difference Rewards for Distributed Intrusion Response
abstract
Distributed denial of service attacks constitute a rapidly evolving threat in the current Internet. Multiagent Router Throttling is a novel approach to respond to such attacks. We demonstrate that our approach can significantly scale-up using hierarchical communication and coordinated team learning. Furthermore, we incorporate a form of reward shaping called difference rewards and show that the scalability of our system is significantly improved in experiments involving over 100 reinforcement learning agents. We also demonstrate that difference rewards constitute an ideal online learning mechanism for network intrusion response. We compare our proposed approach against a popular state-of-the-art router throttling technique from the network security literature, and we show that our proposed approach significantly outperforms it. We note that our approach can be useful in other related multiagent domains.
Kleanthis Malialis, Sam Devlin, Daniel Kudenko
ECAI2
2014 A Phylogenetic Classification of the Video-Game Industry's Business Model Ecosystem
Nikolaos Goumagias, Ignazio Cabras, Kiran Jude Fernandes 0001, Feng Li 0021, Alberto Nucciarelli, Peter I. Cowling, Sam Devlin, Daniel Kudenko
PRO-VE7
2014 Preface to the special issue: Adaptive Learning Agents, Part 1
abstract
1. An Adaptive Learning Agent is capable of adapting its behaviour in order to react to changes in its environment and using previous experience to improve its performance with respect to some eval...
Sam Devlin, Daniel Hennes, Enda Howley
Connect. Sci.1
2014 Preface to the special issue: Adaptive Learning Agents Part 2
abstract
1. An adaptive learning agent is capable of adapting its behaviour in order to react to changes in its environment and using previous experience to improve its performance with respect to some eval...
Sam Devlin, Daniel Hennes, Enda Howley
Connect. Sci.1