VLDB 2026 Research / reviewers in the wild / expert
Matthew Stephenson 0001
dblp:190/8740-1
· DBLP profile ↗
29ranked-venue papers
9as first author
14since 2021 · last 2025
0000-0002-3867-5842ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 16 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | General Game Playing Hyper-Agents for LudiiabstractThis paper explores the viability and effectiveness of hyper-agent approaches for the Ludii general game system. These hyper-agents utilise trained machine learning models to predict the optimal sub-agent and heuristics for previously unseen board games, based on automatically detectable game parameters (ludemes and concepts). Several hyper-agents based on portfolio and ensemble design approaches were implemented within the Ludii system. Each hyper-agent was trained on 460 games with known sub-agent and heuristic performances, with evaluations being performed on a representative set of 50 new games. Our best performing hyper-agent approach demonstrated a statistically significant win-rate improvement over all of the individual sub-agents utilised in its training corpus. Nicholas Thompson, Matthew Stephenson 0001 |
CoG | 2 |
| 2025 | NovPhy: A Physical Reasoning Benchmark for Open-World AI Systems Author Links Open Overlay Panel (Abstract Reprint)abstractDue to the emergence of AI systems that interact with the physical environment, there is an increased interest in incorporating physical reasoning capabilities into those AI systems. But is it enough to only have physical reasoning capabilities to operate in a real physical environment? In the real world, we constantly face novel situations we have not encountered before. As humans, we are competent at successfully adapting to those situations. Similarly, an agent needs to have the ability to function under the impact of novelties in order to properly operate in an open-world physical environment. To facilitate the development of such AI systems, we propose a new benchmark, NovPhy, that requires an agent to reason about physical scenarios in the presence of novelties and take actions accordingly. The benchmark consists of tasks that require agents to detect and adapt to novelties in physical scenarios. To create tasks in the benchmark, we develop eight novelties representing a diverse novelty space and apply them to five commonly encountered scenarios in a physical environment, related to applying forces and motions such as rolling, falling, and sliding of objects. According to our benchmark design, we evaluate two capabilities of an agent: the performance on a novelty when it is applied to different physical scenarios and the performance on a physical scenario when different novelties are applied to it. We conduct a thorough evaluation with human players, learning agents, and heuristic agents. Our evaluation shows that humans' performance is far beyond the agents' performance. Some agents, even with good normal task performance, perform significantly worse when there is a novelty, and the agents that can adapt to novelties typically adapt slower than humans. We promote the development of intelligent agents capable of performing at the human level or above when operating in open-world physical environments. Benchmark website: https://github.com/phy-q/novphy Vimukthini Pinto, Chathura Nagoda Gamage, Cheng Xue 0008, Peng Zhang 0021, Ekaterina Nikonova, Matthew Stephenson 0001, Jochen Renz |
IJCAI | 6 |
| 2025 | Physics-Based Novel Task Generation Through Disrupting and Constructing Causal InteractionsabstractIn response to the growing demand for AI systems that can operate the physical world, there has been an increasing interest in enhancing their physical reasoning capabilities. Equally crucial is the ability to handle unseen novel situations, as such situations frequently arise in real-world environments. To facilitate the development of AI systems with those abilities, researchers have developed testbeds with specialized tasks to evaluate agents' adaptation to novelty in physical environments. In this paper, we propose a method for generating physics-based tasks with incorporated novelties to assess agents' novelty adaptation capabilities. The tasks are defined as causal sequences of physical interactions between objects, and novelties are strategically introduced to disrupt existing causal relationships and construct new ones. This approach ensures that agents must adapt to the effects of novelties to perform those tasks, enabling confident measurement of their novelty adaptation capabilities using task performance. Moreover, our methodology eliminates the need for manual task creation, unlike existing novelty-centric testbeds. The proposed method is demonstrated and evaluated using 12 physical scenarios in the Angry Birds domain. The evaluated metrics include generation time, physical stability, intended solvability, intended unsolvability, and accidental solvability of the tasks, and they yielded favourable results compared to the literature. Chathura Nagoda Gamage, Matthew Stephenson 0001, Jochen Renz |
IEEE Trans. Games | 2 |
| 2024 | Prompt Engineering ChatGPT for CodenamesabstractThe word association game Codenames challenges the AI community with its requirements for multimodal language understanding, theory of mind, and epistemic reasoning. Previous attempts to develop AI agents for the game have focused on word embedding techniques, which while good with other models using the same technique, can sometimes suffer from brittle performance when paired with other models. Recently, Large Language Models (LLMs) have demonstrated enhanced capabilities, excelling in complex cognitive tasks, including symbolic and common sense reasoning. In this paper, we compare a range of recent prompt engineering techniques for GPT-based Codenames agents. While there was no significant game score improvement over the baseline agent, we did observe qualitative changes in agents’ strategies suggesting that further refinement has potential for score improvement. We also propose a revised Codenames AI competition specifically focusing on the use of LLM agents. Matthew Sidji, Matthew Stephenson 0001 |
CoG | 2 |
| 2024 | The Ludii Game Description Language is UniversalabstractThere are several different game description languages (GDLs), each intended to allow wide ranges of arbitrary games (i.e., general games) to be described in a single higher-level language than general-purpose programming languages. Games described in such formats can subsequently be presented as challenges for automated general game playing agents, which are expected to be capable of playing any arbitrary game described in such a language without prior knowledge about the games to be played. The language used by the Ludii general game system was previously shown to be capable of representing equivalent games for any arbitrary, finite, deterministic, fully observable extensive-form game. In this paper, we prove its universality by extending this to include finite non-deterministic and imperfect-information games. Dennis J. N. J. Soemers, Éric Piette, Matthew Stephenson 0001, Cameron Browne |
CoG | 3 |
| 2024 | GAVEL: Generating Games via Evolution and Language ModelsabstractAutomatically generating novel and interesting games is a complex task. Challenges include representing game rules in a computationally workable form, searching through the large space of potential games under most such representations, and accurately evaluating the originality and quality of previously unseen games. Prior work in automated game generation has largely focused on relatively restricted rule representations and relied on domain-specific heuristics. In this work, we explore the generation of novel games in the comparatively expansive Ludii game description language, which encodes the rules of over 1000 board games in a variety of styles and modes of play. We draw inspiration from recent advances in large language models and evolutionary computation in order to train a model that intelligently mutates and recombines games and mechanics expressed as code. We demonstrate both quantitatively and qualitatively that our approach is capable of generating new and interesting games, including in regions of the potential rules space not covered by existing games in the Ludii dataset. Graham Todd, Alexander Padula, Matthew Stephenson 0001, Éric Piette, Dennis J. N. J. Soemers, Julian Togelius |
NeurIPS | 3 |
| 2024 | NovPhy: A physical reasoning benchmark for open-world AI systemsabstractDue to the emergence of AI systems that interact with the physical environment, there is an increased interest in incorporating physical reasoning capabilities into those AI systems. But is it enough to only have physical reasoning capabilities to operate in a real physical environment? In the real world, we constantly face novel situations we have not encountered before. As humans, we are competent at successfully adapting to those situations. Similarly, an agent needs to have the ability to function under the impact of novelties in order to properly operate in an open-world physical environment. To facilitate the development of such AI systems, we propose a new benchmark, NovPhy, that requires an agent to reason about physical scenarios in the presence of novelties and take actions accordingly. The benchmark consists of tasks that require agents to detect and adapt to novelties in physical scenarios. To create tasks in the benchmark, we develop eight novelties representing a diverse novelty space and apply them to five commonly encountered scenarios in a physical environment, related to applying forces and motions such as rolling, falling, and sliding of objects. According to our benchmark design, we evaluate two capabilities of an agent: the performance on a novelty when it is applied to different physical scenarios and the performance on a physical scenario when different novelties are applied to it. We conduct a thorough evaluation with human players, learning agents, and heuristic agents. Our evaluation shows that humans' performance is far beyond the agents' performance. Some agents, even with good normal task performance, perform significantly worse when there is a novelty, and the agents that can adapt to novelties typically adapt slower than humans. We promote the development of intelligent agents capable of performing at the human level or above when operating in open-world physical environments. Benchmark website: https://github.com/phy-q/novphy. Vimukthini Pinto, Chathura Nagoda Gamage, Cheng Xue 0008, Peng Zhang 0021, Ekaterina Nikonova, Matthew Stephenson 0001, Jochen Renz |
Artif. Intell. | 6 |
| 2023 | Spatial state-action features for general gamesabstractIn many board games and other abstract games, patterns have been used as features that can guide automated game-playing agents. Such patterns or features often represent particular configurations of pieces, empty positions, etc., which may be relevant for a game's strategies. Their use has been particularly prevalent in the game of Go, but also many other games used as benchmarks for AI research. In this paper, we formulate a design and efficient implementation of spatial state-action features for general games. These are patterns that can be trained to incentivise or disincentivise actions based on whether or not they match variables of the state in a local area around action variables. We provide extensive details on several design and implementation choices, with a primary focus on achieving a high degree of generality to support a wide variety of different games using different board geometries or other graphs. Secondly, we propose an efficient approach for evaluating active features for any given set of features. In this approach, we take inspiration from heuristics used in problems such as SAT to optimise the order in which parts of patterns are matched and prune unnecessary evaluations. This approach is defined for a highly general and abstract description of the problem—phrased as optimising the order in which propositions of formulas in disjunctive normal form are evaluated—and may therefore also be of interest to other types of problems than board games. An empirical evaluation on 33 distinct games in the Ludii general game system demonstrates the efficiency of this approach in comparison to a naive baseline, as well as a baseline based on prefix trees, and demonstrates that the additional efficiency significantly improves the playing strength of agents using the features to guide search. Dennis J. N. J. Soemers, Éric Piette, Matthew Stephenson 0001, Cameron Browne |
Artif. Intell. | 3 |
| 2023 | Extracting tactics learned from self-play in general gamesabstractLocal, spatial state-action features can be used to effectively train linear policies from self-play in a wide variety of board games. Such policies can play games directly, or be used to bias tree search agents. However, the resulting feature sets can be large, with a significant amount of overlap and redundancies between features. This is a problem for two reasons. Firstly, large feature sets can be computationally expensive, which reduces the playing strength of agents based on them. Secondly, redundancies and correlations between features impair the ability for humans to analyse, interpret, or understand tactics learned by the policies. We look towards decision trees for their ability to perform feature selection, and serve as interpretable models. Previous work on distilling policies into decision trees uses states as inputs, and distributions over the complete action space as outputs. In contrast, we propose and evaluate a variety of decision tree types, which take state-action pairs as inputs, and provide various different types of outputs on a per-action basis. An empirical evaluation over 43 different board games is presented, and two of those games are used as case studies where we attempt to interpret the discovered features. Dennis J. N. J. Soemers, Spyridon Samothrakis, Éric Piette, Matthew Stephenson 0001 |
Inf. Sci. | 4 |
| 2021 | Deceptive Level Generation for Angry BirdsabstractThe Angry Birds AI competition has been held over many years to encourage the development of AI agents that can play Angry Birds game levels better than human players. Many different agents with various approaches have been employed over the competition's lifetime to solve this task. Even though the performance of these agents has increased significantly over the past few years, they still show major drawbacks in playing deceptive levels. This is because most of the current agents try to identify the best next shot rather than planning an effective sequence of shots. In order to encourage advancements in such agents, we present an automated methodology to generate deceptive game levels for Angry Birds. Even though there are many existing content generators for Angry Birds, they do not focus on generating deceptive levels. In this paper, we propose a procedure to generate deceptive levels for six deception categories that can fool the state-of-the-art Angry Birds playing AI agents. Our results show that generated deceptive levels exhibit similar characteristics of human-created deceptive levels. Additionally, we define metrics to measure the stability, solvability, and degree of deception of the generated levels. Chathura Nagoda Gamage, Vimukthini Pinto, Jochen Renz, Matthew Stephenson 0001 |
CoG | 4 |
| 2021 | Novelty Generation Framework for AI Agents in Angry Birds Style Physics GamesabstractHandling novel situations is a critical capability of Artificial Intelligence (AI) agents when working in open-world physical environments. To develop and evaluate these agents, we need realistic and meaningful novelties, that is, novelties that are detectable and learnable. However, there is a lack of research in the area of creating novelties for AI agents in physical environments. Physics-based video games are popular among AI researchers due to the ability to create realistic and controllable physical environments. In this paper, we present a systematic novelty generation framework for physics-based video games. This framework allows the user to define a specific objective when generating novel content that ensures detectability. We instantiate the proposed framework for the video game Angry Birds and conduct experiments to show that the generated novel content is consistent with the user-defined objectives. Furthermore, we use a reinforcement learning agent to experiment with the learnability of the generated novel content. Chathura Nagoda Gamage, Vimukthini Pinto, Cheng Xue 0008, Matthew Stephenson 0001, Peng Zhang 0021, Jochen Renz |
CoG | 4 |
| 2021 | General Board Game ConceptsabstractMany games often share common ideas or aspects between them, such as their rules, controls, or playing area. However, in the context of General Game Playing (GGP) for board games, this area remains under-explored. We propose to formalise the notion of “game concept”, inspired by terms generally used by game players and designers. Through the Ludii General Game System, we describe concepts for several levels of abstraction, such as the game itself, the moves played, or the states reached. This new GGP feature associated with the ludeme representation of games opens many new lines of research. The creation of a hyper-agent selector, the transfer of AI learning between games, or explaining AI techniques using game terms, can all be facilitated by the use of game concepts. Other applications which can benefit from game concepts are also discussed, such as the generation of plausible reconstructed rules for incomplete ancient games, or the implementation of a board game recommender system. Éric Piette, Matthew Stephenson 0001, Dennis J. N. J. Soemers, Cameron Browne |
CoG | 2 |
| 2021 | General Game Heuristic Prediction Based on Ludeme DescriptionsabstractThis paper investigates the performance of different general-game-playing heuristics for games in the Ludii general game system. Based on these results, we train several regression learning models to predict the performance of these heuristics based on each game's description file. We also provide a condensed analysis of the games available in Ludii, and the different ludemes that define them. Matthew Stephenson 0001, Dennis J. N. J. Soemers, Éric Piette, Cameron Browne |
CoG | 1 |
| 2021 | Generating Stable Building Block Structures From SketchesabstractThis paper presents a structure generation algorithm, which converts rough human drawings into stable structures comprising rectangular blocks, suitable for physics-based 2-D environments. Generating viable structures for a physics-based environment imposes many additional requirements above those of most traditional sketch-based domains. Our method is sophisticated enough to deal with these requirements, while still ensuring that the generated structure accurately represents the original sketch. We describe and implement a framework for this process, allowing inexperienced users to create complex structures with ease. Multiple structure possibilities are identified for a single drawing and are then compared based on their similarity to the original sketch using a heuristic value. We evaluate our approach by investigating its ability to replicate structures for the video game Angry Birds, based on human drawn sketches of the original levels. Matthew Stephenson 0001, Jochen Renz, Xiaoyu Ge, Peng Zhang 0021 |
IEEE Trans. Games | 1 |
| 2020 | A Continuous Information Gain Measure to Find the Most Discriminatory Problems for AI BenchmarkingabstractThis paper introduces an information-theoretic method for selecting a subset of problems which gives the most information about a group of problem-solving algorithms. This method was tested on the games in the General Video Game AI (GVGAI) framework, allowing us to identify a smaller set of games that still gives a large amount of information about the abilities of different game-playing agents. This approach can be used to make agent testing more efficient. We can achieve almost as good discriminatory accuracy when testing on only a handful of games as when testing on more than a hundred games, something which is often computationally infeasible. Furthermore, this method can be extended to study the dimensions of the effective variance in game design between these games, allowing us to identify which games differentiate between agents in the most complementary ways. Matthew Stephenson 0001, Damien Anderson, Ahmed Khalifa 0001, John Levine, Jochen Renz, Julian Togelius, Christoph Salge |
CEC | 1 |
| 2020 | Manipulating the Distributions of Experience used for Self-Play Learning in Expert IterationabstractExpert Iteration (ExIt) is an effective framework for learning game-playing policies from self-play. ExIt involves training a policy to mimic the search behaviour of a tree search algorithm -- such as Monte-Carlo tree search -- and using the trained policy to guide it. The policy and the tree search can then iteratively improve each other, through experience gathered in self-play between instances of the guided tree search algorithm. This paper outlines three different approaches for manipulating the distribution of data collected from self-play, and the procedure that samples batches for learning updates from the collected data. Firstly, samples in batches are weighted based on the durations of the episodes in which they were originally experienced. Secondly, Prioritized Experience Replay is applied within the ExIt framework, to prioritise sampling experience from which we expect to obtain valuable training signals. Thirdly, a trained exploratory policy is used to diversify the trajectories experienced in self-play. This paper summarises the effects of these manipulations on training performance evaluated in fourteen different board games. We find major improvements in early training performance in some games, and minor improvements averaged over fourteen games. Dennis J. N. J. Soemers, Éric Piette, Matthew Stephenson 0001, Cameron Browne |
CoG | 3 |
| 2020 | Ludii - The Ludemic General Game SystemabstractAccepted at ECAI 2020 Éric Piette, Dennis J. N. J. Soemers, Matthew Stephenson 0001, Chiara F. Sironi, Mark H. M. Winands, Cameron Browne |
ECAI | 3 |
| 2020 | The Computational Complexity of Angry Birds (Extended Abstract)abstractIn this paper we present several proofs for the computational complexity of the physics-based video game Angry Birds. We are able to demonstrate that solving levels for different versions of Angry Birds is either NP-hard, PSPACE-hard, PSPACE-complete or EXPTIME-hard, depending on the maximum number of birds available and whether the game engine is deterministic or stochastic. We believe that this is the first time that a single-player video game has been proven EXPTIME-hard. Matthew Stephenson 0001, Jochen Renz, Xiaoyu Ge |
IJCAI | 1 |
| 2020 | The computational complexity of Angry Birds
Matthew Stephenson 0001, Jochen Renz, Xiaoyu Ge |
Artif. Intell. | 1 |
| 2019 | "Did You Hear That?" Learning to Play Video Games from Audio CuesabstractGame-playing AI research has focused for a long time on learning to play video games from visual input or symbolic information. However, humans benefit from a wider array of sensors which we utilise in order to navigate the world around us. In particular, sounds and music are key to how many of us perceive the world and influence the decisions we make. In this paper, we present initial experiments on game-playing agents learning to play video games solely from audio cues. We expand the Video Game Description Language to allow for audio specification, and the General Video Game AI framework to provide new audio games and an API for learning agents to make use of audio observations. We analyse the games and the audio game design process, include initial results with simple Q-Learning agents, and encourage further research in this area. Raluca D. Gaina, Matthew Stephenson 0001 |
CoG | 2 |
| 2019 | Using Restart Heuristics to Improve Agent Performance in Angry BirdsabstractOver the past few years the Angry Birds AI competition has been held in an attempt to develop intelligent agents that can successfully and efficiently solve levels for the video game Angry Birds. Many different agents and strategies have been developed to solve the complex and challenging physical reasoning problems associated with such a game. However none of these agents attempt one of the key strategies which humans employ to solve Angry Birds levels, which is restarting levels. Restarting is important in Angry Birds because sometimes the level is no longer solvable or some given shot made has little to no benefit towards the ultimate goal of the game. This paper proposes a framework and experimental evaluation for when to restart levels in Angry Birds. We demonstrate that restarting is a viable strategy to improve agent performance in many cases. Tommy Liu, Jochen Renz, Peng Zhang 0021, Matthew Stephenson 0001 |
CoG | 4 |
| 2019 | Ludii and XCSP: Playing and Solving Logic PuzzlesabstractMany of the famous single-player games, commonly called puzzles, can be shown to be NP-Complete. Indeed, this class of complexity contains hundreds of puzzles, since people particularly appreciate completing an intractable puzzle, such as Sudoku, but also enjoy the ability to check their solution easily once it's done. For this reason, using constraint programming is naturally suited to solve them. In this paper, we focus on logic puzzles described in the Ludii general game system and we propose using the XCSP formalism in order to solve them with any CSP solver. Cédric Piette, Éric Piette, Matthew Stephenson 0001, Dennis J. N. J. Soemers, Cameron Browne |
CoG | 3 |
| 2019 | An Empirical Evaluation of Two General Game Systems: Ludii and RBGabstractAlthough General Game Playing (GGP) systems can facilitate useful research in Artificial Intelligence (AI) for game-playing, they are often computationally inefficient and somewhat specialised to a specific class of games. However, since the start of this year, two General Game Systems have emerged that provide efficient alternatives to the academic state of the art - the Game Description Language (GDL). In order of publication, these are the Regular Boardgames language (RBG), and the Ludii system. This paper offers an experimental evaluation of Ludii. Here, we focus mainly on a comparison between the two new systems in terms of two key properties for any GGP system: simplicity/clarity (e.g. human-readability), and efficiency. Éric Piette, Matthew Stephenson 0001, Dennis J. N. J. Soemers, Cameron Browne |
CoG | 2 |
| 2019 | Learning Policies from Self-Play with Policy Gradients and MCTS Value EstimatesabstractIn recent years, state-of-the-art game-playing agents often involve policies that are trained in self-playing processes where Monte Carlo tree search (MCTS) algorithms and trained policies iteratively improve each other. The strongest results have been obtained when policies are trained to mimic the search behaviour of MCTS by minimising a cross-entropy loss. Because MCTS, by design, includes an element of exploration, policies trained in this manner are also likely to exhibit a similar extent of exploration. In this paper, we are interested in learning policies for a project with future goals including the extraction of interpretable strategies, rather than state-of-the-art game-playing performance. For these goals, we argue that such an extent of exploration is undesirable, and we propose a novel objective function for training policies that are not exploratory. We derive a policy gradient expression for maximising this objective function, which can be estimated using MCTS value estimates, rather than MCTS visit counts. We empirically evaluate various properties of resulting policies, in a variety of board games. Dennis J. N. J. Soemers, Éric Piette, Matthew Stephenson 0001, Cameron Browne |
CoG | 3 |
| 2019 | Ludii as a Competition PlatformabstractLudii is a general game system being developed as part of the ERC-funded Digital Ludeme Project (DLP). While its primary aim is to model, play, and analyse the full range of traditional strategy games, Ludii also has the potential to support a wide range of AI research topics and competitions. This paper describes some of the future competitions and challenges that we intend to run using the Ludii system, highlighting some of its most important aspects that can potentially lead to many algorithm improvements and new avenues of research. We compare and contrast our proposed competition motivations, goals and frameworks against those of existing general game playing competitions, addressing the strengths and weaknesses of each platform. Matthew Stephenson 0001, Éric Piette, Dennis J. N. J. Soemers, Cameron Browne |
CoG | 1 |
| 2019 | An Overview of the Ludii General Game SystemabstractThe Digital Ludeme Project (DLP) aims to reconstruct and analyse over 1000 traditional strategy games using modern techniques. One of the key aspects of this project is the development of Ludii, a general game system that will be able to model and play the complete range of games required by this project. Such an undertaking will create a wide range of possibilities for new AI challenges. In this paper we describe many of the features of Ludii that can be used. This includes designing and modifying games using the Ludii game description language, creating agents capable of playing these games, and several advantages the system has over prior general game software. Matthew Stephenson 0001, Éric Piette, Dennis J. N. J. Soemers, Cameron Browne |
CoG | 1 |
| 2019 | The 2017 AIBIRDS Level Generation CompetitionabstractThis paper presents an overview of the second AIBIRDS level generation competition, held jointly at the 2017 IEEE Conference on Computational Intelligence and Games and the 26th International Joint Conference on Artificial Intelligence. This competition tasked entrants with developing a level generator for the physics-based puzzle game Angry Birds. Submitted generators were required to deal with many physical reasoning constraints caused by the realistic nature of the game's environment, in addition to ensuring that the created levels were fun, challenging, and solvable. This year's competition was a significant improvement over the previous year, with a greater number of participants and more advanced generators. In this paper, we describe the framework, rules, submitted generators, and results for this competition. We also provide some background information on related research and other video game AI competitions and discuss what can be learned from this year's competition. There are several game and real-world applications for this type of research, and we provide some examples of the types of levels we would like future competition entries to generate. Matthew Stephenson 0001, Jochen Renz, Xiaoyu Ge, Lucas Ferreira, Julian Togelius, Peng Zhang 0021 |
IEEE Trans. Games | 1 |
| 2018 | Deceptive Games
Damien Anderson, Matthew Stephenson 0001, Julian Togelius, Christoph Salge, John Levine, Jochen Renz |
EvoApplications | 2 |
| 2018 | Deceptive angry birds: towards smarter game-playing agentsabstractOver the past few years the Angry Birds AI competition has been held in an attempt to develop intelligent agents that can successfully and efficiently solve levels for the video game Angry Birds. Many different agents and strategies have been proposed to solve the complex and challenging physical reasoning problems associated with such a game. The performance of these agents has increased significantly over the competition's lifetime thanks to the different approaches and improved techniques employed. However, there still exist key flaws within the designs of these agents that can often lead them to make illogical or very poor choices. Most of the current approaches try to identify the best or a good next shot, but do not attempt to plan an effective sequence of shots. While this might be due to the difficulty in predicting the exact outcome of a shot, this capability is precisely what is needed to succeed, both in games like Angry Birds, but also in the real world where physical reasoning capabilities are essential. In order to encourage development of such techniques, we can create levels where selecting a seemingly good next shot will lead to a worse outcome. In this paper we present several categories of deception to fool the current state-of-the-art agents. By evaluating the performance of the most recent Angry Birds agents on specific level examples that contain these deceptive elements, we can show how certain AI techniques can be tricked or exploited. We also propose some ways that future agents could help deal with these deceptive levels to increase their overall performance and generality. Matthew Stephenson 0001, Jochen Renz |
FDG | 1 |