Alessandro Sestini

dblp:280/0287 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-5496-5770ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 12 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Improving Sample Efficiency in Multi-Agent Reinforcement Learning for Simulated Football Games via Exploration
abstract
Multi-agent reinforcement learning has shown promise in learning cooperative behaviors in team-based environments. However, such methods often demand extensive training time, which inhibits their application for game-AI in standard game development. For instance, the state-of-the-art method TiZero takes 40 days to train high-quality policies for a football environment. In this paper, we hypothesize that better exploration mechanisms can improve the sample efficiency of multi-agent methods. Thereby, we propose utilizing a random network distillation bonus within the multi-agent TiZero framework, aiming to promote exploration. Additionally, we introduce architectural modifications to the original algorithm to enhance TiZero’s computational efficiency. We evaluate the sample efficiency of our approach against original TiZero through extensive experiments. Our results show that random network distillation improves the sample efficiency per training phase by 13.3% compared with the original TiZero, enhancing generalization and adaptability to previously difficult scenarios. This highlights the better applicability of our variant in practical game development settings. Lastly, we qualitatively evaluate the gameplay of the produced models against a heuristic AI. We find that random network distillation leads to a higher accuracy in shooting, and it achieves higher behavioral stability as shown by the lower standard deviation achieved in gameplay metrics. The code is available at https://github.com/electronicarts/marling.
Amir Baghi, Jens Sjölund, Joakim Bergdahl, Linus Gisslén, Alessandro Sestini
FDG5
2025 A Call for Deeper Collaboration Between Robotics and Game Development
abstract
While robotics and game development have independently achieved significant progress in creating interactive and intelligent systems, a deeper collaboration between these fields could be mutually beneficial. This paper argues for more collaboration, highlighting current limited interactions and proposing directions for future research. We discuss shared foundations such as Artificial Intelligence, Extended Reality, and the increasing use of common tools and standards. We then propose opportunities where game development methodologies can advance robotics (e.g., gamified data collection and richer simulation environments) and where robotics research can contribute to games (e.g., improved NPC autonomy and embodied intelligence). This cross-disciplinary interaction can accelerate innovation and lead to more intelligent and usercentered technologies in both domains.
Iolanda Leite, William Ahlberg, André Pereira 0001, Alessandro Sestini, Linus Gisslén, Konrad Tollmar
CoG4
2025 Real-Time Diffusion Policies for Games: Enhancing Consistency Policies with Q-Ensembles
abstract
Diffusion models have shown impressive performance in capturing complex and multi-modal action distributions for game agents, but their slow inference speed prevents practical deployment in real-time game environments. While consistency models offer a promising approach for one-step generation, they often suffer from training instability and performance degradation when applied to policy learning. In this paper, we present CPQE (Consistency Policy with Q-Ensembles), which combines consistency models with Q-ensembles to address these challenges. CPQE leverages uncertainty estimation through Q-ensembles to provide more reliable value function approximations, resulting in better training stability and improved performance compared to classic double Q-network methods. Our extensive experiments across multiple game scenarios demonstrate that CPQE achieves inference speeds of up to 60 Hz - a significant improvement over state-of-the-art diffusion policies that operate at only 20 Hz - while maintaining comparable performance to multi-step diffusion approaches. CPQE consistently outperforms state-of-theart consistency model approaches, showing both higher rewards and enhanced training stability throughout the learning process. These results indicate that CPQE offers a practical solution for deploying diffusion-based policies in games and other real-time applications where both multi-modal behavior modeling and rapid inference are critical requirements.
Ruoqi Zhang, Ziwei Luo 0002, Jens Sjölund, Per Mattsson, Linus Gisslén, Alessandro Sestini
CoG6
2024 Improving Generalization in Game Agents with Data Augmentation in Imitation Learning
abstract
Imitation learning is an effective approach for training game-playing agents and, consequently, for efficient game production. However, generalization-the ability to perform well in related but unseen scenarios-is an essential requirement that remains an unsolved challenge for game AI. Generalization is difficult for imitation learning agents because it requires the algorithm to take meaningful actions outside of the training distribution. In this paper we propose a solution to this challenge. Inspired by the success of data augmentation in supervised learning, we augment the training data so the distribution of states and actions in the dataset better represents the real state-action distribution. This study evaluates methods for combining and applying data augmentations to observations, to improve generalization of imitation learning agents. It also provides a performance benchmark of these augmentations across several 3D environments. These results demonstrate that data augmentation is a promising framework for improving generalization in imitation learning agents.
Derek Yadgaroff, Alessandro Sestini, Konrad Tollmar, Ayça Özçelikkale, Linus Gisslén
CEC2
2024 Reinforcement Learning for High-Level Strategic Control in Tower Defense Games
abstract
In strategy games, one of the most important aspects of game design is maintaining a sense of challenge for players. Many mobile titles feature quick gameplay loops that allow players to progress steadily, requiring an abundance of levels and puzzles to prevent them from reaching the end too quickly. As with any content creation, testing and validation are essential to ensure engaging gameplay mechanics, enjoyable game assets, and playable levels. In this paper, we propose an automated approach that can be leveraged for gameplay testing and validation that combines traditional scripted methods with reinforcement learning, reaping the benefits of both approaches while adapting to new situations similarly to how a human player would. We test our solution on a popular tower defense game, Plants vs. Zombies. The results show that combining a learned approach, such as reinforcement learning, with a scripted AI produces a higher-performing and more robust agent than using only heuristic AI, achieving a $57.12 \%$ success rate compared to 47.95% in a set of 40 levels. Moreover, the results demonstrate the difficulty of training a general agent for this type of puzzle-like game.
Joakim Bergdahl, Alessandro Sestini, Linus Gisslén
CoG2
2024 A Benchmark Environment for Offline Reinforcement Learning in Racing Games
abstract
Offline Reinforcement Learning (ORL) is a promising approach to reduce the high sample complexity of traditional Reinforcement Learning (RL) by eliminating the need for continuous environmental interactions. ORL exploits a dataset of precollected transitions and thus expands the range of application of RL to tasks in which the excessive environment queries increase training time and decrease efficiency, such as in modern AAA games. This paper introduces OfflineMania a novel environment for ORL research. It is inspired by the iconic TrackMania series and developed using the Unity 3D game engine. The environment simulates a single-agent racing game in which the objective is to complete the track through optimal navigation. We provide a variety of datasets to assess ORL performance. These datasets, created from policies of varying ability and in different sizes, aim to offer a challenging testbed for algorithm development and evaluation. We further establish a set of baselines for a range of Online RL, ORL, and hybrid Offline to Online RL approaches using our environment.
Girolamo Macaluso, Alessandro Sestini, Andrew D. Bagdanov
CoG2
2024 Leveraging Large Language Models for Efficient Failure Analysis in Game Development
abstract
In games, and more generally in the field of software development, early detection of bugs is vital to maintain a high quality of the final product. Automated tests are a powerful tool that can catch a problem earlier in development by executing periodically. As an example, when new code is submitted to the code base, a new automated test verifies these changes. However, identifying the specific change responsible for a test failure becomes harder when dealing with batches of changes especially in the case of a large-scale project such as a AAA game, where thousands of people contribute to a single code base. This paper proposes a new approach to automatically identify which change in the code caused a test to fail. The method leverages Large Language Models (LLMs) to associate error messages with the corresponding code changes causing the failure. We investigate the effectiveness of our approach with quantitative and qualitative evaluations. Our approach reaches an accuracy of $71 \%$ in our newly created dataset, which comprises issues reported by developers at EA over a period of one year. We further evaluated our model through a user study to assess the utility and usability of the tool from a developer perspective, resulting in a significant reduction in time - up to $60 \%$ - spent investigating issues.
Leonardo Marini, Linus Gisslén, Alessandro Sestini
CoG3
2024 Improving Conditional Level Generation Using Automated Validation in Match-3 Games
abstract
Generative models for level generation have shown great potential in game production. However, they often provide limited control over the generation, and the validity of the generated levels is unreliable. Despite this fact, only a few approaches that learn from existing data provide the users with ways of controlling the generation, simultaneously addressing the generation of unsolvable levels. This article proposes autovalidated level generation, a novel method to improve models that learn from existing level designs using difficulty statistics extracted from gameplay. In particular, we use a conditional variational autoencoder to generate layouts for match-3 levels, conditioning the model on precollected statistics, such as game mechanics like difficulty, and relevant visual features, such as size and symmetry. Our method is general enough that multiple approaches could potentially be used to generate these statistics. We quantitatively evaluate our approach by comparing it to an ablated model without difficulty conditioning. In addition, we analyze both quantitatively and qualitatively whether the style of the dataset is preserved in the generated levels. Our approach generates more valid levels than the same method without difficulty conditioning.
Monica Villanueva Aylagas, Joakim Bergdahl, Jonas Gillberg, Alessandro Sestini, Theodor Tolstoy, Linus Gisslén
IEEE Trans. Games4
2024 Automated Gameplay Testing and Validation With Curiosity-Conditioned Proximal Trajectories
abstract
This article proposes a novel deep reinforcement learning algorithm to perform automated analysis and detection of gameplay issues in complex 3-D navigation environments. The curiosity-conditioned proximal trajectories (CCPT) method combines curiosity and imitation learning to train agents that methodically explore in the proximity of known trajectories derived from expert demonstrations. We show how our new algorithm can explore complex environments, discovering gameplay issues, and design oversights in the process, and recognize and highlight them directly to game designers. We also propose a visual analytics interface to aid interpretation of results from the method. This interface transforms information from complex models into interpretable and interactive visual forms. We further demonstrate the effectiveness of the algorithm in a novel 3-D navigation environment, which reflects the complexity of modern video games. Our results show a higher level of coverage and bug discovery than baseline methods, demonstrating that our method can be a useful tool for game designers to automatically identify design issues. Moreover, our experiments show that the visual explanations provided by the analytics interface result in a significant increase in user trust and acceptance of automated playtesting and increased confidence in the use of machine learning techniques for video game development.
Alessandro Sestini, Linus Gisslén, Joakim Bergdahl, Konrad Tollmar, Andrew D. Bagdanov
IEEE Trans. Games1
2023 Generating Personas for Games with Multimodal Adversarial Imitation Learning
abstract
Reinforcement learning has been widely successful in producing agents capable of playing games at a human level. However, this requires complex reward engineering, and the agent’s resulting policy is often unpredictable. Going beyond reinforcement learning is necessary to model a wide range of human playstyles, which can be difficult to represent with a reward function. This paper presents a novel imitation learning approach to generate multiple persona policies for playtesting. Multimodal Generative Adversarial Imitation Learning (Multi-GAIL) uses an auxiliary input parameter to learn distinct personas using a single-agent model. MultiGAIL is based on generative adversarial imitation learning and uses multiple dis-criminators as reward models, inferring the environment reward by comparing the agent and distinct expert policies. The reward from each discriminator is weighted according to the auxiliary input. Our experimental analysis demonstrates the effectiveness of our technique in two environments with continuous and discrete action spaces.
William Ahlberg, Alessandro Sestini, Konrad Tollmar, Linus Gisslén
CoG2
2023 Technical Challenges of Deploying Reinforcement Learning Agents for Game Testing in AAA Games
abstract
Going from research to production, especially for large and complex software systems, is fundamentally a hard problem. In large-scale game production, one of the main reasons is that the development environment can be very different from the final product. In this technical paper we describe an effort to add an experimental reinforcement learning system to an existing automated game testing solution based on scripted bots in order to increase its capacity. We report on how this reinforcement learning system was integrated with the aim to increase test coverage similar to [1] in a set of AAA games including Battlefield 2042 and Dead Space (2023). The aim of this technical paper is to show a use-case of leveraging reinforcement learning in game production and cover some of the largest time sinks anyone who wants to make the same journey for their game may encounter. Furthermore, to help the game industry to adopt this technology faster, we propose a few research directions that we believe will be valuable and necessary for making machine learning, and especially reinforcement learning, an effective tool in game production.
Jonas Gillberg, Joakim Bergdahl, Alessandro Sestini, Andy Eakins, Linus Gisslén
CoG3
2023 Efficient Ground Vehicle Path Following in Game AI
abstract
This short paper presents an efficient path following solution for ground vehicles tailored to game AI. Our focus is on adapting established techniques to design simple solutions with parameters that are easily tunable for an efficient benchmark path follower. Our solution pays particular attention to computing a target speed which uses quadratic Bézier curves to estimate the path curvature. The performance of the proposed path follower is evaluated through a variety of test scenarios in a first-person shooter game, demonstrating its effectiveness and robustness in handling different types of paths and vehicles. We achieved a 70% decrease in the total number of stuck events compared to an existing path following solution.
Rodrigue de Schaetzen, Alessandro Sestini
CoG2
2023 Towards Informed Design and Validation Assistance in Computer Games Using Imitation Learning
abstract
In games, as in many other domains, design validation and testing is a significant challenge as systems are growing in size and manual testing is becoming infeasible. In this position paper we outline an approach to automated game validation based on an imitation learning technique, and provide an analysis of the potential benefits to automated game testing. The method leverages a data-driven technique, which requires little effort and time and no knowledge of machine learning or programming, that designers can use to efficiently train game testing agents. We evaluate the validity of our claim by conducting a user study with industry experts. The survey results presented in this paper demonstrate the potential of a data-driven approach to reduce effort and enhance the quality of game testing. Moreover, the survey reveals several open challenges. To this end, we analyze the identified challenges and provide a basis for further research and discussion, as well as to help guide the development of imitation learning for game testing.
Alessandro Sestini, Joakim Bergdahl, Konrad Tollmar, Andrew D. Bagdanov, Linus Gisslén
CoG1
2021 Demonstration-Efficient Inverse Reinforcement Learning in Procedurally Generated Environments
abstract
Deep Reinforcement Learning achieves very good results in domains where reward functions can be manually engineered. At the same time, there is growing interest within the community in using games based on Procedurally Content Generation (PCG) as benchmark environments since this type of environment is perfect for studying overfitting and generalization of agents under domain shift. Inverse Reinforcement Learning (IRL) can instead extrapolate reward functions from expert demonstrations, with good results even on high-dimensional problems, however there are no examples of applying these techniques to procedurally-generated environments. This is mostly due to the number of demonstrations needed to find a good reward model. We propose a technique based on Adversarial Inverse Reinforcement Learning which can significantly decrease the need for expert demonstrations in PCG games. Through the use of an environment with a limited set of initial seed levels, plus some modifications to stabilize training, we show that our approach, DE-AIRL, is demonstration-efficient and still able to extrapolate reward functions which generalize to the fully procedural domain. We demonstrate the effectiveness of our technique on two procedural environments, MiniGrid and DeepCrawl, for a variety of tasks.
Alessandro Sestini, Andrew D. Bagdanov
CoG1
2021 Policy Fusion for Adaptive and Customizable Reinforcement Learning Agents
abstract
In this article we study the problem of training intelligent agents using Reinforcement Learning for the purpose of game development. Unlike systems built to replace human players and to achieve super-human performance, our agents aim to produce meaningful interactions with the player, and at the same time demonstrate behavioral traits as desired by game designers. We show how to combine distinct behavioral policies to obtain a meaningful “fusion” policy which comprises all these behaviors. To this end, we propose four different policy fusion methods for combining pre-trained policies. We further demonstrate how these methods can be used in combination with Inverse Reinforcement Learning in order to create intelligent agents with specific behavioral styles as chosen by game designers, without having to define many and possibly poorly-designed reward functions. Experiments on two different environments indicate that entropy-weighted policy fusion significantly outperforms all others. We provide several practical examples and use-cases for how these methods are indeed useful for video game production and designers.
Alessandro Sestini, Andrew D. Bagdanov
CoG1