EDBT 2026 Demo / reviewers in the wild / expert
Patrick Mannion
dblp:167/7672
· DBLP profile ↗
20ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0002-7951-878XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert GuidanceabstractAdversarial Inverse Reinforcement Learning (AIRL) addresses sparse rewards by inferring dense reward functions from expert demonstrations, but its performance in complex, imperfect-information settings is underexplored.We evaluate AIRL in Heads-Up Limit Hold'em (HULHE) poker and observe that it faces challenges producing sufficiently informative rewards.To address this, we introduce Hybrid-AIRL (H-AIRL), which improves reward inference and policy learning using a partially supervised loss from expert data and stochastic regularization.Experiments on Gymnasium benchmarks and HULHE poker show that H-AIRL improves sample efficiency and training stability, highlighting the value of supervised signals in inverse RL. Bram Silue, Santiago Amaya-Corredor, Patrick Mannion, Lander Willem, Pieter Libin |
ESANN | 3 |
| 2026 | MOMA-AC: A preference-driven actor-critic framework for continuous multi-objective multi-agent reinforcement learning
Adam Callaghan, Karl Mason, Patrick Mannion |
Neurocomputing | 3 |
| 2025 | Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning
Willem Röpke, Mathieu Reymond, Patrick Mannion, Diederik M. Roijers, Ann Nowé, Roxana Radulescu |
AAMAS | 3 |
| 2025 | Extending Evolution-Guided Policy Gradient Learning into the multi-objective domainabstractMulti-Objective Reinforcement Learning (MORL) poses significant challenges, primarily due to the necessity of balancing conflicting objectives—a limitation that traditional single-objective approaches fail to address. This paper introduces Multi-Objective Evolutionary Reinforcement Learning (MO-ERL), the first adaptation of Evolutionary Reinforcement Learning (ERL) specifically designed to address the complexities of the multi-objective domain effectively. MO-ERL integrates policy gradient-based reinforcement learning (RL), which optimizes expected utility, with evolutionary algorithms (EAs) that maintain diversity across the Pareto front. This combination leverages RL’s strength in exploitation and EAs’ proficiency in exploration, enabling MO-ERL to effectively navigate the trade-offs inherent in multi-objective optimization problems. Evaluation on multi-objective continuous control tasks using the MuJoCo physics engine demonstrates that MO-ERL outperforms state-of-the-art baselines, achieving up to 62.71% higher hypervolume and 196.28% greater expected utility. These results validate MO-ERL’s ability to balance solution diversity and optimality, setting a new benchmark for solving MORL tasks. • First extension of Evolutionary Reinforcement Learning to multi-objective domain. • Outperforms state-of-the-art (CAPQL, PCN) on multi-objective continuous control. • Up to 62.7% higher hypervolume, 196.3% utility on high-dimensional control tasks. • MO-GA hypervolume approximation reduces computational cost, maintaining performance. • Frequent MORL injections into MO-GA yield smoother learning and faster convergence. Adam Callaghan, Karl Mason, Patrick Mannion |
Neurocomputing | 3 |
| 2025 | Expected scalarised returns dominance: a new solution concept for multi-objective decision makingabstractAbstract In many real-world scenarios, the utility of a user is derived from a single execution of a policy. In this case, to apply multi-objective reinforcement learning, the expected utility of the returns must be optimised. Various scenarios exist where a user’s preferences over objectives (also known as the utility function) are unknown or difficult to specify. In such scenarios, a set of optimal policies must be learned. However, settings where the expected utility must be maximised have been largely overlooked by the multi-objective reinforcement learning community and, as a consequence, a set of optimal solutions has yet to be defined. In this work, we propose first-order stochastic dominance as a criterion to build solution sets to maximise expected utility. We also define a new dominance criterion, known as expected scalarised returns (ESR) dominance, that extends first-order stochastic dominance to allow a set of optimal policies to be learned in practice. Additionally, we define a new solution concept called the ESR set, which is a set of policies that are ESR dominant. Finally, we present a new multi-objective tabular distributional reinforcement learning (MOTDRL) algorithm to learn the ESR set in multi-objective multi-armed bandit settings. Conor F. Hayes, Timothy Verstraeten, Diederik M. Roijers, Enda Howley, Patrick Mannion |
Neural Comput. Appl. | 5 |
| 2024 | A Meta-Learning Approach for Multi-Objective Reinforcement Learning in Sustainable Home Energy ManagementabstractEffective residential appliance scheduling is crucial for sustainable living. While multi-objective reinforcement learning (MORL) has proven effective in balancing user preferences in appliance scheduling, traditional MORL struggles with limited data in non-stationary residential settings characterized by renewable generation variations. Significant context shifts in the environment can invalidate previously learned policies. To address this, we extend state-of-the-art MORL algorithms with the meta-learning paradigm, enabling rapid, few-shot adaptation to shifting contexts. Additionally, we employ an auto-encoder (AE)-based unsupervised method to detect shifts in environmental context. We have also developed a residential energy environment to evaluate our method using real-world data from London residential settings. This study not only assesses the application of MORL in residential appliance scheduling but also underscores the effectiveness of meta-learning in energy management. Our top-performing method significantly surpasses the best baseline, while the trained model saves 3.28% on electricity bills, a 2.74% increase in user comfort, and a 5.9% improvement in expected utility. Additionally, it reduces the sparsity of solutions by 62.44%. Remarkably, these gains were accomplished using 96.71% less training data and 61.1% fewer training steps. Junlin Lu, Patrick Mannion, Karl Mason |
ECAI | 2 |
| 2024 | Learning in Multi-Objective Public Goods Games with Non-Linear UtilitiesabstractAddressing the question of how to achieve optimal decision-making under risk and uncertainty is crucial for enhancing the capabilities of artificial agents that collaborate with or support humans. In this work, we address this question in the context of Public Goods Games. We study learning in a novel multi-objective version of the Public Goods Game where agents have different risk preferences, by means of multi-objective reinforcement learning. We introduce a parametric non-linear utility function to model risk preferences at the level of individual agents, over the collective and individual reward components of the game. We study the interplay between such preference modelling and environmental uncertainty on the incentive alignment level in the game. We demonstrate how different combinations of individual preferences and environmental uncertainty sustain the emergence of cooperative patterns in non-cooperative environments (i.e., where competitive strategies are dominant), while others sustain competitive patterns in cooperative environments (i.e., where cooperative strategies are dominant). Nicole Orzan, Erman Acar, Davide Grossi, Patrick Mannion, Roxana Radulescu |
ECAI | 4 |
| 2024 | Exploring the Pareto front of multi-objective COVID-19 mitigation policies using reinforcement learningabstractInfectious disease outbreaks can have a disruptive impact on public health and societal processes. As decision-making in the context of epidemic mitigation is multi-dimensional hence complex, reinforcement learning in combination with complex epidemic models provides a methodology to design refined prevention strategies. Current research focuses on optimizing policies with respect to a single objective, such as the pathogen’s attack rate. However, as the mitigation of epidemics involves distinct, and possibly conflicting, criteria (i.a., mortality, morbidity, economic cost, well-being), a multi-objective decision approach is warranted to obtain balanced policies. To enhance future decision-making, we propose a deep multi-objective reinforcement learning approach by building upon a state-of-the-art algorithm called Pareto Conditioned Networks (PCN) to obtain a set of solutions for distinct outcomes of the decision problem. We consider different deconfinement strategies after the first Belgian lockdown within the COVID-19 pandemic and aim to minimize both COVID-19 cases (i.e., infections and hospitalizations) and the societal burden induced by the mitigation measures. As such, we connected a multi-objective Markov decision process with a stochastic compartment model designed to approximate the Belgian COVID-19 waves and explore reactive strategies. As these social mitigation measures are implemented in a continuous action space that modulates the contact matrix of the age-structured epidemic model, we extend PCN to this setting. We evaluate the solution set that PCN returns, and observe that it explored the whole range of possible social restrictions, leading to high-quality trade-offs, as it captured the problem dynamics. In this work, we demonstrate that multi-objective reinforcement learning adds value to epidemiological modeling and provides essential insights to balance mitigation policies. Mathieu Reymond, Conor F. Hayes, Lander Willem, Roxana Radulescu, Steven Abrams, Diederik M. Roijers, Enda Howley, Patrick Mannion, Niel Hens, Ann Nowé, Pieter Libin |
Expert Syst. Appl. | 8 |
| 2024 | Inferring preferences from demonstrations in multi-objective reinforcement learning
Junlin Lu, Patrick Mannion, Karl Mason |
Neural Comput. Appl. | 2 |
| 2023 | Know Your Enemy: Identifying Adversarial Behaviours in Deep Reinforcement Learning Agents (Student Abstract)abstractIt has been shown that an agent can be trained with an adversarial policy which achieves high degrees of success against a state-of-the-art DRL victim despite taking unintuitive actions. This prompts the question: is this adversarial behaviour detectable through the observations of the victim alone? We find that widely used classification methods such as random forests are only able to achieve a maximum of ≈71% test set accuracy when classifying an agent for a single timestep. However, when the classifier inputs are treated as time-series data, test set classification accuracy is increased significantly to ≈98%. This is true for both classification of episodes as a whole, and for “live” classification at each timestep in an episode. These classifications can then be used to “react” to incoming attacks and increase the overall win rate against Adversarial opponents by approximately 17%. Classification of the victim’s own internal activations in response to the adversary is shown to achieve similarly impressive accuracy while also offering advantages like increased transferability to other domains. Seán Caulfield Curley, Karl Mason, Patrick Mannion |
AAAI | 3 |
| 2023 | Distributional Multi-Objective Decision MakingabstractFor effective decision support in scenarios with conflicting objectives, sets of potentially optimal solutions can be presented to the decision maker. We explore both what policies these sets should contain and how such sets can be computed efficiently. With this in mind, we take a distributional approach and introduce a novel dominance criterion relating return distributions of policies directly. Based on this criterion, we present the distributional undominated set and show that it contains optimal policies otherwise ignored by the Pareto front. In addition, we propose the convex distributional undominated set and prove that it comprises all policies that maximise expected utility for multivariate risk-averse decision makers. We propose a novel algorithm to learn the distributional undominated set and further contribute pruning operators to reduce the set to the convex distributional undominated set. Through experiments, we demonstrate the feasibility and effectiveness of these methods, making this a valuable new approach for decision support in real-world problems. Willem Röpke, Conor F. Hayes, Patrick Mannion, Enda Howley, Ann Nowé, Diederik M. Roijers |
IJCAI | 3 |
| 2023 | Monte Carlo tree search algorithms for risk-aware and multi-objective reinforcement learningabstractAbstract In many risk-aware and multi-objective reinforcement learning settings, the utility of the user is derived from a single execution of a policy. In these settings, making decisions based on the average future returns is not suitable. For example, in a medical setting a patient may only have one opportunity to treat their illness. Making decisions using just the expected future returns–known in reinforcement learning as the value–cannot account for the potential range of adverse or positive outcomes a decision may have. Therefore, we should use the distribution over expected future returns differently to represent the critical information that the agent requires at decision time by taking both the future and accrued returns into consideration. In this paper, we propose two novel Monte Carlo tree search algorithms. Firstly, we present a Monte Carlo tree search algorithm that can compute policies for nonlinear utility functions (NLU-MCTS) by optimising the utility of the different possible returns attainable from individual policy executions, resulting in good policies for both risk-aware and multi-objective settings. Secondly, we propose a distributional Monte Carlo tree search algorithm (DMCTS) which extends NLU-MCTS. DMCTS computes an approximate posterior distribution over the utility of the returns, and utilises Thompson sampling during planning to compute policies in risk-aware and multi-objective settings. Both algorithms outperform the state-of-the-art in multi-objective reinforcement learning for the expected utility of the returns. Conor F. Hayes, Mathieu Reymond, Diederik M. Roijers, Enda Howley, Patrick Mannion |
Auton. Agents Multi Agent Syst. | 5 |
| 2022 | A practical guide to multi-objective reinforcement learning and planningabstractAbstract Real-world sequential decision-making tasks are generally complex, requiring trade-offs between multiple, often conflicting, objectives. Despite this, the majority of research in reinforcement learning and decision-theoretic planning either assumes only a single objective, or that multiple objectives can be adequately handled via a simple linear combination. Such approaches may oversimplify the underlying problem and hence produce suboptimal results. This paper serves as a guide to the application of multi-objective methods to difficult problems, and is aimed at researchers who are already familiar with single-objective reinforcement learning and planning methods who wish to adopt a multi-objective perspective on their research, as well as practitioners who encounter multi-objective decision problems in practice. It identifies the factors that may influence the nature of the desired solution, and illustrates by example how these influence the design of multi-objective decision-making systems for complex problems. Conor F. Hayes, Roxana Radulescu, Eugenio Bargiacchi, Johan Källström, Matthew Macfarlane, Mathieu Reymond, Timothy Verstraeten, Luisa M. Zintgraf, Richard Dazeley, Fredrik Heintz, Enda Howley, Athirai Aravazhi Irissappane, Patrick Mannion, Ann Nowé, Gabriel de Oliveira Ramos, Marcello Restelli, Peter Vamplew 0001, Diederik M. Roijers |
Auton. Agents Multi Agent Syst. | 13 |
| 2022 | Scalar reward is not enough: a response to Silver, Singh, Precup and Sutton (2021)abstractAbstract The recent paper “Reward is Enough” by Silver, Singh, Precup and Sutton posits that the concept of reward maximisation is sufficient to underpin all intelligence, both natural and artificial, and provides a suitable basis for the creation of artificial general intelligence. We contest the underlying assumption of Silver et al. that such reward can be scalar-valued. In this paper we explain why scalar rewards are insufficient to account for some aspects of both biological and computational intelligence, and argue in favour of explicitly multi-objective models of reward maximisation. Furthermore, we contend that even if scalar reward functions can trigger intelligent behaviour in specific cases, this type of reward is insufficient for the development of human-aligned artificial general intelligence due to unacceptable risks of unsafe or unethical behaviour. Peter Vamplew 0001, Benjamin J. Smith, Johan Källström, Gabriel de Oliveira Ramos, Roxana Radulescu, Diederik M. Roijers, Conor F. Hayes, Fredrik Heintz, Patrick Mannion, Pieter Libin, Richard Dazeley, Cameron Foale |
Auton. Agents Multi Agent Syst. | 9 |
| 2022 | Opponent learning awareness and modelling in multi-objective normal form games
Roxana Radulescu, Timothy Verstraeten, Patrick Mannion, Diederik M. Roijers, Ann Nowé |
Neural Comput. Appl. | 4 |
| 2022 | Special issue on adaptive and learning agents 2020
Felipe Leno da Silva, Patrick MacAlpine, Roxana Radulescu, Fernando P. Santos 0001, Patrick Mannion |
Neural Comput. Appl. | 5 |
| 2022 | Deep Reinforcement Learning for Autonomous Driving: A SurveyabstractWith the development of deep representation learning, the domain of reinforcement learning (RL) has become a powerful learning framework now capable of learning complex policies in high dimensional environments. This review summarises deep reinforcement learning (DRL) algorithms and provides a taxonomy of automated driving tasks where (D)RL methods have been employed, while addressing key computational challenges in real world deployment of autonomous driving agents. It also delineates adjacent domains such as behavior cloning, imitation learning, inverse reinforcement learning that are related but are not classical RL algorithms. The role of simulators in training agents, methods to validate, test and robustify existing solutions in RL are discussed. Bangalore Ravi Kiran, Ibrahim Sobh, Victor Talpaert, Patrick Mannion, Ahmad A. Al Sallab, Senthil Kumar Yogamani, Patrick Pérez |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | Multi-objective multi-agent decision making: a utility-based analysis and survey
Roxana Radulescu, Patrick Mannion, Diederik M. Roijers, Ann Nowé |
Auton. Agents Multi Agent Syst. | 2 |
| 2017 | Policy invariance under reward transformations for multi-objective reinforcement learning
Patrick Mannion, Sam Devlin, Karl Mason, Jim Duggan, Enda Howley |
Neurocomputing | 1 |
| 2006 | Integrating UWB in consumer devices
Patrick Mannion |
CCNC | 1 |