EDBT 2026 Demo / reviewers in the wild / expert
Diederik M. Roijers
dblp:116/9425 · also Diederik Marijn Roijers
· DBLP profile ↗
40ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-2825-2491ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalized policy improvement for efficient and robust multi-objective reinforcement learningabstractMulti-objective reinforcement learning (MORL) algorithms tackle sequential decision problems where agents may have different preferences over (possibly conflicting) reward functions. These algorithms often learn a set of policies, each optimized for a particular agent preference, that are later reused when optimizing policies for different preferences. We introduce a novel algorithm that builds upon Generalized Policy Improvement (GPI) to construct principled, formally-derived prioritization schemes that improve sample efficiency. These correspond to active-learning strategies by which the agent can identify (i) the most promising preferences/objectives to train on at each moment; and (ii) the most relevant previous experiences to learn policies for new agent preferences through a novel Dyna-style MORL method. We prove our algorithm is guaranteed to always converge to an optimal solution in a finite number of steps, or an $$\epsilon $$ -optimal solution (for a bounded $$\epsilon $$ ) if the agent can only identify sub-optimal policies. Our method monotonically improves the quality of its partial solutions while learning. We also introduce a bound that characterizes the maximum utility loss (with respect to the optimal solution) incurred by intermediate policies identified by our method during learning. Finally, we propose a novel epistemic uncertainty-aware extension of GPI that exploits high-confidence lower bounds to mitigate the impact of unreliable action-value estimates in GPI policies, and prove that it provides tighter performance bounds than the current state of the art. We empirically show that our method outperforms state-of-the-art MORL algorithms in challenging multi-objective tasks. Lucas Nunes Alegre, Ana L. C. Bazzan, Diederik M. Roijers, Ann Nowé, Bruno C. da Silva 0001 |
Auton. Agents Multi Agent Syst. | 3 |
| 2026 | Scalable Multi-Objective Reinforcement Learning with Fairness Guarantees using Lorenz DominanceabstractMulti-Objective Reinforcement Learning (MORL) aims to learn a set of policies that optimize trade-offs between multiple, often conflicting objectives. MORL is computationally more complex than single-objective RL, particularly as the number of objectives increases. Additionally, when objectives involve the preferences of agents or groups, incorporating fairness becomes both important and socially desirable. This paper introduces a principled algorithm that incorporates fairness into MORL while improving scalability to many-objective problems. We propose using Lorenz dominance to identify policies with equitable reward distributions and introduce λ-Lorenz dominance to enable flexible fairness preferences. We release a new, large-scale real-world transport planning environment and demonstrate that our method encourages the discovery of fair policies, showing improved scalability in two large cities (Xi’an and Amsterdam). Our methods outperform common multi-objective approaches, particularly in high-dimensional objective spaces. Dimitrios Michailidis, Willem Röpke, Diederik M. Roijers, Sennay Ghebreab, Fernando P. Santos 0001 |
J. Artif. Intell. Res. | 3 |
| 2025 | Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning
Willem Röpke, Mathieu Reymond, Patrick Mannion, Diederik M. Roijers, Ann Nowé, Roxana Radulescu |
AAMAS | 4 |
| 2025 | Expected scalarised returns dominance: a new solution concept for multi-objective decision makingabstractAbstract In many real-world scenarios, the utility of a user is derived from a single execution of a policy. In this case, to apply multi-objective reinforcement learning, the expected utility of the returns must be optimised. Various scenarios exist where a user’s preferences over objectives (also known as the utility function) are unknown or difficult to specify. In such scenarios, a set of optimal policies must be learned. However, settings where the expected utility must be maximised have been largely overlooked by the multi-objective reinforcement learning community and, as a consequence, a set of optimal solutions has yet to be defined. In this work, we propose first-order stochastic dominance as a criterion to build solution sets to maximise expected utility. We also define a new dominance criterion, known as expected scalarised returns (ESR) dominance, that extends first-order stochastic dominance to allow a set of optimal policies to be learned in practice. Additionally, we define a new solution concept called the ESR set, which is a set of policies that are ESR dominant. Finally, we present a new multi-objective tabular distributional reinforcement learning (MOTDRL) algorithm to learn the ESR set in multi-objective multi-armed bandit settings. Conor F. Hayes, Timothy Verstraeten, Diederik M. Roijers, Enda Howley, Patrick Mannion |
Neural Comput. Appl. | 3 |
| 2025 | Deep multi-objective reinforcement learning for utility-based infrastructural maintenance optimizationabstractIn this paper, we introduce multi-objective deep centralized multi-agent actor-critic (MO-DCMAC), a multi-objective reinforcement learning method for infrastructural maintenance optimization, an area traditionally dominated by single-objective reinforcement learning (RL) approaches. Previous single-objective RL methods combine multiple objectives, such as probability of collapse and cost, into a singular reward signal through reward-shaping. In contrast, MO-DCMAC can optimize a policy for multiple objectives directly, even when the utility function is nonlinear. We evaluated MO-DCMAC using two utility functions, which use probability of collapse and cost as input. The first utility function is the threshold utility, in which MO-DCMAC should minimize cost so that the probability of collapse is never above the threshold. The second is based on the failure mode, effects, and criticality analysis methodology used by asset managers to assess maintenance plans. We evaluated MO-DCMAC, with both utility functions, in multiple maintenance environments, including ones based on a case study of the historical quay walls of Amsterdam. The performance of MO-DCMAC was compared against multiple rule-based policies based on heuristics currently used for constructing maintenance plans. Our results demonstrate that MO-DCMAC outperforms traditional rule-based policies across various environments and utility functions. Jesse van Remmerden, Maurice Kenter, Diederik M. Roijers, Charalampos Andriotis, Yingqian Zhang 0001, Zaharah Bukhsh |
Neural Comput. Appl. | 3 |
| 2025 | Preference communication in multi-objective normal-form games
Willem Röpke, Diederik M. Roijers, Ann Nowé, Roxana Radulescu |
Neural Comput. Appl. | 2 |
| 2024 | The Wasserstein Believer: Learning Belief Updates for Partially Observable Environments through Reliable Latent Space ModelsabstractPartially Observable Markov Decision Processes (POMDPs) are used to model environments where the state cannot be perceived, necessitating reasoning based on past observations and actions. However, remembering the full history is generally intractable due to the exponential growth in the history space. Maintaining a probability distribution that models the belief over the current state can be used as a sufficient statistic of the history, but its computation requires access to the model of the environment and is often intractable. While SOTA algorithms use Recurrent Neural Networks to compress the observation-action history aiming to learn a sufficient statistic, they lack guarantees of success and can lead to sub-optimal policies. To overcome this, we propose the Wasserstein Belief Updater, an RL algorithm that learns a latent model of the POMDP and an approximation of the belief update under the assumption that the state is observable during training. Our approach comes with theoretical guarantees on the quality of our approximation ensuring that our latent beliefs allow for learning the optimal value function. Raphaël Avalos, Florent Delgrange, Ann Nowé, Guillermo A. Pérez, Diederik M. Roijers |
ICLR | 5 |
| 2024 | Exploring the Pareto front of multi-objective COVID-19 mitigation policies using reinforcement learningabstractInfectious disease outbreaks can have a disruptive impact on public health and societal processes. As decision-making in the context of epidemic mitigation is multi-dimensional hence complex, reinforcement learning in combination with complex epidemic models provides a methodology to design refined prevention strategies. Current research focuses on optimizing policies with respect to a single objective, such as the pathogen’s attack rate. However, as the mitigation of epidemics involves distinct, and possibly conflicting, criteria (i.a., mortality, morbidity, economic cost, well-being), a multi-objective decision approach is warranted to obtain balanced policies. To enhance future decision-making, we propose a deep multi-objective reinforcement learning approach by building upon a state-of-the-art algorithm called Pareto Conditioned Networks (PCN) to obtain a set of solutions for distinct outcomes of the decision problem. We consider different deconfinement strategies after the first Belgian lockdown within the COVID-19 pandemic and aim to minimize both COVID-19 cases (i.e., infections and hospitalizations) and the societal burden induced by the mitigation measures. As such, we connected a multi-objective Markov decision process with a stochastic compartment model designed to approximate the Belgian COVID-19 waves and explore reactive strategies. As these social mitigation measures are implemented in a continuous action space that modulates the contact matrix of the age-structured epidemic model, we extend PCN to this setting. We evaluate the solution set that PCN returns, and observe that it explored the whole range of possible social restrictions, leading to high-quality trade-offs, as it captured the problem dynamics. In this work, we demonstrate that multi-objective reinforcement learning adds value to epidemiological modeling and provides essential insights to balance mitigation policies. Mathieu Reymond, Conor F. Hayes, Lander Willem, Roxana Radulescu, Steven Abrams, Diederik M. Roijers, Enda Howley, Patrick Mannion, Niel Hens, Ann Nowé, Pieter Libin |
Expert Syst. Appl. | 6 |
| 2023 | What Lies beyond the Pareto Front? A Survey on Decision-Support Methods for Multi-Objective OptimizationabstractWe present a review that unifies decision-support methods for exploring the solutions produced by multi-objective optimization (MOO) algorithms. As MOO is applied to solve diverse problems, approaches for analyzing the trade-offs offered by these algorithms are scattered across fields. We provide an overview of the current advances on this topic, including methods for visualization, mining the solution set, and uncertainty exploration as well as emerging research directions, including interactivity, explainability, and support on ethical aspects. We synthesize these methods drawing from different fields of research to enable building a unified approach, independent of the application. Our goals are to reduce the entry barrier for researchers and practitioners on using MOO algorithms and to provide novel research directions. Zuzanna Osika, Jazmin Zatarain Salazar, Diederik M. Roijers, Frans A. Oliehoek, Pradeep K. Murukannaiah |
IJCAI | 3 |
| 2023 | Distributional Multi-Objective Decision MakingabstractFor effective decision support in scenarios with conflicting objectives, sets of potentially optimal solutions can be presented to the decision maker. We explore both what policies these sets should contain and how such sets can be computed efficiently. With this in mind, we take a distributional approach and introduce a novel dominance criterion relating return distributions of policies directly. Based on this criterion, we present the distributional undominated set and show that it contains optimal policies otherwise ignored by the Pareto front. In addition, we propose the convex distributional undominated set and prove that it comprises all policies that maximise expected utility for multivariate risk-averse decision makers. We propose a novel algorithm to learn the distributional undominated set and further contribute pruning operators to reduce the set to the convex distributional undominated set. Through experiments, we demonstrate the feasibility and effectiveness of these methods, making this a valuable new approach for decision support in real-world problems. Willem Röpke, Conor F. Hayes, Patrick Mannion, Enda Howley, Ann Nowé, Diederik M. Roijers |
IJCAI | 6 |
| 2023 | Monte Carlo tree search algorithms for risk-aware and multi-objective reinforcement learningabstractAbstract In many risk-aware and multi-objective reinforcement learning settings, the utility of the user is derived from a single execution of a policy. In these settings, making decisions based on the average future returns is not suitable. For example, in a medical setting a patient may only have one opportunity to treat their illness. Making decisions using just the expected future returns–known in reinforcement learning as the value–cannot account for the potential range of adverse or positive outcomes a decision may have. Therefore, we should use the distribution over expected future returns differently to represent the critical information that the agent requires at decision time by taking both the future and accrued returns into consideration. In this paper, we propose two novel Monte Carlo tree search algorithms. Firstly, we present a Monte Carlo tree search algorithm that can compute policies for nonlinear utility functions (NLU-MCTS) by optimising the utility of the different possible returns attainable from individual policy executions, resulting in good policies for both risk-aware and multi-objective settings. Secondly, we propose a distributional Monte Carlo tree search algorithm (DMCTS) which extends NLU-MCTS. DMCTS computes an approximate posterior distribution over the utility of the returns, and utilises Thompson sampling during planning to compute policies in risk-aware and multi-objective settings. Both algorithms outperform the state-of-the-art in multi-objective reinforcement learning for the expected utility of the returns. Conor F. Hayes, Mathieu Reymond, Diederik M. Roijers, Enda Howley, Patrick Mannion |
Auton. Agents Multi Agent Syst. | 3 |
| 2023 | Actor-critic multi-objective reinforcement learning for non-linear utility functions
Mathieu Reymond, Conor F. Hayes, Denis Steckelmacher, Diederik M. Roijers, Ann Nowé |
Auton. Agents Multi Agent Syst. | 4 |
| 2022 | A practical guide to multi-objective reinforcement learning and planningabstractAbstract Real-world sequential decision-making tasks are generally complex, requiring trade-offs between multiple, often conflicting, objectives. Despite this, the majority of research in reinforcement learning and decision-theoretic planning either assumes only a single objective, or that multiple objectives can be adequately handled via a simple linear combination. Such approaches may oversimplify the underlying problem and hence produce suboptimal results. This paper serves as a guide to the application of multi-objective methods to difficult problems, and is aimed at researchers who are already familiar with single-objective reinforcement learning and planning methods who wish to adopt a multi-objective perspective on their research, as well as practitioners who encounter multi-objective decision problems in practice. It identifies the factors that may influence the nature of the desired solution, and illustrates by example how these influence the design of multi-objective decision-making systems for complex problems. Conor F. Hayes, Roxana Radulescu, Eugenio Bargiacchi, Johan Källström, Matthew Macfarlane, Mathieu Reymond, Timothy Verstraeten, Luisa M. Zintgraf, Richard Dazeley, Fredrik Heintz, Enda Howley, Athirai Aravazhi Irissappane, Patrick Mannion, Ann Nowé, Gabriel de Oliveira Ramos, Marcello Restelli, Peter Vamplew 0001, Diederik M. Roijers |
Auton. Agents Multi Agent Syst. | 18 |
| 2022 | On nash equilibria in normal-form games with vectorial payoffs
Willem Röpke, Diederik M. Roijers, Ann Nowé, Roxana Radulescu |
Auton. Agents Multi Agent Syst. | 2 |
| 2022 | Scalar reward is not enough: a response to Silver, Singh, Precup and Sutton (2021)abstractAbstract The recent paper “Reward is Enough” by Silver, Singh, Precup and Sutton posits that the concept of reward maximisation is sufficient to underpin all intelligence, both natural and artificial, and provides a suitable basis for the creation of artificial general intelligence. We contest the underlying assumption of Silver et al. that such reward can be scalar-valued. In this paper we explain why scalar rewards are insufficient to account for some aspects of both biological and computational intelligence, and argue in favour of explicitly multi-objective models of reward maximisation. Furthermore, we contend that even if scalar reward functions can trigger intelligent behaviour in specific cases, this type of reward is insufficient for the development of human-aligned artificial general intelligence due to unacceptable risks of unsafe or unethical behaviour. Peter Vamplew 0001, Benjamin J. Smith, Johan Källström, Gabriel de Oliveira Ramos, Roxana Radulescu, Diederik M. Roijers, Conor F. Hayes, Fredrik Heintz, Patrick Mannion, Pieter Libin, Richard Dazeley, Cameron Foale |
Auton. Agents Multi Agent Syst. | 6 |
| 2022 | Opponent learning awareness and modelling in multi-objective normal form games
Roxana Radulescu, Timothy Verstraeten, Patrick Mannion, Diederik M. Roijers, Ann Nowé |
Neural Comput. Appl. | 5 |
| 2021 | Deep Learning-Based Energy Disaggregation and On/Off Detection of Household AppliancesabstractEnergy disaggregation, a.k.a. Non-Intrusive Load Monitoring, aims to separate the energy consumption of individual appliances from the readings of a mains power meter measuring the total energy consumption of, e.g., a whole house. Energy consumption of individual appliances can be useful in many applications, e.g., providing appliance-level feedback to the end users to help them understand their energy consumption and ultimately save energy. Recently, with the availability of large-scale energy consumption datasets, various neural network models such as convolutional neural networks and recurrent neural networks have been investigated to solve the energy disaggregation problem. Neural network models can learn complex patterns from large amounts of data and have been shown to outperform the traditional machine learning methods such as variants of hidden Markov models. However, current neural network methods for energy disaggregation are either computational expensive or are not capable of handling long-term dependencies. In this article, we investigate the application of the recently developed WaveNet models for the task of energy disaggregation. Based on a real-world energy dataset collected from 20 households over 2 years, we show that WaveNet models outperforms the state-of-the-art deep learning methods proposed in the literature for energy disaggregation in terms of both error measures and computational cost. On the basis of energy disaggregation, we then investigate the performance of two deep-learning based frameworks for the task of on/off detection which aims at estimating whether an appliance is in operation or not. The first framework obtains the on/off states of an appliance by binarising the predictions of a regression model trained for energy disaggregation, while the second framework obtains the on/off states of an appliance by directly training a binary classifier with binarised energy readings of the appliance serving as the target values. Based on the same dataset, we show that for the task of on/off detection the second framework, i.e., directly training a binary classifier, achieves better performance in terms of F1 score. Jie Jiang 0011, Qiuqiang Kong, Mark D. Plumbley, G. Nigel Gilbert, Mark Hoogendoorn, Diederik M. Roijers |
ACM Trans. Knowl. Discov. Data | 6 |
| 2020 | Interactive Multi-objective Reinforcement Learning in Multi-armed Bandits with Gaussian Process Utility Models
Diederik M. Roijers, Luisa M. Zintgraf, Pieter Libin, Mathieu Reymond, Eugenio Bargiacchi, Ann Nowé |
ECML/PKDD (3) | 1 |
| 2020 | Multi-objective multi-agent decision making: a utility-based analysis and survey
Roxana Radulescu, Patrick Mannion, Diederik M. Roijers, Ann Nowé |
Auton. Agents Multi Agent Syst. | 3 |
| 2020 | AI-Toolbox: A C++ library for Reinforcement Learning and Planning (with Python Bindings)abstractThis paper describes AI-Toolbox, a C++ software library that contains reinforcement learning and planning algorithms, and supports both single and multi agent problems, as well as partial observability. It is designed for simplicity and clarity, and contains extensive documentation of its API and code. It supports Python to enable users not comfortable with C++ to take advantage of the library's speed and functionality. AI-Toolbox is free software, and is hosted online at https://github.com/Svalorzen/AI-Toolbox. Eugenio Bargiacchi, Diederik M. Roijers, Ann Nowé |
J. Mach. Learn. Res. | 2 |
| 2020 | Avatars of a Feather Flock Together: Gender Homophily in Online Video Games Revealed via Exponential Random Graph ModelingabstractIncreasingly more scholars have regarded the virtual worlds of massively multiplayer online games (MMOGs) as a social laboratory, and have paid research attention to the online interactions between its large number of players. In the present study, we examine a widely observed real-life phenomenon gender homophily (i.e., people of the same gender flocking together) in this virtual context. Specifically, we investigate how collaboration networks in an MMOG (Destiny) are shaped by the adopted gender of the avatar characters. This focus is interesting, as avatar gender in video games is generally a choice that is less constrained than it is in real life. To investigate the effect of gender in online video games, while controlling for the effects of several confounding factors, we employed a technique called exponential random graph modeling. Mirroring how interpersonal relationships are often gendered in real life, and despite common phenomenon such as gender swapping, we found evidence supporting gender homophily in the MMOG environment. Sander Bakkes, Diederik M. Roijers, Pieter Spronck |
IEEE Trans. Games | 3 |
| 2019 | Dynamic Weights in Multi-Objective Deep Reinforcement LearningabstractMany real-world decision problems are characterized by multiple conflicting objectives which must be balanced based on their relative importance. In the dynamic weights setting the relative importance changes over time and specialized algorithms that deal with such change, such as a tabular Reinforcement Learning (RL) algorithm by Natarajan and Tadepalli (2005), are required. However, this earlier work is not feasible for RL settings that necessitate the use of function approximators. We generalize across weight changes and high-dimensional inputs by proposing a multi-objective Q-network whose outputs are conditioned on the relative importance of objectives and we introduce Diverse Experience Replay (DER) to counter the inherent non-stationarity of the Dynamic Weights setting. We perform an extensive experimental evaluation and compare our methods to adapted algorithms from Deep Multi-Task/Multi-Objective Reinforcement Learning and show that our proposed network in combination with DER dominates these adapted algorithms across weight change scenarios and problem domains. Axel Abels, Diederik M. Roijers, Tom Lenaerts, Ann Nowé, Denis Steckelmacher |
ICML | 2 |
| 2019 | Bayesian Anytime m-top ExplorationabstractWe introduce Boundary Focused Thompson sampling (BFTS), a new Bayesian algorithm to solve the anytime m-top exploration problem, where the objective is to identify the m best arms in a multi-armed bandit. First, we consider a set of existing benchmark problems that consider sub-Gaussian reward distributions (i.e., Gaussian with fixed variance and categorical reward). Next, we introduce a new environment inspired by a real world decision problem concerning insect control for organic agriculture. This new environment encodes a Poisson rewards distribution. For all these benchmarks, we experimentally show that BFTS consistently outperforms AT-LUCB, the current state of the art algorithm. Pieter Libin, Timothy Verstraeten, Diederik M. Roijers, Kristof Theys, Ann Nowé |
ICTAI | 3 |
| 2019 | Sample-Efficient Model-Free Reinforcement Learning with Off-Policy Critics
Denis Steckelmacher, Hélène Plisnier, Diederik M. Roijers, Ann Nowé |
ECML/PKDD (3) | 3 |
| 2018 | Reinforcement Learning in POMDPs With Memoryless Options and Option-Observation Initiation SetsabstractMany real-world reinforcement learning problems have a hierarchical nature, and often exhibit some degree of partial observability. While hierarchy and partial observability are usually tackled separately (for instance by combining recurrent neural networks and options), we show that addressing both problems simultaneously is simpler and more efficient in many cases. More specifically, we make the initiation set of options conditional on the previously-executed option, and show that options with such Option-Observation Initiation Sets (OOIs) are at least as expressive as Finite State Controllers (FSCs), a state-of-the-art approach for learning in POMDPs. OOIs are easy to design based on an intuitive description of the task, lead to explainable policies and keep the top-level and option policies memoryless. Our experiments show that OOIs allow agents to learn optimal policies in challenging POMDPs, while being much more sample-efficient than a recurrent neural network over options. Denis Steckelmacher, Diederik M. Roijers, Anna Harutyunyan, Peter Vrancx, Hélène Plisnier, Ann Nowé |
AAAI | 2 |
| 2018 | Real-Time Robot Vision on Low-Performance Computing HardwareabstractSmall robots have numerous interesting applications in domains like industry, education, scientific research, and services. For most applications vision is important, however, the limitations of the computing hardware make this a challenging task. In this paper, we address the problem of real-time object recognition and propose the Fast Regions of Interest Search (FROIS) algorithm to quickly find the ROIs of the objects in small robots with low-performance hardware. Subsequently, we use two methods to analyze the ROIs. First, we develop a Convolutional Neural Network on a desktop and deploy it onto the low-performance hardware for object recognition. Second, we adopt the Histogram of Oriented Gradients descriptor and linear Support Vector Machines classifier and optimize the HOG component for faster speed. The experimental results show that the methods work well on our small robots with Raspberry Pi 3 embedded 1.2 GHz ARM CPUs to recognize the objects. Furthermore, we obtain valuable insights about the trade-offs between speed and accuracy. Gongjin Lan, Jesús Benito-Picazo, Diederik M. Roijers, Enrique Domínguez, A. E. Eiben |
ICARCV | 3 |
| 2018 | Learning to Coordinate with Coordination Graphs in Repeated Single-Stage Multi-Agent Decision ProblemsabstractLearning to coordinate between multiple agents is an important problem in many reinforcement learning problems. Key to learning to coordinate is exploiting loose couplings, i.e., conditional independences between agents. In this paper we study learning in repeated fully cooperative games, multi-agent multi-armed bandits (MAMABs), in which the expected rewards can be expressed as a coordination graph. We propose multi-agent upper confidence exploration (MAUCE), a new algorithm for MAMABs that exploits loose couplings, which enables us to prove a regret bound that is logarithmic in the number of arm pulls and only linear in the number of agents. We empirically compare MAUCE to sparse cooperative Q-learning, and a state-of-the-art combinatorial bandit approach, and show that it performs much better on a variety of settings, including learning control policies for wind farms. Eugenio Bargiacchi, Timothy Verstraeten, Diederik M. Roijers, Ann Nowé, Hado van Hasselt |
ICML | 3 |
| 2018 | Bayesian Best-Arm Identification for Selecting Influenza Mitigation Strategies
Pieter Libin, Timothy Verstraeten, Diederik M. Roijers, Jelena Grujic, Kristof Theys, Philippe Lemey, Ann Nowé |
ECML/PKDD (3) | 3 |
| 2018 | Directed Locomotion for Modular Robots with Evolvable Morphologies
Gongjin Lan, Milan Jelisavcic, Diederik M. Roijers, Evert Haasdijk, A. E. Eiben |
PPSN (1) | 3 |
| 2016 | Solving Transition-Independent Multi-Agent MDPs with Sparse InteractionsabstractIn cooperative multi-agent sequential decision making under uncertainty, agents must coordinate to find an optimal joint policy that maximises joint value. Typical algorithms exploit additive structure in the value function, but in the fully-observable multi-agent MDP (MMDP) setting such structure is not present. We propose a new optimal solver for transition-independent MMDPs, in which agents can only affect their own state but their reward depends on joint transitions. We represent these dependencies compactly in conditional return graphs (CRGs). Using CRGs the value of a joint policy and the bounds on partially specified joint policies can be efficiently computed. We propose CoRe, a novel branch-and-bound policy search algorithm building on CRGs. CoRe typically requires less runtime than available alternatives and finds solutions to previously unsolvable problems. Joris Scharpff, Diederik M. Roijers, Frans A. Oliehoek, Matthijs T. J. Spaan, Mathijs de Weerdt |
AAAI | 2 |
| 2016 | Structure in the Value Function of Two-Player Zero-Sum Games of Incomplete InformationabstractIn this paper, we introduce a new formulation for the value function of a zero-sum Partially Observable Stochastic Game (zs-POSG) in terms of a ‘plan-time sufficient statistic’, a distribution over joint sets of information. We prove that this value function exhibits concavity and convexity with respect to appropriately chosen subspaces of the statistic space. We anticipate that this result is a key pre-cursor for developing solution methods that exploit such structure. Finally, we show that the formulation allow us to reduce a finite zs-POSG to a ‘centralized’ model with shared observations, thereby transferring results for the latter (narrower) class of games to games with individual observations. Auke J. Wiggers, Frans A. Oliehoek, Diederik M. Roijers |
ECAI | 3 |
| 2016 | Balancing Relevance Criteria through Multi-Objective OptimizationabstractOffline evaluation of information retrieval systems typically focuses on a single effectiveness measure that models the utility for a typical user. Such a measure usually combines a behavior-based rank discount with a notion of document utility that captures the single relevance criterion of topicality. However, for individual users relevance criteria such as credibility, reputability or readability can strongly impact the utility. Also, for different information needs the utility can be a different mixture of these criteria. Because of the focus on single metrics, offline optimization of IR systems does not account for different preferences in balancing relevance criteria. Joost van Doorn, Daan Odijk, Diederik M. Roijers, Maarten de Rijke |
SIGIR | 3 |
| 2015 | Pareto Local Search for MOMDP Planning
Chiel Kooijman, Maarten de Waard, Maarten Inja, Diederik M. Roijers, Shimon Whiteson |
ESANN | 4 |
| 2015 | Grammar-based Procedural Content Generation from Designer-provided Difficulty Curves
Mircea Traichioiu, Sander Bakkes, Diederik M. Roijers |
FDG | 3 |
| 2015 | Efficient Methods for Multi-Objective Decision-Theoretic Planning
Diederik M. Roijers |
IJCAI | 1 |
| 2015 | Point-Based Planning for Multi-Objective POMDPs
Diederik M. Roijers, Shimon Whiteson, Frans A. Oliehoek |
IJCAI | 1 |
| 2015 | Computing Convex Coverage Sets for Faster Multi-objective CoordinationabstractIn this article, we propose new algorithms for multi-objective coordination graphs (MO-CoGs). Key to the efficiency of these algorithms is that they compute a convex coverage set (CCS) instead of a Pareto coverage set (PCS). Not only is a CCS a sufficient solution set for a large class of problems, it also has important characteristics that facilitate more efficient solutions. We propose two main algorithms for computing a CCS in MO-CoGs. Convex multi-objective variable elimination (CMOVE) computes a CCS by performing a series of agent eliminations, which can be seen as solving a series of local multi-objective subproblems. Variable elimination linear support (VELS) iteratively identifies the single weight vector, w, that can lead to the maximal possible improvement on a partial CCS and calls variable elimination to solve a scalarized instance of the problem for w. VELS is faster than CMOVE for small and medium numbers of objectives and can compute an ε-approximate CCS in a fraction of the runtime. In addition, we propose variants of these methods that employ AND/OR tree search instead of variable elimination to achieve memory efficiency. We analyze the runtime and space complexities of these methods, prove their correctness, and compare them empirically against a naive baseline and an existing PCS method, both in terms of memory-usage and runtime. Our results show that, by focusing on the CCS, these methods achieve much better scalability in the number of agents than the current state of the art. Diederik M. Roijers, Shimon Whiteson, Frans A. Oliehoek |
J. Artif. Intell. Res. | 1 |
| 2014 | Queued Pareto Local Search for Multi-Objective Optimization
Maarten Inja, Chiel Kooijman, Maarten de Waard, Diederik M. Roijers, Shimon Whiteson |
PPSN | 4 |
| 2013 | A Survey of Multi-Objective Sequential Decision-MakingabstractSequential decision-making problems with multiple objectives arise naturally in practice and pose unique challenges for research in decision-theoretic planning and learning, which has largely focused on single-objective settings. This article surveys algorithms designed for sequential decision-making problems with multiple objectives. Though there is a growing body of literature on this subject, little of it makes explicit under what circumstances special methods are needed to solve multi-objective problems. Therefore, we identify three distinct scenarios in which converting such a problem to a single-objective one is impossible, infeasible, or undesirable. Furthermore, we propose a taxonomy that classifies multi-objective methods according to the applicable scenario, the nature of the scalarization function (which projects multi-objective values to scalar ones), and the type of policies considered. We show how these factors determine the nature of an optimal solution, which can be a single policy, a convex hull, or a Pareto front. Using this taxonomy, we survey the literature on multi-objective methods for planning and learning. Finally, we discuss key applications of such methods and outline opportunities for future work. Diederik M. Roijers, Peter Vamplew 0001, Shimon Whiteson, Richard Dazeley |
J. Artif. Intell. Res. | 1 |
| 2012 | Probability estimation and a competence model for rule based e-tutoring systemsabstractIn this paper, we present a student model for rule based e-tutoring systems. This model describes both properties of rewrite rules (difficulty and discriminativity) and of students (start competence and learning speed). The model is an extension of the two-parameter logistic ogive function of Item Response Theory. We show that the model can be applied even to relatively small datasets. We gather data from students working on problems in the logic domain, and show that the model estimates of rule difficulty correspond well to expert opinions. We also show that the estimated start competence corresponds well to our expectations based on the previous experience of the students in the logic domain. We point out that this model can be used to inform students about their competence and learning, and teachers about the students and the difficulty and discriminativity of the rules. Diederik M. Roijers, Johan Jeuring, A. J. Feelders |
LAK | 1 |