Roxana Radulescu

dblp:180/3142 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-1446-5514ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Multi-objective reinforcement learning for provably incentivising alignment with value systems
abstract
This paper addresses the problem of ensuring that autonomous learning agents align with multiple moral values. Specifically, we present the theoretical principles and algorithmic tools necessary for creating an environment where we ensure that the agent learns a behaviour aligned with multiple moral values while striving to achieve its individual objective. To address this value alignment problem, we adopt the Multi-Objective Reinforcement Learning framework and propose a novel algorithm that combines techniques from Multi-Objective Reinforcement Learning and Linear Programming. In addition, we illustrate our value alignment process with an example involving an autonomous vehicle. Here, we demonstrate that the agent learns to behave in alignment with the ethical values of safety, achievement, and comfort, with achievement representing the agent’s individual objective. Such ethical behaviour differs depending on the ordering between values. We also use a synthetic multi-objective environment to evaluate the computational costs of guaranteeing ethical learning as the number of values increases.
Manel Rodriguez-Soto, Roxana Radulescu, Filippo Bistaffa, Oriol Ricart, Arnau Mayoral-Macau, Maite López-Sánchez, Juan A. Rodríguez-Aguilar, Ann Nowé
Artif. Intell.2
2025 Multi-Objective Reinforcement Learning for Water Management
Zuzanna Osika, Roxana Radulescu, Jazmin Zatarain Salazar, Frans A. Oliehoek, Pradeep K. Murukannaiah
AAMAS2
2025 Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning
Willem Röpke, Mathieu Reymond, Patrick Mannion, Diederik M. Roijers, Ann Nowé, Roxana Radulescu
AAMAS6
2025 Learning in public goods games: the effects of uncertainty and communication on cooperation
abstract
Communication is a widely used mechanism to promote cooperation in multi-agent systems. In the field of emergent communication, agents are typically trained in specific environments: cooperative, competitive or mixed-motive. Motivated by the idea that real-world settings are characterized by incomplete information and that humans face daily interactions under a wide spectrum of incentives, we aim to explore the role of emergent communication when simultaneously exploited across all these contexts. In this work, we pursue this line of research by focusing on social dilemmas. To do this, we developed an extended version of the Public Goods Game, which allows us to train independent reinforcement learning agents simultaneously in different scenarios where incentives are (mis)aligned to various extents. Additionally, agents experience uncertainty in terms of the alignment of their incentives with those of others. We equip agents with the ability to learn a communication policy and study the impact of emergent communication in the face of uncertainty among agents. Our findings show that in settings where all agents have the same level of uncertainty, communication can enhance the cooperation of the whole group. However, in cases of asymmetric uncertainty, the agents that do not face uncertainty learn to use communication to deceive and exploit their uncertain peers.
Nicole Orzan, Erman Acar, Davide Grossi, Roxana Radulescu
Neural Comput. Appl.4
2025 Preference communication in multi-objective normal-form games
Willem Röpke, Diederik M. Roijers, Ann Nowé, Roxana Radulescu
Neural Comput. Appl.4
2024 Learning in Multi-Objective Public Goods Games with Non-Linear Utilities
abstract
Addressing the question of how to achieve optimal decision-making under risk and uncertainty is crucial for enhancing the capabilities of artificial agents that collaborate with or support humans. In this work, we address this question in the context of Public Goods Games. We study learning in a novel multi-objective version of the Public Goods Game where agents have different risk preferences, by means of multi-objective reinforcement learning. We introduce a parametric non-linear utility function to model risk preferences at the level of individual agents, over the collective and individual reward components of the game. We study the interplay between such preference modelling and environmental uncertainty on the incentive alignment level in the game. We demonstrate how different combinations of individual preferences and environmental uncertainty sustain the emergence of cooperative patterns in non-cooperative environments (i.e., where competitive strategies are dominant), while others sustain competitive patterns in cooperative environments (i.e., where cooperative strategies are dominant).
Nicole Orzan, Erman Acar, Davide Grossi, Patrick Mannion, Roxana Radulescu
ECAI5
2024 The World is a Multi-Objective Multi-Agent System: Now What?
abstract
Most complex problems of social relevance, such as climate change mitigation, traffic management, taxation policy design, or infrastructure management, involve both multiple stakeholders and multiple potentially conflicting objectives. In a nutshell, the majority of real world problems are multi-agent and multi-objective in nature. Artificial intelligence (AI) is a pivotal tool in designing solutions for such critical domains that come with high impact and ramifications across many dimensions, from societal and economic well-being, to ethical, political, and legal levels. Given the current theoretical and algorithmic developments in AI, it is an opportune moment to take a holistic approach and design decision-support tools that: (i) tackle all the prominent challenges of such problems and consider both the multi-agent and multi-objective aspects; (ii) exhibit vital characteristics, such as explainability and transparency, in order to enhance user agency and alignment. These are the challenges that I will discuss during the Frontiers in AI session at ECAI 2024, together with a brief overview of my work and next steps for this field. This paper summarises my contribution to the session.
Roxana Radulescu
ECAI1
2024 Exploring the Pareto front of multi-objective COVID-19 mitigation policies using reinforcement learning
abstract
Infectious disease outbreaks can have a disruptive impact on public health and societal processes. As decision-making in the context of epidemic mitigation is multi-dimensional hence complex, reinforcement learning in combination with complex epidemic models provides a methodology to design refined prevention strategies. Current research focuses on optimizing policies with respect to a single objective, such as the pathogen’s attack rate. However, as the mitigation of epidemics involves distinct, and possibly conflicting, criteria (i.a., mortality, morbidity, economic cost, well-being), a multi-objective decision approach is warranted to obtain balanced policies. To enhance future decision-making, we propose a deep multi-objective reinforcement learning approach by building upon a state-of-the-art algorithm called Pareto Conditioned Networks (PCN) to obtain a set of solutions for distinct outcomes of the decision problem. We consider different deconfinement strategies after the first Belgian lockdown within the COVID-19 pandemic and aim to minimize both COVID-19 cases (i.e., infections and hospitalizations) and the societal burden induced by the mitigation measures. As such, we connected a multi-objective Markov decision process with a stochastic compartment model designed to approximate the Belgian COVID-19 waves and explore reactive strategies. As these social mitigation measures are implemented in a continuous action space that modulates the contact matrix of the age-structured epidemic model, we extend PCN to this setting. We evaluate the solution set that PCN returns, and observe that it explored the whole range of possible social restrictions, leading to high-quality trade-offs, as it captured the problem dynamics. In this work, we demonstrate that multi-objective reinforcement learning adds value to epidemiological modeling and provides essential insights to balance mitigation policies.
Mathieu Reymond, Conor F. Hayes, Lander Willem, Roxana Radulescu, Steven Abrams, Diederik M. Roijers, Enda Howley, Patrick Mannion, Niel Hens, Ann Nowé, Pieter Libin
Expert Syst. Appl.4
2022 A practical guide to multi-objective reinforcement learning and planning
abstract
Abstract Real-world sequential decision-making tasks are generally complex, requiring trade-offs between multiple, often conflicting, objectives. Despite this, the majority of research in reinforcement learning and decision-theoretic planning either assumes only a single objective, or that multiple objectives can be adequately handled via a simple linear combination. Such approaches may oversimplify the underlying problem and hence produce suboptimal results. This paper serves as a guide to the application of multi-objective methods to difficult problems, and is aimed at researchers who are already familiar with single-objective reinforcement learning and planning methods who wish to adopt a multi-objective perspective on their research, as well as practitioners who encounter multi-objective decision problems in practice. It identifies the factors that may influence the nature of the desired solution, and illustrates by example how these influence the design of multi-objective decision-making systems for complex problems.
Conor F. Hayes, Roxana Radulescu, Eugenio Bargiacchi, Johan Källström, Matthew Macfarlane, Mathieu Reymond, Timothy Verstraeten, Luisa M. Zintgraf, Richard Dazeley, Fredrik Heintz, Enda Howley, Athirai Aravazhi Irissappane, Patrick Mannion, Ann Nowé, Gabriel de Oliveira Ramos, Marcello Restelli, Peter Vamplew 0001, Diederik M. Roijers
Auton. Agents Multi Agent Syst.2
2022 On nash equilibria in normal-form games with vectorial payoffs
Willem Röpke, Diederik M. Roijers, Ann Nowé, Roxana Radulescu
Auton. Agents Multi Agent Syst.4
2022 Scalar reward is not enough: a response to Silver, Singh, Precup and Sutton (2021)
abstract
Abstract The recent paper “Reward is Enough” by Silver, Singh, Precup and Sutton posits that the concept of reward maximisation is sufficient to underpin all intelligence, both natural and artificial, and provides a suitable basis for the creation of artificial general intelligence. We contest the underlying assumption of Silver et al. that such reward can be scalar-valued. In this paper we explain why scalar rewards are insufficient to account for some aspects of both biological and computational intelligence, and argue in favour of explicitly multi-objective models of reward maximisation. Furthermore, we contend that even if scalar reward functions can trigger intelligent behaviour in specific cases, this type of reward is insufficient for the development of human-aligned artificial general intelligence due to unacceptable risks of unsafe or unethical behaviour.
Peter Vamplew 0001, Benjamin J. Smith, Johan Källström, Gabriel de Oliveira Ramos, Roxana Radulescu, Diederik M. Roijers, Conor F. Hayes, Fredrik Heintz, Patrick Mannion, Pieter Libin, Richard Dazeley, Cameron Foale
Auton. Agents Multi Agent Syst.5
2022 Opponent learning awareness and modelling in multi-objective normal form games
Roxana Radulescu, Timothy Verstraeten, Patrick Mannion, Diederik M. Roijers, Ann Nowé
Neural Comput. Appl.1
2022 Special issue on adaptive and learning agents 2020
Felipe Leno da Silva, Patrick MacAlpine, Roxana Radulescu, Fernando P. Santos 0001, Patrick Mannion
Neural Comput. Appl.3
2020 Multi-objective multi-agent decision making: a utility-based analysis and survey
Roxana Radulescu, Patrick Mannion, Diederik M. Roijers, Ann Nowé
Auton. Agents Multi Agent Syst.1