Maris F. L. Galesloot

dblp:365/4403 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Planning, search and constraint satisfaction · 52% Reinforcement learning · 26% Trustworthy machine learning · 11%
Theoretical computer science
1 paper
Automated reasoning and model checking · 100%

Topics — the 7 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
partially observable markov decision process
2.532025
On Evaluating Policies for Robust POMDPs · NeurIPS 2025
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs · IJCAI 2025
Factored Online Planning in Many-Agent POMDPs · AAAI 2024
Machine learning › Reinforcement learning
policy evaluation
0.912025
On Evaluating Policies for Robust POMDPs · NeurIPS 2025
Machine learning › Trustworthy machine learning › robustness › model robustness
robust policy
0.912025
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs · IJCAI 2025
Machine learning › Reinforcement learning › robust reinforcement learning
robust policy optimization
0.912025
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs · IJCAI 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › domain-independent planning
factored planning
0.812024
Factored Online Planning in Many-Agent POMDPs · AAAI 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent decision making
0.812024
Factored Online Planning in Many-Agent POMDPs · AAAI 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
online planning
0.812024
Factored Online Planning in Many-Agent POMDPs · AAAI 2024

Methods — techniques the papers use, named apart from their topics

worst-case POMDP computation · 1.7subgradient ascent · 1.7formal verification · 1.7upper bounds · 0.9MDP reduction · 0.9particle filtering · 0.8monte carlo tree search · 0.8coordination graph · 0.8
YearPublicationVenuePosition
2025 Pessimistic Iterative Planning with RNNs for Robust POMDPs
abstract
Robust POMDPs extend classical POMDPs to incorporate model uncertainty using so-called uncertainty sets on the transition and observation functions, effectively defining ranges of probabilities. Policies for robust POMDPs must be (1) memory-based to account for partial observability and (2) robust against model uncertainty to account for the worst-case probability instances from the uncertainty sets. To compute such robust memory-based policies, we propose the pessimistic iterative planning (PIP) framework, which alternates between (1) selecting pessimistic POMDPs via worst-case probability instances from the uncertainty sets, and (2) computing finite-state controllers (FSCs) for these pessimistic POMDPs. Within PIP, we propose the RFSCNET algorithm, which optimizes a recurrent neural network to compute the FSCs. The empirical evaluation shows that RFSCNET can compute better-performing robust policies than several baselines and a state-of-the-art robust POMDP solver.
Maris F. L. Galesloot, Marnix Suilen, Thiago D. Simão, Steven Carr 0002, Matthijs T. J. Spaan, Ufuk Topcu, Nils Jansen 0001
ECAI1
2025 Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
abstract
Partially observable Markov decision processes (POMDPs) model specific environments in sequential decision-making under uncertainty. Critically, optimal policies for POMDPs may not be robust against perturbations in the environment. Hidden-model POMDPs (HM-POMDPs) capture sets of different environment models, that is, POMDPs with a shared action and observation space. The intuition is that the true model is hidden among a set of potential models, and it is unknown which model will be the environment at execution time. A policy is robust for a given HM-POMDP if it achieves sufficient performance for each of its POMDPs. We compute such robust policies by combining two orthogonal techniques: (1) a deductive formal verification technique that supports tractable robust policy evaluation by computing a worst-case POMDP within the HM-POMDP, and (2) subgradient ascent to optimize the candidate policy for a worst-case POMDP. The empirical evaluation shows that, compared to various baselines, our approach (1) produces policies that are more robust and generalize better to unseen POMDPs, and (2) scales to HM-POMDPs that consist of over a hundred thousand environments.
Maris F. L. Galesloot, Roman Andriushchenko, Milan Ceska 0002, Sebastian Junges, Nils Jansen 0001
IJCAI1
2025 On Evaluating Policies for Robust POMDPs
abstract
Robust partially observable Markov decision processes (RPOMDPs) model sequential decision-making problems under partial observability, where an agent must be robust against a range of dynamics. RPOMDPs can be viewed as a two-player game between an agent, who selects actions, and nature, who adversarially selects the dynamics. Evaluating an agent policy requires finding an adversarial nature policy, which is computationally challenging. In this paper, we advance the evaluation of agent policies for RPOMDPs in three ways. First, we discuss suitable benchmarks. We observe that for some RPOMDPs, an optimal agent policy can be found by considering only subsets of nature policies, making them easier to solve. We formalize this concept of solvability and construct three benchmarks that are only solvable for expressive sets of nature policies. Second, we describe a new method to evaluate agent policies for RPOMDPs by solving an equivalent MDP. Third, we lift two well-known upper bounds from POMDPs to RPOMDPs, which can be used to efficiently approximate the optimality gap of a policy and serve as baselines. Our experimental evaluation shows that (1) our proposed benchmarks cannot be solved by assuming naive nature policies, (2) our method of evaluating policies is accurate, and (3) the upper bounds provide solid baselines for evaluation.
Merlijn Krale, Eline M. Bovy, Maris F. L. Galesloot, Thiago D. Simão, Nils Jansen 0001
NeurIPS3
2024 Factored Online Planning in Many-Agent POMDPs
abstract
In centralized multi-agent systems, often modeled as multi-agent partially observable Markov decision processes (MPOMDPs), the action and observation spaces grow exponentially with the number of agents, making the value and belief estimation of single-agent online planning ineffective. Prior work partially tackles value estimation by exploiting the inherent structure of multi-agent settings via so-called coordination graphs. Additionally, belief estimation methods have been improved by incorporating the likelihood of observations into the approximation. However, the challenges of value estimation and belief estimation have only been tackled individually, which prevents existing methods from scaling to settings with many agents. Therefore, we address these challenges simultaneously. First, we introduce weighted particle filtering to a sample-based online planner for MPOMDPs. Second, we present a scalable approximation of the belief. Third, we bring an approach that exploits the typical locality of agent interactions to novel online planning algorithms for MPOMDPs operating on a so-called sparse particle filter tree. Our experimental evaluation against several state-of-the-art baselines shows that our methods (1) are competitive in settings with only a few agents and (2) improve over the baselines in the presence of many agents.
Maris F. L. Galesloot, Thiago D. Simão, Sebastian Junges, Nils Jansen 0001
AAAI1