EDBT 2026 Demo / reviewers in the wild / expert
Maris F. L. Galesloot
dblp:365/4403
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Planning, search and constraint satisfaction · 52% Reinforcement learning · 26% Trustworthy machine learning · 11% | |
| Theoretical computer science
1 paper |
Automated reasoning and model checking · 100% |
Topics — the 7 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
partially observable markov decision process |
2.5 | 3 | 2025 | On Evaluating Policies for Robust POMDPs · NeurIPS 2025 Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs · IJCAI 2025 Factored Online Planning in Many-Agent POMDPs · AAAI 2024 |
Machine learning › Reinforcement learning
policy evaluation |
0.9 | 1 | 2025 | On Evaluating Policies for Robust POMDPs · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › robustness › model robustness
robust policy |
0.9 | 1 | 2025 | Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs · IJCAI 2025 |
Machine learning › Reinforcement learning › robust reinforcement learning
robust policy optimization |
0.9 | 1 | 2025 | Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs · IJCAI 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › domain-independent planning
factored planning |
0.8 | 1 | 2024 | Factored Online Planning in Many-Agent POMDPs · AAAI 2024 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent decision making |
0.8 | 1 | 2024 | Factored Online Planning in Many-Agent POMDPs · AAAI 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
online planning |
0.8 | 1 | 2024 | Factored Online Planning in Many-Agent POMDPs · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
worst-case POMDP computation · 1.7subgradient ascent · 1.7formal verification · 1.7upper bounds · 0.9MDP reduction · 0.9particle filtering · 0.8monte carlo tree search · 0.8coordination graph · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pessimistic Iterative Planning with RNNs for Robust POMDPsabstractRobust POMDPs extend classical POMDPs to incorporate model uncertainty using so-called uncertainty sets on the transition and observation functions, effectively defining ranges of probabilities. Policies for robust POMDPs must be (1) memory-based to account for partial observability and (2) robust against model uncertainty to account for the worst-case probability instances from the uncertainty sets. To compute such robust memory-based policies, we propose the pessimistic iterative planning (PIP) framework, which alternates between (1) selecting pessimistic POMDPs via worst-case probability instances from the uncertainty sets, and (2) computing finite-state controllers (FSCs) for these pessimistic POMDPs. Within PIP, we propose the RFSCNET algorithm, which optimizes a recurrent neural network to compute the FSCs. The empirical evaluation shows that RFSCNET can compute better-performing robust policies than several baselines and a state-of-the-art robust POMDP solver. Maris F. L. Galesloot, Marnix Suilen, Thiago D. Simão, Steven Carr 0002, Matthijs T. J. Spaan, Ufuk Topcu, Nils Jansen 0001 |
ECAI | 1 |
| 2025 | Robust Finite-Memory Policy Gradients for Hidden-Model POMDPsabstractPartially observable Markov decision processes (POMDPs) model specific environments in sequential decision-making under uncertainty. Critically, optimal policies for POMDPs may not be robust against perturbations in the environment. Hidden-model POMDPs (HM-POMDPs) capture sets of different environment models, that is, POMDPs with a shared action and observation space. The intuition is that the true model is hidden among a set of potential models, and it is unknown which model will be the environment at execution time. A policy is robust for a given HM-POMDP if it achieves sufficient performance for each of its POMDPs. We compute such robust policies by combining two orthogonal techniques: (1) a deductive formal verification technique that supports tractable robust policy evaluation by computing a worst-case POMDP within the HM-POMDP, and (2) subgradient ascent to optimize the candidate policy for a worst-case POMDP. The empirical evaluation shows that, compared to various baselines, our approach (1) produces policies that are more robust and generalize better to unseen POMDPs, and (2) scales to HM-POMDPs that consist of over a hundred thousand environments. Maris F. L. Galesloot, Roman Andriushchenko, Milan Ceska 0002, Sebastian Junges, Nils Jansen 0001 |
IJCAI | 1 |
| 2025 | On Evaluating Policies for Robust POMDPsabstractRobust partially observable Markov decision processes (RPOMDPs) model sequential decision-making problems under partial observability, where an agent must be robust against a range of dynamics. RPOMDPs can be viewed as a two-player game between an agent, who selects actions, and nature, who adversarially selects the dynamics. Evaluating an agent policy requires finding an adversarial nature
policy, which is computationally challenging. In this paper, we advance the evaluation of agent policies for RPOMDPs in three ways. First, we discuss suitable benchmarks. We observe that for some RPOMDPs, an optimal agent policy can be found by considering only subsets of nature policies, making them easier to solve. We formalize this concept of solvability and construct three benchmarks that are only solvable for expressive sets of nature policies. Second, we describe a new method to evaluate agent policies for RPOMDPs by solving an equivalent MDP. Third, we lift two well-known upper bounds from POMDPs to RPOMDPs, which can be used to efficiently approximate the optimality gap of a policy and serve as baselines. Our experimental evaluation shows that (1) our proposed benchmarks cannot be solved by assuming naive nature policies, (2) our method of evaluating policies is accurate, and (3) the upper bounds provide solid baselines for evaluation. Merlijn Krale, Eline M. Bovy, Maris F. L. Galesloot, Thiago D. Simão, Nils Jansen 0001 |
NeurIPS | 3 |
| 2024 | Factored Online Planning in Many-Agent POMDPsabstractIn centralized multi-agent systems, often modeled as multi-agent partially observable Markov decision processes (MPOMDPs), the action and observation spaces grow exponentially with the number of agents, making the value and belief estimation of single-agent online planning ineffective. Prior work partially tackles value estimation by exploiting the inherent structure of multi-agent settings via so-called coordination graphs. Additionally, belief estimation methods have been improved by incorporating the likelihood of observations into the approximation. However, the challenges of value estimation and belief estimation have only been tackled individually, which prevents existing methods from scaling to settings with many agents. Therefore, we address these challenges simultaneously. First, we introduce weighted particle filtering to a sample-based online planner for MPOMDPs. Second, we present a scalable approximation of the belief. Third, we bring an approach that exploits the typical locality of agent interactions to novel online planning algorithms for MPOMDPs operating on a so-called sparse particle filter tree. Our experimental evaluation against several state-of-the-art baselines shows that our methods (1) are competitive in settings with only a few agents and (2) improve over the baselines in the presence of many agents. Maris F. L. Galesloot, Thiago D. Simão, Sebastian Junges, Nils Jansen 0001 |
AAAI | 1 |