VLDB 2026 Research / reviewers in the wild / expert
Roland Stolz
dblp:35/1128
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0002-0653-572XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 94% Motion planning and robot control · 6% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › policy optimization
policy gradient |
1.8 | 2 | 2026 | Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions · AAAI 2026 Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024 |
Machine learning › Reinforcement learning › constrained reinforcement learning
action-constrained reinforcement learning |
1.0 | 1 | 2026 | Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions · AAAI 2026 |
Machine learning › Reinforcement learning
constrained reinforcement learning |
1.0 | 1 | 2026 | Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions · AAAI 2026 |
Machine learning › Reinforcement learning › policy optimization › policy gradient
stochastic policy gradient |
1.0 | 1 | 2026 | Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions · AAAI 2026 |
Machine learning › Reinforcement learning › safe reinforcement learning
action masking |
0.8 | 1 | 2024 | Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024 |
Machine learning › Reinforcement learning
action space design |
0.8 | 1 | 2024 | Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024 |
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization |
0.8 | 1 | 2024 | Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024 |
Robotics › Motion planning and robot control
robot control |
0.2 | 1 | 2024 | Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024 |
Robotics › Motion planning and robot control › robot control
safe control |
0.2 | 1 | 2024 | Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
truncated normal distribution · 1.0numerical approximation · 1.0policy gradient · 0.8action masking · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Stochastic Action-Constrained Reinforcement Learning via Truncated DistributionsabstractIn reinforcement learning (RL), it is often advantageous to consider additional constraints on the action space to ensure safety or action relevance. Existing work on such action-constrained RL faces challenges regarding effective policy updates, computational efficiency, and predictable runtime. Recent work proposes to use truncated normal distributions for stochastic policy gradient methods. However, the computation of key characteristics, such as the entropy, log-probability, and their gradients, becomes intractable under complex constraints. Hence, prior work approximates these using the non-truncated distributions, which severely degrades performance. We argue that accurate estimation of these characteristics is crucial in the action-constrained RL setting, and propose efficient numerical approximations for them. We also provide an efficient sampling strategy for truncated policy distributions and validate our approach on three benchmark environments, which demonstrate significant performance improvements when using accurate estimations. Roland Stolz, Michael Eichelbeck, Matthias Althoff |
AAAI | 1 |
| 2026 | No More Traffic Tickets: A Tutorial to Ensure Traffic-Rule Compliance of Automated VehiclesabstractImagine an automated vehicle violating a traffic rule and, by that, causing an accident. This would not only be devastating for a responsible operator, but more importantly, each such incident would erode the trust of the public in automated vehicles. Fortunately, compliance with traffic rules can be fully controlled, unless other traffic participants breach them—this causality obviously makes it possible for responsible operators to avoid liability claims. Traffic rules can be seen as guardrails for automated driving and should take center stage. Unfortunately, this is currently not the case. Many traffic rules are typically implicitly embedded in various fragments of the software stack of automated vehicles. Instead, the considered traffic rules should be explicitly and centrally provided. Adherence to traffic rules should also be ensured by formal methods to gain the trust needed in the public. This article provides all the steps required to achieve this goal. Due to the interdisciplinary nature of this topic (law, engineering, and computer science), we aim to address a broad audience by focusing on the governing principles of the presented methods, and we refer to more technical works for details. Matthias Althoff, Sebastian Maierhofer, Gerald Würsching, Yuanfei Lin, Florian Lercher, Roland Stolz |
Proc. IEEE | 6 |
| 2024 | Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action MaskingabstractContinuous action spaces in reinforcement learning (RL) are commonly defined as multidimensional intervals. While intervals usually reflect the action boundaries for tasks well, they can be challenging for learning because the typically large global action space leads to frequent exploration of irrelevant actions. Yet, little task knowledge can be sufficient to identify significantly smaller state-specific sets of relevant actions. Focusing learning on these relevant actions can significantly improve training efficiency and effectiveness. In this paper, we propose to focus learning on the set of relevant actions and introduce three continuous action masking methods for exactly mapping the action space to the state-dependent set of relevant actions. Thus, our methods ensure that only relevant actions are executed, enhancing the predictability of the RL agent and enabling its use in safety-critical applications. We further derive the implications of the proposed methods on the policy gradient. Using proximal policy optimization ( PPO), we evaluate our methods on four control tasks, where the relevant action set is computed based on the system dynamics and a relevant state set. Our experiments show that the three action masking methods achieve higher final rewards and converge faster than the baseline without action masking. Roland Stolz, Hanna Krasowski, Jakob Thumm, Michael Eichelbeck, Philipp Gassert, Matthias Althoff |
NeurIPS | 1 |