Roland Stolz

dblp:35/1128 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0002-0653-572XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 94% Motion planning and robot control · 6%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › policy optimization
policy gradient
1.822026
Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions · AAAI 2026
Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024
Machine learning › Reinforcement learning › constrained reinforcement learning
action-constrained reinforcement learning
1.012026
Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions · AAAI 2026
Machine learning › Reinforcement learning
constrained reinforcement learning
1.012026
Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions · AAAI 2026
Machine learning › Reinforcement learning › policy optimization › policy gradient
stochastic policy gradient
1.012026
Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions · AAAI 2026
Machine learning › Reinforcement learning › safe reinforcement learning
action masking
0.812024
Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024
Machine learning › Reinforcement learning
action space design
0.812024
Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization
0.812024
Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024
Robotics › Motion planning and robot control
robot control
0.212024
Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024
Robotics › Motion planning and robot control › robot control
safe control
0.212024
Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

truncated normal distribution · 1.0numerical approximation · 1.0policy gradient · 0.8action masking · 0.8
YearPublicationVenuePosition
2026 Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions
abstract
In reinforcement learning (RL), it is often advantageous to consider additional constraints on the action space to ensure safety or action relevance. Existing work on such action-constrained RL faces challenges regarding effective policy updates, computational efficiency, and predictable runtime. Recent work proposes to use truncated normal distributions for stochastic policy gradient methods. However, the computation of key characteristics, such as the entropy, log-probability, and their gradients, becomes intractable under complex constraints. Hence, prior work approximates these using the non-truncated distributions, which severely degrades performance. We argue that accurate estimation of these characteristics is crucial in the action-constrained RL setting, and propose efficient numerical approximations for them. We also provide an efficient sampling strategy for truncated policy distributions and validate our approach on three benchmark environments, which demonstrate significant performance improvements when using accurate estimations.
Roland Stolz, Michael Eichelbeck, Matthias Althoff
AAAI1
2026 No More Traffic Tickets: A Tutorial to Ensure Traffic-Rule Compliance of Automated Vehicles
abstract
Imagine an automated vehicle violating a traffic rule and, by that, causing an accident. This would not only be devastating for a responsible operator, but more importantly, each such incident would erode the trust of the public in automated vehicles. Fortunately, compliance with traffic rules can be fully controlled, unless other traffic participants breach them—this causality obviously makes it possible for responsible operators to avoid liability claims. Traffic rules can be seen as guardrails for automated driving and should take center stage. Unfortunately, this is currently not the case. Many traffic rules are typically implicitly embedded in various fragments of the software stack of automated vehicles. Instead, the considered traffic rules should be explicitly and centrally provided. Adherence to traffic rules should also be ensured by formal methods to gain the trust needed in the public. This article provides all the steps required to achieve this goal. Due to the interdisciplinary nature of this topic (law, engineering, and computer science), we aim to address a broad audience by focusing on the governing principles of the presented methods, and we refer to more technical works for details.
Matthias Althoff, Sebastian Maierhofer, Gerald Würsching, Yuanfei Lin, Florian Lercher, Roland Stolz
Proc. IEEE6
2024 Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking
abstract
Continuous action spaces in reinforcement learning (RL) are commonly defined as multidimensional intervals. While intervals usually reflect the action boundaries for tasks well, they can be challenging for learning because the typically large global action space leads to frequent exploration of irrelevant actions. Yet, little task knowledge can be sufficient to identify significantly smaller state-specific sets of relevant actions. Focusing learning on these relevant actions can significantly improve training efficiency and effectiveness. In this paper, we propose to focus learning on the set of relevant actions and introduce three continuous action masking methods for exactly mapping the action space to the state-dependent set of relevant actions. Thus, our methods ensure that only relevant actions are executed, enhancing the predictability of the RL agent and enabling its use in safety-critical applications. We further derive the implications of the proposed methods on the policy gradient. Using proximal policy optimization ( PPO), we evaluate our methods on four control tasks, where the relevant action set is computed based on the system dynamics and a relevant state set. Our experiments show that the three action masking methods achieve higher final rewards and converge faster than the baseline without action masking.
Roland Stolz, Hanna Krasowski, Jakob Thumm, Michael Eichelbeck, Philipp Gassert, Matthias Althoff
NeurIPS1