Jakob Thumm

dblp:315/0679 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-0282-2908ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A General Safety Framework for Autonomous Manipulation in Human Environments
abstract
Autonomous robots are projected to significantly augment the manual workforce, especially in repetitive and hazardous tasks. For a successful deployment of such robots in human environments, it is crucial to guarantee human safety. State-of-the-art approaches to ensure human safety are either too conservative to permit a natural human-robot collaboration or make strong assumptions that do not hold for autonomous robots, e.g., knowledge of a pre-defined trajectory. Therefore, we propose the shield for Safe Autonomous human-robot collaboration through Reachability Analysis (SARA shield). This novel power and force limiting framework provides formal safety guarantees for manipulation in human environments while realizing fast robot speeds. As unconstrained contacts allow for significantly higher contact forces than constrained contacts (also known as clamping), we use reachability analysis to classify potential contacts by their type in a formally correct way. For each contact type, we formally verify that the kinetic energy of the robot is below pain and injury thresholds for the respective human body part in contact. Our experiments show that SARA shield satisfies the contact safety constraints while significantly improving the robot performance in comparison to state-of-the-art approaches.
Jakob Thumm, Julian Balletshofer, Leonardo Maglanoc, Luis Muschal, Matthias Althoff
IEEE Trans. Robotics1
2025 Multi-Objective Causal Bayesian Optimization
abstract
In decision-making problems, the outcome of an intervention often depends on the causal relationships between system components and is highly costly to evaluate. In such settings, causal Bayesian optimization (CBO) exploits the causal relationships between the system variables and sequentially performs interventions to approach the optimum with minimal data. Extending CBO to the multi-outcome setting, we propose *multi-objective Causal Bayesian optimization* (MO-CBO), a paradigm for identifying Pareto-optimal interventions within a known multi-target causal graph. Our methodology first reduces the search space by discarding sub-optimal interventions based on the structure of the given causal graph. We further show that any MO-CBO problem can be decomposed into several traditional multi-objective optimization tasks. Our proposed MO-CBO algorithm is designed to identify Pareto-optimal interventions by iteratively exploring these underlying tasks, guided by relative hypervolume improvement. Experiments on synthetic and real-world causal graphs demonstrate the superiority of our approach over non-causal multi-objective Bayesian optimization in settings where causal information is available.
Shriya Bhatija, Paul-David Joshua Zuercher, Jakob Thumm, Thomas Bohné
ICML3
2024 Human-Robot Gym: Benchmarking Reinforcement Learning in Human-Robot Collaboration
abstract
Deep reinforcement learning (RL) has shown promising results in robot motion planning with first attempts in human-robot collaboration (HRC). However, a fair comparison of RL approaches in HRC under the constraint of guaranteed safety is yet to be made. We, therefore, present human-robot gym, a benchmark suite for safe RL in HRC. We provide challenging, realistic HRC tasks in a modular simulation framework. Most importantly, human-robot gym is the first benchmark suite that includes a safety shield to provably guarantee human safety. This bridges a critical gap between theoretic RL research and its real-world deployment. Our evaluation of six tasks led to three key results: (a) the diverse nature of the tasks offered by human-robot gym creates a challenging benchmark for state-of-the-art RL methods, (b) by leveraging expert knowledge in form of an action imitation reward, the RL agent can outperform the expert, and (c) our agents negligibly overfit to training data.
Jakob Thumm, Felix Trost, Matthias Althoff
ICRA1
2024 Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking
abstract
Continuous action spaces in reinforcement learning (RL) are commonly defined as multidimensional intervals. While intervals usually reflect the action boundaries for tasks well, they can be challenging for learning because the typically large global action space leads to frequent exploration of irrelevant actions. Yet, little task knowledge can be sufficient to identify significantly smaller state-specific sets of relevant actions. Focusing learning on these relevant actions can significantly improve training efficiency and effectiveness. In this paper, we propose to focus learning on the set of relevant actions and introduce three continuous action masking methods for exactly mapping the action space to the state-dependent set of relevant actions. Thus, our methods ensure that only relevant actions are executed, enhancing the predictability of the RL agent and enabling its use in safety-critical applications. We further derive the implications of the proposed methods on the policy gradient. Using proximal policy optimization ( PPO), we evaluate our methods on four control tasks, where the relevant action set is computed based on the system dynamics and a relevant state set. Our experiments show that the three action masking methods achieve higher final rewards and converge faster than the baseline without action masking.
Roland Stolz, Hanna Krasowski, Jakob Thumm, Michael Eichelbeck, Philipp Gassert, Matthias Althoff
NeurIPS3
2023 Approximate First-Passage Time Distributions for Gaussian Motion and Transportation Models
abstract
We aim to approximate the distribution of the first-passage time of a particle moving according to a Gaussian process with increasing trend, i. e., the distribution of the first time a particle described, e.g., by a state-space model such as a constant-velocity or constant-acceleration model, arrives at a fixed location. Since the known approaches from the literature either consider processes from different families or lead to highly complex approximations, we seek a fast-to-compute method for the problem. Motivated by an engineering particle transport task for which we can assume that once a particle has arrived at this 10-cation it cannot move back, we derive an analytic approximation for the first-passage time probabilities and calculate its inverse cumulative distribution function analytically and the moments numerically. Furthermore, we propose a Gaussian approximation based on a linearization approach. The strengths and limitations of our methods are discussed and by comparison with Monte Carlo simulations, we show that in particular, the first one satisfies the requirements of engineering problems in terms of accuracy and computation time.
Marcel Reith-Braun, Florian Pfaff, Jakob Thumm, Uwe D. Hanebeck
FUSION3
2023 Reducing Safety Interventions in Provably Safe Reinforcement Learning
abstract
Deep Reinforcement Learning (RL) has shown promise in addressing complex robotic challenges. In real-world applications, RL is often accompanied by failsafe controllers as a last resort to avoid catastrophic events. While necessary for safety, these interventions can result in undesirable behaviors, such as abrupt braking or aggressive steering. This paper proposes two safety intervention reduction methods: proactive replacement and proactive projection, which change the action of the agent if it leads to a potential failsafe intervention. These approaches are compared to state-of-the-art constrained RL on the OpenAI safety gym benchmark and a human-robot collab-oration task. Our study demonstrates that the combination of our method with provably safe RL leads to high-performing policies with zero safety violations and a low number of failsafe interventions. Our versatile method can be applied to a wide range of real-world robotic tasks, while effectively improving safety without sacrificing task performance.
Jakob Thumm, Guillaume Pelat, Matthias Althoff
IROS1
2022 SaRA: A Tool for Safe Human-Robot Coexistence and Collaboration through Reachability Analysis
abstract
Current safety mechanisms implementing industry standards for human-robot coexistence separate humans and robots through caging. Other approaches allowing humans to enter the workspace of manipulators do not provide formal safety guarantees. Thus, this study aims to facilitate the widespread adoption of collaborative robots by presenting SaRA, an extensible tool that performs set-based reachability analysis and formally guarantees safety. Our experimental results show that the set-based prediction of a human can be computed in a few microseconds, using SaRA, allowing for real-time consideration of many surrounding humans in an environment.
Sven R. Schepp, Jakob Thumm, Stefan B. Liu, Matthias Althoff
ICRA2
2022 Provably Safe Deep Reinforcement Learning for Robotic Manipulation in Human Environments
abstract
Deep reinforcement learning (RL) has shown promising results in the motion planning of manipulators. However, no method guarantees the safety of highly dynamic obstacles, such as humans, in RL-based manipulator control. This lack of formal safety assurances prevents the application of RL for manipulators in real-world human environments. Therefore, we propose a shielding mechanism that ensures ISO- verified human safety while training and deploying RL algorithms on manipulators. We utilize a fast reachability analysis of humans and manipulators to guarantee that the manipulator comes to a complete stop before a human is within its range. Our proposed method guarantees safety and significantly improves the RL performance by preventing episode-ending collisions. We demonstrate the performance of our proposed method in simulation using human motion capture data.
Jakob Thumm, Matthias Althoff
ICRA1
2022 Mixture of Experts of Neural Networks and Kalman Filters for Optical Belt Sorting
abstract
In optical sorting of bulk material, the composition of particles may frequently change. State-of-the-art sorting approaches rely on tuning physical models of the particle motion. The aim of this work is to increase the prediction accuracy in complex fast-changing sorting scenarios with data-driven approaches. In this article, we propose two neural network (NN) experts for accurate prediction ofa prioriknown particle types. To handle the large variety of particle types that can occur in real-world sorting scenarios, we introduce a simple but effective mixture of experts’ approach that combines NNs with hand-crafted motion models. Our new method not only improves the prediction accuracy for bulk material consisting of many particle classes, but also proves to be very adaptive and robust to new particle types.
Jakob Thumm, Marcel Reith-Braun, Florian Pfaff, Uwe D. Hanebeck, Merle Flitter, Georg Maier, Robin Gruna, Thomas Längle, Albert Bauer, Harald Kruggel-Emden
IEEE Trans. Ind. Informatics1