VLDB 2026 Research / reviewers in the wild / expert
Steffen Udluft
dblp:34/3546
· DBLP profile ↗
36ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-5767-2591ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Point-wise Q-value maximization for converging Q-learning in continuous state-spacesabstractThis paper introduces a novel Q-learning framework to address instabilities in offline reinforcement learning with continuous state spaces.We identify the recurring collapse of Q-value targets as core challenge and propose a stabilization technique that replaces the iteration-wise targets with their point-wise maximum across iterations.This approach enforces convergence and fully mitigates recursive errors.We show that a performance metric linking Q-values to policy performance is directly available.Our findings represent a first step toward stabilizing Q-learning in challenging settings and highlight the potential of model-based approaches. Philipp Wissmann, Daniel Hein 0001, Steffen Udluft, Thomas A. Runkler |
ESANN | 3 |
| 2026 | Efficient and Resilient Machine Learning for Industrial ApplicationsabstractMachine learning is rapidly transforming industrial landscapes, yet it faces significant hurdles related to efficiency and resilience.This paper discusses industrial challenges and provides a structured overview of current approaches, encompassing data-centric methodologies, efficient training for reliable solutions, hardware-optimized deployment, and the emerging role of foundation models. * This work was Philipp Wissmann, Philip Naumann, Daniel Hein 0001, Steffen Udluft, Marc Weber, Simon Leszek, Thomas A. Runkler |
ESANN | 4 |
| 2025 | TEA: Trajectory Encoding Augmentation for Robust and Transferable Policies in Offline Reinforcement LearningabstractIn this paper, we investigate offline reinforcement learning (RL) with the goal of training a single robust policy that generalizes effectively across environments with unseen dynamics.We propose a novel approach, Trajectory Encoding Augmentation (TEA), which extends the state space by integrating latent representations of environmental dynamics obtained from sequence encoders, such as autoencoders.Our findings show that incorporating these encodings with TEA improves the transferability of a single policy to novel environments with new dynamics, surpassing methods that rely solely on unmodified states.These results indicate that TEA captures critical, environment-specific characteristics, enabling RL agents to generalize effectively across dynamic conditions. Batikan Bora Ormanci, Phillip Swazinna, Steffen Udluft, Thomas A. Runkler |
ESANN | 3 |
| 2025 | Is Q-learning an Ill-posed Problem?abstractThis paper investigates the instability of Q-learning in continuous environments, a challenge frequently encountered by practitioners.Traditionally, this instability is attributed to bootstrapping and regression model errors.Using a representative reinforcement learning benchmark, we systematically examine the effects of bootstrapping and model inaccuracies by incrementally eliminating these potential error sources.Our findings reveal that even in relatively simple benchmarks, the fundamental task of Q-learning -iteratively learning a Q-function from policy-specific target values -can be inherently ill-posed and prone to failure.These insights cast doubt on the reliability of Q-learning as a universal solution for reinforcement learning problems. Philipp Wissmann, Daniel Hein 0001, Steffen Udluft, Thomas A. Runkler |
ESANN | 3 |
| 2024 | Why long model-based rollouts are no reason for bad Q-value estimatesabstractThis paper explores the use of model-based offline reinforcement learning with long model rollouts.While some literature criticizes this approach due to compounding errors, many practitioners have found success in real-world applications.The paper aims to demonstrate that long rollouts do not necessarily result in exponentially growing errors and can actually produce better Q-value estimates than model-free methods.These findings can potentially enhance reinforcement learning techniques. Philipp Wissmann, Daniel Hein 0001, Steffen Udluft, Volker Tresp |
ESANN | 3 |
| 2023 | Automatic Trade-off Adaptation in Offline RLabstractRecently, offline RL algorithms have been proposed that remain adaptive at runtime.For example, the LION algorithm [1] provides the user with an interface to set the trade-off between behavior cloning and optimality w.r.t. the estimated return at runtime.Experts can then use this interface to adapt the policy behavior according to their preferences and find a good trade-off between conservatism and performance optimization.Since expert time is precious, we extend the methodology with an autopilot that automatically finds the best parameterization of the trade-off, yielding a new algorithm which we term AutoLION. Phillip Swazinna, Steffen Udluft, Thomas A. Runkler |
ESANN | 2 |
| 2023 | User-Interactive Offline Reinforcement Learning
Phillip Swazinna, Steffen Udluft, Thomas A. Runkler |
ICLR | 2 |
| 2022 | Safe Policy Improvement Approaches on Discrete Markov Decision ProcessesabstractSafe Policy Improvement (SPI) aims at provable guarantees that a learned policy is at least approximately as good as a given baseline policy. Building on SPI with Soft Baseline Bootstrapping (Soft-SPIBB) by Nadjahi et al., we identify theoretical issues in their approach, provide a corrected theory, and derive a new algorithm that is provably safe on finite Markov Decision Processes (MDP). Additionally, we provide a heuristic algorithm that exhibits the best performance among many state of the art SPI algorithms on two different benchmarks. Furthermore, we introduce a taxonomy of SPI algorithms and empirically show an interesting property of two classes of SPI algorithms: while the mean performance of algorithms that incorporate the uncertainty as a penalty on the action-value is higher, actively restricting the set of policies more consistently produces good policies and is, thus, safer. Philipp Scholl 0003, Felix Dietrich, Clemens Otte, Steffen Udluft |
ICAART (2) | 4 |
| 2021 | Behavior Constraining in Weight Space for Offline Reinforcement LearningabstractIn offline reinforcement learning, a policy needs to be learned from a single pre-collected dataset.Typically, policies are thus regularized during training to behave similarly to the data generating policy, by adding a penalty based on a divergence between action distributions of generating and trained policy.We propose a new algorithm, which constrains the policy directly in its weight space instead, and demonstrate its effectiveness in experiments. *The project this paper is based on was supported with funds from the German Federal Ministry of Phillip Swazinna, Steffen Udluft, Daniel Hein 0001, Thomas A. Runkler |
ESANN | 2 |
| 2021 | Overcoming model bias for robust offline deep reinforcement learning
Phillip Swazinna, Steffen Udluft, Thomas A. Runkler |
Eng. Appl. Artif. Intell. | 2 |
| 2018 | Sensitivity analysis for predictive uncertainty
Stefan Depeweg, José Miguel Hernández-Lobato, Steffen Udluft, Thomas A. Runkler |
ESANN | 3 |
| 2018 | Decomposition of Uncertainty in Bayesian Deep Learning for Efficient and Risk-sensitive LearningabstractBayesian neural networks with latent variables are scalable and flexible probabilistic models: they account for uncertainty in the estimation of the network weights and, by making use of latent variables, can capture complex noise patterns in the data. Using these models we show how to perform and utilize a decomposition of uncertainty in aleatoric and epistemic components for decision making purposes. This allows us to successfully identify informative points for active learning of functions with heteroscedastic and bimodal noise. Using the decomposition we further define a novel risk-sensitive criterion for reinforcement learningto identify policies that balance expected cost, model-bias and noise aversion. Stefan Depeweg, José Miguel Hernández-Lobato, Finale Doshi-Velez, Steffen Udluft |
ICML | 4 |
| 2018 | Interpretable policies for reinforcement learning by genetic programming
Daniel Hein 0001, Steffen Udluft, Thomas A. Runkler |
Eng. Appl. Artif. Intell. | 2 |
| 2017 | Learning and Policy Search in Stochastic Dynamical Systems with Bayesian Neural Networks
Stefan Depeweg, José Miguel Hernández-Lobato, Finale Doshi-Velez, Steffen Udluft |
ICLR (Poster) | 4 |
| 2017 | Batch reinforcement learning on the industrial benchmark: First experiencesabstractThe Particle Swarm Optimization Policy (PSO-P) has been recently introduced and proven to produce remarkable results on interacting with academic reinforcement learning benchmarks in an off-policy, batch-based setting. To further investigate the properties and feasibility on real-world applications, this paper investigates PSO-P on the so-called Industrial Benchmark (IB), a novel reinforcement learning (RL) benchmark that aims at being realistic by including a variety of aspects found in industrial applications, such as continuous state and action spaces, a high dimensional, partially observable state space, delayed effects, and complex stochasticity. The experimental results of PSO-P on IB are compared to results of closed-form control policies derived from the model-based Recurrent Control Neural Network (RCNN) and the model-free Neural Fitted Q-Iteration (NFQ). Experiments show that PSO-P is not only of interest for academic benchmarks, but also for real-world industrial applications, since it also yielded the best performing policy in our IB setting. Compared to other well established RL techniques, PSO-P produced outstanding results in performance and robustness, requiring only a relatively low amount of effort in finding adequate parameters or making complex design decisions. Daniel Hein 0001, Steffen Udluft, Michel Tokic, Alexander Hentschel, Thomas A. Runkler, Volkmar Sterzing |
IJCNN | 2 |
| 2017 | Particle swarm optimization for generating interpretable fuzzy reinforcement learning policies
Daniel Hein 0001, Alexander Hentschel, Thomas A. Runkler, Steffen Udluft |
Eng. Appl. Artif. Intell. | 4 |
| 2015 | Exploiting similarity in system identification tasks with recurrent neural networks
Sigurd Spieckermann, Siegmund Düll, Steffen Udluft, Alexander Hentschel, Thomas A. Runkler |
Neurocomputing | 3 |
| 2014 | Exploiting similarity in system identification tasks with recurrent neural networks
Sigurd Spieckermann, Siegmund Düll, Steffen Udluft, Alexander Hentschel, Thomas A. Runkler |
ESANN | 3 |
| 2014 | Regularized Recurrent Neural Networks for Data Efficient Dual-Task Learning
Sigurd Spieckermann, Siegmund Düll, Steffen Udluft, Thomas A. Runkler |
ICANN | 3 |
| 2013 | Ensembles for Continuous Actions in Reinforcement Learning
Siegmund Düll, Steffen Udluft |
ESANN | 2 |
| 2012 | Recurrent Neural State Estimation in Domains with Long-Term Dependencies
Siegmund Düll, Lina Weichbrodt, Alexander Hans, Steffen Udluft |
ESANN | 4 |
| 2011 | Agent self-assessment: Determining policy quality without executionabstractWith the development of data-efficient reinforcement learning (RL) methods, a promising data-driven solution for optimal control of complex technical systems has become available. For the application of RL to a technical system, it is usually required to evaluate a policy before actually applying it to ensure it operates the system safely and within required performance bounds. In benchmark applications one can use the system dynamics directly to measure the policy quality. In real applications, however, this might be too expensive or even impossible. Being unable to evaluate the policy without using the actual system hinders the application of RL to autonomous controllers. As a first step toward agent self-assessment, we deal with discrete MDPs in this paper. We propose to use the value function along with its uncertainty to assess a policy's quality and show that, when dealing with an MDP estimated from observations, the value function itself can be misleading. We address this problem by determining the value function's uncertainty through uncertainty propagation and evaluate the approach using a number of benchmark applications. Alexander Hans, Siegmund Düll, Steffen Udluft |
ADPRL | 3 |
| 2011 | Ensemble Usage for More Reliable Policy Identification in Reinforcement Learning
Alexander Hans, Steffen Udluft |
ESANN | 2 |
| 2010 | Uncertainty Propagation for Efficient Exploration in Reinforcement Learning
Alexander Hans, Steffen Udluft |
ECAI | 2 |
| 2010 | The Markov Decision Process Extraction Network
Siegmund Düll, Alexander Hans, Steffen Udluft |
ESANN | 3 |
| 2010 | Ensembles of Neural Networks for Robust Reinforcement LearningabstractReinforcement learning algorithms that employ neural networks as function approximators have proven to be powerful tools for solving optimal control problems. However, their training and the validation of final policies can be cumbersome as neural networks can suffer from problems like local minima or over fitting. When using iterative methods, such as neural fitted Q-iteration, the problem becomes even more pronounced since the network has to be trained multiple times and the training process in one iteration builds on the network trained in the previous iteration. Therefore errors can accumulate. In this paper we propose to use ensembles of networks to make the learning process more robust and produce near-optimal policies more reliably. We name various ways of combining single networks to an ensemble that results in a final ensemble policy and show the potential of the approach using a benchmark application. Our experiments indicate that majority voting is superior to Q-averaging and using heterogeneous ensembles (different network topologies) is advisable. Alexander Hans, Steffen Udluft |
ICMLA | 2 |
| 2009 | Efficient Uncertainty Propagation for Reinforcement Learning with Limited Data
Alexander Hans, Steffen Udluft |
ICANN (1) | 2 |
| 2008 | Safe exploration for reinforcement learning
Alexander Hans, Daniel Schneegaß, Anton Maximilian Schäfer, Steffen Udluft |
ESANN | 4 |
| 2008 | Uncertainty propagation for quality assurance in Reinforcement LearningabstractIn this paper we address the reliability of policies derived by Reinforcement Learning on a limited amount of observations. This can be done in a principled manner by taking into account the derived Q-functionpsilas uncertainty, which stems from the uncertainty of the estimators used for the MDPpsilas transition probabilities and the reward function. We apply uncertainty propagation parallelly to the Bellman iteration and achieve confidence intervals for the Q-function. In a second step we change the Bellman operator as to achieve a policy guaranteeing the highest minimum performance with a given probability. We demonstrate the functionality of our method on artificial examples and show that, for an important problem class even an enhancement of the expected performance can be obtained. Finally we verify this observation on an application to gas turbine control. Daniel Schneegaß, Steffen Udluft, Thomas Martinetz |
IJCNN | 2 |
| 2008 | Learning long-term dependencies with recurrent neural networks
Anton Maximilian Schäfer, Steffen Udluft, Hans-Georg Zimmermann |
Neurocomputing | 2 |
| 2007 | The Recurrent Control Neural Network
Anton Maximilian Schäfer, Steffen Udluft, Hans-Georg Zimmermann |
ESANN | 2 |
| 2007 | Neural Rewards Regression for near-optimal policy identification in Markovian and partial observable environments
Daniel Schneegaß, Steffen Udluft, Thomas Martinetz |
ESANN | 2 |
| 2007 | Explicit Kernel Rewards Regression for data-efficient near-optimal policy identification
Daniel Schneegaß, Steffen Udluft, Thomas Martinetz |
ESANN | 2 |
| 2007 | Improving Optimality of Neural Rewards Regression for Data-Efficient Batch Near-Optimal Policy Identification
Daniel Schneegaß, Steffen Udluft, Thomas Martinetz |
ICANN (1) | 2 |
| 2007 | A Neural Reinforcement Learning Approach to Gas Turbine ControlabstractIn this paper a new neural network based approach to control a gas turbine for stable operation on high load is presented. A combination of recurrent neural networks (RNN) and reinforcement learning (RL) is used. The authors start by applying an RNN to identify the minimal state space of a gas turbine's dynamics. Based on this the optimal control policy is determined by standard RL methods. The authors proceed to the recurrent control neural network, which combines these two steps into one integrated neural network. This approach has the advantage that by using neural networks one can easily deal with the high dimensions of a gas turbine. Due to the high system-identification quality of RNN one can further cope with the only limited amount of available data. The proposed methods are demonstrated on an exemplary gas turbine model where, compared to standard controllers, it strongly improves the performance. Anton Maximilian Schäfer, Daniel Schneegaß, Volkmar Sterzing, Steffen Udluft |
IJCNN | 4 |
| 2006 | Learning Long Term Dependencies with Recurrent Neural Networks
Anton Maximilian Schäfer, Steffen Udluft, Hans-Georg Zimmermann |
ICANN (1) | 2 |