EDBT 2026 Demo / reviewers in the wild / expert
Moritz Werling
dblp:70/2570
· DBLP profile ↗
16ranked-venue papers
3as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 5 since 2021Systems, architecture and hardware · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Test-Driven Inverse Reinforcement Learning Using Scenario-Based TestingabstractAutomated vehicles require carefully designed cost functions, which are challenging to specify due to the complexity of the behavior they need to cover. Inverse reinforcement learning is a principled methodology for deriving cost functions, but it requires high-quality expert demonstrations, which are expensive to obtain. Recently, scenario-based testing has emerged as a promising approach for validation of driving behavior. In this paper, we introduce a novel methodology that circumvents the need for costly expert driving demonstrations by harnessing scenario-based testing. Our Test-Driven Inverse Reinforcement Learning approach leverages Bayesian inference, utilizing the outcomes of scenario tests as observations to infer cost functions. We rigorously evaluate our method on simulated and real-world scenarios and demonstrate its ability to learn cost functions that successfully pass the respective scenario tests. We also show that the learned cost function generalizes well by also passing scenario tests from an unseen validation set and illustrate that few scenario tests are sufficient to learn meaningful cost functions. This innovative framework not only streamlines the cost function specification process but also offers a cost-effective and practical solution for advancing automated driving systems. Johannes Fischer 0007, Moritz Werling, Martin Lauer, Christoph Stiller |
IV | 2 |
| 2022 | Deep Surrogate Q-Learning for Autonomous DrivingabstractOpen challenges for deep reinforcement learning systems are their adaptivity to changing environments and their efficiency w.r.t. computational resources and data. In the application of learning lane-change behavior for autonomous driving, the number of required transitions imposes a bottleneck, since test drivers cannot perform an arbitrary amount of lane changes in the real world. In the off-policy setting, additional information on solving the task can be gained by observing actions from others. While in the classical RL setup this knowledge remains unused, we use other drivers as surrogates to learn the agent's value function more efficiently. We propose Surrogate Q-learning that deals with the aforementioned problems and reduces the required driving time drastically. We further propose an efficient implementation based on a permutation equivariant deep neural network architecture of the Q-function to estimate action-values for a variable number of vehicles in sensor range. We evaluate our method in the open traffic simulator SUMO and learn well performing driving policies on the real highD dataset. Maria Kalweit, Gabriel Kalweit, Moritz Werling, Joschka Boedecker |
ICRA | 3 |
| 2021 | Amortized Q-learning with Model-based Action Proposals for Autonomous Driving on HighwaysabstractWell-established optimization-based methods can guarantee an optimal trajectory for a short optimization horizon, typically no longer than a few seconds. As a result, choosing the optimal trajectory for this short horizon may still result in a sub-optimal long-term solution. At the same time, the resulting short-term trajectories allow for effective, comfortable and provable safe maneuvers in a dynamic traffic environment. In this work, we address the question of how to ensure an optimal long-term driving strategy, while keeping the benefits of classical trajectory planning. We introduce a Reinforcement Learning based approach that coupled with a trajectory planner, learns an optimal long-term decision-making strategy for driving on highways. By online generating locally optimal maneuvers as actions, we balance between the infinite low-level continuous action space, and the limited flexibility of a fixed number of predefined standard lane-change actions. We evaluated our method on realistic scenarios in the open-source traffic simulator SUMO and were able to achieve better performance than the 4 benchmark approaches we compared against, including a random action selecting agent, greedy agent, high-level, discrete actions agent and an IDM-based SUMO-controlled agent. Branka Mirchevska, Maria Hügle, Gabriel Kalweit, Moritz Werling, Joschka Boedecker |
ICRA | 4 |
| 2021 | Sampling-based Inverse Reinforcement Learning Algorithms with Safety ConstraintsabstractPlanning for robotic systems is frequently formulated as an optimization problem. Instead of manually tweaking the parameters of the cost function, they can be learned from human demonstrations by Inverse Reinforcement Learning (IRL). Common IRL approaches employ a maximum entropy trajectory distribution that can be learned with soft reinforcement learning, where the reward maximization is regularized with an entropy objective. The consideration of safety constraints is of paramount importance for human-robot collaboration. For this reason, our work addresses maximum entropy IRL in constrained environments. Our contribution to this research area is threefold: (1) We propose Constrained Soft Reinforcement Learning (CSRL), an extension of soft reinforcement learning to Constrained Markov Decision Processes (CMDPs). (2) We transfer maximum entropy IRL to CMDPs based on CSRL. (3) We show that using importance sampling in maximum entropy IRL in constrained environments introduces a bias and fails to achieve feature matching. In our evaluation we consider the tactical lane change decision of an autonomous vehicle in a highway scenario modeled in the SUMO traffic simulation. Johannes Fischer 0007, Christoph Eyberg, Moritz Werling, Martin Lauer |
IROS | 3 |
| 2021 | Q-learning with Long-term Action-space Shaping to Model Complex Behavior for Autonomous Lane ChangesabstractIn autonomous driving applications, reinforcement learning agents often have to perform complex behavior, which can translate into optimizing multiple objectives while following certain rules. Encoding traffic rules and desires such as safety and comfort via classical methods based on reward shaping (i.e. a weighted combination of different objectives in the reward signal) or Lagrangian methods (including auxiliary losses in the optimization) can be very hard and cumbersome. In this work, we propose to instead shape the action-space at the maximization step of Q-learning. We further introduce a formulation for fixed-horizon estimation of auxiliary costs under the current target-policy based on truncated value- functions to encode the desire of comfortable driving ensuring interpretable behavior. We compare our algorithm to reward shaping and Lagrangian methods in the application of high- level decision making in autonomous driving, considering rules for safety, keeping right and comfort. We train and evaluate our agent in the open-source simulator SUMO on a variety of scenarios with different driver types and traffic situations. Additionally, we apply our method on the real HighD data set, showing the real-world applicability and simplicity of Q- learning with Action-space Shaping. Gabriel Kalweit, Maria Hügle, Moritz Werling, Joschka Boedecker |
IROS | 3 |
| 2020 | Dynamic Interaction-Aware Scene Understanding for Reinforcement Learning in Autonomous DrivingabstractThe common pipeline in autonomous driving systems is highly modular and includes a perception component which extracts lists of surrounding objects and passes these lists to a high-level decision component. In this case, leveraging the benefits of deep reinforcement learning for high-level decision making requires special architectures to deal with multiple variable-length sequences of different object types, such as vehicles, lanes or traffic signs. At the same time, the architecture has to be able to cover interactions between traffic participants in order to find the optimal action to be taken. In this work, we propose the novel Deep Scenes architecture, that can learn complex interaction-aware scene representations based on extensions of either 1) Deep Sets or 2) Graph Convolutional Networks. We present the Graph-Q and DeepScene-Q off-policy reinforcement learning algorithms, both outperforming state-ofthe-art methods in evaluations with the publicly available traffic simulator SUMO. Maria Hügle, Gabriel Kalweit, Moritz Werling, Joschka Boedecker |
ICRA | 3 |
| 2020 | Deep Inverse Q-learning with ConstraintsabstractPopular Maximum Entropy Inverse Reinforcement Learning approaches require the computation of expected state visitation frequencies for the optimal policy under an estimate of the reward function. This usually requires intermediate value estimation in the inner loop of the algorithm, slowing down convergence considerably. In this work, we introduce a novel class of algorithms that only needs to solve the MDP underlying the demonstrated behavior once to recover the expert policy. This is possible through a formulation that exploits a probabilistic behavior assumption for the demonstrations within the structure of Q-learning. We propose Inverse Action-value Iteration which is able to fully recover an underlying reward of an external agent in closed-form analytically. We further provide an accompanying class of sampling-based variants which do not depend on a model of the environment. We show how to extend this class of algorithms to continuous state-spaces via function approximation and how to estimate a corresponding action-value function, leading to a policy as close as possible to the policy of the external agent, while optionally satisfying a list of predefined hard constraints. We evaluate the resulting algorithms called Inverse Action-value Iteration, Inverse Q-learning and Deep Inverse Q-learning on the Objectworld benchmark, showing a speedup of up to several orders of magnitude compared to (Deep) Max-Entropy algorithms. We further apply Deep Constrained Inverse Q-learning on the task of learning autonomous lane-changes in the open-source simulator SUMO achieving competent driving after training on data corresponding to 30 minutes of demonstrations. Gabriel Kalweit, Maria Hügle, Moritz Werling, Joschka Boedecker |
NeurIPS | 3 |
| 2019 | Dynamic Input for Deep Reinforcement Learning in Autonomous DrivingabstractIn many real-world decision making problems, reaching an optimal decision requires taking into account a variable number of objects around the agent. Autonomous driving is a domain in which this is especially relevant, since the number of cars surrounding the agent varies considerably over time and affects the optimal action to be taken. Classical methods that process object lists can deal with this requirement. However, to take advantage of recent high-performing methods based on deep reinforcement learning in modular pipelines, special architectures are necessary. For these, a number of options exist, but a thorough comparison of the different possibilities is missing. In this paper, we elaborate limitations of fully-connected neural networks and other established approaches like convolutional and recurrent neural networks in the context of reinforcement learning problems that have to deal with variable sized inputs. We employ the structure of Deep Sets in off-policy reinforcement learning for high-level decision making, highlight their capabilities to alleviate these limitations, and show that Deep Sets not only yield the best overall performance but also offer better generalization to unseen situations than the other approaches. Maria Hügle, Gabriel Kalweit, Branka Mirchevska, Moritz Werling, Joschka Boedecker |
IROS | 4 |
| 2017 | Estimation of collective maneuvers through cooperative multi-agent planningabstractIn order to determine a cooperative driving strategy, it is beneficial for an autonomous vehicle to incorporate the intended motion of surrounding vehicles within its own motion planning. However, as intentions cannot be measured directly and the motion of multiple vehicles often are highly interdependent, this incorporation has proven challenging. In this paper, the problem of maneuver estimation is addressed, focusing on situations with close interaction between traffic participants. Therefore, we define collective maneuvers based on trajectory homotopy, describing the relative motion of multiple vehicles in a scene. Representing maneuvers by sample trajectories, maneuver-dependent prediction models of the vehicle states can be defined. This allows for a Bayesian estimation of maneuver probabilities given observations of the real motion. The approach is evaluated by simulation in overtaking scenarios with oncoming traffic and merging scenarios at an intersection. Jens Schulz, Kira Hirsenkorn, Julian Löchner, Moritz Werling, Darius Burschka |
Intelligent Vehicles Symposium | 4 |
| 2017 | Lateral Vehicle Trajectory Optimization Using Constrained Linear Time-Varying MPCabstractIn this paper, a trajectory optimization algorithm is proposed, which formulates the lateral vehicle guidance task along a reference curve as a constrained optimal control problem. The optimization problem is solved by means of a linear time-varying model predictive control scheme that generates trajectories for path following under consideration of various time-varying system constraints in a receding horizon fashion. Formulating the system dynamics linearly in combination with a quadratic cost function has two great advantages. First, the system constraints can be set up not only to achieve collision avoidance with both static and dynamic obstacles, but also aspects of human driving behavior can be considered. Second, the optimization problem can be solved very efficiently, such that the algorithm can be run with little computational effort. In addition, due to an elaborate problem formulation, reference curves with discontinuous, high curvatures will be effortlessly smoothed out by the algorithm. This makes the proposed algorithm applicable to different traffic scenarios, such as parking or highway driving. Experimental results are presented for different real-world scenarios to demonstrate the algorithm's abilities. Benjamin Gutjahr, Lutz Gröll, Moritz Werling |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2014 | Automatic collision avoidance during parking and maneuvering - An optimal control approachabstractIn order to reduce the great number of parking incidences and other collisions in low speed scenarios, an obstacle avoidance algorithm is proposed. Since collision avoidance can be achieved by sole braking when driving slowly this algorithm's objective is a comfort orientated braking routine. Therefore, an optimization problem is formulated leading to jerk and time optimal trajectories for braking. In addition to that, based on the prediction of the future vehicle motion, this algorithm compensates for a significant actuator time delay. Using an occupancy grid for representing the static vehicle environment and the current driving state, possible collision points are determined not only for the vehicle front or rear, but also for both vehicle sides, where no sensors are located. The algorithm's performance is demonstrated in a real-world scenario. Benjamin Gutjahr, Moritz Werling |
Intelligent Vehicles Symposium | 2 |
| 2014 | Reversing the General One-Trailer System: Asymptotic Curvature Stabilization and Path TrackingabstractBacking up a trailer can be a challenge, particularly for inexperienced recreational drivers. We therefore develop two feedback controllers, which support the driver with automatic steering inputs in various situations. Based on the kinematics of the general one-trailer system, we first derive an input/output-linearizing control law that asymptotically stabilizes a given curvature for the trailer. This enables the driver to directly steer the trailer, e.g., by means of a turning knob, such that the trailer will automatically be prevented from jackknifing. The control task is then modified and solved so that the vehicle can also take over the complete stabilization task along given paths. In combination with a path-planning algorithm, this enables automated parallel parking for example. The complete system is implemented on a rapid-prototyping environment and evaluated in real-world scenarios. Moritz Werling, Philipp Reinisch, Michael Heidingsfeld, Klaus Gresser |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2012 | Lane-based safety assessment of road scenes using Inevitable Collision StatesabstractThis paper presents a method for reasoning about the safety of traffic situations. More precisely, the problem of safety assessment for partial trajectories for vehicles is addressed. Therefore, the Inevitable Collision States (ICS) as well as its probabilistic generalization the Probabilistic Collision States (PCS) are used. Thereby, the assessment is performed for an infinite time horizon. For solving the ICS computation nonlinear programming is applied. In addition to the safety assessment an evaluation of the disturbance of the other traffic participants by the ego vehicle is presented. The results are integrated into an optimal control based planning approach that generates minimum jerk trajectories. An example implementation of the proposed framework is applied to simulation scenarios that demonstrates the necessity of the presented method for guaranteeing motion safety. Daniel Althoff, Moritz Werling, Nico Kaempchen, Dirk Wollherr, Martin Buss |
Intelligent Vehicles Symposium | 2 |
| 2011 | Towards fully autonomous driving: Systems and algorithmsabstractIn order to achieve autonomous operation of a vehicle in urban situations with unpredictable traffic, several realtime systems must interoperate, including environment perception, localization, planning, and control. In addition, a robust vehicle platform with appropriate sensors, computational hardware, networking, and software infrastructure is essential. We previously published an overview of Junior, Stanford's entry in the 2007 DARPA Urban Challenge. This race was a closed-course competition which, while historic and inciting much progress in the field, was not fully representative of the situations that exist in the real world. In this paper, we present a summary of our recent research towards the goal of enabling safe and robust autonomous operation in more realistic situations. First, a trio of unsupervised algorithms automatically calibrates our 64-beam rotating LIDAR with accuracy superior to tedious hand measurements. We then generate high-resolution maps of the environment which are subsequently used for online localization with centimeter accuracy. Improved perception and recognition algorithms now enable Junior to track and classify obstacles as cyclists, pedestrians, and vehicles; traffic lights are detected as well. A new planning system uses this incoming data to generate thousands of candidate trajectories per second, choosing the optimal path dynamically. The improved controller continuously selects throttle, brake, and steering actuations that maximize comfort and minimize trajectory error. All of these algorithms work in sun or rain and during the day or night. With these systems operating together, Junior has successfully logged hundreds of miles of autonomous operation in a variety of real-life conditions. Jesse Levinson, Jake Askeland, Jan Becker, Jennifer Dolson, David Held, Sören Kammel, J. Zico Kolter, Dirk Langer, Oliver Pink, Vaughan R. Pratt, Michael Sokolsky, Ganymed Stanek, David Stavens, Alex Teichman, Moritz Werling, Sebastian Thrun |
Intelligent Vehicles Symposium | 15 |
| 2010 | Optimal trajectory generation for dynamic street scenarios in a Frenét FrameabstractSafe handling of dynamic highway and inner city scenarios with autonomous vehicles involves the problem of generating traffic-adapted trajectories. In order to account for the practical requirements of the holistic autonomous system, we propose a semi-reactive trajectory generation method, which can be tightly integrated into the behavioral layer. The method realizes long-term objectives such as velocity keeping, merging, following, stopping, in combination with a reactive collision avoidance by means of optimal-control strategies within the Frenét-Frame of the street. The capabilities of this approach are demonstrated in the simulation of a typical high-speed highway scenario. Moritz Werling, Julius Ziegler, Sören Kammel, Sebastian Thrun |
ICRA | 1 |
| 2010 | Invariant Trajectory Tracking With a Full-Size Autonomous Road VehicleabstractSafe handling of dynamic inner-city scenarios with autonomous road vehicles involves the problem of stabilization of precalculated state trajectories. In order to account for the practical requirements of the holistic autonomous system, we propose two complementary nonlinear Lyapunov-based tracking-control laws to solve the problem for speeds between ±6 m/s. Their designs are based on an extended kinematic one-track model, and they provide a smooth, singularity-free stopping transient. With regard to autonomous test applications, the proposed tracking law without orientation control performs much better with respect to control effort and steering-input saturation than the one with orientation control but needs to be prudently combined with the latter for backward driving. The controller performance is illustrated with a full-size test vehicle. Moritz Werling, Lutz Gröll, Georg Bretthauer |
IEEE Trans. Robotics | 1 |