Martin A. Riedmiller

dblp:r/MartinARiedmiller · DBLP profile ↗
← Back
79ranked-venue papers
10as first author
11since 2021 · last 2025
0000-0002-8465-5690ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 72 · 7 first-author · 11 since 2021Systems, architecture and hardware · 8 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Learning from negative feedback, or positive feedback or both
abstract
Existing preference optimization methods often assume scenarios where paired preference feedback (preferred/positive vs. dis-preferred/negative examples) is available. This requirement limits their applicability in scenarios where only unpaired feedback—for example, either positive or negative— is available. To address this, we introduce a novel approach that decouples learning from positive and negative feedback. This decoupling enables control over the influence of each feedback type and, importantly, allows learning even when only one feedback type is present. A key contribution is demonstrating stable learning from negative feedback alone, a capability not well-addressed by current methods. Our approach builds upon the probabilistic framework introduced in (Dayan and Hinton, 1997), which uses expectation-maximization (EM) to directly optimize the probability of positive outcomes (as opposed to classic expected reward maximization). We address a key limitation in current EM-based methods: they solely maximize the likelihood of positive examples, while neglecting negative ones. We show how to extend EM algorithms to explicitly incorporate negative examples, leading to a theoretically grounded algorithm that offers an intuitive and versatile way to learn from both positive and negative feedback. We evaluate our approach for training language models based on human feedback as well as training policies for sequential decision-making problems, where learned value functions are available.
Abbas Abdolmaleki, Bilal Piot, Bobak Shahriari, Jost Tobias Springenberg, Tim Hertweck, Michael Bloesch, Rishabh Joshi, Thomas Lampe, Junhyuk Oh, Nicolas Heess, Jonas Buchli, Martin A. Riedmiller
ICLR12
2025 DemoStart: Demonstration-Led Auto-Curriculum Applied to Sim-to-Real with Multi-Fingered Robots
abstract
We present DemoStart, a novel auto-curriculum reinforcement learning method capable of learning complex manipulation behaviors on an arm equipped with a three- fingered robotic hand, from only a sparse reward and a handful of demonstrations in simulation. Learning from simulation drastically reduces the development cycle of behavior generation, and domain randomization techniques are leveraged to achieve successful zero-shot sim-to- real transfer. Transferred policies are learned directly from raw pixels from multiple cameras and robot proprioception. Our approach outperforms policies learned from demonstrations on the real robot and requires 100 times fewer demonstrations, collected in simulation. More details and videos in sites.google.com/view/demostart.
Maria Bauzá 0001, Jose Enriaue Chen, Valentin Dalibard, Nimrod Gileadi, Roland Hafner, Murilo Fernandes Martins, Joss Moore, Rugile Pevceviciute, Antoine Laurens, Dushyant Rao, Martina Zambelli, Martin A. Riedmiller, Jonathan Scholz, Konstantinos Bousmalis, Francesco Nori, Nicolas Heess
ICRA12
2024 Replay across Experiments: A Natural Extension of Off-Policy RL
abstract
Replaying data is a principal mechanism underlying the stability and data efficiency of off-policy reinforcement learning (RL). We present an effective yet simple framework to extend the use of replays across multiple experiments, minimally adapting the RL workflow for sizeable improvements in controller performance and research iteration times. At its core, Replay across Experiments (RaE) involves reusing experience from previous experiments to improve exploration and bootstrap learning while reducing required changes to a minimum in comparison to prior work. We empirically show benefits across a number of RL algorithms and challenging control domains spanning both locomotion and manipulation, including hard exploration tasks from egocentric vision. Through comprehensive ablations, we demonstrate robustness to the quality and amount of data available and various hyperparameter choices. Finally, we discuss how our approach can be applied more broadly across research life cycles and can increase resilience by reloading data across random seeds or hyperparameter variations.
Dhruva Tirumala, Thomas Lampe, José Enrique Chen, Tuomas Haarnoja, Sandy H. Huang, Guy Lever, Ben Moran, Tim Hertweck, Leonard Hasenclever, Martin A. Riedmiller, Nicolas Heess, Markus Wulfmeier
ICLR10
2024 Offline Actor-Critic Reinforcement Learning Scales to Large Models
abstract
We show that offline actor-critic reinforcement learning can scale to large models - such as transformers - and follows similar scaling laws as supervised learning. We find that offline actor-critic algorithms can outperform strong, supervised, behavioral cloning baselines for multi-task training on a large dataset; containing both sub-optimal and expert behavior on 132 continuous control tasks. We introduce a Perceiver-based actor-critic model and elucidate the key features needed to make offline RL work with self- and cross-attention modules. Overall, we find that: i) simple offline actor critic algorithms are a natural choice for gradually moving away from the currently predominant paradigm of behavioral cloning, and ii) via offline RL it is possible to learn multi-task policies that master many domains simultaneously, including real robotics tasks, from sub-optimal demonstrations or self-generated data.
Jost Tobias Springenberg, Abbas Abdolmaleki, Jingwei Zhang 0001, Oliver Groth, Michael Bloesch, Thomas Lampe, Philemon Brakel, Sarah Bechtle, Steven Kapturowski, Roland Hafner, Nicolas Heess, Martin A. Riedmiller
ICML12
2024 Mastering Stacking of Diverse Shapes with Large-Scale Iterative Reinforcement Learning on Real Robots
abstract
Reinforcement learning solely from an agent’s self-generated data is often believed to be infeasible for learning on real robots, due to the amount of data needed. However, if done right, agents learning from real data can be surprisingly efficient through re-using previously collected sub-optimal data. In this paper we demonstrate how the increased understanding of off-policy learning methods and their embedding in an iterative online/offline scheme ("collect and infer") can drastically improve data-efficiency by using all the collected experience, which empowers learning from real robot experience only. Moreover, the resulting policy improves significantly over the state of the art on a recently proposed real robot manipulation benchmark. Our approach learns end-to-end, directly from pixels, and does not rely on additional human domain knowledge such as a simulator or demonstrations.
Thomas Lampe, Abbas Abdolmaleki, Sarah Bechtle, Sandy H. Huang, Jost Tobias Springenberg, Michael Bloesch, Oliver Groth, Roland Hafner, Tim Hertweck, Michael Neunert, Markus Wulfmeier, Jingwei Zhang 0001, Francesco Nori, Nicolas Heess, Martin A. Riedmiller
ICRA15
2024 Imitating Language via Scalable Inverse Reinforcement Learning
abstract
The majority of language model training builds on imitation learning. It covers pretraining, supervised fine-tuning, and affects the starting conditions for reinforcement learning from human feedback (RLHF). The simplicity and scalability of maximum likelihood estimation (MLE) for next token prediction led to its role as predominant paradigm. However, the broader field of imitation learning can more effectively utilize the sequential structure underlying autoregressive generation. We focus on investigating the inverse reinforcement learning (IRL) perspective to imitation, extracting rewards and directly optimizing sequences instead of individual token likelihoods and evaluate its benefits for fine-tuning large language models. We provide a new angle, reformulating inverse soft-Q-learning as a temporal difference regularized extension of MLE. This creates a principled connection between MLE and IRL and allows trading off added complexity with increased performance and diversity of generations in the supervised fine-tuning (SFT) setting. We find clear advantages for IRL-based imitation, in particular for retaining diversity while maximizing task performance, rendering IRL a strong alternative on fixed SFT datasets even without online data generation. Our analysis of IRL-extracted reward functions further indicates benefits for more robust reward functions via tighter integration of supervised and preference-based LLM post-training.
Markus Wulfmeier, Michael Bloesch, Nino Vieillard, Arun Ahuja, Jörg Bornschein, Sandy H. Huang, Artem Sokolov 0001, Matt Barnes 0001, Guillaume Desjardins, Alex Bewley, Sarah Bechtle, Jost Tobias Springenberg, Nikola Momchev, Olivier Bachem, Matthieu Geist, Martin A. Riedmiller
NeurIPS16
2023 Solving Continuous Control via Q-learning
Tim Seyde, Peter Werner, Wilko Schwarting, Igor Gilitschenski, Martin A. Riedmiller, Daniela Rus, Markus Wulfmeier
ICLR5
2022 Evaluating Model-Based Planning and Planner Amortization for Continuous Control
Arunkumar Byravan, Leonard Hasenclever, Piotr Trochim, Mehdi Mirza, Alessandro Davide Ialongo, Yuval Tassa, Jost Tobias Springenberg, Abbas Abdolmaleki, Nicolas Heess, Josh Merel, Martin A. Riedmiller
ICLR11
2021 Data-efficient Hindsight Off-policy Option Learning
abstract
We introduce Hindsight Off-policy Options (HO2), a data-efficient option learning algorithm. Given any trajectory, HO2 infers likely option choices and backpropagates through the dynamic programming inference procedure to robustly train all policy components off-policy and end-to-end. The approach outperforms existing option learning methods on common benchmarks. To better understand the option framework and disentangle benefits from both temporal and action abstraction, we evaluate ablations with flat policies and mixture policies with comparable optimization. The results highlight the importance of both types of abstraction as well as off-policy training and trust-region constraints, particularly in challenging, simulated 3D robot manipulation tasks from raw pixel inputs. Finally, we intuitively adapt the inference step to investigate the effect of increased temporal abstraction on training with pre-trained options and from scratch.
Markus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe, Abbas Abdolmaleki, Tim Hertweck, Michael Neunert, Dhruva Tirumala, Noah Y. Siegel, Nicolas Heess, Martin A. Riedmiller
ICML11
2021 Representation Matters: Improving Perception and Exploration for Robotics
abstract
Projecting high-dimensional environment observations into lower-dimensional structured representations can considerably improve data-efficiency for reinforcement learning in domains with limited data such as robotics. Can a single generally useful representation be found? In order to answer this question, it is important to understand how the representation will be used by the agent and what properties such a good representation should have. In this paper we systematically evaluate a number of common learnt and hand-engineered representations in the context of three robotics tasks: lifting, stacking and pushing of 3D blocks. The representations are evaluated in two use-cases: as input to the agent, or as a source of auxiliary tasks. Furthermore, the value of each representation is evaluated in terms of three properties: dimensionality, observability and disentanglement. We can significantly improve performance in both use-cases and demonstrate that some representations can perform commensurate to simulator states as agent inputs. Finally, our results challenge common intuitions by demonstrating that: 1) dimensionality strongly matters for task generation, but is negligible for inputs, 2) observability of task-relevant aspects mostly affects the input representation use-case, and 3) disentanglement leads to better auxiliary tasks, but has only limited benefits for input representations. This work serves as a step towards a more systematic understanding of what makes a good representation for control in robotics, enabling practitioners to make more informed choices for developing new learned or hand-engineered representations.
Markus Wulfmeier, Arunkumar Byravan, Tim Hertweck, Irina Higgins, Tejas Kulkarni, Malcolm Reynolds, Denis Teplyashin, Roland Hafner, Thomas Lampe, Martin A. Riedmiller
ICRA11
2021 Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies
abstract
Reinforcement learning (RL) for continuous control typically employs distributions whose support covers the entire action space. In this work, we investigate the colloquially known phenomenon that trained agents often prefer actions at the boundaries of that space. We draw theoretical connections to the emergence of bang-bang behavior in optimal control, and provide extensive empirical evaluation across a variety of recent RL algorithms. We replace the normal Gaussian by a Bernoulli distribution that solely considers the extremes along each action dimension - a bang-bang controller. Surprisingly, this achieves state-of-the-art performance on several continuous control benchmarks - in contrast to robotic hardware, where energy and maintenance cost affect controller choices. Since exploration, learning, and the final solution are entangled in RL, we provide additional imitation learning experiments to reduce the impact of exploration on our analysis. Finally, we show that our observations generalize to environments that aim to model real-world challenges and evaluate factors to mitigate the emergence of bang-bang solutions. Our findings emphasise challenges for benchmarking continuous control algorithms, particularly in light of potential real-world applications.
Tim Seyde, Igor Gilitschenski, Wilko Schwarting, Bartolomeo Stellato, Martin A. Riedmiller, Markus Wulfmeier, Daniela Rus
NeurIPS5
2020 Robust Reinforcement Learning for Continuous Control with Model Misspecification
Daniel J. Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki, Jost Tobias Springenberg, Jackie Kay, Todd Hester, Timothy A. Mann, Martin A. Riedmiller
ICLR10
2020 Keep Doing What Worked: Behavior Modelling Priors for Offline Reinforcement Learning
Noah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, Martin A. Riedmiller
ICLR9
2020 V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W. Rae, Seb Noury, Arun Ahuja, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Daniel Belov, Martin A. Riedmiller, Matt M. Botvinick
ICLR13
2020 A distributional view on multi-objective policy optimization
abstract
Many real-world problems require trading off multiple competing objectives. However, these objectives are often in different units and/or scales, which can make it challenging for practitioners to express numerical preferences over objectives in their native units. In this paper we propose a novel algorithm for multi-objective reinforcement learning that enables setting desired preferences for objectives in a scale-invariant way. We propose to learn an action distribution for each objective, and we use supervised learning to fit a parametric policy to a combination of these distributions. We demonstrate the effectiveness of our approach on challenging high-dimensional real and simulated robotics tasks, and show that setting different preferences in our framework allows us to trace out the space of nondominated solutions.
Abbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert, H. Francis Song, Martina Zambelli, Murilo Fernandes Martins, Nicolas Heess, Raia Hadsell, Martin A. Riedmiller
ICML10
2019 Adaptive long-term control of biological neural networks with Deep Reinforcement Learning
Jan Wülfing, Sreedhar S. Kumar, Joschka Boedecker, Martin A. Riedmiller, Ulrich Egert
Neurocomputing4
2018 Controlling biological neural networks with deep reinforcement learning
Jan Wülfing, Sreedhar S. Kumar, Joschka Boedecker, Martin A. Riedmiller, Ulrich Egert
ESANN4
2018 Maximum a Posteriori Policy Optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rémi Munos, Nicolas Heess, Martin A. Riedmiller
ICLR (Poster)6
2018 Learning an Embedding Space for Transferable Robot Skills
Karol Hausman, Jost Tobias Springenberg, Ziyu Wang 0001, Nicolas Heess, Martin A. Riedmiller
ICLR (Poster)5
2018 Learning by Playing Solving Sparse Reward Tasks from Scratch
abstract
We propose Scheduled Auxiliary Control (SAC-X), a new learning paradigm in the context of Reinforcement Learning (RL). SAC-X enables learning of complex behaviors - from scratch - in the presence of multiple sparse reward signals. To this end, the agent is equipped with a set of general auxiliary tasks, that it attempts to learn simultaneously via off-policy RL. The key idea behind our method is that active (learned) scheduling and execution of auxiliary policies allows the agent to efficiently explore its environment - enabling it to excel at sparse reward RL. Our experiments in several challenging robotic manipulation settings demonstrate the power of our approach.
Martin A. Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Van de Wiele, Volodymyr Mnih, Nicolas Heess, Jost Tobias Springenberg
ICML1
2018 Graph Networks as Learnable Physics Engines for Inference and Control
abstract
Understanding and interacting with everyday physical scenes requires rich knowledge about the structure of the world, represented either implicitly in a value or policy function, or explicitly in a transition model. Here we introduce a new class of learnable models–based on graph networks–which implement an inductive bias for object- and relation-centric representations of complex, dynamical systems. Our results show that as a forward model, our approach supports accurate predictions from real and simulated data, and surprisingly strong and efficient generalization, across eight distinct physical systems which we varied parametrically and structurally. We also found that our inference model can perform system identification. Our models are also differentiable, and support online planning via gradient-based trajectory optimization, as well as offline policy optimization. Our framework offers new opportunities for harnessing and exploiting rich knowledge about the world, and takes a key step toward building machines with more human-like representations of the world.
Alvaro Sanchez-Gonzalez, Nicolas Heess, Jost Tobias Springenberg, Josh Merel, Martin A. Riedmiller, Raia Hadsell, Peter W. Battaglia
ICML5
2016 Discriminative Unsupervised Feature Learning with Exemplar Convolutional Neural Networks
abstract
Deep convolutional networks have proven to be very successful in learning task specific features that allow for unprecedented performance on various computer vision tasks. Training of such networks follows mostly the supervised learning paradigm, where sufficiently many input-output pairs are required for training. Acquisition of large training sets is one of the key challenges, when approaching a new task. In this paper, we aim for generic feature learning and present an approach for training a convolutional network using only unlabeled data. To this end, we train the network to discriminate between a set of surrogate classes. Each surrogate class is formed by applying a variety of transformations to a randomly sampled 'seed' image patch. In contrast to supervised network training, the resulting feature representation is not class specific. It rather provides robustness to the transformations that have been applied during training. This generic feature representation allows for classification results that outperform the state of the art for unsupervised learning on several popular datasets (STL-10, CIFAR-10, Caltech-101, Caltech-256). While features learned with our approach cannot compete with class specific features from supervised training on a classification task, we show that they are advantageous on geometric matching problems, where they also outperform the SIFT descriptor.
Alexey Dosovitskiy, Philipp Fischer 0001, Jost Tobias Springenberg, Martin A. Riedmiller, Thomas Brox
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Autonomous Optimization of Targeted Stimulation of Neuronal Networks
abstract
Driven by clinical needs and progress in neurotechnology, targeted interaction with neuronal networks is of increasing importance. Yet, the dynamics of interaction between intrinsic ongoing activity in neuronal networks and their response to stimulation is unknown. Nonetheless, electrical stimulation of the brain is increasingly explored as a therapeutic strategy and as a means to artificially inject information into neural circuits. Strategies using regular or event-triggered fixed stimuli discount the influence of ongoing neuronal activity on the stimulation outcome and are therefore not optimal to induce specific responses reliably. Yet, without suitable mechanistic models, it is hardly possible to optimize such interactions, in particular when desired response features are network-dependent and are initially unknown. In this proof-of-principle study, we present an experimental paradigm using reinforcement-learning (RL) to optimize stimulus settings autonomously and evaluate the learned control strategy using phenomenological models. We asked how to (1) capture the interaction of ongoing network activity, electrical stimulation and evoked responses in a quantifiable 'state' to formulate a well-posed control problem, (2) find the optimal state for stimulation, and (3) evaluate the quality of the solution found. Electrical stimulation of generic neuronal networks grown from rat cortical tissue in vitro evoked bursts of action potentials (responses). We show that the dynamic interplay of their magnitudes and the probability to be intercepted by spontaneous events defines a trade-off scenario with a network-specific unique optimal latency maximizing stimulus efficacy. An RL controller was set to find this optimum autonomously. Across networks, stimulation efficacy increased in 90% of the sessions after learning and learned latencies strongly agreed with those predicted from open-loop experiments. Our results show that autonomous techniques can exploit quantitative relationships underlying activity-response interaction in biological neuronal networks to choose optimal actions. Simple phenomenological models can be useful to validate the quality of the resulting controllers.
Sreedhar S. Kumar, Jan Wülfing, Samora Okujeni, Joschka Boedecker, Martin A. Riedmiller, Ulrich Egert
PLoS Comput. Biol.5
2015 Multimodal deep learning for robust RGB-D object recognition
abstract
Robust object recognition is a crucial ingredient of many, if not all, real-world robotics applications. This paper leverages recent progress on Convolutional Neural Networks (CNNs) and proposes a novel RGB-D architecture for object recognition. Our architecture is composed of two separate CNN processing streams - one for each modality - which are consecutively combined with a late fusion network. We focus on learning with imperfect sensor data, a typical problem in real-world robotics tasks. For accurate learning, we introduce a multi-stage training methodology and two crucial ingredients for handling depth data with CNNs. The first, an effective encoding of depth information for CNNs that enables learning without the need for large depth datasets. The second, a data augmentation scheme for robust learning with depth images by corrupting them with realistic noise patterns. We present state-of-the-art results on the RGB-D object dataset [15] and show recognition in challenging RGB-D real-world noisy settings.
Andreas Eitel, Jost Tobias Springenberg, Luciano Spinello, Martin A. Riedmiller, Wolfram Burgard
IROS4
2015 Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images
abstract
We introduce Embed to Control (E2C), a method for model learning and control of non-linear dynamical systems from raw pixel images. E2C consists of a deep generative model, belonging to the family of variational autoencoders, that learns to generate image trajectories from a latent space in which the dynamics is constrained to be locally linear. Our model is derived directly from an optimal control formulation in latent space, supports long-term prediction of image sequences and exhibits strong performance on a variety of complex control problems.
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, Martin A. Riedmiller
NIPS4
2014 Approximate real-time optimal control based on sparse Gaussian process models
abstract
In this paper we present a fully automated approach to (approximate) optimal control of non-linear systems. Our algorithm jointly learns a non-parametric model of the system dynamics - based on Gaussian Process Regression (GPR) - and performs receding horizon control using an adapted iterative LQR formulation. This results in an extremely data-efficient learning algorithm that can operate under real-time constraints. When combined with an exploration strategy based on GPR variance, our algorithm successfully learns to control two benchmark problems in simulation (two-link manipulator, cart-pole) as well as to swing-up and balance a real cart-pole system. For all considered problems learning from scratch, that is without prior knowledge provided by an expert, succeeds in less than 10 episodes of interaction with the system.
Joschka Boedecker, Jost Tobias Springenberg, Jan Wülfing, Martin A. Riedmiller
ADPRL4
2014 Deterministic Policy Gradient Algorithms
abstract
In this paper we consider deterministic policy gradient algorithms for reinforcement learning with continuous actions. The deterministic policy gradient has a particularly appealing form: it is the expected gradient of the action-value function. This simple form means that the deterministic policy gradient can be estimated much more efficiently than the usual stochastic policy gradient. To ensure adequate exploration, we introduce an off-policy actor-critic algorithm that learns a deterministic target policy from an exploratory behaviour policy. Deterministic policy gradient algorithms outperformed their stochastic counterparts in several benchmark problems, particularly in high-dimensional action spaces.
David Silver 0001, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, Martin A. Riedmiller
ICML6
2014 Approximate model-assisted Neural Fitted Q-Iteration
abstract
In this work, we propose an extension to the Neural Fitted Q-Iteration algorithm that utilizes a learned model to generate virtual trajectories which are used for updating the Q-function. Compared to standard NFQ, this combination has the potential to greatly reduce the amount of system interaction required to learn a good policy. At the same time, the approach still maintains the generalization ability of Q-learning. We provide a general formulation for approximate model-assisted fitted Q-learning, and examine the advantages of its neural implementation regarding interaction time and robustness. Its capabilities are illustrated with first results on a benchmark cart-pole regulation task, on which our method turns out to provide more general policies using much less interaction time.
Thomas Lampe, Martin A. Riedmiller
IJCNN2
2014 A brain-computer interface for high-level remote control of an autonomous, reinforcement-learning-based robotic system for reaching and grasping
abstract
We present an Internet-based brain-computer interface (BCI) for controlling an intelligent robotic device with autonomous reinforcement-learning. BCI control was achieved through dry-electrode electroencephalography (EEG) obtained during imaginary movements. Rather than using low-level direct motor control, we employed a high-level control scheme of the robot, acquired via reinforcement learning, to keep the users cognitive load low while allowing control a reaching-grasping task with multiple degrees of freedom. High-level commands were obtained by classification of EEG responses using an artificial neural network approach utilizing time-frequency features and conveyed through an intuitive user interface. The novel ombination of a rapidly operational dry electrode setup, autonomous control and Internet connectivity made it possible to conveniently interface subjects in an EEG laboratory with remote robotic devices in a closed-loop setup with online visual feedback of the robots actions to the subject. The same approach is also suitable to provide home-bound patients with the possibility to control state-of-the-art robotic devices currently confined to a research environment. Thereby, our BCI approach could help severely paralyzed patients by facilitating patient-centered research of new means of communication, mobility and independence.
Thomas Lampe, Lukas Dominique Josef Fiederer, Martin Völker, Alexander Knorr, Martin A. Riedmiller, Tonio Ball
IUI5
2014 Discriminative Unsupervised Feature Learning with Convolutional Neural Networks
Alexey Dosovitskiy, Jost Tobias Springenberg, Martin A. Riedmiller, Thomas Brox
NIPS3
2013 Optimization of Gaussian process hyperparameters using Rprop
Manuel Blum 0002, Martin A. Riedmiller
ESANN2
2013 Acquiring visual servoing reaching and grasping skills using neural reinforcement learning
abstract
In this work we present a reinforcement learning system for autonomous reaching and grasping using visual servoing with a robotic arm. Control is realized in a visual feedback control loop, making it both reactive and robust to noise. The controller is learned from scratch by success or failure without adding information about the task's solution. All of the system's major components are implemented as neural networks. The system is applied to solving a combined reaching and grasping task involving uncertainty directly on a real robotic platform. Its main parts and the conditions for their successful interoperation are described. It will be shown that even with minimal prior knowledge, the system can learn in a short amount of time to reliably perform its task. Furthermore, we describe the control system's ability to react to changes and errors.
Thomas Lampe, Martin A. Riedmiller
IJCNN2
2012 Learn to Swing Up and Balance a Real Pole Based on Raw Visual Input Data
Jan Mattner, Sascha Lange, Martin A. Riedmiller
ICONIP (5)3
2012 Learning Temporal Coherent Features through Life-Time Sparsity
Jost Tobias Springenberg, Martin A. Riedmiller
ICONIP (1)2
2012 A learned feature descriptor for object recognition in RGB-D data
abstract
In this work we address the problem of feature extraction for object recognition in the context of cameras providing RGB and depth information (RGB-D data). We consider this problem in a bag of features like setting and propose a new, learned, local feature descriptor for RGB-D images, the convolutional k-means descriptor. The descriptor is based on recent results from the machine learning community. It automatically learns feature responses in the neighborhood of detected interest points and is able to combine all available information, such as color and depth into one, concise representation. To demonstrate the strength of this approach we show its applicability to different recognition problems. We evaluate the quality of the descriptor on the RGB-D Object Dataset where it is competitive with previously published results and propose an embedding into an image processing pipeline for object recognition and pose estimation.
Manuel Blum 0002, Jost Tobias Springenberg, Jan Wülfing, Martin A. Riedmiller
ICRA4
2012 Autonomous reinforcement learning on raw visual input data in a real world application
abstract
We propose a learning architecture, that is able to do reinforcement learning based on raw visual input data. In contrast to previous approaches, not only the control policy is learned. In order to be successful, the system must also autonomously learn, how to extract relevant information out of a high-dimensional stream of input information, for which the semantics are not provided to the learning system. We give a first proof-of-concept of this novel learning architecture on a challenging benchmark, namely visual control of a racing slot car. The resulting policy, learned only by success or failure, is hardly beaten by an experienced human player.
Sascha Lange, Martin A. Riedmiller, Arne Voigtländer
IJCNN2
2012 Taming the reservoir: Feedforward training for recurrent neural networks
abstract
Recurrent neural networks are successfully used for tasks like time series processing and system identification. Many of the approaches to train these networks, however, are often regarded as too slow, too complicated, or both. Reservoir computing methods like echo state networks or liquid state machines are an alternative to the more traditional approaches. Echo state networks have the appeal that they are simple to train, and that they have shown to be able to produce excellent results for a number of benchmarks and other tasks. One disadvantage of echo state networks, however, is the high variability in their performance due to a randomly connected hidden layer. Ideally, an efficient and more deterministic way to create connections in the hidden layer could be found, with a performance better than randomly connected hidden layers but without excessively iterating over the same training data many times. We present an approach - tamed reservoirs - that makes use of efficient feedforward training methods, and performs better than echo state networks for some time series prediction tasks. Moreover, our approach reduces some of the variability since all recurrent connections in the network are trained.
Oliver Obst, Martin A. Riedmiller
IJCNN2
2011 Improved neural fitted Q iteration applied to a novel computer gaming and learning benchmark
abstract
Neural batch reinforcement learning (RL) algorithms have recently shown to be a powerful tool for model-free reinforcement learning problems. In this paper, we present a novel learning benchmark from the realm of computer games and apply a variant of a neural batch RL algorithm in the scope of this benchmark. Defining the learning problem and appropriately adjusting all relevant parameters is often a tedious task for the researcher who implements and investigates some learning approach. In RL, the suitable choice of the function c of immediate costs is crucial, and, when utilizing multi-layer perceptron neural networks for the purpose of value function approximation, the definition of c must be well aligned with the specific characteristics of this type of function approximator. Determining this alignment is especially tricky, when no a priori knowledge about the task and, hence, about optimal policies is available. To this end, we propose a simple, but effective dynamic scaling heuristic that can be seamlessly integrated into contemporary neural batch RL algorithms. We evaluate the effectiveness of this heuristic in the context of the well-known pole swing-up benchmark as well as in the context of the novel gaming benchmark we are suggesting.
Thomas Gabel, Christian Lutz, Martin A. Riedmiller
ADPRL3
2011 Enhancing the episodic natural actor-critic algorithm by a regularisation term to stabilize learning of control structures
abstract
Incomplete or imprecise models of control systems make it difficult to find an appropriate structure and parameter set for a corresponding control policy. These problems are addressed by reinforcement learning algorithms like policy gradient methods. We describe how to stabilise the policy gradient descent by introducing a regularisation term to enhance the episodic natural actor-critic approach. This allows a more policy independent usage. We used the resulting algorithm to optimise a z-transformed rational function representing the control policy. This representation facilitates simultaneous optimisation of the control structure and its parameters in time space and can be analysed in terms of control theory to predict the control behaviour for arbitrary scenarios. Furthermore we present a solution to the general problem of finding a initial parameter set with the help of a single demonstrated trajectory. The approach is evaluated on a cartpole simulation for demonstrating the expressiveness of the policy. Furthermore, a real soccer robot scenario demonstrates the ability of the proposed approach to deal with real world scenarios.
Andreas Witsch, Roland Reichle, Kurt Geihs, Sascha Lange, Martin A. Riedmiller
ADPRL5
2011 Reinforcement learning in feedback control - Challenges and benchmarks from technical process control
Roland Hafner, Martin A. Riedmiller
Mach. Learn.2
2010 Deep learning of visual control policies
Sascha Lange, Martin A. Riedmiller
ESANN2
2010 Deep auto-encoder neural networks in reinforcement learning
abstract
This paper discusses the effectiveness of deep auto-encoder neural networks in visual reinforcement learning (RL) tasks. We propose a framework for combining the training of deep auto-encoders (for learning compact feature spaces) with recently-proposed batch-mode RL algorithms (for learning policies). An emphasis is put on the data-efficiency of this combination and on studying the properties of the feature spaces automatically constructed by the deep auto-encoders. These feature spaces are empirically shown to adequately resemble existing similarities and spatial relations between observations and allow to learn useful policies. We propose several methods for improving the topology of the feature spaces making use of task-dependent information. Finally, we present first results on successfully learning good control policies directly on synthesized and real images.
Sascha Lange, Martin A. Riedmiller
IJCNN2
2010 On Progress in RoboCup: The Simulation League Showcase
Thomas Gabel, Martin A. Riedmiller
RoboCup2
2009 The Neuro Slot Car Racer: Reinforcement Learning in a Real World Setting
abstract
This paper describes a novel real-world reinforcement learning application: The Neuro Slot Car Racer. In addition to presenting the system and first results based on Neural Fitted Q-Iteration, a standard batch reinforcement learning technique, an extension is proposed that is capable of improving training times and results by allowing for a reduction of samples required for successful training. The Neuralgic Pattern Selection approach achieves this by applying a failure-probability function which emphasizes neuralgic parts of the state space during sampling.
Tim C. Kietzmann, Martin A. Riedmiller
ICMLA2
2008 Learning to dribble on a real robot by success and failure
abstract
Learning directly on real world systems such as autonomous robots is a challenging task, especially if the training signal is given only in terms of success or failure (Reinforcement Learning). However, if successful, the controller has the advantage of being tailored exactly to the system it eventually has to control. Here we describe, how a neural network based RL controller learns the challenging task of ball dribbling directly on our Middle-Size robot. The learned behaviour was actively used throughout the RoboCup world championship tournament 2007 in Atlanta, where we won the first place. This contistutes another important step within our Brainstormers project. The goal of this project is to develop an intelligent control architecture for a soccer playing robot, that is able to learn more and more complex behaviours from scratch.
Martin A. Riedmiller, Roland Hafner, Sascha Lange, Martin Lauer
ICRA1
2008 A Case Study on Improving Defense Behavior in Soccer Simulation 2D: The NeuroHassle Approach
Thomas Gabel, Martin A. Riedmiller, Florian Trost
RoboCup2
2008 Incremental GRLVQ: Learning relevant features for 3D object recognition
Tim C. Kietzmann, Sascha Lange, Martin A. Riedmiller
Neurocomputing3
2007 Safe Q-Learning on Complete History Spaces
Stephan Timmer, Martin A. Riedmiller
ECML2
2007 Reinforcement learning in a nutshell
Verena Heidrich-Meisner, Martin Lauer, Christian Igel, Martin A. Riedmiller
ESANN4
2007 An Analysis of Case-Based Value Function Approximation by Approximating State Transition Graphs
Thomas Gabel, Martin A. Riedmiller
ICCBR2
2007 Neural Reinforcement Learning Controllers for a Real Robot Application
abstract
Accurate and fast control of wheel speeds in the presence of noise and nonlinearities is one of the crucial requirements for building fast mobile robots, as they are required in the MiddleSize League of RoboCup. We will describe, how highly effective speed controllers can be learned from scratch on the real robot directly. The use of our recently developed neural fitted Q iteration scheme allows reinforcement learning of neural controllers with only a limited amount of training data seen. In the described application, less than 5 minutes of interaction with the real robot were sufficient, to learn fast and accurate control to arbitrary target speeds.
Roland Hafner, Martin A. Riedmiller
ICRA2
2006 Reducing policy degradation in neuro-dynamic programming
Thomas Gabel, Martin A. Riedmiller
ESANN2
2006 Appearance-Based Robot Discrimination Using Eigenimages
Sascha Lange, Martin A. Riedmiller
RoboCup2
2005 Neural Fitted Q Iteration - First Experiences with a Data Efficient Neural Reinforcement Learning Method
Martin A. Riedmiller
ECML1
2005 CBR for State Value Function Approximation in Reinforcement Learning
Thomas Gabel, Martin A. Riedmiller
ICCBR2
2005 Calculating the Perfect Match: An Efficient and Accurate Approach for Robot Self-localization
Martin Lauer, Sascha Lange, Martin A. Riedmiller
RoboCup3
2005 Neural reinforcement learning to swing-up and balance a real pole
abstract
This paper proposes a neural network based reinforcement learning controller that is able to learn control policies in a highly data efficient manner. This allows to apply reinforcement learning directly to real plants -neither a transition model nor a simulation model of the plant is needed for training. The only training information provided to the controller are transition experiences collected from interactions with the real plant. By storing these transition experiences explicitly, they can be reconsidered for updating the neural Q-function in every training step. This results in a stable learning process of a neural Q-value function. The algorithm is applied to learn the highly nonlinear and noisy task of swinging-up and balancing a real inverted pendulum. The amount of real time interaction needed to learn a highly effective policy from scratch was less than 14 minutes.
Martin A. Riedmiller
SMC1
2005 Comparing different methods to speed up reinforcement learning in a complex domain
abstract
We introduce a new learning algorithm (semi-DP algorithm) designed for MDPs (Markov decision process) where actions either lead to a deterministic successor state or to the terminal state. The algorithm only needs a finite number of loops to converge exactly to the optimal action-value function. We compare this algorithm and three other methods to speed up or simplify the learning process to ordinary Q-learning in a soccer grid-world. Furthermore, we show that different reward functions can considerably change the convergence time of the learning algorithms even if the optimal policy remains unchanged.
Martin A. Riedmiller, Daniel Withopf
SMC1
2005 Learning policies for abstract state spaces
abstract
Applying Q-learning to multidimensional, real-valued state spaces is time-consuming in most cases. In this article, we deal with the assumption that a coarse partition of the state space is sufficient for learning good or even optimal policies. An algorithm is presented which constructs proper policies for abstract state spaces using an incremental procedure without approximating a Q-function. By combining an approach similar to dynamic programming and a search for policies, we can speed up the learning process. To provide empirical evidence, we use a cart-pole system. Experiments were conducted for a simulated environment as well as for a real plant.
Stephan Timmer, Martin A. Riedmiller
SMC2
2004 Evolution of Computer Vision Subsystems in Robot Navigation and Image Classification Tasks
Sascha Lange, Martin A. Riedmiller
RoboCup2
2004 Fynesse: An architecture for integrating prior knowledge in autonomously learning agents
Ralf Schoknecht, Martin Spott, Martin A. Riedmiller
Soft Comput.3
2003 Learning to Control at Multiple Time Scales
Ralf Schoknecht, Martin A. Riedmiller
ICANN2
2003 The Smaller the Better: Comparison of Two Approaches for Sales Rate Prediction
Martin Lauer, Martin A. Riedmiller, Thomas Ragg, Walter Baum, Michael Wigbers
IDA2
2003 Reinforcement learning on an omnidirectional mobile robot
abstract
With this paper we describe a well suited, scalable problem for reinforcement learning approaches in the field of mobile robots. We show a suitable representation of the problem for a reinforcement approach and present our results with a model based standard algorithm. Two different approximators for the value function are used, a grid based approximator and a neural network based approximator.
Roland Hafner, Martin A. Riedmiller
IROS2
2003 RoboCup: Yesterday, Today, and Tomorrow Workshop of the Executive Committee in Blaubeuren, October 2003
Hans-Dieter Burkhard, Minoru Asada, Andrea Bonarini, Adam Jacoff, Daniele Nardi, Martin A. Riedmiller, Claude Sammut, Elizabeth Sklar, Manuela M. Veloso
RoboCup6
2003 Overview of RoboCup 2003 Competition and Conferences
Enrico Pagello, Emanuele Menegatti, Ansgar Bredenfeld, Thomas Christaller, Adam Jacoff, Martin A. Riedmiller, Alessandro Saffiotti, Takashi Tomoichi
RoboCup8
2003 Reinforcement learning on explicitly specified time scales
Ralf Schoknecht, Martin A. Riedmiller
Neural Comput. Appl.2
2002 Speeding-up Reinforcement Learning with Multi-step Actions
Ralf Schoknecht, Martin A. Riedmiller
ICANN2
2001 Karlsruhe Brainstormers - A Reinforcement Learning Approach to Robotic Soccer
Artur Merke, Martin A. Riedmiller
RoboCup2
2000 An Algorithm for Distributed Reinforcement Learning in Cooperative Multi-Agent Systems
Martin Lauer, Martin A. Riedmiller
ICML2
2000 Learning Situation Dependent Success Rates of Actions in a RoboCup Scenario
Sebastian Buck 0001, Martin A. Riedmiller
PRICAI2
2000 Karlsruhe Brainstormers 2000 Team Description
Martin A. Riedmiller, Artur Merke, David Meier, Andreas Hoffmann 0004, Alex Sinner, Ortwin Thate
RoboCup1
2000 Karlsruhe Brainstormers - A Reinforcement Learning Approach to Robotic Soccer
Martin A. Riedmiller, Artur Merke, David Meier, Andreas Hoffmann 0004, Alex Sinner, Ortwin Thate, R. Ehrmann
RoboCup1
1999 Distributed Value Functions
Jeff G. Schneider, Weng-Keen Wong, Andrew W. Moore 0001, Martin A. Riedmiller
ICML4
1999 A Neural Reinforcement Learning Approach to Learn Local Dispatching Policies in Production Scheduling
Simone C. Riedmiller, Martin A. Riedmiller
IJCAI2
1999 Karlsruhe Brainstormers - Design Principles
Martin A. Riedmiller, Sebastian Buck 0001, Artur Merke, R. Ehrmann, Ortwin Thate, S. Dilger, Alex Sinner, Andreas Hoffmann 0004, Lutz Frommberger
RoboCup1
1999 Concepts and Facilities of a Neural Reinforcement Learning Control Architecture for Technical Process Control
Martin A. Riedmiller
Neural Comput. Appl.1
1997 Application of a self-learning controller with continuous control signals based on the DOE-approach
Martin A. Riedmiller
ESANN1
1996 Fast Network Pruning and Feature Extraction by using the Unit-OBS Algorithm
Achim Stahlberger, Martin A. Riedmiller
NIPS2