Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tuomas Haarnoja

dblp:80/9963 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
2since 2021 · last 2024
0009-0007-2973-9246ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 2 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Reinforcement learning · 62% Legged, aerial and field robots · 10% Motion planning and robot control · 8%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
off-policy reinforcement learning
1.122024
Replay across Experiments: A Natural Extension of Off-Policy RL · ICLR 2024
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor · ICML 2018
Machine learning › Reinforcement learning
maximum entropy reinforcement learning
0.932018
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor · ICML 2018
Latent Space Policies for Hierarchical Reinforcement Learning · ICML 2018
Reinforcement Learning with Deep Energy-Based Policies · ICML 2017
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay
0.812024
Replay across Experiments: A Natural Extension of Off-Policy RL · ICLR 2024
Robotics › Legged, aerial and field robots › legged robots › legged robot locomotion
bipedal locomotion
0.712023
NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields · ICRA 2023
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.712023
NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields · ICRA 2023
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
skill discovery
0.412020
Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery · ICLR 2020
Machine learning › Reinforcement learning
actor-critic methods
0.312018
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor · ICML 2018
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.312018
Latent Space Policies for Hierarchical Reinforcement Learning · ICML 2018
Robotics › Motion planning and robot control › robot learning › robotic reinforcement learning
reinforcement learning for manipulation
0.312018
Composable Deep Reinforcement Learning for Robotic Manipulation · ICRA 2018
Robotics › Motion planning and robot control
robot learning
0.312018
Composable Deep Reinforcement Learning for Robotic Manipulation · ICRA 2018
Machine learning › Reinforcement learning › actor-critic methods
soft actor-critic
0.312018
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor · ICML 2018
Machine learning › Reinforcement learning › value-based reinforcement learning
soft q-learning
0.312017
Reinforcement Learning with Deep Energy-Based Policies · ICML 2017
Machine learning › Deep learning architectures and training
recurrent neural network
0.212016
Backprop KF: Learning Discriminative Deterministic State Estimators · NIPS 2016
Robotics › Robot navigation and mapping
state estimation
0.212016
Backprop KF: Learning Discriminative Deterministic State Estimators · NIPS 2016
Computer vision › 3D vision
neural radiance field
0.212023
NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields · ICRA 2023
Computer vision › 3D vision
novel view synthesis
0.212023
NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields · ICRA 2023
Robotics › Legged, aerial and field robots › bipedal robot
bipedal locomotion control
0.112011
Assessment of limit-cycle-based control on 2D kneed biped · ICRA 2011
Robotics › Legged, aerial and field robots
legged robots
0.112011
Assessment of limit-cycle-based control on 2D kneed biped · ICRA 2011
Robotics › Robot navigation and mapping
visual odometry
0.112016
Backprop KF: Learning Discriminative Deterministic State Estimators · NIPS 2016
Robotics › Legged, aerial and field robots › dynamic walking
limit cycle walking
0.012011
Assessment of limit-cycle-based control on 2D kneed biped · ICRA 2011
Robotics › Motion planning and robot control
robot control
0.012011
Assessment of limit-cycle-based control on 2D kneed biped · ICRA 2011

Methods — techniques the papers use, named apart from their topics

experience replay · 0.8sim2real transfer · 0.7physics simulation · 0.7neural radiance field · 0.7maximum entropy objective · 0.7skill discovery · 0.4dynamical distance learning · 0.4off-policy update · 0.3latent random variable · 0.3energy-based model · 0.3
YearPublicationVenuePosition
2024 Replay across Experiments: A Natural Extension of Off-Policy RL
abstract
Replaying data is a principal mechanism underlying the stability and data efficiency of off-policy reinforcement learning (RL). We present an effective yet simple framework to extend the use of replays across multiple experiments, minimally adapting the RL workflow for sizeable improvements in controller performance and research iteration times. At its core, Replay across Experiments (RaE) involves reusing experience from previous experiments to improve exploration and bootstrap learning while reducing required changes to a minimum in comparison to prior work. We empirically show benefits across a number of RL algorithms and challenging control domains spanning both locomotion and manipulation, including hard exploration tasks from egocentric vision. Through comprehensive ablations, we demonstrate robustness to the quality and amount of data available and various hyperparameter choices. Finally, we discuss how our approach can be applied more broadly across research life cycles and can increase resilience by reloading data across random seeds or hyperparameter variations.
Dhruva Tirumala, Thomas Lampe, José Enrique Chen, Tuomas Haarnoja, Sandy H. Huang, Guy Lever, Ben Moran, Tim Hertweck, Leonard Hasenclever, Martin A. Riedmiller, Nicolas Heess, Markus Wulfmeier
ICLR4
2023 NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields
abstract
We present a system for applying sim2real approaches to “in the wild” scenes with realistic visuals, and to policies which rely on active perception using RGB cameras. Given a short video of a static scene collected using a generic phone, we learn the scene's contact geometry and a function for novel view synthesis using a Neural Radiance Field (NeRF). We augment the NeRF rendering of the static scene by overlaying the rendering of other dynamic objects (e.g. the robot's own body, a ball). A simulation is then created using the rendering engine in a physics simulator which computes contact dynamics from the static scene geometry (estimated from the NeRF vol-ume density) and the dynamic objects' geometry and physical properties (assumed known). We demonstrate that we can use this simulation to learn vision-based whole body navigation and ball pushing policies for a 20 degree-of-freedom humanoid robot with an actuated head-mounted RGB camera, and we successfully transfer these policies to a real robot.
Arunkumar Byravan, Jan Humplik, Leonard Hasenclever, Arthur Brussee, Francesco Nori, Tuomas Haarnoja, Ben Moran, Steven Bohez, Fereshteh Sadeghi, Bojan Vujatovic, Nicolas Heess
ICRA6
2020 Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery
Kristian Hartikainen, Xinyang Geng, Tuomas Haarnoja, Sergey Levine
ICLR3
2018 Latent Space Policies for Hierarchical Reinforcement Learning
abstract
We address the problem of learning hierarchical deep neural network policies for reinforcement learning. In contrast to methods that explicitly restrict or cripple lower layers of a hierarchy to force them to use higher-level modulating signals, each layer in our framework is trained to directly solve the task, but acquires a range of diverse strategies via a maximum entropy reinforcement learning objective. Each layer is also augmented with latent random variables, which are sampled from a prior distribution during the training of that layer. The maximum entropy objective causes these latent variables to be incorporated into the layer’s policy, and the higher level layer can directly control the behavior of the lower layer through this latent space. Furthermore, by constraining the mapping from latent variables to actions to be invertible, higher layers retain full expressivity: neither the higher layers nor the lower layers are constrained in their behavior. Our experimental evaluation demonstrates that we can improve on the performance of single-layer policies on standard benchmark tasks simply by adding additional layers, and that our method can solve more complex sparse-reward tasks by learning higher-level policies on top of high-entropy skills optimized for simple low-level objectives.
Tuomas Haarnoja, Kristian Hartikainen, Pieter Abbeel, Sergey Levine
ICML1
2018 Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
abstract
Model-free deep reinforcement learning (RL) algorithms have been demonstrated on a range of challenging decision making and control tasks. However, these methods typically suffer from two major challenges: very high sample complexity and brittle convergence properties, which necessitate meticulous hyperparameter tuning. Both of these challenges severely limit the applicability of such methods to complex, real-world domains. In this paper, we propose soft actor-critic, an off-policy actor-critic deep RL algorithm based on the maximum entropy reinforcement learning framework. In this framework, the actor aims to maximize expected reward while also maximizing entropy. That is, to succeed at the task while acting as randomly as possible. Prior deep RL methods based on this framework have been formulated as Q-learning methods. By combining off-policy updates with a stable stochastic actor-critic formulation, our method achieves state-of-the-art performance on a range of continuous control benchmark tasks, outperforming prior on-policy and off-policy methods. Furthermore, we demonstrate that, in contrast to other off-policy algorithms, our approach is very stable, achieving very similar performance across different random seeds.
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine
ICML1
2018 Composable Deep Reinforcement Learning for Robotic Manipulation
abstract
Model-free deep reinforcement learning has been shown to exhibit good performance in domains ranging from video games to simulated robotic manipulation and locomotion. However, model-free methods are known to perform poorly when the interaction time with the environment is limited, as is the case for most real-world robotic tasks. In this paper, we study how maximum entropy policies trained using soft Q-learning can be applied to real-world robotic manipulation. The application of this method to real-world manipulation is facilitated by two important features of soft Q-learning. First, soft Q-learning can learn multimodal exploration strategies by learning policies represented by expressive energy-based models. Second, we show that policies learned with soft Q-learning can be composed to create new policies, and that the optimality of the resulting policy can be bounded in terms of the divergence between the composed policies. This compositionality provides an especially valuable tool for real-world manipulation, where constructing new policies by composing existing skills can provide a large gain in efficiency over training from scratch. Our experimental evaluation demonstrates that soft Q-learning is substantially more sample efficient than prior model-free deep reinforcement learning methods, and that compositionality can be performed for both simulated and real-world tasks.
Tuomas Haarnoja, Vitchyr Pong, Aurick Zhou, Murtaza Dalal, Pieter Abbeel, Sergey Levine
ICRA1
2017 Reinforcement Learning with Deep Energy-Based Policies
abstract
We propose a method for learning expressive energy-based policies for continuous states and actions, which has been feasible only in tabular domains before. We apply our method to learning maximum entropy policies, resulting into a new algorithm, called soft Q-learning, that expresses the optimal policy via a Boltzmann distribution. We use the recently proposed amortized Stein variational gradient descent to learn a stochastic sampling network that approximates samples from this distribution. The benefits of the proposed algorithm include improved exploration and compositionality that allows transferring skills between tasks, which we confirm in simulated experiments with swimming and walking robots. We also draw a connection to actor-critic methods, which can be viewed performing approximate inference on the corresponding energy-based model.
Tuomas Haarnoja, Pieter Abbeel, Sergey Levine
ICML1
2016 Backprop KF: Learning Discriminative Deterministic State Estimators
abstract
Generative state estimators based on probabilistic filters and smoothers are one of the most popular classes of state estimators for robots and autonomous vehicles. However, generative models have limited capacity to handle rich sensory observations, such as camera images, since they must model the entire distribution over sensor readings. Discriminative models do not suffer from this limitation, but are typically more complex to train as latent variable models for state estimation. We present an alternative approach where the parameters of the latent state distribution are directly optimized as a deterministic computation graph, resulting in a simple and effective gradient descent algorithm for training discriminative state estimators. We show that this procedure can be used to train state estimators that use complex input, such as raw camera images, which must be processed using expressive nonlinear function approximators such as convolutional neural networks. Our model can be viewed as a type of recurrent neural network, and the connection to probabilistic filtering allows us to design a network architecture that is particularly well suited for state estimation. We evaluate our approach on synthetic tracking task with raw image inputs and on the visual odometry task in the KITTI dataset. The results show significant improvement over both standard generative approaches and regular recurrent neural networks.
Tuomas Haarnoja, Anurag Ajay, Sergey Levine, Pieter Abbeel
NIPS1
2012 An estimator for the eigenvalues of the system matrix of a periodic-reference LMS algorithm
abstract
The convergence analysis of the Least Mean Square (LMS) algorithm has been conventionally based on stochastic signals and describes thus only the average behavior of the algorithm. It has been shown previously that a periodic-reference LMS system can be regarded as a linear time-periodic system whose stability can be determined from the monodromy matrix. Generally, the monodromy matrix can only be solved numerically and does not thus reveal the actual factors behind the dynamics of the system. This paper derives an estimator for the eigenvalues of the monodromy matrix. The estimator is easy to calculate, and it also reveals the underlying reason for the bad convergence of the LMS algorithm in some special cases. The estimator is confirmed by comparing it to the precise eigenvalues of the monodromy matrix. The estimator is found to be accurate for the eigenvalues close to unity.
Tuomas Haarnoja, Kari Tammi, Kai Zenger
ICASSP1
2011 Assessment of limit-cycle-based control on 2D kneed biped
abstract
This work presents an assessment of different control techniques based on Limit Cycle Walking. The study is performed on a two-dimensional kneed bipedal simulator developed with Open Dynamics Engine, which includes realistic configuration of weight allocation, link length, inertia distribution and impacts. The controllers are model-based driven from an analytic 2D kneed biped. Several controllers from the Computed Torque (CT) family together with a Passivity Based controller (PBC) are compared based on their energy efficiency and robustness. The previous controllers are used in their original form, with minor adjustment to handle hybrid systems such as the kneed biped case. PBC proved not suitable for real implementation however a combined strategy between PBC and a soft PD control has shown promising results. PD Computed Torque control also offered high tolerance to disturbances with reasonable energy consumption. Furthermore, the performance of all the controllers showed high dependency on the parameters of the robot.
Jose-Luis Peralta-Cabezas, Tuomas Haarnoja, Tomi Ylikorpi, Aarne Halme
ICRA2
2011 Model-based velocity control for Limit Cycle Walking
abstract
The aim of this study is to increase the versatility of Limit Cycle Walking (LCW) robots by proposing several ideas to control their velocity. The velocity control methods presented are to be used with Computed-Torque Control (CTC). This controller utilizes a passive reference model walking down a gentle slope which provides the reference trajectory to follow. The controllers were tested on a rigid body simulator. The results show that the methods based on scaling the amplitude (step length) of the reference trajectory and changing the reference slope angle, provide an energy efficient and stable way to control the walking speed. Both of these methods also provide continuous trajectories to standing pose and running gait.
Tuomas Haarnoja, Jose-Luis Peralta-Cabezas, Aarne Halme
IROS1