Dhruva Tirumala

dblp:190/7697 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 88% Motion planning and robot control · 7% Multi-agent systems · 4%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 44% Memory systems · 44% Processor architecture and microarchitecture · 13%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
hierarchical reinforcement learning
2.542025
EvoControl: Multi-Frequency Bi-Level Control for High-Frequency Continuous Control · ICML 2025
Behavior Priors for Efficient Reinforcement Learning · J. Mach. Learn. Res. 2022
Learning transferable motor skills with hierarchical latent mixture policies · ICLR 2022
Machine learning › Reinforcement learning
off-policy reinforcement learning
1.322024
Replay across Experiments: A Natural Extension of Off-Policy RL · ICLR 2024
Data-efficient Hindsight Off-policy Option Learning · ICML 2021
Machine learning › Reinforcement learning
continuous control
0.912025
EvoControl: Multi-Frequency Bi-Level Control for High-Frequency Continuous Control · ICML 2025
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay
0.812024
Replay across Experiments: A Natural Extension of Off-Policy RL · ICLR 2024
Machine learning › Reinforcement learning
behavioral prior
0.612022
Behavior Priors for Efficient Reinforcement Learning · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery
0.512021
Data-efficient Hindsight Off-policy Option Learning · ICML 2021
Machine learning › Reinforcement learning › policy optimization
maximum a posteriori policy optimization
0.412020
V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020
Machine learning › Reinforcement learning › policy optimization
on-policy optimization
0.412020
V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020
Machine learning › Reinforcement learning
policy optimization
0.412020
V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020
Robotics › Motion planning and robot control › robot control
hierarchical control
0.412019
Hierarchical Visuomotor Control of Humanoids · ICLR (Poster) 2019
Robotics › Motion planning and robot control
humanoid robot control
0.412019
Hierarchical Visuomotor Control of Humanoids · ICLR (Poster) 2019
Knowledge, reasoning and agents › Multi-agent systems
information asymmetry
0.412019
Information asymmetry in KL-regularized RL · ICLR (Poster) 2019
Machine learning › Reinforcement learning › policy optimization
KL-regularized RL
0.412019
Information asymmetry in KL-regularized RL · ICLR (Poster) 2019
Memory systems
cache coherence
0.312017
POSTER: An Architecture and Programming Model for Accelerating Parallel Commutative Computations via Privatization · PPoPP 2017
Parallel and multicore computing
parallel programming models
0.312017
POSTER: An Architecture and Programming Model for Accelerating Parallel Commutative Computations via Privatization · PPoPP 2017
Machine learning › Reinforcement learning
multi-task reinforcement learning
0.212022
Behavior Priors for Efficient Reinforcement Learning · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning
transfer learning in reinforcement learning
0.212022
Learning transferable motor skills with hierarchical latent mixture policies · ICLR 2022
Robotics › Robot manipulation › dexterous manipulation
dexterous and mobile manipulation
0.112021
Data-efficient Hindsight Off-policy Option Learning · ICML 2021
Machine learning › Reinforcement learning › reward design
reward shaping
0.112019
Information asymmetry in KL-regularized RL · ICLR (Poster) 2019
Processor architecture and microarchitecture
multicore design
0.112017
POSTER: An Architecture and Programming Model for Accelerating Parallel Commutative Computations via Privatization · PPoPP 2017

Methods — techniques the papers use, named apart from their topics

variational inference · 1.1proximal policy optimization · 0.9proportional-derivative control · 0.9evolution strategies · 0.9experience replay · 0.8probabilistic modeling · 0.6mutual information objective · 0.6hierarchical latent mixture policies · 0.6hindsight relabeling · 0.5dynamic programming inference · 0.5synchronization elimination · 0.3privatization · 0.3
YearPublicationVenuePosition
2025 EvoControl: Multi-Frequency Bi-Level Control for High-Frequency Continuous Control
abstract
High-frequency control in continuous action and state spaces is essential for practical applications in the physical world. Directly applying end-to-end reinforcement learning to high-frequency control tasks struggles with assigning credit to actions across long temporal horizons, compounded by the difficulty of efficient exploration. The alternative, learning low-frequency policies that guide higher-frequency controllers (e.g., proportional-derivative (PD) controllers), can result in a limited total expressiveness of the combined control system, hindering overall performance. We introduce EvoControl, a novel bi-level policy learning framework for learning both a slow high-level policy (using PPO) and a fast low-level policy (using Evolution Strategies) for solving continuous control tasks. Learning with Evolution Strategies for the lower-policy allows robust learning for long horizons that crucially arise when operating at higher frequencies. This enables EvoControl to learn to control interactions at a high frequency, benefitting from more efficient exploration and credit assignment than direct high-frequency torque control without the need to hand-tune PD parameters. We empirically demonstrate that EvoControl can achieve a higher evaluation reward for continuous-control tasks compared to existing approaches, specifically excelling in tasks where high-frequency control is needed, such as those requiring safety-critical fast reactions.
Samuel Holt, Todor Davchev, Dhruva Tirumala, Ben Moran, Atil Iscen, Antoine Laurens, Erik Frey, Markus Wulfmeier, Francesco Romano, Nicolas Heess
ICML3
2024 Replay across Experiments: A Natural Extension of Off-Policy RL
abstract
Replaying data is a principal mechanism underlying the stability and data efficiency of off-policy reinforcement learning (RL). We present an effective yet simple framework to extend the use of replays across multiple experiments, minimally adapting the RL workflow for sizeable improvements in controller performance and research iteration times. At its core, Replay across Experiments (RaE) involves reusing experience from previous experiments to improve exploration and bootstrap learning while reducing required changes to a minimum in comparison to prior work. We empirically show benefits across a number of RL algorithms and challenging control domains spanning both locomotion and manipulation, including hard exploration tasks from egocentric vision. Through comprehensive ablations, we demonstrate robustness to the quality and amount of data available and various hyperparameter choices. Finally, we discuss how our approach can be applied more broadly across research life cycles and can increase resilience by reloading data across random seeds or hyperparameter variations.
Dhruva Tirumala, Thomas Lampe, José Enrique Chen, Tuomas Haarnoja, Sandy H. Huang, Guy Lever, Ben Moran, Tim Hertweck, Leonard Hasenclever, Martin A. Riedmiller, Nicolas Heess, Markus Wulfmeier
ICLR1
2022 Learning transferable motor skills with hierarchical latent mixture policies
Dushyant Rao, Fereshteh Sadeghi, Leonard Hasenclever, Markus Wulfmeier, Martina Zambelli, Giulia Vezzani, Dhruva Tirumala, Yusuf Aytar, Josh Merel, Nicolas Heess, Raia Hadsell
ICLR7
2022 Behavior Priors for Efficient Reinforcement Learning
abstract
As we deploy reinforcement learning agents to solve increasingly challenging problems, methods that allow us to inject prior knowledge about the structure of the world and effective solution strategies becomes increasingly important. In this work we consider how information and architectural constraints can be combined with ideas from the probabilistic modeling literature to learn behavior priors that capture the common movement and interaction patterns that are shared across a set of related tasks or contexts. For example the day-to day behavior of humans comprises distinctive locomotion and manipulation patterns that recur across many different situations and goals. We discuss how such behavior patterns can be captured using probabilistic trajectory models and how these can be integrated effectively into reinforcement learning schemes, e.g. to facilitate multi-task and transfer learning. We then extend these ideas to latent variable models and consider a formulation to learn hierarchical priors that capture different aspects of the behavior in reusable modules. We discuss how such latent variable formulations connect to related work on hierarchical reinforcement learning (HRL) and mutual information and curiosity based objectives, thereby offering an alternative perspective on existing ideas. We demonstrate the effectiveness of our framework by applying it to a range of simulated continuous control domains, videos of which can be found at the following url: https://sites.google.com/view/behavior-priors.
Dhruva Tirumala, Alexandre Galashov, Hyeonwoo Noh, Leonard Hasenclever, Razvan Pascanu, Jonathan Schwarz, Guillaume Desjardins, Wojciech Czarnecki 0001, Arun Ahuja, Yee Whye Teh, Nicolas Heess
J. Mach. Learn. Res.1
2021 Data-efficient Hindsight Off-policy Option Learning
abstract
We introduce Hindsight Off-policy Options (HO2), a data-efficient option learning algorithm. Given any trajectory, HO2 infers likely option choices and backpropagates through the dynamic programming inference procedure to robustly train all policy components off-policy and end-to-end. The approach outperforms existing option learning methods on common benchmarks. To better understand the option framework and disentangle benefits from both temporal and action abstraction, we evaluate ablations with flat policies and mixture policies with comparable optimization. The results highlight the importance of both types of abstraction as well as off-policy training and trust-region constraints, particularly in challenging, simulated 3D robot manipulation tasks from raw pixel inputs. Finally, we intuitively adapt the inference step to investigate the effect of increased temporal abstraction on training with pre-trained options and from scratch.
Markus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe, Abbas Abdolmaleki, Tim Hertweck, Michael Neunert, Dhruva Tirumala, Noah Y. Siegel, Nicolas Heess, Martin A. Riedmiller
ICML8
2020 V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W. Rae, Seb Noury, Arun Ahuja, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Daniel Belov, Martin A. Riedmiller, Matt M. Botvinick
ICLR10
2019 Information asymmetry in KL-regularized RL
Alexandre Galashov, Siddhant M. Jayakumar, Leonard Hasenclever, Dhruva Tirumala, Jonathan Schwarz, Guillaume Desjardins, Wojciech Czarnecki 0001, Yee Whye Teh, Razvan Pascanu, Nicolas Heess
ICLR (Poster)4
2019 Hierarchical Visuomotor Control of Humanoids
Josh Merel, Arun Ahuja, Saran Tunyasuvunakool, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Greg Wayne
ICLR (Poster)6
2017 Learning to reinforcement learn
Jane X. Wang, Zeb Kurth-Nelson, Hubert Soyer, Joel Z. Leibo, Dhruva Tirumala, Rémi Munos, Charles Blundell, Dharshan Kumaran, Matt M. Botvinick
CogSci5
2017 POSTER: An Architecture and Programming Model for Accelerating Parallel Commutative Computations via Privatization
abstract
Synchronization and data movement are the key impediments to an efficient parallel execution. To ensure that data shared by multiple threads remain consistent, the programmer must use synchronization (e.g., mutex locks) to serialize threads' accesses to data. This limits parallelism because it forces threads to sequentially access shared resources. Additionally, systems use cache coherence to ensure that processors always operate on the most up-to-date version of a value even in the presence of private caches. Coherence protocol implementations cause processors to serialize their accesses to shared data, further limiting parallelism and performance.
Vignesh Balaji, Dhruva Tirumala, Brandon Lucia
PPoPP2