EDBT 2026 Demo / reviewers in the wild / expert
Dhruva Tirumala
dblp:190/7697
· DBLP profile ↗
10ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 88% Motion planning and robot control · 7% Multi-agent systems · 4% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 44% Memory systems · 44% Processor architecture and microarchitecture · 13% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
2.5 | 4 | 2025 | EvoControl: Multi-Frequency Bi-Level Control for High-Frequency Continuous Control · ICML 2025 Behavior Priors for Efficient Reinforcement Learning · J. Mach. Learn. Res. 2022 Learning transferable motor skills with hierarchical latent mixture policies · ICLR 2022 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
1.3 | 2 | 2024 | Replay across Experiments: A Natural Extension of Off-Policy RL · ICLR 2024 Data-efficient Hindsight Off-policy Option Learning · ICML 2021 |
Machine learning › Reinforcement learning
continuous control |
0.9 | 1 | 2025 | EvoControl: Multi-Frequency Bi-Level Control for High-Frequency Continuous Control · ICML 2025 |
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay |
0.8 | 1 | 2024 | Replay across Experiments: A Natural Extension of Off-Policy RL · ICLR 2024 |
Machine learning › Reinforcement learning
behavioral prior |
0.6 | 1 | 2022 | Behavior Priors for Efficient Reinforcement Learning · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
0.5 | 1 | 2021 | Data-efficient Hindsight Off-policy Option Learning · ICML 2021 |
Machine learning › Reinforcement learning › policy optimization
maximum a posteriori policy optimization |
0.4 | 1 | 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020 |
Machine learning › Reinforcement learning › policy optimization
on-policy optimization |
0.4 | 1 | 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020 |
Machine learning › Reinforcement learning
policy optimization |
0.4 | 1 | 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control · ICLR 2020 |
Robotics › Motion planning and robot control › robot control
hierarchical control |
0.4 | 1 | 2019 | Hierarchical Visuomotor Control of Humanoids · ICLR (Poster) 2019 |
Robotics › Motion planning and robot control
humanoid robot control |
0.4 | 1 | 2019 | Hierarchical Visuomotor Control of Humanoids · ICLR (Poster) 2019 |
Knowledge, reasoning and agents › Multi-agent systems
information asymmetry |
0.4 | 1 | 2019 | Information asymmetry in KL-regularized RL · ICLR (Poster) 2019 |
Machine learning › Reinforcement learning › policy optimization
KL-regularized RL |
0.4 | 1 | 2019 | Information asymmetry in KL-regularized RL · ICLR (Poster) 2019 |
Memory systems
cache coherence |
0.3 | 1 | 2017 | POSTER: An Architecture and Programming Model for Accelerating Parallel Commutative Computations via Privatization · PPoPP 2017 |
Parallel and multicore computing
parallel programming models |
0.3 | 1 | 2017 | POSTER: An Architecture and Programming Model for Accelerating Parallel Commutative Computations via Privatization · PPoPP 2017 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
0.2 | 1 | 2022 | Behavior Priors for Efficient Reinforcement Learning · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning
transfer learning in reinforcement learning |
0.2 | 1 | 2022 | Learning transferable motor skills with hierarchical latent mixture policies · ICLR 2022 |
Robotics › Robot manipulation › dexterous manipulation
dexterous and mobile manipulation |
0.1 | 1 | 2021 | Data-efficient Hindsight Off-policy Option Learning · ICML 2021 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.1 | 1 | 2019 | Information asymmetry in KL-regularized RL · ICLR (Poster) 2019 |
Processor architecture and microarchitecture
multicore design |
0.1 | 1 | 2017 | POSTER: An Architecture and Programming Model for Accelerating Parallel Commutative Computations via Privatization · PPoPP 2017 |
Methods — techniques the papers use, named apart from their topics
variational inference · 1.1proximal policy optimization · 0.9proportional-derivative control · 0.9evolution strategies · 0.9experience replay · 0.8probabilistic modeling · 0.6mutual information objective · 0.6hierarchical latent mixture policies · 0.6hindsight relabeling · 0.5dynamic programming inference · 0.5synchronization elimination · 0.3privatization · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EvoControl: Multi-Frequency Bi-Level Control for High-Frequency Continuous ControlabstractHigh-frequency control in continuous action and state spaces is essential for practical applications in the physical world. Directly applying end-to-end reinforcement learning to high-frequency control tasks struggles with assigning credit to actions across long temporal horizons, compounded by the difficulty of efficient exploration. The alternative, learning low-frequency policies that guide higher-frequency controllers (e.g., proportional-derivative (PD) controllers), can result in a limited total expressiveness of the combined control system, hindering overall performance. We introduce EvoControl, a novel bi-level policy learning framework for learning both a slow high-level policy (using PPO) and a fast low-level policy (using Evolution Strategies) for solving continuous control tasks. Learning with Evolution Strategies for the lower-policy allows robust learning for long horizons that crucially arise when operating at higher frequencies. This enables EvoControl to learn to control interactions at a high frequency, benefitting from more efficient exploration and credit assignment than direct high-frequency torque control without the need to hand-tune PD parameters. We empirically demonstrate that EvoControl can achieve a higher evaluation reward for continuous-control tasks compared to existing approaches, specifically excelling in tasks where high-frequency control is needed, such as those requiring safety-critical fast reactions. Samuel Holt, Todor Davchev, Dhruva Tirumala, Ben Moran, Atil Iscen, Antoine Laurens, Erik Frey, Markus Wulfmeier, Francesco Romano, Nicolas Heess |
ICML | 3 |
| 2024 | Replay across Experiments: A Natural Extension of Off-Policy RLabstractReplaying data is a principal mechanism underlying the stability and data efficiency of off-policy reinforcement learning (RL).
We present an effective yet simple framework to extend the use of replays across multiple experiments, minimally adapting the RL workflow for sizeable improvements in controller performance and research iteration times.
At its core, Replay across Experiments (RaE) involves reusing experience from previous experiments to improve exploration and bootstrap learning while reducing required changes to a minimum in comparison to prior work.
We empirically show benefits across a number of RL algorithms and challenging control domains spanning both locomotion and manipulation, including hard exploration tasks from egocentric vision.
Through comprehensive ablations, we demonstrate robustness to the quality and amount of data available and various hyperparameter choices. Finally, we discuss how our approach can be applied more broadly across research life cycles and can increase resilience by reloading data across random seeds or hyperparameter variations. Dhruva Tirumala, Thomas Lampe, José Enrique Chen, Tuomas Haarnoja, Sandy H. Huang, Guy Lever, Ben Moran, Tim Hertweck, Leonard Hasenclever, Martin A. Riedmiller, Nicolas Heess, Markus Wulfmeier |
ICLR | 1 |
| 2022 | Learning transferable motor skills with hierarchical latent mixture policies
Dushyant Rao, Fereshteh Sadeghi, Leonard Hasenclever, Markus Wulfmeier, Martina Zambelli, Giulia Vezzani, Dhruva Tirumala, Yusuf Aytar, Josh Merel, Nicolas Heess, Raia Hadsell |
ICLR | 7 |
| 2022 | Behavior Priors for Efficient Reinforcement LearningabstractAs we deploy reinforcement learning agents to solve increasingly challenging problems, methods that allow us to inject prior knowledge about the structure of the world and effective solution strategies becomes increasingly important. In this work we consider how information and architectural constraints can be combined with ideas from the probabilistic modeling literature to learn behavior priors that capture the common movement and interaction patterns that are shared across a set of related tasks or contexts. For example the day-to day behavior of humans comprises distinctive locomotion and manipulation patterns that recur across many different situations and goals. We discuss how such behavior patterns can be captured using probabilistic trajectory models and how these can be integrated effectively into reinforcement learning schemes, e.g. to facilitate multi-task and transfer learning. We then extend these ideas to latent variable models and consider a formulation to learn hierarchical priors that capture different aspects of the behavior in reusable modules. We discuss how such latent variable formulations connect to related work on hierarchical reinforcement learning (HRL) and mutual information and curiosity based objectives, thereby offering an alternative perspective on existing ideas. We demonstrate the effectiveness of our framework by applying it to a range of simulated continuous control domains, videos of which can be found at the following url: https://sites.google.com/view/behavior-priors. Dhruva Tirumala, Alexandre Galashov, Hyeonwoo Noh, Leonard Hasenclever, Razvan Pascanu, Jonathan Schwarz, Guillaume Desjardins, Wojciech Czarnecki 0001, Arun Ahuja, Yee Whye Teh, Nicolas Heess |
J. Mach. Learn. Res. | 1 |
| 2021 | Data-efficient Hindsight Off-policy Option LearningabstractWe introduce Hindsight Off-policy Options (HO2), a data-efficient option learning algorithm. Given any trajectory, HO2 infers likely option choices and backpropagates through the dynamic programming inference procedure to robustly train all policy components off-policy and end-to-end. The approach outperforms existing option learning methods on common benchmarks. To better understand the option framework and disentangle benefits from both temporal and action abstraction, we evaluate ablations with flat policies and mixture policies with comparable optimization. The results highlight the importance of both types of abstraction as well as off-policy training and trust-region constraints, particularly in challenging, simulated 3D robot manipulation tasks from raw pixel inputs. Finally, we intuitively adapt the inference step to investigate the effect of increased temporal abstraction on training with pre-trained options and from scratch. Markus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe, Abbas Abdolmaleki, Tim Hertweck, Michael Neunert, Dhruva Tirumala, Noah Y. Siegel, Nicolas Heess, Martin A. Riedmiller |
ICML | 8 |
| 2020 | V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W. Rae, Seb Noury, Arun Ahuja, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Daniel Belov, Martin A. Riedmiller, Matt M. Botvinick |
ICLR | 10 |
| 2019 | Information asymmetry in KL-regularized RL
Alexandre Galashov, Siddhant M. Jayakumar, Leonard Hasenclever, Dhruva Tirumala, Jonathan Schwarz, Guillaume Desjardins, Wojciech Czarnecki 0001, Yee Whye Teh, Razvan Pascanu, Nicolas Heess |
ICLR (Poster) | 4 |
| 2019 | Hierarchical Visuomotor Control of Humanoids
Josh Merel, Arun Ahuja, Saran Tunyasuvunakool, Siqi Liu 0002, Dhruva Tirumala, Nicolas Heess, Greg Wayne |
ICLR (Poster) | 6 |
| 2017 | Learning to reinforcement learn
Jane X. Wang, Zeb Kurth-Nelson, Hubert Soyer, Joel Z. Leibo, Dhruva Tirumala, Rémi Munos, Charles Blundell, Dharshan Kumaran, Matt M. Botvinick |
CogSci | 5 |
| 2017 | POSTER: An Architecture and Programming Model for Accelerating Parallel Commutative Computations via PrivatizationabstractSynchronization and data movement are the key impediments to an efficient parallel execution. To ensure that data shared by multiple threads remain consistent, the programmer must use synchronization (e.g., mutex locks) to serialize threads' accesses to data. This limits parallelism because it forces threads to sequentially access shared resources. Additionally, systems use cache coherence to ensure that processors always operate on the most up-to-date version of a value even in the presence of private caches. Coherence protocol implementations cause processors to serialize their accesses to shared data, further limiting parallelism and performance. Vignesh Balaji, Dhruva Tirumala, Brandon Lucia |
PPoPP | 2 |