VLDB 2026 Research / reviewers in the wild / expert
Akhil Bagaria
dblp:155/9746
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 87% Motion planning and robot control · 4% Graph learning · 4% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
3.6 | 6 | 2025 | Discovering Options That Minimize Average Planning Time · AAAI 2025 Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning · NeurIPS 2023 Scaling Goal-based Exploration via Pruning Proto-goals · IJCAI 2023 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
2.5 | 4 | 2025 | Discovering Options That Minimize Average Planning Time · AAAI 2025 Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning · NeurIPS 2023 Robustly Learning Composable Options in Deep Reinforcement Learning · IJCAI 2021 |
Machine learning › Reinforcement learning
exploration |
1.9 | 3 | 2023 | Scaling Goal-based Exploration via Pruning Proto-goals · IJCAI 2023 Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement Learning · ICML 2023 Optimistic Initialization for Exploration in Continuous Control · AAAI 2022 |
Machine learning › Reinforcement learning › exploration › novelty-based exploration
count-based exploration |
0.7 | 1 | 2023 | Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning › exploration › directed exploration
goal-directed exploration |
0.7 | 1 | 2023 | Scaling Goal-based Exploration via Pruning Proto-goals · IJCAI 2023 |
Machine learning › Reinforcement learning › unsupervised reinforcement learning
goal discovery |
0.7 | 1 | 2023 | Scaling Goal-based Exploration via Pruning Proto-goals · IJCAI 2023 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
initiation set learning |
0.7 | 1 | 2023 | Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning · NeurIPS 2023 |
Machine learning › Reinforcement learning › exploration › optimistic exploration
optimistic initialization |
0.6 | 1 | 2022 | Optimistic Initialization for Exploration in Continuous Control · AAAI 2022 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › search-based planning
graph-based planning |
0.5 | 1 | 2021 | Skill Discovery for Exploration and Planning using Deep Skill Graphs · ICML 2021 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
skill discovery |
0.5 | 1 | 2021 | Skill Discovery for Exploration and Planning using Deep Skill Graphs · ICML 2021 |
Machine learning › Graph learning
skill graph |
0.5 | 1 | 2021 | Skill Discovery for Exploration and Planning using Deep Skill Graphs · ICML 2021 |
Robotics › Motion planning and robot control › robot learning › robot skill learning
skill chaining |
0.4 | 1 | 2020 | Option Discovery using Deep Skill Chaining · ICLR 2020 |
Robotics › Robot manipulation
grasping |
0.2 | 1 | 2023 | Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning · NeurIPS 2023 |
Machine learning › Reinforcement learning
continuous control |
0.2 | 1 | 2022 | Optimistic Initialization for Exploration in Continuous Control · AAAI 2022 |
Robotics › Motion planning and robot control
robot learning |
0.1 | 1 | 2021 | Robustly Learning Composable Options in Deep Reinforcement Learning · IJCAI 2021 |
Methods — techniques the papers use, named apart from their topics
k-median · 0.9approximation algorithm · 0.9supervised learning · 0.7rademacher distribution · 0.7proto-goal pruning · 0.7off-policy value estimation · 0.7intrinsic motivation · 0.7goal-conditioned exploration · 0.7classification · 0.7domain knowledge incorporation · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Discovering Options That Minimize Average Planning TimeabstractWe present an option discovery algorithm that accelerates planning by minimizing the shortest distance between any two states in the MDP. The proposed algorithm produces options that approximately minimize planning time in the multi-goal setting: it is shown to be a worst case (4-alpha, 2)-approximation of the optimal option set, where alpha is the approximation ratio of the k-medians with penalties subroutine. We then present a variation, "Fast Average Options", with improved run-time and describe a general means of producing similar algorithms based on selection of a k-medians subroutine. We empirically evaluate our method on four discrete and two continuous control planning domains and show that it outperforms other leading option discovery algorithms. Akhil Bagaria, George Dimitri Konidaris |
AAAI | 2 |
| 2023 | Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement LearningabstractWe propose a new method for count-based exploration in high-dimensional state spaces. Unlike previous work which relies on density models, we show that counts can be derived by averaging samples from the Rademacher distribution (or coin flips). This insight is used to set up a simple supervised learning objective which, when optimized, yields a state's visitation count. We show that our method is significantly more effective at deducing ground-truth visitation counts than previous work; when used as an exploration bonus for a model-free reinforcement learning algorithm, it outperforms existing approaches on most of 9 challenging exploration tasks, including the Atari game Montezuma's Revenge. Sam Lobel, Akhil Bagaria, George Dimitri Konidaris |
ICML | 2 |
| 2023 | Scaling Goal-based Exploration via Pruning Proto-goalsabstractOne of the gnarliest challenges in reinforcement learning (RL) is exploration that scales to vast domains, where novelty-, or coverage-seeking behaviour falls short. Goal-directed, purposeful behaviours are able to overcome this, but rely on a good goal space. The core challenge in goal discovery is finding the right balance between generality (not hand-crafted) and tractability (useful, not too many). Our approach explicitly seeks the middle ground, enabling the human designer to specify a vast but meaningful proto-goal space, and an autonomous discovery process to refine this to a narrower space of controllable, reachable, novel, and relevant goals. The effectiveness of goal-conditioned exploration with the latter is then demonstrated in three challenging environments. Akhil Bagaria, Tom Schaul |
IJCAI | 1 |
| 2023 | Effectively Learning Initiation Sets in Hierarchical Reinforcement LearningabstractAn agent learning an option in hierarchical reinforcement learning must solve three problems: identify the option's subgoal (termination condition), learn a policy, and learn where that policy will succeed (initiation set). The termination condition is typically identified first, but the option policy and initiation set must be learned simultaneously, which is challenging because the initiation set depends on the option policy, which changes as the agent learns. Consequently, data obtained from option execution becomes invalid over time, leading to an inaccurate initiation set that subsequently harms downstream task performance. We highlight three issues---data non-stationarity, temporal credit assignment, and pessimism---specific to learning initiation sets, and propose to address them using tools from off-policy value estimation and classification. We show that our method learns higher-quality initiation sets faster than existing methods (in MiniGrid and Montezuma's Revenge), can automatically discover promising grasps for robot manipulation (in Robosuite), and improves the performance of a state-of-the-art option discovery method in a challenging maze navigation task in MuJoCo. Akhil Bagaria, Ben Abbatematteo, Omer Gottesman, Matthew Corsaro, Sreehari Rammohan, George Dimitri Konidaris |
NeurIPS | 1 |
| 2022 | Optimistic Initialization for Exploration in Continuous ControlabstractOptimistic initialization underpins many theoretically sound exploration schemes in tabular domains; however, in the deep function approximation setting, optimism can quickly disappear if initialized naively. We propose a framework for more effectively incorporating optimistic initialization into reinforcement learning for continuous control. Our approach uses metric information about the state-action space to estimate which transitions are still unexplored, and explicitly maintains the initial Q-value optimism for the corresponding state-action pairs. We also develop methods for efficiently approximating these training objectives, and for incorporating domain knowledge into the optimistic envelope to improve sample efficiency. We empirically evaluate these approaches on a variety of hard exploration problems in continuous control, where our method outperforms existing exploration techniques. Sam Lobel, Omer Gottesman, Cameron Allen, Akhil Bagaria, George Dimitri Konidaris |
AAAI | 4 |
| 2021 | Skill Discovery for Exploration and Planning using Deep Skill GraphsabstractWe introduce a new skill-discovery algorithm that builds a discrete graph representation of large continuous MDPs, where nodes correspond to skill subgoals and the edges to skill policies. The agent constructs this graph during an unsupervised training phase where it interleaves discovering skills and planning using them to gain coverage over ever-increasing portions of the state-space. Given a novel goal at test time, the agent plans with the acquired skill graph to reach a nearby state, then switches to learning to reach the goal. We show that the resulting algorithm, Deep Skill Graphs, outperforms both flat and existing hierarchical reinforcement learning methods on four difficult continuous control tasks. Akhil Bagaria, Jason K. Senthil, George Dimitri Konidaris |
ICML | 1 |
| 2021 | Robustly Learning Composable Options in Deep Reinforcement LearningabstractHierarchical reinforcement learning (HRL) is only effective for long-horizon problems when high-level skills can be reliably sequentially executed. Unfortunately, learning reliably composable skills is difficult, because all the components of every skill are constantly changing during learning. We propose three methods for improving the composability of learned skills: representing skill initiation regions using a combination of pessimistic and optimistic classifiers; learning re-targetable policies that are robust to non-stationary subgoal regions; and learning robust option policies using model-based RL. We test these improvements on four sparse-reward maze navigation tasks involving a simulated quadrupedal robot. Each method successively improves the robustness of a baseline skill discovery method, substantially outperforming state-of-the-art flat and hierarchical methods. Akhil Bagaria, Jason K. Senthil, Matthew Slivinski, George Dimitri Konidaris |
IJCAI | 1 |
| 2020 | Option Discovery using Deep Skill Chaining
Akhil Bagaria, George Dimitri Konidaris |
ICLR | 1 |