Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Akhil Bagaria

dblp:155/9746 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 87% Motion planning and robot control · 4% Graph learning · 4%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
hierarchical reinforcement learning
3.662025
Discovering Options That Minimize Average Planning Time · AAAI 2025
Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning · NeurIPS 2023
Scaling Goal-based Exploration via Pruning Proto-goals · IJCAI 2023
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery
2.542025
Discovering Options That Minimize Average Planning Time · AAAI 2025
Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning · NeurIPS 2023
Robustly Learning Composable Options in Deep Reinforcement Learning · IJCAI 2021
Machine learning › Reinforcement learning
exploration
1.932023
Scaling Goal-based Exploration via Pruning Proto-goals · IJCAI 2023
Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement Learning · ICML 2023
Optimistic Initialization for Exploration in Continuous Control · AAAI 2022
Machine learning › Reinforcement learning › exploration › novelty-based exploration
count-based exploration
0.712023
Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › exploration › directed exploration
goal-directed exploration
0.712023
Scaling Goal-based Exploration via Pruning Proto-goals · IJCAI 2023
Machine learning › Reinforcement learning › unsupervised reinforcement learning
goal discovery
0.712023
Scaling Goal-based Exploration via Pruning Proto-goals · IJCAI 2023
Machine learning › Reinforcement learning › hierarchical reinforcement learning
initiation set learning
0.712023
Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning › exploration › optimistic exploration
optimistic initialization
0.612022
Optimistic Initialization for Exploration in Continuous Control · AAAI 2022
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › search-based planning
graph-based planning
0.512021
Skill Discovery for Exploration and Planning using Deep Skill Graphs · ICML 2021
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
skill discovery
0.512021
Skill Discovery for Exploration and Planning using Deep Skill Graphs · ICML 2021
Machine learning › Graph learning
skill graph
0.512021
Skill Discovery for Exploration and Planning using Deep Skill Graphs · ICML 2021
Robotics › Motion planning and robot control › robot learning › robot skill learning
skill chaining
0.412020
Option Discovery using Deep Skill Chaining · ICLR 2020
Robotics › Robot manipulation
grasping
0.212023
Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
continuous control
0.212022
Optimistic Initialization for Exploration in Continuous Control · AAAI 2022
Robotics › Motion planning and robot control
robot learning
0.112021
Robustly Learning Composable Options in Deep Reinforcement Learning · IJCAI 2021

Methods — techniques the papers use, named apart from their topics

k-median · 0.9approximation algorithm · 0.9supervised learning · 0.7rademacher distribution · 0.7proto-goal pruning · 0.7off-policy value estimation · 0.7intrinsic motivation · 0.7goal-conditioned exploration · 0.7classification · 0.7domain knowledge incorporation · 0.6
YearPublicationVenuePosition
2025 Discovering Options That Minimize Average Planning Time
abstract
We present an option discovery algorithm that accelerates planning by minimizing the shortest distance between any two states in the MDP. The proposed algorithm produces options that approximately minimize planning time in the multi-goal setting: it is shown to be a worst case (4-alpha, 2)-approximation of the optimal option set, where alpha is the approximation ratio of the k-medians with penalties subroutine. We then present a variation, "Fast Average Options", with improved run-time and describe a general means of producing similar algorithms based on selection of a k-medians subroutine. We empirically evaluate our method on four discrete and two continuous control planning domains and show that it outperforms other leading option discovery algorithms.
Akhil Bagaria, George Dimitri Konidaris
AAAI2
2023 Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement Learning
abstract
We propose a new method for count-based exploration in high-dimensional state spaces. Unlike previous work which relies on density models, we show that counts can be derived by averaging samples from the Rademacher distribution (or coin flips). This insight is used to set up a simple supervised learning objective which, when optimized, yields a state's visitation count. We show that our method is significantly more effective at deducing ground-truth visitation counts than previous work; when used as an exploration bonus for a model-free reinforcement learning algorithm, it outperforms existing approaches on most of 9 challenging exploration tasks, including the Atari game Montezuma's Revenge.
Sam Lobel, Akhil Bagaria, George Dimitri Konidaris
ICML2
2023 Scaling Goal-based Exploration via Pruning Proto-goals
abstract
One of the gnarliest challenges in reinforcement learning (RL) is exploration that scales to vast domains, where novelty-, or coverage-seeking behaviour falls short. Goal-directed, purposeful behaviours are able to overcome this, but rely on a good goal space. The core challenge in goal discovery is finding the right balance between generality (not hand-crafted) and tractability (useful, not too many). Our approach explicitly seeks the middle ground, enabling the human designer to specify a vast but meaningful proto-goal space, and an autonomous discovery process to refine this to a narrower space of controllable, reachable, novel, and relevant goals. The effectiveness of goal-conditioned exploration with the latter is then demonstrated in three challenging environments.
Akhil Bagaria, Tom Schaul
IJCAI1
2023 Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning
abstract
An agent learning an option in hierarchical reinforcement learning must solve three problems: identify the option's subgoal (termination condition), learn a policy, and learn where that policy will succeed (initiation set). The termination condition is typically identified first, but the option policy and initiation set must be learned simultaneously, which is challenging because the initiation set depends on the option policy, which changes as the agent learns. Consequently, data obtained from option execution becomes invalid over time, leading to an inaccurate initiation set that subsequently harms downstream task performance. We highlight three issues---data non-stationarity, temporal credit assignment, and pessimism---specific to learning initiation sets, and propose to address them using tools from off-policy value estimation and classification. We show that our method learns higher-quality initiation sets faster than existing methods (in MiniGrid and Montezuma's Revenge), can automatically discover promising grasps for robot manipulation (in Robosuite), and improves the performance of a state-of-the-art option discovery method in a challenging maze navigation task in MuJoCo.
Akhil Bagaria, Ben Abbatematteo, Omer Gottesman, Matthew Corsaro, Sreehari Rammohan, George Dimitri Konidaris
NeurIPS1
2022 Optimistic Initialization for Exploration in Continuous Control
abstract
Optimistic initialization underpins many theoretically sound exploration schemes in tabular domains; however, in the deep function approximation setting, optimism can quickly disappear if initialized naively. We propose a framework for more effectively incorporating optimistic initialization into reinforcement learning for continuous control. Our approach uses metric information about the state-action space to estimate which transitions are still unexplored, and explicitly maintains the initial Q-value optimism for the corresponding state-action pairs. We also develop methods for efficiently approximating these training objectives, and for incorporating domain knowledge into the optimistic envelope to improve sample efficiency. We empirically evaluate these approaches on a variety of hard exploration problems in continuous control, where our method outperforms existing exploration techniques.
Sam Lobel, Omer Gottesman, Cameron Allen, Akhil Bagaria, George Dimitri Konidaris
AAAI4
2021 Skill Discovery for Exploration and Planning using Deep Skill Graphs
abstract
We introduce a new skill-discovery algorithm that builds a discrete graph representation of large continuous MDPs, where nodes correspond to skill subgoals and the edges to skill policies. The agent constructs this graph during an unsupervised training phase where it interleaves discovering skills and planning using them to gain coverage over ever-increasing portions of the state-space. Given a novel goal at test time, the agent plans with the acquired skill graph to reach a nearby state, then switches to learning to reach the goal. We show that the resulting algorithm, Deep Skill Graphs, outperforms both flat and existing hierarchical reinforcement learning methods on four difficult continuous control tasks.
Akhil Bagaria, Jason K. Senthil, George Dimitri Konidaris
ICML1
2021 Robustly Learning Composable Options in Deep Reinforcement Learning
abstract
Hierarchical reinforcement learning (HRL) is only effective for long-horizon problems when high-level skills can be reliably sequentially executed. Unfortunately, learning reliably composable skills is difficult, because all the components of every skill are constantly changing during learning. We propose three methods for improving the composability of learned skills: representing skill initiation regions using a combination of pessimistic and optimistic classifiers; learning re-targetable policies that are robust to non-stationary subgoal regions; and learning robust option policies using model-based RL. We test these improvements on four sparse-reward maze navigation tasks involving a simulated quadrupedal robot. Each method successively improves the robustness of a baseline skill discovery method, substantially outperforming state-of-the-art flat and hierarchical methods.
Akhil Bagaria, Jason K. Senthil, Matthew Slivinski, George Dimitri Konidaris
IJCAI1
2020 Option Discovery using Deep Skill Chaining
Akhil Bagaria, George Dimitri Konidaris
ICLR1