Utsav Singh

dblp:241/9336 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 80% Learning theory · 12% Learning paradigms · 4%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 54% Distributed and cloud data management · 23% Data models and query languages · 23%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
hierarchical reinforcement learning
2.632026
CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement Learning · AAAI 2026
PEAR: Primitive Enabled Adaptive Relabeling for Boosting Hierarchical Reinforcement Learning · ICLR 2025
PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling · ICML 2024
Machine learning › Reinforcement learning
bilevel reinforcement learning
0.912025
On the Sample Complexity Bounds of Bilevel Reinforcement Learning · NeurIPS 2025
Machine learning › Learning theory
sample complexity
0.912025
On the Sample Complexity Bounds of Bilevel Reinforcement Learning · NeurIPS 2025
Query processing and optimization
analytical query processing
0.912025
Cloudy With a Chance of JSON · Proc. VLDB Endow. 2025
Query processing and optimization › query execution › query operator implementation
columnar execution
0.912025
Cloudy With a Chance of JSON · Proc. VLDB Endow. 2025
Distributed and cloud data management › cloud database
database-as-a-service
0.912025
Cloudy With a Chance of JSON · Proc. VLDB Endow. 2025
Data models and query languages › NoSQL database
document-oriented database
0.912025
Cloudy With a Chance of JSON · Proc. VLDB Endow. 2025
Mathematical optimization
bilevel optimization
0.912025
On the Sample Complexity Bounds of Bilevel Reinforcement Learning · NeurIPS 2025
Mathematical optimization › bilevel optimization
nonconvex bilevel optimization
0.912025
On the Sample Complexity Bounds of Bilevel Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
hindsight relabeling
0.812024
PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling · ICML 2024
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning
0.812024
PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling · ICML 2024
Machine learning › Learning paradigms
curriculum learning
0.312026
CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement Learning · AAAI 2026
Robotics › Motion planning and robot control
robot learning
0.312026
CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement Learning · AAAI 2026

Methods — techniques the papers use, named apart from their topics

polyak-łojasiewicz condition · 1.7hypergradient estimation · 1.7hessian-free first-order optimization · 1.7subgoal prediction · 1.0inverse reinforcement learning · 1.0expert demonstration · 1.0shared-nothing architecture · 0.9off-policy reinforcement learning · 0.9imitation learning · 0.9primitive-informed regularization · 0.8preference-based learning · 0.8
YearPublicationVenuePosition
2026 CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement Learning
abstract
Hierarchical reinforcement learning (HRL) leverages temporal abstraction to efficiently tackle complex long-horizon tasks. However, HRL often collapses because the low-level primitive’s continual updates make earlier sub-goals issued by the high-level policy obsolete, introducing non-stationarity that destabilizes training. We propose CRISP, a curriculum-driven framework that tackles this instability with three key ingredients: (1) primitive-informed parsing (PIP), which adaptively re-labels a handful of expert demonstrations to always generate reachable subgoals by the current low-level primitive; (2) an inverse-reinforcement-learning regularizer that steers the high-level policy toward the expert-induced subgoal distribution and stabilizes learning; and (3) a unified training loop that leverages these components to boost sample efficiency. Across six sparse-reward robotic navigation and manipulation benchmarks, CRISP improves success rates by more than 40% over strong hierarchical and flat baselines and successfully transfers to real-world tasks, demonstrating the promise of curriculum-based HRL for practical scenarios.
Utsav Singh, Vinay P. Namboodiri
AAAI1
2025 PEAR: Primitive Enabled Adaptive Relabeling for Boosting Hierarchical Reinforcement Learning
abstract
Hierarchical reinforcement learning (HRL) has the potential to solve complex long horizon tasks using temporal abstraction and increased exploration. However, hierarchical agents are difficult to train due to inherent non-stationarity. We present primitive enabled adaptive relabeling (PEAR), a two-phase approach where we first perform adaptive relabeling on a few expert demonstrations to generate efficient subgoal supervision, and then jointly optimize HRL agents by employing reinforcement learning (RL) and imitation learning (IL). We perform theoretical analysis to bound the sub-optimality of our approach and derive a joint optimization framework using RL and IL. Since PEAR utilizes only a few expert demonstrations and considers minimal limiting assumptions on the task structure, it can be easily integrated with typical off-policy \RL algorithms to produce a practical HRL approach. We perform extensive experiments on challenging environments and show that PEAR is able to outperform various hierarchical and non-hierarchical baselines and achieve upto 80% success rates in complex sparse robotic control tasks where other baselines typically fail to show significant progress. We also perform ablations to thoroughly analyze the importance of our various design choices. Finally, we perform real world robotic experiments on complex tasks and demonstrate that PEAR consistently outperforms the baselines.
Utsav Singh, Vinay P. Namboodiri
ICLR1
2025 On the Sample Complexity Bounds of Bilevel Reinforcement Learning
abstract
Bilevel reinforcement learning (BRL) has emerged as a powerful framework for aligning generative models, yet its theoretical foundations, especially sample complexity bounds, remain relatively underexplored. In this work, we present the first sample complexity bound for BRL, establishing a rate of $\tilde{\mathcal{O}}(\epsilon^{-3})$ in continuous state-action spaces. Traditional MDP analysis techniques do not extend to BRL due to its nested structure and non-convex lower-level problems. We overcome these challenges by leveraging the Polyak-Łojasiewicz (PL) condition and the MDP structure to obtain closed-form gradients, enabling tight sample complexity analysis. Our analysis also extends to general bi-level optimization settings with non-convex lower levels, where we achieve state-of-the-art sample complexity results of $\tilde{\mathcal{O}}(\epsilon^{-3})$ improving upon existing bounds of $\tilde{\mathcal{O}}(\epsilon^{-6})$. Additionally, we address the computational bottleneck of hypergradient estimation by proposing a fully first-order, Hessian-free algorithm suitable for large-scale problems.
Mudit Gaur, Utsav Singh, Amrit Singh Bedi, Raghu Pasupathy, Vaneet Aggarwal
NeurIPS2
2025 Cloudy With a Chance of JSON
abstract
Couchbase Capella is a scalable document-oriented database service in the cloud. Its existing Capella Operational service is based on a shared-nothing architecture and supports high volumes of low-latency queries and updates for JSON documents. Its new Capella Columnar cloud service complements the Operational service. The Capella Columnar service supports complex analytical queries (e.g., ad hoc joins and aggregations) over large collections of JSON documents that can originate from a variety of Couchbase and non-Couchbase data sources and formats and can either be stored and managed by the Capella Columnar service or externally stored and accessed on demand at query time. This paper describes the new Capella Columnar service, looking both over and under the hood.
Murtadha Al Hubail, Ali Alsuliman, Wail Y. Alkowaileet, Michael Blow, Michael J. Carey 0001, Savyasach Enukonda, Peeyush Gupta, Santosh Hegde, Kamini Jagtiani, Abhishek Jindal, Nawazish Kahn, Mehnaz Tabassum Mahin, Ian Maxon, M. Muralikrishna, Keshav Murthy, Preetham Poluparthi, Ankit Prabhu, Ritik Raj, Vijay Sarathy, Shahrzad Shirazi, Utsav Singh, Hussain Towaileb, Ayush Tripathi, Janhavi Tripurwar, Bo-Chun Wang, Till Westmann
Proc. VLDB Endow.22
2024 PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling
abstract
In this work, we introduce PIPER: Primitive-Informed Preference-based Hierarchical reinforcement learning via Hindsight Relabeling, a novel approach that leverages preference-based learning to learn a reward model, and subsequently uses this reward model to relabel higher-level replay buffers. Since this reward is unaffected by lower primitive behavior, our relabeling-based approach is able to mitigate non-stationarity, which is common in existing hierarchical approaches, and demonstrates impressive performance across a range of challenging sparse-reward tasks. Since obtaining human feedback is typically impractical, we propose to replace the human-in-the-loop approach with our primitive-in-the-loop approach, which generates feedback using sparse rewards provided by the environment. Moreover, in order to prevent infeasible subgoal prediction and avoid degenerate solutions, we propose primitive-informed regularization that conditions higher-level policies to generate feasible subgoals for lower-level policies. We perform extensive experiments to show that PIPER mitigates non-stationarity in hierarchical reinforcement learning and achieves greater than 50$\%$ success rates in challenging, sparse-reward robotic environments, where most other baselines fail to achieve any significant progress.
Utsav Singh, Wesley Suttle, Brian M. Sadler, Vinay P. Namboodiri, Amrit Singh Bedi
ICML1