EDBT 2026 Demo / reviewers in the wild / expert
Utsav Singh
dblp:241/9336
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 80% Learning theory · 12% Learning paradigms · 4% | |
| Databases, data mining, and information retrieval
1 paper |
Query processing and optimization · 54% Distributed and cloud data management · 23% Data models and query languages · 23% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
2.6 | 3 | 2026 | CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement Learning · AAAI 2026 PEAR: Primitive Enabled Adaptive Relabeling for Boosting Hierarchical Reinforcement Learning · ICLR 2025 PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling · ICML 2024 |
Machine learning › Reinforcement learning
bilevel reinforcement learning |
0.9 | 1 | 2025 | On the Sample Complexity Bounds of Bilevel Reinforcement Learning · NeurIPS 2025 |
Machine learning › Learning theory
sample complexity |
0.9 | 1 | 2025 | On the Sample Complexity Bounds of Bilevel Reinforcement Learning · NeurIPS 2025 |
Query processing and optimization
analytical query processing |
0.9 | 1 | 2025 | Cloudy With a Chance of JSON · Proc. VLDB Endow. 2025 |
Query processing and optimization › query execution › query operator implementation
columnar execution |
0.9 | 1 | 2025 | Cloudy With a Chance of JSON · Proc. VLDB Endow. 2025 |
Distributed and cloud data management › cloud database
database-as-a-service |
0.9 | 1 | 2025 | Cloudy With a Chance of JSON · Proc. VLDB Endow. 2025 |
Data models and query languages › NoSQL database
document-oriented database |
0.9 | 1 | 2025 | Cloudy With a Chance of JSON · Proc. VLDB Endow. 2025 |
Mathematical optimization
bilevel optimization |
0.9 | 1 | 2025 | On the Sample Complexity Bounds of Bilevel Reinforcement Learning · NeurIPS 2025 |
Mathematical optimization › bilevel optimization
nonconvex bilevel optimization |
0.9 | 1 | 2025 | On the Sample Complexity Bounds of Bilevel Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning
hindsight relabeling |
0.8 | 1 | 2024 | PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling · ICML 2024 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
0.8 | 1 | 2024 | PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling · ICML 2024 |
Machine learning › Learning paradigms
curriculum learning |
0.3 | 1 | 2026 | CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement Learning · AAAI 2026 |
Robotics › Motion planning and robot control
robot learning |
0.3 | 1 | 2026 | CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement Learning · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
polyak-łojasiewicz condition · 1.7hypergradient estimation · 1.7hessian-free first-order optimization · 1.7subgoal prediction · 1.0inverse reinforcement learning · 1.0expert demonstration · 1.0shared-nothing architecture · 0.9off-policy reinforcement learning · 0.9imitation learning · 0.9primitive-informed regularization · 0.8preference-based learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement LearningabstractHierarchical reinforcement learning (HRL) leverages temporal abstraction to efficiently tackle complex long-horizon tasks. However, HRL often collapses because the low-level primitive’s continual updates make earlier sub-goals issued by the high-level policy obsolete, introducing non-stationarity that destabilizes training. We propose CRISP, a curriculum-driven framework that tackles this instability with three key ingredients: (1) primitive-informed parsing (PIP), which adaptively re-labels a handful of expert demonstrations to always generate reachable subgoals by the current low-level primitive; (2) an inverse-reinforcement-learning regularizer that steers the high-level policy toward the expert-induced subgoal distribution and stabilizes learning; and (3) a unified training loop that leverages these components to boost sample efficiency. Across six sparse-reward robotic navigation and manipulation benchmarks, CRISP improves success rates by more than 40% over strong hierarchical and flat baselines and successfully transfers to real-world tasks, demonstrating the promise of curriculum-based HRL for practical scenarios. Utsav Singh, Vinay P. Namboodiri |
AAAI | 1 |
| 2025 | PEAR: Primitive Enabled Adaptive Relabeling for Boosting Hierarchical Reinforcement LearningabstractHierarchical reinforcement learning (HRL) has the potential to solve complex long horizon tasks using temporal abstraction and increased exploration. However, hierarchical agents are difficult to train due to inherent non-stationarity. We present primitive enabled adaptive relabeling (PEAR), a two-phase approach where we first perform adaptive relabeling on a few expert demonstrations to generate efficient subgoal supervision, and then jointly optimize HRL agents by employing reinforcement learning (RL) and imitation learning (IL). We perform theoretical analysis to bound the sub-optimality of our approach and derive a joint optimization framework using RL and IL. Since PEAR utilizes only a few expert demonstrations and considers minimal limiting assumptions on the task structure, it can be easily integrated with typical off-policy \RL algorithms to produce a practical HRL approach. We perform extensive experiments on challenging environments and show that PEAR is able to outperform various hierarchical and non-hierarchical baselines and achieve upto 80% success rates in complex sparse robotic control tasks where other baselines typically fail to show significant progress. We also perform ablations to thoroughly analyze the importance of our various design choices. Finally, we perform real world robotic experiments on complex tasks and demonstrate that PEAR consistently outperforms the baselines. Utsav Singh, Vinay P. Namboodiri |
ICLR | 1 |
| 2025 | On the Sample Complexity Bounds of Bilevel Reinforcement LearningabstractBilevel reinforcement learning (BRL) has emerged as a powerful framework for aligning generative models, yet its theoretical foundations, especially sample complexity bounds, remain relatively underexplored. In this work, we present the first sample complexity bound for BRL, establishing a rate of $\tilde{\mathcal{O}}(\epsilon^{-3})$ in continuous state-action spaces. Traditional MDP analysis techniques do not extend to BRL due to its nested structure and non-convex lower-level problems. We overcome these challenges by leveraging the Polyak-Łojasiewicz (PL) condition and the MDP structure to obtain closed-form gradients, enabling tight sample complexity analysis. Our analysis also extends to general bi-level optimization settings with non-convex lower levels, where we achieve state-of-the-art sample complexity results of $\tilde{\mathcal{O}}(\epsilon^{-3})$ improving upon existing bounds of $\tilde{\mathcal{O}}(\epsilon^{-6})$. Additionally, we address the computational bottleneck of hypergradient estimation by proposing a fully first-order, Hessian-free algorithm suitable for large-scale problems. Mudit Gaur, Utsav Singh, Amrit Singh Bedi, Raghu Pasupathy, Vaneet Aggarwal |
NeurIPS | 2 |
| 2025 | Cloudy With a Chance of JSONabstractCouchbase Capella is a scalable document-oriented database service in the cloud. Its existing Capella Operational service is based on a shared-nothing architecture and supports high volumes of low-latency queries and updates for JSON documents. Its new Capella Columnar cloud service complements the Operational service. The Capella Columnar service supports complex analytical queries (e.g., ad hoc joins and aggregations) over large collections of JSON documents that can originate from a variety of Couchbase and non-Couchbase data sources and formats and can either be stored and managed by the Capella Columnar service or externally stored and accessed on demand at query time. This paper describes the new Capella Columnar service, looking both over and under the hood. Murtadha Al Hubail, Ali Alsuliman, Wail Y. Alkowaileet, Michael Blow, Michael J. Carey 0001, Savyasach Enukonda, Peeyush Gupta, Santosh Hegde, Kamini Jagtiani, Abhishek Jindal, Nawazish Kahn, Mehnaz Tabassum Mahin, Ian Maxon, M. Muralikrishna, Keshav Murthy, Preetham Poluparthi, Ankit Prabhu, Ritik Raj, Vijay Sarathy, Shahrzad Shirazi, Utsav Singh, Hussain Towaileb, Ayush Tripathi, Janhavi Tripurwar, Bo-Chun Wang, Till Westmann |
Proc. VLDB Endow. | 22 |
| 2024 | PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight RelabelingabstractIn this work, we introduce PIPER: Primitive-Informed Preference-based Hierarchical reinforcement learning via Hindsight Relabeling, a novel approach that leverages preference-based learning to learn a reward model, and subsequently uses this reward model to relabel higher-level replay buffers. Since this reward is unaffected by lower primitive behavior, our relabeling-based approach is able to mitigate non-stationarity, which is common in existing hierarchical approaches, and demonstrates impressive performance across a range of challenging sparse-reward tasks. Since obtaining human feedback is typically impractical, we propose to replace the human-in-the-loop approach with our primitive-in-the-loop approach, which generates feedback using sparse rewards provided by the environment. Moreover, in order to prevent infeasible subgoal prediction and avoid degenerate solutions, we propose primitive-informed regularization that conditions higher-level policies to generate feasible subgoals for lower-level policies. We perform extensive experiments to show that PIPER mitigates non-stationarity in hierarchical reinforcement learning and achieves greater than 50$\%$ success rates in challenging, sparse-reward robotic environments, where most other baselines fail to achieve any significant progress. Utsav Singh, Wesley Suttle, Brian M. Sadler, Vinay P. Namboodiri, Amrit Singh Bedi |
ICML | 1 |