EDBT 2026 Demo / reviewers in the wild / expert
Dilip Arumugam
dblp:165/1303
· DBLP profile ↗
12ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 85% Motion planning and robot control · 5% Planning, search and constraint satisfaction · 5% | |
| Theoretical computer science
5 papers |
Coding theory · 84% Computational complexity · 16% |
Topics — the 13 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state abstraction |
1.3 | 3 | 2022 | Deciding What to Model: Value-Equivalent Sampling for Reinforcement Learning · NeurIPS 2022 State Abstraction as Compression in Apprenticeship Learning · AAAI 2019 State Abstractions for Lifelong Reinforcement Learning · ICML 2018 |
Machine learning › Reinforcement learning
exploration |
1.1 | 2 | 2022 | Planning to the Information Horizon of BAMDPs via Epistemic State Abstraction · NeurIPS 2022 The Value of Information When Deciding What to Learn · NeurIPS 2021 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
1.0 | 2 | 2022 | Deciding What to Model: Value-Equivalent Sampling for Reinforcement Learning · NeurIPS 2022 Flexible and Efficient Long-Range Planning Through Curious Exploration · ICML 2020 |
Coding theory › source coding
rate-distortion theory |
0.9 | 4 | 2022 | Deciding What to Learn: A Rate-Distortion Approach · ICML 2021 Deciding What to Model: Value-Equivalent Sampling for Reinforcement Learning · NeurIPS 2022 The Value of Information When Deciding What to Learn · NeurIPS 2021 |
Machine learning › Reinforcement learning › markov decision process › Bayesian MDP
bayes-adaptive markov decision process |
0.6 | 1 | 2022 | Planning to the Information Horizon of BAMDPs via Epistemic State Abstraction · NeurIPS 2022 |
Machine learning › Reinforcement learning
bayesian reinforcement learning |
0.6 | 1 | 2022 | Planning to the Information Horizon of BAMDPs via Epistemic State Abstraction · NeurIPS 2022 |
Machine learning › Reinforcement learning › model-based reinforcement learning
value equivalence |
0.6 | 1 | 2022 | Deciding What to Model: Value-Equivalent Sampling for Reinforcement Learning · NeurIPS 2022 |
Machine learning › Reinforcement learning › exploration › information-theoretic exploration
information-directed sampling |
0.5 | 1 | 2021 | The Value of Information When Deciding What to Learn · NeurIPS 2021 |
Machine learning › Reinforcement learning › exploration › intrinsic motivation
curiosity-driven exploration |
0.4 | 1 | 2020 | Flexible and Efficient Long-Range Planning Through Curious Exploration · ICML 2020 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-term planning |
0.4 | 1 | 2020 | Flexible and Efficient Long-Range Planning Through Curious Exploration · ICML 2020 |
Robotics › Motion planning and robot control
task and motion planning |
0.4 | 1 | 2020 | Flexible and Efficient Long-Range Planning Through Curious Exploration · ICML 2020 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.4 | 1 | 2019 | State Abstraction as Compression in Apprenticeship Learning · AAAI 2019 |
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning |
0.3 | 1 | 2018 | State Abstractions for Lifelong Reinforcement Learning · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
rate-distortion theory · 3.9state abstraction · 1.1information-theoretic complexity measure · 1.1bayesian regret analysis · 1.1thompson sampling · 1.0information-directed sampling · 1.0blahut-arimoto algorithm · 0.8imitation learning · 0.4deep reinforcement learning · 0.4curiosity-guided sampling · 0.4information bottleneck · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Discovering Hidden Laws in Innovation by Recombination
Bonan Zhao 0001, Elizabeth Mieczkowski, Dilip Arumugam, Natalia Vélez, Thomas L. Griffiths 0001 |
CogSci | 3 |
| 2025 | Trade-Offs Between Tasks Induced by Capacity Constraints Bound the Scope of Intelligence
Cameron Rouse Turner, Dilip Arumugam, Logan Nelson, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2025 | Hindsight Merging: Diverse Data Generation with Language ModelsabstractPre-training a language model equips it with a broad understanding of the world, while fine- tuning refines it into a helpful assistant. However, fine-tuning does not exclusively enhance task- specific behaviors but also suppresses some of the beneficial variability from pre-training. This reduction in diversity is partly due to the optimization process, which theoretically decreases model entropy in exchange for task performance. To counteract this, we introduce hindsight merging, a technique that combines a fine-tuned model with a previous training checkpoint using linear interpolation to restore entropy and improve performance. Hindsight-merged models retain strong instruction-following capabilities and alignment while displaying increased diversity present in the base model. Additionally, this results in improved inference scaling, achieving a consistent 20-50% increase in pass@10 relative to the instruction tuned model across a coding benchmark and series of models. Our findings suggest that hindsight merging is an effective strategy for generating diverse generations that follow instructions. Veniamin Veselovsky, Benedikt Stroebl, Gianluca M. Bencomo, Dilip Arumugam, Lisa Schut, Arvind Narayanan, Thomas L. Griffiths 0001 |
UAI | 4 |
| 2023 | Cultural reinforcement learning: a framework for modeling cumulative culture on a limited channel
Ben Prystawski, Dilip Arumugam, Noah D. Goodman |
CogSci | 2 |
| 2022 | Planning to the Information Horizon of BAMDPs via Epistemic State AbstractionabstractThe Bayes-Adaptive Markov Decision Process (BAMDP) formalism pursues the Bayes-optimal solution to the exploration-exploitation trade-off in reinforcement learning. As the computation of exact solutions to Bayesian reinforcement-learning problems is intractable, much of the literature has focused on developing suitable approximation algorithms. In this work, before diving into algorithm design, we first define, under mild structural assumptions, a complexity measure for BAMDP planning. As efficient exploration in BAMDPs hinges upon the judicious acquisition of information, our complexity measure highlights the worst-case difficulty of gathering information and exhausting epistemic uncertainty. To illustrate its significance, we establish a computationally-intractable, exact planning algorithm that takes advantage of this measure to show more efficient planning. We then conclude by introducing a specific form of state abstraction with the potential to reduce BAMDP complexity and gives rise to a computationally-tractable, approximate planning algorithm. Dilip Arumugam, Satinder Singh 0001 |
NeurIPS | 1 |
| 2022 | Deciding What to Model: Value-Equivalent Sampling for Reinforcement LearningabstractThe quintessential model-based reinforcement-learning agent iteratively refines its estimates or prior beliefs about the true underlying model of the environment. Recent empirical successes in model-based reinforcement learning with function approximation, however, eschew the true model in favor of a surrogate that, while ignoring various facets of the environment, still facilitates effective planning over behaviors. Recently formalized as the value equivalence principle, this algorithmic technique is perhaps unavoidable as real-world reinforcement learning demands consideration of a simple, computationally-bounded agent interacting with an overwhelmingly complex environment, whose underlying dynamics likely exceed the agent's capacity for representation. In this work, we consider the scenario where agent limitations may entirely preclude identifying an exactly value-equivalent model, immediately giving rise to a trade-off between identifying a model that is simple enough to learn while only incurring bounded sub-optimality. To address this problem, we introduce an algorithm that, using rate-distortion theory, iteratively computes an approximately-value-equivalent, lossy compression of the environment which an agent may feasibly target in lieu of the true model. We prove an information-theoretic, Bayesian regret bound for our algorithm that holds for any finite-horizon, episodic sequential decision-making problem. Crucially, our regret bound can be expressed in one of two possible forms, providing a performance guarantee for finding either the simplest model that achieves a desired sub-optimality gap or, alternatively, the best model given a limit on agent capacity. Dilip Arumugam, Benjamin Van Roy |
NeurIPS | 1 |
| 2021 | Deciding What to Learn: A Rate-Distortion ApproachabstractAgents that learn to select optimal actions represent a prominent focus of the sequential decision-making literature. In the face of a complex environment or constraints on time and resources, however, aiming to synthesize such an optimal policy can become infeasible. These scenarios give rise to an important trade-off between the information an agent must acquire to learn and the sub-optimality of the resulting policy. While an agent designer has a preference for how this trade-off is resolved, existing approaches further require that the designer translate these preferences into a fixed learning target for the agent. In this work, leveraging rate-distortion theory, we automate this process such that the designer need only express their preferences via a single hyperparameter and the agent is endowed with the ability to compute its own learning targets that best achieve the desired trade-off. We establish a general bound on expected discounted regret for an agent that decides what to learn in this manner along with computational experiments that illustrate the expressiveness of designer preferences and even show improvements over Thompson sampling in identifying an optimal policy. Dilip Arumugam, Benjamin Van Roy |
ICML | 1 |
| 2021 | The Value of Information When Deciding What to LearnabstractAll sequential decision-making agents explore so as to acquire knowledge about a particular target. It is often the responsibility of the agent designer to construct this target which, in rich and complex environments, constitutes a onerous burden; without full knowledge of the environment itself, a designer may forge a sub-optimal learning target that poorly balances the amount of information an agent must acquire to identify the target against the target's associated performance shortfall. While recent work has developed a connection between learning targets and rate-distortion theory to address this challenge and empower agents that decide what to learn in an automated fashion, the proposed algorithm does not optimally tackle the equally important challenge of efficient information acquisition. In this work, building upon the seminal design principle of information-directed sampling (Russo & Van Roy, 2014), we address this shortcoming directly to couple optimal information acquisition with the optimal design of learning targets. Along the way, we offer new insights into learning targets from the literature on rate-distortion theory before turning to empirical results that confirm the value of information when deciding what to learn. Dilip Arumugam, Benjamin Van Roy |
NeurIPS | 1 |
| 2020 | Value Preserving State-Action AbstractionsabstractAbstraction can improve the sample efficiency of reinforcement learning. However, the process of abstraction inherently discards information, potentially compromising an agent’s ability to represent high-value policies. To mitigate this, we here introduce combinations of state abstractions and options that are guaranteed to preserve representation of near-optimal policies. We first define $\phi$-relative options, a general formalism for analyzing the value loss of options paired with a state abstraction, and present necessary and sufficient conditions for $\phi$-relative options to preserve near-optimal behavior in any finite Markov Decision Process. We further show that, under appropriate assumptions, $\phi$-relative options can be composed to induce hierarchical abstractions that are also guaranteed to represent high-value policies. David Abel, Nate Umbanhowar, Khimya Khetarpal, Dilip Arumugam, Doina Precup, Michael L. Littman |
AISTATS | 4 |
| 2020 | Flexible and Efficient Long-Range Planning Through Curious ExplorationabstractIdentifying algorithms that flexibly and efficiently discover temporally-extended multi-phase plans is an essential step for the advancement of robotics and model-based reinforcement learning. The core problem of long-range planning is finding an efficient way to search through the tree of possible action sequences. Existing non-learned planning solutions from the Task and Motion Planning (TAMP) literature rely on the existence of logical descriptions for the effects and preconditions for actions. This constraint allows TAMP methods to efficiently reduce the tree search problem but limits their ability to generalize to unseen and complex physical environments. In contrast, deep reinforcement learning (DRL) methods use flexible neural-network-based function approximators to discover policies that generalize naturally to unseen circumstances. However, DRL methods struggle to handle the very sparse reward landscapes inherent to long-range multi-step planning situations. Here, we propose the Curious Sample Planner (CSP), which fuses elements of TAMP and DRL by combining a curiosity-guided sampling strategy with imitation learning to accelerate planning. We show that CSP can efficiently discover interesting and complex temporally-extended plans for solving a wide range of physically realistic 3D tasks. In contrast, standard planning and learning methods often fail to solve these tasks at all or do so only with a huge and highly variable number of training samples. We explore the use of a variety of curiosity metrics with CSP and analyze the types of solutions that CSP discovers. Finally, we show that CSP supports task transfer so that the exploration policies learned during experience with one task can help improve efficiency on related tasks. Aidan Curtis, Minjian Xin, Dilip Arumugam, Kevin T. Feigelis, Dan Yamins |
ICML | 3 |
| 2019 | State Abstraction as Compression in Apprenticeship LearningabstractState abstraction can give rise to models of environments that are both compressed and useful, thereby enabling efficient sequential decision making. In this work, we offer the first formalism and analysis of the trade-off between compression and performance made in the context of state abstraction for Apprenticeship Learning. We build on Rate-Distortion theory, the classic Blahut-Arimoto algorithm, and the Information Bottleneck method to develop an algorithm for computing state abstractions that approximate the optimal tradeoff between compression and performance. We illustrate the power of this algorithmic structure to offer insights into effective abstraction, compression, and reinforcement learning through a mixture of analysis, visuals, and experimentation. David Abel, Dilip Arumugam, Kavosh Asadi, Yuu Jinnai, Michael L. Littman, Lawson L. S. Wong |
AAAI | 2 |
| 2018 | State Abstractions for Lifelong Reinforcement LearningabstractIn lifelong reinforcement learning, agents must effectively transfer knowledge across tasks while simultaneously addressing exploration, credit assignment, and generalization. State abstraction can help overcome these hurdles by compressing the representation used by an agent, thereby reducing the computational and statistical burdens of learning. To this end, we here develop theory to compute and use state abstractions in lifelong reinforcement learning. We introduce two new classes of abstractions: (1) transitive state abstractions, whose optimal form can be computed efficiently, and (2) PAC state abstractions, which are guaranteed to hold with respect to a distribution of tasks. We show that the joint family of transitive PAC abstractions can be acquired efficiently, preserve near optimal-behavior, and experimentally reduce sample complexity in simple domains, thereby yielding a family of desirable abstractions for use in lifelong reinforcement learning. Along with these positive results, we show that there are pathological cases where state abstractions can negatively impact performance. David Abel, Dilip Arumugam, Lucas Lehnert, Michael L. Littman |
ICML | 2 |