EDBT 2026 Demo / reviewers in the wild / expert
Martin Klissarov
dblp:206/7079
· DBLP profile ↗
10ranked-venue papers
7as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Reinforcement learning · 86% Graph learning · 7% Language models and text generation · 4% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
1.3 | 3 | 2021 | Flexible Option Learning · NeurIPS 2021 Options of Interest: Temporal Abstraction with Interest Functions · AAAI 2020 When Waiting Is Not an Option: Learning Options With a Deliberation Cost · AAAI 2018 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
1.3 | 3 | 2021 | Flexible Option Learning · NeurIPS 2021 Options of Interest: Temporal Abstraction with Interest Functions · AAAI 2020 When Waiting Is Not an Option: Learning Options With a Deliberation Cost · AAAI 2018 |
Machine learning › Reinforcement learning › reward design
reward shaping |
1.2 | 2 | 2024 | Code as Reward: Empowering Reinforcement Learning with VLMs · ICML 2024 Reward Propagation Using Graph Convolutional Networks · NeurIPS 2020 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | On the Modeling Capabilities of Large Language Models for Sequential Decision Making · ICLR 2025 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.8 | 1 | 2024 | Motif: Intrinsic Motivation from Artificial Intelligence Feedback · ICLR 2024 |
Machine learning › Reinforcement learning › reward learning
vision-language model reward |
0.8 | 1 | 2024 | Code as Reward: Empowering Reinforcement Learning with VLMs · ICML 2024 |
Compilers and program optimization
code generation |
0.8 | 1 | 2024 | Code as Reward: Empowering Reinforcement Learning with VLMs · ICML 2024 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.7 | 1 | 2023 | Deep Laplacian-based Options for Temporally-Extended Exploration · ICML 2023 |
Machine learning › Reinforcement learning
exploration |
0.7 | 1 | 2023 | Deep Laplacian-based Options for Temporally-Extended Exploration · ICML 2023 |
Machine learning › Reinforcement learning › meta-reinforcement learning
meta-gradient reinforcement learning |
0.6 | 1 | 2022 | Adaptive Interest for Emphatic Reinforcement Learning · NeurIPS 2022 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option-critic |
0.5 | 1 | 2021 | Flexible Option Learning · NeurIPS 2021 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.4 | 1 | 2020 | Reward Propagation Using Graph Convolutional Networks · NeurIPS 2020 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
initiation set learning |
0.4 | 1 | 2020 | Options of Interest: Temporal Abstraction with Interest Functions · AAAI 2020 |
Machine learning › Graph learning › graph neural network
message passing |
0.4 | 1 | 2020 | Reward Propagation Using Graph Convolutional Networks · NeurIPS 2020 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
options framework |
0.4 | 1 | 2020 | Options of Interest: Temporal Abstraction with Interest Functions · AAAI 2020 |
Machine learning › Reinforcement learning › reward design › reward shaping
potential-based reward shaping |
0.4 | 1 | 2020 | Reward Propagation Using Graph Convolutional Networks · NeurIPS 2020 |
Knowledge, reasoning and agents › Multi-agent systems
bounded rationality |
0.3 | 1 | 2018 | When Waiting Is Not an Option: Learning Options With a Deliberation Cost · AAAI 2018 |
Machine learning › Trustworthy machine learning › interpretability
explainable reinforcement learning |
0.1 | 1 | 2020 | Options of Interest: Temporal Abstraction with Interest Functions · AAAI 2020 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.5reinforcement learning · 1.6vision-language model · 1.5code generation · 1.5fine-tuning · 0.9graph laplacian · 0.7eigendecomposition · 0.7deep reinforcement learning · 0.7meta-gradient · 0.6off-policy learning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MaestroMotif: Skill Design from Artificial Intelligence FeedbackabstractDescribing skills in natural language has the potential to provide an accessible way to inject human knowledge about decision-making into an AI system. We present MaestroMotif, a method for AI-assisted skill design, which yields high-performing and adaptable agents. MaestroMotif leverages the capabilities of Large Language Models (LLMs) to effectively create and reuse skills. It first uses an LLM's feedback to automatically design rewards corresponding to each skill, starting from their natural language description. Then, it employs an LLM's code generation abilities, together with reinforcement learning, for training the skills and combining them to implement complex behaviors specified in language. We evaluate MaestroMotif using a suite of complex tasks in the NetHack Learning Environment (NLE), demonstrating that it surpasses existing approaches in both performance and usability. Martin Klissarov, Mikael Henaff, Roberta Raileanu, Shagun Sodhani, Pascal Vincent, Amy Zhang 0001, Pierre-Luc Bacon, Doina Precup, Marlos C. Machado, Pierluca D'Oro |
ICLR | 1 |
| 2025 | On the Modeling Capabilities of Large Language Models for Sequential Decision MakingabstractLarge pretrained models are showing increasingly better performance in reasoning and planning tasks across different modalities, opening the possibility to leverage them for complex sequential decision making problems. In this paper, we investigate the capabilities of Large Language Models (LLMs) for reinforcement learning (RL) across a diversity of interactive domains. We evaluate their ability to produce decision-making policies, either directly, by generating actions, or indirectly, by first generating reward models to train an agent with RL. Our results show that, even without task-specific fine-tuning, LLMs excel at reward modeling. In particular, crafting rewards through artificial intelligence (AI) feedback yields the most generally applicable approach and can enhance performance by improving credit assignment and exploration. Finally, in environments with unfamiliar dynamics, we explore how fine-tuning LLMs with synthetic data can significantly improve their reward modeling capabilities while mitigating catastrophic forgetting, further broadening their utility in sequential decision-making tasks. Martin Klissarov, R. Devon Hjelm, Alexander Toshev, Bogdan Mazoure |
ICLR | 1 |
| 2024 | Motif: Intrinsic Motivation from Artificial Intelligence FeedbackabstractExploring rich environments and evaluating one's actions without prior knowledge is immensely challenging. In this paper, we propose Motif, a general method to interface such prior knowledge from a Large Language Model (LLM) with an agent. Motif is based on the idea of grounding LLMs for decision-making without requiring them to interact with the environment: it elicits preferences from an LLM over pairs of captions to construct an intrinsic reward, which is then used to train agents with reinforcement learning. We evaluate Motif's performance and behavior on the challenging, open-ended and procedurally-generated NetHack game. Surprisingly, by only learning to maximize its intrinsic reward, Motif achieves a higher game score than an algorithm directly trained to maximize the score itself. When combining Motif's intrinsic reward with the environment reward, our method significantly outperforms existing approaches and makes progress on tasks where no advancements have ever been made without demonstrations. Finally, we show that Motif mostly generates intuitive human-aligned behaviors which can be steered easily through prompt modifications, while scaling well with the LLM size and the amount of information given in the prompt. Martin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang 0001, Mikael Henaff |
ICLR | 1 |
| 2024 | Code as Reward: Empowering Reinforcement Learning with VLMsabstractPre-trained Vision-Language Models (VLMs) are able to understand visual concepts, describe and decompose complex tasks into sub-tasks, and provide feedback on task completion. In this paper, we aim to leverage these capabilities to support the training of reinforcement learning (RL) agents. In principle, VLMs are well suited for this purpose, as they can naturally analyze image-based observations and provide feedback (reward) on learning progress. However, inference in VLMs is computationally expensive, so querying them frequently to compute rewards would significantly slowdown the training of an RL agent. To address this challenge, we propose a framework named Code as Reward (VLM-CaR). VLM-CaR produces dense reward functions from VLMs through code generation, thereby significantly reducing the computational burden of querying the VLM directly. We show that the dense rewards generated through our approach are very accurate across a diverse set of discrete and continuous environments, and can be more effective in training RL policies than the original sparse environment rewards. David Venuto, Mohammad Sami Nur Islam, Martin Klissarov, Doina Precup, Sherry Yang 0001, Ankit Anand |
ICML | 3 |
| 2023 | Deep Laplacian-based Options for Temporally-Extended ExplorationabstractSelecting exploratory actions that generate a rich stream of experience for better learning is a fundamental challenge in reinforcement learning (RL). An approach to tackle this problem consists in selecting actions according to specific policies for an extended period of time, also known as options. A recent line of work to derive such exploratory options builds upon the eigenfunctions of the graph Laplacian. Importantly, until now these methods have been mostly limited to tabular domains where (1) the graph Laplacian matrix was either given or could be fully estimated, (2) performing eigendecomposition on this matrix was computationally tractable, and (3) value functions could be learned exactly. Additionally, these methods required a separate option discovery phase. These assumptions are fundamentally not scalable. In this paper we address these limitations and show how recent results for directly approximating the eigenfunctions of the Laplacian can be leveraged to truly scale up options-based exploration. To do so, we introduce a fully online deep RL algorithm for discovering Laplacian-based options and evaluate our approach on a variety of pixel-based tasks. We compare to several state-of-the-art exploration methods and show that our approach is effective, general, and especially promising in non-stationary settings. Martin Klissarov, Marlos C. Machado |
ICML | 1 |
| 2022 | Adaptive Interest for Emphatic Reinforcement LearningabstractEmphatic algorithms have shown great promise in stabilizing and improving reinforcement learning by selectively emphasizing the update rule. Although the emphasis fundamentally depends on an interest function which defines the intrinsic importance of each state, most approaches simply adopt a uniform interest over all states (except where a hand-designed interest is possible based on domain knowledge). In this paper, we investigate adaptive methods that allow the interest function to dynamically vary over states and iterations. In particular, we leverage meta-gradients to automatically discover online an interest function that would accelerate the agent’s learning process. Empirical evaluations on a wide range of environments show that adapting the interest is key to provide significant gains. Qualitative analysis indicates that the learned interest function emphasizes states of particular importance, such as bottlenecks, which can be especially useful in a transfer learning setting. Martin Klissarov, Rasool Fakoor, Jonas Mueller 0001, Kavosh Asadi, Taesup Kim, Alexander J. Smola |
NeurIPS | 1 |
| 2021 | Flexible Option LearningabstractTemporal abstraction in reinforcement learning (RL), offers the promise of improving generalization and knowledge transfer in complex environments, by propagating information more efficiently over time. Although option learning was initially formulated in a way that allows updating many options simultaneously, using off-policy, intra-option learning (Sutton, Precup & Singh, 1999) , many of the recent hierarchical reinforcement learning approaches only update a single option at a time: the option currently executing. We revisit and extend intra-option learning in the context of deep reinforcement learning, in order to enable updating all options consistent with current primitive action choices, without introducing any additional estimates. Our method can therefore be naturally adopted in most hierarchical RL frameworks. When we combine our approach with the option-critic algorithm for option discovery, we obtain significant improvements in performance and data-efficiency across a wide variety of domains. Martin Klissarov, Doina Precup |
NeurIPS | 1 |
| 2020 | Options of Interest: Temporal Abstraction with Interest FunctionsabstractTemporal abstraction refers to the ability of an agent to use behaviours of controllers which act for a limited, variable amount of time. The options framework describes such behaviours as consisting of a subset of states in which they can initiate, an internal policy and a stochastic termination condition. However, much of the subsequent work on option discovery has ignored the initiation set, because of difficulty in learning it from data. We provide a generalization of initiation sets suitable for general function approximation, by defining an interest function associated with an option. We derive a gradient-based learning algorithm for interest functions, leading to a new interest-option-critic architecture. We investigate how interest functions can be leveraged to learn interpretable and reusable temporal abstractions. We demonstrate the efficacy of the proposed approach through quantitative and qualitative results, in both discrete and continuous environments. Khimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon, Doina Precup |
AAAI | 2 |
| 2020 | Reward Propagation Using Graph Convolutional NetworksabstractPotential-based reward shaping provides an approach for designing good reward functions, with the purpose of speeding up learning. However, automatically finding potential functions for complex environments is a difficult problem (in fact, of the same difficulty as learning a value function from scratch). We propose a new framework for learning potential functions by leveraging ideas from graph representation learning. Our approach relies on Graph Convolutional Networks which we use as a key ingredient in combination with the probabilistic inference view of reinforcement learning. More precisely, we leverage Graph Convolutional Networks to perform message passing from rewarding states. The propagated messages can then be used as potential functions for reward shaping to accelerate learning. We verify empirically that our approach can achieve considerable improvements in both small and high-dimensional control problems. Martin Klissarov, Doina Precup |
NeurIPS | 1 |
| 2018 | When Waiting Is Not an Option: Learning Options With a Deliberation CostabstractRecent work has shown that temporally extended actions (options) can be learned fully end-to-end as opposed to being specified in advance. While the problem of how to learn options is increasingly well understood, the question of what good options should be has remained elusive. We formulate our answer to what good options should be in the bounded rationality framework (Simon, 1957) through the notion of deliberation cost. We then derive practical gradient-based learning algorithms to implement this objective. Our results in the Arcade Learning Environment (ALE) show increased performance and interpretability. Jean Harb, Pierre-Luc Bacon, Martin Klissarov, Doina Precup |
AAAI | 3 |