Martin Klissarov

dblp:206/7079 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Reinforcement learning · 86% Graph learning · 7% Language models and text generation · 4%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
hierarchical reinforcement learning
1.332021
Flexible Option Learning · NeurIPS 2021
Options of Interest: Temporal Abstraction with Interest Functions · AAAI 2020
When Waiting Is Not an Option: Learning Options With a Deliberation Cost · AAAI 2018
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery
1.332021
Flexible Option Learning · NeurIPS 2021
Options of Interest: Temporal Abstraction with Interest Functions · AAAI 2020
When Waiting Is Not an Option: Learning Options With a Deliberation Cost · AAAI 2018
Machine learning › Reinforcement learning › reward design
reward shaping
1.222024
Code as Reward: Empowering Reinforcement Learning with VLMs · ICML 2024
Reward Propagation Using Graph Convolutional Networks · NeurIPS 2020
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
On the Modeling Capabilities of Large Language Models for Sequential Decision Making · ICLR 2025
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.812024
Motif: Intrinsic Motivation from Artificial Intelligence Feedback · ICLR 2024
Machine learning › Reinforcement learning › reward learning
vision-language model reward
0.812024
Code as Reward: Empowering Reinforcement Learning with VLMs · ICML 2024
Compilers and program optimization
code generation
0.812024
Code as Reward: Empowering Reinforcement Learning with VLMs · ICML 2024
Machine learning › Reinforcement learning
deep reinforcement learning
0.712023
Deep Laplacian-based Options for Temporally-Extended Exploration · ICML 2023
Machine learning › Reinforcement learning
exploration
0.712023
Deep Laplacian-based Options for Temporally-Extended Exploration · ICML 2023
Machine learning › Reinforcement learning › meta-reinforcement learning
meta-gradient reinforcement learning
0.612022
Adaptive Interest for Emphatic Reinforcement Learning · NeurIPS 2022
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option-critic
0.512021
Flexible Option Learning · NeurIPS 2021
Machine learning › Graph learning › graph neural network
graph convolutional network
0.412020
Reward Propagation Using Graph Convolutional Networks · NeurIPS 2020
Machine learning › Reinforcement learning › hierarchical reinforcement learning
initiation set learning
0.412020
Options of Interest: Temporal Abstraction with Interest Functions · AAAI 2020
Machine learning › Graph learning › graph neural network
message passing
0.412020
Reward Propagation Using Graph Convolutional Networks · NeurIPS 2020
Machine learning › Reinforcement learning › hierarchical reinforcement learning
options framework
0.412020
Options of Interest: Temporal Abstraction with Interest Functions · AAAI 2020
Machine learning › Reinforcement learning › reward design › reward shaping
potential-based reward shaping
0.412020
Reward Propagation Using Graph Convolutional Networks · NeurIPS 2020
Knowledge, reasoning and agents › Multi-agent systems
bounded rationality
0.312018
When Waiting Is Not an Option: Learning Options With a Deliberation Cost · AAAI 2018
Machine learning › Trustworthy machine learning › interpretability
explainable reinforcement learning
0.112020
Options of Interest: Temporal Abstraction with Interest Functions · AAAI 2020

Methods — techniques the papers use, named apart from their topics

large language model · 2.5reinforcement learning · 1.6vision-language model · 1.5code generation · 1.5fine-tuning · 0.9graph laplacian · 0.7eigendecomposition · 0.7deep reinforcement learning · 0.7meta-gradient · 0.6off-policy learning · 0.5
YearPublicationVenuePosition
2025 MaestroMotif: Skill Design from Artificial Intelligence Feedback
abstract
Describing skills in natural language has the potential to provide an accessible way to inject human knowledge about decision-making into an AI system. We present MaestroMotif, a method for AI-assisted skill design, which yields high-performing and adaptable agents. MaestroMotif leverages the capabilities of Large Language Models (LLMs) to effectively create and reuse skills. It first uses an LLM's feedback to automatically design rewards corresponding to each skill, starting from their natural language description. Then, it employs an LLM's code generation abilities, together with reinforcement learning, for training the skills and combining them to implement complex behaviors specified in language. We evaluate MaestroMotif using a suite of complex tasks in the NetHack Learning Environment (NLE), demonstrating that it surpasses existing approaches in both performance and usability.
Martin Klissarov, Mikael Henaff, Roberta Raileanu, Shagun Sodhani, Pascal Vincent, Amy Zhang 0001, Pierre-Luc Bacon, Doina Precup, Marlos C. Machado, Pierluca D'Oro
ICLR1
2025 On the Modeling Capabilities of Large Language Models for Sequential Decision Making
abstract
Large pretrained models are showing increasingly better performance in reasoning and planning tasks across different modalities, opening the possibility to leverage them for complex sequential decision making problems. In this paper, we investigate the capabilities of Large Language Models (LLMs) for reinforcement learning (RL) across a diversity of interactive domains. We evaluate their ability to produce decision-making policies, either directly, by generating actions, or indirectly, by first generating reward models to train an agent with RL. Our results show that, even without task-specific fine-tuning, LLMs excel at reward modeling. In particular, crafting rewards through artificial intelligence (AI) feedback yields the most generally applicable approach and can enhance performance by improving credit assignment and exploration. Finally, in environments with unfamiliar dynamics, we explore how fine-tuning LLMs with synthetic data can significantly improve their reward modeling capabilities while mitigating catastrophic forgetting, further broadening their utility in sequential decision-making tasks.
Martin Klissarov, R. Devon Hjelm, Alexander Toshev, Bogdan Mazoure
ICLR1
2024 Motif: Intrinsic Motivation from Artificial Intelligence Feedback
abstract
Exploring rich environments and evaluating one's actions without prior knowledge is immensely challenging. In this paper, we propose Motif, a general method to interface such prior knowledge from a Large Language Model (LLM) with an agent. Motif is based on the idea of grounding LLMs for decision-making without requiring them to interact with the environment: it elicits preferences from an LLM over pairs of captions to construct an intrinsic reward, which is then used to train agents with reinforcement learning. We evaluate Motif's performance and behavior on the challenging, open-ended and procedurally-generated NetHack game. Surprisingly, by only learning to maximize its intrinsic reward, Motif achieves a higher game score than an algorithm directly trained to maximize the score itself. When combining Motif's intrinsic reward with the environment reward, our method significantly outperforms existing approaches and makes progress on tasks where no advancements have ever been made without demonstrations. Finally, we show that Motif mostly generates intuitive human-aligned behaviors which can be steered easily through prompt modifications, while scaling well with the LLM size and the amount of information given in the prompt.
Martin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang 0001, Mikael Henaff
ICLR1
2024 Code as Reward: Empowering Reinforcement Learning with VLMs
abstract
Pre-trained Vision-Language Models (VLMs) are able to understand visual concepts, describe and decompose complex tasks into sub-tasks, and provide feedback on task completion. In this paper, we aim to leverage these capabilities to support the training of reinforcement learning (RL) agents. In principle, VLMs are well suited for this purpose, as they can naturally analyze image-based observations and provide feedback (reward) on learning progress. However, inference in VLMs is computationally expensive, so querying them frequently to compute rewards would significantly slowdown the training of an RL agent. To address this challenge, we propose a framework named Code as Reward (VLM-CaR). VLM-CaR produces dense reward functions from VLMs through code generation, thereby significantly reducing the computational burden of querying the VLM directly. We show that the dense rewards generated through our approach are very accurate across a diverse set of discrete and continuous environments, and can be more effective in training RL policies than the original sparse environment rewards.
David Venuto, Mohammad Sami Nur Islam, Martin Klissarov, Doina Precup, Sherry Yang 0001, Ankit Anand
ICML3
2023 Deep Laplacian-based Options for Temporally-Extended Exploration
abstract
Selecting exploratory actions that generate a rich stream of experience for better learning is a fundamental challenge in reinforcement learning (RL). An approach to tackle this problem consists in selecting actions according to specific policies for an extended period of time, also known as options. A recent line of work to derive such exploratory options builds upon the eigenfunctions of the graph Laplacian. Importantly, until now these methods have been mostly limited to tabular domains where (1) the graph Laplacian matrix was either given or could be fully estimated, (2) performing eigendecomposition on this matrix was computationally tractable, and (3) value functions could be learned exactly. Additionally, these methods required a separate option discovery phase. These assumptions are fundamentally not scalable. In this paper we address these limitations and show how recent results for directly approximating the eigenfunctions of the Laplacian can be leveraged to truly scale up options-based exploration. To do so, we introduce a fully online deep RL algorithm for discovering Laplacian-based options and evaluate our approach on a variety of pixel-based tasks. We compare to several state-of-the-art exploration methods and show that our approach is effective, general, and especially promising in non-stationary settings.
Martin Klissarov, Marlos C. Machado
ICML1
2022 Adaptive Interest for Emphatic Reinforcement Learning
abstract
Emphatic algorithms have shown great promise in stabilizing and improving reinforcement learning by selectively emphasizing the update rule. Although the emphasis fundamentally depends on an interest function which defines the intrinsic importance of each state, most approaches simply adopt a uniform interest over all states (except where a hand-designed interest is possible based on domain knowledge). In this paper, we investigate adaptive methods that allow the interest function to dynamically vary over states and iterations. In particular, we leverage meta-gradients to automatically discover online an interest function that would accelerate the agent’s learning process. Empirical evaluations on a wide range of environments show that adapting the interest is key to provide significant gains. Qualitative analysis indicates that the learned interest function emphasizes states of particular importance, such as bottlenecks, which can be especially useful in a transfer learning setting.
Martin Klissarov, Rasool Fakoor, Jonas Mueller 0001, Kavosh Asadi, Taesup Kim, Alexander J. Smola
NeurIPS1
2021 Flexible Option Learning
abstract
Temporal abstraction in reinforcement learning (RL), offers the promise of improving generalization and knowledge transfer in complex environments, by propagating information more efficiently over time. Although option learning was initially formulated in a way that allows updating many options simultaneously, using off-policy, intra-option learning (Sutton, Precup & Singh, 1999) , many of the recent hierarchical reinforcement learning approaches only update a single option at a time: the option currently executing. We revisit and extend intra-option learning in the context of deep reinforcement learning, in order to enable updating all options consistent with current primitive action choices, without introducing any additional estimates. Our method can therefore be naturally adopted in most hierarchical RL frameworks. When we combine our approach with the option-critic algorithm for option discovery, we obtain significant improvements in performance and data-efficiency across a wide variety of domains.
Martin Klissarov, Doina Precup
NeurIPS1
2020 Options of Interest: Temporal Abstraction with Interest Functions
abstract
Temporal abstraction refers to the ability of an agent to use behaviours of controllers which act for a limited, variable amount of time. The options framework describes such behaviours as consisting of a subset of states in which they can initiate, an internal policy and a stochastic termination condition. However, much of the subsequent work on option discovery has ignored the initiation set, because of difficulty in learning it from data. We provide a generalization of initiation sets suitable for general function approximation, by defining an interest function associated with an option. We derive a gradient-based learning algorithm for interest functions, leading to a new interest-option-critic architecture. We investigate how interest functions can be leveraged to learn interpretable and reusable temporal abstractions. We demonstrate the efficacy of the proposed approach through quantitative and qualitative results, in both discrete and continuous environments.
Khimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon, Doina Precup
AAAI2
2020 Reward Propagation Using Graph Convolutional Networks
abstract
Potential-based reward shaping provides an approach for designing good reward functions, with the purpose of speeding up learning. However, automatically finding potential functions for complex environments is a difficult problem (in fact, of the same difficulty as learning a value function from scratch). We propose a new framework for learning potential functions by leveraging ideas from graph representation learning. Our approach relies on Graph Convolutional Networks which we use as a key ingredient in combination with the probabilistic inference view of reinforcement learning. More precisely, we leverage Graph Convolutional Networks to perform message passing from rewarding states. The propagated messages can then be used as potential functions for reward shaping to accelerate learning. We verify empirically that our approach can achieve considerable improvements in both small and high-dimensional control problems.
Martin Klissarov, Doina Precup
NeurIPS1
2018 When Waiting Is Not an Option: Learning Options With a Deliberation Cost
abstract
Recent work has shown that temporally extended actions (options) can be learned fully end-to-end as opposed to being specified in advance. While the problem of how to learn options is increasingly well understood, the question of what good options should be has remained elusive. We formulate our answer to what good options should be in the bounded rationality framework (Simon, 1957) through the notion of deliberation cost. We then derive practical gradient-based learning algorithms to implement this objective. Our results in the Arcade Learning Environment (ALE) show increased performance and interpretability.
Jean Harb, Pierre-Luc Bacon, Martin Klissarov, Doina Precup
AAAI3