Théo Cachet

dblp:296/9387 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 94% Vision and language · 6%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
goal-conditioned policy
0.812024
Bridging Environments and Language with Rendering Functions and Vision-Language Models · ICML 2024
Machine learning › Reinforcement learning
multi-task reinforcement learning
0.812024
Bridging Environments and Language with Rendering Functions and Vision-Language Models · ICML 2024
Machine learning › Reinforcement learning › reward learning
vision-language model reward
0.812024
Bridging Environments and Language with Rendering Functions and Vision-Language Models · ICML 2024
Machine learning › Reinforcement learning › imitation learning
few-shot imitation learning
0.512021
Demonstration-Conditioned Reinforcement Learning for Few-Shot Imitation · ICML 2021
Machine learning › Reinforcement learning
imitation learning
0.512021
Demonstration-Conditioned Reinforcement Learning for Few-Shot Imitation · ICML 2021
Computer vision › Vision and language
vision-language model
0.212024
Bridging Environments and Language with Rendering Functions and Vision-Language Models · ICML 2024

Methods — techniques the papers use, named apart from their topics

vision-language model · 0.8knowledge distillation · 0.8goal-conditioned policy · 0.8reward inference · 0.5behavior cloning · 0.5
YearPublicationVenuePosition
2024 Bridging Environments and Language with Rendering Functions and Vision-Language Models
abstract
Vision-language models (VLMs) have tremendous potential for grounding language, and thus enabling language-conditioned agents (LCAs) to perform diverse tasks specified with text. This has motivated the study of LCAs based on reinforcement learning (RL) with rewards given by rendering images of an environment and evaluating those images with VLMs. If single-task RL is employed, such approaches are limited by the cost and time required to train a policy for each new task. Multi-task RL (MTRL) is a natural alternative, but requires a carefully designed corpus of training tasks and does not always generalize reliably to new tasks. Therefore, this paper introduces a novel decomposition of the problem of building an LCA: first find an environment configuration that has a high VLM score for text describing a task; then use a (pretrained) goal-conditioned policy to reach that configuration. We also explore several enhancements to the speed and quality of VLM-based LCAs, notably, the use of distilled models, and the evaluation of configurations from multiple viewpoints to resolve the ambiguities inherent in a single 2D view. We demonstrate our approach on the Humanoid environment, showing that it results in LCAs that outperform MTRL baselines in zero-shot generalization, without requiring any textual task descriptions or other forms of environment-specific annotation during training.
Théo Cachet, Christopher R. Dance, Olivier Sigaud
ICML1
2021 Demonstration-Conditioned Reinforcement Learning for Few-Shot Imitation
abstract
In few-shot imitation, an agent is given a few demonstrations of a previously unseen task, and must then successfully perform that task. We propose a novel approach to learning few-shot-imitation agents that we call demonstration-conditioned reinforcement learning (DCRL). Given a training set consisting of demonstrations, reward functions and transition distributions for multiple tasks, the idea is to work with a policy that takes demonstrations as input, and to train this policy to maximize the average of the cumulative reward over the set of training tasks. Relative to previously proposed few-shot imitation methods that use behaviour cloning or infer reward functions from demonstrations, our method has the disadvantage that it requires reward functions at training time. However, DCRL also has several advantages, such as the ability to improve upon suboptimal demonstrations, to operate given state-only demonstrations, and to cope with a domain shift between the demonstrator and the agent. Moreover, we show that DCRL outperforms methods based on behaviour cloning by a large margin, on navigation tasks and on robotic manipulation tasks from the Meta-World benchmark.
Christopher R. Dance, Julien Perez, Théo Cachet
ICML3