EDBT 2026 Demo / reviewers in the wild / expert
Théo Cachet
dblp:296/9387
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 94% Vision and language · 6% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
goal-conditioned policy |
0.8 | 1 | 2024 | Bridging Environments and Language with Rendering Functions and Vision-Language Models · ICML 2024 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
0.8 | 1 | 2024 | Bridging Environments and Language with Rendering Functions and Vision-Language Models · ICML 2024 |
Machine learning › Reinforcement learning › reward learning
vision-language model reward |
0.8 | 1 | 2024 | Bridging Environments and Language with Rendering Functions and Vision-Language Models · ICML 2024 |
Machine learning › Reinforcement learning › imitation learning
few-shot imitation learning |
0.5 | 1 | 2021 | Demonstration-Conditioned Reinforcement Learning for Few-Shot Imitation · ICML 2021 |
Machine learning › Reinforcement learning
imitation learning |
0.5 | 1 | 2021 | Demonstration-Conditioned Reinforcement Learning for Few-Shot Imitation · ICML 2021 |
Computer vision › Vision and language
vision-language model |
0.2 | 1 | 2024 | Bridging Environments and Language with Rendering Functions and Vision-Language Models · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 0.8knowledge distillation · 0.8goal-conditioned policy · 0.8reward inference · 0.5behavior cloning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Bridging Environments and Language with Rendering Functions and Vision-Language ModelsabstractVision-language models (VLMs) have tremendous potential for grounding language, and thus enabling language-conditioned agents (LCAs) to perform diverse tasks specified with text. This has motivated the study of LCAs based on reinforcement learning (RL) with rewards given by rendering images of an environment and evaluating those images with VLMs. If single-task RL is employed, such approaches are limited by the cost and time required to train a policy for each new task. Multi-task RL (MTRL) is a natural alternative, but requires a carefully designed corpus of training tasks and does not always generalize reliably to new tasks. Therefore, this paper introduces a novel decomposition of the problem of building an LCA: first find an environment configuration that has a high VLM score for text describing a task; then use a (pretrained) goal-conditioned policy to reach that configuration. We also explore several enhancements to the speed and quality of VLM-based LCAs, notably, the use of distilled models, and the evaluation of configurations from multiple viewpoints to resolve the ambiguities inherent in a single 2D view. We demonstrate our approach on the Humanoid environment, showing that it results in LCAs that outperform MTRL baselines in zero-shot generalization, without requiring any textual task descriptions or other forms of environment-specific annotation during training. Théo Cachet, Christopher R. Dance, Olivier Sigaud |
ICML | 1 |
| 2021 | Demonstration-Conditioned Reinforcement Learning for Few-Shot ImitationabstractIn few-shot imitation, an agent is given a few demonstrations of a previously unseen task, and must then successfully perform that task. We propose a novel approach to learning few-shot-imitation agents that we call demonstration-conditioned reinforcement learning (DCRL). Given a training set consisting of demonstrations, reward functions and transition distributions for multiple tasks, the idea is to work with a policy that takes demonstrations as input, and to train this policy to maximize the average of the cumulative reward over the set of training tasks. Relative to previously proposed few-shot imitation methods that use behaviour cloning or infer reward functions from demonstrations, our method has the disadvantage that it requires reward functions at training time. However, DCRL also has several advantages, such as the ability to improve upon suboptimal demonstrations, to operate given state-only demonstrations, and to cope with a domain shift between the demonstrator and the agent. Moreover, we show that DCRL outperforms methods based on behaviour cloning by a large margin, on navigation tasks and on robotic manipulation tasks from the Meta-World benchmark. Christopher R. Dance, Julien Perez, Théo Cachet |
ICML | 3 |