VLDB 2026 Research / reviewers in the wild / expert
Clément Romac
dblp:237/9753
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 50% Knowledge representation and reasoning · 17% Learning paradigms · 16% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms
curriculum learning |
0.9 | 1 | 2025 | MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces · ICML 2025 |
Machine learning › Reinforcement learning
learning progress prediction |
0.9 | 1 | 2025 | MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces · ICML 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces · ICML 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
grounding |
0.7 | 1 | 2023 | Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning › online decision making
online reinforcement learning |
0.7 | 1 | 2023 | Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
automatic curriculum learning |
0.5 | 1 | 2021 | TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL · ICML 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
open-world learning |
0.3 | 1 | 2025 | MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces · ICML 2025 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.1 | 1 | 2021 | TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
online RL · 0.9metacognitive monitoring · 0.9online reinforcement learning · 0.7large language model · 0.7procedural task generation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spacesabstractOpen-ended learning agents must efficiently prioritize goals in vast possibility spaces, focusing on those that maximize learning progress (LP). When such autotelic exploration is achieved by LLM agents trained with online RL in high-dimensional and evolving goal spaces, a key challenge for LP prediction is modeling one’s own competence, a form of metacognitive monitoring. Traditional approaches either require extensive sampling or rely on brittle expert-defined goal groupings. We introduce MAGELLAN, a metacognitive framework that lets LLM agents learn to predict their competence and learning progress online. By capturing semantic relationships between goals, MAGELLAN enables sample-efficient LP estimation and dynamic adaptation to evolving goal spaces through generalization. In an interactive learning environment, we show that MAGELLAN improves LP prediction efficiency and goal prioritization, being the only method allowing the agent to fully master a large and evolving goal space. These results demonstrate how augmenting LLM agents with a metacognitive ability for LP predictions can effectively scale curriculum learning to open-ended goal spaces. Loris Gaven, Thomas Carta, Clément Romac, Cédric Colas, Sylvain Lamprier, Olivier Sigaud, Pierre-Yves Oudeyer |
ICML | 3 |
| 2023 | Grounding Large Language Models in Interactive Environments with Online Reinforcement LearningabstractRecent works successfully leveraged Large Language Models’ (LLM) abilities to capture abstract knowledge about world’s physics to solve decision-making problems. Yet, the alignment between LLMs’ knowledge and the environment can be wrong and limit functional competence due to lack of grounding. In this paper, we study an approach (named GLAM) to achieve this alignment through functional grounding: we consider an agent using an LLM as a policy that is progressively updated as the agent interacts with the environment, leveraging online Reinforcement Learning to improve its performance to solve goals. Using an interactive textual environment designed to study higher-level forms of functional grounding, and a set of spatial and navigation tasks, we study several scientific questions: 1) Can LLMs boost sample efficiency for online learning of various RL tasks? 2) How can it boost different forms of generalization? 3) What is the impact of online learning? We study these questions by functionally grounding several variants (size, architecture) of FLAN-T5. Thomas Carta, Clément Romac, Sylvain Lamprier, Olivier Sigaud, Pierre-Yves Oudeyer |
ICML | 2 |
| 2021 | TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RLabstractTraining autonomous agents able to generalize to multiple tasks is a key target of Deep Reinforcement Learning (DRL) research. In parallel to improving DRL algorithms themselves, Automatic Curriculum Learning (ACL) study how teacher algorithms can train DRL agents more efficiently by adapting task selection to their evolving abilities. While multiple standard benchmarks exist to compare DRL agents, there is currently no such thing for ACL algorithms. Thus, comparing existing approaches is difficult, as too many experimental parameters differ from paper to paper. In this work, we identify several key challenges faced by ACL algorithms. Based on these, we present TeachMyAgent (TA), a benchmark of current ACL algorithms leveraging procedural task generation. It includes 1) challenge-specific unit-tests using variants of a procedural Box2D bipedal walker environment, and 2) a new procedural Parkour environment combining most ACL challenges, making it ideal for global performance assessment. We then use TeachMyAgent to conduct a comparative study of representative existing approaches, showcasing the competitiveness of some ACL algorithms that do not use expert knowledge. We also show that the Parkour environment remains an open problem. We open-source our environments, all studied ACL algorithms (collected from open-source code or re-implemented), and DRL students in a Python package available at https://github.com/flowersteam/TeachMyAgent. Clément Romac, Rémy Portelas, Katja Hofmann, Pierre-Yves Oudeyer |
ICML | 1 |