Clément Romac

dblp:237/9753 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 50% Knowledge representation and reasoning · 17% Learning paradigms · 16%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning paradigms
curriculum learning
0.912025
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces · ICML 2025
Machine learning › Reinforcement learning
learning progress prediction
0.912025
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces · ICML 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces · ICML 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
grounding
0.712023
Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › online decision making
online reinforcement learning
0.712023
Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
automatic curriculum learning
0.512021
TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL · ICML 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
open-world learning
0.312025
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces · ICML 2025
Machine learning › Reinforcement learning
deep reinforcement learning
0.112021
TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL · ICML 2021

Methods — techniques the papers use, named apart from their topics

online RL · 0.9metacognitive monitoring · 0.9online reinforcement learning · 0.7large language model · 0.7procedural task generation · 0.5
YearPublicationVenuePosition
2025 MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
abstract
Open-ended learning agents must efficiently prioritize goals in vast possibility spaces, focusing on those that maximize learning progress (LP). When such autotelic exploration is achieved by LLM agents trained with online RL in high-dimensional and evolving goal spaces, a key challenge for LP prediction is modeling one’s own competence, a form of metacognitive monitoring. Traditional approaches either require extensive sampling or rely on brittle expert-defined goal groupings. We introduce MAGELLAN, a metacognitive framework that lets LLM agents learn to predict their competence and learning progress online. By capturing semantic relationships between goals, MAGELLAN enables sample-efficient LP estimation and dynamic adaptation to evolving goal spaces through generalization. In an interactive learning environment, we show that MAGELLAN improves LP prediction efficiency and goal prioritization, being the only method allowing the agent to fully master a large and evolving goal space. These results demonstrate how augmenting LLM agents with a metacognitive ability for LP predictions can effectively scale curriculum learning to open-ended goal spaces.
Loris Gaven, Thomas Carta, Clément Romac, Cédric Colas, Sylvain Lamprier, Olivier Sigaud, Pierre-Yves Oudeyer
ICML3
2023 Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning
abstract
Recent works successfully leveraged Large Language Models’ (LLM) abilities to capture abstract knowledge about world’s physics to solve decision-making problems. Yet, the alignment between LLMs’ knowledge and the environment can be wrong and limit functional competence due to lack of grounding. In this paper, we study an approach (named GLAM) to achieve this alignment through functional grounding: we consider an agent using an LLM as a policy that is progressively updated as the agent interacts with the environment, leveraging online Reinforcement Learning to improve its performance to solve goals. Using an interactive textual environment designed to study higher-level forms of functional grounding, and a set of spatial and navigation tasks, we study several scientific questions: 1) Can LLMs boost sample efficiency for online learning of various RL tasks? 2) How can it boost different forms of generalization? 3) What is the impact of online learning? We study these questions by functionally grounding several variants (size, architecture) of FLAN-T5.
Thomas Carta, Clément Romac, Sylvain Lamprier, Olivier Sigaud, Pierre-Yves Oudeyer
ICML2
2021 TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL
abstract
Training autonomous agents able to generalize to multiple tasks is a key target of Deep Reinforcement Learning (DRL) research. In parallel to improving DRL algorithms themselves, Automatic Curriculum Learning (ACL) study how teacher algorithms can train DRL agents more efficiently by adapting task selection to their evolving abilities. While multiple standard benchmarks exist to compare DRL agents, there is currently no such thing for ACL algorithms. Thus, comparing existing approaches is difficult, as too many experimental parameters differ from paper to paper. In this work, we identify several key challenges faced by ACL algorithms. Based on these, we present TeachMyAgent (TA), a benchmark of current ACL algorithms leveraging procedural task generation. It includes 1) challenge-specific unit-tests using variants of a procedural Box2D bipedal walker environment, and 2) a new procedural Parkour environment combining most ACL challenges, making it ideal for global performance assessment. We then use TeachMyAgent to conduct a comparative study of representative existing approaches, showcasing the competitiveness of some ACL algorithms that do not use expert knowledge. We also show that the Parkour environment remains an open problem. We open-source our environments, all studied ACL algorithms (collected from open-source code or re-implemented), and DRL students in a Python package available at https://github.com/flowersteam/TeachMyAgent.
Clément Romac, Rémy Portelas, Katja Hofmann, Pierre-Yves Oudeyer
ICML1