Rémy Portelas

dblp:251/3113 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 56% Motion planning and robot control · 12% Language models and text generation · 10%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
automatic curriculum learning
1.532022
Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022
TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL · ICML 2021
Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020
Robotics › Motion planning and robot control
robot learning
1.422025
Efficient Active Imitation Learning with Random Network Distillation · ICLR 2025
Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning › imitation learning › interactive imitation learning
active imitation learning
0.912025
Efficient Active Imitation Learning with Random Network Distillation · ICLR 2025
Machine learning › Trustworthy machine learning › fairness › algorithmic bias
bias amplification
0.912025
Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? · EMNLP 2025
Machine learning › Reinforcement learning
imitation learning
0.912025
Efficient Active Imitation Learning with Random Network Distillation · ICLR 2025
Machine learning › Generative modeling
model collapse
0.912025
Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? · EMNLP 2025
Natural language and speech › Language models and text generation
synthetic data
0.912025
Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? · EMNLP 2025
Machine learning › Reinforcement learning
deep reinforcement learning
0.622021
Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020
TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL · ICML 2021
Machine learning › Learning paradigms
curriculum learning
0.612022
Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning
exploration
0.612022
Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning › exploration › directed exploration
goal-directed exploration
0.612022
Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.612022
Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning
sample efficiency
0.412020
Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
domain randomization
0.112020
Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.112020
Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.112020
Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020

Methods — techniques the papers use, named apart from their topics

regression analysis · 0.9random network distillation · 0.9DAgger · 0.9population-based policy · 0.6policy search · 0.6procedural task generation · 0.5domain randomization · 0.4curriculum learning · 0.4
YearPublicationVenuePosition
2025 A Continual Offline Reinforcement Learning Benchmark for Navigation Tasks
abstract
Autonomous agents operating in domains such as robotics or video game simulations must adapt to changing tasks without forgetting about the previous ones. This process called Continual Reinforcement Learning poses non-trivial difficulties, from preventing catastrophic forgetting to ensuring the scalability of the approaches considered. Building on recent advances, we introduce a benchmark providing a suite of video-game navigation scenarios, thus filling a gap in the literature and capturing key challenges: catastrophic forgetting, task adaptation, and memory efficiency. We define a set of various tasks and datasets, evaluation protocols, and metrics to assess the performance of algorithms, including state-of-the-art baselines. Our benchmark is designed not only to foster reproducible research and to accelerate progress in continual reinforcement learning for gaming, but also to provide a reproducible framework for production pipelines helping practitioners to identify and to apply effective approaches. https://sites.google.com/view/continual-nav-bench
Anthony Kobanda, Odalric-Ambrym Maillard, Rémy Portelas
CoG3
2025 Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?
abstract
Large language models (LLMs) are increasingly used in the creation of online content, creating feedback loops as subsequent generations of models will be trained on this synthetic data.Such loops were shown to lead to distribution shifts -models misrepresenting the true underlying distributions of human data (also called model collapse).However, how human data properties affect such shifts remains poorly understood.In this paper, we provide the first empirical examination of the effect of such properties on the outcome of recursive training.We first confirm that using different human datasets leads to distribution shifts of different magnitudes.Through exhaustive manipulation of dataset properties combined with regression analyses, we then identify a set of properties associated with distribution shift magnitudes.Lexical diversity is found to amplify these shifts, while semantic diversity and data quality mitigate them.Furthermore, we find that these influences are highly modular: data scrapped from a given internet domain has little influence on the content generated for another domain.Finally, experiments on political bias reveal that human data properties affect whether the initial bias will be amplified or reduced.Overall, our results portray a novel view, where different parts of internet may undergo different types of distribution shift.
Grgur Kovac, Jérémy Perez, Rémy Portelas, Peter Ford Dominey, Pierre-Yves Oudeyer
EMNLP3
2025 Efficient Active Imitation Learning with Random Network Distillation
abstract
Developing agents for complex and underspecified tasks, where no clear objective exists, remains challenging but offers many opportunities. This is especially true in video games, where simulated players (bots) need to play realistically, and there is no clear reward to evaluate them. While imitation learning has shown promise in such domains, these methods often fail when agents encounter out-of-distribution scenarios during deployment. Expanding the training dataset is a common solution, but it becomes impractical or costly when relying on human demonstrations. This article addresses active imitation learning, aiming to trigger expert intervention only when necessary, reducing the need for constant expert input along training. We introduce Random Network Distillation DAgger (RND-DAgger), a new active imitation learning method that limits expert querying by using a learned state-based out-of-distribution measure to trigger interventions. This approach avoids frequent expert-agent action comparisons, thus making the expert intervene only when it is useful. We evaluate RND-DAgger against traditional imitation learning and other active approaches in 3D video games (racing and third-person navigation) and in a robotic locomotion task and show that RND-DAgger surpasses previous methods by reducing expert queries. https://sites.google.com/view/rnd-dagger
Emilien Biré, Anthony Kobanda, Ludovic Denoyer, Rémy Portelas
ICLR4
2025 Navigation With QPHIL: Quantizing Planner for Hierarchical Implicit Q-Learning
abstract
Offline Reinforcement Learning (RL) has emerged as a powerful alternative to imitation learning for behavior modeling in various domains, particularly in complex long range navigation tasks. An existing challenge with Offline RL is the signal-to-noise ratio, i.e. how to mitigate incorrect policy updates due to errors in value estimates. Towards this, multiple works have demonstrated the advantage of hierarchical offline RL methods, which decouples high-level path planning from low-level path following. In this work, we present a novel hierarchical transformer-based approach leveraging a learned quantizer of the state space to tackle long horizon navigation tasks. This quantization enables the training of a simpler zone-conditioned low-level policy and simplifies planning, which is reduced to discrete autoregressive prediction. Among other benefits, zone-level reasoning in planning enables explicit trajectory stitching rather than implicit stitching based on noisy value function estimates. By combining this transformer-based planner with recent advancements in offline RL, our proposed approach achieves state-of-the-art results in complex long-distance navigation environments.
Alexi Canesse, Mathieu Petitbois, Ludovic Denoyer, Sylvain Lamprier, Rémy Portelas
IJCNN5
2024 Stick to your Role! Stability of Personal Values Expressed in Large Language Models
Grgur Kovac, Rémy Portelas, Masataka Sawayama, Peter Ford Dominey, Pierre-Yves Oudeyer
CogSci2
2022 Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning
abstract
Intrinsically motivated spontaneous exploration is a key enabler of autonomous developmental learning in human children. It enables the discovery of skill repertoires through autotelic learning, i.e. the self-generation, self-selection, self-ordering and self-experimentation of learning goals. We present an algorithmic approach called Intrinsically Motivated Goal Exploration Processes (IMGEP) to enable similar properties of autonomous learning in machines. The IMGEP architecture relies on several principles: 1) self-generation of goals, generalized as parameterized fitness functions; 2) selection of goals based on intrinsic rewards; 3) exploration with incremental goal-parameterized policy search and exploitation with a batch learning algorithm; 4) systematic reuse of information acquired when targeting a goal for improving towards other goals. We present a particularly efficient form of IMGEP, called AMB, that uses a population-based policy and an object-centered spatio-temporal modularity. We provide several implementations of this architecture and demonstrate their ability to automatically generate a learning curriculum within several experimental setups. One of these experiments includes a real humanoid robot exploring multiple spaces of goals with several hundred continuous dimensions and with distractors. While no particular target goal is provided to these autotelic agents, this curriculum allows the discovery of diverse skills that act as stepping stones for learning more complex skills, e.g. nested tool use.
Sébastien Forestier, Rémy Portelas, Yoan Mollard, Pierre-Yves Oudeyer
J. Mach. Learn. Res.2
2021 TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL
abstract
Training autonomous agents able to generalize to multiple tasks is a key target of Deep Reinforcement Learning (DRL) research. In parallel to improving DRL algorithms themselves, Automatic Curriculum Learning (ACL) study how teacher algorithms can train DRL agents more efficiently by adapting task selection to their evolving abilities. While multiple standard benchmarks exist to compare DRL agents, there is currently no such thing for ACL algorithms. Thus, comparing existing approaches is difficult, as too many experimental parameters differ from paper to paper. In this work, we identify several key challenges faced by ACL algorithms. Based on these, we present TeachMyAgent (TA), a benchmark of current ACL algorithms leveraging procedural task generation. It includes 1) challenge-specific unit-tests using variants of a procedural Box2D bipedal walker environment, and 2) a new procedural Parkour environment combining most ACL challenges, making it ideal for global performance assessment. We then use TeachMyAgent to conduct a comparative study of representative existing approaches, showcasing the competitiveness of some ACL algorithms that do not use expert knowledge. We also show that the Parkour environment remains an open problem. We open-source our environments, all studied ACL algorithms (collected from open-source code or re-implemented), and DRL students in a Python package available at https://github.com/flowersteam/TeachMyAgent.
Clément Romac, Rémy Portelas, Katja Hofmann, Pierre-Yves Oudeyer
ICML2
2020 Automatic Curriculum Learning For Deep RL: A Short Survey
abstract
Automatic Curriculum Learning (ACL) has become a cornerstone of recent successes in Deep Reinforcement Learning (DRL). These methods shape the learning trajectories of agents by challenging them with tasks adapted to their capacities. In recent years, they have been used to improve sample efficiency and asymptotic performance, to organize exploration, to encourage generalization or to solve sparse reward problems, among others. To do so, ACL mechanisms can act on many aspects of learning problems. They can optimize domain randomization for Sim2Real transfer, organize task presentations in multi-task robotic settings, order sequences of opponents in multi-agent scenarios, etc. The ambition of this work is dual: 1) to present a compact and accessible introduction to the Automatic Curriculum Learning literature and 2) to draw a bigger picture of the current state of the art in ACL to encourage the cross-breeding of existing concepts and the emergence of new ideas.
Rémy Portelas, Cédric Colas, Lilian Weng, Katja Hofmann, Pierre-Yves Oudeyer
IJCAI1