VLDB 2026 Research / reviewers in the wild / expert
Rémy Portelas
dblp:251/3113
· DBLP profile ↗
8ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 56% Motion planning and robot control · 12% Language models and text generation · 10% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
automatic curriculum learning |
1.5 | 3 | 2022 | Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022 TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL · ICML 2021 Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020 |
Robotics › Motion planning and robot control
robot learning |
1.4 | 2 | 2025 | Efficient Active Imitation Learning with Random Network Distillation · ICLR 2025 Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning › imitation learning › interactive imitation learning
active imitation learning |
0.9 | 1 | 2025 | Efficient Active Imitation Learning with Random Network Distillation · ICLR 2025 |
Machine learning › Trustworthy machine learning › fairness › algorithmic bias
bias amplification |
0.9 | 1 | 2025 | Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? · EMNLP 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.9 | 1 | 2025 | Efficient Active Imitation Learning with Random Network Distillation · ICLR 2025 |
Machine learning › Generative modeling
model collapse |
0.9 | 1 | 2025 | Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? · EMNLP 2025 |
Natural language and speech › Language models and text generation
synthetic data |
0.9 | 1 | 2025 | Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? · EMNLP 2025 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.6 | 2 | 2021 | Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020 TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL · ICML 2021 |
Machine learning › Learning paradigms
curriculum learning |
0.6 | 1 | 2022 | Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning
exploration |
0.6 | 1 | 2022 | Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning › exploration › directed exploration
goal-directed exploration |
0.6 | 1 | 2022 | Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.6 | 1 | 2022 | Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning
sample efficiency |
0.4 | 1 | 2020 | Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020 |
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
domain randomization |
0.1 | 1 | 2020 | Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.1 | 1 | 2020 | Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020 |
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer |
0.1 | 1 | 2020 | Automatic Curriculum Learning For Deep RL: A Short Survey · IJCAI 2020 |
Methods — techniques the papers use, named apart from their topics
regression analysis · 0.9random network distillation · 0.9DAgger · 0.9population-based policy · 0.6policy search · 0.6procedural task generation · 0.5domain randomization · 0.4curriculum learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Continual Offline Reinforcement Learning Benchmark for Navigation TasksabstractAutonomous agents operating in domains such as robotics or video game simulations must adapt to changing tasks without forgetting about the previous ones. This process called Continual Reinforcement Learning poses non-trivial difficulties, from preventing catastrophic forgetting to ensuring the scalability of the approaches considered. Building on recent advances, we introduce a benchmark providing a suite of video-game navigation scenarios, thus filling a gap in the literature and capturing key challenges: catastrophic forgetting, task adaptation, and memory efficiency. We define a set of various tasks and datasets, evaluation protocols, and metrics to assess the performance of algorithms, including state-of-the-art baselines. Our benchmark is designed not only to foster reproducible research and to accelerate progress in continual reinforcement learning for gaming, but also to provide a reproducible framework for production pipelines helping practitioners to identify and to apply effective approaches. https://sites.google.com/view/continual-nav-bench Anthony Kobanda, Odalric-Ambrym Maillard, Rémy Portelas |
CoG | 3 |
| 2025 | Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?abstractLarge language models (LLMs) are increasingly used in the creation of online content, creating feedback loops as subsequent generations of models will be trained on this synthetic data.Such loops were shown to lead to distribution shifts -models misrepresenting the true underlying distributions of human data (also called model collapse).However, how human data properties affect such shifts remains poorly understood.In this paper, we provide the first empirical examination of the effect of such properties on the outcome of recursive training.We first confirm that using different human datasets leads to distribution shifts of different magnitudes.Through exhaustive manipulation of dataset properties combined with regression analyses, we then identify a set of properties associated with distribution shift magnitudes.Lexical diversity is found to amplify these shifts, while semantic diversity and data quality mitigate them.Furthermore, we find that these influences are highly modular: data scrapped from a given internet domain has little influence on the content generated for another domain.Finally, experiments on political bias reveal that human data properties affect whether the initial bias will be amplified or reduced.Overall, our results portray a novel view, where different parts of internet may undergo different types of distribution shift. Grgur Kovac, Jérémy Perez, Rémy Portelas, Peter Ford Dominey, Pierre-Yves Oudeyer |
EMNLP | 3 |
| 2025 | Efficient Active Imitation Learning with Random Network DistillationabstractDeveloping agents for complex and underspecified tasks, where no clear objective exists, remains challenging but offers many opportunities. This is especially true in video games, where simulated players (bots) need to play realistically, and there is no clear reward to evaluate them. While imitation learning has shown promise in such domains, these methods often fail when agents encounter out-of-distribution scenarios during deployment. Expanding the training dataset is a common solution, but it becomes impractical or costly when relying on human demonstrations. This article addresses active imitation learning, aiming to trigger expert intervention only when necessary, reducing the need for constant expert input along training. We introduce Random Network Distillation DAgger (RND-DAgger), a new active imitation learning method that limits expert querying by using a learned state-based out-of-distribution measure to trigger interventions. This approach avoids frequent expert-agent action comparisons, thus making the expert intervene only when it is useful. We evaluate RND-DAgger against traditional imitation learning and other active approaches in 3D video games (racing and third-person navigation) and in a robotic locomotion task and show that RND-DAgger surpasses previous methods by reducing expert queries.
https://sites.google.com/view/rnd-dagger Emilien Biré, Anthony Kobanda, Ludovic Denoyer, Rémy Portelas |
ICLR | 4 |
| 2025 | Navigation With QPHIL: Quantizing Planner for Hierarchical Implicit Q-LearningabstractOffline Reinforcement Learning (RL) has emerged as a powerful alternative to imitation learning for behavior modeling in various domains, particularly in complex long range navigation tasks. An existing challenge with Offline RL is the signal-to-noise ratio, i.e. how to mitigate incorrect policy updates due to errors in value estimates. Towards this, multiple works have demonstrated the advantage of hierarchical offline RL methods, which decouples high-level path planning from low-level path following. In this work, we present a novel hierarchical transformer-based approach leveraging a learned quantizer of the state space to tackle long horizon navigation tasks. This quantization enables the training of a simpler zone-conditioned low-level policy and simplifies planning, which is reduced to discrete autoregressive prediction. Among other benefits, zone-level reasoning in planning enables explicit trajectory stitching rather than implicit stitching based on noisy value function estimates. By combining this transformer-based planner with recent advancements in offline RL, our proposed approach achieves state-of-the-art results in complex long-distance navigation environments. Alexi Canesse, Mathieu Petitbois, Ludovic Denoyer, Sylvain Lamprier, Rémy Portelas |
IJCNN | 5 |
| 2024 | Stick to your Role! Stability of Personal Values Expressed in Large Language Models
Grgur Kovac, Rémy Portelas, Masataka Sawayama, Peter Ford Dominey, Pierre-Yves Oudeyer |
CogSci | 2 |
| 2022 | Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum LearningabstractIntrinsically motivated spontaneous exploration is a key enabler of autonomous developmental learning in human children. It enables the discovery of skill repertoires through autotelic learning, i.e. the self-generation, self-selection, self-ordering and self-experimentation of learning goals. We present an algorithmic approach called Intrinsically Motivated Goal Exploration Processes (IMGEP) to enable similar properties of autonomous learning in machines. The IMGEP architecture relies on several principles: 1) self-generation of goals, generalized as parameterized fitness functions; 2) selection of goals based on intrinsic rewards; 3) exploration with incremental goal-parameterized policy search and exploitation with a batch learning algorithm; 4) systematic reuse of information acquired when targeting a goal for improving towards other goals. We present a particularly efficient form of IMGEP, called AMB, that uses a population-based policy and an object-centered spatio-temporal modularity. We provide several implementations of this architecture and demonstrate their ability to automatically generate a learning curriculum within several experimental setups. One of these experiments includes a real humanoid robot exploring multiple spaces of goals with several hundred continuous dimensions and with distractors. While no particular target goal is provided to these autotelic agents, this curriculum allows the discovery of diverse skills that act as stepping stones for learning more complex skills, e.g. nested tool use. Sébastien Forestier, Rémy Portelas, Yoan Mollard, Pierre-Yves Oudeyer |
J. Mach. Learn. Res. | 2 |
| 2021 | TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RLabstractTraining autonomous agents able to generalize to multiple tasks is a key target of Deep Reinforcement Learning (DRL) research. In parallel to improving DRL algorithms themselves, Automatic Curriculum Learning (ACL) study how teacher algorithms can train DRL agents more efficiently by adapting task selection to their evolving abilities. While multiple standard benchmarks exist to compare DRL agents, there is currently no such thing for ACL algorithms. Thus, comparing existing approaches is difficult, as too many experimental parameters differ from paper to paper. In this work, we identify several key challenges faced by ACL algorithms. Based on these, we present TeachMyAgent (TA), a benchmark of current ACL algorithms leveraging procedural task generation. It includes 1) challenge-specific unit-tests using variants of a procedural Box2D bipedal walker environment, and 2) a new procedural Parkour environment combining most ACL challenges, making it ideal for global performance assessment. We then use TeachMyAgent to conduct a comparative study of representative existing approaches, showcasing the competitiveness of some ACL algorithms that do not use expert knowledge. We also show that the Parkour environment remains an open problem. We open-source our environments, all studied ACL algorithms (collected from open-source code or re-implemented), and DRL students in a Python package available at https://github.com/flowersteam/TeachMyAgent. Clément Romac, Rémy Portelas, Katja Hofmann, Pierre-Yves Oudeyer |
ICML | 2 |
| 2020 | Automatic Curriculum Learning For Deep RL: A Short SurveyabstractAutomatic Curriculum Learning (ACL) has become a cornerstone of recent successes in Deep Reinforcement Learning (DRL). These methods shape the learning trajectories of agents by challenging them with tasks adapted to their capacities. In recent years, they have been used to improve sample efficiency and asymptotic performance, to organize exploration, to encourage generalization or to solve sparse reward problems, among others. To do so, ACL mechanisms can act on many aspects of learning problems. They can optimize domain randomization for Sim2Real transfer, organize task presentations in multi-task robotic settings, order sequences of opponents in multi-agent scenarios, etc. The ambition of this work is dual: 1) to present a compact and accessible introduction to the Automatic Curriculum Learning literature and 2) to draw a bigger picture of the current state of the art in ACL to encourage the cross-breeding of existing concepts and the emergence of new ideas. Rémy Portelas, Cédric Colas, Lilian Weng, Katja Hofmann, Pierre-Yves Oudeyer |
IJCAI | 1 |