EDBT 2026 Demo / reviewers in the wild / expert
Pierre-Yves Oudeyer
dblp:33/5513
· DBLP profile ↗
67ranked-venue papers
8as first author
24since 2021 · last 2025
0000-0002-1277-130XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 6 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 13 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 since 2021Systems, architecture and hardware · 7Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?abstractLarge language models (LLMs) are increasingly used in the creation of online content, creating feedback loops as subsequent generations of models will be trained on this synthetic data.Such loops were shown to lead to distribution shifts -models misrepresenting the true underlying distributions of human data (also called model collapse).However, how human data properties affect such shifts remains poorly understood.In this paper, we provide the first empirical examination of the effect of such properties on the outcome of recursive training.We first confirm that using different human datasets leads to distribution shifts of different magnitudes.Through exhaustive manipulation of dataset properties combined with regression analyses, we then identify a set of properties associated with distribution shift magnitudes.Lexical diversity is found to amplify these shifts, while semantic diversity and data quality mitigate them.Furthermore, we find that these influences are highly modular: data scrapped from a given internet domain has little influence on the content generated for another domain.Finally, experiments on political bias reveal that human data properties affect whether the initial bias will be amplified or reduced.Overall, our results portray a novel view, where different parts of internet may undergo different types of distribution shift. Grgur Kovac, Jérémy Perez, Rémy Portelas, Peter Ford Dominey, Pierre-Yves Oudeyer |
EMNLP | 5 |
| 2025 | When LLMs Play the Telephone Game: Cultural Attractors as Conceptual Tools to Evaluate LLMs in Multi-turn SettingsabstractAs large language models (LLMs) start interacting with each other and generating an increasing amount of text online, it becomes crucial to better understand how information is transformed as it passes from one LLM to the next. While significant research has examined individual LLM behaviors, existing studies have largely overlooked the collective behaviors and information distortions arising from iterated LLM interactions. Small biases, negligible at the single output level, risk being amplified in iterated interactions, potentially leading the content to evolve towards attractor states. In a series of _telephone game experiments_, we apply a transmission chain design borrowed from the human cultural evolution literature: LLM agents iteratively receive, produce, and transmit texts from the previous to the next agent in the chain. By tracking the evolution of text _toxicity_, _positivity_, _difficulty_, and _length_ across transmission chains, we uncover the existence of biases and attractors, and study their dependence on the initial text, the instructions, language model, and model size. For instance, we find that more open-ended instructions lead to stronger attraction effects compared to more constrained tasks. We also find that different text properties display different sensitivity to attraction effects, with _toxicity_ leading to stronger attractors than _length_. These findings highlight the importance of accounting for multi-step transmission dynamics and represent a first step towards a more comprehensive understanding of LLM cultural dynamics. Jérémy Perez, Grgur Kovac, Corentin Léger, Cédric Colas, Gaia Molinaro, Maxime Derex, Pierre-Yves Oudeyer, Clément Moulin-Frier |
ICLR | 7 |
| 2025 | PhyloLM: Inferring the Phylogeny of Large Language Models and Predicting their Performances in BenchmarksabstractThis paper introduces PhyloLM, a method adapting phylogenetic algorithms to Large Language Models (LLMs) to explore whether and how they relate to each other and to predict their performance characteristics. Our method calculates a phylogenetic distance metric based on the similarity of LLMs' output. The resulting metric is then used to construct dendrograms, which satisfactorily capture known relationships across a set of 111 open-source and 45 closed models. Furthermore, our phylogenetic distance predicts performance in standard benchmarks, thus demonstrating its functional validity and paving the way for a time and cost-effective estimation of LLM capabilities. To sum up, by translating population genetic concepts to machine learning, we propose and validate a tool to evaluate LLM development, relationships and capabilities, even in the absence of transparent training information. Nicolas Yax, Pierre-Yves Oudeyer, Stefano Palminteri |
ICLR | 2 |
| 2025 | MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spacesabstractOpen-ended learning agents must efficiently prioritize goals in vast possibility spaces, focusing on those that maximize learning progress (LP). When such autotelic exploration is achieved by LLM agents trained with online RL in high-dimensional and evolving goal spaces, a key challenge for LP prediction is modeling one’s own competence, a form of metacognitive monitoring. Traditional approaches either require extensive sampling or rely on brittle expert-defined goal groupings. We introduce MAGELLAN, a metacognitive framework that lets LLM agents learn to predict their competence and learning progress online. By capturing semantic relationships between goals, MAGELLAN enables sample-efficient LP estimation and dynamic adaptation to evolving goal spaces through generalization. In an interactive learning environment, we show that MAGELLAN improves LP prediction efficiency and goal prioritization, being the only method allowing the agent to fully master a large and evolving goal space. These results demonstrate how augmenting LLM agents with a metacognitive ability for LP predictions can effectively scale curriculum learning to open-ended goal spaces. Loris Gaven, Thomas Carta, Clément Romac, Cédric Colas, Sylvain Lamprier, Olivier Sigaud, Pierre-Yves Oudeyer |
ICML | 7 |
| 2025 | Self-Improving Language Models for Evolutionary Program Synthesis: A Case Study on ARC-AGIabstractMany program synthesis tasks prove too challenging for even state-of-the-art language models to solve in single attempts. Search-based evolutionary methods offer a promising alternative by exploring solution spaces iteratively, but their effectiveness remain limited by the fixed capabilities of the underlying generative model. We propose SOAR, a method that learns program synthesis by integrating language models into a self-improving evolutionary loop. SOAR alternates between (1) an evolutionary search that uses an LLM to sample and refine candidate solutions, and (2) a hindsight learning phase that converts search attempts into valid problem-solution pairs used to fine-tune the LLM’s sampling and refinement capabilities—enabling increasingly effective search in subsequent iterations. On the challenging ARC-AGI benchmark, SOAR achieves significant performance gains across model scales and iterations, leveraging positive transfer between the sampling and refinement finetuning tasks. These improvements carry over to test-time adaptation, enabling SOAR to solve 52% of the public test set. Julien Pourcel, Cédric Colas, Pierre-Yves Oudeyer |
ICML | 3 |
| 2025 | Flow-Lenia: Emergent Evolutionary Dynamics in Mass Conservative Continuous Cellular AutomataabstractCentral to the Artificial Life endeavor is the creation of artificial systems that spontaneously generate properties found in the living world, such as autopoiesis, self-replication, evolution, and open-endedness. Though numerous models and paradigms have been proposed, cellular automata (CA) have taken a very important place in the field, notably because they enable the study of phenomena like self-reproduction and autopoiesis. Continuous CA like Lenia have been shown to produce lifelike patterns reminiscent, from both aesthetic and ontological points of view, of biological organisms we call "creatures." We propose Flow-Lenia, a mass conservative extension of Lenia. We present experiments demonstrating its effectiveness in generating spatially localized patterns with complex behaviors and show that the update rule parameters can be optimized to generate complex creatures showing behaviors of interest. Furthermore, we show that Flow-Lenia allows us to embed the parameters of the model, defining the properties of the emerging patterns, within its own dynamics, thus allowing for multispecies simulation. Using the evolutionary activity framework and other metrics, we shed light on the emergent evolutionary dynamics taking place in this system. Erwan Plantec, Gautier Hamon, Mayalen Etcheverry, Bert Wang-Chak Chan, Pierre-Yves Oudeyer, Clément Moulin-Frier |
Artif. Life | 5 |
| 2024 | Stick to your Role! Stability of Personal Values Expressed in Large Language Models
Grgur Kovac, Rémy Portelas, Masataka Sawayama, Peter Ford Dominey, Pierre-Yves Oudeyer |
CogSci | 5 |
| 2024 | Latent Learning Progress Drives Autonomous Goal Selection in Human Reinforcement LearningabstractHumans are autotelic agents who learn by setting and pursuing their own goals. However, the precise mechanisms guiding human goal selection remain unclear. Learning progress, typically measured as the observed change in performance, can provide a valuable signal for goal selection in both humans and artificial agents. We hypothesize that human choices of goals may also be driven by _latent learning progress_, which humans can estimate through knowledge of their actions and the environment – even without experiencing immediate changes in performance. To test this hypothesis, we designed a hierarchical reinforcement learning task in which human participants (N = 175) repeatedly chose their own goals and learned goal-conditioned policies. Our behavioral and computational modeling results confirm the influence of latent learning progress on goal selection and uncover inter-individual differences, partially mediated by recognition of the task's hierarchical structure. By investigating the role of latent learning progress in human goal selection, we pave the way for more effective and personalized learning experiences as well as the advancement of more human-like autotelic machines. Gaia Molinaro, Cédric Colas, Pierre-Yves Oudeyer, Anne Gabrielle Eva Collins |
NeurIPS | 3 |
| 2024 | ACES: Generating a Diversity of Challenging Programming Puzzles with Autotelic Generative ModelsabstractThe ability to invent novel and interesting problems is a remarkable feature of human intelligence that drives innovation, art, and science. We propose a method that aims to automate this process by harnessing the power of state-of-the-art generative models to produce a diversity of challenging yet solvable problems, here in the context of Python programming puzzles. Inspired by the intrinsically motivated literature, Autotelic CodE Search (ACES) jointly optimizes for the diversity and difficulty of generated problems. We represent problems in a space of LLM-generated semantic descriptors describing the programming skills required to solve them (e.g. string manipulation, dynamic programming, etc.) and measure their difficulty empirically as a linearly decreasing function of the success rate of \textit{Llama-3-70B}, a state-of-the-art LLM problem solver. ACES iteratively prompts a large language model to generate difficult problems achieving a diversity of target semantic descriptors (goal-directed exploration) using previously generated problems as in-context examples. ACES generates problems that are more diverse and more challenging than problems produced by baseline methods and three times more challenging than problems found in existing Python programming benchmarks on average across 11 state-of-the-art code LLMs. Julien Pourcel, Cédric Colas, Gaia Molinaro, Pierre-Yves Oudeyer, Laetitia Teodorescu |
NeurIPS | 4 |
| 2023 | Interactive environments for training children's curiosity through the practice of metacognitive skills : a pilot studyabstractCuriosity-driven learning has shown significant positive effects on students’ learning experiences and outcomes. But despite this importance, reports show that children lack this skill, especially in formal educational settings. Rania Abdelghani, Edith Law, Chloé Desvaux, Pierre-Yves Oudeyer, Hélène Sauzéon |
IDC | 4 |
| 2023 | Grounding Large Language Models in Interactive Environments with Online Reinforcement LearningabstractRecent works successfully leveraged Large Language Models’ (LLM) abilities to capture abstract knowledge about world’s physics to solve decision-making problems. Yet, the alignment between LLMs’ knowledge and the environment can be wrong and limit functional competence due to lack of grounding. In this paper, we study an approach (named GLAM) to achieve this alignment through functional grounding: we consider an agent using an LLM as a policy that is progressively updated as the agent interacts with the environment, leveraging online Reinforcement Learning to improve its performance to solve goals. Using an interactive textual environment designed to study higher-level forms of functional grounding, and a set of spatial and navigation tasks, we study several scientific questions: 1) Can LLMs boost sample efficiency for online learning of various RL tasks? 2) How can it boost different forms of generalization? 3) What is the impact of online learning? We study these questions by functionally grounding several variants (size, architecture) of FLAN-T5. Thomas Carta, Clément Romac, Sylvain Lamprier, Olivier Sigaud, Pierre-Yves Oudeyer |
ICML | 6 |
| 2022 | Learning to Guide and to be Guided in the Architect-Builder Problem
Paul Barde, Tristan Karch, Derek Nowrouzezahrai, Clément Moulin-Frier, Christopher Joseph Pal, Pierre-Yves Oudeyer |
ICLR | 6 |
| 2022 | Language-biased image classification: evaluation based on semantic representations
Yoann Lemesle, Masataka Sawayama, Guillermo Valle Pérez, Maxime Adolphe, Hélène Sauzéon, Pierre-Yves Oudeyer |
ICLR | 6 |
| 2022 | Asking for Knowledge (AFK): Training RL Agents to Query External Knowledge Using LanguageabstractTo solve difficult tasks, humans ask questions to acquire knowledge from external sources. In contrast, classical reinforcement learning agents lack such an ability and often resort to exploratory behavior. This is exacerbated as few present-day environments support querying for knowledge. In order to study how agents can be taught to query external knowledge via language, we first introduce two new environments: the grid-world-based Q-BabyAI and the text-based Q-TextWorld. In addition to physical interactions, an agent can query an external knowledge source specialized for these environments to gather information. Second, we propose the ‘Asking for Knowledge’ (AFK) agent, which learns to generate language commands to query for meaningful knowledge that helps solve the tasks. AFK leverages a non-parametric memory, a pointer mechanism and an episodic exploration bonus to tackle (1) irrelevant information, (2) a large query language space, (3) delayed reward for making meaningful queries. Extensive experiments demonstrate that the AFK agent outperforms recent baselines on the challenging Q-BabyAI and Q-TextWorld environments. Iou-Jen Liu, Xingdi Yuan, Marc-Alexandre Côté, Pierre-Yves Oudeyer, Alexander G. Schwing |
ICML | 4 |
| 2022 | EAGER: Asking and Answering Questions for Automatic Reward Shaping in Language-guided RLabstractReinforcement learning (RL) in long horizon and sparse reward tasks is notoriously difficult and requires a lot of training steps. A standard solution to speed up the process is to leverage additional reward signals, shaping it to better guide the learning process.In the context of language-conditioned RL, the abstraction and generalisation properties of the language input provide opportunities for more efficient ways of shaping the reward.In this paper, we leverage this idea and propose an automated reward shaping method where the agent extracts auxiliary objectives from the general language goal. These auxiliary objectives use a question generation (QG) and a question answering (QA) system: they consist of questions leading the agent to try to reconstruct partial information about the global goal using its own trajectory.When it succeeds, it receives an intrinsic reward proportional to its confidence in its answer. This incentivizes the agent to generate trajectories which unambiguously explain various aspects of the general language goal.Our experimental study using various BabyAI environments shows that this approach, which does not require engineer intervention to design the auxiliary objectives, improves sample efficiency by effectively directing the exploration. Thomas Carta, Pierre-Yves Oudeyer, Olivier Sigaud, Sylvain Lamprier |
NeurIPS | 2 |
| 2022 | Conversational agents for fostering curiosity-driven learning in children
Rania Abdelghani, Pierre-Yves Oudeyer, Edith Law, Catherine de Vulpillières, Hélène Sauzéon |
Int. J. Hum. Comput. Stud. | 2 |
| 2022 | Autotelic Agents with Intrinsically Motivated Goal-Conditioned Reinforcement Learning: A Short SurveyabstractBuilding autonomous machines that can explore open-ended environments, discover possible interactions and build repertoires of skills is a general objective of artificial intelligence. Developmental approaches argue that this can only be achieved by autotelic agents: intrinsically motivated learning agents that can learn to represent, generate, select and solve their own problems. In recent years, the convergence of developmental approaches with deep reinforcement learning (RL) methods has been leading to the emergence of a new field: developmental reinforcement learning. Developmental RL is concerned with the use of deep RL algorithms to tackle a developmental problem— the intrinsically motivated acquisition of open-ended repertoires of skills. The self-generation of goals requires the learning of compact goal encodings as well as their associated goal-achievement functions. This raises new challenges compared to standard RL algorithms originally designed to tackle pre-defined sets of goals using external reward signals. The present paper introduces developmental RL and proposes a computational framework based on goal-conditioned RL to tackle the intrinsically motivated skills acquisition problem. It proceeds to present a typology of the various goal representations used in the literature, before reviewing existing methods to learn to represent and prioritize goals in autonomous systems. We finally close the paper by discussing some open challenges in the quest of intrinsically motivated skills acquisition. Cédric Colas, Tristan Karch, Olivier Sigaud, Pierre-Yves Oudeyer |
J. Artif. Intell. Res. | 4 |
| 2022 | Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum LearningabstractIntrinsically motivated spontaneous exploration is a key enabler of autonomous developmental learning in human children. It enables the discovery of skill repertoires through autotelic learning, i.e. the self-generation, self-selection, self-ordering and self-experimentation of learning goals. We present an algorithmic approach called Intrinsically Motivated Goal Exploration Processes (IMGEP) to enable similar properties of autonomous learning in machines. The IMGEP architecture relies on several principles: 1) self-generation of goals, generalized as parameterized fitness functions; 2) selection of goals based on intrinsic rewards; 3) exploration with incremental goal-parameterized policy search and exploitation with a batch learning algorithm; 4) systematic reuse of information acquired when targeting a goal for improving towards other goals. We present a particularly efficient form of IMGEP, called AMB, that uses a population-based policy and an object-centered spatio-temporal modularity. We provide several implementations of this architecture and demonstrate their ability to automatically generate a learning curriculum within several experimental setups. One of these experiments includes a real humanoid robot exploring multiple spaces of goals with several hundred continuous dimensions and with distractors. While no particular target goal is provided to these autotelic agents, this curriculum allows the discovery of diverse skills that act as stepping stones for learning more complex skills, e.g. nested tool use. Sébastien Forestier, Rémy Portelas, Yoan Mollard, Pierre-Yves Oudeyer |
J. Mach. Learn. Res. | 4 |
| 2021 | Intrinsic Rewards in Human Curiosity-Driven Exploration: An Empirical Study
Alexandr Ten, Jacqueline Gottlieb, Pierre-Yves Oudeyer |
CogSci | 3 |
| 2021 | Grounding Language to Autonomously-Acquired Skills via Goal Generation
Ahmed Akakzia, Cédric Colas, Pierre-Yves Oudeyer, Mohamed Chetouani, Olivier Sigaud |
ICLR | 3 |
| 2021 | TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RLabstractTraining autonomous agents able to generalize to multiple tasks is a key target of Deep Reinforcement Learning (DRL) research. In parallel to improving DRL algorithms themselves, Automatic Curriculum Learning (ACL) study how teacher algorithms can train DRL agents more efficiently by adapting task selection to their evolving abilities. While multiple standard benchmarks exist to compare DRL agents, there is currently no such thing for ACL algorithms. Thus, comparing existing approaches is difficult, as too many experimental parameters differ from paper to paper. In this work, we identify several key challenges faced by ACL algorithms. Based on these, we present TeachMyAgent (TA), a benchmark of current ACL algorithms leveraging procedural task generation. It includes 1) challenge-specific unit-tests using variants of a procedural Box2D bipedal walker environment, and 2) a new procedural Parkour environment combining most ACL challenges, making it ideal for global performance assessment. We then use TeachMyAgent to conduct a comparative study of representative existing approaches, showcasing the competitiveness of some ACL algorithms that do not use expert knowledge. We also show that the Parkour environment remains an open problem. We open-source our environments, all studied ACL algorithms (collected from open-source code or re-implemented), and DRL students in a Python package available at https://github.com/flowersteam/TeachMyAgent. Clément Romac, Rémy Portelas, Katja Hofmann, Pierre-Yves Oudeyer |
ICML | 4 |
| 2021 | Grounding Spatio-Temporal Language with TransformersabstractLanguage is an interface to the outside world. In order for embodied agents to use it, language must be grounded in other, sensorimotor modalities. While there is an extended literature studying how machines can learn grounded language, the topic of how to learn spatio-temporal linguistic concepts is still largely uncharted. To make progress in this direction, we here introduce a novel spatio-temporal language grounding task where the goal is to learn the meaning of spatio-temporal descriptions of behavioral traces of an embodied agent. This is achieved by training a truth function that predicts if a description matches a given history of observations. The descriptions involve time-extended predicates in past and present tense as well as spatio-temporal references to objects in the scene. To study the role of architectural biases in this task, we train several models including multimodal Transformer architectures; the latter implement different attention computations between words and objects across space and time. We test models on two classes of generalization: 1) generalization to new sentences, 2) generalization to grammar primitives. We observe that maintaining object identity in the attention computation of our Transformers is instrumental to achieving good performance on generalization overall, and that summarizing object traces in a single token has little influence on performance. We then discuss how this opens new perspectives for language-guided autonomous embodied agents. Tristan Karch, Laetitia Teodorescu, Katja Hofmann, Clément Moulin-Frier, Pierre-Yves Oudeyer |
NeurIPS | 5 |
| 2021 | EpidemiOptim: A Toolbox for the Optimization of Control Policies in Epidemiological ModelsabstractModeling the dynamics of epidemics helps to propose control strategies based on pharmaceuticaland non-pharmaceutical interventions (contact limitation, lockdown, vaccination,etc). Hand-designing such strategies is not trivial because of the number of possibleinterventions and the difficulty to predict long-term effects. This task can be cast as an optimization problem where state-of-the-art machine learning methods such as deep reinforcement learning might bring significant value. However, the specificity of each domain|epidemic modeling or solving optimization problems|requires strong collaborationsbetween researchers from different fields of expertise. This is why we introduce EpidemiOptim, a Python toolbox that facilitates collaborations between researchers inepidemiology and optimization. EpidemiOptim turns epidemiological models and cost functions into optimization problems via a standard interface commonly used by optimization practitioners (OpenAI Gym). Reinforcement learning algorithms based on QLearning with deep neural networks (DQN) and evolutionary algorithms (NSGA-II) are already implemented. We illustrate the use of EpidemiOptim to find optimal policies fordynamical on-o lockdown control under the optimization of the death toll and economic recess using a Susceptible-Exposed-Infectious-Removed (SEIR) model for COVID-19. Using EpidemiOptim and its interactive visualization platform in Jupyter notebooks, epidemiologists, optimization practitioners and others (e.g. economists) can easily compare epidemiological models, costs functions and optimization algorithms to address important choicesto be made by health decision-makers. Trained models can be explored by experts and non-experts via a web interface. This article is part of the special track on AI and COVID-19. Cédric Colas, Boris P. Hejblum, Sébastien Rouillon, Rodolphe Thiébaut, Pierre-Yves Oudeyer, Clément Moulin-Frier, Mélanie Prague |
J. Artif. Intell. Res. | 5 |
| 2021 | Transflower: probabilistic autoregressive dance generation with multimodal attentionabstractDance requires skillful composition of complex movements that follow rhythmic, tonal and timbral features of music. Formally, generating dance conditioned on a piece of music can be expressed as a problem of modelling a high-dimensional continuous motion signal, conditioned on an audio signal. In this work we make two contributions to tackle this problem. First, we present a novel probabilistic autoregressive architecture that models the distribution over future poses with a normalizing flow conditioned on previous poses as well as music context, using a multimodal transformer encoder. Second, we introduce the currently largest 3D dance-motion dataset, obtained with a variety of motion-capture technologies, and including both professional and casual dancers. Using this dataset, we compare our new model against two baselines, via objective metrics and a user study, and show that both the ability to model a probability distribution, as well as being able to attend over a large motion and music context are necessary to produce interesting, diverse, and realistic dance that matches the music. Guillermo Valle Pérez, Gustav Eje Henter, Jonas Beskow, Andre Holzapfel, Pierre-Yves Oudeyer, Simon Alexanderson |
ACM Trans. Graph. | 5 |
| 2020 | Pedagogical Agents for Fostering Question-Asking Skills in ChildrenabstractQuestion asking is an important tool for constructing academic knowledge, and a self-reinforcing driver of curiosity. However, research has found that question asking is infrequent in the classroom and children's questions are often superficial, lacking deep reasoning. In this work, we developed a pedagogical agent that encourages children to ask divergent-thinking questions, a more complex form of questions that is associated with curiosity. We conducted a study with 95 fifth grade students, who interacted with an agent that encourages either convergent-thinking or divergent-thinking questions. Results showed that both interventions increased the number of divergent-thinking questions and the fluency of question asking, while they did not significantly alter children's perception of curiosity despite their high intrinsic motivation scores. In addition, children's curiosity trait has a mediating effect on question asking under the divergent-thinking agent, suggesting that question-asking interventions must be personalized to each student based on their tendency to be curious. Mehdi Alaimi, Edith Law, Kevin Daniel Pantasdo, Pierre-Yves Oudeyer, Hélène Sauzéon |
CHI | 4 |
| 2020 | Intrinsically Motivated Discovery of Diverse Patterns in Self-Organizing Systems
Chris Reinke, Mayalen Etcheverry, Pierre-Yves Oudeyer |
ICLR | 3 |
| 2020 | Automatic Curriculum Learning For Deep RL: A Short SurveyabstractAutomatic Curriculum Learning (ACL) has become a cornerstone of recent successes in Deep Reinforcement Learning (DRL). These methods shape the learning trajectories of agents by challenging them with tasks adapted to their capacities. In recent years, they have been used to improve sample efficiency and asymptotic performance, to organize exploration, to encourage generalization or to solve sparse reward problems, among others. To do so, ACL mechanisms can act on many aspects of learning problems. They can optimize domain randomization for Sim2Real transfer, organize task presentations in multi-task robotic settings, order sequences of opponents in multi-agent scenarios, etc. The ambition of this work is dual: 1) to present a compact and accessible introduction to the Automatic Curriculum Learning literature and 2) to draw a bigger picture of the current state of the art in ACL to encourage the cross-breeding of existing concepts and the emergence of new ideas. Rémy Portelas, Cédric Colas, Lilian Weng, Katja Hofmann, Pierre-Yves Oudeyer |
IJCAI | 5 |
| 2020 | User-in-the-loop adaptive intent detection for instructable digital assistantabstractPeople are becoming increasingly comfortable using Digital Assistants (DAs) to interact with services or connected objects. However, for non-programming users, the available possibilities for customizing their DA are limited and do not include the possibility of teaching the assistant new tasks. To make the most of the potential of DAs, users should be able to customize assistants by instructing them through Natural Language (NL). To provide such functionalities, NL interpretation in traditional assistants should be improved: (1) The intent identification system should be able to recognize new forms of known intents, and to acquire new intents as they are expressed by the user. (2) In order to be adaptive to novel intents, the Natural Language Understanding module should be sample efficient, and should not rely on a pretrained model. Rather, the system should continuously collect the training data as it learns new intents from the user. In this work, we propose AidMe (Adaptive Intent Detection in Multi-Domain Environments), a user-in-the-loop adaptive intent detection framework that allows the assistant to adapt to its user by learning his intents as their interaction progresses. AidMe builds its repertoire of intents and collects data to train a model of semantic similarity evaluation that can discriminate between the learned intents and autonomously discover new forms of known intents. AidMe addresses two major issues - intent learning and user adaptation - for instructable digital assistants. We demonstrate the capabilities of AidMe as a standalone system by comparing it with a one-shot learning system and a pretrained NLU module through simulations of interactions with a user. We also show how AidMe can smoothly integrate to an existing instructable digital assistant. Nicolas Lair, Clément Delgrange, David Mugisha, Jean-Michel Dussoux, Pierre-Yves Oudeyer, Peter Ford Dominey |
IUI | 5 |
| 2020 | Language as a Cognitive Tool to Imagine Goals in Curiosity Driven ExplorationabstractDevelopmental machine learning studies how artificial agents can model the way children learn open-ended repertoires of skills. Such agents need to create and represent goals, select which ones to pursue and learn to achieve them. Recent approaches have considered goal spaces that were either fixed and hand-defined or learned using generative models of states. This limited agents to sample goals within the distribution of known effects. We argue that the ability to imagine out-of-distribution goals is key to enable creative discoveries and open-ended learning. Children do so by leveraging the compositionality of language as a tool to imagine descriptions of outcomes they never experienced before, targeting them as goals during play. We introduce IMAGINE, an intrinsically motivated deep reinforcement learning architecture that models this ability. Such imaginative agents, like children, benefit from the guidance of a social peer who provides language descriptions. To take advantage of goal imagination, agents must be able to leverage these descriptions to interpret their imagined out-of-distribution goals. This generalization is made possible by modularity: a decomposition between learned goal-achievement reward function and policy relying on deep sets, gated attention and object-centered representations. We introduce the Playground environment and study how this form of goal imagination improves generalization and exploration over agents lacking this capacity. In addition, we identify the properties of goal imagination that enable these results and study the impacts of modularity and social interactions. Cédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux, Clément Moulin-Frier, Peter Ford Dominey, Pierre-Yves Oudeyer |
NeurIPS | 7 |
| 2020 | Hierarchically Organized Latent Modules for Exploratory Search in Morphogenetic SystemsabstractSelf-organization of complex morphological patterns from local interactions is a fascinating phenomenon in many natural and artificial systems. In the artificial world, typical examples of such morphogenetic systems are cellular automata. Yet, their mechanisms are often very hard to grasp and so far scientific discoveries of novel patterns have primarily been relying on manual tuning and ad hoc exploratory search. The problem of automated diversity-driven discovery in these systems was recently introduced [26, 62], highlighting that two key ingredients are autonomous exploration and unsupervised representation learning to describe “relevant” degrees of variations in the patterns. In this paper, we motivate the need for what we call Meta-diversity search, arguing that there is not a unique ground truth interesting diversity as it strongly depends on the final observer and its motives. Using a continuous game-of-life system for experiments, we provide empirical evidences that relying on monolithic architectures for the behavioral embedding design tends to bias the final discoveries (both for hand-defined and unsupervisedly-learned features) which are unlikely to be aligned with the interest of a final end-user. To address these issues, we introduce a novel dynamic and modular architecture that enables unsupervised learning of a hierarchy of diverse representations. Combined with intrinsically motivated goal exploration algorithms, we show that this system forms a discovery assistant that can efficiently adapt its diversity search towards preferences of a user using only a very small amount of user feedback. Mayalen Etcheverry, Clément Moulin-Frier, Pierre-Yves Oudeyer |
NeurIPS | 3 |
| 2020 | Towards measuring states of epistemic curiosity through electroencephalographic signalsabstractUnderstanding the neurophysiological mechanisms underlying curiosity and therefore being able to identify the curiosity level of a person, would provide useful information for researchers and designers in numerous fields such as neuroscience, psychology, and computer science. A first step to uncovering the neural correlates of curiosity is to collect neurophysiological signals during states of curiosity, in order to develop signal processing and machine learning (ML) tools to recognize the curious states from the non-curious ones. Thus, we ran an experiment in which we used electroencephalography (EEG) to measure the brain activity of participants as they were induced into states of curiosity, using trivia question and answer chains. We used two ML algorithms, i.e. Filter Bank Common Spatial Pattern (FBCSP) coupled with a Linear Discriminant Algorithm (LDA), as well as a Filter Bank Tangent Space Classifier (FBTSC), to classify the curious EEG signals from the non-curious ones. Global results indicate that both algorithms obtained better performances in the 3-to-5s time windows, suggesting an optimal time window length of 4 seconds (63.09% classification accuracy for the FBTSC, 60.93% classification accuracy for the FBCSP+LDA) to go towards curiosity states estimation based on EEG signals. Aurélien Appriou, Jessy Ceha, Smeety Pramij, Dan Dutartre, Edith Law, Pierre-Yves Oudeyer, Fabien Lotte |
SMC | 6 |
| 2019 | Expression of Curiosity in Social Robots: Design, Perception, and Effects on BehaviourabstractCuriosity-the intrinsic desire for new information-can enhance learning, memory, and exploration. Therefore, understanding how to elicit curiosity can inform the design of educational technologies. In this work, we investigate how a social peer robot's verbal expression of curiosity is perceived, whether it can affect the emotional feeling and behavioural expression of curiosity in students, and how it impacts learning. In a between-subjects experiment, 30 participants played the game LinkIt!, a game we designed for teaching rock classification, with a robot verbally expressing: curiosity, curiosity plus rationale, or no curiosity. Results indicate that participants could recognize the robot's curiosity and that curious robots produced both emotional and behavioural curiosity contagion effects in participants. Jessy Ceha, Nalin Chhibber, Joslin Goh, Corina McDonald, Pierre-Yves Oudeyer, Dana Kulic, Edith Law |
CHI | 5 |
| 2019 | CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement LearningabstractIn open-ended environments, autonomous learning agents must set their own goals and build their own curriculum through an intrinsically motivated exploration. They may consider a large diversity of goals, aiming to discover what is controllable in their environments, and what is not. Because some goals might prove easy and some impossible, agents must actively select which goal to practice at any moment, to maximize their overall mastery on the set of learnable goals. This paper proposes CURIOUS , an algorithm that leverages 1) a modular Universal Value Function Approximator with hindsight learning to achieve a diversity of goals of different kinds within a unique policy and 2) an automated curriculum learning mechanism that biases the attention of the agent towards goals maximizing the absolute learning progress. Agents focus sequentially on goals of increasing complexity, and focus back on goals that are being forgotten. Experiments conducted in a new modular-goal robotic environment show the resulting developmental self-organization of a learning curriculum, and demonstrate properties of robustness to distracting goals, forgetting and changes in body properties. Cédric Colas, Pierre-Yves Oudeyer, Olivier Sigaud, Pierre Fournier, Mohamed Chetouani |
ICML | 2 |
| 2019 | Developmental Autonomous Learning: AI, Cognitive Sciences and Educational TechnologyabstractCurrent approaches to AI and machine learning are still fundamentally limited in comparison with autonomous learning capabilities of children. What is remarkable is not that some children become world champions in certain games or specialties: it is rather their autonomy, flexibility and efficiency at learning many everyday skills under strongly limited resources of time, computation and energy. And they do not need the intervention of an engineer for each new task (e.g. they do not need someone to provide a new task specific reward function). XX I will present a research program that has focused on computational modeling of child development and learning mechanisms in the last decade. I will discuss several developmental forces that guide exploration in large real world spaces, starting from the perspective of how algorithmic models can help us understand better how they work in humans, and in return how this opens new approaches to autonomous machine learning. XX In particular, I will discuss models of curiosity-driven autonomous learning, enabling machines to sample and explore their own goals and their own learning strategies, self-organizing a learning curriculum without any external reward or supervision. XX I will show how this has helped scientists understand better aspects of human development such as the emergence of developmental transitions between object manipulation, tool use and speech. I will also show how the use of real robotic platforms for evaluating these models has led to highly efficient unsupervised learning methods, enabling robots to discover and learn multiple skills in high-dimensions in a handful of hours. I will discuss how these techniques are now being integrated with modern deep learning methods. XX Finally, I will show how these models and techniques can be successfully applied in the domain of educational technologies, enabling to personalize sequences of exercises for human learners, while maximizing both learning efficiency and intrinsic motivation. I will illustrate this with a large-scale experiment recently performed in primary schools, enabling children of all levels to improve their skills and motivation in learning aspects of mathematics. Pierre-Yves Oudeyer |
IVA | 1 |
| 2018 | Complexity Reduction in the Negotiation of New Lexical Conventions
William Schueller, Vittorio Loreto, Pierre-Yves Oudeyer |
CogSci | 3 |
| 2018 | Unsupervised Learning of Goal Spaces for Intrinsically Motivated Goal Exploration
Alexandre Péré, Sébastien Forestier, Olivier Sigaud, Pierre-Yves Oudeyer |
ICLR (Poster) | 4 |
| 2018 | GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning AlgorithmsabstractIn continuous action domains, standard deep reinforcement learning algorithms like DDPG suffer from inefficient exploration when facing sparse or deceptive reward problems. Conversely, evolutionary and developmental methods focusing on exploration like Novelty Search, Quality-Diversity or Goal Exploration Processes explore more robustly but are less efficient at fine-tuning policies using gradient-descent. In this paper, we present the GEP-PG approach, taking the best of both worlds by sequentially combining a Goal Exploration Process and two variants of DDPG . We study the learning performance of these components and their combination on a low dimensional deceptive reward problem and on the larger Half-Cheetah benchmark. We show that DDPG fails on the former and that GEP-PG improves over the best DDPG variant in both environments. Cédric Colas, Olivier Sigaud, Pierre-Yves Oudeyer |
ICML | 3 |
| 2017 | A Unified Model of Speech and Tool Use Early Development
Sébastien Forestier, Pierre-Yves Oudeyer |
CogSci | 2 |
| 2016 | Curiosity-Driven Development of Tool Use Precursors: a Computational Model
Sébastien Forestier, Pierre-Yves Oudeyer |
CogSci | 2 |
| 2016 | A Comparison of Automatic Teaching Strategies for Heterogeneous Student Populations
Benjamin Clément, Pierre-Yves Oudeyer, Manuel Lopes 0001 |
EDM | 2 |
| 2016 | Modular active curiosity-driven discovery of tool useabstractThis article studies algorithms used by a learner to explore high-dimensional structured sensorimotor spaces such as in tool use discovery. In particular, we consider goal babbling architectures that were designed to explore and learn solutions to fields of sensorimotor problems, i.e. to acquire inverse models mapping a space of parameterized sensorimotor problems/effects to a corresponding space of parameterized motor primitives. However, so far these architectures have not been used in high-dimensional spaces of effects. Here, we show the limits of existing goal babbling architectures for efficient exploration in such spaces, and introduce a novel exploration architecture called Model Babbling (MB). MB exploits efficiently a modular representation of the space of parameterized problems/effects. We also study an active version of Model Babbling (the MACOB architecture). These architectures are compared in a simulated experimental setup with an arm that can discover and learn how to move objects using two tools with different properties, embedding structured high-dimensional continuous motor and sensory spaces. Sébastien Forestier, Pierre-Yves Oudeyer |
IROS | 2 |
| 2015 | Multi-Armed Bandits for Intelligent Tutoring Systems
Benjamin Clément, Didier Roy, Pierre-Yves Oudeyer, Manuel Lopes 0001 |
EDM | 3 |
| 2014 | Calibration-Free BCI Based ControlabstractRecent works have explored the use of brain signals to directly control virtual and robotic agents in sequential tasks. So far in such brain-computer interfaces (BCI), an explicit calibration phase was required to build a decoder that translates raw electroencephalography (EEG) signals from the brain of each user into meaningful instructions. This paper proposes a method that removes the calibration phase, and allows a user to control an agent to solve a sequential task. The proposed method assumes a distribution of possible tasks, and infers the interpretation of EEG signals and the task by selecting the hypothesis which best explains the history of interaction. We introduce a measure of uncertainty on the task and on the EEG signal interpretation to act as an exploratory bonus for a planning strategy. This speeds up learning by guiding the system to regions that better disambiguate among task hypotheses. We report experiments where four users use BCI to control an agent on a virtual world to reach a target without any previous calibration process. Jonathan Grizou, Iñaki Iturrate, Luis Montesano, Pierre-Yves Oudeyer, Manuel Lopes 0001 |
AAAI | 4 |
| 2014 | Online Optimization of Teaching Sequences with Multi-Armed Bandits
Benjamin Clément, Pierre-Yves Oudeyer, Didier Roy, Manuel Lopes 0001 |
EDM | 2 |
| 2014 | Interactive Learning from Unlabeled Instructions
Jonathan Grizou, Iñaki Iturrate, Luis Montesano, Pierre-Yves Oudeyer, Manuel Lopes 0001 |
UAI | 4 |
| 2013 | The role of intrinsic motivations in learning sensorimotor vocal mappings: a developmental robotics studyabstractInternational audience Clément Moulin-Frier, Pierre-Yves Oudeyer |
INTERSPEECH | 2 |
| 2013 | The poppy humanoid robot: Leg design for biped locomotionabstractWe introduce a novel humanoid robotic platform designed to jointly address three central goals of humanoid robotics: 1) study the role of morphology in biped locomotion; 2) study full-body compliant physical human-robot interaction; 3) be robust while easy and fast to duplicate to facilitate experimentation. The taken approach relies on functional modeling of certain aspects of human morphology, optimizing materials and geometry, as well as on the use of 3D printing techniques. In this article, we focus on the presentation of the design of specific morphological parts related to biped locomotion: the hip, the thigh, the limb mesh and the knee. We present initial experiments showing properties of the robot when walking with the physical guidance of a human. Matthieu Lapeyre, Pierre Rouanet, Pierre-Yves Oudeyer |
IROS | 3 |
| 2013 | The Impact of Human-Robot Interfaces on the Learning of Visual ObjectsabstractThis paper studies the impact of interfaces, allowing nonexpert users to efficiently and intuitively teach a robot to recognize new visual objects. We present challenges that need to be addressed for real-world deployment of robots capable of learning new visual objects in interaction with everyday users. We argue that in addition to robust machine learning and computer vision methods, well-designed interfaces are crucial for learning efficiency. In particular, we argue that interfaces can be key in helping nonexpert users to collect good learning examples and, thus, improve the performance of the overall learning system. Then, we present four alternative human-robot interfaces: Three are based on the use of a mediating artifact (smartphone, wiimote, wiimote and laser), and one is based on natural human gestures (with a Wizard-of-Oz recognition system). These interfaces mainly vary in the kind of feedback provided to the user, allowing him to understand more or less easily what the robot is perceiving and, thus, guide his way of providing training examples differently. We then evaluate the impact of these interfaces, in terms of learning efficiency, usability, and user's experience, through a real world and large-scale user study. In this experiment, we asked participants to teach a robot 12 different new visual objects in the context of a robotic game. This game happens in a home-like environment and was designed to motivate and engage users in an interaction where using the system was meaningful. We then discuss results that show significant differences among interfaces. In particular, we show that interfaces such as the smartphone interface allows nonexpert users to intuitively provide much better training examples to the robot, which is almost as good as expert users who are trained for this task and are aware of the different visual perception and machine learning issues. We also show that artifact-mediated teaching is significantly more efficient for robot learning, and equally good in terms of usability and user's experience, than teaching thanks to a gesture-based human-like interaction. Pierre Rouanet, Pierre-Yves Oudeyer, Fabien Danieau, David Filliat |
IEEE Trans. Robotics | 2 |
| 2012 | Learning to recognize parallel combinations of human motion primitives with linguistic descriptions using non-negative matrix factorizationabstractWe present an approach, based on non-negative matrix factorization, for learning to recognize parallel combinations of initially unknown human motion primitives, associated with ambiguous sets of linguistic labels during training. In the training phase, the learner observes a human producing complex motions which are parallel combinations of initially unknown motion primitives. Each time the human shows a complex motion, he also provides high-level linguistic descriptions, consisting of a set of labels giving the name of the primitives inside the complex motion. From the observation of multimodal combinations of high-level labels with high-dimensional continuous unsegmented values representing complex motions, the learner must later on be able to recognize, through the production of the adequate set of labels, which are the motion primitives in a novel complex motion produced by a human, even if those combinations were never observed during training. We explain how this problem, as well as natural extensions, can be addressed using non-negative matrix factorization. Then, we show in an experiment in which a learner has to recognize the primitive motions of complex human dance choreographies, that this technique allows the system to infer with good performance the combinatorial structure of parallel combinations of unknown primitives. Olivier Mangin, Pierre-Yves Oudeyer |
IROS | 2 |
| 2012 | Exploration in Model-based Reinforcement Learning by Empirically Estimating Learning ProgressabstractFormal exploration approaches in model-based reinforcement learning estimate the accuracy of the currently learned model without consideration of the empirical prediction error. For example, PAC-MDP approaches such as Rmax base their model certainty on the amount of collected data, while Bayesian approaches assume a prior over the transition dynamics. We propose extensions to such approaches which drive exploration solely based on empirical estimates of the learner's accuracy and learning progress. We provide a ``sanity check'' theoretical analysis, discussing the behavior of our extensions in the standard stationary finite state-action case. We then provide experimental studies demonstrating the robustness of these exploration measures in cases of non-stationary environments or where original approaches are misled by wrong domain assumptions. Manuel Lopes 0001, Tobias Lang 0001, Marc Toussaint, Pierre-Yves Oudeyer |
NIPS | 4 |
| 2012 | Properties for efficient demonstrations to a socially guided intrinsically motivated learnerabstractThe combination of learning by intrinsic motivation and social learning has been shown to improve the learner's performance and gain precision over a wider range of motor skills, with for instance the SGIM-D learning algorithm [1]. Nevertheless, this bootstrapping a-priori depends on the demonstrations made by the teacher. We propose in this paper to examine this dependence: to what extend the quality of the demonstrations can influence the learning performance, and which are the characteristics of a good demonstrator. Results on a fishing experiment highlights the importance of the difficulty of the demonstrated tasks, as well as the structure of the actions demonstrated. Sao Mai Nguyen, Pierre-Yves Oudeyer |
RO-MAN | 2 |
| 2011 | A robotic game to evaluate interfaces used to show and teach visual objects to a robot in real world conditionabstractIn this paper, we present a real world user study of 4 interfaces designed to teach new visual objects to a social robot. This study was designed as a robotic game in order to maintain the user's motivation during the whole experiment. Among the 4 interfaces 3 were based on mediator objects such as an iPhone, a Wiimote and a laser pointer. They also provided the users with different kind of feedback of what the robot is perceiving. The fourth interface was a gesture based interface with a Wizard-of-Oz recognition system added to compare our mediator interfaces with a more natural interaction. Here, we specially studied the impact the interfaces have on the quality of the learning examples and the usability. We showed that providing non-expert users with a feedback of what the robot is perceiving is needed if one is interested in robust interaction. In particular, the iPhone interface allowed non-expert users to provide better learning examples due to its whole visual feedback. Furthermore, we also studied the user's gaming experience and found that in spite of its lower usability, the gestures interface was stated as entertaining as the other interfaces and increases the user's feeling of cooperating with the robot. Thus, we argue that this kind of interface could be well-suited for robotic game. Pierre Rouanet, Fabien Danieau, Pierre-Yves Oudeyer |
HRI | 3 |
| 2011 | Bio-inspired vertebral column, compliance and semi-passive dynamics in a lightweight humanoid robotabstractThis paper presents the humanoid robot Acroban. We study two main issues: 1) Compliance and semi-passive dynamics for locomotion of humanoid robots regarding robustness against unknown external perturbations; 2) The advantages of a bio-inspired multi-articulated vertebral column. We combine mechatronic compliance with structural compliance due to the use of flexible materials. And we explore how these capabilities allow to enforce morphological computation in the design of robust dynamic locomotion. We also investigate the use of compliance to design semi-passive motor primitives using the torso and the arms as a system of accumulation/release of potential/kinetic energy. Olivier Ly, Matthieu Lapeyre, Pierre-Yves Oudeyer |
IROS | 3 |
| 2010 | A study of three interfaces allowing non-expert users to teach new visual objects to a robot and their impact on learning efficiencyabstractWe developed three interfaces to allow non-expert users to teach name for new visual objects and compare them through user's studies in term of learning efficiency. Pierre Rouanet, Pierre-Yves Oudeyer, David Filliat |
HRI | 2 |
| 2010 | Intrinsically motivated goal exploration for active motor learning in robots: A case studyabstractWe introduce the Self-Adaptive Goal Generation - Robust Intelligent Adaptive Curiosity (SAGG-RIAC) algorithm as an intrinsically motivated goal exploration mechanism which allows a redundant robot to efficiently and actively learn its inverse kinematics. The main idea is to push the robot to perform babbling in the goal/operational space, as opposed to motor babbling in the actuator space, by self-generating goals actively and adaptively in regions of the goal space which provide a maximal competence improvement for reaching those goals. Then, a lower level active motor learning algorithm, inspired by the SSA algorithm, is used to allow the robot to locally explore how to reach a given self-generated goal. We present simulated experiments in a 32 dimensional continuous sensorimotor space showing that 1) exploration in the goal space can be a lot faster than exploration in the actuator space for learning the inverse kinematics of a redundant robot; 2) selecting goals based on the maximal improvement heuristics is statistically significantly more efficient than selecting goals randomly. Adrien Baranes, Pierre-Yves Oudeyer |
IROS | 2 |
| 2010 | Incremental local online Gaussian Mixture Regression for imitation learning of multiple tasksabstractGaussian Mixture Regression has been shown to be a powerful and easy-to-tune regression technique for imitation learning of constrained motor tasks in robots. Yet, current formulations are not suited when one wants a robot to learn incrementally and online a variety of new context-dependant tasks whose number and complexity is not known at programming time, and when the demonstrator is not allowed to tell the system when he introduces a new task (but rather the system should infer this from the continuous sensorimotor context). In this paper, we show that this limitation can be addressed by introducing an Incremental, Local and Online variation of Gaussian Mixture Regression (ILO-GMR) which successfully allows a simulated robot to learn incrementally and online new motor tasks through modelling them locally as dynamical systems, and able to use the sensorimotor context to cope with the absence of categorical information both during demonstrations and when a reproduction is asked to the system. Moreover, we integrate a complementary statistical technique which allows the system to incrementally learn various tasks which can be intrinsically defined in different frames of reference, which we call framings, without the need to tell the system which particular framing should be used for each task: this is inferred automatically by the system. Thomas Cederborg, Adrien Baranes, Pierre-Yves Oudeyer |
IROS | 4 |
| 2010 | Acroban the humanoid: Compliance for stabilization and human interactionabstractThis video presents the humanoid robot Acroban which is to our knowledge the first humanoid robot which is able to: 1) demonstrate playful, compliant and intuitive physical interaction with children; 2) at the same time move and walk dynamically while keeping its equilibrium even if unpredicted physical interactions are initiated by humans. Olivier Ly, Pierre-Yves Oudeyer |
IROS | 2 |
| 2009 | A comparison of three interfaces using handheld devices to intuitively drive and show objects to a social robot: the impact of underlying metaphorsabstractIn this paper, we present three human-robot interfaces using a handheld device as a mediator object between a human and a robot, allowing the human to intuitively drive the robot and show it objects. One of the interface is based on a virtual keyboard interface on the iPhone, another is a gesture based interface on the iPhone too and the last one is a Wiimote based interface. They were designed to be easy to use, especially by non-expert domestic users, and to span different metaphors of interaction. In order to compare them and study the impact of the metaphor on the interaction, we designed a user-study based on two obstacle courses. Each of the 25 participants performed two courses with two different interfaces (total of 100 trials). Although the three interfaces were rather equally efficient and all considered as satisfying by the participants, the iPhone gesture interface was largely preferred while the Wiimote was poorly chosen due to the impact of the chosen metaphor. Pierre Rouanet, Jerome Bechu, Pierre-Yves Oudeyer |
RO-MAN | 3 |
| 2007 | Self-organization in the evolution of shared systems of speech sounds: a computational studyabstractHow did culturally shared systems of combinatorial speech sounds initially appear in human evolution? This paper proposes the hypothesis that their bootstrapping may have happened rather easily if one assumes an individual capacity for vocal replication, and thanks to self-organization in the neural coupling of vocal modalities and in the coupling of babbling individuals. This hypothesis is embodied in agent-based computational experiments, that allow to show that crucial phenomena, including structural regularities and diversity of sound systems, can only be accounted if speech is considered as a complex adaptive system. Thus, the second objective of this paper is to show that integrative computational approaches, even if speculative in certain respects, might be key in the understanding of speech and its evolution 1 . Pierre-Yves Oudeyer |
INTERSPEECH | 1 |
| 2007 | Intrinsic Motivation Systems for Autonomous Mental DevelopmentabstractExploratory activities seem to be intrinsically rewarding for children and crucial for their cognitive development. Can a machine be endowed with such an intrinsic motivation system? This is the question we study in this paper, presenting a number of computational systems that try to capture this drive towards novel or curious situations. After discussing related research coming from developmental psychology, neuroscience, developmental robotics, and active learning, this paper presents the mechanism of Intelligent Adaptive Curiosity, an intrinsic motivation system which pushes a robot towards situations in which it maximizes its learning progress. This drive makes the robot focus on situations which are neither too predictable nor too unpredictable, thus permitting autonomous mental development. The complexity of the robot's activities autonomously increases and complex developmental sequences self-organize without being constructed in a supervised manner. Two experiments are presented illustrating the stage-like organization emerging with this mechanism. In one of them, a physical robot is placed on a baby play mat with objects that it can learn to manipulate. Experimental results show that the robot first spends time in situations which are easy to learn, then shifts its attention progressively to situations of increasing difficulty, avoiding situations in which nothing can be learned. Finally, these various results are discussed in relation to more complex forms of behavioral organization and data coming from developmental psychology. Pierre-Yves Oudeyer, Frédéric Kaplan, Verena V. Hafner |
IEEE Trans. Evol. Comput. | 1 |
| 2006 | Modeling interaction strategies using POS: An application to soccer robots
Jean-Luc Koning, Pierre-Yves Oudeyer |
Appl. Intell. | 2 |
| 2006 | Discovering communicationabstractWhat kind of motivation drives child language development? This article presents a computational model and a robotic experiment to articulate the hypothesis that children discover communication as a result of exploring and playing with their environment. The considered robotic agent is intrinsically motivated towards situations in which it optimally progresses in learning. To experience optimal learning progress, it must avoid situations already familiar but also situations where nothing can be learned. The robot is placed in an environment in which both communicating and non-communicating objects are present. As a consequence of its intrinsic motivation, the robot explores this environment in an organized manner focussing first on non-communicative activities and then discovering the learning potential of certain types of interactive behavior. In this experiment, the agent ends up being interested by communication through vocal interactions without having a specific drive for communication. Pierre-Yves Oudeyer, Frédéric Kaplan |
Connect. Sci. | 1 |
| 2005 | The self-organization of combinatoriality and phonotactics in vocalization systemsabstractThis paper shows how a society of agents can self-organize a shared vocalization system that is\ndiscrete, combinatorial and has a form of primitive phonotactics, starting from holistic inarticulate\nvocalizations. The originality of the system is that: (1) it does not include any explicit pressure for\ncommunication; (2) agents do not possess capabilities of coordinated interactions, in particular they\ndo not play language games; (3) agents possess no specific linguistic capacities; and (4) initially\nthere exists no convention that agents can use. As a consequence, the system shows how a primitive\nspeech code may bootstrap in the absence of a communication system between agents, i.e. before the\nappearance of language. Pierre-Yves Oudeyer |
Connect. Sci. | 1 |
| 2005 | Erratum to: "The production and recognition of emotions in speech: features and algorithms": [Int. J. Hum.-Comput. Stud 59 (2003) 157]
Pierre-Yves Oudeyer |
Int. J. Hum. Comput. Stud. | 1 |
| 2003 | The production and recognition of emotions in speech: features and algorithms
Pierre-Yves Oudeyer |
Int. J. Hum. Comput. Stud. | 1 |
| 2001 | Coupled Neural Maps for the Origins of Vowel Systems
Pierre-Yves Oudeyer |
ICANN | 1 |
| 2001 | Introduction to POS: A Protocol Operational SemanticsabstractIn this paper, we propose a system for representing interaction protocols called POS which is both Turing complete and determine a complete semantics of protocols. This work is inspired by the Structured Operational Semantics in programming languages. We precisely define POS and illustrate its power on an extended example. Jean-Luc Koning, Pierre-Yves Oudeyer |
Int. J. Cooperative Inf. Syst. | 2 |