Olivier Sigaud

dblp:50/5522 · DBLP profile ↗
← Back
49ranked-venue papers
4as first author
17since 2021 · last 2025
0000-0002-8544-0229ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 4 first-author · 16 since 2021Systems, architecture and hardware · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
abstract
Open-ended learning agents must efficiently prioritize goals in vast possibility spaces, focusing on those that maximize learning progress (LP). When such autotelic exploration is achieved by LLM agents trained with online RL in high-dimensional and evolving goal spaces, a key challenge for LP prediction is modeling one’s own competence, a form of metacognitive monitoring. Traditional approaches either require extensive sampling or rely on brittle expert-defined goal groupings. We introduce MAGELLAN, a metacognitive framework that lets LLM agents learn to predict their competence and learning progress online. By capturing semantic relationships between goals, MAGELLAN enables sample-efficient LP estimation and dynamic adaptation to evolving goal spaces through generalization. In an interactive learning environment, we show that MAGELLAN improves LP prediction efficiency and goal prioritization, being the only method allowing the agent to fully master a large and evolving goal space. These results demonstrate how augmenting LLM agents with a metacognitive ability for LP predictions can effectively scale curriculum learning to open-ended goal spaces.
Loris Gaven, Thomas Carta, Clément Romac, Cédric Colas, Sylvain Lamprier, Olivier Sigaud, Pierre-Yves Oudeyer
ICML6
2025 RT-HCP: Dealing with Inference Delays and Sample Efficiency to Learn Directly on Robotic Platforms
abstract
Learning a controller directly on the robot requires extreme sample efficiency. Model-based reinforcement learning (RL) methods are the most sample efficient, but they often suffer from a too long inference time to meet the robot control frequency requirements. In this paper, we address the sample efficiency and inference time challenges with two contributions. First, we define a general framework to deal with inference delays where the slow inference robot controller provides a sequence of actions to feed the control-hungry robotic platform without execution gaps. Then, we compare several RL algorithms in the light of this framework and propose RT-HCP, an algorithm that offers an excellent trade-off between performance, sample efficiency and inference time. We validate the superiority of RT-HCP with experiments where we learn a controller directly on a simple but high frequency FURUTA pendulum platform. Code: github.com/elasriz/RTHCP
Zakariae El Asri, Ibrahim Laiche, Clément Rambour, Olivier Sigaud, Nicolas Thome
IROS4
2025 Imagine Beyond ! Distributionally Robust Autoencoding for State Space Coverage in Online Reinforcement Learning
abstract
Goal-Conditioned Reinforcement Learning (GCRL) enables agents to autonomously acquire diverse behaviors, but faces major challenges in visual environments due to high-dimensional, semantically sparse observations. In the online setting, where agents learn representations while exploring, the latent space evolves with the agent's policy, to capture newly discovered areas of the environment. However, without incentivization to maximize state coverage in the representation, classical approaches based on auto-encoders may converge to latent spaces that over-represent a restricted set of states frequently visited by the agent. This is exacerbated in an intrinsic motivation setting, where the agent uses the distribution encoded in the latent space to sample the goals it learns to master. To address this issue, we propose to progressively enforce distributional shifts towards a uniform distribution over the full state space, to ensure a full coverage of skills that can be learned in the environment. We introduce DRAG (Distributionally Robust Auto-Encoding for GCRL), a method that combines the $\beta$-VAE framework with Distributionally Robust Optimization (DRO). DRAG leverage an adversarial neural weighter of training states of the VAE, to account for the mismatch between the current data distribution and unseen parts of the environment. This allows the agent to construct semantically meaningful latent spaces beyond its immediate experience. Our approach improves state space coverage and downstream control performance on hard exploration environments such as mazes and robotic control involving walls to bypass, without relying on pre-training nor prior environment knowledge.
Nicolas Castanet, Olivier Sigaud, Sylvain Lamprier
NeurIPS2
2024 Bridging Environments and Language with Rendering Functions and Vision-Language Models
abstract
Vision-language models (VLMs) have tremendous potential for grounding language, and thus enabling language-conditioned agents (LCAs) to perform diverse tasks specified with text. This has motivated the study of LCAs based on reinforcement learning (RL) with rewards given by rendering images of an environment and evaluating those images with VLMs. If single-task RL is employed, such approaches are limited by the cost and time required to train a policy for each new task. Multi-task RL (MTRL) is a natural alternative, but requires a carefully designed corpus of training tasks and does not always generalize reliably to new tasks. Therefore, this paper introduces a novel decomposition of the problem of building an LCA: first find an environment configuration that has a high VLM score for text describing a task; then use a (pretrained) goal-conditioned policy to reach that configuration. We also explore several enhancements to the speed and quality of VLM-based LCAs, notably, the use of distilled models, and the evaluation of configurations from multiple viewpoints to resolve the ambiguities inherent in a single 2D view. We demonstrate our approach on the Humanoid environment, showing that it results in LCAs that outperform MTRL baselines in zero-shot generalization, without requiring any textual task descriptions or other forms of environment-specific annotation during training.
Théo Cachet, Christopher R. Dance, Olivier Sigaud
ICML3
2023 Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning
abstract
Recent works successfully leveraged Large Language Models’ (LLM) abilities to capture abstract knowledge about world’s physics to solve decision-making problems. Yet, the alignment between LLMs’ knowledge and the environment can be wrong and limit functional competence due to lack of grounding. In this paper, we study an approach (named GLAM) to achieve this alignment through functional grounding: we consider an agent using an LLM as a policy that is progressively updated as the agent interacts with the environment, leveraging online Reinforcement Learning to improve its performance to solve goals. Using an interactive textual environment designed to study higher-level forms of functional grounding, and a set of spatial and navigation tasks, we study several scientific questions: 1) Can LLMs boost sample efficiency for online learning of various RL tasks? 2) How can it boost different forms of generalization? 3) What is the impact of online learning? We study these questions by functionally grounding several variants (size, architecture) of FLAN-T5.
Thomas Carta, Clément Romac, Sylvain Lamprier, Olivier Sigaud, Pierre-Yves Oudeyer
ICML5
2023 Stein Variational Goal Generation for adaptive Exploration in Multi-Goal Reinforcement Learning
abstract
In multi-goal Reinforcement Learning, an agent can share experience between related training tasks, resulting in better generalization for new tasks at test time. However, when the goal space has discontinuities and the reward is sparse, a majority of goals are difficult to reach. In this context, a curriculum over goals helps agents learn by adapting training tasks to their current capabilities. In this work, we propose Stein Variational Goal Generation (SVGG), which samples goals of intermediate difficulty for the agent, by leveraging a learned predictive model of its goal reaching capabilities. The distribution of goals is modeled with particles that are attracted in areas of appropriate difficulty using Stein Variational Gradient Descent. We show that SVGG outperforms state-of-the-art multi-goal Reinforcement Learning methods in terms of success coverage in hard exploration problems, and demonstrate that it is endowed with a useful recovery property when the environment changes.
Nicolas Castanet, Olivier Sigaud, Sylvain Lamprier
ICML2
2023 Combining Evolution and Deep Reinforcement Learning for Policy Search: A Survey
abstract
Deep neuroevolution and deep Reinforcement Learning have received a lot of attention over the past few years. Some works have compared them, highlighting their pros and cons, but an emerging trend combines them so as to benefit from the best of both worlds. In this article, we provide a survey of this emerging trend by organizing the literature into related groups of works and casting all the existing combinations in each group into a generic framework. We systematically cover all easily available papers irrespective of their publication status, focusing on the combination mechanisms rather than on the experimental results. In total, we cover 45 algorithms more recent than 2017. We hope this effort will favor the growth of the domain by facilitating the understanding of the relationships between the methods, leading to deeper analyses, outlining missing useful comparisons and suggesting new combinations of mechanisms.
Olivier Sigaud
ACM Trans. Evol. Learn. Optim.1
2022 Diversity policy gradient for sample efficient quality-diversity optimization
abstract
A fascinating aspect of nature lies in its ability to produce a large and diverse collection of organisms that are all high-performing in their niche. By contrast, most AI algorithms focus on finding a single eficient solution to a given problem. Aiming for diversity in addition to performance is a convenient way to deal with the exploration-exploitation trade-off that plays a central role in learning. It also allows for increased robustness when the returned collection contains several working solutions to the considered problem, making it well-suited for real applications such as robotics. Quality-Diversity (QD) methods are evolutionary algorithms designed for this purpose. This paper proposes a novel algorithm, qd-pg, which combines the strength of Policy Gradient algorithms and Quality Diversity approaches to produce a collection of diverse and high-performing neural policies in continuous control environments. The main contribution of this work is the introduction of a Diversity Policy Gradient (DPG) that exploits information at the time-step level to drive policies towards more diversity in a sample-efficient manner. Specifically, qd-pg selects neural controllers from a map-elites grid and uses two gradient-based mutation operators to improve both quality and diversity. Our results demonstrate that qd-pg is significantly more sample-eficient than its evolutionary competitors.
Thomas Pierrot, Valentin Macé, Félix Chalumeau, Arthur Flajolet, Geoffrey Cideron, Karim Beguir, Antoine Cully, Olivier Sigaud, Nicolas Perrin-Gilbert
GECCO8
2022 Neural Architecture Search for Fracture Classification
abstract
The adoption by radiologists of deep-learning based solutions to the bone fracture problem has helped improved diagnostic performances and patient care. The base models behind these tools were initially designed to solve problems on natural images, favoring transfer learning between standard image datasets and sets of radiographs. Those architectures could yet be made more specific to radiographs using neural architecture search (NAS). Unfortunately, current NAS approaches do not benefit from transfer learning. In this paper, we introduce an efficient scheme to exploit transfer learning when performing NAS. Using our approach, we validate the architecture tailoring paradigm to radiographs. On a custom fracture classification task, we find a new model with improved performances and reduced computational overhead over its counterparts pre-trained on ImageNet.
Aloïs Pourchot, Kevin Bailly, Alexis Ducarouge, Olivier Sigaud
ICIP4
2022 Divide & Conquer Imitation Learning
abstract
International audience
Alexandre Chenu, Nicolas Perrin-Gilbert, Olivier Sigaud
IROS3
2022 EAGER: Asking and Answering Questions for Automatic Reward Shaping in Language-guided RL
abstract
Reinforcement learning (RL) in long horizon and sparse reward tasks is notoriously difficult and requires a lot of training steps. A standard solution to speed up the process is to leverage additional reward signals, shaping it to better guide the learning process.In the context of language-conditioned RL, the abstraction and generalisation properties of the language input provide opportunities for more efficient ways of shaping the reward.In this paper, we leverage this idea and propose an automated reward shaping method where the agent extracts auxiliary objectives from the general language goal. These auxiliary objectives use a question generation (QG) and a question answering (QA) system: they consist of questions leading the agent to try to reconstruct partial information about the global goal using its own trajectory.When it succeeds, it receives an intrinsic reward proportional to its confidence in its answer. This incentivizes the agent to generate trajectories which unambiguously explain various aspects of the general language goal.Our experimental study using various BabyAI environments shows that this approach, which does not require engineer intervention to design the auxiliary objectives, improves sample efficiency by effectively directing the exploration.
Thomas Carta, Pierre-Yves Oudeyer, Olivier Sigaud, Sylvain Lamprier
NeurIPS3
2022 Pragmatically Learning from Pedagogical Demonstrations in Multi-Goal Environments
abstract
Learning from demonstration methods usually leverage close to optimal demonstrations to accelerate training. By contrast, when demonstrating a task, human teachers deviate from optimal demonstrations and pedagogically modify their behavior by giving demonstrations that best disambiguate the goal they want to demonstrate. Analogously, human learners excel at pragmatically inferring the intent of the teacher, facilitating communication between the two agents. These mechanisms are critical in the few demonstrations regime, where inferring the goal is more difficult. In this paper, we implement pedagogy and pragmatism mechanisms by leveraging a Bayesian model of Goal Inference from demonstrations. We highlight the benefits of this model in multi-goal teacher-learner setups with two artificial agents that learn with goal-conditioned Reinforcement Learning. We show that combining BGI-agents (a pedagogical teacher and a pragmatic learner) results in faster learning and reduced goal ambiguity over standard learning from demonstrations, especially in the few demonstrations regime.
Hugo Caselles-Dupré, Olivier Sigaud, Mohamed Chetouani
NeurIPS2
2022 An extensive appraisal of weight-sharing on the NAS-Bench-101 benchmark
Aloïs Pourchot, Kevin Bailly, Alexis Ducarouge, Olivier Sigaud
Neurocomputing4
2022 Autotelic Agents with Intrinsically Motivated Goal-Conditioned Reinforcement Learning: A Short Survey
abstract
Building autonomous machines that can explore open-ended environments, discover possible interactions and build repertoires of skills is a general objective of artificial intelligence. Developmental approaches argue that this can only be achieved by autotelic agents: intrinsically motivated learning agents that can learn to represent, generate, select and solve their own problems. In recent years, the convergence of developmental approaches with deep reinforcement learning (RL) methods has been leading to the emergence of a new field: developmental reinforcement learning. Developmental RL is concerned with the use of deep RL algorithms to tackle a developmental problem— the intrinsically motivated acquisition of open-ended repertoires of skills. The self-generation of goals requires the learning of compact goal encodings as well as their associated goal-achievement functions. This raises new challenges compared to standard RL algorithms originally designed to tackle pre-defined sets of goals using external reward signals. The present paper introduces developmental RL and proposes a computational framework based on goal-conditioned RL to tackle the intrinsically motivated skills acquisition problem. It proceeds to present a typology of the various goal representations used in the literature, before reviewing existing methods to learn to represent and prioritize goals in autonomous systems. We finally close the paper by discussing some open challenges in the quest of intrinsically motivated skills acquisition.
Cédric Colas, Tristan Karch, Olivier Sigaud, Pierre-Yves Oudeyer
J. Artif. Intell. Res.3
2021 Selection-Expansion: A Unifying Framework for Motion-Planning and Diversity Search Algorithms
Alexandre Chenu, Nicolas Perrin-Gilbert, Stéphane Doncieux, Olivier Sigaud
ICANN (4)4
2021 First-Order and Second-Order Variants of the Gradient Descent in a Unified Framework
Thomas Pierrot, Nicolas Perrin-Gilbert, Olivier Sigaud
ICANN (2)3
2021 Grounding Language to Autonomously-Acquired Skills via Goal Generation
Ahmed Akakzia, Cédric Colas, Pierre-Yves Oudeyer, Mohamed Chetouani, Olivier Sigaud
ICLR5
2020 PBCS: Efficient Exploration and Exploitation Using a Synergy Between Reinforcement Learning and Motion Planning
abstract
The exploration-exploitation trade-off is at the heart of reinforcement learning (RL). However, most continuous control benchmarks used in recent RL research only require local exploration. This led to the development of algorithms that have basic exploration capabilities, and behave poorly in benchmarks that require more versatile exploration. For instance, as demonstrated in our empirical study, state-of-the-art RL algorithms such as DDPG and TD3 are unable to steer a point mass in even small 2D mazes. In this paper, we propose a new algorithm called "Plan, Backplay, Chain Skills" (PBCS) that combines motion planning and reinforcement learning to solve hard exploration environments. In a first phase, a motion planning algorithm is used to find a single good trajectory, then an RL algorithm is trained using a curriculum derived from the trajectory, by combining a variant of the Backplay algorithm and skill chaining. We show that this method outperforms state-of-the-art RL algorithms in 2D maze environments of various sizes, and is able to improve on the trajectory obtained by the motion planning phase.
Guillaume Matheron, Nicolas Perrin-Gilbert, Olivier Sigaud
ICANN (2)3
2020 Understanding Failures of Deterministic Actor-Critic with Continuous Action Spaces and Sparse Rewards
Guillaume Matheron, Nicolas Perrin-Gilbert, Olivier Sigaud
ICANN (2)3
2020 TIRL: Enriching Actor-Critic RL with non-expert human teachers and a Trust Model
abstract
Reinforcement learning (RL) algorithms have been demonstrated to be very attractive tools to train agents to achieve sequential tasks. However, these algorithms require too many training data to converge to be efficiently applied to physical robots. By using a human teacher, the learning process can be made faster and more robust, but the overall performance heavily depends on the quality and availability of teacher demonstrations or instructions. In particular, when these teaching signals are inadequate, the agent may fail to learn an optimal policy. In this paper, we introduce a trust-based interactive task learning approach. We propose an RL architecture able to learn both from environment rewards and from various sparse teaching signals provided by non-expert teachers, using an actor-critic agent, a human model and a trust model. We evaluate the performance of this architecture on 4 different setups using a maze environment with different simulated teachers and show that the benefits of the trust model.
Félix Rutard, Olivier Sigaud, Mohamed Chetouani
RO-MAN2
2020 Interactively shaping robot behaviour with unlabeled human instructions
Anis Najar, Olivier Sigaud, Mohamed Chetouani
Auton. Agents Multi Agent Syst.2
2019 CEM-RL: Combining evolutionary and gradient-based methods for policy search
Aloïs Pourchot, Olivier Sigaud
ICLR (Poster)2
2019 CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning
abstract
In open-ended environments, autonomous learning agents must set their own goals and build their own curriculum through an intrinsically motivated exploration. They may consider a large diversity of goals, aiming to discover what is controllable in their environments, and what is not. Because some goals might prove easy and some impossible, agents must actively select which goal to practice at any moment, to maximize their overall mastery on the set of learnable goals. This paper proposes CURIOUS , an algorithm that leverages 1) a modular Universal Value Function Approximator with hindsight learning to achieve a diversity of goals of different kinds within a unique policy and 2) an automated curriculum learning mechanism that biases the attention of the agent towards goals maximizing the absolute learning progress. Agents focus sequentially on goals of increasing complexity, and focus back on goals that are being forgotten. Experiments conducted in a new modular-goal robotic environment show the resulting developmental self-organization of a learning curriculum, and demonstrate properties of robustness to distracting goals, forgetting and changes in body properties.
Cédric Colas, Pierre-Yves Oudeyer, Olivier Sigaud, Pierre Fournier, Mohamed Chetouani
ICML3
2019 Learning Compositional Neural Programs with Recursive Tree Search and Planning
abstract
We propose a novel reinforcement learning algorithm, AlphaNPI, that incorpo- rates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero. NPI contributes structural biases in the form of modularity, hierarchy and recursion, which are helpful to reduce sample complexity, improve generalization and in- crease interpretability. AlphaZero contributes powerful neural network guided search algorithms, which we augment with recursion. AlphaNPI only assumes a hierarchical program specification with sparse rewards: 1 when the program execution satisfies the specification, and 0 otherwise. This specification enables us to overcome the need for strong supervision in the form of execution traces and consequently train NPI models effectively with reinforcement learning. The experiments show that AlphaNPI can sort as well as previous strongly supervised NPI variants. The AlphaNPI agent is also trained on a Tower of Hanoi puzzle with two disks and is shown to generalize to puzzles with an arbitrary number of disks. The experiments also show that when deploying our neural network policies, it is advantageous to do planning with guided Monte Carlo tree search.
Thomas Pierrot, Guillaume Ligner, Scott E. Reed, Olivier Sigaud, Nicolas Perrin-Gilbert, Alexandre Laterre, David Kas, Karim Beguir, Nando de Freitas
NeurIPS4
2019 Policy search in continuous action domains: An overview
Olivier Sigaud, Freek Stulp
Neural Networks1
2018 Unsupervised Learning of Goal Spaces for Intrinsically Motivated Goal Exploration
Alexandre Péré, Sébastien Forestier, Olivier Sigaud, Pierre-Yves Oudeyer
ICLR (Poster)3
2018 GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
abstract
In continuous action domains, standard deep reinforcement learning algorithms like DDPG suffer from inefficient exploration when facing sparse or deceptive reward problems. Conversely, evolutionary and developmental methods focusing on exploration like Novelty Search, Quality-Diversity or Goal Exploration Processes explore more robustly but are less efficient at fine-tuning policies using gradient-descent. In this paper, we present the GEP-PG approach, taking the best of both worlds by sequentially combining a Goal Exploration Process and two variants of DDPG . We study the learning performance of these components and their combination on a low dimensional deceptive reward problem and on the larger Half-Cheetah benchmark. We show that DDPG fails on the former and that GEP-PG improves over the best DDPG variant in both environments.
Cédric Colas, Olivier Sigaud, Pierre-Yves Oudeyer
ICML2
2017 Tensor Based Knowledge Transfer Across Skill Categories for Robot Control
abstract
Advances in hardware and learning for control are enabling robots to perform increasingly dextrous and dynamic control tasks. These skills typically require a prohibitive amount of exploration for reinforcement learning, and so are commonly achieved by imitation learning from manual demonstration. The costly non-scalable nature of manual demonstration has motivated work into skill generalisation, e.g., through contextual policies and options. Despite good results, existing work along these lines is limited to generalising across variants of one skill such as throwing an object to different locations. In this paper we go significantly further and investigate generalisation across qualitatively different classes of control skills. In particular, we introduce a class of neural network controllers that can realise four distinct skill classes: reaching, object throwing, casting, and ball-in-cup. By factorising the weights of the neural network, we are able to extract transferrable latent skills, that enable dramatic acceleration of learning in cross-task transfer. With a suitable curriculum, this allows us to learn challenging dextrous control tasks like ball-in-cup from scratch with pure reinforcement learning.
Chenyang Zhao 0007, Timothy M. Hospedales, Freek Stulp, Olivier Sigaud
IJCAI4
2016 Training a robot with evaluative feedback and unlabeled guidance signals
abstract
In this paper, we present a new method for training a robot by natural interaction using evaluative feedback and unlabeled guidance signals. Feedback signals are directly mapped to reward values and used for learning both the task and the meaning of the guidance signals. The learned guidance signals are used in return to bootstrap task learning. We propose to use unlabeled guidance signals as an alternative solution to preprogrammed guidance. We evaluate our method both in simulation and on a real robot.
Anis Najar, Olivier Sigaud, Mohamed Chetouani
RO-MAN2
2015 Variance modulated task prioritization in Whole-Body Control
abstract
Whole-Body Control methods offer the potential to execute several tasks on highly redundant robots, such as humanoids. Unfortunately, task combinations often result in incompatibilities which generate undesirable behaviors. Prioritization techniques can prevent tasks from perturbing one another but often to the detriment of the lower precedence tasks. For many tasks, static prioritization is not necessary or even appropriate because tasks can often be achieved in variable ways, as in reaching. In this paper, we show that such task variability can be used to modulate task priorities during execution, to temporarily deviate certain tasks as needed, in the presence of incompatibilities. We first present a method for mapping from task variance to task priority and then provide an approach for computing task variance. Through three common conflict scenarios, we demonstrate that mapping from task variance to priorities reactively solves a number of task incompatibilities.
Ryan Lober, Vincent Padois, Olivier Sigaud
IROS3
2015 Many regression algorithms, one unified model: A review
Freek Stulp, Olivier Sigaud
Neural Networks2
2014 Modelling Individual Differences in the Form of Pavlovian Conditioned Approach Responses: A Dual Learning Systems Approach with Factored Representations
abstract
Reinforcement Learning has greatly influenced models of conditioning, providing powerful explanations of acquired behaviour and underlying physiological observations. However, in recent autoshaping experiments in rats, variation in the form of Pavlovian conditioned responses (CRs) and associated dopamine activity, have questioned the classical hypothesis that phasic dopamine activity corresponds to a reward prediction error-like signal arising from a classical Model-Free system, necessary for Pavlovian conditioning. Over the course of Pavlovian conditioning using food as the unconditioned stimulus (US), some rats (sign-trackers) come to approach and engage the conditioned stimulus (CS) itself - a lever - more and more avidly, whereas other rats (goal-trackers) learn to approach the location of food delivery upon CS presentation. Importantly, although both sign-trackers and goal-trackers learn the CS-US association equally well, only in sign-trackers does phasic dopamine activity show classical reward prediction error-like bursts. Furthermore, neither the acquisition nor the expression of a goal-tracking CR is dopamine-dependent. Here we present a computational model that can account for such individual variations. We show that a combination of a Model-Based system and a revised Model-Free system can account for the development of distinct CRs in rats. Moreover, we show that revising a classical Model-Free system to individually process stimuli by using factored representations can explain why classical dopaminergic patterns may be observed for some rats and not for others depending on the CR they develop. In addition, the model can account for other behavioural and pharmacological results obtained using the same, or similar, autoshaping procedures. Finally, the model makes it possible to draw a set of experimental predictions that may be verified in a modified experimental protocol. We suggest that further investigation of factored representations in computational neuroscience studies may be useful.
Florian Lesaint, Olivier Sigaud, Shelly B. Flagel, Terry E. Robinson, Mehdi Khamassi
PLoS Comput. Biol.2
2013 Gated Autoencoders with Tied Input Weights
abstract
The semantic interpretation of images is one of the core applications of deep learning. Several techniques have been recently proposed to model the relation between two images, with application to pose estimation, action recognition or invariant object recognition. Among these techniques, higher-order Boltzmann machines or relational autoencoders consider projections of the images on different subspaces and intermediate layers act as transformation specific detectors. In this work, we extend the mathematical study of (Memisevic, 2012b) to show that it is possible to use a unique projection for both images in a way that turns intermediate layers as spectrum encoders of transformations. We show that this results in networks that are easier to tune and have greater generalization capabilities.
Alain Droniou, Olivier Sigaud
ICML (2)2
2012 Path Integral Policy Improvement with Covariance Matrix Adaptation
Freek Stulp, Olivier Sigaud
ICML2
2012 Autonomous online learning of velocity kinematics on the iCub: A comparative study
abstract
In the last years, several regression algorithms have been proposed to learn accurate mechanical models of robots. Comparisons are proposed at the conceptual level or through the use of recorded databases, but they deliver limited conclusions with respect to the real performance of these algorithms in their true context of use, i.e. online learning on the real robot interacting with its environment, within a feedback control loop. In this paper, we provide an empirical study of three state-of-the-art regression methods through online learning on the iCub robot holding a tool. We show that they can effectively learn a visuo-motor kinematic model for a simple visual servoing task in a very limited time (few minutes), without making any a priori hypothesis on the geometry of the robot and its tool. Furthermore, we can draw from the results some stronger conclusions about the comparison of the algorithms than previous studies based on databases.
Alain Droniou, Serena Ivaldi, Vincent Padois, Olivier Sigaud
IROS4
2011 Learning cost-efficient control policies with XCSF: generalization capabilities and further improvement
abstract
In this paper we present a method based on the "learning from demonstration" paradigm to get a cost-efficient control policy in a continuous state and action space. The controlled plant is a two degrees-of-freedom planar arm actuated by six muscles. We learn a parametric control policy with XCSF from a few near-optimal trajectories, and we study its capability to generalize over the whole reachable space. Furthermore, we show that an additional Cross-Entropy Policy Search method can improve the global performance of the parametric controller.
Didier Marin, Jérémie Decock, Lionel Rigoux, Olivier Sigaud
GECCO4
2009 Transfer of knowledge for a climbing Virtual Human: A reinforcement learning approach
abstract
In the reinforcement learning literature, transfer is the capability to reuse on a new problem what has been learnt from previous experiences on similar problems. Adapting transfer properties for robotics is a useful challenge because it can reduce the time spent in the first exploration phase on a new problem. In this paper we present a transfer framework adapted to the case of a climbing virtual human (VH). We show that our VH learns faster to climb a wall after having learnt on a different previous wall.
Benoit Libeau, Alain Micaelli, Olivier Sigaud
ICRA3
2009 Control of redundant robots using learned models: An operational space control approach
abstract
We present an adaptive control approach combining forward kinematics model learning methods with the operational space control approach. This combination endows the robot with the ability to realize hierarchically organised learned tasks in parallel, using tasks null space projectors built upon the learned models. We illustrate the proposed method on a simulated 3 degrees of freedom planar robot. This system is used as a benchmark to compare our method to an alternative approach based on learning an extended Jacobian. We show the better versatility of the retained approach with respect to the latter.
Camille Salaün, Vincent Padois, Olivier Sigaud
IROS3
2009 Considering Unseen States as Impossible in Factored Reinforcement Learning
Olga Kozlova, Olivier Sigaud, Pierre-Henri Wuillemin, Christophe Meyer
ECML/PKDD (1)2
2008 A comparison between ATNoSFERES and Learning Classifier Systems on non-Markov problems
Samuel Landau, Olivier Sigaud
Inf. Sci.2
2007 Learning classifier systems: a survey
Olivier Sigaud, Stewart W. Wilson
Soft Comput.1
2006 Learning the structure of Factored Markov Decision Processes in reinforcement learning problems
abstract
Recent decision-theoric planning algorithms are able to find optimal solutions in large problems, using Factored Markov Decision Processes (FMDPs). However, these algorithms need a perfect knowledge of the structure of the problem. In this paper, we propose SDYNA, a general framework for addressing large reinforcement learning problems by trial-and-error and with no initial knowledge of their structure. SDYNA integrates incremental planning algorithms based on FMDPs with supervised learning techniques building structured representations of the problem. We describe SPITI, an instantiation of SDYNA, that uses incremental decision tree induction to learn the structure of a problem combined with an incremental version of the Structured Value Iteration algorithm. We show that SPITI can build a factored representation of a reinforcement learning problem and may improve the policy faster than tabular reinforcement learning algorithms by exploiting the generalization property of decision tree induction algorithms.
Thomas Degris, Olivier Sigaud, Pierre-Henri Wuillemin
ICML2
2006 Chi-square Tests Driven Method for Learning the Structure of Factored MDPs
Thomas Degris, Olivier Sigaud, Pierre-Henri Wuillemin
UAI2
2005 ATNoSFERES revisited
abstract
International audience
Samuel Landau, Olivier Sigaud, Marc Schoenauer
GECCO2
2004 Improving MACS Thanks to a Comparison with 2TBNs
Olivier Sigaud, Thierry Gourdin, Pierre-Henri Wuillemin
GECCO (2)1
2004 Rapid response of head direction cells to reorienting visual cues: a computational model
Thomas Degris, Olivier Sigaud, Sidney I. Wiener, Angelo Arleo
Neurocomputing2
2003 Designing Efficient Exploration with MACS: Modules and Function Approximation
Pierre Gérard, Olivier Sigaud
GECCO2
2002 A Comparison Between ATNoSFERES And XCSM
Samuel Landau, Sébastien Picault, Olivier Sigaud, Pierre Gérard
GECCO3
2002 YACS: a new learning classifier system using anticipation
Pierre Gérard, Wolfgang Stolzmann, Olivier Sigaud
Soft Comput.3