Alban Laflaquière

dblp:123/6900 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
2since 2021 · last 2024
0000-0002-3842-0858ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 2 since 2021Systems, architecture and hardware · 5 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 60% Robot navigation and mapping · 17% Representation and self-supervised learning · 17%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
exploration
0.412020
Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020
Machine learning › Reinforcement learning › exploration
novelty-based exploration
0.412020
Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020
Machine learning › Reinforcement learning
policy learning
0.412020
Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020
Robotics › Robot navigation and mapping
spatial representation
0.412019
Unsupervised Emergence of Egocentric Spatial Structure from Sensorimotor Prediction · NeurIPS 2019
Robotics › Motion planning and robot control
robot learning
0.112019
Unsupervised Emergence of Egocentric Spatial Structure from Sensorimotor Prediction · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

population-based divergent search · 0.4autoencoder · 0.4sensorimotor prediction · 0.4predictive learning · 0.4
YearPublicationVenuePosition
2024 Discovering and Exploiting Sparse Rewards in a Learned Behavior Space
abstract
Learning optimal policies in sparse rewards settings is difficult as the learning agent has little to no feedback on the quality of its actions. In these situations, a good strategy is to focus on exploration, hopefully leading to the discovery of a reward signal to improve on. A learning algorithm capable of dealing with this kind of setting has to be able to (1) explore possible agent behaviors and (2) exploit any possible discovered reward. Exploration algorithms have been proposed that require the definition of a low-dimension behavior space, in which the behavior generated by the agent's policy can be represented. The need to design a priori this space such that it is worth exploring is a major limitation of these algorithms. In this work, we introduce STAX, an algorithm designed to learn a behavior space on-the-fly and to explore it while optimizing any reward discovered (see Figure 1). It does so by separating the exploration and learning of the behavior space from the exploitation of the reward through an alternating two-step process. In the first step, STAX builds a repertoire of diverse policies while learning a low-dimensional representation of the high-dimensional observations generated during the policies evaluation. In the exploitation step, emitters optimize the performance of the discovered rewarding solutions. Experiments conducted on three different sparse reward environments show that STAX performs comparably to existing baselines while requiring much less prior information about the task as it autonomously builds the behavior space it explores.
Giuseppe Paolo, Miranda Coninx, Alban Laflaquière, Stéphane Doncieux
Evol. Comput.3
2021 Sparse reward exploration via novelty search and emitters
abstract
Reward-based optimization algorithms require both exploration, to find rewards, and exploitation, to maximize performance. The need for efficient exploration is even more significant in sparse reward settings, in which performance feedback is given sparingly, thus rendering it unsuitable for guiding the search process. In this work, we introduce the SparsE Reward Exploration via Novelty and Emitters (SERENE) algorithm, capable of efficiently exploring a search space, as well as optimizing rewards found in potentially disparate areas. Contrary to existing emitters-based approaches, SERENE separates the search space exploration and reward exploitation into two alternating processes. The first process performs exploration through Novelty Search, a divergent search algorithm. The second one exploits discovered reward areas through emitters, i.e. local instances of population-based optimization algorithms. A meta-scheduler allocates a global computational budget by alternating between the two processes, ensuring the discovery and efficient exploitation of disjoint reward areas. SERENE returns both a collection of diverse solutions covering the search space and a collection of high-performing solutions for each distinct reward area. We evaluate SERENE on various sparse reward environments and show it compares favorably to existing baselines.
Giuseppe Paolo, Miranda Coninx, Stéphane Doncieux, Alban Laflaquière
GECCO4
2020 Novelty search makes evolvability inevitable
abstract
Evolvability is an important feature that impacts the ability of evolutionary processes to find interesting novel solutions and to deal with changing conditions of the problem to solve. The estimation of evolvability is not straight-forward and is generally too expensive to be directly used as selective pressure in the evolutionary process. Indirectly promoting evolvability as a side effect of other easier and faster to compute selection pressures would thus be advantageous. In an unbounded behavior space, it has already been shown that evolvable individuals naturally appear and tend to be selected as they are more likely to invade empty behavior niches. Evolvability is thus a natural byproduct of the search in this context. However, practical agents and environments often impose limits on the reachable behavior space. How do these boundaries impact evolvability? In this context, can evolvability still be promoted without explicitly rewarding it? We show that Novelty Search implicitly creates a pressure for high evolvability even in bounded behavior spaces, and explore the reasons for such a behavior. More precisely we show that, throughout the search, the dynamic evaluation of novelty rewards individuals which are very mobile in the behavior space, which in turn promotes evolvability.
Stéphane Doncieux, Giuseppe Paolo, Alban Laflaquière, Miranda Coninx
GECCO3
2020 Unsupervised Learning and Exploration of Reachable Outcome Space
abstract
Performing Reinforcement Learning in sparse rewards settings, with very little prior knowledge, is a challenging problem since there is no signal to properly guide the learning process. In such situations, a good search strategy is fundamental. At the same time, not having to adapt the algorithm to every single problem is very desirable. Here we introduce TAXONS, a Task Agnostic eXploration of Outcome spaces through Novelty and Surprise algorithm. Based on a population-based divergent-search approach, it learns a set of diverse policies directly from high-dimensional observations, without any task-specific information. TAXONS builds a repertoire of policies while training an autoencoder on the high-dimensional observation of the final state of the system to build a low-dimensional outcome space. The learned outcome space, combined with the reconstruction error, is used to drive the search for new policies. Results show that TAXONS can find a diverse set of controllers, covering a good part of the ground-truth outcome space, while having no information about such space.
Giuseppe Paolo, Alban Laflaquière, Miranda Coninx, Stéphane Doncieux
ICRA2
2019 Novelty search: a theoretical perspective
abstract
Novelty Search is an exploration algorithm driven by the novelty of a behavior. The same individual evaluated at different generations has different fitness values. The corresponding fitness landscape is thus constantly changing and if, at the scale of a single generation, the metaphor of a fitness landscape with peaks and valleys still holds, this is not the case anymore at the scale of the whole evolutionary process. How does this kind of algorithms behave? Is it possible to define a model that would help understand how it works? This understanding is critical to analyse existing Novelty Search variants and design new and potentially more efficient ones. We assert that Novelty Search asymptotically behaves like a uniform random search process in the behavior space. This is an interesting feature, as it is not possible to directly sample in this space: the algorithm has a direct access to the genotype space only, whose relationship to the behavior space is complex. We describe the model and check its consistency on a classical Novelty Search experiment. We also show that it sheds a new light on results of the literature and suggests future research work.
Stéphane Doncieux, Alban Laflaquière, Miranda Coninx
GECCO2
2019 A Bi-directional Multiple Timescales LSTM Model for Grounding of Actions and Verbs
abstract
In this paper we present a neural architecture to learn a bi-directional mapping between actions and language. We implement a Multiple Timescale Long Short-Term Memory (MT-LSTM) network comprised of 7 layers with different timescale factors, to connect actions to language without explicitly learning an intermediate representation. Instead, the model self-organizes such representations at the level of a slow-varying latent layer, linking action branch and language branch at the center. We train the model in a bi-directional way, learning how to produce a sentence from a certain action sequence input and, simultaneously, how to generate an action sequence given a sentence as input. Furthermore we show this model preserves some of the generalization behaviour of Multiple Timescale Recurrent Neural Networks (MTRNN) in generating sentences and actions that were not explicitly trained. We compare this model with a number of different baseline models, confirming the importance of both the bi-directional training and the multiple timescales architecture. Finally, the network was evaluated on motor actions performed by an iCub robot and their corresponding letter-based description. The results of these experiments are presented at the end of the paper.
Alexandre Antunes, Alban Laflaquière, Tetsuya Ogata, Angelo Cangelosi
IROS2
2019 Unsupervised Emergence of Egocentric Spatial Structure from Sensorimotor Prediction
abstract
Despite its omnipresence in robotics application, the nature of spatial knowledge and the mechanisms that underlie its emergence in autonomous agents are still poorly understood. Recent theoretical works suggest that the Euclidean structure of space induces invariants in an agent’s raw sensorimotor experience. We hypothesize that capturing these invariants is beneficial for sensorimotor prediction and that, under certain exploratory conditions, a motor representation capturing the structure of the external space should emerge as a byproduct of learning to predict future sensory experiences. We propose a simple sensorimotor predictive scheme, apply it to different agents and types of exploration, and evaluate the pertinence of these hypotheses. We show that a naive agent can capture the topology and metric regularity of its sensor’s position in an egocentric spatial frame without any a priori knowledge, nor extraneous supervision.
Alban Laflaquière, Michaël Garcia Ortiz
NeurIPS1
2018 Online Learning of Body Orientation Control on a Humanoid Robot Using Finite Element Goal Babbling
abstract
How can high dimensional robots learn general sets of skills from experience in the real world? Many previous approaches focus on maximizing a single utility function and require large datasets of experience to do this, something that is not possible to collect outside of simulation as every data point is expensive both in time and in a potential wear down of the robot. This paper addresses this question using a newly developed framework called Finite Element Goal Babbling (FEGB). FEGB is an online learning method that aims at providing general control over some measurable feature, in contrast to optimizing it to some given utility function. It generalizes standard goal babbling by breaking down the full learning problem into local sub-problems, and combining it with a planner that learns how to navigate between these subproblems. We test FEGB using a real humanoid robot Nao, and find that it could quickly learn to robustly control its body orientation. After only 20-30 minutes of training, the robot could freely move into any body orientation between lying on either side and on its back. Rapid learning of body orientation control in high dimensional real robots is largely an unexplored field of robotics, and although many challenges remain, FEGB shows a feasible approach to the problem.
Pontus Loviken, Nikolas Hemion, Alban Laflaquière, Michael Spranger, Angelo Cangelosi
IROS3
2018 Discovering space - Grounding spatial topology and metric regularity in a naive agent's sensorimotor experience
Alban Laflaquière, J. Kevin O'Regan, Bruno Gas, Alexander V. Terekhov
Neural Networks1
2017 Grounding the experience of a visual field through sensorimotor contingencies
Alban Laflaquière
Neurocomputing1
2016 Grounding the experience of a visual field through sensorimotor contingencies
Alban Laflaquière, Michaël Garcia Ortiz, Ahmed Faraz Khan
ESANN1
2013 Learning an internal representation of the end-effector configuration space
abstract
Current machine learning techniques proposed to automatically discover a robot's kinematics usually rely on a priori information about the robot's structure, sensor properties or end-effector position. This paper proposes a method to estimate a certain aspect of the forward kinematics model with no such information. An internal representation of the end-effector configuration is generated from unstructured proprioceptive and exteroceptive data flow under very limited assumptions. A mapping from the proprioceptive space to this representational space can then be used to control the robot.
Alban Laflaquière, Alexander V. Terekhov, Bruno Gas, J. Kevin O'Regan
IROS1
2012 A non-linear approach to space dimension perception by a naive agent
abstract
Developmental Robotics offers a new approach to numerous AI features that are often taken as granted. Traditionally, perception is supposed to be an inherent capacity of the agent. Moreover, it largely relies on models built by the system's designer. A new approach is to consider perception as an experimentally acquired ability that is learned exclusively through the analysis of the agent's sensorimotor flow. Previous works, based on H.Poincaré's intuitions and the sensorimotor contingencies theory, allow a simulated agent to extract the dimension of geometrical space in which it is immersed without any a priori knowledge. Those results are limited to infinitesimal movement's amplitude of the system. In this paper, a non-linear dimension estimation method is proposed to push back this limitation.
Alban Laflaquière, Sylvain Argentieri, Olivia Breysse, Stéphane Genet, Bruno Gas
IROS1