EDBT 2026 Demo / reviewers in the wild / expert
Alban Laflaquière
dblp:123/6900
· DBLP profile ↗
13ranked-venue papers
6as first author
2since 2021 · last 2024
0000-0002-3842-0858ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 2 since 2021Systems, architecture and hardware · 5 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 60% Robot navigation and mapping · 17% Representation and self-supervised learning · 17% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
exploration |
0.4 | 1 | 2020 | Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020 |
Machine learning › Reinforcement learning › exploration
novelty-based exploration |
0.4 | 1 | 2020 | Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020 |
Machine learning › Reinforcement learning
policy learning |
0.4 | 1 | 2020 | Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020 |
Robotics › Robot navigation and mapping
spatial representation |
0.4 | 1 | 2019 | Unsupervised Emergence of Egocentric Spatial Structure from Sensorimotor Prediction · NeurIPS 2019 |
Robotics › Motion planning and robot control
robot learning |
0.1 | 1 | 2019 | Unsupervised Emergence of Egocentric Spatial Structure from Sensorimotor Prediction · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
population-based divergent search · 0.4autoencoder · 0.4sensorimotor prediction · 0.4predictive learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Discovering and Exploiting Sparse Rewards in a Learned Behavior SpaceabstractLearning optimal policies in sparse rewards settings is difficult as the learning agent has little to no feedback on the quality of its actions. In these situations, a good strategy is to focus on exploration, hopefully leading to the discovery of a reward signal to improve on. A learning algorithm capable of dealing with this kind of setting has to be able to (1) explore possible agent behaviors and (2) exploit any possible discovered reward. Exploration algorithms have been proposed that require the definition of a low-dimension behavior space, in which the behavior generated by the agent's policy can be represented. The need to design a priori this space such that it is worth exploring is a major limitation of these algorithms. In this work, we introduce STAX, an algorithm designed to learn a behavior space on-the-fly and to explore it while optimizing any reward discovered (see Figure 1). It does so by separating the exploration and learning of the behavior space from the exploitation of the reward through an alternating two-step process. In the first step, STAX builds a repertoire of diverse policies while learning a low-dimensional representation of the high-dimensional observations generated during the policies evaluation. In the exploitation step, emitters optimize the performance of the discovered rewarding solutions. Experiments conducted on three different sparse reward environments show that STAX performs comparably to existing baselines while requiring much less prior information about the task as it autonomously builds the behavior space it explores. Giuseppe Paolo, Miranda Coninx, Alban Laflaquière, Stéphane Doncieux |
Evol. Comput. | 3 |
| 2021 | Sparse reward exploration via novelty search and emittersabstractReward-based optimization algorithms require both exploration, to find rewards, and exploitation, to maximize performance. The need for efficient exploration is even more significant in sparse reward settings, in which performance feedback is given sparingly, thus rendering it unsuitable for guiding the search process. In this work, we introduce the SparsE Reward Exploration via Novelty and Emitters (SERENE) algorithm, capable of efficiently exploring a search space, as well as optimizing rewards found in potentially disparate areas. Contrary to existing emitters-based approaches, SERENE separates the search space exploration and reward exploitation into two alternating processes. The first process performs exploration through Novelty Search, a divergent search algorithm. The second one exploits discovered reward areas through emitters, i.e. local instances of population-based optimization algorithms. A meta-scheduler allocates a global computational budget by alternating between the two processes, ensuring the discovery and efficient exploitation of disjoint reward areas. SERENE returns both a collection of diverse solutions covering the search space and a collection of high-performing solutions for each distinct reward area. We evaluate SERENE on various sparse reward environments and show it compares favorably to existing baselines. Giuseppe Paolo, Miranda Coninx, Stéphane Doncieux, Alban Laflaquière |
GECCO | 4 |
| 2020 | Novelty search makes evolvability inevitableabstractEvolvability is an important feature that impacts the ability of evolutionary processes to find interesting novel solutions and to deal with changing conditions of the problem to solve. The estimation of evolvability is not straight-forward and is generally too expensive to be directly used as selective pressure in the evolutionary process. Indirectly promoting evolvability as a side effect of other easier and faster to compute selection pressures would thus be advantageous. In an unbounded behavior space, it has already been shown that evolvable individuals naturally appear and tend to be selected as they are more likely to invade empty behavior niches. Evolvability is thus a natural byproduct of the search in this context. However, practical agents and environments often impose limits on the reachable behavior space. How do these boundaries impact evolvability? In this context, can evolvability still be promoted without explicitly rewarding it? We show that Novelty Search implicitly creates a pressure for high evolvability even in bounded behavior spaces, and explore the reasons for such a behavior. More precisely we show that, throughout the search, the dynamic evaluation of novelty rewards individuals which are very mobile in the behavior space, which in turn promotes evolvability. Stéphane Doncieux, Giuseppe Paolo, Alban Laflaquière, Miranda Coninx |
GECCO | 3 |
| 2020 | Unsupervised Learning and Exploration of Reachable Outcome SpaceabstractPerforming Reinforcement Learning in sparse rewards settings, with very little prior knowledge, is a challenging problem since there is no signal to properly guide the learning process. In such situations, a good search strategy is fundamental. At the same time, not having to adapt the algorithm to every single problem is very desirable. Here we introduce TAXONS, a Task Agnostic eXploration of Outcome spaces through Novelty and Surprise algorithm. Based on a population-based divergent-search approach, it learns a set of diverse policies directly from high-dimensional observations, without any task-specific information. TAXONS builds a repertoire of policies while training an autoencoder on the high-dimensional observation of the final state of the system to build a low-dimensional outcome space. The learned outcome space, combined with the reconstruction error, is used to drive the search for new policies. Results show that TAXONS can find a diverse set of controllers, covering a good part of the ground-truth outcome space, while having no information about such space. Giuseppe Paolo, Alban Laflaquière, Miranda Coninx, Stéphane Doncieux |
ICRA | 2 |
| 2019 | Novelty search: a theoretical perspectiveabstractNovelty Search is an exploration algorithm driven by the novelty of a behavior. The same individual evaluated at different generations has different fitness values. The corresponding fitness landscape is thus constantly changing and if, at the scale of a single generation, the metaphor of a fitness landscape with peaks and valleys still holds, this is not the case anymore at the scale of the whole evolutionary process. How does this kind of algorithms behave? Is it possible to define a model that would help understand how it works? This understanding is critical to analyse existing Novelty Search variants and design new and potentially more efficient ones. We assert that Novelty Search asymptotically behaves like a uniform random search process in the behavior space. This is an interesting feature, as it is not possible to directly sample in this space: the algorithm has a direct access to the genotype space only, whose relationship to the behavior space is complex. We describe the model and check its consistency on a classical Novelty Search experiment. We also show that it sheds a new light on results of the literature and suggests future research work. Stéphane Doncieux, Alban Laflaquière, Miranda Coninx |
GECCO | 2 |
| 2019 | A Bi-directional Multiple Timescales LSTM Model for Grounding of Actions and VerbsabstractIn this paper we present a neural architecture to learn a bi-directional mapping between actions and language. We implement a Multiple Timescale Long Short-Term Memory (MT-LSTM) network comprised of 7 layers with different timescale factors, to connect actions to language without explicitly learning an intermediate representation. Instead, the model self-organizes such representations at the level of a slow-varying latent layer, linking action branch and language branch at the center. We train the model in a bi-directional way, learning how to produce a sentence from a certain action sequence input and, simultaneously, how to generate an action sequence given a sentence as input. Furthermore we show this model preserves some of the generalization behaviour of Multiple Timescale Recurrent Neural Networks (MTRNN) in generating sentences and actions that were not explicitly trained. We compare this model with a number of different baseline models, confirming the importance of both the bi-directional training and the multiple timescales architecture. Finally, the network was evaluated on motor actions performed by an iCub robot and their corresponding letter-based description. The results of these experiments are presented at the end of the paper. Alexandre Antunes, Alban Laflaquière, Tetsuya Ogata, Angelo Cangelosi |
IROS | 2 |
| 2019 | Unsupervised Emergence of Egocentric Spatial Structure from Sensorimotor PredictionabstractDespite its omnipresence in robotics application, the nature of spatial knowledge and the mechanisms that underlie its emergence in autonomous agents are still poorly understood. Recent theoretical works suggest that the Euclidean structure of space induces invariants in an agent’s raw sensorimotor experience. We hypothesize that capturing these invariants is beneficial for sensorimotor prediction and that, under certain exploratory conditions, a motor representation capturing the structure of the external space should emerge as a byproduct of learning to predict future sensory experiences. We propose a simple sensorimotor predictive scheme, apply it to different agents and types of exploration, and evaluate the pertinence of these hypotheses. We show that a naive agent can capture the topology and metric regularity of its sensor’s position in an egocentric spatial frame without any a priori knowledge, nor extraneous supervision. Alban Laflaquière, Michaël Garcia Ortiz |
NeurIPS | 1 |
| 2018 | Online Learning of Body Orientation Control on a Humanoid Robot Using Finite Element Goal BabblingabstractHow can high dimensional robots learn general sets of skills from experience in the real world? Many previous approaches focus on maximizing a single utility function and require large datasets of experience to do this, something that is not possible to collect outside of simulation as every data point is expensive both in time and in a potential wear down of the robot. This paper addresses this question using a newly developed framework called Finite Element Goal Babbling (FEGB). FEGB is an online learning method that aims at providing general control over some measurable feature, in contrast to optimizing it to some given utility function. It generalizes standard goal babbling by breaking down the full learning problem into local sub-problems, and combining it with a planner that learns how to navigate between these subproblems. We test FEGB using a real humanoid robot Nao, and find that it could quickly learn to robustly control its body orientation. After only 20-30 minutes of training, the robot could freely move into any body orientation between lying on either side and on its back. Rapid learning of body orientation control in high dimensional real robots is largely an unexplored field of robotics, and although many challenges remain, FEGB shows a feasible approach to the problem. Pontus Loviken, Nikolas Hemion, Alban Laflaquière, Michael Spranger, Angelo Cangelosi |
IROS | 3 |
| 2018 | Discovering space - Grounding spatial topology and metric regularity in a naive agent's sensorimotor experience
Alban Laflaquière, J. Kevin O'Regan, Bruno Gas, Alexander V. Terekhov |
Neural Networks | 1 |
| 2017 | Grounding the experience of a visual field through sensorimotor contingencies
Alban Laflaquière |
Neurocomputing | 1 |
| 2016 | Grounding the experience of a visual field through sensorimotor contingencies
Alban Laflaquière, Michaël Garcia Ortiz, Ahmed Faraz Khan |
ESANN | 1 |
| 2013 | Learning an internal representation of the end-effector configuration spaceabstractCurrent machine learning techniques proposed to automatically discover a robot's kinematics usually rely on a priori information about the robot's structure, sensor properties or end-effector position. This paper proposes a method to estimate a certain aspect of the forward kinematics model with no such information. An internal representation of the end-effector configuration is generated from unstructured proprioceptive and exteroceptive data flow under very limited assumptions. A mapping from the proprioceptive space to this representational space can then be used to control the robot. Alban Laflaquière, Alexander V. Terekhov, Bruno Gas, J. Kevin O'Regan |
IROS | 1 |
| 2012 | A non-linear approach to space dimension perception by a naive agentabstractDevelopmental Robotics offers a new approach to numerous AI features that are often taken as granted. Traditionally, perception is supposed to be an inherent capacity of the agent. Moreover, it largely relies on models built by the system's designer. A new approach is to consider perception as an experimentally acquired ability that is learned exclusively through the analysis of the agent's sensorimotor flow. Previous works, based on H.Poincaré's intuitions and the sensorimotor contingencies theory, allow a simulated agent to extract the dimension of geometrical space in which it is immersed without any a priori knowledge. Those results are limited to infinitesimal movement's amplitude of the system. In this paper, a non-linear dimension estimation method is proposed to push back this limitation. Alban Laflaquière, Sylvain Argentieri, Olivia Breysse, Stéphane Genet, Bruno Gas |
IROS | 1 |