EDBT 2026 Demo / reviewers in the wild / expert
Stéphane Doncieux
dblp:28/1564
· DBLP profile ↗
36ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0003-1541-054XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 5 first-author · 12 since 2021Systems, architecture and hardware · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Extract-QD Framework: A Generic Approach for Quality-Diversity in Noisy, Stochastic or Uncertain DomainsabstractQuality-Diversity (QD) has demonstrated potential in discovering collections of diverse solutions to optimisation problems. Originally designed for deterministic environments, QD has been extended to noisy, stochastic, or uncertain domains through various Uncertain-QD (UQD) methods. However, the large number of UQD methods, each with unique constraints, makes selecting the most suitable one challenging. To remedy this situation, we present two contributions: first, the Extract-QD Framework (EQD Framework), and second, Extract-MAP-Elites (EME), a new method derived from it. The EQD Framework unifies existing approaches within a modular view, and facilitates developing novel methods by interchanging modules. We use it to derive EME, a novel method that consistently outperforms or matches the best existing methods on standard benchmarks, while previous methods show varying performance. In a second experiment, we show how our EQD Framework can be used to augment existing QD algorithms and, in particular, the well-established Policy-Gradient-Assisted-ME method, and demonstrate improved performance in uncertain domains at no additional evaluation cost. For any new uncertain task, our contributions now provide EME as a reliable "first guess" method, and the EQD Framework as a tool for developing task-specific approaches. Together, these contributions aim to lower the cost of adopting UQD insights in QD applications. Manon Flageat, Johann Huber, François Hélénon, Stéphane Doncieux, Antoine Cully |
GECCO | 4 |
| 2025 | Qdgset: a Large Scale Grasping Dataset Generated With Quality-DiversityabstractRecent advances in AI have led to significant results in robotic learning, but skills like grasping remain partially solved. Many recent works exploit synthetic grasping datasets to learn to grasp unknown objects. However, those datasets were generated using simple grasp sampling methods using priors. Recently, Quality-Diversity (QD) algorithms have been proven to make grasp sampling significantly more efficient. In this work, we extend QDG-6DoF, a QD framework for generating object-centric grasps, to scale up the production of synthetic grasping datasets. We propose a data augmentation method that combines the transformation of object meshes with transfer learning from previous grasping repertoires. The conducted experiments show that this approach reduces the number of required evaluations per discovered robust grasp by up to 20 %. We used this approach to generate QDGset, a dataset of 6 DoF grasp poses that contains about 3.5 and 4.5 times more grasps and objects, respectively, than the previous state-of-the-art. Our method allows anyone to easily generate data, eventually contributing to a large-scale collaborative dataset of synthetic grasps. Johann Huber, François Hélénon, Mathilde Kappel, Ignacio de Loyola Páez-Ubieta, Santiago T. Puente Méndez, Pablo Gil, Faïz Ben Amar, Stéphane Doncieux |
ICRA | 8 |
| 2025 | Learning to Explore when Mistakes are Not Allowed
Charly Pecqueux-Guézénec, Stéphane Doncieux, Nicolas Perrin-Gilbert |
AAMAS | 2 |
| 2025 | Task-Aware Robotic Grasping by evaluating Quality Diversity Solutions through Foundation ModelsabstractTask-aware robotic grasping is a challenging problem that requires the integration of semantic understanding and geometric reasoning. This paper proposes a novel framework that leverages Large Language Models (LLMs) and Quality Diversity (QD) algorithms to enable zero-shot task-conditioned grasp synthesis. The framework segments objects into meaningful subparts and labels each subpart semantically, creating structured representations that can be used to prompt an LLM. By coupling semantic and geometric representations of an object’s structure, the LLM’s knowledge about tasks and which parts to grasp can be applied in the physical world. The QD-generated grasp archive provides a diverse set of grasps, allowing us to select the most suitable grasp based on the task. We evaluated the proposed method on a subset of the YCB dataset with a Franka Emika robot. A consolidated ground truth for task-specific grasp regions is established through a survey. Our work achieves a weighted intersection over union (IoU) of 73.6% in predicting task-conditioned grasp regions in 65 task-object combinations. An end-to-end validation study on a smaller subset further confirms the effectiveness of our approach, with 88% of responses favoring the task-aware grasp over the control group. A binomial test shows that participants significantly prefer the task-aware grasp. Aurel Appius, Émiland Garrabé, François Hélénon, Mahdi Khoramshahi, Mohamed Chetouani, Stéphane Doncieux |
IROS | 6 |
| 2025 | Enhancing Robustness in Language-Driven Robotics: A Modular Approach to Failure ReductionabstractRecent advances in large language models (LLMs) have led to significant progress in robotics, enabling embodied agents to understand and execute open-ended tasks. However, existing LLM-based approaches face limitations in grounding their outputs within the physical environment and aligning with the capabilities of the robot. While fine-tuning is an attractive approach to addressing these issues, the required data can be expensive to collect, especially when using very large language models. Smaller language models, while more computationally efficient, are less robust in task planning and execution, leading to a difficult trade-off between performance and tractability. In this paper, we present a novel, modular architecture designed to enhance the robustness of locally-executable LLMs in the context of robotics by addressing these grounding and alignment issues. We formalize the task planning problem within a goal-conditioned POMDP framework, identify key failure modes in LLM-driven planning, and propose targeted design principles to mitigate these issues. Our architecture introduces an "expected outcomes" module to prevent mischaracterization of subgoals and a feedback mechanism to enable real-time error recovery. Experimental results, both in simulation and on physical robots, demonstrate that our approach leads to significant improvements in success rates for pick-and-place and manipulation tasks, surpassing baselines using larger models. Through hardware experiments, we also demonstrate how our architecture can be run efficiently and locally. This work highlights the potential of smaller, locally-executable LLMs in robotics and provides a scalable, efficient solution for robust task execution and data collection.1 Émiland Garrabé, Pierre Teixeira, Mahdi Khoramshahi, Stéphane Doncieux |
IROS | 4 |
| 2024 | Domain Randomization for Sim2real Transfer of Automatically Generated Grasping DatasetsabstractRobotic grasping refers to making a robotic system pick an object by applying forces and torques on its surface. Many recent studies use data-driven approaches to address grasping, but the sparse reward nature of this task made the learning process challenging to bootstrap. To avoid constraining the operational space, an increasing number of works propose grasping datasets to learn from. But most of them are limited to simulations. The present paper investigates how automatically generated grasps can be exploited in the real world. More than 7000 reach-and-grasp trajectories have been generated with Quality-Diversity (QD) methods on 3 different arms and grippers, including parallel fingers and a dexterous hand, and tested in the real world. Conducted analysis on the collected measure shows correlations between several Domain Randomization-based quality criteria and sim-to-real transferability. Key challenges regarding the reality gap for grasping have been identified, stressing matters on which researchers on grasping should focus in the future. A QD approach has finally been proposed for making grasps more robust to domain randomization, resulting in a transfer ratio of 84% on the Franka Research 3 arm. Johann Huber, François Hélénon, Hippolyte Watrelot, Faïz Ben Amar, Stéphane Doncieux |
ICRA | 5 |
| 2024 | Speeding up 6-DoF Grasp Sampling with Quality-DiversityabstractRecent advances in AI have led to significant results in robotic learning, including natural language-conditioned planning and efficient optimization of controllers using generative models. However, the interaction data remains the bottleneck for generalization. Getting data for grasping is a critical challenge, as this skill is required to complete many manipulation tasks. Quality-Diversity (QD) algorithms optimize a set of solutions to get diverse, high-performing solutions to a given problem. This paper investigates how QD can be combined with priors to speed up the generation of diverse grasps poses in simulation compared to standard 6-DoF grasp sampling schemes. Experiments conducted on 4 grippers with 2-to-5 fingers on standard objects show that QD outperforms commonly used methods by a large margin. Further experiments show that QD optimization automatically finds some efficient priors that are usually hard coded. The deployment of generated grasps on a 2-finger gripper and an Allegro hand shows that the diversity produced maintains sim-to-real transferability. We believe these results to be a significant step toward the generation of large datasets that can lead to robust and generalizing robotic grasping policies. Johann Huber, François Hélénon, Mathilde Kappel, Elie Chelly, Mahdi Khoramshahi, Faïz Ben Amar, Stéphane Doncieux |
IROS | 7 |
| 2024 | Discovering and Exploiting Sparse Rewards in a Learned Behavior SpaceabstractLearning optimal policies in sparse rewards settings is difficult as the learning agent has little to no feedback on the quality of its actions. In these situations, a good strategy is to focus on exploration, hopefully leading to the discovery of a reward signal to improve on. A learning algorithm capable of dealing with this kind of setting has to be able to (1) explore possible agent behaviors and (2) exploit any possible discovered reward. Exploration algorithms have been proposed that require the definition of a low-dimension behavior space, in which the behavior generated by the agent's policy can be represented. The need to design a priori this space such that it is worth exploring is a major limitation of these algorithms. In this work, we introduce STAX, an algorithm designed to learn a behavior space on-the-fly and to explore it while optimizing any reward discovered (see Figure 1). It does so by separating the exploration and learning of the behavior space from the exploitation of the reward through an alternating two-step process. In the first step, STAX builds a repertoire of diverse policies while learning a low-dimensional representation of the high-dimensional observations generated during the policies evaluation. In the exploitation step, emitters optimize the performance of the discovered rewarding solutions. Experiments conducted on three different sparse reward environments show that STAX performs comparably to existing baselines while requiring much less prior information about the task as it autonomously builds the behavior space it explores. Giuseppe Paolo, Miranda Coninx, Alban Laflaquière, Stéphane Doncieux |
Evol. Comput. | 4 |
| 2024 | Adaptive Asynchronous Control Using Meta-Learned Neural Ordinary Differential EquationsabstractModel-based reinforcement learning and control have demonstrated great potential in various sequential decision making problem domains, including in robotics settings. However, real-world robotics systems often present challenges that limit the applicability of those methods. In particular, we note two problems that jointly happen in many industrial systems: first, irregular/asynchronous observations and actions and, second, dramatic changes in environment dynamics from an episode to another (e.g.$,$varying payload inertial properties). We propose a general framework that overcomes those difficulties by meta-learning adaptive dynamics models for continuous-time prediction and control. The proposed approach is task-agnostic and can be adapted to new tasks in a straight-forward manner. We present evaluations in two different robot simulations and on a real industrial robot. Achkan Salehi, Steffen Rühl, Stéphane Doncieux |
IEEE Trans. Robotics | 3 |
| 2022 | Automatic Acquisition of a Repertoire of Diverse Grasping Trajectories through Behavior Shaping and Novelty SearchabstractGrasping a particular object may require a dedicated grasping movement that may also be specific to the robot end-effector. No generic and autonomous method does exist to generate these movements without making hypotheses on the robot or on the object. Learning methods could help to autonomously discover relevant grasping movements, but they face an important issue: grasping movements are so rare that a learning method based on exploration has little chance to ever observe an interesting movement, thus creating a bootstrap issue. We introduce an approach to generate diverse grasping movements in order to solve this problem. The movements are generated in simulation, for particular object positions. We test it on several simulated robots: Baxter, Pepper and a Kuka Iiwa arm. Although we show that generated movements actually work on a real Baxter robot, the aim is to use this method to create a large dataset to bootstrap deep learning methods. Aurélien Morel, Yakumo Kunimoto, Miranda Coninx, Stéphane Doncieux |
ICRA | 4 |
| 2021 | Sparse reward exploration via novelty search and emittersabstractReward-based optimization algorithms require both exploration, to find rewards, and exploitation, to maximize performance. The need for efficient exploration is even more significant in sparse reward settings, in which performance feedback is given sparingly, thus rendering it unsuitable for guiding the search process. In this work, we introduce the SparsE Reward Exploration via Novelty and Emitters (SERENE) algorithm, capable of efficiently exploring a search space, as well as optimizing rewards found in potentially disparate areas. Contrary to existing emitters-based approaches, SERENE separates the search space exploration and reward exploitation into two alternating processes. The first process performs exploration through Novelty Search, a divergent search algorithm. The second one exploits discovered reward areas through emitters, i.e. local instances of population-based optimization algorithms. A meta-scheduler allocates a global computational budget by alternating between the two processes, ensuring the discovery and efficient exploitation of disjoint reward areas. SERENE returns both a collection of diverse solutions covering the search space and a collection of high-performing solutions for each distinct reward area. We evaluate SERENE on various sparse reward environments and show it compares favorably to existing baselines. Giuseppe Paolo, Miranda Coninx, Stéphane Doncieux, Alban Laflaquière |
GECCO | 3 |
| 2021 | BR-NS: an archive-less approach to novelty searchabstractAs open-ended learning based on divergent search algorithms such as Novelty Search (NS) draws more and more attention from the research community, it is natural to expect that its application to increasingly complex real-world problems will require the exploration to operate in higher dimensional Behavior Spaces (BSs) which will not necessarily be Euclidean. Novelty Search traditionally relies on k-nearest neighbours search and an archive of previously visited behavior descriptors which are assumed to live in a Euclidean space. This is problematic because of a number of issues. On one hand, Euclidean distance and Nearest-neighbour search are known to behave differently and become less meaningful in high dimensional spaces. On the other hand, the archive has to be bounded since, memory considerations aside, the computational complexity of finding nearest neighbours in that archive grows linearithmically with its size. A sub-optimal bound can result in "cycling" in the behavior space, which inhibits the progress of the exploration. Furthermore, the performance of NS depends on a number of algorithmic choices and hyperparameters, such as the strategies to add or remove elements to the archive and the number of neighbours to use in k-nn search. In this paper, we discuss an alternative approach to novelty estimation, dubbed Behavior Recognition based Novelty Search (BR-NS), which does not require an archive, makes no assumption on the metrics that can be defined in the behavior space and does not rely on nearest neighbours search. We conduct experiments to gain insight into its feasibility and dynamics as well as potential advantages over archive-based NS in terms of time complexity. Achkan Salehi, Miranda Coninx, Stéphane Doncieux |
GECCO | 3 |
| 2021 | Selection-Expansion: A Unifying Framework for Motion-Planning and Diversity Search Algorithms
Alexandre Chenu, Nicolas Perrin-Gilbert, Stéphane Doncieux, Olivier Sigaud |
ICANN (4) | 3 |
| 2020 | Novelty search makes evolvability inevitableabstractEvolvability is an important feature that impacts the ability of evolutionary processes to find interesting novel solutions and to deal with changing conditions of the problem to solve. The estimation of evolvability is not straight-forward and is generally too expensive to be directly used as selective pressure in the evolutionary process. Indirectly promoting evolvability as a side effect of other easier and faster to compute selection pressures would thus be advantageous. In an unbounded behavior space, it has already been shown that evolvable individuals naturally appear and tend to be selected as they are more likely to invade empty behavior niches. Evolvability is thus a natural byproduct of the search in this context. However, practical agents and environments often impose limits on the reachable behavior space. How do these boundaries impact evolvability? In this context, can evolvability still be promoted without explicitly rewarding it? We show that Novelty Search implicitly creates a pressure for high evolvability even in bounded behavior spaces, and explore the reasons for such a behavior. More precisely we show that, throughout the search, the dynamic evaluation of novelty rewards individuals which are very mobile in the behavior space, which in turn promotes evolvability. Stéphane Doncieux, Giuseppe Paolo, Alban Laflaquière, Miranda Coninx |
GECCO | 1 |
| 2020 | Unsupervised Learning and Exploration of Reachable Outcome SpaceabstractPerforming Reinforcement Learning in sparse rewards settings, with very little prior knowledge, is a challenging problem since there is no signal to properly guide the learning process. In such situations, a good search strategy is fundamental. At the same time, not having to adapt the algorithm to every single problem is very desirable. Here we introduce TAXONS, a Task Agnostic eXploration of Outcome spaces through Novelty and Surprise algorithm. Based on a population-based divergent-search approach, it learns a set of diverse policies directly from high-dimensional observations, without any task-specific information. TAXONS builds a repertoire of policies while training an autoencoder on the high-dimensional observation of the final state of the system to build a low-dimensional outcome space. The learned outcome space, combined with the reconstruction error, is used to drive the search for new policies. Results show that TAXONS can find a diverse set of controllers, covering a good part of the ground-truth outcome space, while having no information about such space. Giuseppe Paolo, Alban Laflaquière, Miranda Coninx, Stéphane Doncieux |
ICRA | 4 |
| 2020 | The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research CommunitiesabstractEvolution provides a creative fount of complex and subtle adaptations that often surprise the scientists who discover them. However, the creativity of evolution is not limited to the natural world: Artificial organisms evolving in computational environments have also elicited surprise and wonder from the researchers studying them. The process of evolution is an algorithmic process that transcends the substrate in which it occurs. Indeed, many researchers in the field of digital evolution can provide examples of how their evolving algorithms and organisms have creatively subverted their expectations or intentions, exposed unrecognized bugs in their code, produced unexpectedly adaptations, or engaged in behaviors and outcomes, uncannily convergent with ones found in nature. Such stories routinely reveal surprise and creativity by evolution in these digital worlds, but they rarely fit into the standard scientific narrative. Instead they are often treated as mere obstacles to be overcome, rather than results that warrant study in their own right. Bugs are fixed, experiments are refocused, and one-off surprises are collapsed into a single data point. The stories themselves are traded among researchers through oral tradition, but that mode of information transmission is inefficient and prone to error and outright loss. Moreover, the fact that these stories tend to be shared only among practitioners means that many natural scientists do not realize how interesting and lifelike digital organisms are and how natural their evolution can be. To our knowledge, no collection of such anecdotes has been published before. This article is the crowd-sourced product of researchers in the fields of artificial life and evolutionary computation who have provided first-hand accounts of such cases. It thus serves as a written, fact-checked collection of scientifically important and even entertaining stories. In doing so we also present here substantial evidence that the existence and importance of evolutionary surprises extends beyond the natural world, and may indeed be a universal property of all complex evolving systems. Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Lee Altenberg, Julie Beaulieu, Peter J. Bentley, Samuel Bernard, Guillaume Beslon, David M. Bryson, Nicholas Cheney, Patryk Chrabaszcz, Antoine Cully, Stéphane Doncieux, Fred C. Dyer, Kai Olav Ellefsen, Robert Feldt, Stephan Fischer 0002, Stephanie Forrest, Antoine Frénoy, Christian Gagné 0001, Leni K. Le Goff, Laura M. Grabowski, Babak Hodjat, Frank Hutter, Laurent Keller, Carole Knibbe, Peter Krcah, Richard E. Lenski, Hod Lipson, Robert MacCurdy, Carlos Maestre, Risto Miikkulainen, Sara Mitri, David E. Moriarty, Jean-Baptiste Mouret, Anh Totti Nguyen, Charles Ofria, Marc Parizeau, David P. Parsons, Robert T. Pennock, William F. Punch, Thomas S. Ray, Marc Schoenauer, Eric Schulte, Karl Sims, Kenneth O. Stanley, François Taddei, Danesh Tarapore, Simon Thibault, Richard A. Watson, Westley Weimer, Jason Yosinski |
Artif. Life | 14 |
| 2019 | Novelty search: a theoretical perspectiveabstractNovelty Search is an exploration algorithm driven by the novelty of a behavior. The same individual evaluated at different generations has different fitness values. The corresponding fitness landscape is thus constantly changing and if, at the scale of a single generation, the metaphor of a fitness landscape with peaks and valleys still holds, this is not the case anymore at the scale of the whole evolutionary process. How does this kind of algorithms behave? Is it possible to define a model that would help understand how it works? This understanding is critical to analyse existing Novelty Search variants and design new and potentially more efficient ones. We assert that Novelty Search asymptotically behaves like a uniform random search process in the behavior space. This is an interesting feature, as it is not possible to directly sample in this space: the algorithm has a direct access to the genotype space only, whose relationship to the behavior space is complex. We describe the model and check its consistency on a classical Novelty Search experiment. We also show that it sheds a new light on results of the literature and suggests future research work. Stéphane Doncieux, Alban Laflaquière, Miranda Coninx |
GECCO | 1 |
| 2013 | Behavioral diversity with multiple behavioral distancesabstractRecent results in evolutionary robotics show that explicitly encouraging the behavioral diversity of candidate solutions drastically improves the convergence of many experiments. The performance of this technique depends, however, on the choice of a behavioral similarity measure (BSM). Here we propose that the experimenter does not actually need to choose: provided that several similarity measures are conceivable, using them all could lead to better results than choosing a single one. Values computed by several BSM can be averaged, which is computationally expensive because it requires the computation of all the BSM at each generation, or randomly switched at a user-chosen frequency, which is a cheaper alternative. We compare these two approaches in two experimental setups - a ball collecting task and hexapod locomotion - with five different BSMs. Results show that (1) using several BSM in a single run increases the performance while avoiding the need to choose the most appropriate BSM and (2) switching between BSMs leads to better results than taking the mean behavioral diversity, while requiring less computational power. Stéphane Doncieux, Jean-Baptiste Mouret |
IEEE Congress on Evolutionary Computation | 1 |
| 2013 | The Transferability Approach: Crossing the Reality Gap in Evolutionary RoboticsabstractThe reality gap, which often makes controllers evolved in simulation inefficient once transferred onto the physical robot, remains a critical issue in evolutionary robotics (ER). We hypothesize that this gap highlights a conflict between the efficiency of the solutions in simulation and their transferability from simulation to reality: the most efficient solutions in simulation often exploit badly modeled phenomena to achieve high fitness values with unrealistic behaviors. This hypothesis leads to the transferability approach, a multiobjective formulation of ER in which two main objectives are optimized via a Pareto-based multiobjective evolutionary algorithm: 1) the fitness; and 2) the transferability, estimated by a simulation-to-reality (STR) disparity measure. To evaluate this second objective, a surrogate model of the exact STR disparity is built during the optimization. This transferability approach has been compared to two reality-based optimization methods, a noise-based approach inspired from Jakobi's minimal simulation methodology and a local search approach. It has been validated on two robotic applications: 1) a navigation task with an e-puck robot; and 2) a walking task with a 8-DOF quadrupedal robot. For both experimental setups, our approach successfully finds efficient and well-transferable controllers only with about ten experiments on the physical robot. Sylvain Koos, Jean-Baptiste Mouret, Stéphane Doncieux |
IEEE Trans. Evol. Comput. | 3 |
| 2012 | With a little help from selection pressures: evolution of memory in robot controllersabstractEvolutionary robotics (ER) have successfully built robot controllers presenting a reactive behavior. However, the evolution of cognitive controllers is still a challenge. We hypothesize here that a fitness function which rewards the fulfillment of a task requiring cognitive abilities does not necessarily reward the stepping stones that lead to cognitive controllers. In other words, our hypothesis is that evolving cognitive abilities is a deceptive problem, and that the selective pressures driving the evolutionary search are of critical importance. This paper presents some experiments to confirm this hypothesis and addresses this selective pressure problem by introducing a new helper-objective that rewards controllers with a memory. This is potentially useful for the design of controllers in which an internal representation of some data is required to solve a task. It does not assume how the memory is stored in the controller, therefore reducing the bias towards a particular solution. The new objective is tested in a multi-objective scheme on a T-maze ER task — a task involving both navigation and working memory. The efficiency of the helperobjective is studied, as well as its effects on the overall performance and generalization ability of the controller. Charles Ollion, Tony Pinville, Stéphane Doncieux |
ALIFE | 3 |
| 2012 | Encouraging Behavioral Diversity in Evolutionary Robotics: An Empirical StudyabstractEvolutionary robotics (ER) aims at automatically designing robots or controllers of robots without having to describe their inner workings. To reach this goal, ER researchers primarily employ phenotypes that can lead to an infinite number of robot behaviors and fitness functions that only reward the achievement of the task-and not how to achieve it. These choices make ER particularly prone to premature convergence. To tackle this problem, several papers recently proposed to explicitly encourage the diversity of the robot behaviors, rather than the diversity of the genotypes as in classic evolutionary optimization. Such an approach avoids the need to compute distances between structures and the pitfalls of the noninjectivity of the phenotype/behavior relation; however, it also introduces new questions: how to compare behavior? should this comparison be task specific? and what is the best way to encourage diversity in this context? In this paper, we review the main published approaches to behavioral diversity and benchmark them in a common framework. We compare each approach on three different tasks and two different genotypes. The results show that fostering behavioral diversity substantially improves the evolutionary process in the investigated experiments, regardless of genotype or task. Among the benchmarked approaches, multi-objective methods were the most efficient and the generic, Hamming-based, behavioral distance was at least as efficient as task specific behavioral metrics. Jean-Baptiste Mouret, Stéphane Doncieux |
Evol. Comput. | 2 |
| 2011 | Why and how to measure exploration in behavioral spaceabstractExploration and exploitation are two complementary aspects of Evolutionary Algorithms. Exploration, in particular, is promoted by specific diversity keeping mechanisms generally relying on the genotype or the fitness value. Recent works suggest that, in the case of Evolutionary Robotics or more generally behavioral system evolution, promoting exploration directly in the behavioral space is of critical importance. In this work an exploration indicator is proposed, based on the sparseness of the population in the behavioral space. This exploration measure is used on two challenging neuro-evolution experiments and validated by showing the dependence of the fitness at the end of the run on the exploration measure during the very first generations. Such a prediction ability could be used to design parameter settings algorithms or selection algorithms dedicated to the evolution of behavioral systems. Several other potential uses of this measure are also proposed and discussed. Charles Ollion, Stéphane Doncieux |
GECCO | 2 |
| 2011 | How to promote generalisation in evolutionary robotics: the ProGAb approachabstractIn Evolutionary Robotics (ER), controllers are assessed in a single or a few environments. As a consequence, good performances in new different contexts are not guaranteed. While a lot of ER works deal with robustness, i.e. the ability to perform well on new contexts close to the ones used for evaluation, no current approach is able to promote broader generalisation abilities without any assumption on the new contexts. In this paper, we introduce the ProGAb approach, which is based on the standard three data sets methodology of supervised machine learning, and compare it to state-of-the-art ER methods on two simulated robotic tasks: a navigation task in a T-maze and a more complex ball-collecting task in an arena. In both applications, the ProGAb approach: (1) produced controllers with better generalisation abilities than the other methods; (2) needed two to three times fewer evaluations to discover such solutions. Tony Pinville, Sylvain Koos, Jean-Baptiste Mouret, Stéphane Doncieux |
GECCO | 4 |
| 2010 | Behavioral diversity measures for Evolutionary RoboticsabstractIn Evolutionary Robotics (ER), explicitly rewarding for behavioral diversity recently revealed to generate efficient results without recourse to complex fitness functions. The principle of such approaches is to explicitly encourage diversity in the robot behavior space instead of in the space of genotypes (the space explored by the evolutionary algorithm) or the space of phenotypes (the space of robot controllers and morphologies). To implement such approaches, a similarity between behaviors needs to be evaluated but, up to now, used similarity measures are problem-specific. The goal of this work is to explore generic behavioral similarity measures that only rely on sensori-motor values. With such a measure, we managed to evolve the topology and the parameters of neuro-controllers that make a simulated robot go towards a ball, take it, find a basket, put the ball into the basket, perform a half-turn, search and take another ball, put it into the basket, etc. In this experiment, two objectives were simultaneously optimized with NSGA-II: the number of collected balls and the generic behavioral diversity objective. Several generic behavioral measures are compared. To confirm the interpretation of behavioral diversity objective and in an attempt to characterize behavioral similarity measures, they are also compared to human-made behavioral similarity evaluations. They reveal to classify behaviors globally as humans did, but with no clear correlation between the closeness to human classification and the efficiency within an evolutionary run. Stéphane Doncieux, Jean-Baptiste Mouret |
IEEE Congress on Evolutionary Computation | 1 |
| 2010 | Sferesv2: Evolvin' in the multi-core worldabstractThis paper introduces and benchmarks Sferesv2, a C++ framework designed to help researchers in evolutionary computation to make their code run as fast as possible on a multi-core computer. It is based on three main concepts: (1) including multi-core optimizations from the start of the design process; (2) providing state-of-the art implementations of well-selected current evolutionary algorithms (EA), and especially multiobjective EAs; (3) being based on modern (template-based) C++ techniques to be both abstract and efficient. Benchmark results show that when a single core is used, running time of classic EAs included in Sferesv2(NSGA-2 and CMA-ES) are of the same order of magnitude than specialized C code. When n cores are used, typical speed-ups range from 0.75n to 0.9n; however, parallelization efficiency critically depends on the time to evaluate the fitness function. Jean-Baptiste Mouret, Stéphane Doncieux |
IEEE Congress on Evolutionary Computation | 2 |
| 2010 | Crossing the reality gap in evolutionary robotics by promoting transferable controllersabstractThe reality gap, that often makes controllers evolved in simulation inefficient once transferred onto the real system, remains a critical issue in Evolutionary Robotics (ER); it prevents ER application to real-world problems. We hypothesize that this gap mainly stems from a conflict between the efficiency of the solutions in simulation and their transferability from simulation to reality: best solutions in simulation often rely on bad simulated phenomena (e.g. the most dynamic ones). This hypothesis leads to a multi-objective formulation of ER in which two main objectives are optimized via a Pareto-based Multi-Objective Evolutionary Algorithm: (1) the fitness and (2) the transferability. To evaluate this second objective, a simulation-to-reality disparity value is approximated for each controller. The proposed method is applied to the evolution of walking controllers for a real 8-DOF quadrupedal robot. It successfully finds efficient and well-transferable controllers with only a few experiments in reality. Sylvain Koos, Jean-Baptiste Mouret, Stéphane Doncieux |
GECCO | 3 |
| 2010 | Importing the computational neuroscience toolbox into neuro-evolution-application to basal gangliaabstractNeuro-evolution and computational neuroscience are two scientific domains that produce surprisingly different artificial neural networks. Inspired by the "toolbox" used by neuroscientists to create their models, this paper argues two main points: (1) neural maps (spatially-organized identical neurons) should be the building blocks to evolve neural networks able to perform cognitive functions and (2) well-identified modules of the brain for which there exists computational neuroscience models provide well-defined benchmarks for neuro-evolution. Jean-Baptiste Mouret, Stéphane Doncieux, Benoît Girard 0001 |
GECCO | 2 |
| 2009 | Automatic system identification based on coevolution of models and testsabstractIn evolutionary robotics, controllers are often designed in simulation, then transferred onto the real system. Nevertheless, when no accurate model is available, controller transfer from simulation to reality means potential performance loss. It is the reality gap problem. Unmanned aerial vehicles are typical systems where it may arise. Their locomotion dynamics may be hard to model because of a limited knowledge about the underlying physics. Moreover, a batch identification approach is difficult to use due to costly and time consuming experiments. An automatic identification method is then needed that builds a relevant local model of the system concerning a target issue. This paper deals with such an approach that is based on coevolution of models and tests. It aims at improving both modeling and control of a given system with a limited number of manipulations carried out on it. Experiments conducted with a simulated quadrotor helicopter show promising initial results about test learning and control improvement. Sylvain Koos, Jean-Baptiste Mouret, Stéphane Doncieux |
IEEE Congress on Evolutionary Computation | 3 |
| 2009 | Overcoming the bootstrap problem in evolutionary robotics using behavioral diversityabstractThe bootstrap problem is often recognized as one of the main challenges of evolutionary robotics: if all individuals from the first randomly generated population perform equally poorly, the evolutionary process won't generate any interesting solution. To overcome this lack of fitness gradient, we propose to efficiently explore behaviors until the evolutionary process finds an individual with a non-minimal fitness. To that aim, we introduce an original diversity-preservation mechanism, called behavioral diversity, that relies on a distance between behaviors (instead of genotypes or phenotypes) and multi-objective evolutionary optimization. This approach has been successfully tested and compared to a recently published incremental evolution method (multi-subgoal evolution) on the evolution of a neuro-controller for a light-seeking mobile robot. Results obtained with these two approaches are qualitatively similar although the introduced one is less directed than multi-subgoal evolution. Jean-Baptiste Mouret, Stéphane Doncieux |
IEEE Congress on Evolutionary Computation | 2 |
| 2009 | Evolving modular neural-networks through exaptationabstractDespite their success as optimization methods, evolutionary algorithms face many difficulties to design artifacts with complex structures. According to paleontologists, living organisms evolved by opportunistically co-opting characters adapted to a function to solve new problems, a phenomenon called exaptation. In this paper, we draw the hypotheses (1) that exaptation requires the presence of multiple selection pressures, (2) that Pareto-based multi-objective evolutionary algorithms (MOEA) can create such pressures and (3) that the modularity of the genotype is a key to enable exaptation. To explore these hypotheses, we designed an evolutionary process to find the structure and the parameters of neural networks to compute a Boolean function with a modular structure. We then analyzed the role of each component using a Shapley value analysis. Our results show that: (1) the proposed method is efficient to evolve neural networks to solve this task; (2) genotypic modules and multiple selections gradients needed to be aligned to converge faster than the control experiments. This prominent role of multiple selection pressures contradicts the basic assumption that underlies most published modular methods for the evolution of neural networks, in which only the modularity of the genotype is considered. Jean-Baptiste Mouret, Stéphane Doncieux |
IEEE Congress on Evolutionary Computation | 2 |
| 2009 | Single step evolution of robot controllers for sequential tasksabstractThe generation of robot controllers for a task requiring a sequence of elementary behaviors is still a challenge. If these behaviors are known, intermediate steps can be given to help bootstrap the search, thus leading to task decomposition or incremental approaches. The goal of this paper is to propose an alternative, within which such behaviors do not need to be known. The proposed approach relies on a classical multi-objective evolutionary algorithm and consists in designing objectives dedicated to the enhancement of evolutionary search abilities. These objectives are to be used in addition to performance objectives rewarding the efficiency, robustness, or whatever aspect a robot designer might be interested in. Two different kinds of objectives are proposed, tested and compared on a ball collecting problem. Both rely on states that can be directly extracted from the sensors and are completely independent from the genotype and phenotype. They show promising results, even with a simple direct neural network encoding. Stéphane Doncieux, Jean-Baptiste Mouret |
GECCO | 1 |
| 2009 | Using behavioral exploration objectives to solve deceptive problems in neuro-evolutionabstractEncouraging exploration, typically by preserving the diversity within the population, is one of the most common method to improve the behavior of evolutionary algorithms with deceptive fitness functions. Most of the published approaches to stimulate exploration rely on a distance between genotypes or phenotypes; however, such distances are difficult to compute when evolving neural networks due to (1) the algorithmic complexity of graph similarity measures, (2) the competing conventions problem and (3) the complexity of most neural-network encodings. Jean-Baptiste Mouret, Stéphane Doncieux |
GECCO | 2 |
| 2009 | Visual topological SLAM and global localizationabstractVisual localization and mapping for mobile robots has been achieved with a large variety of methods. Among them, topological navigation using vision has the advantage of offering a scalable representation, and of relying on a common and affordable sensor. In previous work, we developed such an incremental and real-time topological mapping and localization solution, without using any metrical information, and by relying on a Bayesian visual loop-closure detection algorithm. In this paper, we propose an extension of this work by integrating metrical information from robot odometry in the topological map, so as to obtain a globally consistent environment model. Also, we demonstrate the performance of our system on the global localization task, where the robot has to determine its position in a map acquired beforehand. Adrien Angeli, Stéphane Doncieux, Jean-Arcady Meyer, David Filliat |
ICRA | 2 |
| 2008 | Real-time visual loop-closure detectionabstractIn robotic applications of visual simultaneous localization and mapping, loop-closure detection and global localization are two issues that require the capacity to recognize a previously visited place from current camera measurements. We present an online method that makes it possible to detect when an image comes from an already perceived scene using local shape information. Our approach extends the bag of visual words method used in image recognition to incremental conditions and relies on Bayesian filtering to estimate loop-closure probability. We demonstrate the efficiency of our solution by real-time loop-closure detection under strong perceptual aliasing conditions in an indoor image sequence taken with a handheld camera. Adrien Angeli, Stéphane Doncieux, Jean-Arcady Meyer, David Filliat |
ICRA | 2 |
| 2008 | Incremental vision-based topological SLAMabstractIn robotics, appearance-based topological map building consists in infering the topology of the environment explored by a robot from its sensor measurements. In this paper, we propose a vision-based framework that considers this data association problem from a loop-closure detection perspective in order to correctly assign each measurement to its location. Our approach relies on the visual bag of words paradigm to represent the images and on a discrete Bayes filter to compute the probability of loop-closure. We demonstrate the efficiency of our solution by incremental and real-time consistent map building in an indoor environment and under strong perceptual aliasing conditions using a single monocular wide-angle camera. Adrien Angeli, Stéphane Doncieux, Jean-Arcady Meyer, David Filliat |
IROS | 2 |
| 2008 | Fast and Incremental Method for Loop-Closure Detection Using Bags of Visual WordsabstractIn robotic applications of visual simultaneous localization and mapping techniques, loop-closure detection and global localization are two issues that require the capacity to recognize a previously visited place from current camera measurements. We present an online method that makes it possible to detect when an image comes from an already perceived scene using local shape and color information. Our approach extends the bag-of-words method used in image classification to incremental conditions and relies on Bayesian filtering to estimate loop-closure probability. We demonstrate the efficiency of our solution by real-time loop-closure detection under strong perceptual aliasing conditions in both indoor and outdoor image sequences taken with a handheld camera. Adrien Angeli, David Filliat, Stéphane Doncieux, Jean-Arcady Meyer |
IEEE Trans. Robotics | 3 |