Jean-Baptiste Mouret

dblp:26/122 · DBLP profile ↗
← Back
52ranked-venue papers
8as first author
5since 2021 · last 2025
0000-0002-2513-027XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 8 first-author · 5 since 2021Systems, architecture and hardware · 9 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Bayesian Optimization for Quality Diversity Search With Coupled Descriptor Functions
abstract
Quality Diversity (QD) algorithms such as MAP-Elites are a class of optimisation techniques that attempt to find many high performing points that all behave differently according to a user-defined behavioural metric. In this paper we propose the Bayesian Optimisation of Elites (BOP-Elites) algorithm. Designed for problems with expensive black-box objective and behaviour functions, it is able to return a QD solution-set after a relatively small number of samples. BOP-Elites models both objective and behavioural descriptors with Gaussian Process surrogate models and uses Bayesian Optimisation strategies for choosing points to evaluate in order to solve the quality-diversity problem. In addition, BOP-Elites produces high quality surrogate models which can be used after convergence to predict solutions with any behaviour in a continuous range. An empirical comparison shows that BOP-Elites significantly outperforms other state-of-the-art algorithms without the need for problem-specific parameter tuning.
Paul Kent, Adam Gaier, Jean-Baptiste Mouret, Jürgen Branke
IEEE Trans. Evol. Comput.3
2024 Parametric-Task MAP-Elites
abstract
Optimizing a set of functions simultaneously by leveraging their similarity is called multi-task optimization. Current black-box multi-task algorithms only solve a finite set of tasks, even when the tasks originate from a continuous space. In this paper, we introduce Parametric-Task MAP-Elites (PT-ME), a new black-box algorithm for continuous multi-task optimization problems. This algorithm (1) solves a new task at each iteration, effectively covering the continuous space, and (2) exploits a new variation operator based on local linear regression. The resulting dataset of solutions makes it possible to create a function that maps any task parameter to its optimal solution. We show that PT-ME outperforms all baselines, including the deep reinforcement learning algorithm PPO on two parametric-task toy problems and a robotic problem in simulation.
Timothée Anne, Jean-Baptiste Mouret
GECCO2
2023 Learning Height for Top-Down Grasps with the DIGIT Sensor
abstract
We address the problem of grasping unknown objects identified from top-down images with a parallel gripper. When no object 3D model is available, the state-of-the-art grasp generators identify the best candidate locations for planar grasps using the RGBD image. However, while they generate the Cartesian location and orientation of the gripper, the height of the grasp center is often determined by heuristics based on the highest point in the depth map, which leads to unsuccessful grasps when the objects are not thick, or have transparencies or curved shapes. In this paper, we propose to learn a regressor that predicts the best grasp height based from the image. We train this regressor with a dataset that is automatically acquired thanks to the DIGIT optical tactile sensors, which can evaluate grasp success and stability. Using our predictor, the grasping success is improved by 6% for all objects, by 16% on average on difficult objects, and by 40% for objects that are notably very difficult to grasp (e.g., transparent, curved, thin).
Thais Bernardi, Yoann Fleytoux, Jean-Baptiste Mouret, Serena Ivaldi
ICRA3
2022 Data-efficient learning of object-centric grasp preferences
abstract
Grasping made impressive progress during the last few years thanks to deep learning. However, there are many objects for which it is not possible to choose a grasp by only looking at an RGB-D image, might it be for physical reasons (e.g., a hammer with uneven mass distribution) or task constraints (e.g., food that should not be spoiled). In such situations, the preferences of experts need to be taken into account. In this paper, we introduce a data-efficient grasping pipeline (Latent Space GP Selector - LGPS) that learns grasp prefer-ences with only a few labels per object (typically 1 to 4) and generalizes to new views of this object. Our pipeline is based on learning a latent space of grasps with a dataset generated with any state-of-the-art grasp generator (e.g., Dex-Net). This latent space is then used as a low-dimensional input for a Gaussian process classifier that selects the preferred grasp among those proposed by the generator. The results show that our method outperforms both GR-ConvNet and GG-CNN (two state-of-the-art methods that are also based on labeled grasps) on the Cornell dataset, especially when only a few labels are used: only 80 labels are enough to correctly choose 80% of the grasps (885 scenes, 244 objects). Results are similar on our dataset (91 scenes, 28 objects).
Yoann Fleytoux, Anji Ma, Serena Ivaldi, Jean-Baptiste Mouret
ICRA4
2022 Automatic Tuning and Selection of Whole-Body Controllers
abstract
Designing controllers for complex robots such as humanoids is not an easy task. Often, researchers hand-tune controllers, but this is a time-consuming approach that yields a single controller which cannot generalize well to varied tasks. This work presents a method which uses the NSGA-II multi-objective optimization algorithm with various training trajectories to output a diverse Pareto set of well-functioning controller weights and gains. The best of these are shown to also work well on the real Talos robot. The learned Pareto front is then used in a Bayesian optimization (BO) algorithm both as a search space and as a source of prior information in the initial mean estimate. This combined learning approach, leveraging the two optimization methods together, finds a suitable parameter set for a new trajectory within 20 trials and outperforms both BO in the continuous parameter search space and random search along the precomputed Pareto front. The few trials required for this formulation of BO suggest that it could feasibly be applied on the physical robot using a Pareto front generated in simulation.
Evelyn D'Elia, Jean-Baptiste Mouret, Jens Kober, Serena Ivaldi
IROS2
2020 Learning a Behavioral Repertoire from Demonstrations
abstract
Imitation Learning (IL) is a machine learning approach to learn a policy from a set of demonstrations. IL can be useful to kick-start learning before applying reinforcement learning (RL) but it can also be useful on its own, e.g. to learn to imitate human players in video games. Despite the success of systems that use IL and RL, how such systems can adapt in-between game rounds is a neglected area of study but an important aspect of many strategy games. In this paper, we present a new approach called Behavioral Repertoire Imitation Learning (BRIL) that learns a repertoire of behaviors from a set of demonstrations by augmenting the state-action pairs with behavioral descriptions. The outcome of this approach is a single neural network policy conditioned on a behavior description that can be precisely modulated. We apply this approach to train a policy on 7,777 human demonstrations for the build-order planning task in StarCraft II. Dimensionality reduction is applied to construct a low-dimensional behavioral space from a high-dimensional description of the army unit composition of each human replay. The results demonstrate that the learned policy can be effectively manipulated to express distinct behaviors. Additionally, by applying the UCB1 algorithm, the policy can adapt its behavior - in-between games - to reach a performance beyond that of the traditional IL baseline approach.
Niels Justesen, Miguel González Duque, Daniel Cabarcas Jaramillo, Jean-Baptiste Mouret, Sebastian Risi
CoG4
2020 Learning behaviour-performance maps with meta-evolution
abstract
The MAP-Elites quality-diversity algorithm has been successful in robotics because it can create a behaviorally diverse set of solutions that later can be used for adaptation, for instance to unanticipated damages. In MAP-Elites, the choice of the behaviour space is essential for adaptation, the recovery of performance in unseen environments, since it defines the diversity of the solutions. Current practice is to hand-code a set of behavioural features, however, given the large space of possible behaviour-performance maps, the designer does not know a priori which behavioural features maximise a map's adaptation potential. We introduce a new meta-evolution algorithm that discovers those behavioural features that maximise future adaptations. The proposed method applies Covariance Matrix Adaptation Evolution Strategy to evolve a population of behaviour-performance maps to maximise a meta-fitness function that rewards adaptation. The method stores solutions found by MAP-Elites in a database which allows to rapidly construct new behaviour-performance maps on-the-fly. To evaluate this system, we study the gait of the RHex robot as it adapts to a range of damages sustained on its legs. When compared to MAP-Elites with user-defined behaviour spaces, we demonstrate that the meta-evolution system learns high-performing gaits with or without damages injected to the robot.
David M. Bossens, Jean-Baptiste Mouret, Danesh Tarapore
GECCO2
2020 Discovering representations for black-box optimization
abstract
The encoding of solutions in black-box optimization is a delicate, handcrafted balance between expressiveness and domain knowledge --- between exploring a wide variety of solutions, and ensuring that those solutions are useful. Our main insight is that this process can be automated by generating a dataset of high-performing solutions with a quality diversity algorithm (here, MAP-Elites), then learning a representation with a generative model (here, a Variational Autoencoder) from that dataset. Our second insight is that this representation can be used to scale quality diversity optimization to higher dimensions --- but only if we carefully mix solutions generated with the learned representation and those generated with traditional variation operators. We demonstrate these capabilities by learning an low-dimensional encoding for the inverse kinematics of a thousand joint planar arm. The results show that learned representations make it possible to solve high-dimensional problems with orders of magnitude fewer evaluations than the standard MAP-Elites, and that, once solved, the produced encoding can be used for rapid optimization of novel, but similar, tasks. The presented techniques not only scale up quality diversity algorithms to high dimensions, but show that black-box optimization encodings can be automatically learned, rather than hand designed.
Adam Gaier, Alexander Asteroth, Jean-Baptiste Mouret
GECCO3
2020 Quality diversity for multi-task optimization
abstract
Quality Diversity (QD) algorithms are a recent family of optimization algorithms that search for a large set of diverse but high-performing solutions. In some specific situations, they can solve multiple tasks at once. For instance, they can find the joint positions required for a robotic arm to reach a set of points, which can also be solved by running a classic optimizer for each target point. However, they cannot solve multiple tasks when the fitness needs to be evaluated independently for each task (e.g., optimizing policies to grasp many different objects). In this paper, we propose an extension of the MAP-Elites algorithm, called Multi-task MAP-Elites, that solves multiple tasks when the fitness function depends on the task. We evaluate it on a simulated parametrized planar arm (10-dimensional search space; 5000 tasks) and on a simulated 6-legged robot with legs of different lengths (36-dimensional search space; 2000 tasks). The results show that in both cases our algorithm outperforms the optimization of each task separately with the CMA-ES algorithm.
Jean-Baptiste Mouret, Glenn Maguire
GECCO1
2020 Fast Online Adaptation in Robotics through Meta-Learning Embeddings of Simulated Priors
abstract
Meta-learning algorithms can accelerate the model-based reinforcement learning (MBRL) algorithms by finding an initial set of parameters for the dynamical model such that the model can be trained to match the actual dynamics of the system with only a few data-points. However, in the real world, a robot might encounter any situation starting from motor failures to finding itself in a rocky terrain where the dynamics of the robot can be significantly different from one another. In this paper, first, we show that when meta-training situations (the prior situations) have such diverse dynamics, using a single set of meta-trained parameters as a starting point still requires a large number of observations from the real system to learn a useful model of the dynamics. Second, we propose an algorithm called FAMLE that mitigates this limitation by meta-training several initial starting points (i.e., initial parameters) for training the model and allows robots to select the most suitable starting point to adapt the model to the current situation with only a few gradient steps. We compare FAMLE to MBRL, MBRL with a meta-trained model with MAML, and model-free policy search algorithm PPO for various simulated and real robotic tasks, and show that FAMLE allows robots to adapt to novel damages in significantly fewer time-steps than the baselines.
Rituraj Kaushik, Timothée Anne, Jean-Baptiste Mouret
IROS3
2020 The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities
abstract
Evolution provides a creative fount of complex and subtle adaptations that often surprise the scientists who discover them. However, the creativity of evolution is not limited to the natural world: Artificial organisms evolving in computational environments have also elicited surprise and wonder from the researchers studying them. The process of evolution is an algorithmic process that transcends the substrate in which it occurs. Indeed, many researchers in the field of digital evolution can provide examples of how their evolving algorithms and organisms have creatively subverted their expectations or intentions, exposed unrecognized bugs in their code, produced unexpectedly adaptations, or engaged in behaviors and outcomes, uncannily convergent with ones found in nature. Such stories routinely reveal surprise and creativity by evolution in these digital worlds, but they rarely fit into the standard scientific narrative. Instead they are often treated as mere obstacles to be overcome, rather than results that warrant study in their own right. Bugs are fixed, experiments are refocused, and one-off surprises are collapsed into a single data point. The stories themselves are traded among researchers through oral tradition, but that mode of information transmission is inefficient and prone to error and outright loss. Moreover, the fact that these stories tend to be shared only among practitioners means that many natural scientists do not realize how interesting and lifelike digital organisms are and how natural their evolution can be. To our knowledge, no collection of such anecdotes has been published before. This article is the crowd-sourced product of researchers in the fields of artificial life and evolutionary computation who have provided first-hand accounts of such cases. It thus serves as a written, fact-checked collection of scientifically important and even entertaining stories. In doing so we also present here substantial evidence that the existence and importance of evolutionary surprises extends beyond the natural world, and may indeed be a universal property of all complex evolving systems.
Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Lee Altenberg, Julie Beaulieu, Peter J. Bentley, Samuel Bernard, Guillaume Beslon, David M. Bryson, Nicholas Cheney, Patryk Chrabaszcz, Antoine Cully, Stéphane Doncieux, Fred C. Dyer, Kai Olav Ellefsen, Robert Feldt, Stephan Fischer 0002, Stephanie Forrest, Antoine Frénoy, Christian Gagné 0001, Leni K. Le Goff, Laura M. Grabowski, Babak Hodjat, Frank Hutter, Laurent Keller, Carole Knibbe, Peter Krcah, Richard E. Lenski, Hod Lipson, Robert MacCurdy, Carlos Maestre, Risto Miikkulainen, Sara Mitri, David E. Moriarty, Jean-Baptiste Mouret, Anh Totti Nguyen, Charles Ofria, Marc Parizeau, David P. Parsons, Robert T. Pennock, William F. Punch, Thomas S. Ray, Marc Schoenauer, Eric Schulte, Karl Sims, Kenneth O. Stanley, François Taddei, Danesh Tarapore, Simon Thibault, Richard A. Watson, Westley Weimer, Jason Yosinski
Artif. Life36
2020 Robust Reinforcement Learning with Bayesian Optimisation and Quadrature
abstract
Bayesian optimisation has been successfully applied to a variety of reinforcement learning problems. However, the traditional approach for learning optimal policies in simulators does not utilise the opportunity to improve learning by adjusting certain environment variables: state features that are unobservable and randomly determined by the environment in a physical setting but are controllable in a simulator. This article considers the problem of finding a robust policy while taking into account the impact of environment variables. We present Alternating Optimisation and Quadrature (ALOQ), which uses Bayesian optimisation and Bayesian quadrature to address such settings. We also present Transferable ALOQ (TALOQ), for settings where simulator inaccuracies lead to difficulty in transferring the learnt policy to the physical system. We show that our algorithms are robust to the presence of significant rare events, which may not be observable under random sampling but play a substantial role in determining the optimal policy. Experimental results across different domains show that our algorithms learn robust policies efficiently.
Supratik Paul, Konstantinos Chatzilygeroudis, Kamil Ciosek, Jean-Baptiste Mouret, Michael A. Osborne, Shimon Whiteson
J. Mach. Learn. Res.4
2020 A Survey on Policy Search Algorithms for Learning Robot Controllers in a Handful of Trials
abstract
Most policy search (PS) algorithms require thousands of training episodes to find an effective policy, which is often infeasible with a physical robot. This survey article focuses on the extreme other end of the spectrum: how can a robot adapt with only a handful of trials (a dozen) and a few minutes? By analogy with the word “big-data,” we refer to this challenge as “micro-data reinforcement learning.” In this article, we show that a first strategy is to leverage prior knowledge on the policy structure (e.g., dynamic movement primitives), on the policy parameters (e.g., demonstrations), or on the dynamics (e.g., simulators). A second strategy is to create data-driven surrogate models of the expected reward (e.g., Bayesian optimization) or the dynamical model (e.g., model-based PS), so that the policy optimizer queries the model instead of the real system. Overall, all successful micro-data algorithms combine these two strategies by varying the kind of model and prior knowledge. The current scientific challenges essentially revolve around scaling up to complex robots, designing generic priors, and optimizing the computing time.
Konstantinos Chatzilygeroudis, Vassilis Vassiliades, Freek Stulp, Sylvain Calinon, Jean-Baptiste Mouret
IEEE Trans. Robotics5
2018 Alternating Optimisation and Quadrature for Robust Control
abstract
Bayesian optimisation has been successfully applied to a variety of reinforcement learning problems. However, the traditional approach for learning optimal policies in simulators does not utilise the opportunity to improve learning by adjusting certain environment variables: state features that are unobservable and randomly determined by the environment in a physical setting but are controllable in a simulator. This paper considers the problem of finding a robust policy while taking into account the impact of environment variables. We present Alternating Optimisation and Quadrature (ALOQ), which uses Bayesian optimisation and Bayesian quadrature to address such settings. ALOQ is robust to the presence of significant rare events, which may not be observable under random sampling, but play a substantial role in determining the optimal policy. Experimental results across different domains show that ALOQ can learn more efficiently and robustly than existing methods.
Supratik Paul, Konstantinos Chatzilygeroudis, Kamil Ciosek, Jean-Baptiste Mouret, Michael A. Osborne, Shimon Whiteson
AAAI4
2018 Data-efficient neuroevolution with kernel-based surrogate models
abstract
Surrogate-assistance approaches have long been used in computationally expensive domains to improve the data-efficiency of optimization algorithms. Neuroevolution, however, has so far resisted the application of these techniques because it requires the surrogate model to make fitness predictions based on variable topologies, instead of a vector of parameters. Our main insight is that we can sidestep this problem by using kernel-based surrogate models, which require only the definition of a distance measure between individuals. Our second insight is that the well-established Neuroevolution of Augmenting Topologies (NEAT) algorithm provides a computationally efficient distance measure between dissimilar networks in the form of "compatibility distance", initially designed to maintain topological diversity. Combining these two ideas, we introduce a surrogate-assisted neuroevolution algorithm that combines NEAT and a surrogate model built using a compatibility distance kernel. We demonstrate the data-efficiency of this new algorithm on the low dimensional cart-pole swing-up problem, as well as the higher dimensional half-cheetah running task. In both tasks the surrogate-assisted variant achieves the same or better results with several times fewer function evaluations as the original NEAT.
Adam Gaier, Alexander Asteroth, Jean-Baptiste Mouret
GECCO3
2018 Discovering the elite hypervolume by leveraging interspecies correlation
abstract
Evolution has produced an astonishing diversity of species, each filling a different niche. Algorithms like MAP-Elites mimic this divergent evolutionary process to find a set of behaviorally diverse but high-performing solutions, called the elites. Our key insight is that species in nature often share a surprisingly large part of their genome, in spite of occupying very different niches; similarly the elites are likely to be concentrated in a specific "elite hypervolume" whose shape is defined by their common features. In this paper, we first introduce the elite hypervolume concept and propose two metrics to characterize it: the genotypic spread and the genotypic similarity. We then introduce a new variation operator, called "directional variation", that exploits interspecies (or inter-elites) correlations to accelerate the MAP-Elites algorithm. We demonstrate the effectiveness of this operator in three problems (a toy function, a redundant robotic arm, and a hexapod robot).
Vassilis Vassiliades, Jean-Baptiste Mouret
GECCO2
2018 Using Parameterized Black-Box Priors to Scale Up Model-Based Policy Search for Robotics
abstract
The most data-efficient algorithms for reinforcement learning in robotics are model-based policy search algorithms, which alternate between learning a dynamical model of the robot and optimizing a policy to maximize the expected return given the model and its uncertainties. Among the few proposed approaches, the recently introduced Black-DROPS algorithm exploits a black-box optimization algorithm to achieve both high data-efficiency and good computation times when several cores are used; nevertheless, like all model-based policy search approaches, Black-DROPS does not scale to high dimensional state/action spaces. In this paper, we introduce a new model learning procedure in Black-DROPS that leverages parameterized black-box priors to (1) scale up to high-dimensional systems, and (2) be robust to large inaccuracies of the prior information. We demonstrate the effectiveness of our approach with the “pendubot” swing-up task in simulation and with a physical hexapod robot (48D state space, 18D action space) that has to walk forward as fast as possible. The results show that our new algorithm is more data-efficient than previous model-based policy search algorithms (with and without priors) and that it can allow a physical 6-legged robot to learn new gaits in only 16 to 30 seconds of interaction time.
Konstantinos Chatzilygeroudis, Jean-Baptiste Mouret
ICRA2
2018 Bayesian Optimization with Automatic Prior Selection for Data-Efficient Direct Policy Search
abstract
One of the most interesting features of Bayesian optimization for direct policy search is that it can leverage priors (e.g., from simulation or from previous tasks) to accelerate learning on a robot. In this paper, we are interested in situations for which several priors exist but we do not know in advance which one fits best the current situation. We tackle this problem by introducing a novel acquisition function, called Most Likely Expected Improvement (MLEI), that combines the likelihood of the priors and the expected improvement. We evaluate this new acquisition function on a transfer learning task for a 5-DOF planar arm and on a possibly damaged, 6-legged robot that has to learn to walk on flat ground and on stairs, with priors corresponding to different stairs and different kinds of damages. Our results show that MLEI effectively identifies and exploits the priors, even when there is no obvious match between the current situations and the priors.
Rémi Pautrat, Konstantinos Chatzilygeroudis, Jean-Baptiste Mouret
ICRA3
2018 Data-Efficient Design Exploration through Surrogate-Assisted Illumination
abstract
Design optimization techniques are often used at the beginning of the design process to explore the space of possible designs. In these domains illumination algorithms, such as MAP-Elites, are promising alternatives to classic optimization algorithms because they produce diverse, high-quality solutions in a single run, instead of only a single near-optimal solution. Unfortunately, these algorithms currently require a large number of function evaluations, limiting their applicability. In this article, we introduce a new illumination algorithm, Surrogate-Assisted Illumination (SAIL), that leverages surrogate modeling techniques to create a map of the design space according to user-defined features while minimizing the number of fitness evaluations. On a two-dimensional airfoil optimization problem, SAIL produces hundreds of diverse but high-performing designs with several orders of magnitude fewer evaluations than MAP-Elites or CMA-ES. We demonstrate that SAIL is also capable of producing maps of high-performing designs in realistic three-dimensional aerodynamic tasks with an accurate flow simulation. Data-efficient design exploration with SAIL can help designers understand what is possible, beyond what is optimal, by considering more than pure objective-based optimization.
Adam Gaier, Alexander Asteroth, Jean-Baptiste Mouret
Evol. Comput.3
2018 Using Centroidal Voronoi Tessellations to Scale Up the Multidimensional Archive of Phenotypic Elites Algorithm
abstract
The recently introduced multidimensional archive of phenotypic elites (MAP-Elites) is an evolutionary algorithm capable of producing a large archive of diverse, high-performing solutions in a single run. It works by discretizing a continuous feature space into unique regions according to the desired discretization per dimension. While simple, this algorithm has a main drawback: it cannot scale to high-dimensional feature spaces since the number of regions increase exponentially with the number of dimensions. In this paper, we address this limitation by introducing a simple extension of MAP-Elites that has a constant, predefined number of regions irrespective of the dimensionality of the feature space. Our main insight is that methods from computational geometry could partition a high-dimensional space into well-spread geometric regions. In particular, our algorithm uses a centroidal Voronoi tessellation (CVT) to divide the feature space into a desired number of regions; it then places every generated individual in its closest region, replacing a less fit one if the region is already occupied. We demonstrate the effectiveness of the new “CVT-MAP-Elites” algorithm in high-dimensional feature spaces through comparisons against MAP-Elites in maze navigation and hexapod locomotion tasks.
Vassilis Vassiliades, Konstantinos Chatzilygeroudis, Jean-Baptiste Mouret
IEEE Trans. Evol. Comput.3
2017 Data-efficient exploration, optimization, and modeling of diverse designs through surrogate-assisted illumination
abstract
The MAP-Elites algorithm produces a set of high-performing solutions that vary according to features defined by the user. This technique to 'illuminate' the problem space through the lens of chosen features has the potential to be a powerful tool for exploring design spaces, but is limited by the need for numerous evaluations. The Surrogate-Assisted Illumination (SAIL) algorithm, introduced here, integrates approximative models and intelligent sampling of the objective function to minimize the number of evaluations required by MAP-Elites.
Adam Gaier, Alexander Asteroth, Jean-Baptiste Mouret
GECCO3
2017 Black-box data-efficient policy search for robotics
abstract
The most data-efficient algorithms for reinforcement learning (RL) in robotics are based on uncertain dynamical models: after each episode, they first learn a dynamical model of the robot, then they use an optimization algorithm to find a policy that maximizes the expected return given the model and its uncertainties. It is often believed that this optimization can be tractable only if analytical, gradient-based algorithms are used; however, these algorithms require using specific families of reward functions and policies, which greatly limits the flexibility of the overall approach. In this paper, we introduce a novel model-based RL algorithm, called Black-DROPS (Black-box Data-efficient RObot Policy Search) that: (1) does not impose any constraint on the reward function or the policy (they are treated as black-boxes), (2) is as data-efficient as the state-of-the-art algorithm for data-efficient RL in robotics, and (3) is as fast (or faster) than analytical approaches when several cores are available. The key idea is to replace the gradient-based optimization algorithm with a parallel, black-box algorithm that takes into account the model uncertainties. We demonstrate the performance of our new algorithm on two standard control benchmark problems (in simulation) and a low-cost robotic manipulator (with a real robot).
Konstantinos Chatzilygeroudis, Roberto Rama, Rituraj Kaushik, Dorian Goepp, Vassilis Vassiliades, Jean-Baptiste Mouret
IROS6
2017 Introduction to the Evolution of Physical Systems Special Issue
abstract
We are delighted to introduce this Special Issue on the Evolution of Physical Systems, the culmination of a series of workshops organized around the topic that began with ALIFE XIII (2012) and ran through ECAL 2013, ALIFE XIV (2014), and ECAL 2015.Inspired by our mutual interests, we coined the term evolution of physical systems (EPS) to describe evolutionary approaches that occur in real-world physical substrates rather than in simulation. We deliberately chose this term to be broad enough to encompass both parallel embodied evolution [19], in which evolution is distributed across a population of robots, and more classical evolutionary robotics work, where evaluation is serialized on a single robot, as by Floreano and Mondada [5].The evolution of physical systems has its roots in the embodiment philosophy of Rodney Brooks, who famously said “the world is its own best model” [3]. Brooks' emphatic critique of the symbol system hypothesis argued that relatively simple systems, developed and grounded in suitably complex real-world environments, can lead to the emergence of complex behaviors. Contemporary approaches to the EPS aim at automating Brooks' vision by exploiting evolutionary algorithms on systems that “live” and “evolve” in the real world, that is, whose behavior (and possibly their form) are as grounded as possible in the environment.Embodied evolutionary algorithms first gained prominence with the “Sussex approach” of Harvey et al. [7] in the 1990s. The Sussex gantry-based robot [4, 6], driven by a neural network, was capable of robust behaviors in a noisy real-world environment. Their motivation for using the real-world rather than a simulated environment stemmed largely from the fact that in that era of limited computational power, it was faster and more efficient to test networks in situ than it was to develop and run a simulator that realistically modeled sensor noise. This notion, that it is sometimes more practical (and computationally efficient) to dispense with simulation entirely and evolve behaviors directly in the real world, persists at the heart of the EPS.Two other results of the Sussex group are particularly noteworthy. First is Adrian Thompson's silicon evolution of a field-programmable gate array (FPGA) [17, 18]. The chosen task was to evolve a circuit capable of discriminating between 1- and 10-kHz signals without the use of a clock. By evolving in silicon, rather than in simulation, Thompson was able to generate nearly inscrutable solutions that exploited the analogue (and non-simulable) nature of FPGAs. This idea of using embodiment to arrive at “novel surprise” solutions is a further guiding principle of the EPS. The second is Jakobi's essential work on the reality gap [8], which highlighted the difficulty in transferring results evolved in simulation into the real world. Jakobi notes that the best results emerge when “the noise levels of the simulation have similar amplitudes to those observed in reality,” and points out that these approaches become less feasible as environments and sensors become more complex.The practice of evolving outside of simulators grew from there. Floreano and Mondada [5] were able to evolve robust wall-following behavior in a tethered Khepera robot through embodiment. Watson et al. [19] took embodiment a step further by simultaneously embedding an entire population of evolving TupperBots into a shared arena—and demonstrated how the behaviors that emerged through this process were qualitatively different, and more effective, than those evolved in a simulated environment (their use of an electrified floor as an alternative to tethers was particularly innovative). Zykov et al. evolved dynamic open-loop gaits on a large hexapod robot, remarking upon the amount of labor required to reset the large robot between physical trials [21].Several of our own research areas have been in this realm of problems that are easier to test physically than they are to simulate realistically. In their work on “evolutionary fabrication” [16, 12], Rieffel and colleagues ran an embodied evolutionary algorithm on a 3D printer, learning to produce a variety of shapes out of viscous extruded materials such as silicone. These efforts led to novel shapes (and novel means of producing the shapes) that no human would have designed. Similarly, and more recently, Rieffel's work on the evolution of vibration-based gaits for tensegrity robots [15] relied purely on physical trials, and led not only to fast and dynamically complex gaits, but also to a state-machine-based controller for a tensegrity robot able to chase targets [9].While not using a simulator at all has many benefits, evolutionary robotics can take advantage of both simulators and real experiments to accelerate the evolutionary process while keeping a direct link with the real world. A first idea is to start evolution in simulation and then finish it on the real robot(s); this is, for instance, what Lipson and Pollack did to evolve the morphology and controllers of 3D-printed robots [13]. However, this approach means that designs for the real world are variants of those evolved in simulation, that is, robots are more grounded in the simulation than in the real world. A second, classic idea is to use data collected from tests on the real robot to improve the simulator, ideally to make the reality gap so small that it becomes unnoticeable [2, 20]. Nevertheless, to do so, simulators need to be generic enough to model almost anything that could happen, which, in turn, requires acquiring a considerable amount of data in the real world. More recently, a weaker version of this idea has emerged: Instead of trying to correct the simulator, one could use a few tests in the real world to learn to avoid the behaviors that exploit phenomena that are inaccurately simulated. Following this concept, Koos et al. managed to evolve controllers for six-legged robots and Khepera-like robots with only a dozen tests in the real world [10, 11, 14].Despite numerous advances in both robotics and technology, Rodney Brooks' prescient words about the crucial role of embodiment in robotics remain as true as ever, and the EPS is as exciting as ever.We the editors are delighted to present the articles that were selected to be published in this issue. Collectively they provide an inspiring snapshot of the state of the evolution of physical systems.Jacobsen and colleagues describe an experimental setup that removes the burden of manually setting the environment (e.g., moving the robot(s) and object(s) back to their initial positions and orientations) from the human supervisor. This is particularly relevant in evolutionary robotics, where a large number of trials are generally required, all starting from similar initial conditions. The authors present a proof of concept, tested with a real robotics setup in an arena of 1 by 0.5 m, which makes use of an industrial robot arm, a camera, and fiducial markers that can be stuck on top of robots and objects alike. These fiducial markers are used both for monitoring the experiments, by tracking positions and orientations, and as visual anchors for the robot arm to grab so as to move back robots and objects to their initial positions. Beyond evolutionary robotics, the proposed setup may also be used for other robotics setups involving multiple trials (e.g., reinforcement learning and embodied evolution), whether for resetting the experimental setup, as demonstrated in this article, or for other, more dynamical interventions (e.g., by conditionally moving objects around or transporting robots during the course of an experiment). To ensure that their system can be adopted and easily deployed by others, the authors have included the source code and a detailed description of the technical aspects of their setup.In their contribution, Scheper and de Croon describe how choosing an appropriate level of abstraction can improve the transfer of evolved solutions across the reality gap. The task domain they have have chosen is a decentralized formation flying task for a group of three small airborne vehicles controlled by neural networks. They compare the performance of a low-level controller that has precise control over vehicle dynamics (pitch rate, pitch angle, etc.) against a higher level controller that uses a more abstract velocity setpoint. After evolution (in simulation), the more abstract-level controller was better able to survive transfer into the real world. Scheper and de Croon suggest that the higher-level controller is better able to perform a kind of environmental exploration by inherently modeling the behavior of the other vehicles in the swarm. As they point out, an advantage of more abstracted controllers is that it reduces the necessity of high-fidelity simulation—an important advantage for evolutionary robotics tasks. We were impressed by the insights of this article in regard to the tradeoffs between simulator fidelity and real-world behavior. The flight videos on YouTube are also quite impressive.The work of Vujovic and colleagues is very much in the spirit of this special issue. They introduce a novel robot-building robot composed of a hot-melt extruder and an articulated arm, which acts as a sort of 3D printer, capable of fabricating the legs of simple modular robots. An evolutionary developmental approach is used to “grow” the designs of the legs, and once printed, the locomotive abilities of the printed robots are automatically evaluated in situ. This type of physically embodied integrated design-fabricate-test loop, in which one robot evolves and assembles other robots, is very much in the vein of Lipson and Pollack's GOLEM robots [13], moving the field one step closer to 3D-printed robots that can walk themselves off the printer.Preen and Bull present an interesting theoretical analysis followed up with an actual physical case study. Evaluation of physical artefacts is commonly a time-consuming aspect of the EPS, so a thorough analysis of methods to minimize the number of required evaluations can greatly boost applicability. Preen and Bull employ the abstract NKCS model of coevolution to achieve this: It enables informed sampling of candidate solutions in a coevolutionary scenario, where various components of a physical system interact in combination with a neural-net-based surrogate model. A successful case study to develop a heterogeneous set of vertical wind turbines for wind farms shows the applicability and relevance of their method to the evolution of designs for physical artefacts without relying on simulation.For obvious practical reasons, most of the work in the EPS has been focused on the evolution of controllers for existing robots, that is, for robots with a fixed morphology. Nevertheless, most of those who are interested in the field dream of a much more ambitious vision: robots whose morphology and brains can evolve on their own, much in the same ways as animals evolve.In their contribution, Jelisavcic et al. investigate how such a vision could be achieved with a population of robots and 3D printing technology. Using the RoboGen system of 3D-printed modules [1], they first designed and 3D-printed two robots. In a second step, they put these two robots in an arena so that they could exchange their genome to create the genotype of a third robot, which they also 3D-printed. This first step of an evolutionary experiment is an interesting stepping stone to identify the challenges of setting up experiments in which a population of robots would evolve autonomously.Overall, these articles provide valuable contributions to the field of the evolution of physical systems, demonstrating that many of the current challenges of the field are more technological than scientific: To design machines that evolve in the physical world, we first need machines that can self-reproduce or, at least, autonomous machines that can in turn help other machines to evolve. And until high-fidelity computationally efficient simulators are developed, the role of physical embodiment in evolutionary robotics remains essential.
John Rieffel, Jean-Baptiste Mouret, Nicolas Bredèche, Evert Haasdijk
Artif. Life2
2016 Does Aligning Phenotypic and Genotypic Modularity Improve the Evolution of Neural Networks?
abstract
Many argue that to evolve artificial intelligence that rivals that of natural animals, we need to evolve neural networks that are structurally organized in that they exhibit modularity, regularity, and hierarchy. It was recently shown that a cost for network connections, which encourages the evolution of modularity, can be combined with an indirect encoding, which encourages the evolution of regularity, to evolve networks that are both modular and regular. However, the bias towards regularity from indirect encodings may prevent evolution from independently optimizing different modules to perform different functions, unless modularity in the phenotype is aligned with modularity in the genotype. We test this hypothesis on two multi-modal problems---a pattern recognition task and a robotics task---that each require different phenotypic modules. In general, we find that performance is improved only when genotypic and phenotypic modularity are encouraged simultaneously, though the role of alignment remains unclear. In addition, intuitive manual decompositions fail to provide the performance benefits of automatic methods on the more challenging robotics problem, emphasizing the importance of automatic, rather than manual, decomposition methods. These results suggest encouraging modularity in both the genotype and phenotype as an important step towards solving large-scale multi-modal problems, but also indicate that more research is required before we can evolve structurally organized networks to solve tasks that require multiple, different neural modules.
Joost Huizinga, Jean-Baptiste Mouret, Jeff Clune
GECCO2
2016 How do Different Encodings Influence the Performance of the MAP-Elites Algorithm?
abstract
The recently introduced Intelligent Trial and Error algorithm (IT&E) both improves the ability to automatically generate controllers that transfer to real robots, and enables robots to creatively adapt to damage in less than 2 minutes. A key component of IT&E is a new evolutionary algorithm called MAP-Elites, which creates a behavior-performance map that is provided as a set of "creative" ideas to an online learning algorithm. To date, all experiments with MAP-Elites have been performed with a directly encoded list of parameters: it is therefore unknown how MAP-Elites would behave with more advanced encodings, like HyperNeat and SUPG. In addition, because we ultimately want robots that respond to their environments via sensors, we investigate the ability of MAP-Elites to evolve closed-loop controllers, which are more complicated, but also more powerful. Our results show that the encoding critically impacts the quality of the results of MAP-Elites, and that the differences are likely linked to the locality of the encoding (the likelihood of generating a similar behavior after a single mutation). Overall, these results improve our understanding of both the dynamics of the MAP-Elites algorithm and how to best harness MAP-Elites to evolve effective and adaptable robotic controllers.
Danesh Tarapore, Jeff Clune, Antoine Cully, Jean-Baptiste Mouret
GECCO4
2016 Evolving a Behavioral Repertoire for a Walking Robot
abstract
Numerous algorithms have been proposed to allow legged robots to learn to walk. However, most of these algorithms are devised to learn walking in a straight line, which is not sufficient to accomplish any real-world mission. Here we introduce the Transferability-based Behavioral Repertoire Evolution algorithm (TBR-Evolution), a novel evolutionary algorithm that simultaneously discovers several hundreds of simple walking controllers, one for each possible direction. By taking advantage of solutions that are usually discarded by evolutionary processes, TBR-Evolution is substantially faster than independently evolving each controller. Our technique relies on two methods: (1) novelty search with local competition, which searches for both high-performing and diverse solutions, and (2) the transferability approach, which combines simulations and real tests to evolve controllers for a physical robot. We evaluate this new technique on a hexapod robot. Results show that with only a few dozen short experiments performed on the robot, the algorithm learns a repertoire of controllers that allows the robot to reach every point in its reachable space. Overall, TBR-Evolution introduced a new kind of learning algorithm that simultaneously optimizes all the achievable behaviors of a robot.
Antoine Cully, Jean-Baptiste Mouret
Evol. Comput.2
2016 The Evolutionary Origins of Hierarchy
abstract
Hierarchical organization-the recursive composition of sub-modules-is ubiquitous in biological networks, including neural, metabolic, ecological, and genetic regulatory networks, and in human-made systems, such as large organizations and the Internet. To date, most research on hierarchy in networks has been limited to quantifying this property. However, an open, important question in evolutionary biology is why hierarchical organization evolves in the first place. It has recently been shown that modularity evolves because of the presence of a cost for network connections. Here we investigate whether such connection costs also tend to cause a hierarchical organization of such modules. In computational simulations, we find that networks without a connection cost do not evolve to be hierarchical, even when the task has a hierarchical structure. However, with a connection cost, networks evolve to be both modular and hierarchical, and these networks exhibit higher overall performance and evolvability (i.e. faster adaptation to new environments). Additional analyses confirm that hierarchy independently improves adaptability after controlling for modularity. Overall, our results suggest that the same force-the cost of connections-promotes the evolution of both hierarchy and modularity, and that these properties are important drivers of network performance and adaptability. In addition to shedding light on the emergence of hierarchy across the many domains in which it appears, these findings will also accelerate future research into evolving more complex, intelligent computational brains in the fields of artificial intelligence and robotics.
Henok Mengistu, Joost Huizinga, Jean-Baptiste Mouret, Jeff Clune
PLoS Comput. Biol.3
2015 Evolvability signatures of generative encodings: Beyond standard performance benchmarks
Danesh Tarapore, Jean-Baptiste Mouret
Inf. Sci.2
2015 Neural Modularity Helps Organisms Evolve to Learn New Skills without Forgetting Old Skills
abstract
A long-standing goal in artificial intelligence is creating agents that can learn a variety of different skills for different problems. In the artificial intelligence subfield of neural networks, a barrier to that goal is that when agents learn a new skill they typically do so by losing previously acquired skills, a problem called catastrophic forgetting. That occurs because, to learn the new task, neural learning algorithms change connections that encode previously acquired skills. How networks are organized critically affects their learning dynamics. In this paper, we test whether catastrophic forgetting can be reduced by evolving modular neural networks. Modularity intuitively should reduce learning interference between tasks by separating functionality into physically distinct modules in which learning can be selectively turned on or off. Modularity can further improve learning by having a reinforcement learning module separate from sensory processing modules, allowing learning to happen only in response to a positive or negative reward. In this paper, learning takes place via neuromodulation, which allows agents to selectively change the rate of learning for each neural connection based on environmental stimuli (e.g. to alter learning in specific locations based on the task at hand). To produce modularity, we evolve neural networks with a cost for neural connections. We show that this connection cost technique causes modularity, confirming a previous result, and that such sparsely connected, modular networks have higher overall performance because they learn new skills faster while retaining old skills more and because they have a separate reinforcement learning module. Our results suggest (1) that encouraging modularity in neural networks may help us overcome the long-standing barrier of networks that cannot learn new skills without forgetting old ones, and (2) that one benefit of the modularity ubiquitous in the brains of natural animals might be to alleviate the problem of catastrophic forgetting.
Kai Olav Ellefsen, Jean-Baptiste Mouret, Jeff Clune
PLoS Comput. Biol.2
2014 Summary of "The Evolutionary Origins of Modularity"
abstract
A long-standing, open question in biology is how populations are capable of rapidly adapting to novel environments, a trait called evolvability. A major contributor to evolvability is the fact that many biological entities are modular, especially the many biological processes and structures that can be modeled as networks, such as metabolic pathways, gene regulation, protein interactions, and animal brains. Networks are modular if they contain highly connected clusters of nodes that are sparsely connected to nodes in other clusters [4, 2]. Despite its importance and decades of research, there is no agreement on why modularity evolves [4]. Intuitively, modular systems seem more adaptable, a lesson well-known to human engineers, because it is easier to rewire a modular network with functional subunits than an entangled, monolithic network [1]. However, because this evolvability only provides a selective advantage over the long-term, such selection is at best indirect and may not be strong enough to explain the level of modularity in the natural world [4]. Modularity is likely caused by multiple forces acting to various degrees in different contexts [4], and a comprehensive understanding of the evolutionary origins of modularity involves identifying those multiple forces and their relative contributions. The leading hypothesis is that modularity mainly emerges due to rapidly changing environments that have common subproblems, but different overall problems [1]. It is unknown how much natural modularity MVG can explain, however, because it unclear if biological environments change modularly, and whether they change at a high enough frequency for this force to play a significant role. We investigate an alternate hypothesis that has been suggested, but heretofore untested, which is that modularity evolves not because it conveys evolvability, but as a byproduct from selection to reduce connection costs in a network [3].
Jeff Clune, Jean-Baptiste Mouret, Hod Lipson
ALIFE2
2014 Learning to Walk in Every Direction with the TBR-Learning algorithm
abstract
Legged robots are versatile machines that can outperform wheeled robots on rough terrain (Raibert, 1986), for instance in exploration or rescue missions. Their versatility is, however, tempered by their mechanical and control complexity, which makes them prone to mechanical damages and difficult to control robustly (Raibert, 1986; Bongard et al., 2006; Koos et al., 2013a). A promising way to compensate for these two weaknesses is to let robots discover on their own the best way to move in the current situation. A legged robot can thus cope with an unexpected terrain or with mechanical damages by learning a new walking gait (Bongard et al., 2006; Koos et al., 2013a), in the same way as animals can learn to limp with a sprained ankle. Reinforcement learning (Kohl and Stone, 2004; Tedrake et al., 2005) and evolutionary algorithms (Zykov et al., 2004; Chernova and Veloso, 2004; Hornby et al., 2005) have been investigated to discover walking gaits for physical robots. Nevertheless, most of these investigations are limited to straight, forward walking, whereas a robot that only walks along a straight line is obviously unable to accomplish any mission. Only a handful of works deal with controllers able to turn or to change the walking speed. In these cases, controllers are successively evaluated on each possible direction (Mouret et al., 2006), or learned with an incremental process (Kodjabachian and Meyer, 1998). Compared to learning a simple controller, these two approaches significantly increase the learning time and the complexity of the search process. In the present paper, we describe the Transferabilitybased Behavioral Repertoire Evolution (TBR-Evolution), a new learning algorithm that allows a robot to learn to walk in every direction in a single run of evolutionary algorithm. This algorithm combines the BR-Evolution algorithm (Cully and Mouret, 2013), which creates a behavioral repertoire in a single run, with the transferability approach (Koos et al., 2013b), which minimizes the number of evaluations on a physical robot when evolving controllers thanks to a simulator. A behavioral repertoire is a collection of simple controllers, where each of them reaches one position. An exter1 2 3 4 5
Antoine Cully, Jean-Baptiste Mouret
ALIFE2
2014 Abstract of: "Fast Damage Recovery in Robotics with the T-Resilience Algorithm"
abstract
of: Fast Damage Recovery in Robotics with the T-Resilience Algorithm Sylvain Koos, Antoine Cully and Jean-Baptiste Mouret Sorbonne Universites, UPMC Univ Paris 06, UMR 722, ISIR, F-75005, Paris, France CNRS, UMR 7222, ISIR, F-75005, Paris, France [email protected] Damage recovery is critical for autonomous robots that need to operate for a long time without assistance. Most current methods are complex and costly because they require anticipating each potential damage in order to have a contingency plan ready and diagnosis procedures. An alternative line of thought is to let the robot learn on its own the best behavior for the current situation. If the learning process is open enough, then the robot should be able to discover new compensatory behaviors in situations that have not been foreseen by its designers. Classic reinforcement learning algorithms are hard to apply to low-level robotic problems (Togelius et al., 2009), but evolutionary algorithms (EAs) are good candidates to find original solutions because they can optimize in the continuous domain and work on the structure of controllers (e.g. neural networks). When evolving controllers for robots, EAs are reported to require many hundreds of trials on the robot and to last from two to tens of hours (e.g. (Hornby et al., 2005; Yosinski et al., 2011)). These EAs spend most of their running time in evaluating the quality of controllers by testing them on the target robot. Since, contrary to simulation, reality cannot be sped up, their running time can only be improved by finding strategies to evaluate fewer candidate solutions on the robot. By first learning a self-model for the robot, then evolving a controller with this simulation, Bongard et al. (Bongard et al., 2006) designed an algorithm for resilience that makes an important step in this direction. Nevertheless, this algorithm has a few important shortcomings. First, actions and models are undirected: the algorithm can “waste” a lot of time to improve parts of the self-model that are irrelevant for the task. Second, the diagnosis may be wrong, which leads to a useless contingency plan. Third, there is often a “reality gap” between a behavior learned in simulation and the same behavior on the target robot (Jakobi et al., 1995), but nothing is included in Bongard’s algorithm to prevent such gap to happen: the controller learned in the simulation may not work well on the real robot, even if the self-model is accurate. This paper is an extended abstract of Koos et al. (2013a). Figure 1: (A) The hexapod robot is not damaged. (B) The left middle leg is no longer powered. (C) The terminal part of the front right leg is shortened by half. Our algorithm is inspired by the “transferability approach” (Koos et al., 2013b), whose original purpose is to cross the “reality gap” that separates behaviors optimized in simulation to those observed on the target robot. The main proposition of this approach is to make the optimization algorithm aware of the limits of the simulation. To this end, a few controllers are transferred during the optimization and a regression algorithm (here a SVM) is used to approximate the function that maps behaviors in simulation to the difference of performance between simulation and reality. To use this approximated transferability function, the singleobjective optimization problem is transformed into a multiobjective optimization in which both performance in simulation and transferability are maximized. This optimization is performed with a a multi-objective evolutionary algorithm (NSGA-II, Deb et al. (2002)). The same concepts can be applied to design a fast adaptation algorithm for resilient robotics, leading to a new algorithm that we called “T-Resilience” (for Transferabilitybased resilience). If a damaged robot embeds a simulation of itself, then behaviors that rely on damaged parts will not be transferable: they will perform very differently in the self-model and in reality. During the adaptation process, the robot will thus create an approximated transferability function that classifies behaviors as “working as expected” and “not working as expected”. Hence the robot will possess an “intuition” of the damages but it will not explicitly represent or identify them. By optimizing both the transferability and ALIFE 14: Proceedings of the Fourteenth International Conference on the Synthesis and Simulation of Living Systems −0.50 −0.25 0.00 0.25 0.50 0.75 1.00
Sylvain Koos, Antoine Cully, Jean-Baptiste Mouret
ALIFE3
2014 Abstract of: "On the Relationship between Generative Encodings, Regularity, and Learning Abilities when Evolving Plastic Artificial Networks"
Jean-Baptiste Mouret, Paul Tonelli
ALIFE1
2014 Comparing the Evolvability of Generative Encoding Schemes
abstract
International audience
Danesh Tarapore, Jean-Baptiste Mouret
ALIFE2
2014 Evolving neural networks that are both modular and regular: HyperNEAT plus the connection cost technique
abstract
One of humanity's grand scientific challenges is to create artificially intelligent robots that rival natural animals in intelligence and agility. A key enabler of such animal complexity is the fact that animal brains are structurally organized in that they exhibit modularity and regularity, amongst other attributes. Modularity is the localization of function within an encapsulated unit. Regularity refers to the compressibility of the information describing a structure, and typically involves symmetries and repetition. These properties improve evolvability, but they rarely emerge in evolutionary algorithms without specific techniques to encourage them. It has been shown that (1) modularity can be evolved in neural networks by adding a cost for neural connections and, separately, (2) that the HyperNEAT algorithm produces neural networks with complex, functional regularities. In this paper we show that adding the connection cost technique to HyperNEAT produces neural networks that are significantly more modular, regular, and higher performing than HyperNEAT without a connection cost, even when compared to a variant of HyperNEAT that was specifically designed to encourage modularity. Our results represent a stepping stone towards the goal of producing artificial neural networks that share key organizational properties with the brains of natural animals.
Joost Huizinga, Jeff Clune, Jean-Baptiste Mouret
GECCO3
2013 Behavioral diversity with multiple behavioral distances
abstract
Recent results in evolutionary robotics show that explicitly encouraging the behavioral diversity of candidate solutions drastically improves the convergence of many experiments. The performance of this technique depends, however, on the choice of a behavioral similarity measure (BSM). Here we propose that the experimenter does not actually need to choose: provided that several similarity measures are conceivable, using them all could lead to better results than choosing a single one. Values computed by several BSM can be averaged, which is computationally expensive because it requires the computation of all the BSM at each generation, or randomly switched at a user-chosen frequency, which is a cheaper alternative. We compare these two approaches in two experimental setups - a ball collecting task and hexapod locomotion - with five different BSMs. Results show that (1) using several BSM in a single run increases the performance while avoiding the need to choose the most appropriate BSM and (2) switching between BSMs leads to better results than taking the mean behavioral diversity, while requiring less computational power.
Stéphane Doncieux, Jean-Baptiste Mouret
IEEE Congress on Evolutionary Computation2
2013 Behavioral repertoire learning in robotics
abstract
Learning in robotics typically involves choosing a simple goal (e.g. walking) and assessing the performance of each controller with regard to this task (e.g. walking speed). However, learning advanced, input-driven controllers (e.g. walking in each direction) requires testing each controller on a large sample of the possible input signals. This costly process makes difficult to learn useful low-level controllers in robotics. Here we introduce BR-Evolution, a new evolutionary learning technique that generates a behavioral repertoire by taking advantage of the candidate solutions that are usually discarded. Instead of evolving a single, general controller, BR-evolution thus evolves a collection of simple controllers, one for each variant of the target behavior; to distinguish similar controllers, it uses a performance objective that allows it to produce a collection of diverse but high-performing behaviors. We evaluated this new technique by evolving gait controllers for a simulated hexapod robot. Results show that a single run of the EA quickly finds a collection of controllers that allows the robot to reach each point of the reachable space. Overall, BR-Evolution opens a new kind of learning algorithm that simultaneously optimizes all the achievable behaviors of a robot.
Antoine Cully, Jean-Baptiste Mouret
GECCO2
2013 The Transferability Approach: Crossing the Reality Gap in Evolutionary Robotics
abstract
The reality gap, which often makes controllers evolved in simulation inefficient once transferred onto the physical robot, remains a critical issue in evolutionary robotics (ER). We hypothesize that this gap highlights a conflict between the efficiency of the solutions in simulation and their transferability from simulation to reality: the most efficient solutions in simulation often exploit badly modeled phenomena to achieve high fitness values with unrealistic behaviors. This hypothesis leads to the transferability approach, a multiobjective formulation of ER in which two main objectives are optimized via a Pareto-based multiobjective evolutionary algorithm: 1) the fitness; and 2) the transferability, estimated by a simulation-to-reality (STR) disparity measure. To evaluate this second objective, a surrogate model of the exact STR disparity is built during the optimization. This transferability approach has been compared to two reality-based optimization methods, a noise-based approach inspired from Jakobi's minimal simulation methodology and a local search approach. It has been validated on two robotic applications: 1) a navigation task with an e-puck robot; and 2) a walking task with a 8-DOF quadrupedal robot. For both experimental setups, our approach successfully finds efficient and well-transferable controllers only with about ten experiments on the physical robot.
Sylvain Koos, Jean-Baptiste Mouret, Stéphane Doncieux
IEEE Trans. Evol. Comput.2
2012 Encouraging Behavioral Diversity in Evolutionary Robotics: An Empirical Study
abstract
Evolutionary robotics (ER) aims at automatically designing robots or controllers of robots without having to describe their inner workings. To reach this goal, ER researchers primarily employ phenotypes that can lead to an infinite number of robot behaviors and fitness functions that only reward the achievement of the task-and not how to achieve it. These choices make ER particularly prone to premature convergence. To tackle this problem, several papers recently proposed to explicitly encourage the diversity of the robot behaviors, rather than the diversity of the genotypes as in classic evolutionary optimization. Such an approach avoids the need to compute distances between structures and the pitfalls of the noninjectivity of the phenotype/behavior relation; however, it also introduces new questions: how to compare behavior? should this comparison be task specific? and what is the best way to encourage diversity in this context? In this paper, we review the main published approaches to behavioral diversity and benchmark them in a common framework. We compare each approach on three different tasks and two different genotypes. The results show that fostering behavioral diversity substantially improves the evolutionary process in the investigated experiments, regardless of genotype or task. Among the benchmarked approaches, multi-objective methods were the most efficient and the generic, Hamming-based, behavioral distance was at least as efficient as task specific behavioral metrics.
Jean-Baptiste Mouret, Stéphane Doncieux
Evol. Comput.1
2011 How to promote generalisation in evolutionary robotics: the ProGAb approach
abstract
In Evolutionary Robotics (ER), controllers are assessed in a single or a few environments. As a consequence, good performances in new different contexts are not guaranteed. While a lot of ER works deal with robustness, i.e. the ability to perform well on new contexts close to the ones used for evaluation, no current approach is able to promote broader generalisation abilities without any assumption on the new contexts. In this paper, we introduce the ProGAb approach, which is based on the standard three data sets methodology of supervised machine learning, and compare it to state-of-the-art ER methods on two simulated robotic tasks: a navigation task in a T-maze and a more complex ball-collecting task in an arena. In both applications, the ProGAb approach: (1) produced controllers with better generalisation abilities than the other methods; (2) needed two to three times fewer evaluations to discover such solutions.
Tony Pinville, Sylvain Koos, Jean-Baptiste Mouret, Stéphane Doncieux
GECCO3
2011 On the relationships between synaptic plasticity and generative systems
abstract
The present paper analyzes the mutual relationships between generative and developmental systems (GDS) and synaptic plasticity when evolving plastic artificial neural networks (ANNs) in reward-based scenarios. We first introduce the concept of synaptic Transitive Learning Abilities (sTLA), which reflects how well an evolved plastic ANN can cope with learning scenarios not encountered during the evolution process. We subsequently report results of a set of experiments designed to check that (1) synaptic plasticity can help a GDS to fine-tune synaptic weights and (2) that with the investigated generative encoding (EvoNeuro), only a few learning scenarios are necessary to evolve a general learning system, which can adapt itself to reward-based scenarios not tested during the fitness evaluation.
Paul Tonelli, Jean-Baptiste Mouret
GECCO2
2011 Stochastic optimization of a chain sliding mode controller for the mobile robot maneuvering
abstract
In this study we present a chain sliding mode controller for the control of a four wheeled autonomous mobile robot performing aggressive turning maneuver to 90 degrees on a slippery surface. The controller consists of a set of local sliding mode controllers and the hyperplanes of switching between them. The parameters of the sliding mode controllers and the hyperplanes are obtained using methods of multiobjective stochastic optimization applied to a model of the robot. The obtained controller is used to drive the mobile robot. The results show that the controller allowed the robot to execute the aggressive maneuver. Moreover, the turn radius obtained with the controller was twice less than the minimal turn radius admitted by the robot's geometry and the steering system.
Alexander V. Terekhov, Jean-Baptiste Mouret, Christophe Grand
IROS2
2010 Behavioral diversity measures for Evolutionary Robotics
abstract
In Evolutionary Robotics (ER), explicitly rewarding for behavioral diversity recently revealed to generate efficient results without recourse to complex fitness functions. The principle of such approaches is to explicitly encourage diversity in the robot behavior space instead of in the space of genotypes (the space explored by the evolutionary algorithm) or the space of phenotypes (the space of robot controllers and morphologies). To implement such approaches, a similarity between behaviors needs to be evaluated but, up to now, used similarity measures are problem-specific. The goal of this work is to explore generic behavioral similarity measures that only rely on sensori-motor values. With such a measure, we managed to evolve the topology and the parameters of neuro-controllers that make a simulated robot go towards a ball, take it, find a basket, put the ball into the basket, perform a half-turn, search and take another ball, put it into the basket, etc. In this experiment, two objectives were simultaneously optimized with NSGA-II: the number of collected balls and the generic behavioral diversity objective. Several generic behavioral measures are compared. To confirm the interpretation of behavioral diversity objective and in an attempt to characterize behavioral similarity measures, they are also compared to human-made behavioral similarity evaluations. They reveal to classify behaviors globally as humans did, but with no clear correlation between the closeness to human classification and the efficiency within an evolutionary run.
Stéphane Doncieux, Jean-Baptiste Mouret
IEEE Congress on Evolutionary Computation2
2010 Sferesv2: Evolvin' in the multi-core world
abstract
This paper introduces and benchmarks Sferesv2, a C++ framework designed to help researchers in evolutionary computation to make their code run as fast as possible on a multi-core computer. It is based on three main concepts: (1) including multi-core optimizations from the start of the design process; (2) providing state-of-the art implementations of well-selected current evolutionary algorithms (EA), and especially multiobjective EAs; (3) being based on modern (template-based) C++ techniques to be both abstract and efficient. Benchmark results show that when a single core is used, running time of classic EAs included in Sferesv2(NSGA-2 and CMA-ES) are of the same order of magnitude than specialized C code. When n cores are used, typical speed-ups range from 0.75n to 0.9n; however, parallelization efficiency critically depends on the time to evaluate the fitness function.
Jean-Baptiste Mouret, Stéphane Doncieux
IEEE Congress on Evolutionary Computation1
2010 Crossing the reality gap in evolutionary robotics by promoting transferable controllers
abstract
The reality gap, that often makes controllers evolved in simulation inefficient once transferred onto the real system, remains a critical issue in Evolutionary Robotics (ER); it prevents ER application to real-world problems. We hypothesize that this gap mainly stems from a conflict between the efficiency of the solutions in simulation and their transferability from simulation to reality: best solutions in simulation often rely on bad simulated phenomena (e.g. the most dynamic ones). This hypothesis leads to a multi-objective formulation of ER in which two main objectives are optimized via a Pareto-based Multi-Objective Evolutionary Algorithm: (1) the fitness and (2) the transferability. To evaluate this second objective, a simulation-to-reality disparity value is approximated for each controller. The proposed method is applied to the evolution of walking controllers for a real 8-DOF quadrupedal robot. It successfully finds efficient and well-transferable controllers with only a few experiments in reality.
Sylvain Koos, Jean-Baptiste Mouret, Stéphane Doncieux
GECCO2
2010 Importing the computational neuroscience toolbox into neuro-evolution-application to basal ganglia
abstract
Neuro-evolution and computational neuroscience are two scientific domains that produce surprisingly different artificial neural networks. Inspired by the "toolbox" used by neuroscientists to create their models, this paper argues two main points: (1) neural maps (spatially-organized identical neurons) should be the building blocks to evolve neural networks able to perform cognitive functions and (2) well-identified modules of the brain for which there exists computational neuroscience models provide well-defined benchmarks for neuro-evolution.
Jean-Baptiste Mouret, Stéphane Doncieux, Benoît Girard 0001
GECCO1
2010 Stochastic optimization of a neural network-based controller for aggressive maneuvers on loose surfaces
abstract
In this study we develop a feedback controller for a four wheeled autonomous mobile robot. The purpose of the controller is to guarantee robust performance of an aggressive maneuver (90 degrees turn) at high velocity (about 10 m/s) on a loose surface (dirty road). To tackle this highly nonlinear control problem, we employ multi-objective evolutionary algorithms to explore and optimize the parameters of a neural network-based controller. The obtained controller is shown to be robust with respect to uncertainties of the robot parameters, speed of the maneuver and properties of the ground. The controller is tested using two mathematical models of significantly different complexity and accuracy.
Alexander V. Terekhov, Jean-Baptiste Mouret, Christophe Grand
IROS2
2009 Automatic system identification based on coevolution of models and tests
abstract
In evolutionary robotics, controllers are often designed in simulation, then transferred onto the real system. Nevertheless, when no accurate model is available, controller transfer from simulation to reality means potential performance loss. It is the reality gap problem. Unmanned aerial vehicles are typical systems where it may arise. Their locomotion dynamics may be hard to model because of a limited knowledge about the underlying physics. Moreover, a batch identification approach is difficult to use due to costly and time consuming experiments. An automatic identification method is then needed that builds a relevant local model of the system concerning a target issue. This paper deals with such an approach that is based on coevolution of models and tests. It aims at improving both modeling and control of a given system with a limited number of manipulations carried out on it. Experiments conducted with a simulated quadrotor helicopter show promising initial results about test learning and control improvement.
Sylvain Koos, Jean-Baptiste Mouret, Stéphane Doncieux
IEEE Congress on Evolutionary Computation2
2009 Overcoming the bootstrap problem in evolutionary robotics using behavioral diversity
abstract
The bootstrap problem is often recognized as one of the main challenges of evolutionary robotics: if all individuals from the first randomly generated population perform equally poorly, the evolutionary process won't generate any interesting solution. To overcome this lack of fitness gradient, we propose to efficiently explore behaviors until the evolutionary process finds an individual with a non-minimal fitness. To that aim, we introduce an original diversity-preservation mechanism, called behavioral diversity, that relies on a distance between behaviors (instead of genotypes or phenotypes) and multi-objective evolutionary optimization. This approach has been successfully tested and compared to a recently published incremental evolution method (multi-subgoal evolution) on the evolution of a neuro-controller for a light-seeking mobile robot. Results obtained with these two approaches are qualitatively similar although the introduced one is less directed than multi-subgoal evolution.
Jean-Baptiste Mouret, Stéphane Doncieux
IEEE Congress on Evolutionary Computation1
2009 Evolving modular neural-networks through exaptation
abstract
Despite their success as optimization methods, evolutionary algorithms face many difficulties to design artifacts with complex structures. According to paleontologists, living organisms evolved by opportunistically co-opting characters adapted to a function to solve new problems, a phenomenon called exaptation. In this paper, we draw the hypotheses (1) that exaptation requires the presence of multiple selection pressures, (2) that Pareto-based multi-objective evolutionary algorithms (MOEA) can create such pressures and (3) that the modularity of the genotype is a key to enable exaptation. To explore these hypotheses, we designed an evolutionary process to find the structure and the parameters of neural networks to compute a Boolean function with a modular structure. We then analyzed the role of each component using a Shapley value analysis. Our results show that: (1) the proposed method is efficient to evolve neural networks to solve this task; (2) genotypic modules and multiple selections gradients needed to be aligned to converge faster than the control experiments. This prominent role of multiple selection pressures contradicts the basic assumption that underlies most published modular methods for the evolution of neural networks, in which only the modularity of the genotype is considered.
Jean-Baptiste Mouret, Stéphane Doncieux
IEEE Congress on Evolutionary Computation1
2009 Single step evolution of robot controllers for sequential tasks
abstract
The generation of robot controllers for a task requiring a sequence of elementary behaviors is still a challenge. If these behaviors are known, intermediate steps can be given to help bootstrap the search, thus leading to task decomposition or incremental approaches. The goal of this paper is to propose an alternative, within which such behaviors do not need to be known. The proposed approach relies on a classical multi-objective evolutionary algorithm and consists in designing objectives dedicated to the enhancement of evolutionary search abilities. These objectives are to be used in addition to performance objectives rewarding the efficiency, robustness, or whatever aspect a robot designer might be interested in. Two different kinds of objectives are proposed, tested and compared on a ball collecting problem. Both rely on states that can be directly extracted from the sensors and are completely independent from the genotype and phenotype. They show promising results, even with a simple direct neural network encoding.
Stéphane Doncieux, Jean-Baptiste Mouret
GECCO2
2009 Using behavioral exploration objectives to solve deceptive problems in neuro-evolution
abstract
Encouraging exploration, typically by preserving the diversity within the population, is one of the most common method to improve the behavior of evolutionary algorithms with deceptive fitness functions. Most of the published approaches to stimulate exploration rely on a distance between genotypes or phenotypes; however, such distances are difficult to compute when evolving neural networks due to (1) the algorithmic complexity of graph similarity measures, (2) the competing conventions problem and (3) the complexity of most neural-network encodings.
Jean-Baptiste Mouret, Stéphane Doncieux
GECCO1