EDBT 2026 Demo / reviewers in the wild / expert
Miranda Coninx
dblp:40/10833 · also Alex Coninx, Alexandre Coninx
· DBLP profile ↗
11ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0001-7992-8183ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 70% Robot manipulation · 30% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-robot interaction · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation
grasping |
0.6 | 1 | 2022 | Automatic Acquisition of a Repertoire of Diverse Grasping Trajectories through Behavior Shaping and Novelty Search · ICRA 2022 |
Machine learning › Reinforcement learning
exploration |
0.4 | 1 | 2020 | Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020 |
Machine learning › Reinforcement learning › exploration
novelty-based exploration |
0.4 | 1 | 2020 | Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020 |
Machine learning › Reinforcement learning
policy learning |
0.4 | 1 | 2020 | Unsupervised Learning and Exploration of Reachable Outcome Space · ICRA 2020 |
Human-robot interaction
child-robot interaction |
0.2 | 1 | 2014 | Behavioral accommodation towards a dance robot tutor · HRI 2014 |
Human-robot interaction › educational robotics
social robot tutoring |
0.1 | 1 | 2014 | Behavioral accommodation towards a dance robot tutor · HRI 2014 |
Methods — techniques the papers use, named apart from their topics
novelty search · 0.6deep learning · 0.6behavior shaping · 0.6population-based divergent search · 0.4autoencoder · 0.4longitudinal study · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Discovering and Exploiting Sparse Rewards in a Learned Behavior SpaceabstractLearning optimal policies in sparse rewards settings is difficult as the learning agent has little to no feedback on the quality of its actions. In these situations, a good strategy is to focus on exploration, hopefully leading to the discovery of a reward signal to improve on. A learning algorithm capable of dealing with this kind of setting has to be able to (1) explore possible agent behaviors and (2) exploit any possible discovered reward. Exploration algorithms have been proposed that require the definition of a low-dimension behavior space, in which the behavior generated by the agent's policy can be represented. The need to design a priori this space such that it is worth exploring is a major limitation of these algorithms. In this work, we introduce STAX, an algorithm designed to learn a behavior space on-the-fly and to explore it while optimizing any reward discovered (see Figure 1). It does so by separating the exploration and learning of the behavior space from the exploitation of the reward through an alternating two-step process. In the first step, STAX builds a repertoire of diverse policies while learning a low-dimensional representation of the high-dimensional observations generated during the policies evaluation. In the exploitation step, emitters optimize the performance of the discovered rewarding solutions. Experiments conducted on three different sparse reward environments show that STAX performs comparably to existing baselines while requiring much less prior information about the task as it autonomously builds the behavior space it explores. Giuseppe Paolo, Miranda Coninx, Alban Laflaquière, Stéphane Doncieux |
Evol. Comput. | 2 |
| 2022 | Automatic Acquisition of a Repertoire of Diverse Grasping Trajectories through Behavior Shaping and Novelty SearchabstractGrasping a particular object may require a dedicated grasping movement that may also be specific to the robot end-effector. No generic and autonomous method does exist to generate these movements without making hypotheses on the robot or on the object. Learning methods could help to autonomously discover relevant grasping movements, but they face an important issue: grasping movements are so rare that a learning method based on exploration has little chance to ever observe an interesting movement, thus creating a bootstrap issue. We introduce an approach to generate diverse grasping movements in order to solve this problem. The movements are generated in simulation, for particular object positions. We test it on several simulated robots: Baxter, Pepper and a Kuka Iiwa arm. Although we show that generated movements actually work on a real Baxter robot, the aim is to use this method to create a large dataset to bootstrap deep learning methods. Aurélien Morel, Yakumo Kunimoto, Miranda Coninx, Stéphane Doncieux |
ICRA | 3 |
| 2021 | Sparse reward exploration via novelty search and emittersabstractReward-based optimization algorithms require both exploration, to find rewards, and exploitation, to maximize performance. The need for efficient exploration is even more significant in sparse reward settings, in which performance feedback is given sparingly, thus rendering it unsuitable for guiding the search process. In this work, we introduce the SparsE Reward Exploration via Novelty and Emitters (SERENE) algorithm, capable of efficiently exploring a search space, as well as optimizing rewards found in potentially disparate areas. Contrary to existing emitters-based approaches, SERENE separates the search space exploration and reward exploitation into two alternating processes. The first process performs exploration through Novelty Search, a divergent search algorithm. The second one exploits discovered reward areas through emitters, i.e. local instances of population-based optimization algorithms. A meta-scheduler allocates a global computational budget by alternating between the two processes, ensuring the discovery and efficient exploitation of disjoint reward areas. SERENE returns both a collection of diverse solutions covering the search space and a collection of high-performing solutions for each distinct reward area. We evaluate SERENE on various sparse reward environments and show it compares favorably to existing baselines. Giuseppe Paolo, Miranda Coninx, Stéphane Doncieux, Alban Laflaquière |
GECCO | 2 |
| 2021 | BR-NS: an archive-less approach to novelty searchabstractAs open-ended learning based on divergent search algorithms such as Novelty Search (NS) draws more and more attention from the research community, it is natural to expect that its application to increasingly complex real-world problems will require the exploration to operate in higher dimensional Behavior Spaces (BSs) which will not necessarily be Euclidean. Novelty Search traditionally relies on k-nearest neighbours search and an archive of previously visited behavior descriptors which are assumed to live in a Euclidean space. This is problematic because of a number of issues. On one hand, Euclidean distance and Nearest-neighbour search are known to behave differently and become less meaningful in high dimensional spaces. On the other hand, the archive has to be bounded since, memory considerations aside, the computational complexity of finding nearest neighbours in that archive grows linearithmically with its size. A sub-optimal bound can result in "cycling" in the behavior space, which inhibits the progress of the exploration. Furthermore, the performance of NS depends on a number of algorithmic choices and hyperparameters, such as the strategies to add or remove elements to the archive and the number of neighbours to use in k-nn search. In this paper, we discuss an alternative approach to novelty estimation, dubbed Behavior Recognition based Novelty Search (BR-NS), which does not require an archive, makes no assumption on the metrics that can be defined in the behavior space and does not rely on nearest neighbours search. We conduct experiments to gain insight into its feasibility and dynamics as well as potential advantages over archive-based NS in terms of time complexity. Achkan Salehi, Miranda Coninx, Stéphane Doncieux |
GECCO | 2 |
| 2020 | Novelty search makes evolvability inevitableabstractEvolvability is an important feature that impacts the ability of evolutionary processes to find interesting novel solutions and to deal with changing conditions of the problem to solve. The estimation of evolvability is not straight-forward and is generally too expensive to be directly used as selective pressure in the evolutionary process. Indirectly promoting evolvability as a side effect of other easier and faster to compute selection pressures would thus be advantageous. In an unbounded behavior space, it has already been shown that evolvable individuals naturally appear and tend to be selected as they are more likely to invade empty behavior niches. Evolvability is thus a natural byproduct of the search in this context. However, practical agents and environments often impose limits on the reachable behavior space. How do these boundaries impact evolvability? In this context, can evolvability still be promoted without explicitly rewarding it? We show that Novelty Search implicitly creates a pressure for high evolvability even in bounded behavior spaces, and explore the reasons for such a behavior. More precisely we show that, throughout the search, the dynamic evaluation of novelty rewards individuals which are very mobile in the behavior space, which in turn promotes evolvability. Stéphane Doncieux, Giuseppe Paolo, Alban Laflaquière, Miranda Coninx |
GECCO | 4 |
| 2020 | Unsupervised Learning and Exploration of Reachable Outcome SpaceabstractPerforming Reinforcement Learning in sparse rewards settings, with very little prior knowledge, is a challenging problem since there is no signal to properly guide the learning process. In such situations, a good search strategy is fundamental. At the same time, not having to adapt the algorithm to every single problem is very desirable. Here we introduce TAXONS, a Task Agnostic eXploration of Outcome spaces through Novelty and Surprise algorithm. Based on a population-based divergent-search approach, it learns a set of diverse policies directly from high-dimensional observations, without any task-specific information. TAXONS builds a repertoire of policies while training an autoencoder on the high-dimensional observation of the final state of the system to build a low-dimensional outcome space. The learned outcome space, combined with the reconstruction error, is used to drive the search for new policies. Results show that TAXONS can find a diverse set of controllers, covering a good part of the ground-truth outcome space, while having no information about such space. Giuseppe Paolo, Alban Laflaquière, Miranda Coninx, Stéphane Doncieux |
ICRA | 3 |
| 2019 | Novelty search: a theoretical perspectiveabstractNovelty Search is an exploration algorithm driven by the novelty of a behavior. The same individual evaluated at different generations has different fitness values. The corresponding fitness landscape is thus constantly changing and if, at the scale of a single generation, the metaphor of a fitness landscape with peaks and valleys still holds, this is not the case anymore at the scale of the whole evolutionary process. How does this kind of algorithms behave? Is it possible to define a model that would help understand how it works? This understanding is critical to analyse existing Novelty Search variants and design new and potentially more efficient ones. We assert that Novelty Search asymptotically behaves like a uniform random search process in the behavior space. This is an interesting feature, as it is not possible to directly sample in this space: the algorithm has a direct access to the genotype space only, whose relationship to the behavior space is complex. We describe the model and check its consistency on a classical Novelty Search experiment. We also show that it sheds a new light on results of the literature and suggests future research work. Stéphane Doncieux, Alban Laflaquière, Miranda Coninx |
GECCO | 3 |
| 2017 | Semantic-based interaction for teaching robot behavior compositionsabstractAllowing humans to teach robot behaviors will facilitate acceptability as well as long-term interactions. Humans would mainly use speech to transfer knowledge or to teach highlevel behaviors. In this paper, we propose a proof-of-concept application allowing a Pepper robot to learn behaviors from their natural-language-based description, provided by naive human users. In our model, natural language input is provided by grammar-free speech recognition, and is then processed to produce semantic knowledge, grounded in language and primitive behaviors. The same semantic knowledge is used to represent any kind of perceived input as well as actions the robot can perform. The experiment shows that the system can work independently from the domain of application, but also that it has limitations. Progress in semantic extraction, behavior planning and interaction scenario could stretch these limits. Victor Paleologue, Jocelyn Martin, Amit Kumar Pandey, Miranda Coninx, Mohamed Chetouani |
RO-MAN | 4 |
| 2017 | Quick and energy-efficient Bayesian computing of binocular disparity using stochastic digital signals
Miranda Coninx, Pierre Bessière, Jacques Droulez |
Int. J. Approx. Reason. | 1 |
| 2016 | Towards long-term social child-robot interaction: using multi-activity switching to engage young usersabstractSocial robots have the potential to provide support in a number of practical domains, such as learning and behaviour change. This potential is particularly relevant for children, who have proven receptive to interactions with social robots. To reach learning and therapeutic goals, a number of issues need to be investigated, notably the design of an effective child-robot interaction (cHRI) to ensure the child remains engaged in the relationship and that educational goals are met. Typically, current cHRI research experiments focus on a single type of interaction activity (e.g. a game). However, these can suffer from a lack of adaptation to the child, or from an increasingly repetitive nature of the activity and interaction. In this paper, we motivate and propose a practicable solution to this issue: an adaptive robot able to switch between multiple activities within single interactions. We describe a system that embodies this idea, and present a case study in which diabetic children collaboratively learn with the robot about various aspects of managing their condition. We demonstrate the ability of our system to induce a varied interaction and show the potential of this approach both as an educational tool and as a research method for long-term cHRI. Miranda Coninx, Paul Baxter 0001, Elettra Oleari, Sara Bellini, Bert P. B. Bierman, Olivier A. Blanson Henkemans, Lola Cañamero, Piero Cosi, Valentin Enescu, Raquel Ros, Antoine Hiolle, Rémi Humbert, Bernd Kiefer, Ivana Kruijff-Korbayová, Rosemarijn Looije, Marco Mosconi, Mark A. Neerincx, Giulio Paci, Yorgos Patsis, Clara Pozzi, Francesca Sacchitelli, Hichem Sahli, Alberto Sanna, Giacomo Sommavilla, Fabio Tesser, Yiannis Demiris, Tony Belpaeme |
J. Hum. Robot Interact. | 1 |
| 2014 | Behavioral accommodation towards a dance robot tutorabstractWe report first results on children adaptive behavior towards a dance tutoring robot. We can observe that children behavior rapidly evolves through few sessions in order to accommodate with the robotic tutor rhythm and instructions. Raquel Ros, Miranda Coninx, Yiannis Demiris, Yorgos Patsis, Valentin Enescu, Hichem Sahli |
HRI | 2 |