EDBT 2026 Demo / reviewers in the wild / expert
David Filliat
dblp:13/5289
· DBLP profile ↗
47ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-5739-1618ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 3 first-author · 14 since 2021Systems, architecture and hardware · 15 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking XAI Explanations with Human-Aligned EvaluationsabstractWe introduce PASTA (Perceptual Assessment System for explanaTion of Artificial Intelligence), a novel human-centric framework for evaluating eXplainable AI (XAI) techniques in computer vision. Our first contribution is the creation of the PASTA-dataset, the first large-scale benchmark that spans a diverse set of models and both saliency-based and concept-based explanation methods. This dataset enables robust, comparative analysis of XAI techniques based on human judgment. Our second contribution is an automated, data-driven benchmark that predicts human preferences using the PASTA-dataset. This scoring called PASTA-score method offers scalable, reliable, and consistent evaluation aligned with human perception. Additionally, our benchmark allows for comparisons between explanations across different modalities, an aspect previously unaddressed. We then propose to apply our scoring method to probe the interpretability of existing models and to build more human interpretable XAI methods. Rémi Kazmierczak, Steve Azzolin, Eloïse Berthier, Anna Hedström, Patricia Delhomme, David Filliat, Nicolas Bousquet 0001, Goran Frehse, Massimiliano Mancini, Baptiste Caramiaux, Andrea Passerini, Gianni Franchi |
AAAI | 6 |
| 2025 | Improved Monocular Depth Prediction Using Distance Transform Over Pre-semantic Contours with Self-supervised Neural NetworksabstractMonocular depth estimation (MDE) with self-supervised training approaches struggles in low-texture areas, where photometric losses may lead to ambiguous depth predictions. To address this, we propose a novel technique that enhances spatial information by applying a distance transform over pre-semantic contours, augmenting discriminative power in low texture regions. Our approach jointly estimates pre-semantic contours, depth and ego-motion. The pre-semantic contours are leveraged to produce new input images, with variance augmented by the distance transform in uniform areas. This approach results in more effective loss functions, enhancing the training process for depth and ego-motion. We demonstrate theoretically that the distance transform is the optimal variance-augmenting technique in this context. Through extensive experiments on KITTI, Cityscapes, Waymo, NYUv2 and ScanNet our model demonstrates robust performance, surpassing competing self-supervised methods in MDE. Marwane Hariat, Antoine Manzanera, David Filliat |
CVPR | 3 |
| 2025 | NAMOUnc: Navigation Among Movable Obstacles with Decision Making on Uncertainty IntervalabstractInternational audience Eric Lucet, Julien Alexandre Dit Sandretto, Shoubin Chen, David Filliat |
ICINCO (2) | 5 |
| 2025 | A Simple yet Effective Test-Time Adaptation for Zero-Shot Monocular Metric Depth EstimationabstractThe recent development of foundation models for monocular depth estimation such as Depth Anything paved the way to zero-shot monocular depth estimation. Since it returns an affine-invariant disparity map, the favored technique to recover the metric depth consists in fine-tuning the model. However, this stage is not straightforward, it can be costly and time-consuming because of the training and the creation of the dataset. The latter must contain images captured by the camera that will be used at test time and the corresponding ground truth. Moreover, the fine-tuning may also degrade the generalizing capacity of the original model. Instead, we propose in this paper a new method to rescale Depth Anything predictions using 3D points provided by sensors or techniques such as low-resolution LiDAR or structure-from-motion with poses given by an IMU. This approach avoids fine-tuning and preserves the generalizing power of the original depth estimation model while being robust to the noise of the sparse depth, of the camera-LiDAR calibration or of the depth model. Our experiments highlight enhancements relative to zero-shot monocular metric depth estimation methods, competitive results compared to fine-tuned approaches and a better robustness than depth completion approaches. Code available at github.com/ENSTA-U2IS-AI/depth-rescaling. Rémi Marsal, Alexandre Chapoutot, Philippe Xu, David Filliat |
IROS | 4 |
| 2025 | Hierarchical Light Transformer Ensembles for Multimodal Trajectory ForecastingabstractAccurate trajectory forecasting is crucial for the performance of various systems, such as advanced driver-assistance systems and self-driving vehicles. These forecasts allow us to anticipate events that lead to collisions and, therefore, to mitigate them. Deep Neural Networks have excelled in motion forecasting, but overconfidence and weak uncertainty quantification persist. Deep Ensembles address these concerns, yet applying them to multimodal distributions remains challenging. In this paper, we propose a novel approach named Hierarchical Light Transformer Ensembles (HLT-Ens) aimed at efficiently training an ensemble of Transformer architectures using a novel hierarchical loss function. HLT-Ens leverages grouped fully connected layers, inspired by grouped convolution techniques, to capture multimodal distributions effectively. We demonstrate that HLT-Ens achieves state-of-the-art performance levels through extensive experimentation, offering a promising avenue for improving trajectory forecasting techniques. We make our code available at github.com/alafage/hlt-ens. Adrien Lafage, Mathieu Barbier, Gianni Franchi, David Filliat |
WACV | 4 |
| 2025 | Robust trajectory forecasting in autonomous systems using mixtures of Student's T-distributions with T-DistNet
Adrien Lafage, Gianni Franchi, Mathieu Barbier, David Filliat |
Pattern Recognit. | 4 |
| 2024 | On Double Descent in Reinforcement Learning with LSTD and Random FeaturesabstractTemporal Difference (TD) algorithms are widely used in Deep Reinforcement Learning (RL). Their performance is heavily influenced by the size of the neural network. While in supervised learning, the regime of over-parameterization and its benefits are well understood, the situation in RL is much less clear. In this paper, we present a theoretical analysis of the influence of network size and $l_2$-regularization on performance. We identify the ratio between the number of parameters and the number of visited states as a crucial factor and define over-parameterization as the regime when it is larger than one. Furthermore, we observe a double descent phenomenon, i.e., a sudden drop in performance around the parameter/state ratio of one. Leveraging random features and the lazy training regime, we study the regularized Least-Square Temporal Difference (LSTD) algorithm in an asymptotic regime, as both the number of parameters and states go to infinity, maintaining a constant ratio. We derive deterministic limits of both the empirical and the true Mean-Squared Bellman Error (MSBE) that feature correction terms responsible for the double descent. Correction terms vanish when the $l_2$-regularization is increased or the number of unvisited states goes to zero. Numerical experiments with synthetic and small real-world environments closely match the theoretical predictions. David Brellmann, Eloïse Berthier, David Filliat, Goran Frehse |
ICLR | 3 |
| 2024 | A probabilistic approach for learning and adapting shared control skills with the human in the loopabstractAssistive robots promise to be of great help to wheelchair users with motor impairments, for example for activities of daily living. Using shared control to provide task-specific assistance – for instance with the Shared Control Templates (SCT) framework – facilitates user control, even with low-dimensional input signals. However, designing SCTs is a laborious task requiring robotic expertise. To facilitate their design, we propose a method to learn one of their core components – active constraints – from demonstrated end-effector trajectories. We use a probabilistic model, Kernelized Movement Primitives, which additionally allows adaptation from user commands to improve the shared control skills, during both design and execution. We demonstrate that the SCTs so acquired can be successfully used to pick up an object, as well as adjusted for new environmental constraints, with our assistive robot EDAN. Gabriel Quere, Freek Stulp, David Filliat, João Silvério |
ICRA | 3 |
| 2024 | Open the Chests: An Environment for Activity Recognition and Sequential Decision Problems Using Temporal LogicabstractInternational audience Ivelina Stoyanova, Nicolas Museux, Sao Mai Nguyen, David Filliat |
TIME | 4 |
| 2024 | InfraParis: A multi-modal and multi-task autonomous driving datasetabstractCurrent deep neural networks (DNNs) for autonomous driving computer vision are typically trained on specific datasets that only involve a single type of data and urban scenes. Consequently, these models struggle to handle new objects, noise, nighttime conditions, and diverse scenarios, which is essential for safety-critical applications. Despite ongoing efforts to enhance the resilience of computer vision DNNs, progress has been sluggish, partly due to the absence of benchmarks featuring multiple modalities. We introduce a novel and versatile dataset named InfraParis that supports multiple tasks across three modalities: RGB, depth, and infrared. We assess various state-of-the-art baseline techniques, encompassing models for the tasks of semantic segmentation, object detection, and depth estimation. More visualizations and the download link for InfraParis are available at https://enstau2is.github.io/infraParis/. Gianni Franchi, Marwane Hariat, Xuanlong Yu, Nacim Belkhir, Antoine Manzanera, David Filliat |
WACV | 6 |
| 2023 | Navigation Among Movable Obstacles Using Machine Learning Based Total Time Cost OptimizationabstractMost navigation approaches treat obstacles as static objects and choose to bypass them. However, the detour could be costly or could lead to failures in indoor environments. The recently developed navigation among movable obstacles (NAMO) methods prefer to remove all the movable obstacles blocking the way, which might be not the best choice when planning and moving obstacles takes a long time. We propose a pipeline where the robot solves the NAMO problems by optimizing the total time to reach the goal. This is achieved by a supervised learning approach that can predict the time of planning and performing obstacle motion before actually doing it if this leads to faster goal reaching. Besides, a pose generator based on reinforcement learning is proposed to decide where the robot can move the obstacle. The method is evaluated in two kinds of simulation environments and the results demonstrate its advantages compared to the classical bypass and obstacle removal strategies. Eric Lucet, Julien Alexandre Dit Sandretto, David Filliat |
IROS | 4 |
| 2023 | Rebalancing gradient to improve self-supervised co-training of depth, odometry and optical flow predictionsabstractWe present CoopNet, an approach that improves the co-operation of co-trained networks by dynamically adapting the apportionment of gradient, to ensure equitable learning progress. It is applied to motion-aware self-supervised prediction of depth maps, by introducing a new hybrid loss, based on a distribution model of photo-metric reconstruction errors made by, on the one hand the depth + odometry paired networks, and on the other hand the optical flow network. This model essentially assumes that the pixels from moving objects (that must be discarded for training depth and odometry), correspond to those where the two reconstructions strongly disagree. We justify this model by theoretical considerations and experimental evidences. A comparative evaluation on KITTI and CityScapes datasets shows that CoopNet improves or is comparable to the state-of-the-art in depth, odometry and optical flow predictions. Our code is available here: https://github.com/mhariat/CoopNet. Marwane Hariat, Antoine Manzanera, David Filliat |
WACV | 3 |
| 2022 | MUAD: Multiple Uncertainties for Autonomous Driving, a benchmark for multiple uncertainty types and tasks
Gianni Franchi, Xuanlong Yu, Andrei Bursuc, Ángel Tena, Rémi Kazmierczak, Séverine Dubuisson, Emanuel Aldea, David Filliat |
BMVC | 8 |
| 2022 | Latent Discriminant Deterministic Uncertainty
Gianni Franchi, Xuanlong Yu, Andrei Bursuc, Emanuel Aldea, Séverine Dubuisson, David Filliat |
ECCV (12) | 6 |
| 2022 | Task and Motion Planning Methods: Applications and LimitationsabstractInternational audience Eric Lucet, Julien Alexandre Dit Sandretto, Selma Kchir, David Filliat |
ICINCO | 5 |
| 2021 | S-TRIGGER: Continual State Representation Learning via Self-Triggered Generative ReplayabstractWe consider the problem of building a state representation model for control, in a continual learning setting. As the environment changes, the aim is to efficiently compress the sensory state information without losing past knowledge, and then use Reinforcement Learning on the resulting features for efficient policy learning. To this end, we propose S-TRIGGER, a general method for Continual State Representation Learning applicable to Variational Auto-Encoders and its many variants. The method is based on Generative Replay, i.e. the use of generated samples to maintain past knowledge. It comes along with a statistically sound method for environment change detection, which self-triggers the Generative Replay. Our experiments on VAEs show that S-TRIGGER learns state representations that allows fast and high-performing Reinforcement Learning, while avoiding catastrophic forgetting. The resulting system has a bounded size and is capable of autonomously learning new information without using past data. Hugo Caselles-Dupré, Michaël Garcia Ortiz, David Filliat |
IJCNN | 3 |
| 2021 | Evaluating Robustness over High Level Driving Instruction for Autonomous DrivingabstractIn recent years, we have witnessed increasingly high performance in the field of autonomous end-to-end driving. In particular, more and more research is being done on driving in urban environments, where the car has to follow high level commands to navigate. However, few evaluations are made on the ability of these agents to react in an unexpected situation. Specifically, no evaluations are conducted on the robustness of driving agents in the event of a bad high-level command. We propose here an evaluation method, namely a benchmark that allows to assess the robustness of an agent, and to appreciate its understanding of the environment through its ability to keep a safe behavior, regardless of the instruction. Florence Carton, David Filliat, Jaonary Rabarisoa, Quoc Cuong Pham |
IV | 2 |
| 2020 | Trajectory Prediction of Traffic Agents: Incorporating context into machine learning approachesabstractFor a vehicle to navigate autonomously, it needs to perceive its surroundings and estimate the future state of the relevant traffic-agents with which it might interact as it navigates across public road networks. Predicting the future state of the perceived entities is a challenge, as these might appear to move in a stochastic manner. However, their motion is constrained to an extent by context, in particular the road network structure. Conventional machine learning methods are mainly trained using data from the perceived entities without considering roads, as a result trajectory prediction is difficult. In this paper, the notion of maps representing the road structure are included into the machine learning process. For this purpose, 3D LiDAR points and maps in the form of binary masks are used. These are used on a recurrent artificial neural network, the LSTM encoder-decoder based architecture to predict the motion of the interacting traffic agents. A comparison between the proposed solution with one that is only sensor driven (LiDAR) is included. For this purpose, NuScenes dataset is utilised, that includes annotated 3D point clouds. The results have demonstrated the importance of context to enhance our prediction performance as well as the capability of our machine learning framework to incorporate map information. Vyshakh Palli-Thazha, David Filliat, Javier Ibañez-Guzmán |
VTC Spring | 2 |
| 2019 | Marginal Replay vs Conditional Replay for Continual Learning
Timothée Lesort, Alexander Gepperth, Andrei Stoian, David Filliat |
ICANN (2) | 4 |
| 2019 | Training Discriminative Models to Evaluate Generative Ones
Timothée Lesort, Andrei Stoian, Jean-François Goudou, David Filliat |
ICANN (3) | 4 |
| 2019 | Generative Models from the perspective of Continual LearningabstractWhich generative model is the most suitable for Continual Learning? This paper aims at evaluating and comparing generative models on disjoint sequential image generation tasks. We investigate how several models learn and forget, considering various strategies: rehearsal, regularization, generative replay and fine-tuning. We used two quantitative metrics to estimate the generation quality and memory ability. We experiment with sequential tasks on three commonly used benchmarks for Continual Learning (MNIST, Fashion MNIST and CIFAR10). We found that among all models, the original GAN performs best and among Continual Learning strategies, generative replay outperforms all other methods. Even if we found satisfactory combinations on MNIST and Fashion MNIST, training generative models sequentially on CIFAR10 is particularly instable, and remains a challenge. Our code is available online1. Timothée Lesort, Hugo Caselles-Dupré, Michaël Garcia Ortiz, Andrei Stoian, David Filliat |
IJCNN | 5 |
| 2019 | Deep unsupervised state representation learning with robotic priors: a robustness analysisabstractOur understanding of the world depends highly on our capacity to produce intuitive and simplified representations which can be easily used to solve problems. We reproduce this simplification process using a neural network to build a low dimensional state representation of the world from images acquired by a robot. As in Jonschkowski et al. 2015, we learn in an unsupervised way using prior knowledge about the world as loss functions called robotic priors and extend this approach to high dimension richer images to learn a 3D representation of the hand position of a robot from RGB images. We propose a quantitative evaluation metric of the learned representation that uses nearest neighbors in the state space and allows to assess its quality and show both the potential and limitations of robotic priors in realistic environments. We augment image size, add distractors and domain randomization, all crucial components to achieve transfer learning to real robots. Finally, we also contribute a new prior to improve the robustness of the representation. The applications of such low dimensional state representation range from easing reinforcement learning (RL) and knowledge transfer across tasks, to facilitating learning from raw data with more efficient and compact high level representations. The results show that the robotic prior approach is able to extract high level representation as the 3D position of an arm and organize it into a compact and coherent space of states in a challenging dataset. Timothée Lesort, Mathieu Seurin, Natalia Díaz Rodríguez, David Filliat |
IJCNN | 5 |
| 2019 | Symmetry-Based Disentangled Representation Learning requires Interaction with EnvironmentsabstractFinding a generally accepted formal definition of a disentangled representation in the context of an agent behaving in an environment is an important challenge towards the construction of data-efficient autonomous agents. Higgins et al. recently proposed Symmetry-Based Disentangled Representation Learning, a definition based on a characterization of symmetries in the environment using group theory. We build on their work and make observations, theoretical and empirical, that lead us to argue that Symmetry-Based Disentangled Representation Learning cannot only be based on static observations: agents should interact with the environment to discover its symmetries. Our experiments can be reproduced in Colab and the code is available on GitHub. Hugo Caselles-Dupré, Michaël Garcia Ortiz, David Filliat |
NeurIPS | 3 |
| 2018 | Experimental Validation of a Multirobot Distributed Receding Horizon Motion Planning ApproachabstractThis paper addresses the problem of motion planning for a multirobot system in a partially known environment where conditions such as uncertainty about robots' positions and communication delays are real. In particular, we detail the use of a Distributed Receding Horizon Approach that guarantees collision avoidance with static obstacles and between robots communicating with each other. Underlying optimization problems are solved by using a Sequential Least Squares Programming algorithm. Experiments with real nonholonomic mobile platforms are performed. The proposed framework is compared with the Dynamic Window approach to motion planning in a single robot setup. A second experiment shows results for a multirobot case using two robots where collision is avoided even in presence of significant localization uncertainties. José M. Mendes Filho, Eric Lucet, David Filliat |
ICARCV | 3 |
| 2018 | State representation learning for control: An overview
Timothée Lesort, Natalia Díaz Rodríguez, Jean-François Goudou, David Filliat |
Neural Networks | 4 |
| 2017 | Real-time distributed receding horizon motion planning and control for mobile multi-robot dynamic systemsabstractThis paper proposes an improvement of a motion planning approach and a modified model predictive control (MPC) for solving the navigation problem of a team of dynamical wheeled mobile robots in the presence of obstacles in a realistic environment. Planning is performed by a distributed receding horizon algorithm where constrained optimization problems are numerically solved for each prediction time-horizon. This approach allows distributed motion planning for a multi-robot system with asynchronous communication while avoiding collisions and minimizing the travel time of each robot. However, the robots dynamics prevents the planned motion to be applied directly to the robots. Using unicycle-like vehicles in a dynamic simulation, we show that deviations from the planned motion caused by the robots dynamics can be overcome by modifying the optimization problem underlying the planning algorithm and by adding an MPC for trajectory tracking. Results also indicate that this approach can be used in systems subjected to real-time constraint. José M. Mendes Filho, Eric Lucet, David Filliat |
ICRA | 3 |
| 2016 | Environment exploration for object-based visual saliency learningabstractSearching for objects in an indoor environment can be drastically improved if a task-specific visual saliency is available. We describe a method to incrementally learn such an object-based visual saliency directly on a robot, using an environment exploration mechanism. We first define saliency based on a geometrical criterion and use this definition to segment salient elements given an attentive but costly and restrictive observation of the environment. These elements are used to train a fast classifier that predicts salient objects given large-scale visual features. In order to get a better and faster learning, we use an exploration strategy based on intrinsic motivation to drive our displacement in order to get relevant observations. Our approach has been tested on a robot in indoor environments as well as on publicly available RGB-D images sequences. We demonstrate that the approach outperforms several state-of-the-art methods in the case of indoor object detection and that the exploration strategy can drastically decrease the time required for learning saliency. Céline Craye, David Filliat, Jean-François Goudou |
ICRA | 2 |
| 2016 | RL-IAC: An exploration policy for online saliency learning on an autonomous mobile robotabstractIn the context of visual object search and localization, saliency maps provide an efficient way to find object candidates in images. Unlike most approaches, we propose a way to learn saliency maps directly on a robot, by exploring the environment, discovering salient objects using geometric cues, and learning their visual aspects. More importantly, we provide an autonomous exploration strategy able to drive the robot for the task of learning saliency. For that, we describe the Reinforcement Learning-Intelligent Adaptive Curiosity algorithm (RL-IAC), a mechanism based on IAC (Intelligent Adaptive Curiosity) able to guide the robot through areas of the space where learning progress is high, while minimizing the time spent to move in its environment without learning. We demonstrate first that our saliency approach is an efficient tool to generate relevant object boxes proposal in the input image and significantly outperforms the state-of-the-art EdgeBoxes algorithm. Second, we show that RL-IAC can drastically decrease the required time for learning saliency compared to random exploration. Céline Craye, David Filliat, Jean-François Goudou |
IROS | 2 |
| 2016 | A Bayesian framework for preventive assistance at road intersectionsabstractModern vehicles embed an increasing number of Advanced Driving Assistance Systems (ADAS). Whilst such systems showed their capability to improve comfort and safety, most of them provide assistance only as a last resort, that is, they alert the driver or trigger automatic braking only when collision is imminent. This limitation is mainly due to the difficulty to accurately anticipate risk situations in order to provide the driver with preventive assistance, i.e. assistance allowing for comfortable reaction. This paper presents a Bayesian framework which aims to detect risk situations sufficiently early to trigger conventional curative assistance as well as preventive assistance. By taking into consideration the context, the vehicle state, the driver actuation and the manner how the driver usually negotiates given situations, the framework allows to infer which type of assistance is the most pertinent to be provided to the driver. The principles of this framework are applied to a fundamental case study, the arrival to a stop intersection. Results obtained from data recorded under controlled conditions are presented. They show that the framework allows to coherently detect risk situations and to identify what assistance, including preventive assistance, is the most appropriate for the situation. Alexandre Armand, David Filliat, Javier Ibañez-Guzmán |
Intelligent Vehicles Symposium | 2 |
| 2015 | Asynchronous Event-Based Multikernel Algorithm for High-Speed Visual Features TrackingabstractThis paper presents a number of new methods for visual tracking using the output of an event-based asynchronous neuromorphic dynamic vision sensor. It allows the tracking of multiple visual features in real time, achieving an update rate of several hundred kilohertz on a standard desktop PC. The approach has been specially adapted to take advantage of the event-driven properties of these sensors by combining both spatial and temporal correlations of events in an asynchronous iterative framework. Various kernels, such as Gaussian, Gabor, combinations of Gabor functions, and arbitrary user-defined kernels, are used to track features from incoming events. The trackers described in this paper are capable of handling variations in position, scale, and orientation through the use of multiple pools of trackers. This approach avoids the N(2) operations per event associated with conventional kernel-based convolution operations with N × N kernels. The tracking performance was evaluated experimentally for each type of kernel in order to demonstrate the robustness of the proposed solution. Xavier Lagorce, Cedric Meyer, Sio-Hoi Ieng, David Filliat, Ryad Benosman |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2014 | Neural network based 2D/3D fusion for robotic object recognition
Louis-Charles Caron, David Filliat, Alexander Gepperth |
ESANN | 3 |
| 2014 | Unsupervised and online non-stationary obstacle discovery and modeling using a laser range finderabstractUsing laser range finders has shown its efficiency to perform mapping and navigation for mobile robots. However, most of existing methods assume a mostly static world and filter away dynamic aspects while those dynamic aspects are often caused by non-stationary objects which may be important for the robot task. We propose an approach that makes it possible to detect, learn and recognize these objects through a multi-view model, using only a planar laser range finder. We show using a supervised approach that despite the limited information provided by the sensor, it is possible to recognize efficiently up to 22 different object, with a low computing cost while taking advantage of the large field of view of the sensor. We also propose an online, incremental and unsupervised approach that make it possible to continuously discover and learn all kind of dynamic elements encountered by the robot including people and objects. Guillaume Duceux, David Filliat |
IROS | 2 |
| 2014 | Ontology-based context awareness for driving assistance systemsabstractWithin a vehicle driving space, different entities such as vehicles and vulnerable road users are in constant interaction which governs their behaviour. Whilst smart sensors provide information about the state of the perceived objects, considering the spatio-temporal relationships between them with respect to the subject vehicle remains a challenge. This paper proposes to fill this gap by using contextual information to infer how perceived entities are expected to behave, and thus what are the consequences of these behaviours on the subject vehicle. For this purpose, an ontology is formulated about the vehicle, perceived entities and context (map information) to provide a conceptual description of all road entities with their interaction. It allows for inferences of knowledge about the situation of the subject vehicle with respect to the environment in which it is navigating. The framework is applied to the navigation of a vehicle as it approaches road intersections, to demonstrate its applicability. Results from the real-time implementation on a vehicle operating under controlled conditions are included. They show that the proposed ontology allows for a coherent understanding of the interactions between the perceived entities and contextual data. Further, it can be used to improve the situation awareness of an ADAS (Advanced Driving Assistance System), by determining which entities are the most relevant for the subject vehicle navigation. Alexandre Armand, David Filliat, Javier Ibañez-Guzmán |
Intelligent Vehicles Symposium | 2 |
| 2014 | A framework for proactive assistance: SummaryabstractAdvanced Driving Assistance Systems usually provide assistance to drivers only once a high risk situation has been detected. Indeed, it is difficult for an embedded system to understand driving situations, and to predict early enough that it is to become uncomfortable or dangerous. Most of ADAS work assume that interactions between road entities do not exist (or are limited), and that all drivers react in the same manner in similar conditions. We propose a framework that enables to fill these gaps. On one hand, an ontology which is a conceptual description of entities present in driving spaces is used to understand how all the perceived entities interact together with the subject vehicle, and govern its behavior. On the other hand, a dynamic Bayesian Network enables to estimate the driver situation awareness with regard to the perceived objects, based on the ontology inferences, map information, driver actuation and driving style. Alexandre Armand, David Filliat, Javier Ibañez-Guzmán |
SMC | 2 |
| 2013 | Appearance-based segmentation of indoors/outdoors sequences of spherical viewsabstractNavigating in large scale, complex and dynamic environments requires reliable representations able to capture metric, topological and semantic aspects of the scene for supporting path planing and real time motion control. In a previous work [11], we addressed metric and topological representations thanks to a multi-cameras system which allows building of dense visual maps of large scale 3D environments. The map is a set of locally accurate spherical panoramas related by 6d of poses graph. The work presented here is a further step toward a semantic representation. We aim at detecting the changes in the structural properties of the scene during navigation. Structural properties are estimated online using a global descriptor relying on spherical harmonics which are particularly well-fitted to capture properties in spherical views. A change-point detection algorithm based on a statistical Neyman-Pearson test allows us to find optimal transitions between topological places. Results are presented and discussed both for indoors and outdoors experiments. Alexandre Chapoulie, Patrick Rives, David Filliat |
IROS | 3 |
| 2013 | The Impact of Human-Robot Interfaces on the Learning of Visual ObjectsabstractThis paper studies the impact of interfaces, allowing nonexpert users to efficiently and intuitively teach a robot to recognize new visual objects. We present challenges that need to be addressed for real-world deployment of robots capable of learning new visual objects in interaction with everyday users. We argue that in addition to robust machine learning and computer vision methods, well-designed interfaces are crucial for learning efficiency. In particular, we argue that interfaces can be key in helping nonexpert users to collect good learning examples and, thus, improve the performance of the overall learning system. Then, we present four alternative human-robot interfaces: Three are based on the use of a mediating artifact (smartphone, wiimote, wiimote and laser), and one is based on natural human gestures (with a Wizard-of-Oz recognition system). These interfaces mainly vary in the kind of feedback provided to the user, allowing him to understand more or less easily what the robot is perceiving and, thus, guide his way of providing training examples differently. We then evaluate the impact of these interfaces, in terms of learning efficiency, usability, and user's experience, through a real world and large-scale user study. In this experiment, we asked participants to teach a robot 12 different new visual objects in the context of a robotic game. This game happens in a home-like environment and was designed to motivate and engage users in an interaction where using the system was meaningful. We then discuss results that show significant differences among interfaces. In particular, we show that interfaces such as the smartphone interface allows nonexpert users to intuitively provide much better training examples to the robot, which is almost as good as expert users who are trained for this task and are aware of the different visual perception and machine learning issues. We also show that artifact-mediated teaching is significantly more efficient for robot learning, and equally good in terms of usability and user's experience, than teaching thanks to a gesture-based human-like interaction. Pierre Rouanet, Pierre-Yves Oudeyer, Fabien Danieau, David Filliat |
IEEE Trans. Robotics | 4 |
| 2012 | Developmental approach for interactive object discoveryabstractWe present a visual system for a humanoid robot that supports an efficient online learning and recognition of various elements of the environment. Taking inspiration from child's perception and following the principles of developmental robotics, our algorithm does not require image databases, predefined objects nor face/skin detectors. The robot explores the visual space from interactions with people and its own experiments. The object detection is based on the hypothesis of coherent motion and appearance during manipulations. A hierarchical object representation is constructed from SURF points and color of superpixels that are grouped in local geometric structures and form the basis of a multiple-view object model. The learning algorithm accumulates the statistics of feature occurrences and identifies objects using a maximum likelihood approach and temporal coherency. The proposed visual system is implemented on the iCub robot and shows 85% average recognition rate for 10 objects after 30 minutes of interaction. Natalia Lyubova, David Filliat |
IJCNN | 2 |
| 2012 | Topological segmentation of indoors/outdoors sequences of spherical viewsabstractTopological navigation consists for a robot in navigating in a topological graph which nodes are topological places. Either for indoor or outdoor environments, segmentation into topological places is a challenging issue. In this paper, we propose a common approach for indoor and outdoor environment segmentation without elaborating a complete topological navigation system. The approach is novel in that environment sensing is performed using spherical images. Environment structure estimation is performed by a global structure descriptor specially adapted to the spherical representation. This descriptor is processed by a custom designed algorithm which detects change-points defining the segmentation between topological places. Alexandre Chapoulie, Patrick Rives, David Filliat |
IROS | 3 |
| 2011 | Incremental topo-metric SLAM using vision and robot odometryabstractWe address the problem of simultaneous localization and mapping by combining visual loop-closure detection with metrical information given by the robot odometry. The proposed algorithm builds in real-time topo-metric maps of an unknown environment, with a monocular or omnidirectional camera and odometry gathered by motors encoders. A dedicated improved version of our previous work on purely appearance-based loop-closure detection [1] is used to extract potential loop-closure locations. Potential locations are then verified and classified using a new validation stage. The main contributions we bring are the generalization of the validation method for the use of monocular and omnidirectional camera with the removal of the camera calibration stage, the inclusion of an odometry-based evolution model in the Bayesian filter which improves accuracy and responsiveness, and the addition of a consistent metric position estimation. This new SLAM method does not require any calibration or learning stage (i.e. no a priori information about environment). It is therefore fully incremental and generates maps usable for global localization and planned navigation. This algorithm is moreover well suited for remote processing and can be used on toy robots with very small computational power. Stéphane Bazeille, David Filliat |
ICRA | 2 |
| 2010 | A study of three interfaces allowing non-expert users to teach new visual objects to a robot and their impact on learning efficiencyabstractWe developed three interfaces to allow non-expert users to teach name for new visual objects and compare them through user's studies in term of learning efficiency. Pierre Rouanet, Pierre-Yves Oudeyer, David Filliat |
HRI | 3 |
| 2009 | Visual topological SLAM and global localizationabstractVisual localization and mapping for mobile robots has been achieved with a large variety of methods. Among them, topological navigation using vision has the advantage of offering a scalable representation, and of relying on a common and affordable sensor. In previous work, we developed such an incremental and real-time topological mapping and localization solution, without using any metrical information, and by relying on a Bayesian visual loop-closure detection algorithm. In this paper, we propose an extension of this work by integrating metrical information from robot odometry in the topological map, so as to obtain a globally consistent environment model. Also, we demonstrate the performance of our system on the global localization task, where the robot has to determine its position in a map acquired beforehand. Adrien Angeli, Stéphane Doncieux, Jean-Arcady Meyer, David Filliat |
ICRA | 4 |
| 2008 | Real-time visual loop-closure detectionabstractIn robotic applications of visual simultaneous localization and mapping, loop-closure detection and global localization are two issues that require the capacity to recognize a previously visited place from current camera measurements. We present an online method that makes it possible to detect when an image comes from an already perceived scene using local shape information. Our approach extends the bag of visual words method used in image recognition to incremental conditions and relies on Bayesian filtering to estimate loop-closure probability. We demonstrate the efficiency of our solution by real-time loop-closure detection under strong perceptual aliasing conditions in an indoor image sequence taken with a handheld camera. Adrien Angeli, Stéphane Doncieux, Jean-Arcady Meyer, David Filliat |
ICRA | 4 |
| 2008 | Incremental vision-based topological SLAMabstractIn robotics, appearance-based topological map building consists in infering the topology of the environment explored by a robot from its sensor measurements. In this paper, we propose a vision-based framework that considers this data association problem from a loop-closure detection perspective in order to correctly assign each measurement to its location. Our approach relies on the visual bag of words paradigm to represent the images and on a discrete Bayes filter to compute the probability of loop-closure. We demonstrate the efficiency of our solution by incremental and real-time consistent map building in an indoor environment and under strong perceptual aliasing conditions using a single monocular wide-angle camera. Adrien Angeli, Stéphane Doncieux, Jean-Arcady Meyer, David Filliat |
IROS | 4 |
| 2008 | Interactive learning of visual topological navigationabstractWe present a topological navigation system that is able to visually recognize the different rooms of an apartment and guide a robot between them. Specifically tailored for small entertainment robots, the system relies on vision only and learns its navigation capabilities incrementally by interacting with a user. This continuous learning strategy makes the system particularly adaptable to environmental lighting and structure modifications. From the computer vision point of view, the system uses a purely appearance-based image representation called bag of visual words, without any metric information. This representation was adapted to the incremental context of robotics and supplemented by active perception to enhance performances. Empirical validation on real robots and on the publicly available INDECS image database are presented. David Filliat |
IROS | 1 |
| 2008 | Fast and Incremental Method for Loop-Closure Detection Using Bags of Visual WordsabstractIn robotic applications of visual simultaneous localization and mapping techniques, loop-closure detection and global localization are two issues that require the capacity to recognize a previously visited place from current camera measurements. We present an online method that makes it possible to detect when an image comes from an already perceived scene using local shape and color information. Our approach extends the bag-of-words method used in image classification to incremental conditions and relies on Bayesian filtering to estimate loop-closure probability. We demonstrate the efficiency of our solution by real-time loop-closure detection under strong perceptual aliasing conditions in both indoor and outdoor image sequences taken with a handheld camera. Adrien Angeli, David Filliat, Stéphane Doncieux, Jean-Arcady Meyer |
IEEE Trans. Robotics | 2 |
| 2007 | A visual bag of words method for interactive qualitative localization and mappingabstractLocalization for low cost humanoid or animal-like personal robots has to rely on cheap sensors and has to be robust to user manipulations of the robot. We present a visual localization and map-learning system that relies on vision only and that is able to incrementally learn to recognize the different rooms of an apartment from any robot position. This system is inspired by visual categorization algorithms called bag of words methods that we modified to make fully incremental and to allow a user-interactive training. Our system is able to reliably recognize the room in which the robot is after a short training time and is stable for long term use. Empirical validation on a real robot and on an image database acquired in real environments are presented. David Filliat |
ICRA | 1 |
| 1999 | Evolution of Neural Controllers for Locomotion and Obstacle Avoidance in a Six-legged RobotabstractThis article describes how the SGOCE paradigm has been used within the context of a 'minimal simulation' strategy to evolve neural networks controlling locomotion and obstacle avoidance in a six-legged robot. A standard genetic algorithm has been used to evolve developmental programs according to which recurrent networks of leaky-integrator neurons were grown in a user-provided developmental substrate and were connected to the robot's sensors and actuators. Specific grammars have been used to limit the complexity of the developmental programs and of the corresponding neural controllers. Such controllers were first evolved through simulation and then successfully downloaded on the real robot. David Filliat, Jérôme Kodjabachian, Jean-Arcady Meyer |
Connect. Sci. | 1 |