Fabien Moutarde

dblp:69/3569 · DBLP profile ↗
← Back
25ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0003-4799-7285ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 2 since 2021Systems, architecture and hardware · 5 · 3 since 2021
YearPublicationVenuePosition
2026 J-Neus: Joint Field Optimization for Neural Surface Reconstruction in Urban Scenes With Limited Image Overlap
abstract
Reconstructing the surrounding surface geometry from recorded driving sequences poses a significant challenge due to the limited image overlap and complex topology of urban environments. SoTA neural implicit surface reconstruction methods often struggle in such setting, either failing due to small vision overlap or exhibiting suboptimal performance in accurately reconstructing both the surface and fine structures. To address these limitations, we introduce J-NeuS, a novel hybrid implicit surface reconstruction method for large driving sequences with outward facing camera poses. J-NeuS leverages cross-representation uncertainty estimation to tackle ambiguous geometry caused by limited observations. Our method performs joint optimization of two radiance fields in addition to guided sampling achieving accurate reconstruction of large areas along with fine structures in complex urban scenarios. Extensive evaluation on major driving datasets demonstrates the superiority of our approach in reconstructing large driving sequences with limited image overlap, outperforming concurrent SoTA methods.
Fusang Wang, Hala Djeghim, Fabien Moutarde, Desire Sidibé
3DV3
2026 Future-Interactions-Aware Trajectory Prediction via Braid Theory
abstract
International audience
Caio Azevedo, Stefano Sabatini, Sascha Hornauer, Fabien Moutarde
IV4
2026 An Efficient Multi-Estimation-Based Parameter Centroid Decision via Linear Regression Approach
abstract
We propose a novel post-processing approach for the local optimization of Locally Optimized RANdom SAmple Consensus (LO-RANSAC), called the Multi-Estimation-based Parameter Centroid (MEPC) decision. It is observed that the optimal thresholds for hypothesis generation and evaluation differ in local optimization with the inner RANSAC. Instead of binary labeling for inliers and outliers, a new ternary labeling for inliers, midliers, and outliers is introduced, using two thresholds. Our experimental results show that the highest-scoring model measured by the ternary method is closer to the real model than that measured by the existing binary method. However, it should be noted that the highest score still does not correspond to the best model due to inaccurate evaluation by data noise. We introduce a new linear model centroid decision method to compensate for the highest-scoring model distorted by noise. In this process, an efficient method for measuring the similarity between two hypotheses is introduced, and candidates close to the real model are found by comparing their similarity with the highest-scoring model. Our approach determines a representative model of the multiple candidate hypotheses, which is defined as the geometric centroid of hyperplanes. We test on various datasets for homography, fundamental, and essential matrices, demonstrating that applying MEPC to existing RANSAC algorithms achieves more accurate and stable model estimation. Moreover, additional experiments on vanishing point detection show the potential of our approach for various model estimation applications.
Yeongyu Choi, Fabien Moutarde, Ju H. Park 0001, Ho-Youl Jung
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields
abstract
Sound plays a major role in human perception. Along with vision, it provides essential information for understanding our surroundings. Despite advances in neural implicit representations, learning acoustics that align with visual scenes remains a challenge. We propose NeRAF, a method that jointly learns acoustic and radiance fields. NeRAF synthesizes both novel views and spatialized room impulse responses (RIR) at new positions by conditioning the acoustic field on 3D scene geometric and appearance priors from the radiance field. The generated RIR can be applied to auralize any audio signal. Each modality can be rendered independently and at spatially distinct positions, offering greater versatility. We demonstrate that NeRAF generates high-quality audio on SoundSpaces and RAF datasets, achieving significant performance improvements over prior methods while being more data-efficient. Additionally, NeRAF enhances novel view synthesis of complex scenes trained with sparse data through cross-modal learning. NeRAF is designed as a Nerfstudio module, providing convenient access to realistic audio-visual generation. Project page: https://amandinebtto.github.io/NeRAF
Amandine Brunetto, Sascha Hornauer, Fabien Moutarde
ICLR3
2025 S2BEV: Lightweight, Robust, and Precise SLAM-Oriented Segmentation Bird Eye's View Mapping Approach
abstract
As modern agriculture progresses, the swift deployment of accurate maps becomes essential for the autonomous navigation and operation of orchard robots. Traditional mapping techniques often fall short in addressing the challenges posed by orchards, which are characterized by unstructured, dynamically changing environments with complex spatial and temporal dynamics due to seasonal and continuous operations. This paper proposes a new approach to orchard map construction that merges topological maps with semantic SLAM. This integration enables the creation, optimization, and rapid deployment of maps that are not only lightweight and robust but also precise. To evaluate the effectiveness of our method, we performed navigation tests in orchard environments using the newly developed maps. The experimental outcomes demonstrated a significant reduction in CPU usage, with maximum and average reductions of 7.6% and 4.5%, respectively. This approach not only enhances navigation efficiency but also facilitates quicker map deployment, effectively freeing computational resources for other critical tasks.
Yefeng Sun, Jialing Dai, Bishu Gao, Jinghan Cai, Gengjie Lin, Fabien Moutarde, Chengliang Liu 0001
ICRA7
2025 DOC-Depth: A Novel Approach for Dense Depth Ground Truth Generation
abstract
Accurate depth information is essential for many computer vision applications. Yet, no available dataset recording method allows for fully dense accurate depth estimation in a large-scale dynamic environment. In this paper, we introduce DOC-Depth, a novel, efficient and easy-to-deploy approach for dense depth generation from any LiDAR sensor. After reconstructing consistent dense 3D environment using LiDAR odometry, we address dynamic objects occlusions automatically thanks to DOC, our state-of-the art dynamic object classification method. Additionally, DOC-Depth is fast and scalable, allowing for the creation of unbounded datasets in terms of size and time. We demonstrate the effectiveness of our approach on the KITTI dataset, improving its density from 16.1 % to 71.2% and release this new fully dense depth annotation, to facilitate future research in the domain. We also showcase results using various LiDAR sensors and in multiple environments. All software components are publicly available for the research community at https://simondemoreau.github.io/DOC-Depth/.
Simon De Moreau, Mathias Corsia, Hassan Bouchiba, Yasser Almehio, Andrei Bursuc, Hafid El-Idrissi, Fabien Moutarde
IV7
2024 MBAPPE: MCTS-Built-Around Prediction for Planning Explicitly
abstract
We present MBAPPE, a novel approach to motion planning for autonomous driving combining tree search with a partially-learned model of the environment. Leveraging the inherent explainable exploration and optimization capabilities of the Monte-Carlo Tree Search (MCTS), our method addresses complex decision-making in a dynamic environment. We propose a framework that combines MCTS with supervised learning, enabling the autonomous vehicle to effectively navigate through diverse scenarios. Experimental results demonstrate the effectiveness and adaptability of our approach, showcasing improved real-time decision-making and collision avoidance. This paper contributes to the field by providing a robust solution for motion planning in autonomous driving systems, enhancing their explainability and reliability. Code is available under https://github.com/raphychek/mbappe-nuplan.
Raphaël Chekroun, Thomas Gilles, Marin Toromanoff, Sascha Hornauer, Fabien Moutarde
IV5
2023 The Audio-Visual BatVision Dataset for Research on Sight and Sound
abstract
Vision research showed remarkable success in understanding our world, propelled by datasets of images and videos. Sensor data from radar, LiDAR and cameras supports research in robotics and autonomous driving for at least a decade. However, while visual sensors may fail in some conditions, sound has recently shown potential to complement sensor data. Simulated room impulse responses (RIR) in 3D apartment-models became a benchmark dataset for the community, fostering a range of audiovisual research. In simulation, depth is predictable from sound, by learning bat-like perception with a neural network. Concurrently, the same was achieved in reality by using RGB-D images and echoes of chirping sounds. Biomimicking bat perception is an exciting new direction but needs dedicated datasets to explore the potential. Therefore, we collected the BatVision dataset to provide large-scale echoes in complex real-world scenes to the community. We equipped a robot with a speaker to emit chirps and a binaural microphone to record their echoes. Synchronized RGB-D images from the same perspective provide visual labels of traversed spaces. We sampled modern US office spaces to historic French university grounds, indoor and outdoor with large architectural variety. This dataset will allow research on robot echolocation, general audio-visual tasks and sound phænomena unavailable in simulated data. We show promising results for audio-only depth prediction and show how state-of-the-art work developed for simulated data can also succeed on our dataset. Project page: https://amandinebtto.github.io/Batvision-Dataset/
Amandine Brunetto, Sascha Hornauer, Stella X. Yu, Fabien Moutarde
IROS4
2023 TSGN: Temporal Scene Graph Neural Networks with Projected Vectorized Representation for Multi-Agent Motion Prediction
abstract
Predicting future motions of nearby agents is essential for an autonomous vehicle to take safe and effective actions. In this paper, we propose TSGN, a framework using Temporal Scene Graph Neural Networks with projected vectorized representations for multi-agent trajectory prediction. Projected vectorized representation models the traffic scene as a graph which is constructed by a set of vectors. These vectors represent agents, road network, and their spatial relative relationships. All relative features under this representation are both translation-and rotation-invariant. Based on this representation, TSGN captures the spatial-temporal features across agents, road network, interactions among them, and temporal dependencies of temporal traffic scenes. TSGN can predict multimodal future trajectories for all agents simultaneously, plausibly, and accurately. Mean-while, we propose a Hierarchical Lane Transformer for capturing interactions between agents and road network, which filters the surrounding road network and only keeps the most probable lane segments which could have an impact on the future behavior of the target agent. Without sacrificing the prediction performance, this greatly reduces the computational burden. Experiments show TSGN achieves state-of-the-art performance on the Argoverse motion forecasting benchmark.
Yunong Wu, Thomas Gilles, Bogdan Stanciulescu, Fabien Moutarde
IV4
2022 THOMAS: Trajectory Heatmap Output with learned Multi-Agent Sampling
Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu, Fabien Moutarde
ICLR5
2022 GOHOME: Graph-Oriented Heatmap Output for future Motion Estimation
abstract
In this paper, we propose GOHOME, a method leveraging graph representations of the High Definition Map and sparse projections to generate a heatmap output representing the future position probability distribution for a given agent in a traffic scene. This heatmap output yields an unconstrained 2D grid representation of agent future possible locations, allowing inherent multimodality and a measure of the uncertainty of the prediction. Our graph-oriented model avoids the high computation burden of representing the surrounding context as squared images and processing it with classical CNNs, but focuses instead only on the most probable lanes where the agent could end up in the immediate future. GOHOME reaches 2nd on Argoverse Motion Forecasting Benchmark on the Misskate6metric while achieving significant speed-up and memory burden diminution compared to Argoverse 1stplace method HOME. We also highlight that heatmap output enables multimodal ensembling and improve 1stplace MissRate6by more than 15% with our best ensemble on Argoverse. Finally, we evaluate and reach state-of-the-art performance on the other trajectory prediction datasets nuScenes and Interaction, demonstrating the generalizability of our method.
Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu, Fabien Moutarde
ICRA5
2022 Assessing Cross-dataset Generalization of Pedestrian Crossing Predictors
abstract
Pedestrian crossing prediction has been a topic of active research, resulting in many new algorithmic solutions. While measuring the overall progress of those solutions over time tends to be more and more established due to the new publicly available benchmark and standardized evaluation procedures, knowing how well existing predictors react to unseen data remains an unanswered question. This evaluation is imperative as serviceable crossing behavior predictors should be set to work in various scenarios without compromising pedestrian safety due to misprediction. To this end, we conduct a study based on direct cross-dataset evaluation. Our experiments show that current state-of-the-art pedestrian behavior predictors generalize poorly in cross-dataset evaluation scenarios, regardless of their robustness during a direct training-test set evaluation setting. In the light of what we observe, we argue that the future of pedestrian crossing prediction, e.g. reliable and generalizable implementations, should not be about tailoring models, trained with very little available data, and tested in a classical train-test scenario with the will to infer anything about their behavior in real life. It should be about evaluating models in a cross-dataset setting while considering their uncertainty estimates under domain shift.
Joseph Gesnouin, Steve Pechberti, Bogdan Stanciulescu, Fabien Moutarde
IV4
2021 TrouSPI-Net: Spatio-temporal attention on parallel atrous convolutions and U-GRUs for skeletal pedestrian crossing prediction
abstract
Understanding the behaviors and intentions of pedestrians is still one of the main challenges for vehicle autonomy, as accurate predictions of their intentions can guarantee their safety and driving comfort of vehicles. In this paper, we address pedestrian crossing prediction in urban traffic environments by linking the dynamics of a pedestrian's skeleton to a binary crossing intention. We introduce TrouSPI-Net: a context-free, lightweight, multi-branch predictor. TrouSPI-Net extracts spatio-temporal features for different time resolutions by encoding pseudo-images sequences of skeletal joints' positions and processes them with parallel attention modules and atrous convolutions. The proposed approach is then enhanced by processing features such as relative distances of skeletal joints, bounding box positions, or ego-vehicle speed with U-GRUs. Using the newly proposed evaluation procedures for two large public naturalistic data sets for studying pedestrian behavior in traffic: JAAD and PIE, we evaluate TrouSPI-Net and analyze its performance. Experimental results show that TrouSPI-Net achieved 76% F1 score on JAAD and 80% F1 score on PIE, therefore outperforming current state-of-the-art while being lightweight and context-free.
Joseph Gesnouin, Steve Pechberti, Bogdan Stanciulescu, Fabien Moutarde
FG4
2020 End-to-End Model-Free Reinforcement Learning for Urban Driving Using Implicit Affordances
abstract
Reinforcement Learning (RL) aims at learning an optimal behavior policy from its own experiments and not rule-based control methods. However, there is no RL algorithm yet capable of handling a task as difficult as urban driving. We present a novel technique, coined implicit affordances, to effectively leverage RL for urban driving thus including lane keeping, pedestrians and vehicles avoidance, and traffic light detection. To our knowledge we are the first to present a successful RL agent handling such a complex task especially regarding the traffic light detection. Furthermore, we have demonstrated the effectiveness of our method by winning the Camera Only track of the CARLA challenge.
Marin Toromanoff, Émilie Wirbel, Fabien Moutarde
CVPR3
2019 Real-Time Gestural Control of Robot Manipulator Through Deep Learning Human-Pose Inference
Jesus Bujalance Martin, Fabien Moutarde
ICVS2
2019 Urban Localization with Street Views using a Convolutional Neural Network for End-to-End Camera Pose Regression
abstract
This paper presents an end-to-end real-time monocular absolute localization approach that uses Google Street View panoramas as a prior source of information to train a Convolutional Neural Network (CNN). We propose an adaptation of the PoseNet architecture [8] to a sparse database of panoramas. We show that we can expand the latter by synthesizing new images and consequently improve the accuracy of the pose regressor. The main advantage of our method is that it does not require a first passage of an equipped vehicle to build a map. Moreover, the offline data generation and CNN training are automatic and does not require the input of an operator. In the online phase, the approach only uses one camera for localization and regresses poses in a global frame. The conducted experiments show that augmenting the training set as presented in this paper drastically improves the accuracy of the CNN. The results, when compared to a handcrafted-feature-based approach, are less accurate (around 7.5 to 8 m against 2.5 to 3 m) but also less dependent on the position of the camera inside the vehicle. Furthermore, our CNN-based method computes the pose approximately 40 times faster (75 ms per image instead of 3 s) than the handcrafted approach.
Guillaume Bresson, Cyril Joly, Fabien Moutarde
IV4
2018 Deep Learning for Hand Gesture Recognition on Skeletal Data
abstract
In this paper, we introduce a new 3D hand gesture recognition approach based on a deep learning model. We propose a new Convolutional Neural Network (CNN) where sequences of hand-skeletal joints' positions are processed by parallel convolutions; we then investigate the performance of this model on hand gesture sequence classification tasks. Our model only uses hand-skeletal data and no depth image. Experimental results show that our approach achieves a state-of-the-art performance on a challenging dataset (DHG dataset from the SHREC 2017 3D Shape Retrieval Contest), when compared to other published approaches. Our model achieves a 91.28% classification accuracy for the 14 gesture classes case and an 84.35% classification accuracy for the 28 gesture classes case.
Guillaume Devineau, Fabien Moutarde
FG2
2018 End to End Vehicle Lateral Control Using a Single Fisheye Camera
abstract
Convolutional neural networks are commonly used to control the steering angle for autonomous cars. Most of the time, multiple long range cameras are used to generate lateral failure cases. In this paper we present a novel model to generate this data and label augmentation using only one short range fisheye camera. We present our simulator and how it can be used as a consistent metric for lateral end-to-end control evaluation. Experiments are conducted on a custom dataset corresponding to more than 10000 km and 200 hours of open road driving. Finally we evaluate this model on real world driving scenarios, open road and a custom test track with challenging obstacle avoidance and sharp turns. In our simulator based on real-world videos, the final model was capable of more than 99% autonomy on urban road.
Marin Toromanoff, Émilie Wirbel, Frédéric Wilhelm, Camilo Vejarano, Xavier Perrotton, Fabien Moutarde
IROS6
2017 Topological localization using Wi-Fi and vision merged into FABMAP framework
abstract
This paper introduces a topological localization algorithm that uses visual and Wi-Fi data. Its main contribution is a novel way of merging data from these sensors. By making Wi-Fi signature suited to FABMAP algorithm, it develops an early-fusion framework that solves global localization and kidnapped robot problem. The resulting algorithm is tested and compared to FABMAP visual localization, over data acquired by a Pepper robot in an office building. Several constraints were applied during acquisition to make the experiment fitted to real-life scenarios. Without any tuning, early-fusion surpasses the performances of visual localization by a significant margin: 94% of estimated localizations are less than 5m away from ground truth compared to 81% with visual localization.
Mathieu Nowakowski, Cyril Joly, Sébastien Dalibard, Fabien Moutarde
IROS5
2016 A distributed MPC framework for road-following formation control of car-like vehicles
abstract
This work presents a novel framework for the formation control of multiple autonomous ground vehicles in an on-road environment. Unique challenges of this problem lie in 1) the design of collision avoidance strategies with obstacles and with other vehicles in a highly structured environment, 2) dynamic reconfiguration of the formation to handle different task specifications. In this paper, we design a local MPC-based trajectory planner for each individual vehicle to follow a reference trajectory while satisfying the various kinematic and dynamic constraints of the vehicles as well as collision avoidance and formation-keeping requirements. The reference trajectory of a vehicle is computed from its leader's trajectory, based on a predefined formation tree. We use logic rules to organize the collision avoidance behaviors of member vehicles. Moreover, we propose a methodology to safely reconfigure the formation on-the-fly. The proposed framework has been validated using high-fidelity simulations.
Xiangjun Qian, Florent Altché, Arnaud de La Fortelle, Fabien Moutarde
ICARCV4
2016 Monocular urban localization using street view
abstract
This paper presents a metric global localization in the urban environment only with a monocular camera and the Google Street View database. We fully leverage the abundant sources from the Street View and benefits from its topo-metric structure to build a coarse-to-fine positioning, namely a topological place recognition process and then a metric pose estimation by local bundle adjustment. Our method is tested on a 3 km urban environment and demonstrates both sub-meter accuracy and robustness to viewpoint changes, illumination and occlusion. To our knowledge, this is the first work that studies the global urban localization simply with a single camera and Street View.
Cyril Joly, Guillaume Bresson, Fabien Moutarde
ICARCV4
2016 A hierarchical Model Predictive Control framework for on-road formation control of autonomous vehicles
abstract
This paper presents an approach for the formation control of autonomous vehicles traversing along a multi-lane road with obstacles and traffic. A major challenge in this problem is a requirement for integrating individual vehicle behaviors such as lane-keeping and collision avoidance with a global formation maintenance behavior. We propose a hierarchical Model Predictive Control (MPC) approach. The desired formation is modeled as a virtual structure evolving curvilinearly along a centerline, and vehicle configurations are expressed as curvilinear relative longitudinal and lateral offsets from the virtual center. At high-level, the trajectory generation of the virtual center is achieved through an MPC framework, which allows various on-road driving constraints to be considered in the optimization. At low-level, a local MPC controller computes the vehicle inputs in order to track the desired trajectory, taking into account more personalized driving constraints. High-fidelity simulations show that the proposed approach drives vehicles to the desired formation while retains some freedom for individual vehicle behaviors.
Xiangjun Qian, Arnaud de La Fortelle, Fabien Moutarde
Intelligent Vehicles Symposium3
2013 Recognition of supplementary signs for correct interpretation of traffic signs
abstract
Traffic Sign Recognition (TSR) is now relatively well-handled by several approaches. However, traffic signs are often completed by one (or several) supplementary sign(s) placed below. They are essential for correct interpretation of main sign, as they specify its applicability scope. The main difficulty of supplementary sub-sign recognition is the potentially infinite number of classes, as nearly any information can be written on them. In this paper, we propose and evaluate a hierarchical approach for recognition of supplementary signs, in which the “meta-class” of the sub-sign (Arrow, Pictogram, Text or Mixed) is first determined. The classification is based on the pyramid-HOG feature, completed by dark area proportion measured on the same pyramid. Evaluation on a large database of images with and without supplementary signs shows that the classification accuracy of our approach reaches 95% precision and recall. When used on output of our sub-sign specific detection algorithm, the global correct detection and recognition rate is 91%.
Anne-Sophie Puthon, Fabien Moutarde, Fawzi Nashashibi
Intelligent Vehicles Symposium2
2008 A robot behavior-learning experiment using Particle Swarm Optimization for training a neural-based animat
abstract
We investigate the use of particle swarm optimization (PSO), and compare with genetic algorithms (GA), for a particular robot behavior-learning task: the training of an animat behavior totally determined by a fully-recurrent neural network, and with which we try to fulfill a simple exploration and food foraging task. The target behavior is simple, but the learning task is challenging because of the dynamic complexity of fully-recurrent neural networks. We show that standard PSO yield very good results for this learning problem, and appears to be much more effective than simple GA.
Fabien Moutarde
ICARCV1
2004 Fast semi-automatic segmentation algorithm for Self-Organizing Maps
David Opolon, Fabien Moutarde
ESANN2