Vincent Frémont

dblp:95/3310 · DBLP profile ↗
← Back
37ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-1627-2241ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 since 2021Systems, architecture and hardware · 6Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Multi-view Projection for Unsupervised Domain Adaptation in 3D Semantic Segmentation
Andrew Caunes, Thierry Chateau, Vincent Frémont
ICPR (11)3
2025 Mixed Signals: A Diverse Point Cloud Dataset for Heterogeneous LiDAR V2X Collaboration
abstract
Vehicle-to-everything (V2X) collaborative perception has emerged as a promising solution to address the limitations of single-vehicle perception systems. However, existing V2X datasets are limited in scope, diversity, and quality. To address these gaps, we present Mixed Signals, a comprehensive V2X dataset featuring 45.1k point clouds and 240.6k bounding boxes collected from three connected autonomous vehicles (CAVs) equipped with two different configurations of LiDAR sensors, plus a roadside unit with dual LiDARs. Our dataset provides point clouds and bounding box annotations across 10 classes, ensuring reliable data for perception training. We provide detailed statistical analysis on the quality of our dataset and extensively benchmark existing V2X methods on it. The Mixed Signals dataset is ready-to-use, with precise alignment and consistent annotations across time and viewpoints. Dataset website is available at https://mixedsignalsdataset.cs.cornell.edu/.
Katie Luo, Minh-Quan Dao, Mark E. Campbell, Wei-Lun Chao, Kilian Q. Weinberger, Ezio Malis, Vincent Frémont, Bharath Hariharan, Mao Shan, Stewart Worrall 0002, Julie Stephany Berrio
ICCV8
2025 Risk-Aware Nonlinear Model Predictive Control for Autonomous Navigation: Confidence-Based Obstacle Constraints and Time-to-Collision Safe Trajectory Planning
abstract
Nonlinear Model Predictive Control (NMPC) has established itself as a robust framework for autonomous navigation problems by effectively handling nonlinear dynamics and complex constraints in real-time. Despite its success in various applications, challenges remain in integrating robust obstacle avoidance while navigating in dynamic and uncertain environment. To enhance safety, risk assessment metrics like Time-To-Collision (TTC) are widely used for evaluating collision risks. Extensions to these metrics allow their application across diverse road scenarios, making them valuable as cost function components or constraints within risk-aware NMPC formulations. In this work, we propose novel obstacle constraint formulations that account for uncertainties in obstacle positions using confidence thresholds, along with TTC-based constraints to ensure safe trajectory planning. Our approaches are validated through extensive simulations in urban scenarios, including overtaking maneuvers and intersections, demonstrating their effectiveness in mitigating risks. Detailed implementations and results are publicly available to support further research and development.
Charlotte Beaune, Elwan Héry, Vincent Frémont
IV3
2024 Active Collaborative Visual SLAM Exploiting ORB Features
abstract
In autonomous robotics, a significant challenge involves devising robust solutions for Active Collaborative SLAM (AC-SLAM). This process requires multiple robots to cooperatively explore and map an unknown environment by intelligently coordinating their movements and sensor data acquisition. In this article, we present an efficient visual AC-SLAM method using aerial and ground robots for environment exploration and mapping. We propose an efficient frontiers filtering method that takes into account the common IoU map frontiers and reduces the frontiers for each robot. Additionally, we present an approach to guide robots to previously visited goal positions to promote loop closure and reduce SLAM uncertainty. The proposed method is implemented in ROS and evaluated through simulations on publicly available datasets and similar methods, achieving an accumulative average of 59% increase in area coverage.
Muhammad Farhan Ahmed, Vincent Frémont, Isabelle Fantoni
ICARCV2
2024 3D Can Be Explored In 2D : Pseudo-Label Generation for LiDAR Point Clouds Using Sensor-Intensity-Based 2D Semantic Segmentation
abstract
Semantic segmentation of 3D LiDAR point clouds, essential for autonomous driving and infrastructure management, is best achieved by supervised learning, which demands extensive annotated datasets and faces the problem of domain shifts. We introduce a new 3D semantic segmentation pipeline that leverages aligned scenes and state-of-the-art 2D segmentation methods, avoiding the need for direct 3D annotation or reliance on additional modalities such as camera images at inference time. Our approach generates 2D views from LiDAR scans colored by sensor intensity and applies 2D semantic segmentation to these views using a camera-domain pretrained model. The segmented 2D outputs are then back-projected onto the 3D points, with a simple voting-based estimator that merges the labels associated to each 3D point. Our main contribution is a global pipeline for 3D semantic segmentation requiring no prior 3D annotation and not other modality for inference, which can be used for pseudo-label generation. We conduct a thorough ablation study and demonstrate the potential of the generated pseudo-labels for the Unsupervised Domain Adaptation task.
Andrew Caunes, Thierry Chateau, Vincent Frémont
IV3
2024 Practical Collaborative Perception: A Framework for Asynchronous and Multi-Agent 3D Object Detection
abstract
Occlusion is a major challenge for LiDAR-based object detection methods as it renders regions of interest unobservable to the ego vehicle. A proposed solution to this problem comes from collaborative perception via Vehicle-to-Everything (V2X) communication, which leverages a diverse perspective thanks to the presence of connected agents (vehicles and intelligent roadside units) at multiple locations to form a complete scene representation. The major challenge of V2X collaboration is the performance-bandwidth tradeoff which presents two questions (i) which information should be exchanged over the V2X network, and (ii) how the exchanged information is fused. The current state-of-the-art resolves to the mid-collaboration approach where Birds-Eye View (BEV) images of point clouds are communicated to enable a deep interaction among connected agents while reducing bandwidth consumption. While achieving strong performance, the real-world deployment of most mid-collaboration approaches are hindered by their overly complicated architectures and unrealistic assumptions about inter-agent synchronization. In this work, we devise a simple yet effective collaboration method based on exchanging the outputs from each agent that achieves a better bandwidth-performance tradeoff while minimising the required changes to the single-vehicle detection models. Moreover, we relax the assumptions used in existing state-of-the-art approaches about inter-agent synchronization to only require a common time reference among connected agents, which can be achieved in practice using GPS time. Experiments on the V2X-Sim dataset show that our collaboration method reaches 76.72 mean average precision which is 99% the performance of the early collaboration method while consuming as much bandwidth as the late collaboration (0.01 MB on average). The code is released in https://github.com/quan-dao/practical-collab-perception.
Minh-Quan Dao, Julie Stephany Berrio, Vincent Frémont, Mao Shan, Elwan Héry, Stewart Worrall 0002
IV3
2024 Label-Efficient 3D Object Detection For Road-Side Units
abstract
Occlusion presents a significant challenge for safety-critical applications such as autonomous driving. Collaborative perception has recently attracted a large research interest thanks to the ability to enhance the perception of autonomous vehicles via deep information fusion with intelligent roadside units (RSU), thus minimizing the impact of occlusion. While significant advancement has been made, the data-hungry nature of these methods creates a major hurdle for their realworld deployment, particularly due to the need for annotated RSU data. Manually annotating the vast amount of RSU data required for training is prohibitively expensive, given the sheer number of intersections and the effort involved in annotating point clouds. We address this challenge by devising a label-efficient object detection method for RSU based on unsupervised object discovery. Our paper introduces two new modules: one for object discovery based on a spatial temporal aggregation of point clouds, and another for refinement. Furthermore, we demonstrate that fine-tuning on a small portion of annotated data allows our object discovery models to narrow the performance gap with, or even surpass, fully supervised models. Extensive experiments are carried out in simulated and real-world datasets to evaluate our method†.
Minh-Quan Dao, Holger Caesar, Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002, Vincent Frémont, Ezio Malis
IV6
2024 Practical Collaborative Perception: A Framework for Asynchronous and Multi-Agent 3D Object Detection
abstract
Occlusion is a major challenge for LiDAR-based object detection methods as it renders regions of interest unobservable to the ego vehicle. A proposed solution to this problem comes from collaborative perception via Vehicle-to-Everything (V2X) communication, which leverages a diverse perspective thanks to the presence of connected agents (vehicles and intelligent roadside units) at multiple locations to form a complete scene representation. The major challenge of V2X collaboration is the performance-bandwidth tradeoff which presents two questions 1) which information should be exchanged over the V2X network and 2) how the exchanged information is fused. The current state-of-the-art resolves to the mid-collaboration approach where Birds-Eye View (BEV) images of point clouds are communicated to enable a deep interaction among connected agents while reducing bandwidth consumption. While achieving strong performance, the real-world deployment of most mid-collaboration approaches are hindered by their overly complicated architectures and unrealistic assumptions about inter-agent synchronization. In this work, we devise a simple yet effective collaboration method based on exchanging the outputs from each agent that achieves a better bandwidth-performance tradeoff while minimising the required changes to the single-vehicle detection models. Moreover, we relax the assumptions used in existing state-of-the-art approaches about inter-agent synchronization to only require a common time reference among connected agents, which can be achieved in practice using GPS time. Experiments on the V2X-Sim dataset show that our collaboration method reaches 76.72 mean average precision which is 99% the performance of the early collaboration method while consuming as much bandwidth as the late collaboration (0.01 MB on average). The code will be released in https://github.com/quan-dao/practical-collab-perception.
Minh-Quan Dao, Julie Stephany Berrio, Vincent Frémont, Mao Shan, Elwan Héry, Stewart Worrall 0002
IEEE Trans. Intell. Transp. Syst.3
2023 Aligning Bird-Eye View Representation of Point Cloud Sequences using Scene Flow
abstract
Low-resolution point clouds are challenging for object detection methods due to their sparsity. Densifying the present point cloud by concatenating it with its predecessors is a popular solution to this challenge. Such concatenation is possible thanks to the removal of ego vehicle motion using its odometry. This method is called Ego Motion Compensation (EMC). Thanks to the added points, EMC significantly improves the performance of single-frame detectors. However, it suffers from the shadow effect that manifests in dynamic objects’ points scattering along their trajectories. This effect results in a misalignment between feature maps and objects’ locations, thus limiting performance improvement to stationary and slow-moving objects only. Scene flow allows aligning point clouds in 3D space, thus naturally resolving the misalignment in feature spaces. By observing that scene flow computation shares several components with 3D object detection pipelines, we develop a plug-in module that enables single-frame detectors to compute scene flow to rectify their Bird-Eye View representation. Experiments on the NuScenes dataset show that our module leads to a significant increase (up to 16%) in the Average Precision of large vehicles, which interestingly demonstrates the most severe shadow effect.
Minh-Quan Dao, Vincent Frémont, Elwan Héry
IV2
2022 Attention-based Proposals Refinement for 3D Object Detection
abstract
Recent advances in 3D object detection are made by developing the refinement stage for voxel-based Region Proposal Networks (RPN) to better strike the balance between accuracy and efficiency. A popular approach among state-of-the-art frameworks is to divide proposals, or Regions of Interest (ROI), into grids and extract features for each grid location before synthesizing them to form ROI features. While achieving impressive performances, such an approach involves several hand-crafted components (e.g. grid sampling, set abstraction) which requires expert knowledge to be tuned correctly. This paper proposes a data-driven approach to ROI feature computing named APRO3D-Net which consists of a voxel-based RPN and a refinement stage made of Vector Attention. Unlike the original multi-head attention, Vector Attention assigns different weights to different channels within a point feature, thus being able to capture a more sophisticated relation between pooled points and ROI. Our method achieves a competitive performance of 84.85 AP for class Car at moderate difficulty on validation set of KITTI and 47.03 mAP (average over 10 classes) on NuScenes while having the least parameters compared to closely related methods and attaining an inference speed at 15 FPS on NVIDIA V100 GPU. The code is released1.1https://github.com/quan-dao/APRO3D-Net
Minh-Quan Dao, Elwan Héry, Vincent Frémont
IV3
2022 3D-FlowNet: Event-based optical flow estimation with 3D representation
abstract
Event-based cameras can overpass frame-based cameras’ limitations for important tasks such as high-speed motion detection during self-driving cars’ navigation in low illumination conditions. The event cameras’ high temporal resolution and high dynamic range allow them to work in fast motion and extreme light scenarios. However, conventional computer vision methods, such as Deep Neural Networks, are not well adapted to work with event data as they are asynchronous and discrete. A popular approach among state-of-the-art is accumulating the events into 2D image-like matrix. Such 2D-encoding representation is compatible with the traditional convolution neural network but sacrifices the time resolution. In this paper, we present a 3D encoding representation for the event data that can better preserve the temporal distribution of the events compared to the 2D-encoding representation. We then propose 3D-FlowNet, a novel network architecture that processes the 3D input representation with the 3D convolution layer based on the new representation. The 3D features are then flattened into 2D so they can be processed by the 2D convolution layer and generate the optical flow estimation. A self-supervised training strategy is adopted due to the lack of labeled datasets for the event-based camera. Finally, the proposed network is trained and evaluated with the Multi-Vehicle Stereo Event Camera (MVSEC) dataset. The results show that our 3D-FlowNet outperforms state-of-the-art approaches with less training epoch (30 compared to 100 of Spike-FlowNet). The code is released in https://github.com/adosum/3D-FlowNet.
Minh-Quan Dao, Vincent Frémont
IV3
2022 A Deep Learning Approach for LiDAR Resolution-Agnostic Object Detection
abstract
Existing neural network-based object detection approaches process LiDAR point clouds trained from one kind of LiDAR sensor. In the case of a different point cloud input, the trained network performs with less efficiency, especially when the given point cloud has low resolution. In this paper, we propose a new object detection approach, which is more resilient to variations in point cloud resolution. Firstly, layers from the point cloud are randomly discarded during the training phase in order to increase the variability of the data processed by the network. Secondly, the obstacles are described as Gaussian functions, grouping multiple parameters into a single representation. A Bhattacharyya distance is used as a loss function. This approach is tested on a LiDAR-based network and on an architecture using camera and LiDAR sensors. The networks are trained exclusively on the KITTI dataset and tested on Pandaset and the nuScenes Mini dataset. Experiments show that our method improves the performance of the tested networks on low-resolution point clouds without decreasing the ability to process high-resolution data.
Ruddy Théodose, Dieumet Denis, Thierry Chateau, Vincent Frémont, Paul Checchin
IEEE Trans. Intell. Transp. Syst.4
2021 A Loosely Coupled Vision-LiDAR Odometry using Covariance Intersection Filtering
abstract
This paper presents a loosely-coupled sensor fusion approach, which efficiently combines complementary visual and range sensor information to estimate the vehicle ego-motion. Descriptor-based and distance-based matching strategies are respectively applied to visual and range measurements for feature tracking. Nonlinear optimization optimally estimates the relative pose across consecutive frames and an uncertainty analysis using forward and backward covariance propagation is made to model the estimation accuracy. Covariance intersection filter paves the way for us to loosely couple stereo vision and LiDAR odometry considering respective uncertainties. We evaluate our approach with KITTI dataset which shows its effectiveness to fierce rotational motion and temporary absence of visual features, achieving the average relative translation error of 0.84% for the challenging 01 sequence on the highway.
Song-Ming Chen, Vincent Frémont
IV2
2020 Dense Decentralized Multi-robot SLAM based on locally consistent TSDF submaps
abstract
This article introduces a decentralized multi-robot algorithm for Simultaneous Localization And Mapping (SLAM) inspired from previous work on collaborative mapping [1]. This method makes robots jointly build and exchange i) a collection of 3D dense locally consistent submaps, based on a Truncated Signed Distance Field (TSDF) representation of the environment, and ii) a pose-graph representation which encodes the relative pose constraints between the TSDF submaps and the trajectory keyframes, derived from the odometry, inter-robot observations and loop closures. Such loop closures are spotted by aligning and fusing the TSDF submaps. The performances of this method have been evaluated on multi-robot scenarios built from the EuRoC dataset [2].
Rodolphe Dubois, Alexandre Eudes, Julien Moras, Vincent Frémont
IROS4
2020 Longitudinal Dynamics Model Identification of an Electric Car Based on Real Response Approximation
abstract
Obtaining a realistic and accurate model of the longitudinal dynamics is key for a good speed control of a self-driving car. It is also useful to simulate the longitudinal behavior of the vehicle with high fidelity. In this paper, a straightforward and generic method for obtaining the friction, braking and propulsion forces as a function of speed, throttle input and brake input is proposed. Experimental data is recorded during tests over the full speed range to estimate the forces, to which the corresponding curves are adjusted. A simple and direct balance of forces in the direction tangent to the ground is used to obtain an estimation of the real forces involved. Then a model composed of approximate spline curves that fit the results is proposed. Using splines to model the dynamic response has the advantage of being quick and accurate, avoiding the complexity of parameter identification and tuning of non-linear responses embedding the internal functionalities of the car, like ABS or regenerative brake. This methodology has been applied to LS2N's electric Renault Zoe but can be applied to any other electric car. As shown in the experimental section, a comparison between the estimated acceleration of the car using the model and the real one for a normal driving over a wide range of speeds along a trip of about 10 km/h reveals only 0.35 m/s2of error standard deviation in a range of ±2m/s2which is very encouraging.
Salvador Dominguez, Gaëtan Garcia, Arnaud Hamon, Vincent Frémont
IV4
2020 Adaptive Visual Assistance System for Enhancing the Driver Awareness of Pedestrians
abstract
In the past decade, Pedestrian Collision Warning Systems have been proposed to detect pedestrians and warn drivers of imminent collision. However, such systems are often limited by eye-off-road and cognitive overload problems. Head Up displays with augmented reality are being considered as a key technology for changing drivers’ user experiences. In this paper, we propose a new visual assistance system that can enhance drivers’ perception by dynamically directing attention to pedestrians to avoid collisions using Augmented Reality cues. The proposed system takes into account driver behaviors through vehicle driving signals analysis, in order to warn it at the right moments. To that end, we statistically model correct and incorrect driver behaviors in situations with pedestrians. Based on this model, a warning visual metaphor is displayed if unawareness is detected and a driving simulator was used to evaluate the concept. The experimental results suggest that our proposed adaptive visual aids can enhance driver awareness of pedestrians in critical situations.
Vincent Frémont, Minh Tien Phan, Indira Thouvenin
Int. J. Hum. Comput. Interact.1
2019 On Data Sharing Strategy for Decentralized Collaborative Visual-Inertial Simultaneous Localization And Mapping
abstract
This article introduces and evaluates two decentralized data sharing algorithms for multi-robot visual-inertial simultaneous localization and mapping (VI-SLAM): Factor Sparsification for Visual-Inertial Packets (FS-VIP) and Min-K-Cover Selection for Visual-Inertial Packets (MKCS-VIP). Both methods make robots regularly build and exchange data packets which describe the successive portions of their map, but rely on distinct paradigms. While FS-VIP builds on consistent marginalization and sparsification techniques, MKCSVIP selects raw visual and inertial information which can best help to perform a faithful and consistent re-estimation while reducing the communication cost. Performances in terms of accuracy and communication loads are evaluated on multi-robot scenarios built on both available (EUROC) and custom datasets (SOTTEVILLE).
Rodolphe Dubois, Alexandre Eudes, Vincent Frémont
IROS3
2018 Improving semantic segmentation in urban scenes with a cartographic information
abstract
This paper presents three different approaches to inject a location information in semantic segmentation Convolutional Neural Networks (CNN) applied to urban scenes. The assumption that a location information would improve semantic segmentation performance emerges from the idea that some elements of urban scenes are located in a predictable manner. This assumption is confronted to realistic data on the CARLA autonomous driving simulator, which is used to create our own synthetic dataset with images, depth maps and bird-eye-view cartographic images. Simulators circumvent the difficulties due to the scarcity of publicly available synchronous labeled images and location information. We encode the location information as a cartographic image to process it as a camera image. We assess the relevance of injecting the cartographic information in three different manners: as a Conditional Random Field potential, as an additional task and as an additional encoder input of a CNN. The three methods are evaluated and compared with a state of the art CNN with regards to the pixel-wise accuracy, mean intersection over union and intersection over union of some important classes. The multi-encoder approach improves the intersection over union of the pedestrians, vehicles and traffic signs classes by respectively 4%, 1.6% and 9%.
Abdelhak Loukkal, Vincent Frémont, Yves Grandvalet, You Li 0005
ICARCV2
2017 Problem-based band selection for hyperspectral images
abstract
This paper addresses the band selection of a hyperspectral image. Considering a binary classification, we devise a method to choose the more discriminating bands for the separation of the two classes involved, by using a simple algorithm: single-layer neural network. After that, the most discriminative bands are selected, and the resulting reduced data set is used in a more powerful classifier, namely, stacked denoising autoencoder. Besides its simplicity, the advantage of this method is that the selection of features is made by an algorithm similar to the classifier to be used, and not focused only on the separability measures of the data set. Results indicate the decrease of overfitting for the reduced data set, when compared to the full data architecture.
Mateus Habermann 0001, Vincent Frémont, Elcio Hideiti Shiguemori
IGARSS2
2017 Mono-vision based moving object detection in complex traffic scenes
abstract
International audience
Vincent Frémont, Sergio Alberto Rodriguez Florez, Bihao Wang
Intelligent Vehicles Symposium1
2017 Moving object detection and segmentation in urban environments from a moving platform
Dingfu Zhou, Vincent Frémont, Benjamin Quost, Yuchao Dai, Hongdong Li
Image Vis. Comput.2
2016 Exploiting fully convolutional neural networks for fast road detection
abstract
Road detection is a crucial task in autonomous navigation systems. It is responsible for delimiting the road area and hence the free and valid space for maneuvers. In this paper, we consider the visual road detection problem where, given an image, the objective is to classify every of its pixels into road or non-road. We address this task by proposing a convolutional neural network architecture. We are especially interested in a model that takes advantage of a large contextual window while maintaining a fast inference. We achieve this by using a Network-in-Network (NiN) architecture and by converting the model into a fully convolutional network after training. Experiments have been conducted to evaluate the effects of different contextual window sizes (the amount of contextual information) and also to evaluate the NiN aspect of the proposed architecture. Finally, we evaluated our approach using the KITTI road detection benchmark achieving results in line with other state-of-the-art methods while maintaining real-time inference. The benchmark results also reveal that the inference time of our approach is unique at this level of accuracy, being two orders of magnitude faster than other methods with similar performance.
Caio C. T. Mendes, Vincent Frémont, Denis F. Wolf
ICRA2
2016 Fast Depth Video Compression for Mobile RGB-D Sensors
abstract
We propose a new method, called 3-D image warping-based depth video compression (IW-DVC), for fast and efficient compression of depth images captured by mobile RGB-D sensors. The emergence of low-cost RGB-D sensors has created opportunities to find new solutions for a number of computer vision and networked robotics problems, such as 3-D map building, immersive telepresence, or remote sensing. However, efficient transmission and storage of depth data still presents a challenging task to the research community in these applications. Image/video compression has been comprehensively studied and several methods have already been developed. However, these methods result in unacceptably suboptimal outcomes when applied to the depth images. We have designed the IW-DVC method to exploit the special properties of the depth data to achieve a high compression ratio while preserving the quality of the captured depth images. Our solution combines the egomotion estimation and 3-D image warping techniques and includes a lossless coding scheme that is capable of adapting to depth data with a high dynamic range. IW-DVC operates at a high speed, suitable for real-time applications, and is able to attain an enhanced motion compensation accuracy compared with the conventional approaches. In addition, it removes the existing redundant information between the depth frames to further increase compression efficiency. Our experiments show that IW-DVC attains a very high performance yielding significant compression ratios without sacrificing image quality.
Y. Ahmet Sekercioglu, Tom Drummond, Enrico Natalizio, Isabelle Fantoni, Vincent Frémont
IEEE Trans. Circuits Syst. Video Technol.6
2015 Estimation of driver awareness of pedestrian based on Hidden Markov Model
abstract
Understanding driver behaviors is an important need for the Advanced Driver Assistance Systems. In particular, the pedestrian detection systems become extremely distracting and annoying when they inform the driver with unnecessary warning messages. In this paper, we propose to study the driver behaviors whenever a pedestrian appears in front of the vehicle. A method based on the driving actions and the Hidden Markov Model (HMM) algorithm is developed to classify the driver awareness of pedestrian and the driver unawareness of pedestrian. The method is successfully validated using the collected data from the experiments that are conducted on a driving simulator. Furthermore, two simple methods based on the static parameters such as the Time-To-Collision and the Required Deceleration Parameter are also applied to our problem and are compared to the proposed method. The result shows a significant improvement of the HMM-based method compared to the simple ones.
Minh Tien Phan, Vincent Frémont, Indira Thouvenin, Mohamed Sallak, Véronique Berge-Cherfaoui
Intelligent Vehicles Symposium2
2014 Deformable parts model for people detection in heavy machines applications
abstract
In this paper we focus on the evaluation of the deformable part model (DPM) proposed by Felzenszwalb et al. [IO] in the context of vision-based people detection in heavy machines applications. The proposed system uses a single fisheye camera to provide a wide field-of-view (FOV) at low cost. However, the fisheye optical distortions present several difficulties for image processing and object recognition. The DPM approach shows important flexibility when dealing with varying object's form. It gives good performances on people detection when images present strong fisheye distortions. Base on the analysis of DPM in the context of fisheye image, we proposed an adaptive detector which is more suitable.
Manh-Tuan Bui, Vincent Frémont, Djamal Boukerroui, Pierrick Letort
ICARCV2
2014 Multiple obstacle detection and tracking using stereo vision: Application and analysis
abstract
Vision systems provide a large functional spectrum for perception applications and, in recent years, they have demonstrated to be essential in the development of Advanced Driver Assistance Systems (ADAS) and Autonomous Vehicles. In this context, this paper presents an on-road objects detection approach improved by our previous work in defining the traffic area and new strategy in obstacle extraction from U-disparity. Then, a modified particle filtering is proposed for multiple object tracking. The perception strategy of the proposed vision-only detection system is structured as follows : First, a method based on illuminant invariant image is employed at an early stage for free road space detection. A convex hull is then constructed to generate a region of interest (ROI) which includes the main traffic road area. Based on this ROI, an U-disparity map is built to characterize on-road obstacles. In this approach, connected regions extraction is applied for obstacles detection instead of standard Hough Transform. Finally, a modified particle filter framework is employed for multiple targets tracking based on the former detection results. Besides, multiple cues, such as obstacle's size verification and combination of redundant detections, are embedded in the system to improve its accuracy. Our experimental findings demonstrates that the system is effective and reliable when applied on different traffic video sequences from a public database.
Bihao Wang, Sergio Alberto Rodriguez Florez, Vincent Frémont
ICARCV3
2014 Soft label based semi-supervised boosting for classification and object recognition
abstract
Supervised classification algorithms such as Boosting and SVM have achieved significant success in the field of computer vision for classification and object recognition. However, the performance of the classifier decreases rapidly if there are insufficient labeled training samples. In this paper, a semi-supervised boosting algorithm is proposed to overcome this limitation. First, a few labeled instances are use to estimate probabilistic class labels for unlabeled samples using Gaussian Mixture Models after a dimension reduction step performed via Principal Component Analysis. Then, we apply a boosting strategy on decision stumps trained using the soft labeled instances thus obtained. The performances of our strategy are evaluated on several state-of-the-art classification datasets, as well as on a pedestrian detection and recognition problem. Experimental results demonstrate the interest of taking into account additional data in the training process.
Dingfu Zhou, Benjamin Quost, Vincent Frémont
ICARCV3
2014 Color-based road detection and its evaluation on the KITTI road benchmark
abstract
Road detection is one of the key issues of scene understanding for Advanced Driving Assistance Systems (ADAS). Recent approaches has addressed this issue through the use of different kinds of sensors, features and algorithms. KITTI-ROAD benchmark has provided an open-access dataset and standard evaluation mean for road area detection. In this paper, we propose an improved road detection algorithm that provides a pixel-level confidence map. The proposed approach is inspired from our former work based on road feature extraction using illuminant intrinsic image and plane extraction from v-disparity map segmentation. In the former research, detection results of road area are represented by binary map. The novelty of this improved algorithm is to introduce likelihood theory to build a confidence map of road detection. Such a strategy copes better with ambiguous environments, compared to a simple binary map. Evaluations and comparisons of both, binary map and confidence map, have been done using the KITTI-ROAD benchmark.
Bihao Wang, Vincent Frémont, Sergio Alberto Rodriguez Florez
Intelligent Vehicles Symposium2
2014 On modeling ego-motion uncertainty for moving object detection from a mobile platform
abstract
In this paper, we propose an effective approach for moving object detection based on modeling the ego-motion uncertainty and using a graph-cut based motion segmentation. First, the relative camera pose is estimated by minimizing the sum of reprojection errors and its covariance matrix is calculated using a first-order errors propagation method. Next, a motion likelihood for each pixel is obtained by propagating the uncertainty of the ego-motion to the Residual Image Motion Flow (RIMF). Finally, the motion likelihood and the depth gradient are used in a graph-cut based approach as region and boundary terms respectively, in order to obtain the moving objects segmentation. Experimental results on real-world data show that our approach can detect dynamic objects which move on the epipolar plane or that are partially occluded in complex urban traffic scenes.
Dingfu Zhou, Vincent Frémont, Benjamin Quost, Bihao Wang
Intelligent Vehicles Symposium2
2014 Multi-modal object detection and localization for high integrity driving assistance
Sergio Alberto Rodriguez Florez, Vincent Frémont, Philippe Bonnifait, Véronique Berge-Cherfaoui
Mach. Vis. Appl.2
2013 Mapping and localization using GPS, lane markings and proprioceptive sensors
abstract
Estimating the pose in real-time is a primary function for intelligent vehicle navigation. Whilst different solutions exist, most of them rely on the use of high-end sensors. This paper proposes a solution that exploits an automotive type L1-GPS receiver, features extracted by low-cost perception sensors and vehicle proprioceptive information. A key idea is to use the lane detection function of a video camera to retrieve accurate lateral and orientation information with respect to road lane markings. To this end, lane markings are mobile-mapped by the vehicle itself during a first stage by using an accurate localizer. Then, the resulting map allows for the exploitation of camera-detected features for autonomous real-time localization. The results are then combined with GPS estimates and dead-reckoning sensors in order to provide localization information with high availability. As L1-GPS errors can be large and are time correlated, we study in the paper several GPS error models that are experimentally tested with shaping filters. The approach demonstrates that the use of low-cost sensors with adequate data-fusion algorithms should lead to computer-controlled guidance functions in complex road networks.
Philippe Bonnifait, Vincent Frémont, Javier Ibañez-Guzmán
IROS3
2013 Fast road detection from color images
abstract
In this paper, we present a method for drivable road detection by extracting its specular intrinsic feature from an image. The resulting detection is then used in a stereo vision-based 3D road parameters extraction algorithm. A substantial representation of the road surface, called axis-calibration, is represented as an angle in logchromaticity space. This feature provides an invariance to road surface under illuminant conditions with shadow or not. We also add a sky removal function in order to eliminate the negative effects of sky light on axis-calibration result. Then, a confidence interval calculation helps the pixels' classification to speed up the detection processing. At last, the approach is combined with a stereovision based method to filter out false detected pixels and to obtain precise 3D road parameters. The experimental results show that the proposed approach can be adapted for real-time ADAS system in various driving conditions.
Bihao Wang, Vincent Frémont
Intelligent Vehicles Symposium2
2012 DAARIA: Driver assistance by augmented reality for intelligent automobile
abstract
Taking into account the drivers' state is a major challenge for designing new advanced driver assistance systems. In this paper we present a driver assistance system strongly coupled to the user. DAARIA1stands for Driver Assistance by Augmented Reality for Intelligent Automobile. It is an augmented reality interface powered by several sensors. The detection has two goals: one is the position of obstacles and the quantification of the danger represented by them. The other is the driver's behavior. A suitable visualization metaphor allows the driver to perceive at any time the location of the relevant hazards while keeping his eyes on the road. First results show that our method could be applied to a vehicle but also to aerospace, fluvial or maritime navigation.
Paul George, Indira Thouvenin, Vincent Frémont, Véronique Berge-Cherfaoui
Intelligent Vehicles Symposium3
2010 UAV altitude estimation by mixed stereoscopic vision
abstract
Altitude is one of the most important parameters to be known for an Unmanned Aerial Vehicle (UAV) especially during critical maneuvers such as landing or steady flight. In this paper, we present mixed stereoscopic vision system made of a fish-eye camera and a perspective camera for altitude estimation. Contrary to classical stereoscopic systems based on feature matching, we propose a plane sweeping approach in order to estimate the altitude and consequently to detect the ground plane. Since there exists a homography between the two views and the sensor being calibrated and the attitude estimated by the fish-eye camera, the algorithm consists then in searching the altitude which verifies this homography. We show that this approach is robust and accurate, and a CPU implementation allows a real time estimation. Experimental results on real sequences of a small UAV demonstrate the effectiveness of the approach.
Damien Eynard, Pascal Vasseur, Cédric Demonceaux, Vincent Frémont
IROS4
2010 An embedded multi-modal system for object localization and tracking
abstract
Reliable obstacle detection and localization is a key issue for driving assistance systems particularly in urban environments. In this article, a multi-modal perception approach is studied in order to enhance vehicle localization and dynamic objects tracking, in a world-centric map. 3D ego-localization is done by merging a stereo vision system and proprioceptive information coming from vehicle sensors. Mobile objects are detected using a multi-layer lidar that is simultaneously used to characterize a zone of interest in order to reduce the complexity of the perception process. Objects localization and tracking is then performed in the fixed frame which simplifies the scene analysis and understanding. Real experimental results are reported to evaluate the performance of the multi-modal system.
Sergio Alberto Rodriguez Florez, Vincent Frémont, Philippe Bonnifait, Véronique Berge-Cherfaoui
Intelligent Vehicles Symposium2
2009 Automatic Camera-Based Microscope Calibration for a Telemicromanipulation System Using a Virtual Pattern
abstract
In the context of virtualized-reality-based telemicromanipulation, this paper presents a visual calibration technique for an optical microscope coupled to a charge-coupled device (CCD) camera. The accuracy and flexibility of the proposed automatic virtual calibration method, based on parallel single-plane properties, are outlined. In contrast to standard approaches, a 3-D virtual calibration pattern is constructed using the micromanipulator tip with subpixel-order localization in the image frame. The proposed procedure leads to a linear system whose solution provides directly both the intrinsic and extrinsic parameters of the geometrical model. Computer simulations and real data have been used to test the proposed technique, and promising results have been obtained. Based on the proposed calibration techniques, a 3-D virtual microenvironment of the workspace is reconstructed through the real-time imaging of two perpendicular optical microscopes. Our method provides a flexible, easy-to-use technical alternative to the classical techniques used in micromanipulation systems.
Mehdi Ammi, Vincent Frémont, Antoine Ferreira
IEEE Trans. Robotics2
2005 Flexible Microscope Calibration using Virtual Pattern for 3-D Telemicromanipulation
abstract
In the context of virtualized reality based telemicromanipulation, we present in this paper a visual calibration technique for optical microscope coupled with a CCD camera. In contrast to previous approaches, a virtual calibration pattern is constructed using the micromanipulator with a sub-pixel localization in the image. We also present a new camera calibration algorithm based on Parallel Single-Plane properties. The proposed procedure leads to a linear system from which the solution gives directly both intrinsic and extrinsic parameters of the geometrical model. Both computer simulation and real data have been used to test the proposed technique, and very good results have been obtained. Compared with classical techniques, our method provides an alternative technical solution, easy to use and flexible in the context of micromanipulation and virtual reality.
Mehdi Ammi, Vincent Frémont, Antoine Ferreira
ICRA2