Steve Bourgeois

dblp:34/4622 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0006-1150-7941ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 15 · 5 since 2021Systems, architecture and hardware · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 42% Segmentation and scene understanding · 19% Vision and language · 17%
Computer graphics and multimedia
4 papers
Virtual and augmented reality · 59% Geometric modeling and processing · 30% Computational photography and imaging · 11%

Topics — the 30 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
neural radiance field
1.622025
DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF · ICCV 2025
RING-NeRF : Rethinking Inductive Biases for Versatile and Efficient Neural Fields · ECCV (35) 2024
Computer vision › Vision and language › 3d vision and language
3d language grounding
1.012026
LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMs · AAAI 2026
Computer vision › Vision and language › 3d vision and language
3d question answering
1.012026
LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMs · AAAI 2026
Computer vision › 3D vision
3d scene understanding
1.012026
LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMs · AAAI 2026
Computer vision › Segmentation and scene understanding
3d semantic segmentation
0.912025
DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF · ICCV 2025
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.912025
DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF · ICCV 2025
Machine learning › Representation and self-supervised learning
inductive biases
0.812024
RING-NeRF : Rethinking Inductive Biases for Versatile and Efficient Neural Fields · ECCV (35) 2024
Computer vision › Image recognition and object detection › object localization
multi-scale object localization
0.812024
Introducing CEA-IMSOLD: an Industrial Multi-Scale Object Localization Dataset · ICRA 2024
Computer vision › 3D vision › implicit neural representation
neural field
0.812024
RING-NeRF : Rethinking Inductive Biases for Versatile and Efficient Neural Fields · ECCV (35) 2024
Virtual and augmented reality
augmented reality
0.432014
Markerless augmented reality solution for industrial manufacturing · ISMAR 2014
An interactive Augmented Reality system: A prototype for industrial maintenance training applications · ISMAR 2012
Augmented reality in large environments: Application to aided navigation in urban context · ISMAR 2010
Computer vision › Segmentation and scene understanding › 3d semantic segmentation
unsupervised 3d semantic segmentation
0.312025
DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF · ICCV 2025
Computer vision › Segmentation and scene understanding › image segmentation
unsupervised segmentation
0.312025
DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF · ICCV 2025
Robotics › Robot manipulation
grasping
0.212024
Introducing CEA-IMSOLD: an Industrial Multi-Scale Object Localization Dataset · ICRA 2024
Robotics › Robot navigation and mapping › SLAM › visual SLAM
monocular SLAM
0.222010
Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models · CVPR 2010
Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009
Robotics › Robot navigation and mapping
SLAM
0.222010
Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models · CVPR 2010
Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009
Geometric modeling and processing › computer-aided design
CAD model processing
0.212014
Markerless augmented reality solution for industrial manufacturing · ISMAR 2014
Virtual and augmented reality › augmented reality › augmented reality applications
industrial augmented reality
0.212014
Markerless augmented reality solution for industrial manufacturing · ISMAR 2014
Computer vision › 3D vision › 3d localization
6-dof localization
0.212013
Fast and automatic city-scale environment modeling for an accurate 6DOF vehicle localization · ISMAR 2013
Virtual and augmented reality
immersive interaction
0.112012
An interactive Augmented Reality system: A prototype for industrial maintenance training applications · ISMAR 2012
Computational photography and imaging
camera localization
0.112011
NonLinear refinement of structure from motion reconstruction by taking advantage of a partial knowledge of the environment · CVPR 2011
Geometric modeling and processing › 3d reconstruction
structure from motion
0.112011
NonLinear refinement of structure from motion reconstruction by taking advantage of a partial knowledge of the environment · CVPR 2011
Robotics › Robot navigation and mapping › localization › vision-based localization
monocular localization
0.112010
Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models · CVPR 2010
Robotics › Robot navigation and mapping › visual odometry
scale drift correction
0.112010
Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models · CVPR 2010
Robotics › Robot navigation and mapping › localization
vehicle localization
0.112010
Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models · CVPR 2010
Computer vision › 3D vision
visual localization
0.112010
Augmented reality in large environments: Application to aided navigation in urban context · ISMAR 2010
Computer vision › 3D vision
3d reconstruction
0.112009
Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009
Computer vision › 3D vision › structure from motion
bundle adjustment
0.112009
Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009
Robotics › Robot navigation and mapping
drift reduction
0.112009
Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009
Computer vision › 3D vision
geometric constraints
0.112009
Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009
Geometric modeling and processing
3d reconstruction
0.012011
NonLinear refinement of structure from motion reconstruction by taking advantage of a partial knowledge of the environment · CVPR 2011

Methods — techniques the papers use, named apart from their topics

vision-language model · 1.0multi-view 3d reconstruction · 1.0weak open-vocabulary guidance · 0.9unsupervised clustering · 0.9neural field · 0.9neural radiance field · 0.8inductive biases · 0.8BOP format · 0.8markerless tracking · 0.3GIS · 0.3bundle adjustment · 0.2CAD overlay · 0.2structure from motion · 0.2optical see-through HMD · 0.1nonlinear refinement · 0.1markov random field · 0.1graph cuts · 0.1belief propagation · 0.1
YearPublicationVenuePosition
2026 LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMs
abstract
Developing a multi-modal language model capable of understanding 3D scenes remains challenging due to the limited availability of 3D training data, in contrast to the abundance of 2D datasets used for vision-language models (VLMs). As an alternative, we introduce LLaVA³ (pronounced LLaVA Cube), a novel method that improves the 3D scene understanding capabilities of VLMs using only multi-view 2D images, and without requiring any fine-tuning. Inspired by Cubist painters, who represented multiple viewpoints of a 3D object within a single 2D picture, we propose to describe the 3D scene for the VLM through omnidirectional visual representations of each object. These representations are derived from an intermediate multi-view 3D reconstruction of the scene. Extensive experiments on 3D visual question answering and 3D language grounding show that our approach significantly outperforms previous 2D-based VLM solutions.
Doriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot, Loïc Barthe
AAAI2
2026 BOP-Distrib: Revisiting 6D Pose Estimation Benchmarks for Better Evaluation under Visual Ambiguities
abstract
6D pose estimation aims at determining the object pose that best explains the camera observation. The unique solution for non-ambiguous objects can turn into a multimodal pose distribution for symmetrical objects or when occlusions of symmetry-breaking elements happen, depending on the viewpoint. Currently, 6D pose estimation methods are benchmarked on datasets that consider, for their ground truth annotations, visual ambiguities as only related to global object symmetries, whereas they should be defined per-image to account for the camera viewpoint. We thus first propose an automatic method to re-annotate those datasets with a 6D pose distribution specific to each image, taking into account the object surface visibility in the image to correctly determine the visual ambiguities. Second, given this improved ground truth, we re-evaluate the state-of-the-art single pose methods and show that this greatly modifies the ranking of these methods. Third, as some recent works focus on estimating the complete set of solutions, we derive a precision/recall formulation to evaluate them against our image-wise distribution ground truth, making it the first benchmark for pose distribution methods on real images.
Boris Meden, Asma Brazi, Fabrice Mayran de Chamisso, Steve Bourgeois, Vincent Lepetit
WACV4
2025 DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF
abstract
3D semantic segmentation provides high-level scene understanding for applications in robotics, autonomous systems, \textit{etc}. Traditional methods adapt exclusively to either task-specific goals (open-vocabulary segmentation) or scene content (unsupervised semantic segmentation). We propose DiSCO-3D, the first method addressing the broader problem of 3D Open-Vocabulary Sub-concepts Discovery, which aims to provide a 3D semantic segmentation that adapts to both the scene and user queries. We build DiSCO-3D on Neural Fields representations, combining unsupervised segmentation with weak open-vocabulary guidance. Our evaluations demonstrate that DiSCO-3D achieves effective performance in Open-Vocabulary Sub-concepts Discovery and exhibits state-of-the-art results in the edge cases of both open-vocabulary and unsupervised segmentation.
Doriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot, Loïc Barthe
ICCV2
2024 RING-NeRF : Rethinking Inductive Biases for Versatile and Efficient Neural Fields
Doriand Petit, Steve Bourgeois, Dumitru Pavel, Vincent Gay-Bellile, Florian Chabot, Loïc Barthe
ECCV (35)2
2024 Introducing CEA-IMSOLD: an Industrial Multi-Scale Object Localization Dataset
abstract
We introduce the CEA Industrial Multi-Scale Object Localization Dataset (CEA-IMSOLD), a new BOP format dataset for 6-DoF object localization, crucial for robotics. This dataset aims to evaluate the current localization methods with respect to a new difficulty: large variations in observation distance and, consequently, large variations in image appearance. Compared to the other publicly available datasets, our dataset provides both images with objects small and completely visible in the image, and images where objects are observed close enough so they appear larger than the field of view of the camera. We also propose to consider the observation distance in the evaluation process and introduce new metrics to do so. Finally, our dataset contains a large variety of industrial objects, from small and simple objects such as bolts to sizable and complex ones such as large car parts. We provide baseline results and the dataset is made publicly available to support the community at https://cea-list.github.io/CEA-IMSOLD/.
Boris Meden, Pablo Vega, Fabrice Mayran de Chamisso, Steve Bourgeois
ICRA4
2023 MagHT: A Magnetic Hough Transform for Fast Indoor Place Recognition
abstract
This article proposes a novel indoor magnetic field-based place recognition algorithm that is accurate and fast to compute. For that, we modified the generalized “Hough Transform” to process magnetic data (MagHT). It takes as input a sequence of magnetic measures whose relative positions are recovered by an odometry system and recognizes the places in the magnetic map where they were acquired. It also returns the global transformation from the coordinate frame of the input magnetic data to the magnetic map reference frame. Experimental results on several real datasets in large indoor environments demonstrate that the obtained localization error, recall, and precision are similar to or are better than state-of-the-art methods while improving the runtime by several orders of magnitude. Moreover, unlike magnetic sequence matching-based solutions such as DTW, our approach is independent of the path taken during the magnetic map creation.
Iad Abdul Raouf, Vincent Gay-Bellile, Steve Bourgeois, Cyril Joly, Alexis Paljic
IROS3
2018 Localization of 3D objects using model-constrained SLAM
Angélique Loesch, Steve Bourgeois, Vincent Gay-Bellile, Olivier Gomez, Michel Dhome
Mach. Vis. Appl.2
2017 Large-scale, drift-free SLAM using highly robustified building model constraints
abstract
Constrained key-frame based local bundle adjustment is at the core of many recent systems that address the problem of large-scale, georeferenced SLAM based on a monocular camera and on data from inexpensive sensors and/or databases. The majority of these methods, however, impose constraints that result from proprioceptive sensors (e.g. IMUs, GPS, Odometry) while ignoring the possibility of explicitly constraining the structure (e.g. point cloud) resulting from the reconstruction process. Moreover, research on on-line interactions between SLAM and deep learning methods remains scarce, and as a result, few SLAM systems take advantage of deep architectures. We explore both these areas in this work: we use a fast deep neural network to infer semantic and structural information about the environment, and using a Bayesian framework, inject the results into a bundle adjustment process that constrains the 3d point cloud to texture-less 3d building models.
Achkan Salehi, Vincent Gay-Bellile, Steve Bourgeois, Nicolas Allezard, Frédéric Chausse
IROS3
2017 A hybrid bundle adjustment/pose-graph approach to VSLAM/GPS fusion for low-capacity platforms
abstract
We focus on the real-time fusion of monocular visual SLAM with GPS data in order to obtain city-scale, georeferenced pose estimations and reconstructions. Recently, GPS/VSLAM fusion through constrained local key-frame based Bundle Adjustment (BA) using Barrier Term Optimization (BTO) has proven to be (to the best of our knowledge) the most robust and accurate method. However, this approach requires a higher number of cameras to be considered in the optimization: in practice, more than 30 cameras are necessary, while a typical vision-only BA can succeed with as few as 10 cameras. This problem dimensionality makes the method unsuitable for autonomous embedded platforms of low computational capacity (e.g. MAVs). In this paper, we present a hybrid constrained BA/pose-graph approach using BTO, which is motivated by theoretical observations about covariance changes as a function of the gauge. We show that our method has desirable properties that allows its successful use in a BTO context, and present two different formulations. The experimental validation of our method shows that both our formulations reduce the computational cost in comparison with constrained BA using BTO, without any significant loss of precision. In particular, our first formulation yields a 60% reduction in execution time.
Achkan Salehi, Vincent Gay-Bellile, Steve Bourgeois, Frédéric Chausse
Intelligent Vehicles Symposium3
2016 A Hybrid Structure/Trajectory Constraint for Visual SLAM
abstract
This paper presents a hybrid structure/trajectory constraint, that uses output camera poses of a model-based tracker, for object localization with SLAM algorithm. This constraint takes into account the structure information given by a CAD model while relying on the formalism of trajectory constraints. It has the advantages to be compact in memory and to accelerate the SLAM optimization process. The accuracy and robustness of the resulting localization as well as the memory and time gains are evaluated on synthetic and real data. Videos are available as supplementary material.
Angélique Loesch, Steve Bourgeois, Vincent Gay-Bellile, Michel Dhome
3DV2
2016 The constrained SLAM framework for non-instrumented augmented reality - Application to industrial training
Mohamed Tamaazousti, Sylvie Naudet-Collette, Vincent Gay-Bellile, Steve Bourgeois, Bassem Besbes, Michel Dhome
Multim. Tools Appl.4
2016 Fast and automatic city-scale environment modelling using hard and/or weak constrained bundle adjustments
Dorra Larnaout, Vincent Gay-Bellile, Steve Bourgeois, Michel Dhome
Mach. Vis. Appl.3
2015 Generic edgelet-based tracking of 3D objects in real-time
abstract
This paper addresses the challenging issue of real-time camera localization relative to any object that have texture or not, sharp edges or occluding contours. 3D contour points, dynamically extracted from a CAD model by Analysis-by-Synthesis on the graphics hardware, are combined with a keyframe-based SLAM algorithm to estimate camera poses. Our tracking solution is accurate, robust to sudden motions and to occlusions, as demonstrated on synthetic and real data. This solution is also easy to deploy since it only uses an RGB camera and a CAD model of the object of interest, requires no manual intervention on this model and runs on a consumer tablet at a frequency of 40Hz on a HD video-stream. Videos are available as supplemental material.
Angélique Loesch, Steve Bourgeois, Vincent Gay-Bellile, Michel Dhome
IROS2
2015 Noise modelling in time-of-flight sensors with application to depth noise removal and uncertainty estimation in three-dimensional measurement
abstract
Time‐of‐flight (TOF) sensors provide real‐time depth information at high frame‐rates. One issue with TOF sensors is the usual high level of noise (i.e . the depth measure's repeatability within a static setting). However, until now, TOF sensors’ noise has not been well studied. The authors show that the commonly agreed hypothesis that noise depends only on the amplitude information is not valid in practice. They empirically establish that the noise follows a signal‐dependent Gaussian distribution and varies according to pixel position, depth and integration time. They thus consider all these factors to model noise in two new noise models. Both models are evaluated, compared and used in the two following applications: depth noise removal by depth filtering and uncertainty (repeatability) estimation in three‐dimensional measurement.
Amira Belhedi, Adrien Bartoli, Steve Bourgeois, Vincent Gay-Bellile, Kamel Hamrouni, Patrick Sayd
IET Comput. Vis.3
2014 Vision-Based Differential GPS: Improving VSLAM / GPS Fusion in Urban Environment with 3D Building Models
abstract
We improve in this paper the localization accuracy of visual SLAM (VSLAM) / GPS fusion in dense urban area by using 3D building models provided by Geographic Information System (GIS). GPS inaccuracies are corrected by comparison of the reconstruction resulting from the VSLAM / GPS fusion with 3D building models. These corrected GPS data are thereafter re-injected in the fusion process. Experimental results demonstrate the accuracy improvements achieved through our proposed solution.
Dorra Larnaout, Vincent Gay-Bellile, Steve Bourgeois, Michel Dhome
3DV3
2014 Markerless augmented reality solution for industrial manufacturing
abstract
We present a comprehensive augmented reality solution to efficiently perform different verification pocedures during industrial assembly-line manufacturing using CAD data from product lifecycle management (PLM) systems. The demonstration focuses on a variety of industrial use-cases that have to go through different control processes, assisted by augmented reality tools. Thus, the experience has to be precise, robust to natural movements and easy to realise for an assembly-line technician.
Boris Meden, Sebastian Knödel, Steve Bourgeois
ISMAR3
2013 Vehicle 6-DoF localization based on SLAM constrained by GPS and digital elevation model information
abstract
Vehicle geo-localization based on monocular visual Simultaneous Localization And Mapping (SLAM) remains a challenging issue mainly due to the accumulation errors and scale factor drift. To tackle these limitations, a common solution is to introduce geo-referenced information into the visual SLAM algorithm. In this paper, we propose two different bundle adjustment processes that merge both GPS measurements and “Digital Elevation Model” (DEM) data. Proposed solutions are devoted to ensure an accurate and robust geo-localization in both rural and urban environment. Experiments on synthetic and large scale real sequences show that, in addition to the real-time (i.e. about 30 Hz) performances, we obtain an accurate 6DoF localization.
Dorra Larnaout, Vincent Gay-Bellile, Steve Bourgeois, Michel Dhome
ICIP3
2013 Fast and automatic city-scale environment modeling for an accurate 6DOF vehicle localization
abstract
To provide high quality augmented reality service in a car navigation system, accurate 6DoF localization is required. To ensure such accuracy, most of current vision-based solutions rely on an off-line large scale modeling of the environment. Nevertheless, while existing solutions require expensive equipments and/or a prohibitive computation time, we propose in this paper a complete framework that automatically builds an accurate city scale database using only a standard camera, a GPS and Geographic Information System (GIS). As illustrated in the experiments, only few minutes are required to model large scale environments. The resulting databases can then be used during a localization algorithm for high quality Augmented Reality experiences.
Dorra Larnaout, Vincent Gay-Bellile, Steve Bourgeois, Benjamin Labbé, Michel Dhome
ISMAR3
2012 Depth Correction for Depth Camera From Planarity
abstract
Depth cameras open new possibilities in fields such as 3D reconstruction, Augmented Reality and video-surveillance since they provide depth information at high frame-rates. However, like any sensor, they have limitations related to their technology. One of them is depth distortion. In this paper, we present a method to estimate depth correction for depth cameras. The proposed method is based on two steps. The first one is a nonplanarity correction that needs depth measurement of different plane views. The second one is an affinity correction that,contrary to state of the art approaches, requires a very small set of ground truth measurements. Thus, it is more easy to use compared to other methods and does not need a large set of accurate ground truth that is extremely difficult to obtain in practice. Experiments on both simulated and real data show that the proposed approach improve also the depth accuracy compare to state of the art methods.
Amira Belhedi, Adrien Bartoli, Vincent Gay-Bellile, Steve Bourgeois, Patrick Sayd, Kamel Hamrouni
BMVC4
2012 Non-parametric depth calibration of a TOF camera
abstract
Time-of-Flight (TOF) cameras measure, in real-time, the distance between the camera and objects in the scene. This opens new perspectives in different application fields: 3D reconstruction, Augmented Reality, video-surveillance, etc. However, like any sensor, TOF cameras have limitations related to their technology. One of them is distance distortion. In this paper, we present a new depth calibration method (estimation of distance distortion) for TOF cameras. Our approach has several advantages. First, it is based on a non-parametric model, contrary to most of the other methods. Second, it models under the same formalism the distortion variation according to the distance and the pixel position in the image. This improves calibration accuracy even at the image boundaries which are typically more distorted than the image center. A comparison with two state of the art parametric methods is presented.
Amira Belhedi, Steve Bourgeois, Vincent Gay-Bellile, Patrick Sayd, Adrien Bartoli, Kamel Hamrouni
ICIP2
2012 An interactive Augmented Reality system: A prototype for industrial maintenance training applications
abstract
In this paper, we present an innovative Augmented Reality prototype designed for industrial education and training applications. The system uses an Optical See-Through HMD integrating a calibrated camera and a laser pointer to interactively augment an industrial object with virtual sequences designed to train a user for specific maintenance tasks. The training leverages user interactions by simply pointing on a specific object component. The architecture of our prototype involves two main vision-based modules : camera localization and user-interaction handling. The first module includes markerless trackers for camera localization, which can deal with partial occlusions and specular reflections on the metallic object surfaces. In the second module, we developed fast image processing methods for red laser dot tracking. By combining these processing elements, the proposed system is able to interactively augment in real time an industrial object making the learning process more interesting and intuitive.
Bassem Besbes, Sylvie Naudet-Collette, Mohamed Tamaazousti, Steve Bourgeois, Vincent Gay-Bellile
ISMAR4
2011 NonLinear refinement of structure from motion reconstruction by taking advantage of a partial knowledge of the environment
abstract
We address the challenging issue of camera localization in a partially known environment, i.e. for which a geometric 3D model that covers only a part of the observed scene is available. When this scene is static, both known and unknown parts of the environment provide constraints on the camera motion. This paper proposes a nonlinear refinement process of an initial SfM reconstruction that takes advantage of these two types of constraints. Compare to those that exploit only the model constraints i.e. the known part of the scene, including the unknown part of the environment in the optimization process yields a faster, more accurate and robust refinement. It also presents a much larger convergence basin. This paper will demonstrate these statements on varied synthetic and real sequences for both 3D object tracking and outdoor localization applications.
Mohamed Tamaazousti, Vincent Gay-Bellile, Sylvie Naudet-Collette, Steve Bourgeois, Michel Dhome
CVPR4
2010 Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models
abstract
In this system paper, we propose a real-time car localisation process in dense urban areas by using a single perspective camera and a priori on the environment. To tackle this problem, it is necessary to solve two well-known monocular SLAM limitations: scale factor drift and error accumulation. The proposed idea is to combine a monocular SLAM process based on bundle adjustment with simple knowledge, i.e. the position and orientation of the camera with regard to the road and a coarse 3D model of the environment, as those provided by GIS database. First, we show that, thanks to specific SLAM-based constraints, the road homography can be expressed only with respect to the scale factor parameter. This allows the scale factor to be robustly and frequently estimated. Then, we propose to use the global information brought by 3D city models in order to correct the monocular SLAM error accumulation. Even with coarse 3D models, turnings give enough geometrical constraints to allow fitting the reconstructed 3D point cloud with the 3D model. Experiments on large-scale sequences (several kilometres) show that the entire process permits the real-time localisation of a car in city centre, even in real traffic condition.
Pierre Lothe, Steve Bourgeois, Eric Royer, Michel Dhome, Sylvie Naudet-Collette
CVPR2
2010 Augmented reality in large environments: Application to aided navigation in urban context
abstract
This paper addresses the challenging issue of vision-based localization in urban context. It briefly describes our contributions in large environments modeling and accurate camera localization. The efficiency of the resulting system is illustrated through Augmented Reality results on large trajectory of several hundred meters.
Vincent Gay-Bellile, Pierre Lothe, Steve Bourgeois, Eric Royer, Sylvie Naudet-Collette
ISMAR3
2009 Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization
abstract
In the past few years, lots of works were achieved on Simultaneous Localization and Mapping (SLAM). It is now possible to follow in real time the trajectory of a moving camera in an unknown environment. However, current SLAM methods are still prone to drift errors, which prevent their use in large-scale applications. In this paper, we propose a solution to reduce those errors a posteriori. Our solution is based on a postprocessing algorithm that exploits additional geometric constraints, relative to the environment, to correct both the reconstructed geometry and the camera trajectory. These geometric constraints are obtained through a coarse 3D modelisation of the environment, similar to those provided by GIS database. First, we propose an original articulated transformation model in order to roughly align the SLAM reconstruction with this 3D model through a non-rigid ICP step. Then, to refine the reconstruction, we introduce a new bundle adjustment cost function that includes, in a single term, the usual 3D point/ID observation consistency constraint as well as the geometric constraints provided by the 3D model. Results on large-scale synthetic and real sequences show that our method successfully improves SLAM reconstructions. Besides, experiments prove that the resulting reconstruction is accurate enough to be directly used for global relocalization applications.
Pierre Lothe, Steve Bourgeois, Fabien Dekeyser, Eric Royer, Michel Dhome
CVPR2
2005 A Practical Guide to Marker Based and Hybrid Visual Registration for AR Industrial Applications
Steve Bourgeois, Hanna Martinsson, Quoc Cuong Pham, Sylvie Naudet
CAIP1