EDBT 2026 Demo / reviewers in the wild / expert
Steve Bourgeois
dblp:34/4622
· DBLP profile ↗
26ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0006-1150-7941ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 15 · 5 since 2021Systems, architecture and hardware · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 4
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
3D vision · 42% Segmentation and scene understanding · 19% Vision and language · 17% | |
| Computer graphics and multimedia
4 papers |
Virtual and augmented reality · 59% Geometric modeling and processing · 30% Computational photography and imaging · 11% |
Topics — the 30 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
neural radiance field |
1.6 | 2 | 2025 | DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF · ICCV 2025 RING-NeRF : Rethinking Inductive Biases for Versatile and Efficient Neural Fields · ECCV (35) 2024 |
Computer vision › Vision and language › 3d vision and language
3d language grounding |
1.0 | 1 | 2026 | LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMs · AAAI 2026 |
Computer vision › Vision and language › 3d vision and language
3d question answering |
1.0 | 1 | 2026 | LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMs · AAAI 2026 |
Computer vision › 3D vision
3d scene understanding |
1.0 | 1 | 2026 | LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMs · AAAI 2026 |
Computer vision › Segmentation and scene understanding
3d semantic segmentation |
0.9 | 1 | 2025 | DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF · ICCV 2025 |
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation |
0.9 | 1 | 2025 | DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF · ICCV 2025 |
Machine learning › Representation and self-supervised learning
inductive biases |
0.8 | 1 | 2024 | RING-NeRF : Rethinking Inductive Biases for Versatile and Efficient Neural Fields · ECCV (35) 2024 |
Computer vision › Image recognition and object detection › object localization
multi-scale object localization |
0.8 | 1 | 2024 | Introducing CEA-IMSOLD: an Industrial Multi-Scale Object Localization Dataset · ICRA 2024 |
Computer vision › 3D vision › implicit neural representation
neural field |
0.8 | 1 | 2024 | RING-NeRF : Rethinking Inductive Biases for Versatile and Efficient Neural Fields · ECCV (35) 2024 |
Virtual and augmented reality
augmented reality |
0.4 | 3 | 2014 | Markerless augmented reality solution for industrial manufacturing · ISMAR 2014 An interactive Augmented Reality system: A prototype for industrial maintenance training applications · ISMAR 2012 Augmented reality in large environments: Application to aided navigation in urban context · ISMAR 2010 |
Computer vision › Segmentation and scene understanding › 3d semantic segmentation
unsupervised 3d semantic segmentation |
0.3 | 1 | 2025 | DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF · ICCV 2025 |
Computer vision › Segmentation and scene understanding › image segmentation
unsupervised segmentation |
0.3 | 1 | 2025 | DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRF · ICCV 2025 |
Robotics › Robot manipulation
grasping |
0.2 | 1 | 2024 | Introducing CEA-IMSOLD: an Industrial Multi-Scale Object Localization Dataset · ICRA 2024 |
Robotics › Robot navigation and mapping › SLAM › visual SLAM
monocular SLAM |
0.2 | 2 | 2010 | Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models · CVPR 2010 Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009 |
Robotics › Robot navigation and mapping
SLAM |
0.2 | 2 | 2010 | Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models · CVPR 2010 Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009 |
Geometric modeling and processing › computer-aided design
CAD model processing |
0.2 | 1 | 2014 | Markerless augmented reality solution for industrial manufacturing · ISMAR 2014 |
Virtual and augmented reality › augmented reality › augmented reality applications
industrial augmented reality |
0.2 | 1 | 2014 | Markerless augmented reality solution for industrial manufacturing · ISMAR 2014 |
Computer vision › 3D vision › 3d localization
6-dof localization |
0.2 | 1 | 2013 | Fast and automatic city-scale environment modeling for an accurate 6DOF vehicle localization · ISMAR 2013 |
Virtual and augmented reality
immersive interaction |
0.1 | 1 | 2012 | An interactive Augmented Reality system: A prototype for industrial maintenance training applications · ISMAR 2012 |
Computational photography and imaging
camera localization |
0.1 | 1 | 2011 | NonLinear refinement of structure from motion reconstruction by taking advantage of a partial knowledge of the environment · CVPR 2011 |
Geometric modeling and processing › 3d reconstruction
structure from motion |
0.1 | 1 | 2011 | NonLinear refinement of structure from motion reconstruction by taking advantage of a partial knowledge of the environment · CVPR 2011 |
Robotics › Robot navigation and mapping › localization › vision-based localization
monocular localization |
0.1 | 1 | 2010 | Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models · CVPR 2010 |
Robotics › Robot navigation and mapping › visual odometry
scale drift correction |
0.1 | 1 | 2010 | Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models · CVPR 2010 |
Robotics › Robot navigation and mapping › localization
vehicle localization |
0.1 | 1 | 2010 | Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city models · CVPR 2010 |
Computer vision › 3D vision
visual localization |
0.1 | 1 | 2010 | Augmented reality in large environments: Application to aided navigation in urban context · ISMAR 2010 |
Computer vision › 3D vision
3d reconstruction |
0.1 | 1 | 2009 | Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009 |
Computer vision › 3D vision › structure from motion
bundle adjustment |
0.1 | 1 | 2009 | Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009 |
Robotics › Robot navigation and mapping
drift reduction |
0.1 | 1 | 2009 | Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009 |
Computer vision › 3D vision
geometric constraints |
0.1 | 1 | 2009 | Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localization · CVPR 2009 |
Geometric modeling and processing
3d reconstruction |
0.0 | 1 | 2011 | NonLinear refinement of structure from motion reconstruction by taking advantage of a partial knowledge of the environment · CVPR 2011 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 1.0multi-view 3d reconstruction · 1.0weak open-vocabulary guidance · 0.9unsupervised clustering · 0.9neural field · 0.9neural radiance field · 0.8inductive biases · 0.8BOP format · 0.8markerless tracking · 0.3GIS · 0.3bundle adjustment · 0.2CAD overlay · 0.2structure from motion · 0.2optical see-through HMD · 0.1nonlinear refinement · 0.1markov random field · 0.1graph cuts · 0.1belief propagation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMsabstractDeveloping a multi-modal language model capable of understanding 3D scenes remains challenging due to the limited availability of 3D training data, in contrast to the abundance of 2D datasets used for vision-language models (VLMs). As an alternative, we introduce LLaVA³ (pronounced LLaVA Cube), a novel method that improves the 3D scene understanding capabilities of VLMs using only multi-view 2D images, and without requiring any fine-tuning. Inspired by Cubist painters, who represented multiple viewpoints of a 3D object within a single 2D picture, we propose to describe the 3D scene for the VLM through omnidirectional visual representations of each object. These representations are derived from an intermediate multi-view 3D reconstruction of the scene. Extensive experiments on 3D visual question answering and 3D language grounding show that our approach significantly outperforms previous 2D-based VLM solutions. Doriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot, Loïc Barthe |
AAAI | 2 |
| 2026 | BOP-Distrib: Revisiting 6D Pose Estimation Benchmarks for Better Evaluation under Visual Ambiguitiesabstract6D pose estimation aims at determining the object pose that best explains the camera observation. The unique solution for non-ambiguous objects can turn into a multimodal pose distribution for symmetrical objects or when occlusions of symmetry-breaking elements happen, depending on the viewpoint. Currently, 6D pose estimation methods are benchmarked on datasets that consider, for their ground truth annotations, visual ambiguities as only related to global object symmetries, whereas they should be defined per-image to account for the camera viewpoint. We thus first propose an automatic method to re-annotate those datasets with a 6D pose distribution specific to each image, taking into account the object surface visibility in the image to correctly determine the visual ambiguities. Second, given this improved ground truth, we re-evaluate the state-of-the-art single pose methods and show that this greatly modifies the ranking of these methods. Third, as some recent works focus on estimating the complete set of solutions, we derive a precision/recall formulation to evaluate them against our image-wise distribution ground truth, making it the first benchmark for pose distribution methods on real images. Boris Meden, Asma Brazi, Fabrice Mayran de Chamisso, Steve Bourgeois, Vincent Lepetit |
WACV | 4 |
| 2025 | DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-Vocabulary Queries in NeRFabstract3D semantic segmentation provides high-level scene understanding for applications in robotics, autonomous systems, \textit{etc}. Traditional methods adapt exclusively to either task-specific goals (open-vocabulary segmentation) or scene content (unsupervised semantic segmentation). We propose DiSCO-3D, the first method addressing the broader problem of 3D Open-Vocabulary Sub-concepts Discovery, which aims to provide a 3D semantic segmentation that adapts to both the scene and user queries. We build DiSCO-3D on Neural Fields representations, combining unsupervised segmentation with weak open-vocabulary guidance. Our evaluations demonstrate that DiSCO-3D achieves effective performance in Open-Vocabulary Sub-concepts Discovery and exhibits state-of-the-art results in the edge cases of both open-vocabulary and unsupervised segmentation. Doriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot, Loïc Barthe |
ICCV | 2 |
| 2024 | RING-NeRF : Rethinking Inductive Biases for Versatile and Efficient Neural Fields
Doriand Petit, Steve Bourgeois, Dumitru Pavel, Vincent Gay-Bellile, Florian Chabot, Loïc Barthe |
ECCV (35) | 2 |
| 2024 | Introducing CEA-IMSOLD: an Industrial Multi-Scale Object Localization DatasetabstractWe introduce the CEA Industrial Multi-Scale Object Localization Dataset (CEA-IMSOLD), a new BOP format dataset for 6-DoF object localization, crucial for robotics. This dataset aims to evaluate the current localization methods with respect to a new difficulty: large variations in observation distance and, consequently, large variations in image appearance. Compared to the other publicly available datasets, our dataset provides both images with objects small and completely visible in the image, and images where objects are observed close enough so they appear larger than the field of view of the camera. We also propose to consider the observation distance in the evaluation process and introduce new metrics to do so. Finally, our dataset contains a large variety of industrial objects, from small and simple objects such as bolts to sizable and complex ones such as large car parts. We provide baseline results and the dataset is made publicly available to support the community at https://cea-list.github.io/CEA-IMSOLD/. Boris Meden, Pablo Vega, Fabrice Mayran de Chamisso, Steve Bourgeois |
ICRA | 4 |
| 2023 | MagHT: A Magnetic Hough Transform for Fast Indoor Place RecognitionabstractThis article proposes a novel indoor magnetic field-based place recognition algorithm that is accurate and fast to compute. For that, we modified the generalized “Hough Transform” to process magnetic data (MagHT). It takes as input a sequence of magnetic measures whose relative positions are recovered by an odometry system and recognizes the places in the magnetic map where they were acquired. It also returns the global transformation from the coordinate frame of the input magnetic data to the magnetic map reference frame. Experimental results on several real datasets in large indoor environments demonstrate that the obtained localization error, recall, and precision are similar to or are better than state-of-the-art methods while improving the runtime by several orders of magnitude. Moreover, unlike magnetic sequence matching-based solutions such as DTW, our approach is independent of the path taken during the magnetic map creation. Iad Abdul Raouf, Vincent Gay-Bellile, Steve Bourgeois, Cyril Joly, Alexis Paljic |
IROS | 3 |
| 2018 | Localization of 3D objects using model-constrained SLAM
Angélique Loesch, Steve Bourgeois, Vincent Gay-Bellile, Olivier Gomez, Michel Dhome |
Mach. Vis. Appl. | 2 |
| 2017 | Large-scale, drift-free SLAM using highly robustified building model constraintsabstractConstrained key-frame based local bundle adjustment is at the core of many recent systems that address the problem of large-scale, georeferenced SLAM based on a monocular camera and on data from inexpensive sensors and/or databases. The majority of these methods, however, impose constraints that result from proprioceptive sensors (e.g. IMUs, GPS, Odometry) while ignoring the possibility of explicitly constraining the structure (e.g. point cloud) resulting from the reconstruction process. Moreover, research on on-line interactions between SLAM and deep learning methods remains scarce, and as a result, few SLAM systems take advantage of deep architectures. We explore both these areas in this work: we use a fast deep neural network to infer semantic and structural information about the environment, and using a Bayesian framework, inject the results into a bundle adjustment process that constrains the 3d point cloud to texture-less 3d building models. Achkan Salehi, Vincent Gay-Bellile, Steve Bourgeois, Nicolas Allezard, Frédéric Chausse |
IROS | 3 |
| 2017 | A hybrid bundle adjustment/pose-graph approach to VSLAM/GPS fusion for low-capacity platformsabstractWe focus on the real-time fusion of monocular visual SLAM with GPS data in order to obtain city-scale, georeferenced pose estimations and reconstructions. Recently, GPS/VSLAM fusion through constrained local key-frame based Bundle Adjustment (BA) using Barrier Term Optimization (BTO) has proven to be (to the best of our knowledge) the most robust and accurate method. However, this approach requires a higher number of cameras to be considered in the optimization: in practice, more than 30 cameras are necessary, while a typical vision-only BA can succeed with as few as 10 cameras. This problem dimensionality makes the method unsuitable for autonomous embedded platforms of low computational capacity (e.g. MAVs). In this paper, we present a hybrid constrained BA/pose-graph approach using BTO, which is motivated by theoretical observations about covariance changes as a function of the gauge. We show that our method has desirable properties that allows its successful use in a BTO context, and present two different formulations. The experimental validation of our method shows that both our formulations reduce the computational cost in comparison with constrained BA using BTO, without any significant loss of precision. In particular, our first formulation yields a 60% reduction in execution time. Achkan Salehi, Vincent Gay-Bellile, Steve Bourgeois, Frédéric Chausse |
Intelligent Vehicles Symposium | 3 |
| 2016 | A Hybrid Structure/Trajectory Constraint for Visual SLAMabstractThis paper presents a hybrid structure/trajectory constraint, that uses output camera poses of a model-based tracker, for object localization with SLAM algorithm. This constraint takes into account the structure information given by a CAD model while relying on the formalism of trajectory constraints. It has the advantages to be compact in memory and to accelerate the SLAM optimization process. The accuracy and robustness of the resulting localization as well as the memory and time gains are evaluated on synthetic and real data. Videos are available as supplementary material. Angélique Loesch, Steve Bourgeois, Vincent Gay-Bellile, Michel Dhome |
3DV | 2 |
| 2016 | The constrained SLAM framework for non-instrumented augmented reality - Application to industrial training
Mohamed Tamaazousti, Sylvie Naudet-Collette, Vincent Gay-Bellile, Steve Bourgeois, Bassem Besbes, Michel Dhome |
Multim. Tools Appl. | 4 |
| 2016 | Fast and automatic city-scale environment modelling using hard and/or weak constrained bundle adjustments
Dorra Larnaout, Vincent Gay-Bellile, Steve Bourgeois, Michel Dhome |
Mach. Vis. Appl. | 3 |
| 2015 | Generic edgelet-based tracking of 3D objects in real-timeabstractThis paper addresses the challenging issue of real-time camera localization relative to any object that have texture or not, sharp edges or occluding contours. 3D contour points, dynamically extracted from a CAD model by Analysis-by-Synthesis on the graphics hardware, are combined with a keyframe-based SLAM algorithm to estimate camera poses. Our tracking solution is accurate, robust to sudden motions and to occlusions, as demonstrated on synthetic and real data. This solution is also easy to deploy since it only uses an RGB camera and a CAD model of the object of interest, requires no manual intervention on this model and runs on a consumer tablet at a frequency of 40Hz on a HD video-stream. Videos are available as supplemental material. Angélique Loesch, Steve Bourgeois, Vincent Gay-Bellile, Michel Dhome |
IROS | 2 |
| 2015 | Noise modelling in time-of-flight sensors with application to depth noise removal and uncertainty estimation in three-dimensional measurementabstractTime‐of‐flight (TOF) sensors provide real‐time depth information at high frame‐rates. One issue with TOF sensors is the usual high level of noise (i.e . the depth measure's repeatability within a static setting). However, until now, TOF sensors’ noise has not been well studied. The authors show that the commonly agreed hypothesis that noise depends only on the amplitude information is not valid in practice. They empirically establish that the noise follows a signal‐dependent Gaussian distribution and varies according to pixel position, depth and integration time. They thus consider all these factors to model noise in two new noise models. Both models are evaluated, compared and used in the two following applications: depth noise removal by depth filtering and uncertainty (repeatability) estimation in three‐dimensional measurement. Amira Belhedi, Adrien Bartoli, Steve Bourgeois, Vincent Gay-Bellile, Kamel Hamrouni, Patrick Sayd |
IET Comput. Vis. | 3 |
| 2014 | Vision-Based Differential GPS: Improving VSLAM / GPS Fusion in Urban Environment with 3D Building ModelsabstractWe improve in this paper the localization accuracy of visual SLAM (VSLAM) / GPS fusion in dense urban area by using 3D building models provided by Geographic Information System (GIS). GPS inaccuracies are corrected by comparison of the reconstruction resulting from the VSLAM / GPS fusion with 3D building models. These corrected GPS data are thereafter re-injected in the fusion process. Experimental results demonstrate the accuracy improvements achieved through our proposed solution. Dorra Larnaout, Vincent Gay-Bellile, Steve Bourgeois, Michel Dhome |
3DV | 3 |
| 2014 | Markerless augmented reality solution for industrial manufacturingabstractWe present a comprehensive augmented reality solution to efficiently perform different verification pocedures during industrial assembly-line manufacturing using CAD data from product lifecycle management (PLM) systems. The demonstration focuses on a variety of industrial use-cases that have to go through different control processes, assisted by augmented reality tools. Thus, the experience has to be precise, robust to natural movements and easy to realise for an assembly-line technician. Boris Meden, Sebastian Knödel, Steve Bourgeois |
ISMAR | 3 |
| 2013 | Vehicle 6-DoF localization based on SLAM constrained by GPS and digital elevation model informationabstractVehicle geo-localization based on monocular visual Simultaneous Localization And Mapping (SLAM) remains a challenging issue mainly due to the accumulation errors and scale factor drift. To tackle these limitations, a common solution is to introduce geo-referenced information into the visual SLAM algorithm. In this paper, we propose two different bundle adjustment processes that merge both GPS measurements and “Digital Elevation Model” (DEM) data. Proposed solutions are devoted to ensure an accurate and robust geo-localization in both rural and urban environment. Experiments on synthetic and large scale real sequences show that, in addition to the real-time (i.e. about 30 Hz) performances, we obtain an accurate 6DoF localization. Dorra Larnaout, Vincent Gay-Bellile, Steve Bourgeois, Michel Dhome |
ICIP | 3 |
| 2013 | Fast and automatic city-scale environment modeling for an accurate 6DOF vehicle localizationabstractTo provide high quality augmented reality service in a car navigation system, accurate 6DoF localization is required. To ensure such accuracy, most of current vision-based solutions rely on an off-line large scale modeling of the environment. Nevertheless, while existing solutions require expensive equipments and/or a prohibitive computation time, we propose in this paper a complete framework that automatically builds an accurate city scale database using only a standard camera, a GPS and Geographic Information System (GIS). As illustrated in the experiments, only few minutes are required to model large scale environments. The resulting databases can then be used during a localization algorithm for high quality Augmented Reality experiences. Dorra Larnaout, Vincent Gay-Bellile, Steve Bourgeois, Benjamin Labbé, Michel Dhome |
ISMAR | 3 |
| 2012 | Depth Correction for Depth Camera From PlanarityabstractDepth cameras open new possibilities in fields such as 3D reconstruction, Augmented Reality and video-surveillance since they provide depth information at high frame-rates. However, like any sensor, they have limitations related to their technology. One of them is depth distortion. In this paper, we present a method to estimate depth correction for depth cameras. The proposed method is based on two steps. The first one is a nonplanarity correction that needs depth measurement of different plane views. The second one is an affinity correction that,contrary to state of the art approaches, requires a very small set of ground truth measurements. Thus, it is more easy to use compared to other methods and does not need a large set of accurate ground truth that is extremely difficult to obtain in practice. Experiments on both simulated and real data show that the proposed approach improve also the depth accuracy compare to state of the art methods. Amira Belhedi, Adrien Bartoli, Vincent Gay-Bellile, Steve Bourgeois, Patrick Sayd, Kamel Hamrouni |
BMVC | 4 |
| 2012 | Non-parametric depth calibration of a TOF cameraabstractTime-of-Flight (TOF) cameras measure, in real-time, the distance between the camera and objects in the scene. This opens new perspectives in different application fields: 3D reconstruction, Augmented Reality, video-surveillance, etc. However, like any sensor, TOF cameras have limitations related to their technology. One of them is distance distortion. In this paper, we present a new depth calibration method (estimation of distance distortion) for TOF cameras. Our approach has several advantages. First, it is based on a non-parametric model, contrary to most of the other methods. Second, it models under the same formalism the distortion variation according to the distance and the pixel position in the image. This improves calibration accuracy even at the image boundaries which are typically more distorted than the image center. A comparison with two state of the art parametric methods is presented. Amira Belhedi, Steve Bourgeois, Vincent Gay-Bellile, Patrick Sayd, Adrien Bartoli, Kamel Hamrouni |
ICIP | 2 |
| 2012 | An interactive Augmented Reality system: A prototype for industrial maintenance training applicationsabstractIn this paper, we present an innovative Augmented Reality prototype designed for industrial education and training applications. The system uses an Optical See-Through HMD integrating a calibrated camera and a laser pointer to interactively augment an industrial object with virtual sequences designed to train a user for specific maintenance tasks. The training leverages user interactions by simply pointing on a specific object component. The architecture of our prototype involves two main vision-based modules : camera localization and user-interaction handling. The first module includes markerless trackers for camera localization, which can deal with partial occlusions and specular reflections on the metallic object surfaces. In the second module, we developed fast image processing methods for red laser dot tracking. By combining these processing elements, the proposed system is able to interactively augment in real time an industrial object making the learning process more interesting and intuitive. Bassem Besbes, Sylvie Naudet-Collette, Mohamed Tamaazousti, Steve Bourgeois, Vincent Gay-Bellile |
ISMAR | 4 |
| 2011 | NonLinear refinement of structure from motion reconstruction by taking advantage of a partial knowledge of the environmentabstractWe address the challenging issue of camera localization in a partially known environment, i.e. for which a geometric 3D model that covers only a part of the observed scene is available. When this scene is static, both known and unknown parts of the environment provide constraints on the camera motion. This paper proposes a nonlinear refinement process of an initial SfM reconstruction that takes advantage of these two types of constraints. Compare to those that exploit only the model constraints i.e. the known part of the scene, including the unknown part of the environment in the optimization process yields a faster, more accurate and robust refinement. It also presents a much larger convergence basin. This paper will demonstrate these statements on varied synthetic and real sequences for both 3D object tracking and outdoor localization applications. Mohamed Tamaazousti, Vincent Gay-Bellile, Sylvie Naudet-Collette, Steve Bourgeois, Michel Dhome |
CVPR | 4 |
| 2010 | Real-time vehicle global localisation with a single camera in dense urban areas: Exploitation of coarse 3D city modelsabstractIn this system paper, we propose a real-time car localisation process in dense urban areas by using a single perspective camera and a priori on the environment. To tackle this problem, it is necessary to solve two well-known monocular SLAM limitations: scale factor drift and error accumulation. The proposed idea is to combine a monocular SLAM process based on bundle adjustment with simple knowledge, i.e. the position and orientation of the camera with regard to the road and a coarse 3D model of the environment, as those provided by GIS database. First, we show that, thanks to specific SLAM-based constraints, the road homography can be expressed only with respect to the scale factor parameter. This allows the scale factor to be robustly and frequently estimated. Then, we propose to use the global information brought by 3D city models in order to correct the monocular SLAM error accumulation. Even with coarse 3D models, turnings give enough geometrical constraints to allow fitting the reconstructed 3D point cloud with the 3D model. Experiments on large-scale sequences (several kilometres) show that the entire process permits the real-time localisation of a car in city centre, even in real traffic condition. Pierre Lothe, Steve Bourgeois, Eric Royer, Michel Dhome, Sylvie Naudet-Collette |
CVPR | 2 |
| 2010 | Augmented reality in large environments: Application to aided navigation in urban contextabstractThis paper addresses the challenging issue of vision-based localization in urban context. It briefly describes our contributions in large environments modeling and accurate camera localization. The efficiency of the resulting system is illustrated through Augmented Reality results on large trajectory of several hundred meters. Vincent Gay-Bellile, Pierre Lothe, Steve Bourgeois, Eric Royer, Sylvie Naudet-Collette |
ISMAR | 3 |
| 2009 | Towards geographical referencing of monocular SLAM reconstruction using 3D city models: Application to real-time accurate vision-based localizationabstractIn the past few years, lots of works were achieved on Simultaneous Localization and Mapping (SLAM). It is now possible to follow in real time the trajectory of a moving camera in an unknown environment. However, current SLAM methods are still prone to drift errors, which prevent their use in large-scale applications. In this paper, we propose a solution to reduce those errors a posteriori. Our solution is based on a postprocessing algorithm that exploits additional geometric constraints, relative to the environment, to correct both the reconstructed geometry and the camera trajectory. These geometric constraints are obtained through a coarse 3D modelisation of the environment, similar to those provided by GIS database. First, we propose an original articulated transformation model in order to roughly align the SLAM reconstruction with this 3D model through a non-rigid ICP step. Then, to refine the reconstruction, we introduce a new bundle adjustment cost function that includes, in a single term, the usual 3D point/ID observation consistency constraint as well as the geometric constraints provided by the 3D model. Results on large-scale synthetic and real sequences show that our method successfully improves SLAM reconstructions. Besides, experiments prove that the resulting reconstruction is accurate enough to be directly used for global relocalization applications. Pierre Lothe, Steve Bourgeois, Fabien Dekeyser, Eric Royer, Michel Dhome |
CVPR | 2 |
| 2005 | A Practical Guide to Marker Based and Hybrid Visual Registration for AR Industrial Applications
Steve Bourgeois, Hanna Martinsson, Quoc Cuong Pham, Sylvie Naudet |
CAIP | 1 |