Gilles Simon

dblp:11/1177 · DBLP profile ↗
← Back
31ranked-venue papers
11as first author
6since 2021 · last 2026
0000-0002-5909-6524ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 11 first-author · 2 since 2021Artificial intelligence and machine learning · 15 · 6 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 10 · 4 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
3D vision · 76% Robot navigation and mapping · 24%
Computer graphics and multimedia
9 papers
Virtual and augmented reality · 60% Geometric modeling and processing · 18% Computer animation and physical simulation · 10%

Topics — the 21 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
camera pose estimation
1.232023
Perspective-1-Ellipsoid: Formulation, Analysis and Solutions of the Camera Pose Estimation Problem from One Ellipse-Ellipsoid Correspondence · Int. J. Comput. Vis. 2023
Object-Based Visual Camera Pose Estimation From Ellipsoidal Model and 3D-Aware Ellipse Prediction · Int. J. Comput. Vis. 2022
Handling Uncertain Sensor Data in Vision-Based Camera Tracking · ISMAR 2004
Computer vision › 3D vision › visual localization
camera relocalization
1.022022
OA-SLAM: Leveraging Objects for Camera Relocalization in Visual SLAM · ISMAR 2022
Camera Relocalization with Ellipsoidal Abstraction of Objects · ISMAR 2019
Computer vision › 3D vision
3d shape modeling
0.612022
Object-Based Visual Camera Pose Estimation From Ellipsoidal Model and 3D-Aware Ellipse Prediction · Int. J. Comput. Vis. 2022
Robotics › Robot navigation and mapping
SLAM
0.612022
OA-SLAM: Leveraging Objects for Camera Relocalization in Visual SLAM · ISMAR 2022
Robotics › Robot navigation and mapping › SLAM
visual SLAM
0.612022
OA-SLAM: Leveraging Objects for Camera Relocalization in Visual SLAM · ISMAR 2022
Virtual and augmented reality
augmented reality
0.672019
Facade Proposals for Urban Augmented Reality · ISMAR 2017
Camera Relocalization with Ellipsoidal Abstraction of Objects · ISMAR 2019
Handling Uncertain Sensor Data in Vision-Based Camera Tracking · ISMAR 2004
Computer vision › 3D vision
pose estimation
0.422019
Camera Relocalization with Ellipsoidal Abstraction of Objects · ISMAR 2019
A Two-Stage Robust Statistical Method for Temporal Registration from Features of Various Type · ICCV 1998
Computer vision › 3D vision › camera calibration
vanishing point estimation
0.312018
A-Contrario Horizon-First Vanishing Point Detection Using Second-Order Grouping Laws · ECCV (10) 2018
Computer vision › 3D vision › 3d reconstruction
object reconstruction
0.212022
OA-SLAM: Leveraging Objects for Camera Relocalization in Visual SLAM · ISMAR 2022
Computer vision › 3D vision
3d reconstruction
0.212013
In-situ interactive modeling using a single-point laser rangefinder coupled with a new hybrid orientation tracker · ISMAR 2013
Geometric modeling and processing
3d reconstruction
0.122009
Immersive image-based modeling of polyhedral scenes · ISMAR 2009
Reconstructing While Registering: A Novel Approach for Markerless Augmented Reality · ISMAR 2002
Virtual and augmented reality › tracking
camera tracking
0.112011
Tracking-by-synthesis using point features and pyramidal blurring · ISMAR 2011
Computational photography and imaging
image-based modeling
0.112009
Immersive image-based modeling of polyhedral scenes · ISMAR 2009
Computer vision › 3D vision
camera calibration
0.112005
Calibration Errors in Augmented Reality: A Practical Study · ISMAR 2005
Human-robot interaction
natural interaction
0.012013
In-situ interactive modeling using a single-point laser rangefinder coupled with a new hybrid orientation tracker · ISMAR 2013
Robotics › Robot navigation and mapping
sensor fusion
0.012004
Handling Uncertain Sensor Data in Vision-Based Camera Tracking · ISMAR 2004
Virtual and augmented reality › tracking
hybrid tracking
0.012004
Handling Uncertain Sensor Data in Vision-Based Camera Tracking · ISMAR 2004
Visual content generation and editing
image generation
0.012011
Tracking-by-synthesis using point features and pyramidal blurring · ISMAR 2011
Geometric modeling and processing › registration
camera registration
0.012000
Registration with a Moving Zoom Lens Camera for Augmented Reality Applications · ECCV (2) 2000
Computational photography and imaging
camera calibration
0.012000
Registration with a Moving Zoom Lens Camera for Augmented Reality Applications · ECCV (2) 2000
Virtual and augmented reality › tracking
augmented reality tracking
0.011998
A Two-Stage Robust Statistical Method for Temporal Registration from Features of Various Type · ICCV 1998

Methods — techniques the papers use, named apart from their topics

reprojection error estimation · 0.8object detection · 0.8ellipse-ellipsoid correspondence · 0.8YOLO · 0.8object landmark · 0.63d ellipsoid · 0.6laser range finder · 0.3hybrid orientation tracking · 0.3IMU · 0.3object proposals · 0.3CNN-based descriptors · 0.3mipmapping · 0.1harris · 0.1depth response calibration · 0.1FAST · 0.1camera model evaluation · 0.1experimental protocol · 0.1
YearPublicationVenuePosition
2026 Gaussian Splatting Map Registration with Orthographic Bird's-Eye-View Renderings
abstract
Gaussian Splatting (GS) is a promising scene representation for visual localization and SLAM. Recent works have explored loop closure detection via Gaussian registration, improving map consistency and accuracy. However, achieving reliable registration given two GS representations from different acquisitions remains challenging. In this paper, we propose a complete pipeline to perform the matching and registration given two GS maps. The proposed method is grounded in generating orthographic bird’s-eye views (BEVs) of optimized Gaussian models. The proposed approach leverages photometric and geometric information extracted directly from the GS to provide a trade-off of accuracy and invariance to different viewing changes (e.g., as types of GS maps, seasons, or illumination). Unlike existing 3D registration methods, which become inefficient as the number of Gaussians grows, our approach leverages 2D orthographic renders thus considerably reducing the registration complexity. Experiments on two public datasets demonstrate that our method achieves higher accuracy than several existing baselines, while also maintaining better registration results when dealing with GS maps learned by different techniques (e.g., 3DGS to LightGaussian), or GS maps presenting viewing changes such as varying illumination conditions. Source code is available at: https://gitlab.inria.fr/tangram/bev-splatreg
Hugo Leblond, Gilles Simon, Renato Martins, Cédric Demonceaux, Marie-Odile Berger
WACV2
2023 Perspective-1-Ellipsoid: Formulation, Analysis and Solutions of the Camera Pose Estimation Problem from One Ellipse-Ellipsoid Correspondence
Vincent Gaudillière, Gilles Simon, Marie-Odile Berger
Int. J. Comput. Vis.2
2022 Level Set-Based Camera Pose Estimation From Multiple 2D/3D Ellipse-Ellipsoid Correspondences
abstract
In this paper, we propose an object-based camera pose estimation from a single RGB image and a pre-built map of objects, represented with ellipsoidal models. We show that contrary to point correspondences, the definition of a cost function characterizing the projection of a 3D object onto a 2D object detection is not straightforward. We develop an ellipse-ellipse cost based on level sets sampling, demonstrate its nice properties for handling partially visible objects and compare its performance with other common metrics. Finally, we show that the use of a predictive uncertainty on the detected ellipses allows a fair weighting of the contribution of the correspondences which improves the computed pose. The code is released at gitlab.inria.fr/tangram/level-set-based-camera-pose-estimation.
Matthieu Zins, Gilles Simon, Marie-Odile Berger
IROS2
2022 OA-SLAM: Leveraging Objects for Camera Relocalization in Visual SLAM
abstract
In this work, we explore the use of objects in Simultaneous Localization and Mapping in unseen worlds and propose an object-aided system (OA-SLAM). More precisely, we show that, compared to low-level points, the major benefit of objects lies in their higher-level semantic and discriminating power. Points, on the contrary, have a better spatial localization accuracy than the generic coarse models used to represent objects (cuboid or ellipsoid). We show that combining points and objects is of great interest to address the problem of camera pose recovery. Our main contributions are: (1) we improve the relocalization ability of a SLAM system using high-level object landmarks; (2) we build an automatic system, capable of identifying, tracking and reconstructing objects with 3D ellipsoids; (3) we show that object-based localization can be used to reinitialize or resume camera tracking. Our fully automatic system allows on-the-fly object mapping and enhanced pose tracking recovery, which we think, can significantly benefit to the AR community. Our experiments show that the camera can be relocalized from viewpoints where classical methods fail. We demonstrate that this localization allows a SLAM system to continue working despite a tracking loss, which can happen frequently with an uninitiated user. Our code and test data are released at gitlab.inria.fr/tangram/oa-slam.
Matthieu Zins, Gilles Simon, Marie-Odile Berger
ISMAR2
2022 Object-Based Visual Camera Pose Estimation From Ellipsoidal Model and 3D-Aware Ellipse Prediction
Matthieu Zins, Gilles Simon, Marie-Odile Berger
Int. J. Comput. Vis.2
2021 Model-image registration of a building's facade based on dense semantic segmentation
Antoine Fond, Marie-Odile Berger, Gilles Simon
Comput. Vis. Image Underst.3
2020 3D-Aware Ellipse Prediction for Object-Based Camera Pose Estimation
abstract
In this paper, we propose a method for coarse camera pose computation which is robust to viewing conditions and does not require a detailed model of the scene. This method meets the growing need of easy deployment of robotics or augmented reality applications in any environments, especially those for which no accurate 3D model nor huge amount of ground truth data are available. It exploits the ability of deep learning techniques to reliably detect objects regardless of viewing conditions. Previous works have also shown that abstracting the geometry of a scene of objects by an ellipsoid cloud allows to compute the camera pose accurately enough for various application needs. Though promising, these approaches use the ellipses fitted to the detection bounding boxes as an approximation of the imaged objects. In this paper, we go one step further and propose a learning-based method which detects improved elliptic approximations of objects which are coherent with the 3D ellipsoid in terms of perspective projection. Experiments prove that the accuracy of the computed pose significantly increases thanks to our method and is more robust to the variability of the boundaries of the detection boxes. This is achieved with very little effort in terms of training data acquisition - a few hundred calibrated images of which only three need manual object annotation.
Matthieu Zins, Gilles Simon, Marie-Odile Berger
3DV2
2020 Generic Document Image Dewarping by Probabilistic Discretization of Vanishing Points
abstract
Document images dewarping is still a challenge especially when documents are captured with one camera in an uncontrolled environment. In this paper we propose a generic approach based on vanishing points (VP) to reconstruct the 3D shape of document pages. Unlike previous methods we do not need to segment the text included in the documents. Therefore, our approach is less sensitive to pre-processing and segmentation errors. The computation of the VPs is robust and relies on the a-contrario framework, which has only one parameter whose setting is based on probabilistic reasoning instead of experimental tuning. Thus, our method can be applied to any kind of document including text and non-text blocks and extended to other kind of images. Experimental results show that the proposed method is robust to a variety of distortions.
Gilles Simon, Salvatore Tabbone
ICPR1
2019 Camera Pose Estimation with Semantic 3D Model
abstract
In computer vision, estimating camera pose from correspondences between 3D geometric entities and their projections into the image is a widely investigated problem. Although most state-of-the-art methods exploit simple primitives such as points or lines, and thus require dense scene models, the emergence of very effective CNN-based object detectors in the recent years have paved the way to the use of much lighter 3D models composed solely of a few semantically relevant features. In that context, we propose a novel model-based camera pose estimation method in which the scene is modeled by a set of virtual ellipsoids. We show that 6-DoF camera pose can be determined by optimizing only the three orientation parameters, and that at least two correspondences between 3D ellipsoids and their 2D projections are necessary in practice. We validate the approach on both simulated and real environments.
Vincent Gaudillière, Gilles Simon, Marie-Odile Berger
IROS2
2019 Camera Relocalization with Ellipsoidal Abstraction of Objects
abstract
We are interested in AR applications which take place in man-made GPS-denied environments, as industrial or indoor scenes. In such environments, relocalization may fail due to repeated patterns and large changes in appearance which occur even for small changes in viewpoint. We investigate in this paper a new method for relocalization which operates at the level of objects and takes advantage of the impressive progress realized in object detection. Recent works have opened the way towards object oriented reconstruction from elliptic approximation of objects detected in images. We go one step further and propose a new method for pose computation based on ellipse/ellipsoid correspondences. We consider in this paper the practical common case where an initial guess of the rotation matrix of the pose is known, for instance with an inertial sensor or from the estimation of orthogonal vanishing points. Our contributions are twofold: we prove that a closed form estimate of the translation can be computed from one ellipse-ellipsoid correspondence. The accuracy of the method is assessed on the LINEMOD database using only one correspondence. Second, we prove the effectiveness of the method on real scenes from a set of object detections generated by YOLO. A robust framework that is able to choose the best set of hypotheses is proposed and is based on an appropriate estimation of the reprojection error of ellipsoids. Globally, considering pose at the level of object allows us to avoid common failures due to repeated structures. In addition, due to the small combinatory induced by object correspondences, our method is well suited to fast rough localization even in large environments.
Vincent Gaudillière, Gilles Simon, Marie-Odile Berger
ISMAR2
2018 A-Contrario Horizon-First Vanishing Point Detection Using Second-Order Grouping Laws
Gilles Simon, Antoine Fond, Marie-Odile Berger
ECCV (10)1
2018 Region-Based Epipolar and Planar Geometry Estimation in Low─Textured Environments
abstract
Given two views of the same scene, usual correspondence geometry estimation techniques exploit the well-established effectiveness of keypoint descriptors. However, such features have a hard time in poorly textured man-made environments, possibly containing repetitive patterns and/or specularities, such as industrial places. In that paper, we propose a novel method for two-view epipolar and planar geometry estimation that first aims at detecting and matching physical vertical planes frequently present in these environments, before estimating corresponding homographies. Inferred local correspondences are finally used to improve fundamental matrix estimation. The gain in precision is demonstrated on industrial and urban environments.
Vincent Gaudillière, Gilles Simon, Marie-Odile Berger
ICIP2
2017 Facade Proposals for Urban Augmented Reality
abstract
We introduce a novel object proposals method specific to building facades. We define new image cues that measure typical facade characteristics such as semantic, symmetry and repetitions. They are combined to generate a few facade candidates in urban environments fast. We show that our method outperforms state-of-the-art object proposals techniques for this task on the 1000 images of the Zurich Building Database. We demonstrate the interest of this procedure for augmented reality through facade recognition and camera pose initialization. In a very time-efficient pipeline we classify the candidates and match them to a facade references database using CNN-based descriptors. We prove that this approach is more robust to severe changes of viewpoint and occlusions than standard object recognition methods.
Antoine Fond, Marie-Odile Berger, Gilles Simon
ISMAR3
2013 In-situ interactive modeling using a single-point laser rangefinder coupled with a new hybrid orientation tracker
abstract
We present a method for in situ modeling of polygonal scenes, using a laser rangefinder, an IMU and a camera. The main contributions of this work are a well-founded calibration procedure, a new hybrid, driftless orientation tracking method and an easy-to-use interface based on natural interactions.
Christel Leonet, Gilles Simon, Marie-Odile Berger
ISMAR2
2011 Tracking-by-synthesis using point features and pyramidal blurring
abstract
Tracking-by-synthesis is a promising method for markerless vision-based camera tracking, particularly suitable for Augmented Reality applications. In particular, it is drift-free, viewpoint invariant and easy-to-combine with physical sensors such as GPS and inertial sensors. While edge features have been used succesfully within the tracking-by-synthesis framework, point features have, to our knowledge, still never been used. We believe that this is due to the fact that real-time corner detectors are generally weakly repeatable between a camera image and a rendered texture. In this paper, we compare the repeatability of commonly used FAST, Harris and SURF interest point detectors across view synthesis. We show that adding depth blur to the rendered texture can drastically improve the repeatability of FAST and Harris corner detectors (up to 100% in our experiments), which can be very helpful, e.g., to make tracking-by-synthesis running on mobile phones. We propose a method for simulating depth blur on the rendered images using a pre-calibrated depth response curve. In order to fulfil the performance requirements, a pyramidal approach is used based on the well-known MIP mapping technique. We also propose an original method for calibrating the depth response curve, which is suitable for any kind of focus lenses and comes for free in terms of programming effort, once the tracking-by-synthesis algorithm has been implemented.
Gilles Simon
ISMAR1
2011 Interactive building and augmentation of piecewise planar environments using the intersection lines
Gilles Simon, Marie-Odile Berger
Vis. Comput.1
2010 Transitive Closure Based Visual Words for Point Matching in Video Sequence
abstract
We present Transitive Closure based visual word formation technique for obtaining robust object representations from smoothly varying multiple views. Each one of our visual words is represented by a set of feature vectors which is obtained by performing transitive closure operation on SIFT features. We also present range-reducing tree structure to speed up the transitive closure operation. The robustness of our visual word representation is demonstrated for Structure from Motion (SfM) and location identification in video images.
K. K. Srikrishna Bhat, Marie-Odile Berger, Gilles Simon, Frédéric Sur
ICPR3
2009 Immersive image-based modeling of polyhedral scenes
abstract
In this paper, we describe a purely image-based system that allows a user to interactively capture the 3D geometry of a polyhedral scene with the aid of its physical presence. A video camera is used as both an interaction and tracking device. The 3D user interface is intuitive to a non-expert and the mouseless control procedure makes the system particularly suitable for mobile devices such as PDAs and mobile phones. The efficiency and accuracy of the method are demonstrated on a polyhedral scene made of two house-like boxes.
Gilles Simon
ISMAR1
2008 Detection of the intersection lines in multiplanar environments: Application to real-time estimation of the camera-scene geometry
abstract
This paper describes an integrated system for building a multiplanar model of the scene as the camera is localized on the fly. The core of this system is a robust and accurate procedure for detecting the intersection line between two planes. User cues are used to assist the system in the mapping tasks. Synthetic results and a long video demonstrate the relevance of the method.
Gilles Simon, Marie-Odile Berger
ICPR1
2007 Use of inertial sensors to support video tracking
abstract
Abstract One of the biggest obstacles to building effective augmented reality (AR) systems is the lack of accurate sensors that report the location of the user in an environment during arbitrary long periods of movements. In this paper, we present an effective hybrid approach that integrates inertial and vision‐based technologies. This work is motivated by the need to explicitly take into account the relatively poor accuracy of inertial sensors and thus to define an efficient strategy for the collaborative process between the vision‐based system and the sensor. The contributions of this papers are threefold: (i) our collaborative strategy fully integrates the sensitivity error of the sensor: the sensitivity is practically studied and is propagated into the collaborative process, especially in the matching stage (ii) we propose an original online synchronization process between the vision‐based system and the sensor. This process allows us to use the sensor only when needed. (iii) an effective AR system using this hybrid tracking is demonstrated through an e‐commerce application in unprepared environments. Copyright © 2007 John Wiley & Sons, Ltd.
Michael Aron, Gilles Simon, Marie-Odile Berger
Comput. Animat. Virtual Worlds2
2006 Automatic online walls detection for immediate use in AR tasks
abstract
This paper proposes a method to automatically detect and reconstruct planar surfaces for immediate use in AR tasks. Traditional methods for plane detection are typically based on the comparison of transfer errors of a homography, which make them sensitive to the choice of a discrimination threshold. We propose a very different approach: the image is divided into a grid and rectangles that belong to the same planar surface are clustered around the local maxima of a Hough transform. As a result, we simultaneously get clusters of coplanar rectangles and the image of their intersection line with a reference plane, which easily leads to their 3D position and orientation. Results are shown on both synthetic and real data.
Gilles Simon
ISMAR1
2005 Calibration Errors in Augmented Reality: A Practical Study
abstract
This work confronts some theoretical camera models to reality and evaluates the suitability of these models for effective augmented reality (AR). It analyses what level of accuracy can be expected in real situations using a particular camera model and how robust the results are against realistic calibration errors. An experimental protocol is used that consists of taking images of a particular scene from different quality cameras mounted on a 4DOF micro-controlled device. The scene is made of a calibration target and three markers placed at different distances of the target. This protocol enables us to consider assessment criteria specific to AR as alignment error and visual impression, in addition to the classical camera positioning error.
Javier-Flavio Vigueras, Gilles Simon, Marie-Odile Berger
ISMAR2
2004 Handling Uncertain Sensor Data in Vision-Based Camera Tracking
abstract
A hybrid approach for real-time markerless tracking is presented. Robust and accurate tracking is obtained from the coupling of camera and inertial sensor data. Unlike previous approaches, we use sensor information only when the image-based system fails to track the camera. In addition, sensor errors are measured and taken into account at each step of our algorithm. Finally, we address the camera/sensor synchronization problem and propose a method to resynchronize these two devices online. We demonstrate our method in two example sequences that illustrate the behavior and benefits of the new tracking method.
Michael Aron, Gilles Simon, Marie-Odile Berger
ISMAR2
2002 Real time registration of known or recovered multi-planar structures: application to AR
abstract
Colloque avec actes et comité de lecture. internationale.
Gilles Simon, Marie-Odile Berger
BMVC1
2002 Reconstructing While Registering: A Novel Approach for Markerless Augmented Reality
abstract
This paper addresses the registration problem for unprepared multi-planar scenes. An interactive process is proposed to obtain accurate results using only the texture information of planes. In particular, classical preparation steps (camera calibration, scene acquisition) are greatly simplified, since they are included in the on-line registration process. Results are shown on indoor and outdoor scenes. Videos are available at url http://www.loria.fr//spl tilde/gsimon/Ismar.
Gilles Simon, Marie-Odile Berger
ISMAR1
2000 Registration with a Moving Zoom Lens Camera for Augmented Reality Applications
Gilles Simon, Marie-Odile Berger
ECCV (2)1
1999 Mixing synthetic and video images of an outdoor urban environment
Marie-Odile Berger, Brigitte Wrobel-Dautcourt, Sylvain Petitjean, Gilles Simon
Mach. Vis. Appl.4
1998 Robust Image Composition Algorithms for Augmented Reality
Marie-Odile Berger, Gilles Simon
ACCV (2)2
1998 A Two-Stage Robust Statistical Method for Temporal Registration from Features of Various Type
abstract
A model registration system capable of tracking an object, the model of which is known, in an image sequence is presented. It integrates tracking, pose determination and updating of the visible features. The heart of our system is the pose computation method, which handles various features (points, lines and free-form curves) in a very robust way and is able to give a correct estimate of the pose even when tracking errors occur. The reliability of the system is shown on an augmented reality project.
Gilles Simon, Marie-Odile Berger
ICCV1
1996 Mixing synthesis and video images of outdoor environments: application to the bridges of Paris
abstract
Augmented reality is the technique by which real images can be enhanced by addition of computer-generated information. Augmented reality shows great promises in fields where a simulation in situ would be either impossible, not realistic enough or too expensive. We present in this paper an augmented reality loop that uses vision tools (perspective inversion, tracking) and show how it was used to fully enrich a sequence of the bridge of Paris with a model illuminated synthetically.
Marie-Odile Berger, Gilles Simon, Sylvain Petitjean, Brigitte Wrobel-Dautcourt
ICPR2
1996 Compositing Computer and Video Image Sequences: Robust Algorithms for the Reconstruction of the Camera Parameters
abstract
Abstract Augmented reality shows great promises in fields where a simulation in situ would be impossible or too expensive. When mixing synthetic and real objects in the same animated sequence, we must be sure that the geometrical coherence as well as the photometrical coherence is ensured. One major challenge is to compute the camera viewpoint with sufficient accuracy to ensure a satisfactory composition. We especially address this point in this paper using computer vision techniques and robust statistical methods. We prove that such techniques make it possible to compute almost automatically the viewpoint for long video sequences even for bad quality images in outdoor environments. Significant results on the lighting simulation of the bridges of Paris are shown.
Marie-Odile Berger, Christine Chevrier, Gilles Simon
Comput. Graph. Forum3