Guillaume Caron

dblp:55/7736 · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0003-3703-2973ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 8 first-author · 7 since 2021Systems, architecture and hardware · 12 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 FuseDPT: Multi-scale and Multi-projection Model for Learning Depth in 360$^\circ $C
Matheus Paula, Nevrez Imamoglu, Guillaume Caron, Antoine N. André
ICPR (5)3
2026 Temporal Illumination Variation Compensation Using Perpendicular Whisk-Broom Hyperspectral Scans
abstract
This study introduces a novel method to mitigate temporal illumination variations in whisk-broom hyperspectral imaging under varying environmental illumination conditions. Whisk-broom hyperspectral imaging captures high resolution spectra pixel-by-pixel sequentially, a process susceptible to sunlight fluctuations over time, particularly when imaging cultural artifacts outdoors. Despite sunlight's broad spectrum, its variability over time can compromise the quality of hyperspectral images, affecting data analysis. Prior strategies suggested a supplementary single compensating vertical scan alongside the standard row-wise raster scan. However, this strategy fails when the additional single vertical scan is performed near or on a black frame. Building on this, our study proposes using two scans, one performed orthogonally to the traditional row-wise scan, to counteract illumination variations. Furthermore, we formulate the compensation problem in logarithmic space, exploiting the low-dimensional structure of the reflectance and illumination spectra. Using total variation penalization in the cost function enhances smoothness. Our method achieves robust compensation for changes in environmental illumination. We also demonstrate that we can use multiple columns from the column-wise scan without significantly decreasing the compensation quality while reducing acquisition time. We illustrate the application of our methods to hyperspectral images of stained-glass windows of the historic Cathedrale Notre-Dame d'Amiens in France.
Suzan Joseph Kessy, Takuya Funatomi, Takahiro Kushida 0001, Kazuya Kitano, Yuki Fujimura, Guillaume Caron, El Mustapha Mouaddib, Yasuhiro Mukaigawa
Comput. Vis. Media6
2025 Enhancing multimodal-input object goal navigation by leveraging large language models for inferring room-object relationship knowledge
Leyuan Sun, Asako Kanezaki, Guillaume Caron, Yusuke Yoshiyasu
Adv. Eng. Informatics3
2025 UniphorM: A New Uniform Spherical Image Representation for Robotic Vision
abstract
In this article, we present a new spherical image representation, called uniform spherical mapping of omnidirectional images (UniphorM), and show its strong potential in robotic vision. UniphorM provides an accurate and distortion-free representation of a 360-degree image, by relying on multiple subdivisions of an icosahedron and its associated Voronoi diagrams. The geometric mapping procedure is described in detail, and the tradeoff between pixel accuracy and computational complexity is investigated. To demonstrate the benefits of UniphorM in real-world problems, we applied it to direct visual attitude estimation and visual place recognition (VPR), by considering dual-fisheye images captured by a camera mounted on multiple robotic platforms. In the experiments, we measured the impact of the number of subdivision levels of the icosahedron on the attitude estimation error, time efficiency, and size of convergence domain of an existing visual gyroscope, using UniphorM and three competing mapping algorithms. A similar evaluation procedure was carried out for VPR. Finally, two new omnidirectional image datasets, one recorded with a hexacopter, calledSVMIS+, the other based on theMapillaryplatform, have been created and released for the entire research community.
Antoine N. André, Fabio Morbidi, Guillaume Caron
IEEE Trans. Robotics3
2024 On camera model conversions
abstract
On the one hand, cameras of conventional field-of-view usually considered in computer vision and robotics are very often modeled as a pinhole plus possibly a distortion model. On the other hand, there is a large variety of models for panoramic cameras. Many camera models have been proposed for fisheye cameras, catadioptric cameras, and super fisheye cameras. But in both cases, few models offer the possibility of converting them into another model.This paper contributes to filling this gap in, to allow an algorithm designed with a projection model to accept data of a camera calibrated with another model. So, a pre-existing data set can be used without having to recalibrate the camera. We provide the methodology and mathematical developments for three conversions considering three different types of cameras that are evaluated with respect to calibration and within a visual Simultaneous Localization And Mapping benchmark. The source code of the camera model conversions studied in this paper is shared within the libPeR library for Perception in Robotics: https://github.com/PerceptionRobotique/libPeRbase.
Eva Goichon, Guillaume Caron, Pascal Vasseur, Fumio Kanehiro
ICRA2
2024 Direct 3D model-based object tracking with event camera by motion interpolation
abstract
Event cameras are recent sensors that measure intensity changes in each pixel asynchronously. It is being used due to lower latency and higher temporal resolution compared to traditional frame-based camera. We propose a method of 3D model-based object tracking directly from events captured by event camera. To enable reliable and accurate tracking of objects, we use a new event representation and predict brightness increment images with motion interpolation. Results of object tracking show the new methods significantly improves tracking duration and robustness, both for perspective and fisheye cameras. Our implementation succeeds in tracking objects when the camera speed is reaching 2 m/s.
Yufan Kang, Guillaume Caron, Ryoichi Ishikawa, Adrien Escande, Kevin Chappellet, Ryusuke Sagawa, Takeshi Oishi
ICRA2
2024 A mathematical characterization of the convergence domain for Direct Visual Servoing
abstract
Direct Visual Servoing (DVS) is a technique that controls the robot motion by using the pixel intensities captured by a camera. DVS demonstrates high accuracy at convergence, prompting the development of various methods aimed at expanding its convergence domain.In this paper, we propose a mathematical characterization of the DVS convergence domain with closed-form expressions for the controlled degrees of freedom. From these expressions, we concluded that the extent of the convergence domain is related to the presence of isotropic or defocus blur, a phenomenon that had only been observed previously as a trend in empirical experiments.
Meriem Belinda Naamani, Guillaume Caron, Mitsuharu Morisawa, El Mustapha Mouaddib
IROS2
2024 A Survey on Adaptive Cameras
Julien Ducrocq, Guillaume Caron
Int. J. Comput. Vis.2
2024 Humanoid Loco-Manipulations Using Combined Fast Dense 3D Tracking and SLAM With Wide-Angle Depth-Images
abstract
To efficiently achieve complex humanoid loco-manipulation tasks in industrial contexts, we propose a combined vision-based tracker-localization interplay integrated as part of a task-space whole-body optimization control. To achieve good perception complementarity between manipulation and localization, a new fast dense 3D model-based tracking using wide-angle depth image is developed and used in conjunction with a simultaneous localization and mapping software. Our approach allows humanoid robots, targeted for industrial manufacturing, to manipulate and assemble large-scale objects while walking. It is assessed with experiments consisting in rolling and assembling in an unwinder a heavy and wide bobbin using bimanual grasping and bipedal locomotion at a time. This experimental use-case is found in some large-scale manufacturing where bobbins are enrolled with various materials (cables, papers, rubbers, etc.). The same experiments are made using two different humanoid robots of the same family.Note to Practitioners—This paper aims at deploying humanoid robots in large-scale manufacturing industries. We consider non-added value tasks related to transporting large tools or objects such as large bobbins by means of locomanipulation skills, similarly to human workers. We developed a task-space control framework that has been successfully applied in the aircraft industry. In the frame of a current collaboration with other major industrial sectors, we enhanced our control framework to interplay between SLAM and visual tracking to realize robust loco-manipulation tasks. Our approach can be applied and ported to any humanoid robot or bi-manual wheeled mobile robots with minor programming effort as the software is made open. Preliminary experiments with two different humanoids and use-cases suggest that our approach is feasible. In future research, we will address the problem of performance to reach at least human-speed in the execution of locomanipulation tasks in large-scale industry and automation contexts.
Kevin Chappellet, Masaki Murooka, Guillaume Caron, Fumio Kanehiro, Abderrahmane Kheddar
IEEE Trans Autom. Sci. Eng.3
2022 Direct Alignment Of Narrow Field-Of-View Hyperspectral Data And Full-View Rgb Image
abstract
The novelty of this paper is the alignment method of narrow field-of-view hyperspectral images to full-view RGB images. The interest is to locate hyperspectral measurements in an environment described by an equirectangular image. But the very different modalities (3 vs. hundreds of channels) and fields-of-view are challenges for accurate alignment. We solve these problems within a dense direct alignment framework that optimizes the warping parameters together with those of a global illumination difference model. Our alignment code is shared with an example dataset available at github.com/jrl-umi3218/hsrgbalign.
Guillaume Caron, Suzan Joseph Kessy, Yasuhiro Mukaigawa, Takuya Funatomi
ICIP1
2022 Direct visual servoing in the non-linear scale space of camera pose
abstract
This paper proposes to consider direct visual servoing (DVS) for object manipulation by a robot arm. The convergence domain limits of DVS are overcome by introducing the non-linear scale space related to camera pose. Its use in a new, yet little complex, direct cost can enlarge twice the convergence domain of one state-of-the-art DVS, as many experiments of symmetric object orientation control assess.
Guillaume Caron, Yusuke Yoshiyasu
ICPR1
2022 Eliminating Temporal Illumination Variations in Whisk-broom Hyperspectral Imaging
abstract
Abstract We propose a method for eliminating the temporal illumination variations in whisk-broom (point-scan) hyperspectral imaging. Whisk-broom scanning is useful for acquiring a spatial measurement using a pixel-based hyperspectral sensor. However, when it is applied to outdoor cultural heritages, temporal illumination variations become an issue due to the lengthy measurement time. As a result, the incoming illumination spectra vary across the measured image locations because different locations are measured at different times. To overcome this problem, in addition to the standard raster scan, we propose an additional perpendicular scan that traverses the raster scan. We show that this additional scan allows us to infer the illumination variations over the raster scan. Furthermore, the sparse structure in the illumination spectrum is exploited to robustly eliminate these variations. We quantitatively show that a hyperspectral image captured under sunlight is indeed affected by temporal illumination variations, that a Naïve mitigation method suffers from severe artifacts, and that the proposed method can robustly eliminate the illumination variations. Finally, we demonstrate the usefulness of the proposed method by capturing historic stained-glass windows of a French cathedral.
Takuya Funatomi, Takehiro Ogawa, Kenichiro Tanaka, Hiroyuki Kubo, Guillaume Caron, El Mustapha Mouaddib, Yasuyuki Matsushita, Yasuhiro Mukaigawa
Int. J. Comput. Vis.5
2020 Benchmarking Cameras for Open VSLAM Indoors
abstract
In this paper we benchmark different types of cameras and evaluate their performance in terms of reliable localization reliability and precision in Visual Simultaneous Localization and Mapping (vSLAM). Such benchmarking is merely found for visual odometry, but never for vSLAM. Existing studies usually compare several algorithms for a given camera. The evaluation methodology we propose is applied to the recent OpenVSLAM framework. The latter is versatile enough to natively deal with perspective, fisheye, 360 cameras in a monocular or stereoscopic setup, an in RGB or RGB-D modalities. Results in various sequences containing light variation and scenery modifications in the scene assess quantitatively the maximum localization rate for 360 vision. In the contrary, RGB-D vision shows the lowest localization rate, but highest precision when localization is possible. Stereo-fisheye trades-off with localization rates and precision between 360 vision and RGB-D vision. The dataset with ground truth will be made available in open access to allow evaluating other/future vSLAM algorithms with respect to these camera types.
Kevin Chappellet, Guillaume Caron, Fumio Kanehiro, Ken Sakurada, Abderrahmane Kheddar
ICPR2
2020 Photometric Path Planning for Vision-Based Navigation
abstract
We present a vision-based navigation system that uses a visual memory to navigate. Such memory corresponds to a topological map of key images created from moving a virtual camera over a model of the real scene. The advantage of our approach is that it provides a useful insight into the navigability of a visual path without relying on a traditional learning stage. During the navigation stage, the robot is controlled by sequentially comparing the images stored in the memory with the images acquired by the onboard camera.The evaluation is conducted on a robotic arm equipped with a camera and the model of the environment corresponds to a top view image of an urban scene.
Eder Alejandro Rodríguez Martínez, Guillaume Caron, Claude Pégard, David Lara Alabazares
ICRA2
2019 Adaptive Lucas-Kanade tracking
Yassine Ahmine, Guillaume Caron, El Mustapha Mouaddib, Fatima Chouireb
Image Vis. Comput.2
2019 Visual Servoing With Photometric Gaussian Mixtures as Dense Features
abstract
The direct use of the entire photometric image information as dense features for visual servoing brings several advantages. First, it does not require any feature detection, matching, or tracking process. Thanks to the redundancy of visual information, the precision at convergence is highly accurate. However, the corresponding highly nonlinear cost function reduces the convergence domain. In this paper, we propose visual servoing based on the analytical formulation of Gaussian mixtures to enlarge the convergence domain. Pixels are represented by two-dimensional Gaussian functions that denote a “power of attraction.” In addition to the control of the camera velocities during the servoing, we also optimize the Gaussian spreads allowing the camera to precisely converge to a desired pose even from a far initial one. Simulations show that our approach outperforms the state of the art and real experiments show the effectiveness, robustness, and accuracy of our approach.
Nathan Crombez, El Mustapha Mouaddib, Guillaume Caron, François Chaumette
IEEE Trans. Robotics3
2018 Amift: Affine-Mirror Invariant Feature Transform
abstract
In this paper, we propose a descriptor for image matching under multiple mirror reflections. Indeed, existing adaptations of SIFT for the mirror transformation are not successful when object and mirrors orientations are not constrained. Hence, we propose to combine MIFT and Affine-SIFT descriptors as the Affine Mirror Invariant Feature Transform (AMIFT). The experimental results and given evaluation show that our proposed descriptor outperforms MIFT and ASIFT on both synthetic and real images datasets.
Noureddine Mohtaram, Amina Radgui, Guillaume Caron, El Mustapha Mouaddib
ICIP3
2018 Spherical Visual Gyroscope for Autonomous Robots Using the Mixture of Photometric Potentials
abstract
In this paper, we present a new direct omnidirectional visual gyroscope for mobile robotic platforms. The gyroscope estimates the 3D orientation of a camera-robot by comparing the current spherical image with that acquired at a reference pose. By transforming pixel intensities into a Mixture of Photometric Potentials, we introduce a novel image-similarity measure which can be seamlessly integrated into a classical nonlinear least-squares optimization scheme, offering an extended convergence domain. Our method provides accurate and robust attitude estimates, and it is easy-to-use since it involves a single tuning parameter, the width of the photometric potentials (Gaussian functions, in this work) controlling the power of attraction of each pixel. The visual gyroscope has been successfully tested on spherical image sequences generated by a twin-fisheye camera mounted on the end-effector of a robot arm and on a fixed-wing UAV.
Guillaume Caron, Fabio Morbidi
ICRA1
2015 Photometric Gaussian mixtures based visual servoing
abstract
The advantages of using the entire photometric image information as visual feature are: it does not require any feature detections, matching or tracking process. To enlarge the convergence domain, we propose to accomplish visual servoing based on the analytical formulation of Gaussian mixtures to model the images. During the servoing, we consider the optimization of the Gaussian spreads allowing the camera to converge to a desired pose even from a far initial one. Simulation that overcomes the state-of-the-art and real experiments highlight the success of our approach.
Nathan Crombez, Guillaume Caron, El Mustapha Mouaddib
IROS2
2015 Good feature for framing: Saliency-based Gaussian Mixture
abstract
In this paper, we present a new automatic camera control to achieve a relevant information framing. This camera control will be performed using a visual servoing framework in order to reach a salient area in a plane and in space. The relevant framing will be modelled by maximizing the saliency-based Gaussian Mixture Model (GMM) feature in the image. Furthermore, in order to achieve a realistic automatic camera control, we add an obstacles avoidance constraint and we ensure a relevant orientation during the motion. We validate our contribution in different synthetic 2D and 3D environments. Finally, we test our approach on a dense 3D points cloud model and in a real environment with a robot.
Zaynab Habibi, El Mustapha Mouaddib, Guillaume Caron
IROS3
2014 Toward the Adaptive and Context-Aware Serious Game Design
abstract
The context of this research focus on the design approach of Serious Games dedicated to Cultural Heritage. This approach allowed creating and testing several Serious Games. Feedback obtained from the tests of SG, allowed us to refine the design of the games and to design a meta-scenario which can be adapted (depending on parameters), to the needs of the teacher and according to his pedagogical aims.
Dominique Groux, Guillaume Caron
ICALT2
2014 Direct model based visual tracking and pose estimation using mutual information
Guillaume Caron, Amaury Dame, Éric Marchand
Image Vis. Comput.1
2011 Multiple camera types simultaneous stereo calibration
abstract
Calibration is a classical issue in computer vision needed to retrieve 3D information from image measurements. This work presents a calibration approach for hybrid stereo rig involving multiple central camera types (perspective, flsheye, catadioptric). The paper extends the method of monocular perspective camera calibration using virtual visual servoing. The simultaneous intrinsic and extrinsic calibration of central cameras rig, using different models for each camera, is developed. The presented approach is suitable for the calibration of rigs composed by N cameras modelled by N different models. Calibration results, compared with state of the art approaches, and a 3D plane estimation application, allowed by the calibration, show the effectiveness of the approach. A cross-platform software implementing this method is available.
Guillaume Caron, Damien Eynard
ICRA1
2011 Tracking planes in omnidirectional stereovision
abstract
Omnidirectional cameras allow direct tracking and motion estimation of planar regions in images during a long period of time. However, using only one camera leads to plane and trajectory reconstruction up to a scale factor. We propose to develop dense plane tracking based on omnidirectional stereovision to answer this issue. The presented method estimates simultaneously the parameters of several 3D planes along with the camera motion in a spherical model formulation. Results show the efficiency of the approach.
Guillaume Caron, Éric Marchand, El Mustapha Mouaddib
ICRA1
2010 Omnidirectional photometric visual servoing
abstract
Visual servoing has been based on geometric features for a long time. Recent works have highlighted the interest of taking into account the photometric information of the entire image. This approach was tackled with images of perspective cameras. We propose, in this paper, to adapt this technique to central cameras. This generalization allows to apply this kind of method to wide field of view cameras. We also propose to adapt gradient computation to take into account distorsions of such cameras. Several experiments have been successfully done with a fisheye camera.
Guillaume Caron, Éric Marchand, El Mustapha Mouaddib
IROS1
2009 Vertical line matching for omnidirectional stereovision images
abstract
We are investigating the mobile robot indoor localization and environment mapping using an omnidirectional stereovision sensor. It uses four parabolic mirrors and an orthographic camera, giving four images of the same scene. At least, only two mirrors are needed. Using four mirrors gives redundancy. We propose to exploit the images of vertical lines. This paper presents a new method in order to match these lines in the four images. Contrary to existing approaches, we took into account the four sub-images existence in the design of this method, in order to exploit redundancy. This brought an original algorithm combining matching and pose estimation of vertical lines from the 3D environment. Experimental results will be presented to validate this approach.
Guillaume Caron, El Mustapha Mouaddib
ICRA1
2009 3D model based pose estimation for omnidirectional stereovision
abstract
Robot vision has a lot to win as well with wide field of view induced by catadioptric cameras as with redundancy brought by stereovision. Merging these two characteristics in a single sensor is obtained by combining a single camera and multiple mirrors. This paper proposes a 3D model tracking algorithm that allows a robust tracking of 3D objects using stereo catadioptric images given by this sensor. The presented work relies on an adapted virtual visual servoing approach, a non-linear pose computation technique. The model take into account central projection and multiple mirrors. Results show robustness in illumination changes, mistracking and even higher robustness with four mirrors than with two.
Guillaume Caron, Éric Marchand, El Mustapha Mouaddib
IROS1