EDBT 2026 Demo / reviewers in the wild / expert
Cédric Demonceaux
dblp:58/116
· DBLP profile ↗
88ranked-venue papers
7as first author
28since 2021 · last 2026
0000-0001-6916-1273ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 65 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 2 first-author · 16 since 2021Systems, architecture and hardware · 27 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Domain Human Action Recognition from Multiview Motion and Textual Descriptions
Yannick Porto, Renato Martins, Thomas Chalumeau, Cédric Demonceaux |
ICPR (10) | 4 |
| 2026 | Gaussian Splatting Map Registration with Orthographic Bird's-Eye-View RenderingsabstractGaussian Splatting (GS) is a promising scene representation for visual localization and SLAM. Recent works have explored loop closure detection via Gaussian registration, improving map consistency and accuracy. However, achieving reliable registration given two GS representations from different acquisitions remains challenging. In this paper, we propose a complete pipeline to perform the matching and registration given two GS maps. The proposed method is grounded in generating orthographic bird’s-eye views (BEVs) of optimized Gaussian models. The proposed approach leverages photometric and geometric information extracted directly from the GS to provide a trade-off of accuracy and invariance to different viewing changes (e.g., as types of GS maps, seasons, or illumination). Unlike existing 3D registration methods, which become inefficient as the number of Gaussians grows, our approach leverages 2D orthographic renders thus considerably reducing the registration complexity. Experiments on two public datasets demonstrate that our method achieves higher accuracy than several existing baselines, while also maintaining better registration results when dealing with GS maps learned by different techniques (e.g., 3DGS to LightGaussian), or GS maps presenting viewing changes such as varying illumination conditions. Source code is available at: https://gitlab.inria.fr/tangram/bev-splatreg Hugo Leblond, Gilles Simon, Renato Martins, Cédric Demonceaux, Marie-Odile Berger |
WACV | 4 |
| 2025 | Event-Aware Distilled DETR for Object Detection in an Automotive ContextabstractAutonomous driving systems require robust object detection in complex environments. Event cameras outperform RGB cameras under challenging lighting conditions, but face limitations due to the scarcity of available datasets and lack of specialized training. To narrow the gap between RGB- and event-based detection accuracy and avoid the high complexity of real-time RGB-event fusion, in this paper, we propose a knowledge distillation framework. Our approach uses both modalities during training but relies solely on sparse event data at inference and transfers knowledge from a robust RGB-based teacher model. We build on the success of DETR (DEtection TRansformer) and we leverage an event-aware masked knowledge distillation mechanism, to boost event-based detection accuracy. Experiments on the DSEC-DET dataset demonstrate that our method not only excels in challenging driving scenarios where RGB images are unreliable, but also surpasses the state-of-the-art in event-based object detection. Djessy Rossi, Pascal Vasseur, Fabio Morbidi, Cédric Demonceaux, François Rameau |
IV | 4 |
| 2025 | Dense Scene Reconstruction from Light-Field Images Affected by Rolling ShutterabstractThis paper presents a dense depth estimation approach from light-field (LF) images that is able to compensate for strong rolling shutter (RS) effects. Our method estimates RS compensated views and dense RS compensated disparity maps. We present a two-stage method based on a 2D Gaussians Splatting that allows for a “render and compare” strategy with a point cloud formulation. In the first stage, a subset of sub-aperture images is used to estimate an RS agnostic 3D shape that is related to the scene target shape “up to a motion”. In the second stage, the deformation of the 3D shape is computed by estimating an admissible camera motion. We demonstrate the effectiveness and advantages of this approach through several experiments conducted for different scenes and types of motions. Due to lack of suitable datasets for evaluation, we also present a new carefully designed synthetic dataset of RS LF images. The source code, trained models and dataset will be made publicly available at: https://github.com/ICB-Vision-AI/DenseRSLF. Hermes McGriff, Renato Martins, Nicolas Andreff, Cédric Demonceaux |
WACV | 4 |
| 2025 | Tree-Based Personalized Clustered Federated Learning: A Driver Stress Monitoring Through Physiological Data Case StudyabstractRecent advancements in wearable biosensor technology have significantly enhanced the precision and ease of collecting physiological data. This has proven particularly useful in monitoring stress among drivers. However, the sensitive nature of this data raises significant privacy concerns. To tackle these challenges, federated learning (FL) has emerged as a novel solution to safeguard data privacy by decentralizing model training to individual devices, eliminating the need for data sharing. This approach is compliant with strict privacy regulations like the GDPR (EU) and HIPAA (US), significantly reducing data breach risks and data transfer costs. Despite its benefits, FL struggles with nonindependent and identically distributed (non-IID) data. This issue hampers the FL model performance and adaptability by complicating convergence with a generalized model. To address this limitation, we introduce in this article an innovative tree-based personalized clustered FL (TPCFL) approach. TPCFL effectively exploits similarities in drivers’ private data characteristics to assign each driver a personalized model relying on a tree-based clustering approach. Grounded in the realm of individuals clustering in FL to address non-IID data challenges, notably recognized as clustered FL (CFL), TPCFL is augmented with a novel tree-based clustering approach and a tailored cluster selection technique, enabling it to adeptly address core CFL challenges, such as hyperparameter optimization for cluster selection and integration of new unlabeled drivers. Experiments demonstrated the superior performance of the proposed clustering method and highlighted TPCFL’s efficiency in achieving an optimal balance between personalized and generalized learning, showcasing its effectiveness on the two public data sets. Houda Rafi, Yannick Benezeth, Fan Yang 0019, Philippe Reynaud, Emmanuel Arnoux, Cédric Demonceaux |
IEEE Internet Things J. | 6 |
| 2025 | High-Resolution Underwater Creature SegmentationabstractUnderwater creature segmentation (UCS) is critical for marine research and robotics but faces unique challenges: environmental distortions and biological traits that distinguish it from terrestrial segmentation. While deep learning advances exist, current UCS models are constrained to low-resolution inputs, losing critical details when processing high-resolution (HR) imagery and degrading segmentation precision. To bridge this gap, we introduce UCS4K, the first large-scale HR dataset for UCS, containing 4,096 images with pixel-wise annotations. UCS4K offers 4 times higher average resolution than existing datasets, covering diverse species, habitats, and environmental complexities essential for robust model training. Additionally, we propose a Resolution-Asymmetric Dual-branch Alignment and Refinement (RADAR) network to address the efficiency-receptiveness trade-off in HR-UCS. RADAR decouples context and detail processing: a CNN branch preserves HR spatial details, while a Transformer branch models global semantics on downsampled inputs to avoid quadratic complexity. Crucially, it resolves the inherent semantic misalignment issue between branches via the Global Semantic Alignment (GSA) module in the encoder and the Bidirectional Collaborative Refinement (BCR) module-embedded decoder that progressively integrates multi-scale encoding features to sharpen boundaries. This asymmetric design ensures efficient long-range context capture without sacrificing spatial precision. Extensive benchmarks demonstrate that RADAR sets new state-of-the-art performance on UCS4K and other existing datasets. Our contributions establish the first HR benchmark for UCS and deliver a scalable framework for high-precision segmentation. Dataset, code, and models are available at https://github.com/WHYfromNUT/RADAR. Huiyang Wu, Qiuping Jiang, Zongwei Wu, Runmin Cong, Cédric Demonceaux, Yi Yang 0001, Xiangyang Ji |
IEEE Trans. Image Process. | 5 |
| 2024 | SOAC: Spatio-Temporal Overlap-Aware Multi-Sensor Calibration using Neural Radiance FieldsabstractIn rapidly-evolving domains such as autonomous driving, the use of multiple sensors with different modalities is crucial to ensure high operational precision and stability. To correctly exploit the provided information by each sensor in a single common frame, it is essential for these sensors to be accurately calibrated. In this paper, we leverage the ability of Neural Radiance Fields (NeRF) to represent different sensors modalities in a common volumetric representation to achieve robust and accurate spatio-temporal sensor calibration. By designing a partitioning approach based on the visible part of the scene for each sensor, we formulate the calibration problem using only the overlapping areas. This strategy results in a more robust and accurate calibration that is less prone to failure. We demonstrate that our approach works on outdoor urban scenes by validating it on multiple established driving datasets. Results show that our method is able to get better accuracy and robustness compared to existing methods. Quentin Herau, Nathan Piasco, Moussâb Bennehar, Luis Roldão, Dzmitry Tsishkou, Cyrille Migniot, Pascal Vasseur, Cédric Demonceaux |
CVPR | 8 |
| 2024 | 3DGS-Calib: 3D Gaussian Splatting for Multimodal SpatioTemporal CalibrationabstractReliable multimodal sensor fusion algorithms require accurate spatiotemporal calibration. Recently, targetless calibration techniques based on implicit neural representations have proven to provide precise and robust results. Nevertheless, such methods are inherently slow to train given the high computational overhead caused by the large number of sampled points required for volume rendering. With the recent introduction of 3D Gaussian Splatting as a faster alternative to implicit representation methods, we propose to leverage this new rendering approach to achieve faster multi-sensor calibration. We introduce 3DGS-Calib, a new calibration method that relies on the speed and rendering accuracy of 3D Gaussian Splatting to achieve multimodal spatiotemporal calibration that is accurate, robust, and with a substantial speed-up compared to methods relying on implicit neural representations. We demonstrate the superiority of our proposal with experimental results on sequences from KITTI-360, a widely used driving dataset. Quentin Herau, Moussâb Bennehar, Arthur Moreau, Nathan Piasco, Luis Roldão, Dzmitry Tsishkou, Cyrille Migniot, Pascal Vasseur, Cédric Demonceaux |
IROS | 9 |
| 2024 | Joint 3D Shape and Motion Estimation from Rolling Shutter Light-Field ImagesabstractIn this paper, we propose an approach to address the problem of 3D reconstruction of scenes from a single image captured by a light-field camera equipped with a rolling shutter sensor. Our method leverages the 3D information cues present in the light-field and the motion information provided by the rolling shutter effect. We present a generic model for the imaging process of this sensor and a two-stage algorithm that minimizes the re-projection error while considering the position and motion of the camera in a motion-shape bundle adjustment estimation strategy. Thereby, we provide an instantaneous 3D shape-and-pose-and-velocity sensing paradigm. To the best of our knowledge, this is the first study to leverage this type of sensor for this purpose. We also present a new benchmark dataset composed of different light-fields showing rolling shutter effects, which can be used as a common base to improve the evaluation and tracking the progress in the field. We demonstrate the effectiveness and advantages of our approach through several experiments conducted for different scenes and types of motions. The source code and dataset are publicly available at: https://github.com/ICB-Vision-AI/RSLF. Hermes McGriff, Renato Martins, Nicolas Andreff, Cédric Demonceaux |
WACV | 4 |
| 2024 | Transformer fusion for indoor RGB-D semantic segmentationabstractFusing geometric cues with visual appearance is an imperative theme for RGB-D indoor semantic segmentation . Existing methods commonly adopt convolutional modules to aggregate multi-modal features, paying little attention to explicitly leveraging the long-range dependencies in feature fusion . Therefore, it is challenging for existing methods to accurately segment objects with large-scale variations. In this paper, we propose a novel transformer-based fusion scheme, named TransD-Fusion, to better model contextualized awareness. Specifically, TransD-Fusion consists of a self-refinement module, a calibration scheme with cross-interaction, and a depth-guided fusion. The objective is to first improve modality-specific features with self- and cross-attention, and then explore the geometric cues to better segment objects sharing a similar visual appearance. Additionally, our transformer fusion benefits from a semantic-aware position encoding which spatially constrains the attention to neighboring pixels . Extensive experiments on RGB-D benchmarks demonstrate that the proposed method performs well over the state-of-the-art methods by large margins. Zongwei Wu, Zhuyun Zhou, Guillaume Allibert, Christophe Stolz, Cédric Demonceaux, Chao Ma 0004 |
Comput. Vis. Image Underst. | 5 |
| 2024 | N-QGNv2: Predicting the optimum quadtree representation of a depth map from a monocular cameraabstractSelf-supervised monocular depth prediction is a widely researched field that aims to provide a better scene understanding. However, most existing methods prioritize prediction accuracy over computation cost, which can hinder the deployment of these methods in real-world applications. Our objective is to propose a solution that efficiently compresses the depth map while maintaining a high level of accuracy for navigation purpose. The proposed method is an expansion of the work presented in N-QGN, which utilizes a quadtree representation for compression. This approach has already shown promising results, but we aim to improve it further by making it more accurate, faster, and easier to train. Therefore, we introduce a new method that directly predicts the quadtree structure, resulting in a more consistent prediction, and we revise the network architecture to be lighter and produce state-of-the-art accuracy results, depending on the data compression rate. The new implementation is also faster, making it more suitable for real-time applications. Experiments have been conducted on various scene configuration highlighting the capability of the method to efficiently predicting a reliable quadtree depth representation of the scene at low computation cost and high accuracy. Daniel Braun 0008, Olivier Morel, Cédric Demonceaux, Pascal Vasseur |
Pattern Recognit. Lett. | 3 |
| 2023 | Alignment-free HDR Deghosting with Semantics Consistent TransformerabstractHigh dynamic range (HDR) imaging aims to retrieve information from multiple low-dynamic range inputs to generate realistic output. The essence is to leverage the contextual information, including both dynamic and static semantics, for better image generation. Existing methods often focus on the spatial misalignment across input frames caused by the foreground and/or camera motion. However, there is no research on jointly leveraging the dynamic and static context in a simultaneous manner. To delve into this problem, we propose a novel alignment-free network with a Semantics Consistent Transformer (SCTNet) with both spatial and channel attention modules in the network. The spatial attention aims to deal with the intra-image correlation to model the dynamic motion, while the channel attention enables the inter-image intertwining to enhance the semantic consistency across frames. Aside from this, we introduce a novel realistic HDR dataset with more variations in foreground objects, environmental factors, and larger motions. Extensive comparisons on both conventional datasets and ours validate the effectiveness of our method, achieving the best trade-off on the performance and the computational cost. The source code and dataset are available at https://steven-tel.github.io/sctnet/. Steven Tel, Zongwei Wu, Yulun Zhang 0001, Barthélémy Heyrman, Cédric Demonceaux, Radu Timofte, Dominique Ginhac |
ICCV | 5 |
| 2023 | Source-free Depth for Object Pop-outabstractDepth cues are known to be useful for visual perception. However, direct measurement of depth is often impracticable. Fortunately, though, modern learning-based methods offer promising depth maps by inference in the wild. In this work, we adapt such depth inference models for object segmentation using the objects’ "pop-out" prior in 3D. The "pop-out" is a simple composition prior that assumes objects reside on the background surface. Such compositional prior allows us to reason about objects in the 3D space. More specifically, we adapt the inferred depth maps such that objects can be localized using only 3D information. Such separation, however, requires knowledge about contact surface which we learn using the weak supervision of the segmentation mask. Our intermediate representation of contact surface, and thereby reasoning about objects purely in 3D, allows us to better transfer the depth knowledge into semantics. The proposed adaptation method uses only the depth model without needing the source data used for training, making the learning process efficient and practical. Our experiments on eight datasets of two challenging tasks, namely salient object detection and camouflaged object detection, consistently demonstrate the benefit of our method in terms of both performance and generalizability. The source code is publicly available at https://github.com/Zongwei97/PopNet. Zongwei Wu, Danda Pani Paudel, Deng-Ping Fan, Shuo Wang 0010, Cédric Demonceaux, Radu Timofte, Luc Van Gool |
ICCV | 6 |
| 2023 | RGB-Event Fusion for Moving Object Detection in Autonomous DrivingabstractMoving Object Detection (MOD) is a critical vision task for successfully achieving safe autonomous driving. Despite plausible results of deep learning methods, most existing approaches are only frame-based and may fail to reach reasonable performance when dealing with dynamic traffic participants. Recent advances in sensor technologies, especially the Event camera, can naturally complement the conventional camera approach to better model moving objects. However, event-based works often adopt a pre-defined time window for event representation, and simply integrate it to estimate image intensities from events, neglecting much of the rich temporal information from the available asynchronous events. Therefore, from a new perspective, we propose RENet, a novel RGB-Event fusion Network, that jointly exploits the two complementary modalities to achieve more robust MOD under challenging scenarios for autonomous driving. Specifically, we first design a temporal multi-scale aggregation module to fully leverage event frames from both the RGB exposure time and larger intervals. Then we introduce a bi-directional fusion module to attentively calibrate and fuse multi-modal features. To evaluate the performance of our network, we carefully select and annotate a sub-MOD dataset from the commonly used DSEC dataset. Extensive experiments demonstrate that our proposed method performs significantly better than the state-of-the-art RGB-Event fusion alternatives. The source code and dataset are publicly available at: https://github.com/ZZY-Zhou/RENet. Zhuyun Zhou, Zongwei Wu, Rémi Boutteau, Fan Yang 0019, Cédric Demonceaux, Dominique Ginhac |
ICRA | 5 |
| 2023 | MOISST: Multimodal Optimization of Implicit Scene for SpatioTemporal CalibrationabstractWith the recent advances in autonomous driving and the decreasing cost of LiDARs, the use of multimodal sensor systems is on the rise. However, in order to make use of the information provided by a variety of complimentary sensors, it is necessary to accurately calibrate them. We take advantage of recent advances in computer graphics and implicit volumetric scene representation to tackle the problem of multi-sensor spatial and temporal calibration. Thanks to a new formulation of the Neural Radiance Field (NeRF) optimization, we are able to jointly optimize calibration parameters along with scene representation based on radiometric and geometric measurements. Our method enables accurate and robust calibration from data captured in uncontrolled and unstructured urban environments, making our solution more scalable than existing calibration solutions. We demonstrate the accuracy and robustness of our method in urban scenes typically encountered in autonomous driving scenarios. Quentin Herau, Nathan Piasco, Moussâb Bennehar, Luis Roldão, Dzmitry Tsishkou, Cyrille Migniot, Pascal Vasseur, Cédric Demonceaux |
IROS | 8 |
| 2023 | Object Segmentation by Mining Cross-Modal SemanticsabstractMulti-sensor clues have shown promise for object segmentation, but inherent noise in each sensor, as well as the calibration error in practice, may bias the segmentation accuracy. In this paper, we propose a novel approach by mining the Cross-Modal Semantics to guide the fusion and decoding of multimodal features, with the aim of controlling the modal contribution based on relative entropy. We explore semantics among the multimodal inputs in two aspects: the modality-shared consistency and the modality-specific variation. Specifically, we propose a novel network, termed XMSNet, consisting of (1) all-round attentive fusion (AF), (2) coarse-to-fine decoder (CFD), and (3) cross-layer self-supervision. On the one hand, the AF block explicitly dissociates the shared and specific representation and learns to weight the modal contribution by adjusting the proportion, region, and pattern, depending upon the quality. On the other hand, our CFD initially decodes the shared feature and then refines the output through specificity-aware querying. Further, we enforce semantic consistency across the decoding layers to enable interaction across network hierarchies, improving feature discriminability. Exhaustive comparison on eleven datasets with depth or thermal clues, and on two challenging tasks, namely salient and camouflage object segmentation, validate our effectiveness in terms of both performance and robustness. The source code is publicly available at https://github.com/Zongwei97/XMSNet. Zongwei Wu, Zhuyun Zhou, Zhaochong An, Qiuping Jiang, Cédric Demonceaux, Guolei Sun, Radu Timofte |
ACM Multimedia | 6 |
| 2023 | Bidirectional Collaborative Mentoring Network for Marine Organism Detection and BeyondabstractOrganism detection plays a vital role in marine resource exploitation and marine economy. How to accurately locate the target organism object within the camouflaged and dark light oceanic scene has recently drawn great attention in the research community. Existing learning-based works usually leverage local texture details within a neighboring area, with few methods explicitly exploring the usage of contextualized awareness for accurate object detection. From a novel perspective, we in this work present a Bidirectional Collaborative Mentoring Network (BCMNet) which fully explores both texture and context clues during the encoding and decoding stages, making the cross-paradigm interaction bidirectional and improving the scene understanding at all stages. Specifically, we first extract texture and context features through a dual-branch encoder and attentively fuse them through our adjacent feature fusion (AFF) block. Then, we propose a structure-aware module (SAM) and a detail-enhanced module (DEM) to form our two-stage decoding pipeline. On the one hand, our SAM leverages both local and global clues to preserve morphological integrity and generate an initial prediction of the target object. On the other hand, the DEM explicitly explores long-range dependencies to refine the initially predicted object mask further. The combination of SAM and DEM enables better extracting, preserving, and enhancing the object morphology, making it easier to segment the target object from the camouflaged background with sharp contour. Extensive experiments on three benchmark datasets show that our proposed BCMNet performs favorably over state-of-the-art models. The code will be made available athttps://github.com/chasecjg/BCMNet. Jinguang Cheng, Zongwei Wu, Shuo Wang 0010, Cédric Demonceaux, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | HiDAnet: RGB-D Salient Object Detection via Hierarchical Depth AwarenessabstractRGB-D saliency detection aims to fuse multi-modal cues to accurately localize salient regions. Existing works often adopt attention modules for feature modeling, with few methods explicitly leveraging fine-grained details to merge with semantic cues. Thus, despite the auxiliary depth information, it is still challenging for existing models to distinguish objects with similar appearances but at distinct camera distances. In this paper, from a new perspective, we propose a novel Hierarchical Depth Awareness network (HiDAnet) for RGB-D saliency detection. Our motivation comes from the observation that the multi-granularity properties of geometric priors correlate well with the neural network hierarchies. To realize multi-modal and multi-level fusion, we first use a granularity-based attention scheme to strengthen the discriminatory power of RGB and depth features separately. Then we introduce a unified cross dual-attention module for multi-modal and multi-level fusion in a coarse-to-fine manner. The encoded multi-modal features are gradually aggregated into a shared decoder. Further, we exploit a multi-scale loss to take full advantage of the hierarchical information. Extensive experiments on challenging benchmark datasets demonstrate that our HiDAnet performs favorably over the state-of-the-art methods by large margins. The source code can be found in https://github.com/Zongwei97/HIDANet/. Zongwei Wu, Guillaume Allibert, Fabrice Mériaudeau, Chao Ma 0004, Cédric Demonceaux |
IEEE Trans. Image Process. | 5 |
| 2022 | Robust RGB-D Fusion for Saliency DetectionabstractEfficiently exploiting multi-modal inputs for accurate RGB-D saliency detection is a topic of high interest. Most existing works leverage cross-modal interactions to fuse the two streams of RGB-D for intermediate features' enhancement. In this process, a practical aspect of the low quality of the available depths has not been fully considered yet. In this work, we aim for RGB-D saliency detection that is robust to the low-quality depths which primarily appear in two forms: inaccuracy due to noise and the misalignment to RGB. To this end, we propose a robust RGB-D fusion method that benefits from (1) layer-wise, and (2) trident spatial, attention mechanisms. On the one hand, layer-wise attention (LWA) learns the trade-off between early and late fusion of RGB and depth features, depending upon the depth accuracy. On the other hand, trident spatial attention (TSA) aggregates the features from a wider spatial context to address the depth misalignment problem. The proposed LWA and TSA mechanisms allow us to efficiently exploit the multi-modal inputs for saliency detection while being robust against low-quality depths. Our experiments on five bench-mark datasets demonstrate that the proposed fusion method performs consistently better than the state-of-the-art fusion alternatives. The source code is publicly available at: https://github.com/Zongwei97/RFnet. Zongwei Wu, Shriarulmozhivarman Gobichettipalayam, Brahim Tamadazte, Guillaume Allibert, Danda Pani Paudel, Cédric Demonceaux |
3DV | 6 |
| 2022 | Leveraging Semantic Cues from Foundation Vision Models for Enhanced Local Feature Correspondence
Felipe C. Chamone, Guilherme A. Potje, Renato Martins, Cédric Demonceaux, Erickson R. Nascimento |
ACCV (4) | 4 |
| 2022 | Deep Reinforcement Learning with Omnidirectional Images: application to UAV Navigation in ForestsabstractDeep Reinforcement Learning (DRL) is highly efficient for solving complex tasks such as drone obstacle avoidance using cameras. However, these methods are often limited by the camera perception capabilities. In this paper, we demonstrate that point-goal navigation performances can be improved by using cameras with a wider Field-Of-View (FOV). To this end, we present a DRL solution based on equirectangular images and demonstrates its relevance, especially compared to its perspective version. Several visual modalities are compared: ground truth depth, RGB, and depth directly estimated from these$360^{\circ}$RGB images using Deep Learning methods. Next, we propose a spherical adaptation to take into account the spherical distortions of omnidirectional images in the convolutional neural networks (CNNs) used in the actor-critic network and show a significant improvement in navigation performance. Finally, we modify the perspective depth estimation network using this spherical adaptation and demonstrate a further performance improvement. Charles-Olivier Artizzu, Guillaume Allibert, Cédric Demonceaux |
ICARCV | 3 |
| 2022 | N-QGN: Navigation Map from a Monocular Camera using Quadtree Generating NetworksabstractMonocular depth estimation has been a popu-lar area of research for several years, especially since self-supervised networks have shown increasingly good results in bridging the gap with supervised and stereo methods. However, these approaches focus their interest on dense 3D reconstruction and sometimes on tiny details that are superfluous for autonomous navigation. In this paper, we propose to address this issue by estimating the navigation map under a quad tree representation. The objective is to create an adaptive depth map prediction that only extract details that are essential for the obstacle avoidance. Other 3D space which leaves large room for navigation will be provided with approximate distance. Experiment on KITTI dataset shows that our method can significantly reduce the number of output information without major loss of accuracy. Daniel Braun 0008, Olivier Morel, Pascal Vasseur, Cédric Demonceaux |
ICRA | 4 |
| 2022 | Trifocal Tensor and Relative Pose Estimation from 8 Lines and Known Vertical DirectionabstractIn this paper, we present a relative pose estimation algorithm based on lines knowing the vertical direction associated to each image. We demonstrate that a closed-form solution requiring only eight lines between three views is possible. As a linear solution, it is shown that our approach outperforms the standard trifocal estimation based on 13 triplets of lines and can be efficiently inserted into an hypothesize-and-test framework such as RANSAC. We also study our approach on different singular configurations of lines. The method is evaluated on both synthetic data and real-world sequences from KITTI and the Zürich Urban Micro Aerial Vehicle datasets. Our method is compared to 13 lines algorithm as well to points based methods such as 7-points, 5-points and 3-points. Banglei Guan, Pascal Vasseur, Cédric Demonceaux |
IROS | 3 |
| 2021 | FLYBO: A Unified Benchmark Environment for Autonomous Flying RobotsabstractThe use of Micro-Aerial Vehicles (MAVs) equipped with odometry- and depth sensors has become predominant for a wide variety of challenging industrial applications such as the autonomous exploration (i.e., digital mapping), and inspection (i.e., online surface reconstruction) of unknown facilities. However, despite the ongoing attention these topics receive, autonomous exploration systems still lack common evaluation grounds to assess their relative performance in terms of data and experimental tools. We address this deficit by introducing FLYBO, the first unified benchmark environment that focuses on the performance of such flying robots in terms of autonomous exploration and online surface reconstruction. It includes (i) 11 challenging realistic indoor- and outdoor datasets of increasing complexity and size, with ground-truth, (ii) a comprehensive benchmark of 7 of the top-performing autonomous exploration algorithms including methods without publicly available code. (iii) A unified experimental system factorizes the routines shared by autonomous planners in order to fairly and accurately assess their exploration performance in a controlled environment. Anthony Brunel, Amine Bourki, Olivier Strauss, Cédric Demonceaux |
3DV | 4 |
| 2021 | Modality-Guided Subnetwork for Salient Object DetectionabstractRecent RGBD-based models for saliency detection have attracted research attention. The depth clues such as boundary clues, surface normal, shape attribute, etc., contribute to the identification of salient objects with complicated scenarios. However, most RGBD networks require multi-modalities from the input side and feed them separately through a two-stream design, which inevitably results in extra costs on depth sensors and computation. To tackle these inconveniences, we present in this paper a novel fusion design named modality-guided subnetwork (MGSnet). It has the following superior designs: 1) Our model works for both RGB and RGBD data, and dynamically estimates depth if not available. Taking the inner workings of depth-prediction networks into account, we propose to estimate the pseudo-geometry maps from RGB input — essentially mimicking the multi-modality input. 2) Our MGSnet for RGB SOD results in real-time inference but achieves state-of-the-art performance compared to other RGB models. 3) The flexible and lightweight design of MGS facilitates the integration into RGBD two-streaming models. The introduced fusion design enables a cross-modality interaction to enable further progress but with a minimal cost. Zongwei Wu, Guillaume Allibert, Christophe Stolz, Chao Ma 0004, Cédric Demonceaux |
3DV | 5 |
| 2021 | SplatPlanner: Efficient Autonomous Exploration via Permutohedral Frontier FilteringabstractWe address the problem of autonomous exploration of unknown environments using a Micro Aerial Vehicle (MAV) equipped with an active depth sensor. As such, the task consists in mapping the gradually discovered environment while planning the envisioned trajectories in real-time, using on-board computation only. To do so, we present SplatPlanner, an end-to-end autonomous planner that is based on a novel Permutohedral Frontier Filtering (PFF) which relies on a combination of highly efficient operations stemming from bilateral filtering using permutohedral lattices to guide the entire exploration. In particular, our PFF is computationally linear in input size, nearly parameter-free, and aggregates spatial information about frontier-neighborhoods into density scores in one single step. Comparative experiments made on simulated environments of increasing complexity show our method consistently outperforms recent state-of-the-art methods in terms of computational efficiency, exploration speed and qualitative coverage of scenes. Finally, we also display the practical capabilities of our end-to-end system in a challenging real-flight scenario. Anthony Brunel, Amine Bourki, Cédric Demonceaux, Olivier Strauss |
ICRA | 3 |
| 2021 | Improving Image Description with Auxiliary Modality for Visual Localization in Challenging Conditions
Nathan Piasco, Desire Sidibé, Valérie Gouet-Brunet, Cédric Demonceaux |
Int. J. Comput. Vis. | 4 |
| 2021 | Moving Object Detection by 3D Flow Field AnalysisabstractMap-based localization and sensing are one of the key components in autonomous driving technologies, where high quality 3D map reconstruction is fundamentally utmost important. However, due to the highly dynamic and uncontrollable properties of real world environment, building a high quality 3D map is not straightforward and requires several strong assumptions. To address this challenge, we present a complete framework, which detects and extracts the moving objects from a sequence of unordered and texture-less point clouds, to build high quality static maps. To accurately detect the moving objects from data acquired with a possibly fast moving platform, we propose a novel 3D Flow Field Analysis approach in which we inspect the motion behaviour of the registered point sets. The proposed algorithm elegantly models the temporal and spatial displacement of the moving objects. Thus, both small moving objects (e.g. walking pedestrians) and large moving objects (e.g. moving trucks) can be detected effectively. Further, by incorporating the Sparse Subspace Clustering framework, we propose a Sparse Flow Clustering algorithm to group the 3D motion flows under both the constraints of motion similarity and spatial closeness. To this end, the static scene parts and the moving objects can be independently processed to achieve photo-realistic 3D reconstructions. Finally, we show that the proposed 3D Flow Field Analysis algorithm and the Sparse Flow Clustering approach are highly effective for motion detection and segmentation, as exemplified on the KITTI benchmark, and yield high quality reconstructed static-maps as well as rigidly moving objects. Cansen Jiang, Danda Pani Paudel, David Fofi, Yohan D. Fougerolle, Cédric Demonceaux |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | Depth-Adapted CNN for RGB-D Cameras
Zongwei Wu, Guillaume Allibert, Christophe Stolz, Cédric Demonceaux |
ACCV (4) | 4 |
| 2020 | Stratified Autocalibration of Cameras with Euclidean Image Plane
Devesh Adlakha, Adlane Habed, Fabio Morbidi, Cédric Demonceaux, Michel de Mathelin |
BMVC | 4 |
| 2020 | Unsupervised Learning of Category-Specific Symmetric 3D Keypoints from Point Sets
Clara Fernandez-Labrador, Ajad Chhatkuli, Danda Pani Paudel, Josechu J. Guerrero, Cédric Demonceaux, Luc Van Gool |
ECCV (25) | 5 |
| 2020 | OmniFlowNet: a Perspective Neural Network Adaptation for Optical Flow Estimation in Omnidirectional ImagesabstractSpherical cameras and the latest image processing techniques open up new horizons. In particular, methods based on Convolutional Neural Networks (CNNs) now give excellent results for optical flow estimation on perspective images. However, these approaches are highly dependent on their architectures and training datasets. This paper proposes to benefit from years of improvement in perspective images optical flow estimation and to apply it to omnidirectional ones without training on new datasets. Our network, OmniFlowNet, is built on a CNN specialized in perspective images. Its convolution operation is adapted to be consistent with the equirectangular projection. Tested on spherical datasets created with Blender1and several equirectangular videos realized from real indoor and outdoor scenes, OmniFlowNet shows better performance than its original network without extra training. Charles-Olivier Artizzu, Haozhou Zhang, Guillaume Allibert, Cédric Demonceaux |
ICPR | 4 |
| 2020 | What's in my Room? Object Recognition on Indoor Panoramic ImagesabstractIn the last few years, there has been a growing interest in taking advantage of the 360°panoramic images potential, while managing the new challenges they imply. While several tasks have been improved thanks to the contextual information these images offer, object recognition in indoor scenes still remains a challenging problem that has not been deeply investigated. This paper provides an object recognition system that performs object detection and semantic segmentation tasks by using a deep learning model adapted to match the nature of equirectangular images. From these results, instance segmentation masks are recovered, refined and transformed into 3D bounding boxes that are placed into the 3D model of the room. Quantitative and qualitative results support that our method outperforms the state of the art by a large margin and show a complete understanding of the main objects in indoor scenes. Julia Guerrero-Viu, Clara Fernandez-Labrador, Cédric Demonceaux, Josechu J. Guerrero |
ICRA | 3 |
| 2019 | Perspective-n-Learned-Point: Pose Estimation from Relative Depth
Nathan Piasco, Desire Sidibé, Cédric Demonceaux, Valérie Gouet-Brunet |
BMVC | 3 |
| 2019 | QUARCH: A New Quasi-Affine Reconstruction Stratum From Vague Relative Camera Orientation KnowledgeabstractWe present a new quasi-affine reconstruction of a scene and its application to camera self-calibration. We refer to this reconstruction as QUARCH (QUasi-Affine Reconstruction with respect to Camera centers and the Hodographs of horopters). A QUARCH can be obtained by solving a semidefinite programming problem when, (i) the images have been captured by a moving camera with constant intrinsic parameters, and (ii) a vague knowledge of the relative orientation (under or over 120 degrees) between camera pairs is available. The resulting reconstruction comes close enough to an affine one allowing thus an easy upgrade of the QUARCH to its affine and metric counterparts. We also present a constrained Levenberg-Marquardt method for nonlinear optimization subject to Linear Matrix Inequality (LMI) constraints so as to ensure that the QUARCH LMIs are satisfied during optimization. Experiments with synthetic and real data show the benefits of QUARCH in reliably obtaining a metric reconstruction. Devesh Adlakha, Adlane Habed, Fabio Morbidi, Cédric Demonceaux, Michel de Mathelin |
ICCV | 4 |
| 2019 | Geometric Camera Pose Refinement with Learned Depth MapsabstractWe present a new method for image-only camera relocalisation composed of a fast image indexing retrieval step followed by pose refinement based on ICP (Iterative Closest Point). The first step aims to find an initial pose for the query by evaluating images similarity with low dimensional global deep descriptors. Subsequently, we predict with a fully convolutional deep encoder-decoder neural network a dense depth map from the image query. We use this depth map to create a local point cloud and refine the initial query pose using an ICP algorithm.We demonstrate the effectiveness of our new approach on various indoor scenes. Compared to learned pose regression methods, our proposal can be used on multiple scenes without the need of a specific weights-setup for each scene, while showing equivalent results. Nathan Piasco, Desire Sidibé, Cédric Demonceaux, Valérie Gouet-Brunet |
ICIP | 3 |
| 2019 | Learning Scene Geometry for Visual Localization in Challenging ConditionsabstractWe propose a new approach for outdoor large scale image based localization that can deal with challenging scenarios like cross-season, cross-weather, day/night and long-term localization. The key component of our method is a new learned global image descriptor, that can effectively benefit from scene geometry information during training. At test time, our system is capable of inferring the depth map related to the query image and use it to increase localization accuracy. We are able to increase recall@1 performances by 2.15% on cross-weather and long-term localization scenario and by 4.24% points on a challenging winter/summer localization sequence versus state-of-the-art methods. Our method can also use weakly annotated data to localize night images across a reference dataset of daytime images. Nathan Piasco, Desire Sidibé, Valérie Gouet-Brunet, Cédric Demonceaux |
ICRA | 4 |
| 2019 | Robust and Optimal Registration of Image Sets and Structured Scenes via Sum-of-Squares Polynomials
Danda Pani Paudel, Adlane Habed, Cédric Demonceaux, Pascal Vasseur |
Int. J. Comput. Vis. | 3 |
| 2018 | Multimodal 2D Image to 3D Model Registration via a Mutual Alignment of Sparse and Dense Visual FeaturesabstractMany fields of application could benefit from an accurate registration of measurements of different modalities over a known 3D model. However, aligning a 2D image to a 3D model is a challenging task and is even more complex when the two have a different modality. Most of the 2D/3D registration methods are based on either geometric or dense visual features. Both have their own advantages and their own drawbacks. We propose, in this paper, to mutually exploit the advantages of one feature type to reduce the drawbacks of the other one. For this, an hybrid registration framework has been designed to mutually align geometrical and dense visual features in order to obtain an accurate final 2D/3D alignment. We evaluate and compare the proposed registration method on real data acquired by a robot equipped with several visual sensors. The results highlights the robustness of the method and its ability to produce wide convergence domain and a high registration accuracy. Nathan Crombez, Ralph Seulin, Olivier Morel, David Fofi, Cédric Demonceaux |
ICRA | 5 |
| 2018 | Visual Odometry Using a Homography Formulation with Decoupled Rotation and Translation Estimation Using Minimal SolutionsabstractIn this paper we present minimal solutions for two-view relative motion estimation based on a homography formulation. By assuming a known vertical direction (e.g. from an IMU) and assuming a dominant ground plane we demonstrate that rotation and translation estimation can be decoupled. This result allows us to reduce the number of point matches needed to compute a motion hypothesis. We then derive different algorithms based on this decoupling that allow an efficient estimation. We also demonstrate how these algorithms can be used efficiently to compute an optimal inlier set using exhaustive search or histogram voting instead of a traditional RANSAC step. Our methods are evaluated on synthetic data and on the KITTI data set, demonstrating that our methods are well suited for visual odometry in road driving scenarios. Banglei Guan, Pascal Vasseur, Cédric Demonceaux, Friedrich Fraundorfer |
ICRA | 3 |
| 2018 | Attitude Estimation from Polarimetric CamerasabstractIn the robotic field, navigation and path planning applications benefit from a wide range of visual systems (e.g, perspective cameras, depth cameras, catadioptric cameras, etc.). In outdoor conditions, these systems capture information in which sky regions cover a major segment of the images acquired. However, sky regions are discarded and are not considered as visual cue in vision applications. In this paper, we propose to estimate attitude of Unmanned Aerial Vehicle (UAV) from sky information using a polarimetric camera. Theoretically, we provide a framework estimating the attitude from the skylight polarized patterns. We showcase this formulation on both simulated and real-word data sets which proved the benefit of using polarimetric sensors along with other visual sensors in robotic applications. Mojdeh Rastgoo, Cédric Demonceaux, Ralph Seulin, Olivier Morel |
IROS | 2 |
| 2018 | Summarizing Large Scale 3D MeshabstractRecent progress in 3D sensor devices and in semantic mapping allows to build very rich HD 3D maps very useful for autonomous navigation and localization. However, these maps are particularly huge and require important memory capabilities as well computational resources. In this paper, we propose a new method for summarizing a 3D map (Mesh)as a set of compact spheres in order to facilitate its use by systems with limited resources (smartphones, robots, UAVs,...). This vision-based summarizing process is applied in a fully automatic way using jointly photometric, geometric and semantic information of the studied environment. The main contribution of this research is to provide a very compact map that maximizes the significance of its content while maintaining the full visibility of the environment. Experimental results in summarizing large-scale 3D map demonstrate the feasibility of our approach and evaluate the performance of the algorithm. Imeen Ben Salah, Sébastien Kramm, Cédric Demonceaux, Pascal Vasseur |
IROS | 3 |
| 2018 | A survey on Visual-Based Localization: On the benefit of heterogeneous data
Nathan Piasco, Desire Sidibé, Cédric Demonceaux, Valérie Gouet-Brunet |
Pattern Recognit. | 3 |
| 2018 | A hierarchical stereo matching algorithm based on adaptive support region aggregation method
Oussama Zeglazi, Mohammed Rziza, Aouatif Amine, Cédric Demonceaux |
Pattern Recognit. Lett. | 4 |
| 2017 | Static and Dynamic Objects Analysis as a 3D Vector FieldabstractIn the context of scene modelling, understanding, and landmark-based robot navigation, the knowledge of static scene parts and moving objects with their motion behaviours plays a vital role. We present a complete framework to detect and extract the moving objects to reconstruct a high quality static map. For a moving 3D camera setup, we propose a novel 3D Flow Field Analysis approach which accurately detects the moving objects using only 3D point cloud information. Further, we introduce a Sparse Flow Clustering approach to effectively and robustly group the motion flow vectors. Experiments show that the proposed Flow Field Analysis algorithm and Sparse Flow Clustering approach are highly effective for motion detection and segmentation, and yield high quality reconstructed static maps as well as rigidly moving objects of real-world scenarios. Cansen Jiang, Danda Pani Paudel, Yohan D. Fougerolle, David Fofi, Cédric Demonceaux |
3DV | 5 |
| 2017 | High quality reconstruction of dynamic objects using 2D-3D camera fusionabstractIn this paper, we propose a complete pipeline for high quality reconstruction of dynamic objects using 2D-3D camera setup attached to a moving vehicle. Starting from the segmented motion trajectories of individual objects, we compute their precise motion parameters, register multiple sparse point clouds to increase the density, and develop a smooth and textured surface from the dense (but scattered) point cloud. The success of our method relies on the proposed optimization framework for accurate motion estimation between two sparse point clouds. Our formulation for fusing closest-point and consensus based motion estimations, respectively in the absence and presence of motion trajectories, is the key to obtain such accuracy. Several experiments performed on both synthetic and real (KITTI) datasets show that the proposed framework is very robust and accurate. Cansen Jiang, Dennis Christie, Danda Pani Paudel, Cédric Demonceaux |
ICIP | 4 |
| 2017 | Accurate dense stereo matching for road scenesabstractStereo matching task is the core of applications linked to the intelligent vehicles. In this paper, we present a new variant function of the Census Transform (CT) which is more robust against radiometric changes in real road scenes. We demonstrate that the proposed cost function outperforms the conventional cost functions using the KITTI benchmark1. The cost aggregation method is also updated for taking into account the edge information. This enables to improve significantly the aggregated costs especially within homogenous regions. The Winner-Takes-All (WTA) strategy is used to compute disparity values. To further eliminate the remainder matching ambiguities, a post-processing step is performed. Experiments were conducted on the new Middlebury2dataset, as well as on the real road traffic scenes of the KITTI database. Obtained disparity results have demonstrated that the proposed method is promising. Oussama Zeglazi, Mohammed Rziza, Aouatif Amine, Cédric Demonceaux |
ICIP | 4 |
| 2017 | Incomplete 3D motion trajectory segmentation and 2D-to-3D label transfer for dynamic scene analysisabstractThe knowledge of the static scene parts and the moving objects in a dynamic scene plays a vital role for scene modelling, understanding, and landmark-based robot navigation. The key information for these tasks lies on semantic labels of the scene parts and the motion trajectories of the dynamic objects. In this work, we propose a method that segments the 3D feature trajectories based on their motion behaviours, and assigns them semantic labels using 2D-to-3D label transfer. These feature trajectories are constructed by using the proposed trajectory recovery algorithm which takes the loss of feature tracking into account. We introduce a complete framework for static-map and dynamic objects' reconstruction, as well as semantic scene understanding for a calibrated and moving 2D-3D camera setup. Our motion segmentation approach is faster by two orders of magnitude, while performing better than the state-of-the-art 3D motion segmentation methods, and successfully handles the previously discarded incomplete trajectory scenarios. Cansen Jiang, Danda Pani Paudel, Yohan D. Fougerolle, David Fofi, Cédric Demonceaux |
IROS | 5 |
| 2017 | Homography Based Egomotion Estimation with a Common DirectionabstractIn this paper, we explore the different minimal solutions for egomotion estimation of a camera based on homography knowing the gravity vector between calibrated images. These solutions depend on the prior knowledge about the reference plane used by the homography. We then demonstrate that the number of matched points can vary from two to three and that a direct closed-form solution or a Gröbner basis based solution can be derived according to this plane. Many experimental results on synthetic and real sequences in indoor and outdoor environments show the efficiency and the robustness of our approach compared to standard methods. Olivier Saurer, Pascal Vasseur, Rémi Boutteau, Cédric Demonceaux, Marc Pollefeys, Friedrich Fraundorfer |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Line reconstruction using prior knowledge in single non-central view
Jesus Bermudez-Cameo, Cédric Demonceaux, Gonzalo López-Nicolás, Josechu J. Guerrero |
BMVC | 2 |
| 2016 | Fast earth mover's distance computation for catadioptric image sequencesabstractEarth mover's distance is one of the most effective metric for comparing histograms in various image retrieval applications. The main drawback is its computational complexity which hinders its usage in various comparison tasks. We propose fast earth mover's distance computation by providing better initialization to the transportation simplex algorithm. The new approach enables faster EMD computation in Visual Memory (VM) compared to the state of the art methods. The new proposed strategy computes earth mover distance without compromising its accuracy. Omar Tahri, Cédric Demonceaux, David Fofi, Mohamad Mazen Hittawe |
ICIP | 3 |
| 2015 | LMI-based 2D-3D registration: From uncalibrated images to Euclidean sceneabstractThis paper investigates the problem of registering a scanned scene, represented by 3D Euclidean point coordinates, and two or more uncalibrated cameras. An unknown subset of the scanned points have their image projections detected and matched across images. The proposed approach assumes the cameras only known in some arbitrary projective frame and no calibration or autocalibration is required. The devised solution is based on a Linear Matrix Inequality (LMI) framework that allows simultaneously estimating the projective transformation relating the cameras to the scene and establishing 2D-3D correspondences without triangulating image points. The proposed LMI framework allows both deriving triangulation-free LMI cheirality conditions and establishing putative correspondences between 3D volumes (boxes) and 2D pixel coordinates. Two registration algorithms, one exploiting the scene's structure and the other concerned with robustness, are presented. Both algorithms employ the Branch-and-Prune paradigm and guarantee convergence to a global solution under mild initial bound conditions. The results of our experiments are presented and compared against other approaches. Danda Pani Paudel, Adlane Habed, Cédric Demonceaux, Pascal Vasseur |
CVPR | 3 |
| 2015 | Robust and Optimal Sum-of-Squares-Based Point-to-Plane Registration of Image Sets and Structured ScenesabstractThis paper deals with the problem of registering a known structured 3D scene and its metric Structure-from-Motion (SfM) counterpart. The proposed work relies on a prior plane segmentation of the 3D scene and aligns the data obtained from both modalities by solving the point-to-plane assignment problem. An inliers-maximization approach within a Branch-and-Bound (BnB) search scheme is adopted. For the first time in this paper, a Sum-of-Squares optimization theory framework is employed for identifying point-to-plane mismatches (i.e. outliers) with certainty. This allows us to iteratively build potential inliers sets and converge to the solution satisfied by the largest number of point-to-plane assignments. Furthermore, our approach is boosted by new plane visibility conditions which are also introduced in this paper. Using this framework, we solve the registration problem in two cases: (i) a set of putative point-to-plane correspondences (with possibly overwhelmingly many outliers) is given as input and (ii) no initial correspondences are given. In both cases, our approach yields outstanding results in terms of robustness and optimality. Danda Pani Paudel, Adlane Habed, Cédric Demonceaux, Pascal Vasseur |
ICCV | 3 |
| 2015 | Visual Servoing Based on Shifted MomentsabstractOver the past decade, image moments have been exploited in several visual servoing schemes for their ability to represent object regions, objects defined by contours or a set of discrete points. Moments have also been useful to achieve control decoupling properties and to choose a minimal number of features to control the whole degrees of freedom (DOFs) of a camera. However, the choice of moment-based features to control the rotational motions around the x-axis and y-axis simultaneously with the translational motions along the same axis remains a key issue. In this paper, we introduce new visual features computed from low-order “shifted moments invariant.” Importantly, they allow us 1) to define a unique combination of visual features to control the whole six DOFs of an eye-in-hand camera independently from the object shape and (2) to significantly enlarge the convergence domain of the closed-loop system. Omar Tahri, Aurelien Yeremou Tamtsia, Youcef Mezouar, Cédric Demonceaux |
IEEE Trans. Robotics | 4 |
| 2014 | A Homography Formulation to the 3pt Plus a Common Direction Relative Pose Problem
Olivier Saurer, Pascal Vasseur, Cédric Demonceaux, Friedrich Fraundorfer |
ACCV (2) | 3 |
| 2014 | Efficient Pruning LMI Conditions for Branch-and-Prune Rank and Chirality-Constrained Estimation of the Dual Absolute QuadricabstractWe present a new globally optimal algorithm for self-calibrating a moving camera with constant parameters. Our method aims at estimating the Dual Absolute Quadric (DAQ) under the rank-3 and, optionally, camera centers chirality constraints. We employ the Branch-and-Prune paradigm and explore the space of only 5 parameters. Pruning in our method relies on solving Linear Matrix Inequality (LMI) feasibility and Generalized Eigenvalue (GEV) problems that solely depend upon the entries of the DAQ. These LMI and GEV problems are used to rule out branches in the search tree in which a quadric not satisfying the rank and chirality conditions on camera centers is guaranteed not to exist. The chirality LMI conditions are obtained by relying on the mild assumption that the camera undergoes a rotation of no more than 90 between consecutive views. Furthermore, our method does not rely on calculating bounds on any particular cost function and hence can virtually optimize any objective while achieving global optimality in a very competitive running-time. Adlane Habed, Danda Pani Paudel, Cédric Demonceaux, David Fofi |
CVPR | 3 |
| 2014 | Localization of 2D Cameras in a Known Environment Using Direct 2D-3D RegistrationabstractIn this paper we propose a robust and direct 2D-to-3D registration method for localizing 2D cameras in a known 3D environment. Although the 3D environment is known, localizing the cameras remains a challenging problem that is particularly undermined by the unknown 2D-3D correspondences, outliers, scale ambiguities and occlusions. Once the cameras are localized, the Structure-from-Motion reconstruction obtained from image correspondences is refined by means of a constrained nonlinear optimization that benefits from the knowledge of the scene. We also propose a common optimization framework for both localization and refinement steps in which projection errors in one view are minimized while preserving the existing relationships between images. The problem of occlusion and that of missing scene parts are handled by employing a scale histogram while the effect of data inaccuracies is minimized using an M-estimator-based technique. Danda Pani Paudel, Cédric Demonceaux, Adlane Habed, Pascal Vasseur |
ICPR | 2 |
| 2014 | 2D-3D camera fusion for visual odometry in outdoor environmentsabstractAccurate estimation of camera motion is very important for many robotics applications involving SfM and visual SLAM. Such accuracy is attempted by refining the estimated motion through nonlinear optimization. As many modern robots are equipped with both 2D and 3D cameras, it is both highly desirable and challenging to exploit data acquired from both modalities to achieve a better localization. Existing refinement methods, such as Bundle adjustment and loop closing, may be employed only when precise 2D-to-3D correspondences across frames are available. In this paper, we propose a framework for robot localization that benefits from both 2D and 3D information without requiring such accurate correspondences to be established. This is carried out through a 2D-3D based initial motion estimation followed by a constrained nonlinear optimization for motion refinement. The initial motion estimation finds the best possible 2D-to-3D correspondences and localizes the cameras with respect the 3D scene. The refinement step minimizes the projection errors of 3D points while preserving the existing relationships between images. The problems of occlusion and that of missing scene parts are handled by comparing the image-based reconstruction and 3D sensor measurements. The effect of data inaccuracies is minimized using an M-estimator based technique. Our experiments have demonstrated that the proposed framework allows to obtain a good initial motion estimate and a significant improvement through refinement. Danda Pani Paudel, Cédric Demonceaux, Adlane Habed, Pascal Vasseur, In-So Kweon |
IROS | 2 |
| 2014 | Extrinsic calibration of heterogeneous cameras by line images
Sang Ly, Cédric Demonceaux, Pascal Vasseur, Claude Pégard |
Mach. Vis. Appl. | 2 |
| 2013 | Gradient-based time to contact on paracatadioptric cameraabstractThe problem of time to contact or time to collision (TTC) estimation is largely discussed in perspective images. However, a few works have dealt with images of catadioptric sensors despite of their utility in robotics applications. The objective of this paper is to develop a novel model for estimating TTC with catadioptric images relative to a planar surface, and to demonstrate that TTC can be estimated only with derivative brightness and image coordinates. This model, called “gradient based time to contact”, does not need high processing such as explicit estimation of optical flow and feature detection and/or tracking. The proposed method allows to estimate TTC and gives additional information about the orientation of planar surface. It was tested on simulated and real datasets. Fatima Zahra Benamar, Sanaa El Fkihi, Cédric Demonceaux, El Mustapha Mouaddib, Driss Aboutajdine |
ICIP | 3 |
| 2013 | Scale invariant line matching on the sphereabstractThis paper proposes a novel approach of line matching across images captured by different types of cameras, from perspective to omnidirectional ones. Based on the spherical mapping, this method utilizes spherical SIFT point features to boost line matching and searches line correspondences using an affine invariant measure of similarity. It permits to unify the commonest cameras and to process heterogeneous images with the least distortion of visual information. Sang Ly, Cédric Demonceaux, Ralph Seulin, Yohan D. Fougerolle |
ICIP | 2 |
| 2013 | A Branch-and-Bound Approach to Correspondence and Grouping ProblemsabstractData correspondence/grouping under an unknown parametric model is a fundamental topic in computer vision. Finding feature correspondences between two images is probably the most popular application of this research field, and is the main motivation of our work. It is a key ingredient for a wide range of vision tasks, including three-dimensional reconstruction and object recognition. Existing feature correspondence methods are based on either local appearance similarity or global geometric consistency or a combination of both in some heuristic manner. None of these methods is fully satisfactory, especially in the presence of repetitive image textures or mismatches. In this paper, we present a new algorithm that combines the benefits of both appearance-based and geometry-based methods and mathematically guarantees a global optimization. Our algorithm accepts the two sets of features extracted from two images as input, and outputs the feature correspondences with the largest number of inliers, which verify both the appearance similarity and geometric constraints. Specifically, we formulate the problem as a mixed integer program and solve it efficiently by a series of linear programs via a branch-and-bound procedure. We subsequently generalize our framework in the context of data correspondence/grouping under an unknown parametric model and show it can be applied to certain classes of computer vision problems. Our algorithm has been validated successfully on synthesized data and challenging real images. Jean-Charles Bazin, Hongdong Li, In-So Kweon, Cédric Demonceaux, Pascal Vasseur, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Self-calibration of a PTZ Camera Using New LMI Constraints
François Rameau, Adlane Habed, Cédric Demonceaux, Desire Sidibé, David Fofi |
ACCV (4) | 3 |
| 2012 | Globally optimal line clustering and vanishing point estimation in Manhattan worldabstractThe projections of world parallel lines in an image intersect at a single point called the vanishing point (VP). VPs are a key ingredient for various vision tasks including rotation estimation and 3D reconstruction. Urban environments generally exhibit some dominant orthogonal VPs. Given a set of lines extracted from a calibrated image, this paper aims to (1) determine the line clustering, i.e. find which line belongs to which VP, and (2) estimate the associated orthogonal VPs. None of the existing methods is fully satisfactory because of the inherent difficulties of the problem, such as the local minima and the chicken-and-egg aspect. In this paper, we present a new algorithm that solves the problem in a mathematically guaranteed globally optimal manner and can inherently enforce the VP orthogonality. Specifically, we formulate the task as a consensus set maximization problem over the rotation search space, and further solve it efficiently by a branch-and-bound procedure based on the Interval Analysis theory. Our algorithm has been validated successfully on sets of challenging real images as well as synthetic data sets. Jean-Charles Bazin, Yongduek Seo, Cédric Demonceaux, Pascal Vasseur, Katsushi Ikeuchi, In-So Kweon, Marc Pollefeys |
CVPR | 3 |
| 2012 | Time to contact estimation on paracatadioptric cameras
Fatima Zahra Benamar, Cédric Demonceaux, Sanaa El Fkihi, El Mustapha Mouaddib, Driss Aboutajdine |
ICPR | 2 |
| 2012 | A geometrical approach For vision based attitude and altitude estimation for UAVs in dark environmentsabstractThis paper presents a single camera and laser system dedicated to the realtime estimation of attitude and altitude for unmanned aerial vehicles (UAV) under low illumination conditions to dark environments. The fisheye camera allows to cover a large field of view (FOV). The approach, close to structured light systems, uses the geometrical information obtained by the projection of a laser circle onto the ground plane and perceived by the camera. We propose some experiments based on simulated data and real sequences. The results show good agreement with the ground truth values from the commercial sensors in terms of its accuracy and correctness. The results also prove its suitability for autonomous take-off and landing as well as for the case of low altitude manoeuvre in dark, GPS signal deficient unknown environments with no prebuilt map. It also provides room for additional payload to be used for different applications due to it being inexpensive and use of light weight micro-camera and laser system. Ashutosh Natraj, Peter F. Sturm, Cédric Demonceaux, Pascal Vasseur |
IROS | 3 |
| 2011 | Vision based attitude and altitude estimation for UAVs in dark environmentsabstractThis paper presents a system dedicated to the real-time estimation of attitude and altitude for unmanned aerial vehicles (UAV) under low light and dark environment. This system consists in a fisheye camera, which allows to cover a large field of view (FOV), and a laser circle projector mounted on a fixed baseline. The approach, close to structured light systems, uses the geometrical information obtained by the projection of the laser circle onto the ground plane and perceived by the camera. We present a theoretical study of the system in which the camera is modelled as a sphere and show that the estimation of a conic on this sphere allows to obtain the attitude and the altitude of the robot. We propose some experiments based on simulated data and real sequences. The estimated attitude and altitude from our method are comparable with commercial sensors in terms of its accuracy and correctness. The results also prove its suitability for autonomous take-off and landing as well as for the case of low altitude manoeuvre in dark environments. It also provides room for additional payload to be used for different applications due to use of light weight micro-camera and laser system. Ashutosh Natraj, Cédric Demonceaux, Pascal Vasseur, Peter F. Sturm |
IROS | 2 |
| 2011 | Optical flow estimation from multichannel spherical image decomposition
Amina Radgui, Cédric Demonceaux, El Mustapha Mouaddib, Mohammed Rziza, Driss Aboutajdine |
Comput. Vis. Image Underst. | 2 |
| 2011 | Central catadioptric image processing with geodesic metric
Cédric Demonceaux, Pascal Vasseur, Yohan D. Fougerolle |
Image Vis. Comput. | 1 |
| 2010 | Translation estimation for single viewpoint cameras using linesabstractWe present a translation estimation method for single viewpoint (SVP) cameras using line features. Images captured by multiple central cameras such as perspective, central catadioptric and fisheye cameras are mapped to spherical images using the unified projection model. It is possible to recover the camera rotations using vanishing points of parallel line sets. We then estimate the translations from known rotations and line images on the spheres. The algorithm has been validated on simulated data and real images. This vision-based estimation approach can be applied in navigation of autonomous robots besides the conventional devices such as Global Positioning System (GPS) and Inertial Navigation System (INS). It helps vision-based localization of a single robot or recovery of relative positions among multiple robots equipped with different types of cameras. Sang Ly, Cédric Demonceaux, Pascal Vasseur |
ICRA | 2 |
| 2010 | Central catadioptric line matching for robotic applicationsabstractThis paper presents a method for catadioptric line matching across multiple images. While most of previous works deals with vertical lines and planar motion, our approach is able to match any kind of lines between two views separated by a rigid transformation without any prior knowledge of the epipolar geometry. Catadioptric lines are represented by their normals in sphere space and we use only these normals and their relative positions in order to perform the matching. A geometric hashing approach allows in the first image to construct hashing tables based on bases defined by every possible couples of normals. In the second image, a voting scheme permits to select the best corresponding bases and subsequently to match catadioptric lines.We show that the proposed representation is invariant in the case of a pure rotation and quasi-invariant for a combination of rotation and translation. We also propose different experimental results obtained in real time on real outdoor sequences. Pascal Vasseur, Cédric Demonceaux |
ICRA | 2 |
| 2010 | An original approach for automatic plane extraction by omnidirectional visionabstractWhereas some methods for plane extraction have been proposed, this problem still remains an open issue due to the complexity of the task. This paper especially focuses on the extraction of points lying on a plane (such as the ground and buildings walls) in sequences acquired by a central omnidirectional camera. Our approach is based on the epipolar constraint for planar scenes (i.e. homography) on a pair of omnidirectional images to detect some interest points belonging to a plane. Our main contribution is the introduction of a new method, called “2-point algorithm for homography”, that imposes some constraints on the homography using vanishing point (VP) information. Compared to the widely used DLT (4-point) algorithm, experiments on real data demonstrated that the proposed “2-point algorithm for homography” is more robust to noise and false matching, even when the plane to extract is not dominant in the image. Finally, we show that our system provides key clues for ground segmentation by GrabCut. Jean-Charles Bazin, Pierre-Yves Laffont, In-So Kweon, Cédric Demonceaux, Pascal Vasseur |
IROS | 4 |
| 2010 | UAV altitude estimation by mixed stereoscopic visionabstractAltitude is one of the most important parameters to be known for an Unmanned Aerial Vehicle (UAV) especially during critical maneuvers such as landing or steady flight. In this paper, we present mixed stereoscopic vision system made of a fish-eye camera and a perspective camera for altitude estimation. Contrary to classical stereoscopic systems based on feature matching, we propose a plane sweeping approach in order to estimate the altitude and consequently to detect the ground plane. Since there exists a homography between the two views and the sensor being calibrated and the attitude estimated by the fish-eye camera, the algorithm consists then in searching the altitude which verifies this homography. We show that this approach is robust and accurate, and a CPU implementation allows a real time estimation. Experimental results on real sequences of a small UAV demonstrate the effectiveness of the approach. Damien Eynard, Pascal Vasseur, Cédric Demonceaux, Vincent Frémont |
IROS | 3 |
| 2010 | Motion estimation by decoupling rotation and translation in catadioptric vision
Jean-Charles Bazin, Cédric Demonceaux, Pascal Vasseur, In-So Kweon |
Comput. Vis. Image Underst. | 2 |
| 2009 | Particle Filter Approach Adapted to Catadioptric Images for Target Tracking ApplicationabstractInternational audience Jean-Charles Bazin, Kuk-Jin Yoon, In-So Kweon, Cédric Demonceaux, Pascal Vasseur |
BMVC | 4 |
| 2009 | Omnidirectional image processing using geodesic metricabstractDue to distorsions of catadioptric sensors, omnidirectional images can not be treated as classical images. If the equivalence between central catadioptric images and spherical images is now well known and used, spherical analysis often leads to complex methods particularly tricky to employ. In this paper, we propose to derive omnidirectional image treatments by using geodesic metric. We demonstrate that this approach allows to adapt efficiently classical image processing to omnidirectional images. Cédric Demonceaux, Pascal Vasseur |
ICIP | 1 |
| 2009 | Dynamic programming and skyline extraction in catadioptric infrared imagesabstractUnmanned Aerial Vehicles (UAV) are the subject of an increasing interest in many applications and a key requirement for autonomous navigation is the attitude/position stabilization of the vehicle. Some previous works have suggested using catadioptric vision, instead of traditional perspective cameras, in order to gather much more information from the environment and therefore improve the robustness of the UAV attitude/position estimation. This paper belongs to a series of recent publications of our research group concerning catadioptric vision for UAVs. Currently, we focus on the extraction of skyline in catadioptric images since it provides important information about the attitude/position of the UAV. For example, the DEM-based methods can match the extracted skyline with a Digital Elevation Map (DEM) by process of registration, which permits to estimate the attitude and the position of the camera. Like any standard cameras, catadioptric systems cannot work in low luminosity situations because they are based on visible light. To overcome this important limitation, in this paper, we propose using a catadioptric infrared camera and extending one of our methods of skyline detection towards catadioptric infrared images. The task of extracting the best skyline in images is usually converted in an energy minimization problem that can be solved by dynamic programming. The major contribution of this paper is the extension of dynamic programming for catadioptric images using an adapted neighborhood and an appropriate scanning direction. Finally, we present some experimental results to demonstrate the validity of our approach. Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur |
ICRA | 3 |
| 2008 | Improvement of feature matching in catadioptric images using gyroscope dataabstractMost of vision-based algorithms for motion and localization estimation requires matching some interest points in a pair of images. After building feature correspondence, it is possible to estimate camera motion/localization using epipolar geometry. However feature matching is still a challenging problem because of time constraint or image variability for example. In several robotic applications, the camera rotation may be known thanks to a gyroscope or another orientation sensor. Therefore, in this paper, we aim to answer the following question: can the knowledge of rotation from a gyroscope be used to improve feature matching. To analyze this new approach of camera and gyroscope data fusion, we proceed in two steps. First, we rotationally align the images using rotation information of the gyroscope. And second, we compare the quality of feature matching in the original and rotationally aligned images. Experimental results on a real catadioptric sequence show that gyroscope data permits to sensibly improve the number of inliers according to epipolar geometry. Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur |
ICPR | 3 |
| 2008 | UAV Attitude estimation by vanishing points in catadioptric imagesabstractUnmanned aerial vehicles (UAV) are the subject of an increasing interest in many applications and a key requirement is the stabilization of the vehicle. Some previous works have suggested using catadioptric vision, instead of traditional perspective cameras, in order to gather much more information from the environment and therefore improve the robustness of the UAV attitude estimation. This paper belongs to a series of recent publications of our research group concerning catadioptric vision for UAVs. Currently, we focus on the estimation of the complete attitude of a UAV flying in urban environment. In order to avoid the limitations of horizon-based approaches, the difficulties of traditional epipolar methods (such as rotation-translation ambiguity, lack of features, retrieving motion parameters from matrix decomposition, etc..) and improve UAV dynamic control, we suggest computing infinite homography. We show how catadioptric vision plays a key role to: first, extract a large number of lines, second robustly estimate the associated vanishing points and third, track them even during long video sequences. Therefore it is not only possible to estimate the relative rotation between consecutive frames but also compute the absolute rotation between two distant frames without error accumulation. Finally, we present some experimental results with ground truth data to demonstrate the accuracy and the robustness of our method. Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur |
ICRA | 3 |
| 2008 | A robust top-down approach for rotation estimation and vanishing points extraction by catadioptric vision in urban environmentabstractA key requirement for unmanned aerial vehicles (UAV) applications is the attitude stabilization of the aircraft, which requires the knowledge of its orientation. It is now well established that traditional navigation equipments, like GPS or INS, suffer from several disadvantages. That is why some works have suggested a vision-based approach of the problem. Especially, catadioptric vision is more and more used since it permits to gather much more information from the environment, compared to traditional perspective cameras, and therefore the robustness of the UAV attitude estimation is improved. Rotation estimation from conventional and catadioptric images has been extensively studied. Whereas interesting results can be obtained, the existing methods have non-negligible limitations such as difficult features matching (e.g. repeated texture, blurring or illumination changing) or a high computational cost (e.g. vanishing point extraction or analyze in frequency domain). In order to overcome these limitations, this paper presents a top-down approach for estimating the rotation and extracting the vanishing points in catadioptric images. This new framework is accurate and can run in real-time. To obtain the ground truth data, we also calibrate our catadioptric camera with a gyroscope. Finally, experimental results on a real video sequence are presented and compared to the ground truth data obtained by the gyroscope. Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur |
IROS | 3 |
| 2008 | Automatic calibration of catadioptric cameras in urban environmentabstractCamera calibration is an important step for vision-based stabilization of unmanned aerial vehicles (UAV). The goal of this paper is to develop a method for automatic calibration of a catadioptric camera so that it can be easily run before mounting the camera on the UAV or even during the flight to deal with vibrations or shocks. Whereas existing works can provide interesting results, they suffer from several practical limitations (manual line extraction, inaccurate conic fitting, calibration pattern, camera motion, execution time, etc...) and therefore cannot be applied in our application. The proposed algorithm aims to determine the most probable calibration that verifies some geometric constraints induced by catadioptric projection. In order to efficiently maximize this probability, we use a particle filtering approach. Experimental results have demonstrated the effectiveness of the proposed method. Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur |
IROS | 3 |
| 2007 | Rectangle Extraction in Catadioptric ImagesabstractNowadays, robotic systems are more and more equipped with catadioptric cameras. However several problems associated to catadioptric vision have been studied only slightly. Especially algorithms for detecting rectangles in catadioptric images have not yet been developed whereas it is required in diverse applications such as building extraction in aerial images. We show that working in the equivalent sphere provides an appropriate framework to detect lines, parallelism, orthogonality and therefore rectangles. Finally, we present experimental results on synthesized and real data. Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur |
ICCV | 3 |
| 2007 | UAV Attitude Computation by Omnidirectional Vision in Urban EnvironmentabstractAttitude is one of the most important parameters for a UAV during a flight. Attitude computation methods based vision generally use the horizon line as reference. However, the horizon line becomes an inadequate feature in urban environment. We then propose in this paper an omnidirectional vision system based on straight lines (very frequent in urban environment) that is able to compute the roll and pitch angles. The method consists in finding bundles of horizontal and vertical parallel lines in order to obtain an absolute reference for the attitude computation. We also develop here a new and efficient method for line extraction and bundle of parallel line detection. An original method of horizontal and vertical plane detection is also provided. We show experimental results on different images extracted from video sequences. Cédric Demonceaux, Pascal Vasseur, Claude Pégard |
ICRA | 1 |
| 2007 | Proposition and Comparison of Catadioptric Homography Estimation Methods
Christophe Simler, Cédric Demonceaux, Pascal Vasseur |
PSIVT | 2 |
| 2006 | Omnidirectional Vision on UAV for Attitude ComputationabstractUnmanned aerial vehicles (UAVs) are the subject of an increasing interest in many applications. Autonomy is one of the major advantages of these vehicles. It is then necessary to develop particular sensors in order to provide efficient navigation functions. In this paper, we propose a method for attitude computation catadioptric images. We first demonstrate the advantages of the catadioptric vision sensor for this application. In fact, the geometric properties of the sensor permit to compute easily the roll and pitch angles. The method consists in separating the sky from the earth in order to detect the horizon. We propose an adaptation of the Markov random fields for catadioptric images for this segmentation. The second step consists in estimating the parameters of the horizon line thanks to a robust estimation algorithm. We also present the angle estimation algorithm and finally, we show experimental results on synthetic and real images captured from an airplane Cédric Demonceaux, Pascal Vasseur, Claude Pégard |
ICRA | 1 |
| 2006 | Robust Attitude Estimation with Catadioptric VisionabstractAttitude (roll and pitch) is an essential data for the navigation of a UAV. Rather than using inertial sensors, we propose a catadioptric vision system allowing a fast, robust and accurate estimation of these angles. We show that the optimization of a sky/ground partitioning criterion associated with the specific geometric characteristics of the catadioptric sensor provides very interesting results. Experimental results obtained on real sequences are presented and compared with inertial sensor measures Cédric Demonceaux, Pascal Vasseur, Claude Pégard |
IROS | 1 |
| 2006 | Markov random fields for catadioptric image processing
Cédric Demonceaux, Pascal Vasseur |
Pattern Recognit. Lett. | 1 |
| 2004 | Fast motion estimation and motion segmentation using multi-scale approachabstractThe paper investigates a fast method for motion estimation and motion segmentation. We choose to decompose the motion on basis functions. That allows us to compute the velocity of each pixel by only solving a linear system. The motion segmentation is carried out using a Markovian formulation. We minimize the associated energy by a determinist algorithm. However, these choices induce two drawbacks: a temporal aliasing problem; a sensitivity of the optimization algorithm with the initialization. To overcome these two problems, we use a multiresolution method that successively computes the motion and the segmentation at each scale by motion compensation. Thanks to this procedure, we obtain a fast method of motion estimation and segmentation which does not require either initial spacial segmentation or dominant motion in the sequence. Cédric Demonceaux, Djemaa Kachi |
ICIP | 1 |