Raoul de Charette

dblp:30/7749 · DBLP profile ↗
← Back
36ranked-venue papers
2as first author
23since 2021 · last 2026
0000-0003-3738-1962ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 23 · 1 first-author · 17 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 CLIP's Visual Embedding Projector is a Few-shot Cornucopia
abstract
We introduce ProLIP, a simple and architecture-agnostic method for adapting contrastively pretrained vision-language models, such as CLIP [36], to few-shot classification. ProLIP fine-tunes the vision encoder’s projection matrix with Frobenius norm regularization on its deviation from the pretrained weights. It achieves state-of-the-art performance on 11 few-shot classification benchmarks under both "few-shot validation" [23] and "validation-free" [42] settings. Moreover, by rethinking the non-linear CLIP-Adapter [13] through ProLIP’s lens, we design a Regularized Linear Adapter (RLA) that performs better, requires no hyperparameter tuning, is less sensitive to learning rate values, and offers an alternative to ProLIP in black-box scenarios where model weights are inaccessible. Beyond few-shot classification, ProLIP excels in cross-dataset transfer, domain generalization, base-to-new class generalization, and test-time adaptation—where it outperforms prompt tuning while being an order of magnitude faster to train. Code is available at https://github.com/astra-vision/ProLIP.
Mohammad Fahes, Andrei Bursuc, Patrick Pérez, Raoul de Charette
WACV5
2026 Domain Adaptation with a Single Vision-Language Embedding
Mohammad Fahes, Andrei Bursuc, Patrick Pérez, Raoul de Charette
Int. J. Comput. Vis.5
2025 FLOSS: Free Lunch in Open-Vocabulary Semantic Segmentation
abstract
In this paper, we challenge the conventional practice in Open-Vocabulary Semantic Segmentation (OVSS) of using averaged class-wise text embeddings, which are typically obtained by encoding each class name with multiple templates (e.g., a photo of , a sketch of a ). We investigate the impact of templates for OVSS, and find that for each class, there exist single-template classifiers--which we refer to as class-experts--that significantly outperform the conventional averaged classifier. First, to identify these class-experts, we introduce a novel approach that estimates them without any labeled data or training. By leveraging the class-wise prediction entropy of single-template classifiers, we select those yielding the lowest entropy as the most reliable class-experts. Second, we combine the outputs of class-experts in a new fusion process. Our plug-and-play method, coined FLOSS, is orthogonal and complementary to existing OVSS methods, offering an improvement without the need for additional labels or training. Extensive experiments show that FLOSS consistently enhances state-of-the-art OVSS models, generalizes well across datasets with different distribution shifts, and delivers substantial improvements in low-data scenarios where only a few unlabeled images are available. Our code is available at https://github.com/yasserben/FLOSS .
Yasser Benigmim, Mohammad Fahes, Andrei Bursuc, Raoul de Charette
ICCV5
2025 LiDPM: Rethinking Point Diffusion for Lidar Scene Completion
abstract
Training diffusion models that work directly on lidar points at the scale of outdoor scenes is challenging due to the difficulty of generating fine-grained details from white noise over a broad field of view. The latest works addressing scene completion with diffusion models tackle this problem by reformulating the original DDPM as a local diffusion process. It contrasts with the common practice of operating at the level of objects, where vanilla DDPMs are currently used. In this work, we close the gap between these two lines of work. We identify approximations in the local diffusion formulation, show that they are not required to operate at the scene level, and that a vanilla DDPM with a well-chosen starting point is enough for completion. Finally, we demonstrate that our method, LiDPM, leads to better results in scene completion on SemanticKITTI. The project page is https://astra-vision.github.io/LiDPM.
Tetiana Martyniuk, Gilles Puy, Alexandre Boulch, Renaud Marlet, Raoul de Charette
IV5
2025 LATTECLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
abstract
Large-scale vision-language pre-trained (VLP) models (e.g., CLIP [46]) are renowned for their versatility, as they can be applied to diverse applications in a zero-shot setup. However, when these models are used in specific domains, their performance often falls short due to domain gaps or the under-representation of these domains in the training data. While fine-tuning VLP models on custom datasets with human-annotated labels can address this issue, annotating even a small-scale dataset (e.g., 100k samples) can be an expensive endeavor, often requiring expert annotators if the task is complex. To address these challenges, we propose LATTECLIP, an unsupervised method for fine-tuning CLIP models on classification with known class names in custom domains, without relying on human annotations. Our method leverages Large Multimodal Models (LMMs) to generate expressive textual descriptions for both individual images and groups of images. These provide additional contextual information to guide the fine-tuning process in the custom domains. Since LMM-generated descriptions are prone to hallucination or missing details, we introduce a novel strategy to distill only the useful information and stabilise the training. Specifically, we learn rich per-class prototype representations from noisy generated texts and dual pseudo-labels. Our experiments on 10 domain-specific datasets show that LATTECLIP outperforms pre-trained zero-shot methods by an average improvement of +4.74 points in top-1 accuracy and other state-of-the-art unsupervised methods by +3.45 points.
Anh-Quan Cao, Maximilian Jaritz, Matthieu Guillaumin, Raoul de Charette, Loris Bazzani
WACV4
2025 MatSwap: Light-aware material transfers in images
abstract
Abstract We present MatSwap, a method to transfer materials to designated surfaces in an image realistically. Such a task is non‐trivial due to the large entanglement of material appearance, geometry, and lighting in a photograph. In the literature, material editing methods typically rely on either cumbersome text engineering or extensive manual annotations requiring artist knowledge and 3D scene properties that are impractical to obtain. In contrast, we propose to directly learn the relationship between the input material—as observed on a flat surface—and its appearance within the scene, without the need for explicit UV mapping. To achieve this, we rely on a custom light‐ and geometry‐aware diffusion model. We fine‐tune a large‐scale pre‐trained text‐to‐image model for material transfer using our synthetic dataset, preserving its strong priors to ensure effective generalization to real images. As a result, our method seamlessly integrates a desired material into the target location in the photograph while retaining the identity of the scene. MatSwap is evaluated on synthetic and real images showing that it compares favorably to recent works. Our code and data are made publicly available on https://github.com/astra‐vision/MatSwap
Ivan Lopes, Valentin Deschaintre, Yannick Hold-Geoffroy, Raoul de Charette
Comput. Graph. Forum4
2025 Material transforms from disentangled NeRF representations
abstract
Abstract In this paper, we first propose a novel method for transferring material transformations across different scenes. Building on disentangled Neural Radiance Field (NeRF) representations, our approach learns to map Bidirectional Reflectance Distribution Functions (BRDF) from pairs of scenes observed in varying conditions, such as dry and wet. The learned transformations can then be applied to unseen scenes with similar materials, therefore effectively rendering the transformation learned with an arbitrary level of intensity. Extensive experiments on synthetic scenes and real‐world objects validate the effectiveness of our approach, showing that it can learn various transformations such as wetness, painting, coating, etc. Our results highlight not only the versatility of our method but also its potential for practical applications in computer graphics. We publish our method implementation, along with our synthetic/real datasets on https://github.com/astra‐vision/BRDFTransform
Ivan Lopes, Jean-François Lalonde, Raoul de Charette
Comput. Graph. Forum3
2024 PaSCo: Urban 3D Panoptic Scene Completion with Uncertainty Awareness
abstract
We propose the task of Panoptic Scene Completion (PSC) which extends the recently popular Semantic Scene Completion (SSC) task with instance-level information to produce a richer understanding of the 3D scene. Our PSC proposal utilizes a hybrid mask-based technique on the non-empty voxels from sparse multi-scale completions. Whereas the SSC literature overlooks uncertainty which is critical for robotics applications, we instead propose an efficient ensembling to estimate both voxel-wise and instance-wise uncertainties along PSC. This is achieved by building on a multi-input multi-output (MIMO) strategy, while improving performance and yielding better uncertainty for little additional compute. Additionally, we introduce a technique to aggregate permutation-invariant mask predictions. Our experiments demonstrate that our method surpasses all baselines in both Panoptic Scene Completion and uncertainty estimation on three large-scale autonomous driving datasets. Our code and data are available at https://astra-vision.github.ioIPaSCo.
Anh-Quan Cao, Angela Dai, Raoul de Charette
CVPR3
2024 A Simple Recipe for Language-Guided Domain Generalized Segmentation
abstract
Generalization to new domains not seen during training is one of the longstanding challenges in deploying neural networks in real-world applications. Existing generalization techniques either necessitate external images for augmentation, and/or aim at learning invariant representations by imposing various alignment constraints. Largescale pretraining has recently shown promising generalization capabilities, along with the potential of binding different modalities. For instance, the advent of vision-language models like CLIP has opened the doorway for vision models to exploit the textual modality. In this paper, we introduce a simple framework for generalizing semantic segmentation networks by employing language as the source of randomization. Our recipe comprises three key ingredients: (i) the preservation of the intrinsic CLIP robustness through mini-mal fine-tuning, (ii) language-driven local style augmentation, and (iii) randomization by locally mixing the source and augmented styles during training. Extensive experiments report state-of-the-art results on various generalization benchmarks. Code is accessible on the project page11https://astra-vision.github.io/FAMix.
Mohammad Fahes, Andrei Bursuc, Patrick Pérez, Raoul de Charette
CVPR5
2024 Material Palette: Extraction of Materials from a Single Image
abstract
Physically-Based Rendering (PBR) is key to modeling the interaction between light and materials, and finds extensive applications across computer graphics domains. However, acquiring PBR materials is costly and requires special apparatus. In this paper, we propose a method to extract PBR materials from a single real-world image. We do so in two steps: first, we map regions of the image to material concept tokens using a diffusion model, allowing the sampling of texture images resembling each material in the scene. Second, we leverage a separate network to decom-pose the generated textures into spatially varying BRDFs (SVBRDFs), offering us readily usable materials for rendering applications. Our approach relies on existing synthetic material libraries with SVBRDF ground truth. It exploits a diffusion-generated RGB texture dataset to allow generalization to new samples using unsupervised do-main adaptation (UDA). Our contributions are thoroughly evaluated on synthetic and real-world datasets. We further demonstrate the applicability of our method for editing 3D scenes with materials estimated from real photographs. Along with video, we share code and models as open-source on the project page: https://github.com/astra-vision/MaterialPalette.
Ivan Lopes, Fabio Pizzati, Raoul de Charette
CVPR3
2024 UMBRAE: Unified Multimodal Brain Decoding
Weihao Xia 0001, Raoul de Charette, A. Cengiz Öztireli, Jing-Hao Xue
ECCV (7)2
2024 DREAM: Visual Decoding from REversing HumAn Visual SysteM
abstract
In this work we present DREAM, an fMRI-to-image method for reconstructing viewed images from brain activities, grounded on fundamental knowledge of the human visual system. We craft reverse pathways that emulate the hierarchical and parallel nature of how humans perceive the visual world. These tailored pathways are specialized to decipher semantics, color, and depth cues from fMRI data, mirroring the forward pathways from visual stimuli to fMRI recordings. To do so, two components mimic the inverse processes within the human visual system: the Reverse Visual Association Cortex (R-VAC) which reverses pathways of this brain region, extracting semantics from fMRI data; the Reverse Parallel PKM (R-PKM) component simultaneously predicting color and depth from fMRI signals. The experiments indicate that our method outperforms the current state-of-the-art models in terms of the consistency of appearance, structure, and semantics. Code will be available at https://github.com/weihaox/DREAM.
Weihao Xia 0001, Raoul de Charette, A. Cengiz Öztireli, Jing-Hao Xue
WACV2
2023 SceneRF: Self-Supervised Monocular 3D Scene Reconstruction with Radiance Fields
abstract
3D reconstruction from a single 2D image was extensively covered in the literature but relies on depth supervision at training time, which limits its applicability. To relax the dependence to depth we propose SceneRF, a self-supervised monocular scene reconstruction method using only posed image sequences for training. Fueled by the recent progress in neural radiance fields (NeRF) we optimize a radiance field though with explicit depth optimization and a novel probabilistic sampling strategy to efficiently handle large scenes. At inference, a single input image suffices to hallucinate novel depth views which are fused together to obtain 3D scene reconstruction. Thorough experiments demonstrate that we outperform all baselines for novel depth views synthesis and scene reconstruction, on indoor BundleFusion and outdoor SemanticKITTI. Code is available at https://astra-vision.github.io/SceneRF.
Anh-Quan Cao, Raoul de Charette
ICCV2
2023 PØDA: Prompt-driven Zero-shot Domain Adaptation
abstract
Domain adaptation has been vastly investigated in computer vision but still requires access to target images at train time, which might be intractable in some uncommon conditions. In this paper, we propose the task of ‘Prompt-driven Zero-shot Domain Adaptation’, where we adapt a model trained on a source domain using only a general description in natural language of the target domain, i.e., a prompt. First, we leverage a pretrained contrastive vision-language model (CLIP) to optimize affine transformations of source features, steering them towards the target text embedding while preserving their content and semantics. To achieve this, we propose Prompt-driven Instance Normalization (PIN). Second, we show that these prompt-driven augmentations can be used to perform zero-shot domain adaptation for semantic segmentation. Experiments demonstrate that our method significantly outperforms CLIP-based style transfer baselines on several datasets for the downstream task at hand, even surpassing one-shot unsupervised domain adaptation. A similar boost is observed on object detection and image classification. The code is available at https://github.com/astra-vision/PODA.
Mohammad Fahes, Andrei Bursuc, Patrick Pérez, Raoul de Charette
ICCV5
2023 Cross-task Attention Mechanism for Dense Multi-task Learning
abstract
Multi-task learning has recently become a promising solution for comprehensive understanding of complex scenes. With an appropriate design, multi-task models can not only be memory-efficient but also favour the exchange of complementary signals across tasks. In this work, we jointly address 2D semantic segmentation, and two geometry-related tasks, namely dense depth, surface normal estimation as well as edge estimation showing their benefit on several datasets. We propose a novel multi-task learning architecture that exploits pair-wise cross-task exchange through correlation-guided attention and self-attention to enhance the average representation learning for all tasks. We conduct extensive experiments on three multi-task setups, showing the benefit of our proposal in comparison to competitive baselines in both synthetic and real benchmarks. We also extend our method to the novel multi-task unsupervised domain adaptation setting. Our code is available at https://github.com/cv-rits/DenseMTL
Ivan Lopes, Raoul de Charette
WACV3
2023 Cross-Modal Learning for Domain Adaptation in 3D Semantic Segmentation
abstract
Domain adaptation is an important task to enable learning when labels are scarce. While most works focus only on the image modality, there are many important multi-modal datasets. In order to leverage multi-modality for domain adaptation, we propose cross-modal learning, where we enforce consistency between the predictions of two modalities via mutual mimicking. We constrain our network to make correct predictions on labeled data and consistent predictions across modalities on unlabeled target-domain data. Experiments in unsupervised and semi-supervised domain adaptation settings prove the effectiveness of this novel domain adaptation strategy. Specifically, we evaluate on the task of 3D semantic segmentation from either the 2D image, the 3D point cloud or from both. We leverage recent driving datasets to produce a wide variety of domain adaptation scenarios including changes in scene layout, lighting, sensor setup and weather, as well as the synthetic-to-real setup. Our method significantly improves over previous uni-modal adaptation baselines on all adaption scenarios. Code will be made available upon publication.
Maximilian Jaritz, Raoul de Charette, Émilie Wirbel, Patrick Pérez
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Physics-Informed Guided Disentanglement in Generative Networks
abstract
Image-to-image translation (i2i) networks suffer from entanglement effects in presence of physics-related phenomena in target domain (such as occlusions, fog, etc), lowering altogether the translation quality, controllability and variability. In this paper, we propose a general framework to disentangle visual traits in target images. Primarily, we build upon collection of simple physics models, guiding the disentanglement with a physical model that renders some of the target traits, and learning the remaining ones. Because physics allows explicit and interpretable outputs, our physical models (optimally regressed on target) allows generating unseen scenarios in a controllable manner. Secondarily, we show the versatility of our framework to neural-guided disentanglement where a generative network is used in place of a physical model in case the latter is not directly accessible. Altogether, we introduce three strategies of disentanglement being guided from either a fully differentiable physics model, a (partially) non-differentiable physics model, or a neural network. The results show our disentanglement strategies dramatically increase performances qualitatively and quantitatively in several challenging scenarios for image translation.
Fabio Pizzati, Pietro Cerri, Raoul de Charette
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation
Anh-Quan Cao, Raoul de Charette
BMVC3
2022 MonoScene: Monocular 3D Semantic Scene Completion
abstract
MonoScene proposes a 3D Semantic Scene Completion (SSC) framework, where the dense geometry and semantics of a scene are inferred from a single monocular RGB image. Different from the SSC literature, relying on 2.5 or 3D input, we solve the complex problem of 2D to 3D scene reconstruction while jointly inferring its semantics. Our framework relies on successive 2D and 3D UNets, bridged by a novel 2D-3D features projection inspired by optics, and introduces a 3D context relation prior to enforce spatio-semantic consistency. Along with architectural contributions, we introduce novel global scene and local frustums losses. Experiments show we outperform the literature on all metries and datasets while hallucinating plausible scenery even beyond the camera field of view. Our code and trained models are available at https://github.com/cv-rits/MonoScene.
Anh-Quan Cao, Raoul de Charette
CVPR2
2022 ManiFest: Manifold Deformation for Few-Shot Image Translation
Fabio Pizzati, Jean-François Lalonde, Raoul de Charette
ECCV (17)3
2022 3D Semantic Scene Completion: A Survey
Luis Roldão, Raoul de Charette, Anne Verroust-Blondet
Int. J. Comput. Vis.2
2021 CoMoGAN: Continuous Model-Guided Image-to-Image Translation
abstract
CoMoGAN is a continuous GAN relying on the unsupervised reorganization of the target data on a functional manifold. To that matter, we introduce a new Functional Instance Normalization layer and residual mechanism, which together disentangle image content from position on target manifold. We rely on naive physics-inspired models to guide the training while allowing private model/translations features. CoMoGAN can be used with any GAN backbone and allows new types of image translation, such as cyclic image translation like timelapse generation, or detached linear translation. On all datasets, it outperforms the literature. Our code is available in this page: https://github.com/cv-rits/CoMoGAN.
Fabio Pizzati, Pietro Cerri, Raoul de Charette
CVPR3
2021 Rain Rendering for Evaluating and Improving Robustness to Bad Weather
Maxime Tremblay, Shirsendu Sukanta Halder, Raoul de Charette, Jean-François Lalonde
Int. J. Comput. Vis.3
2020 LMSCNet: Lightweight Multiscale 3D Semantic Completion
abstract
We introduce a new approach for multiscale 3Dsemantic scene completion from voxelized sparse 3D LiDAR scans. As opposed to the literature, we use a 2D UNet backbone with comprehensive multiscale skip connections to enhance feature flow, along with 3D segmentation heads. On the SemanticKITTI benchmark, our method performs on par on semantic completion and better on occupancy completion than all other published methods - while being significantly lighter and faster. As suchit provides a great performance/speed trade-off for mobile-robotics applications. The ablation studies demonstrate our method is robust to lower density inputs, and that it enables very high speed semantic completion at the coarsest level. Our code is available at https://github.com/cv-rits/LMSCNet.
Luis Roldão, Raoul de Charette, Anne Verroust-Blondet
3DV2
2020 xMUDA: Cross-Modal Unsupervised Domain Adaptation for 3D Semantic Segmentation
abstract
Unsupervised Domain Adaptation (UDA) is crucial to tackle the lack of annotations in a new domain. There are many multi-modal datasets, but most UDA approaches are uni-modal. In this work, we explore how to learn from multi-modality and propose cross-modal UDA (xMUDA) where we assume the presence of 2D images and 3D point clouds for 3D semantic segmentation. This is challenging as the two input spaces are heterogeneous and can be impacted differently by domain shift. In xMUDA, modalities learn from each other through mutual mimicking, disentangled from the segmentation objective, to prevent the stronger modality from adopting false predictions from the weaker one. We evaluate on new UDA scenarios including day-to-night, country-to-country and dataset-to-dataset, leveraging recent autonomous driving datasets. xMUDA brings large improvements over uni-modal UDA on all tested scenarios, and is complementary to state-of-the-art UDA techniques. Code is available at https://github.com/valeoai/xmuda.
Maximilian Jaritz, Raoul de Charette, Émilie Wirbel, Patrick Pérez
CVPR3
2020 Model-Based Occlusion Disentanglement for Image-to-Image Translation
Fabio Pizzati, Pietro Cerri, Raoul de Charette
ECCV (20)3
2020 RGB-D-E: Event Camera Calibration for Fast 6-DOF object Tracking
abstract
Augmented reality devices require multiple sensors to perform various tasks such as localization and tracking. Currently, popular cameras are mostly frame-based (e.g. RGB and Depth) which impose a high data bandwidth and power usage. With the necessity for low power and more responsive augmented reality systems, using solely frame-based sensors imposes limits to the various algorithms that needs high frequency data from the environement. As such, event-based sensors have become increasingly popular due to their low power, bandwidth and latency, as well as their very high frequency data acquisition capabilities. In this paper, we propose, for the first time, to use an event-based camera to increase the speed of 3D object tracking in 6 degrees of freedom. This application requires handling very high object speed to convey compelling AR experiences. To this end, we propose a new system which combines a recent RGB-D sensor (Kinect Azure) with an event camera (DAVIS346). We develop a deep learning approach, which combines an existing RGB-D network along with a novel event-based network in a cascade fashion, and demonstrate that our approach significantly improves the robustness of a state-of-the-art frame-based 6-DOF object tracker using our RGB-D-E pipeline. Our code and our RGB-D-E evaluation dataset are available at https://github.com/lvsn/rgbde-tracking.
Etienne Dubeau, Mathieu Garon, Benoit Debaque, Raoul de Charette, Jean-François Lalonde
ISMAR4
2020 Domain Bridge for Unpaired Image-to-Image Translation and Unsupervised Domain Adaptation
abstract
Image-to-image translation architectures may have limited effectiveness in some circumstances. For example, while generating rainy scenarios, they may fail to model typical traits of rain as water drops, and this ultimately impacts the synthetic images realism. With our method, called domain bridge, web-crawled data are exploited to reduce the domain gap, leading to the inclusion of previously ignored elements in the generated images. We make use of a network for clear to rain translation trained with the domain bridge to extend our work to Unsupervised Domain Adaptation (UDA). In that context, we introduce an online multimodal style-sampling strategy, where image translation multimodality is exploited at training time to improve performances. Finally, a novel approach for self-supervised learning is presented, and used to further align the domains. With our contributions, we simultaneously increase the realism of the generated images, while reaching on par performances with respect to the UDA state-of-the-art, with a simpler approach.
Fabio Pizzati, Raoul de Charette, Michela Zaccaria, Pietro Cerri
WACV2
2019 Physics-Based Rendering for Improving Robustness to Rain
abstract
To improve the robustness to rain, we present a physically-based rain rendering pipeline for realistically inserting rain into clear weather images. Our rendering relies on a physical particle simulator, an estimation of the scene lighting and an accurate rain photometric modeling to augment images with arbitrary amount of realistic rain or fog. We validate our rendering with a user study, proving our rain is judged 40% more realistic that state-of-the-art. Using our generated weather augmented Kitti and Cityscapes dataset, we conduct a thorough evaluation of deep object detection and semantic segmentation algorithms and show that their performance decreases in degraded weather, on the order of 15% for object detection and 60% for semantic segmentation. Furthermore, we show refining existing networks with our augmented images improves the robustness of both object detection and semantic segmentation algorithms. We experiment on nuScenes and measure an improvement of 15% for object detection and 35% for semantic segmentation compared to original rainy performance. Augmented databases and code are available on the project page.
Shirsendu Sukanta Halder, Jean-François Lalonde, Raoul de Charette
ICCV3
2019 A Cooperative Car-Following/Emergency Braking System With Prediction-Based Pedestrian Avoidance Capabilities
abstract
Urban environments are among the most challenging scenarios for car-following systems, since pedestrians may interfere with the platoon unexpectedly. To address this problem, this paper proposes a cooperative system using vehicle-to-vehicle and vehicle-to-pedestrian communication links. A fractional-order control-based cooperative adaptive cruise control benefits of communication for tighter inter-vehicle distances, while pedestrian communication is fused with LiDAR sensing to allow the detection of occluded pedestrians. The prediction of the pedestrians' trajectories is used to perform a speed reduction or an emergency braking that interrupts the car-following yif necessary. Whenever a platoon decoupling occurs, a gap-closing maneuver is executed so that the ego-vehicle rejoins the platoon in a string stable way. The complete system was tested on experimental platforms at inria facilities, providing encouraging results and demonstrating the correct performance of the integrated systems.
Pierre Merdrignac, Raoul de Charette, Francisco M. Navas Matos, Vicente Milanés Montero, Fawzi Nashashibi
IEEE Trans. Intell. Transp. Syst.3
2018 Sparse and Dense Data with CNNs: Depth Completion and Semantic Segmentation
abstract
Convolutional neural networks are designed for dense data, but vision data is often sparse (stereo depth, point clouds, pen stroke, etc.). We present a method to handle sparse depth data with optional dense RGB, and accomplish depth completion and semantic segmentation changing only the last layer. Our proposal efficiently learns sparse features without the need of an additional validity mask. We show how to ensure network robustness to varying input sparsities. Our method even works with densities as low as 0.8% (8 layer lidar), and outperforms all published state-of-the-art on the Kitti depth completion benchmark.
Maximilian Jaritz, Raoul de Charette, Émilie Wirbel, Xavier Perrotton, Fawzi Nashashibi
3DV2
2018 End-to-End Race Driving with Deep Reinforcement Learning
abstract
We present research using the latest reinforcement learning algorithm for end-to-end driving without any mediated perception (object recognition, scene understanding). The newly proposed reward and learning strategies lead together to faster convergence and more robust driving using only RGB image from a forward facing camera. An Asynchronous Actor Critic (A3C) framework is used to learn the car control in a physically and graphically realistic rally game, with the agents evolving simultaneously on tracks with a variety of road structures (turns, hills), graphics (seasons, location) and physics (road adherence). A thorough evaluation is conducted and generalization is proven on unseen tracks and using legal speed limits. Open loop tests on real sequences of images show some domain adaption capability of our method.
Maximilian Jaritz, Raoul de Charette, Marin Toromanoff, Etienne Perot, Fawzi Nashashibi
ICRA2
2018 Wifi fingerprinting localization for intelligent vehicles in car park
abstract
In this paper, a novel method of WiFi fingerprinting for localizing intelligent vehicles in GPS-denied area, such as car parks, is proposed. Although the method itself is a popular approach for indoor localization application, adapting it to the speed of vehicles requires different treatment. By deploying an ensemble neural network for fingerprinting classification, the method shows a reasonable localization precision at car park speed. Furthermore, a Gaussian Mixture Model (GMM) Particle Filter is applied to increase localization frequency as well as accuracy. Experiments show promising results with average localization error of 0.6m.
Van-Dinh Nguyen, Raoul de Charette, Fawzi Nashashibi, Trung-Kien Dao, Eric Castelli
IPIN2
2012 Fast reactive control for illumination through rain and snow
abstract
During low-light conditions, drivers rely mainly on headlights to improve visibility. But in the presence of rain and snow, headlights can paradoxically reduce visibility due to light reflected off of precipitation back towards the driver. Precipitation also scatters light across a wide range of angles that disrupts the vision of drivers in oncoming vehicles. In contrast to recent computer vision methods that digitally remove rain and snow streaks from captured images, we present a system that will directly improve driver visibility by controlling illumination in response to detected precipitation. The motion of precipitation is tracked and only the space around particles is illuminated using fast dynamic control. Using a physics-based simulator, we show how such a system would perform under a variety of weather conditions. We build and evaluate a proof-of-concept system that can avoid water drops generated in the laboratory.
Raoul de Charette, Robert Tamburo, Peter C. Barnum, Anthony Rowe 0001, Takeo Kanade, Srinivasa G. Narasimhan
ICCP1
2010 Detection of unfocused raindrops on a windscreen using low level image processing
abstract
In a scene, rain produces a complex set of visual effects. Obviously, such effects may infer failures in outdoor vision-based systems which could have important side-effects in terms of security applications. For the sake of these applications, rain detection would be useful to adjust their reliability. In this paper, we introduce the problem (almost unprecedented) of unfocused raindrops. Then, we present a first approach to detect these unfocused raindrops on a transparent screen using a spatio-temporal approach to achieve detection in real-time. We successfully tested our algorithm for Intelligent Transport System (ITS) using an on-board camera and thus, detected the raindrops on the windscreen. Our algorithm differs from the others in that we do not need the focus to be set on the windscreen. Therefore, it means that our algorithm may run on the same camera sensor as the other vision-based algorithms.
Fawzi Nashashibi, Raoul de Charette, Alexandre Lia
ICARCV2
2009 Traffic light recognition using image processing compared to learning processes
abstract
In this paper we introduce a real-time traffic light recognition system for intelligent vehicles. The method proposed is fully based on image processing. Detection step is achieved in grayscale with spot light detection, and recognition is done using our generic ¿adaptive templates¿. The whole process was kept modular which make our TLR capable of recognizing different traffic lights from various countries. To compare our image processing algorithm with standard object recognition methods we also developed several traffic light recognition systems based on learning processes such as cascade classifiers with AdaBoost. Our system was validated in real conditions in our prototype vehicle and also using registered video sequence from various countries (France, China, and U.S.A.). We noticed high rate of correctly recognized traffic lights and few false alarms. Processing is performed in real-time on 640x480 images using a 2.9 GHz single core desktop computer.
Raoul de Charette, Fawzi Nashashibi
IROS1