VLDB 2026 Research / reviewers in the wild / expert
Clément Rambour
dblp:211/2108
· DBLP profile ↗
17ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-9899-3201ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorSystems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Computationally Efficient Text-to-Video Editing for Autonomous Driving Data Augmentation
Thérèse Tisseau des Escotais, Adam Diakite, Bertrand Leroy, Javier Ibanez-Guzman, Clément Rambour, Arnaud Breloy |
IV | 5 |
| 2025 | ViLU: Learning Vision-Language Uncertainties for Failure PredictionabstractReliable Uncertainty Quantification (UQ) and failure prediction remain open challenges for Vision-Language Models (VLMs). We introduce ViLU, a new Vision-Language Uncertainty quantification framework that contextualizes uncertainty estimates by leveraging all task-relevant textual representations. ViLU constructs an uncertainty-aware multi-modal representation by integrating the visual embedding, the predicted textual embedding, and an image-conditioned textual representation via cross-attention. Unlike traditional UQ methods based on loss prediction, ViLU trains an uncertainty predictor as a binary classifier to distinguish correct from incorrect predictions using a weighted binary cross-entropy loss, making it loss-agnostic. In particular, our proposed approach is well-suited for post-hoc settings, where only vision and text embeddings are available without direct access to the model itself. Extensive experiments on diverse datasets show the significant gains of our method compared to state-of-the-art failure prediction methods. We apply our method to standard classification datasets, such as ImageNet-1k, as well as large-scale image-caption datasets like CC12M and LAION-400M. Ablation studies highlight the critical role of our architecture and training in achieving effective uncertainty quantification. Our code is publicly available and can be found here: https://github.com/ykrmm/ViLU. Marc Lafon, Yannis Karmim, Julio Silva-Rodríguez, Paul Couairon, Clément Rambour, Raphaël Fournier-S'niehotta, Ismail Ben Ayed, Jose Dolz, Nicolas Thome |
ICCV | 5 |
| 2025 | RT-HCP: Dealing with Inference Delays and Sample Efficiency to Learn Directly on Robotic PlatformsabstractLearning a controller directly on the robot requires extreme sample efficiency. Model-based reinforcement learning (RL) methods are the most sample efficient, but they often suffer from a too long inference time to meet the robot control frequency requirements. In this paper, we address the sample efficiency and inference time challenges with two contributions. First, we define a general framework to deal with inference delays where the slow inference robot controller provides a sequence of actions to feed the control-hungry robotic platform without execution gaps. Then, we compare several RL algorithms in the light of this framework and propose RT-HCP, an algorithm that offers an excellent trade-off between performance, sample efficiency and inference time. We validate the superiority of RT-HCP with experiments where we learn a controller directly on a simple but high frequency FURUTA pendulum platform. Code: github.com/elasriz/RTHCP Zakariae El Asri, Ibrahim Laiche, Clément Rambour, Olivier Sigaud, Nicolas Thome |
IROS | 3 |
| 2025 | CLIPTTA: Robust Contrastive Vision-Language Test-Time AdaptationabstractVision-language models (VLMs) like CLIP exhibit strong zero-shot capabilities but often fail to generalize under distribution shifts. Test-time adaptation (TTA) allows models to update at inference time without labeled data, typically via entropy minimization. However, this objective is fundamentally misaligned with the contrastive image-text training of VLMs, limiting adaptation performance and introducing failure modes such as pseudo-label drift and class collapse. We propose CLIPTTA, a new gradient-based TTA method for vision-language models that leverages a soft contrastive loss aligned with CLIP’s pre-training objective. We provide a theoretical analysis of CLIPTTA’s gradients, showing how its batch-aware design mitigates the risk of collapse. We further extend CLIPTTA to the open-set setting, where both in-distribution (ID) and out-of-distribution (OOD) samples are encountered, using an Outlier Contrastive Exposure (OCE) loss to improve OOD detection. Evaluated on 75 datasets spanning diverse distribution shifts, CLIPTTA consistently outperforms entropy-based objectives and is highly competitive with state-of-the-art TTA methods, outperforming them on a large number of datasets and exhibiting more stable performance across diverse shifts. Marc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers, Nicolas Thome |
NeurIPS | 3 |
| 2025 | Optimization of Rank Losses for Image RetrievalabstractIn image retrieval, standard evaluation metrics rely on score ranking, e.g. average precision (AP), recall at k (R@k), normalized discounted cumulative gain (NDCG). In this work, we introduce a general framework for robust and decomposable rank losses optimization. It addresses two major challenges for end-to-end training of deep neural networks with rank losses: non-differentiability and non-decomposability. First, we propose a general surrogate for ranking operator, SupRank, that is amenable to stochastic gradient descent. It provides an upperbound for rank losses and ensures robust training. Second, we use a simple yet effective loss function to reduce the decomposability gap between the averaged batch approximation of ranking losses and their values on the whole training set. We apply our framework to two standard metrics for image retrieval: AP and R@k. Additionally, we apply our framework to hierarchical image retrieval. We introduce an extension of AP, the hierarchical average precision $\mathcal {H}{\mathrm -AP}$H- AP , and optimize it as well as the NDCG. Finally, we create the first hierarchical landmarks retrieval dataset. We use a semi-automatic pipeline to create hierarchical labels, extending the large scale Google Landmarks v2 dataset. Elias Ramzi, Nicolas Audebert, Clément Rambour, André Araújo 0001, Xavier Bitot, Nicolas Thome |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | GalLoP: Learning Global and Local Prompts for Vision-Language Models
Marc Lafon, Elias Ramzi, Clément Rambour, Nicolas Audebert, Nicolas Thome |
ECCV (61) | 3 |
| 2024 | MERLIN-Seg: Self-supervised despeckling for label-efficient semantic segmentationabstractRemote sensing satellites acquire a continuous stream of data on a daily basis. As most of those data are unlabeled, the development of algorithms requiring weak supervision is of paramount importance. In this paper, we show that the need for annotation for Synthetic Aperture Radar data can be reduced by coupling a despeckling task (self-supervised) and a segmentation task (supervised). The proposed self-supervised learning framework, called MERLIN-Seg, has been trained for building footprint extraction and achieves favorable performances even with 1% of annotated data. We show that conditioning the network on despeckling without labels is beneficial for supervised segmentation. Our experiments demonstrate that the joint training of the two tasks achieves better performances than a vanilla segmentation network in terms of IoU, F1 score, and accuracy on both simulated and real SAR images. Emanuele Dalsasso, Clément Rambour, Nicolas Trouvé, Nicolas Thome |
Comput. Vis. Image Underst. | 2 |
| 2023 | Hybrid Energy Based Model in the Feature Space for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is a critical requirement for the deployment of deep neural networks. This paper introduces the HEAT model, a new post-hoc OOD detection method estimating the density of in-distribution (ID) samples using hybrid energy-based models (EBM) in the feature space of a pre-trained backbone. HEAT complements prior density estimators of the ID density, e.g. parametric models like the Gaussian Mixture Model (GMM), to provide an accurate yet robust density estimation. A second contribution is to leverage the EBM framework to provide a unified density estimation and to compose several energy terms. Extensive experiments demonstrate the significance of the two contributions. HEAT sets new state-of-the-art OOD detection results on the CIFAR-10 / CIFAR-100 benchmark as well as on the large-scale Imagenet benchmark. The code is available at: https://github.com/MarcLafon/heatood. Marc Lafon, Elias Ramzi, Clément Rambour, Nicolas Thome |
ICML | 3 |
| 2023 | Full Contextual Attention for Multi-resolution Transformers in Semantic SegmentationabstractTransformers have proved to be very effective for visual recognition tasks. In particular, vision transformers construct compressed global representations through self-attention and learnable class tokens. Multi-resolution transformers have shown recent successes in semantic segmentation but can only capture local interactions in high-resolution feature maps. This paper extends the notion of global tokens to build GLobal Attention Multi-resolution (GLAM) transformers. GLAM is a generic module that can be integrated into most existing transformer backbones. GLAM includes learnable global tokens, which unlike previous methods can model interactions between all image regions, and extracts powerful representations during training. Extensive experiments show that GLAM-Swin or GLAM-Swin-Unet exhibit substantially better performances than their vanilla counterparts on ADE20K and Cityscapes. Moreover, GLAM can be used to segment large 3D medical images, and GLAM-nnFormer achieves new state-of-the-art performance on the BCV dataset. Loic Themyr, Clément Rambour, Nicolas Thome, Toby Collins, Alexandre Hostettler |
WACV | 2 |
| 2022 | Complementing Brightness Constancy with Deep Networks for Optical Flow Prediction
Vincent Le Guen, Clément Rambour, Nicolas Thome |
ECCV (21) | 2 |
| 2022 | Hierarchical Average Precision Training for Pertinent Image Retrieval
Elias Ramzi, Nicolas Audebert, Nicolas Thome, Clément Rambour, Xavier Bitot |
ECCV (14) | 4 |
| 2021 | Robust and Decomposable Average Precision for Image RetrievalabstractIn image retrieval, standard evaluation metrics rely on score ranking, e.g. average precision (AP). In this paper, we introduce a method for robust and decomposable average precision (ROADMAP) addressing two major challenges for end-to-end training of deep neural networks with AP: non-differentiability and non-decomposability.Firstly, we propose a new differentiable approximation of the rank function, which provides an upper bound of the AP loss and ensures robust training. Secondly, we design a simple yet effective loss function to reduce the decomposability gap between the AP in the whole training set and its averaged batch approximation, for which we provide theoretical guarantees.Extensive experiments conducted on three image retrieval datasets show that ROADMAP outperforms several recent AP approximation methods and highlight the importance of our two contributions. Finally, using ROADMAP for training deep models yields very good performances, outperforming state-of-the-art results on the three datasets.Code and instructions to reproduce our results will be made publicly available at https://github.com/elias-ramzi/ROADMAP. Elias Ramzi, Nicolas Thome, Clément Rambour, Nicolas Audebert, Xavier Bitot |
NeurIPS | 3 |
| 2020 | Regularized SAR Tomography ApproachesabstractSynthetic Aperture Radar (SAR) tomographic techniques enable the reconstruction of the scene scattering structure along the vertical direction and can provide the temporal evolution of a cloud of reliable points located in the 3D space. The use of Generalized Likelihood Ratio Test approaches have been shown to be effective in selecting reliable multiple scatterers. Recently regularized tomographic methods have been proposed for increasing the density of the recovered scatterers in urban environments. This paper discusses the differences between these two approaches and performs a comparison of reconstruction results obtained from a stack of TerraSAR-X images, in a region of interest located in the city of Paris, France. Alessandra Budillon, Loïc Denis, Clément Rambour, Gilda Schirinzi, Florence Tupin |
IGARSS | 3 |
| 2019 | Urban surface reconstruction in SAR tomography by graph-cuts
Clément Rambour, Loïc Denis, Florence Tupin, Hélène Oriot, Yue Huang 0002, Laurent Ferro-Famil |
Comput. Vis. Image Underst. | 1 |
| 2019 | Introducing Spatial Regularization in SAR Tomography ReconstructionabstractThe resolution achieved by current synthetic aperture radar (SAR) sensors provides a detailed visualization of urban areas. Spaceborne sensors such as TerraSAR-X can be used to analyze large areas at a very high resolution. In addition, repeated passes of the satellite give access to temporal and interferometric information on the scene. Because of the complex 3-D structure of urban surfaces, scatterers located at different heights (ground, building facade, and roof) produce radar echoes that often get mixed within the same radar cells. These echoes must be numerically unmixed in order to get a fine understanding of the radar images. This unmixing is at the core of SAR tomography. SAR tomography reconstruction is generally performed in two steps: 1) reconstruction of the so-called tomogram by vertical focusing, at each radar resolution cell, to extract the complex amplitudes (a 1-D processing) and 2) transformation from radar geometry to ground geometry and extraction of significant scatterers. We propose to perform the tomographic inversion directly in ground geometry in order to enforce spatial regularity in 3-D space. This inversion requires solving a large-scale nonconvex optimization problem. We describe an iterative method based on variable splitting and the augmented Lagrangian technique. Spatial regularizations can easily be included in this generic scheme. We illustrate, on simulated data and a TerraSAR-X tomographic data set, the potential of this approach to produce 3-D reconstructions of urban surfaces. Clément Rambour, Loïc Denis, Florence Tupin, Hélène Oriot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | SAR Tomography of Urban Areas: 3D Regularized Inversion in the Scene GeometryabstractStarting from a stack of co-registered SAR images in interferometric configuration, SAR tomography performs a reconstruction of the reflectivity of scatterers in 3-D. Scatterers seen within the same resolution cell in each SAR image can be separated by jointly unmixing the SAR complex amplitude observed throughout the stack. In urban areas, Compress Sensing (CS) approaches have been applied to achieve super-resolution in the estimation of the position of the scatterers. However, even if all the local information coming from a stack at a given pixel is used, the structural information that is inherent to the image is not directly used to improve the rendering of the scene. This paper addresses the problem of adding structural constraints to sparse tomographic reconstructions of urban areas. We derive an algorithm allowing the inversion of tomographic data under structural constraints and illustrate its performances on a stack of Spotlight TerraSAR-X images. Clément Rambour, Loïc Denis, Florence Tupin, Jean-Marie Nicolas 0002, Hélène Oriot |
IGARSS | 1 |
| 2017 | Similarity criterion for SAR tomography over dense urban areaabstractStarting from a stack of co-registered SAR images in interferometric configuration, SAR tomography performs a reconstruction of the reflectivity of scatterers in 3-D. Several scatterers observed within the same resolution cell of each SAR image can be separated by jointly unmixing the SAR complex amplitude observed throughout the stack. To achieve a reliable tomographic reconstruction, it is necessary to estimate locally the SAR covariance matrix by performing some spatial averaging. This necessary averaging step introduces some resolution loss and can bias the tomographic reconstruction by mistakenly including the response of scatterers located within the averaging area but outside the resolution cell of interest. This paper addresses the problem of identifying pixels corresponding to similar tomographic content, i.e., pixels that can be safely averaged prior to tomographic reconstruction. We derive a similarity criterion adapted to SAR tomography and compare its performance with existing criteria on a stack of Spotlight TerraSAR-X images. Clément Rambour, Loïc Denis, Florence Tupin, Jean-Marie Nicolas 0002, Hélène Oriot, Laurent Ferro-Famil, Charles-Alban Deledalle |
IGARSS | 1 |