VLDB 2026 Research / reviewers in the wild / expert
Gabriela Csurka
dblp:c/GabrielaCsurka
· DBLP profile ↗
57ranked-venue papers
13as first author
16since 2021 · last 2025
0009-0005-1067-1056ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 10 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 8 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MUSt3R: Multi-view Network for Stereo 3D ReconstructionabstractDUSt3R introduced a novel paradigm in geometric computer vision by proposing a model that can provide dense and unconstrained Stereo 3D Reconstruction of arbitrary image collections with no prior information about camera calibration nor viewpoint poses. Under the hood, however, DUSt3R processes image pairs, regressing local 3D reconstructions that need to be aligned in a global coordinate system. The number of pairs, growing quadratically, is an inherent limitation that becomes especially concerning for robust and fast optimization in the case of large image collections. In this paper, we propose an extension of DUSt3R from pairs to multiple views, that addresses all aforementioned concerns. Indeed, we propose a Multi-view Network for Stereo 3D Reconstruction, or MUSt3R, that modifies the DUSt3R architecture by making it symmetric and extending it to directly predict 3D structure for all views in a common coordinate frame. Second, we entail the model with a multi-layer memory mechanism which allows to reduce the computational complexity and to scale the reconstruction to large collections, inferring thousands of 3D pointmaps at high frame-rates with limited added complexity. The framework is designed to perform 3D reconstruction both offline and online, and hence can be seamlessly applied to SfM and visual SLAM scenarios showing state-of-the-art performance on various 3D downstream tasks, including uncalibrated Visual Odometry, relative camera pose, scale and focal estimation, 3D reconstruction and multi-view depth estimation. Yohann Cabon, Lucas Stoffl, Leonid Antsfeld, Gabriela Csurka, Boris Chidlovskii, Jérôme Revaud, Vincent Leroy 0003 |
CVPR | 4 |
| 2025 | Gaussian Splatting Feature Fields for (Privacy-Preserving) Visual LocalizationabstractVisual localization is the task of estimating a camera pose in a known environment. In this paper, we utilize 3D Gaussian Splatting (3DGS)-based representations for accurate and privacy-preserving visual localization. We propose Gaussian Splatting Feature Fields (GSFFs), a scene representation for visual localization that combines an explicit geometry model (3DGS) with an implicit feature field. We leverage the dense geometric information and differentiable rasterization algorithm from 3DGS to learn robust feature representations grounded in 3D. In particular, we align a 3D scale-aware feature field and a 2D feature encoder in a common embedding space through a contrastive framework. Using a 3D structure-informed clustering procedure, we further regularize the representation learning and seamlessly convert the features to segmentations, which can be used for privacy-preserving visual localization. Pose refinement, which involves aligning either feature maps or segmentations from a query image with those rendered from the GSFFs scene representation, is used to achieve localization. The resulting privacy- and non-privacy-preserving localization pipelines, evaluated on multiple real-world datasets, show state-of-the-art performances. Maxime Pietrantoni, Gabriela Csurka, Torsten Sattler |
CVPR | 2 |
| 2025 | PanSt3R: Multi-View Consistent Panoptic SegmentationabstractPanoptic segmentation of 3D scenes, involving the segmentation and classification of object instances in a dense 3D reconstruction of a scene, is a challenging problem, especially when relying solely on unposed 2D images. Existing approaches typically leverage off-the-shelf models to extract per-frame 2D panoptic segmentations, before optimizing an implicit geometric representation (often based on NeRF) to integrate and fuse the 2D predictions. We argue that relying on 2D panoptic segmentation for a problem inherently 3D and multi-view is likely suboptimal as it fails to leverage the full potential of spatial relationships across views. In addition to requiring camera parameters, these approaches also necessitate computationally expensive test-time optimization for each scene. Instead, in this work, we propose a unified and integrated approach PanSt3R, which eliminates the need for test-time optimization by jointly predicting 3D geometry and multi-view panoptic segmentation in a single forward pass. Our approach builds upon recent advances in 3D reconstruction, specifically upon MUSt3R, a scalable multi-view version of DUSt3R, and enhances it with semantic awareness and multi-view panoptic segmentation capabilities. We additionally revisit the standard post-processing mask merging procedure and introduce a more principled approach for multi-view segmentation. We also introduce a simple method for generating novel-view predictions based on the predictions of PanSt3R and vanilla 3DGS. Overall, the proposed PanSt3R is conceptually simple, yet fast and scalable, and achieves state-of-the-art performance on several benchmarks, while being orders of magnitude faster than existing methods. Lojze Zust, Yohann Cabon, Juliette Marrie, Leonid Antsfeld, Boris Chidlovskii, Jérôme Revaud, Gabriela Csurka |
ICCV | 7 |
| 2025 | Test-Time Vocabulary Adaptation for Language-Driven Object DetectionabstractOpen-Vocabulary object detection models allow users to freely specify a class vocabulary in natural language at test time, guiding the detection of desired objects. However, vocabularies can be overly broad or even mis-specified, hampering the overall performance of the detector. In this work, we propose a plug-and-play Vocabulary Adapter (VocAda) to refine the user-defined vocabulary, automatically tailoring it to categories that are relevant for a given image. VocAda does not require any training, it operates at inference time in three steps: i) it uses an image captionner to describe visible objects, ii) it parses nouns from those captions, and iii) it selects relevant classes from the user-defined vocabulary, discarding irrelevant ones. Experiments on COCO and Objects365 with three state-of-the-art detectors show that VocAda consistently improves performance, proving its versatility. The code is open source. Tyler L. Hayes, Massimiliano Mancini, Elisa Ricci 0001, Riccardo Volpi, Gabriela Csurka |
ICIP | 6 |
| 2024 | Self-Supervised Learning of Neural Implicit Feature Fields for Camera Pose RefinementabstractVisual localization techniques rely upon some underlying scene representation to localize against. These representations can be explicit such as 3D SFM map or implicit, such as a neural network that learns to encode the scene. The former requires sparse feature extractors and matchers to build the scene representation. The latter might lack geometric grounding not capturing the 3D structure of the scene well enough. This paper proposes to jointly learn the scene representation along with a 3D dense feature field and a 2D feature extractor whose outputs are embedded in the same metric space. Through a contrastive framework we align this volumetric field with the image-based extractor and regularize the latter with a ranking loss from learned surface information. We learn the underlying geometry of the scene with an implicit field through volumetric rendering and design our feature field to leverage intermediate geometric information encoded in the implicit field. The resulting features are discriminative and robust to viewpoint change while maintaining rich encoded information. Visual localization is then achieved by aligning the image-based features and the rendered volumetric features. We show the effectiveness of our approach on real-world scenes, demonstrating that our approach outperforms prior and concurrent work on leveraging implicit scene representations for localization. Maxime Pietrantoni, Gabriela Csurka, Martin Humenberger, Torsten Sattler |
3DV | 2 |
| 2024 | SHiNe: Semantic Hierarchy Nexus for Open-Vocabulary Object DetectionabstractOpen-vocabulary object detection (OvOD) has transformed detection into a language-guided task, empowering users to freely define their class vocabularies of interest during inference. However, our initial investigation indicates that existing OvOD detectors exhibit significant vari-ability when dealing with vocabularies across various semantic granularities, posing a concern for real-world deployment. To this end, we introduce Semantic Hierarchy Nexus (SHiNe), a novel classifier that uses semantic knowledge from class hierarchies. It runs offline in three steps: i) it retrieves relevant super-/sub-categories from a hierar-chy for each target class; ii) it integrates these categories into hierarchy-aware sentences; iii) it fuses these sentence embeddings to generate the nexus classifier vector. Our evaluation on various detection benchmarks demonstrates that SHiNe enhances robustness across diverse vocabulary granularities, achieving up to +31.9% mAP50 with ground truth hierarchies, while retaining improvements using hierarchies generated by large language models. Moreover, when applied to open-vocabulary classification on ImageNet-1k, SHiNe improves the CLIP zero-shot baseline by +2.8% accuracy. SHiNe is training-free and can be seamlessly integrated with any off-the-shelf OvOD detector, without incurring additional computational overhead during inference. The code is open source. Tyler L. Hayes, Elisa Ricci 0001, Gabriela Csurka, Riccardo Volpi |
CVPR | 4 |
| 2024 | Weatherproofing Retrieval for Localization with Generative AI and Geometric ConsistencyabstractState-of-the-art visual localization approaches generally rely on a first image retrieval step whose role is crucial. Yet, retrieval often struggles when facing varying conditions, due to e.g. weather or time of day, with dramatic consequences on the visual localization accuracy. In this paper, we improve this retrieval step and tailor it to the final localization task. Among the several changes we advocate for, we propose to synthesize variants of the training set images, obtained from generative text-to-image models, in order to automatically expand the training set towards a number of nameable variations that particularly hurt visual localization. After expanding the training set, we propose a training approach that leverages the specificities and the underlying geometry of this mix of real and synthetic images. We experimentally show that those changes translate into large improvements for the most challenging visual localization datasets. Yannis Kalantidis, Mert Bülent Sariyildiz, Rafael S. Rezende, Philippe Weinzaepfel, Diane Larlus, Gabriela Csurka |
ICLR | 6 |
| 2023 | SegLoc: Learning Segmentation-Based Representations for Privacy-Preserving Visual LocalizationabstractInspired by properties of semantic segmentation, in this paper we investigate how to leverage robust image segmentation in the context of privacy-preserving visual localization. We propose a new localization framework, SegLoc, that leverages image segmentation to create robust, compact, and privacy-preserving scene representations, i.e., 3D maps. We build upon the correspondence-supervised, fine-grained segmentation approach from [42], making it more robust by learning a set of cluster labels with discriminative clustering, additional consistency regularization terms and we jointly learn a global image representation along with a dense local representation. In our localization pipeline, the former will be used for retrieving the most similar images, the latter to refine the retrieved poses by minimizing the label inconsistency between the 3D points of the map and their projection onto the query image. In various experiments, we show that our proposed representation allows to achieve (close-to) state-of-the-art pose estimation results while only using a compact 3D map that does not contain enough information about the original images for an attacker to reconstruct personal information. Maxime Pietrantoni, Martin Humenberger, Torsten Sattler, Gabriela Csurka |
CVPR | 4 |
| 2023 | CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowabstractDespite impressive performance for high-level downstream tasks, self-supervised pre-training methods have not yet fully delivered on dense geometric vision tasks such as stereo matching or optical flow. The application of self-supervised concepts, such as instance discrimination or masked image modeling, to geometric tasks is an active area of research. In this work, we build on the recent cross-view completion framework, a variation of masked image modeling that leverages a second view from the same scene which makes it well suited for binocular downstream tasks. The applicability of this concept has so far been limited in at least two ways: (a) by the difficulty of collecting real-world image pairs – in practice only synthetic data have been used – and (b) by the lack of generalization of vanilla transformers to dense downstream tasks for which relative position is more meaningful than absolute position. We explore three avenues of improvement. First, we introduce a method to collect suitable real-world image pairs at large scale. Second, we experiment with relative positional embeddings and show that they enable vision transformers to perform substantially better. Third, we scale up vision transformer based cross-completion architectures, which is made possible by the use of large amounts of data. With these improvements, we show for the first time that state-of-the-art results on stereo matching and optical flow can be reached without using any classical task-specific techniques like correlation volume, iterative estimation, image warping or multi-scale reasoning, thus paving the way towards universal vision models. Philippe Weinzaepfel, Thomas Lucas 0002, Vincent Leroy 0003, Yohann Cabon, Vaibhav Arora, Romain Brégier, Gabriela Csurka, Leonid Antsfeld, Boris Chidlovskii, Jérôme Revaud |
ICCV | 7 |
| 2022 | Deep Visual Geo-localization BenchmarkabstractIn this paper, we propose a new open-source benchmarkingframeworkfor Visual Geo-localization (VG) that allows to build, train, and test a wide range of commonly used ar-chitectures, with the flexibility to change individual components of a geo-localization pipeline. The purpose of this framework is twofold: i) gaining insights into how differ-ent components and design choices in a VG pipeline im-pact the final results, both in terms of performance (re-call@N metric) and system requirements (such as execution time and memory consumption); ii) establish a system-atic evaluation protocol for comparing different methods. Using the proposed framework, we perform a large suite of experiments which provide criteria for choosing back-bone, aggregation and negative mining depending on the use-case and requirements. We also assess the impact of engineering techniques like pre/post-processing, data aug-mentation and image resizing, showing that better performance can be obtained through somewhat simple procedures: for example, downscaling the images' resolution to 80% can lead to similar results with a 36% savings in ex-traction time and dataset storage requirement. Code and trained models are available at dataset storage requirement. https://deep-vg-bench.herokuapp.com/. Gabriele Moreno Berton, Riccardo Mereu, Gabriele Trivigno, Carlo Masone, Gabriela Csurka, Torsten Sattler, Barbara Caputo |
CVPR | 5 |
| 2022 | On the Road to Online Adaptation for Semantic Image SegmentationabstractWe propose a new problem formulation and a corresponding evaluation framework to advance research on unsupervised domain adaptation for semantic image segmentation. The overall goal is fostering the development of adaptive learning systems that will continuously learn, without supervision, in ever-changing environments. Typical protocols that study adaptation algorithms for segmentation models are limited to few domains, adaptation happens offline, and human intervention is generally required, at least to annotate data for hyperparameter tuning. We argue that such constraints are incompatible with algorithms that can continuously adapt to different real-world situations. To address this, we propose a protocol where models need to learn online, from sequences of temporally correlated images, requiring continuous, frame-by-frame adaptation. We accompany this new protocol with a variety of baselines to tackle the proposed formulation, as well as an extensive analysis of their behaviors, which can serve as a starting point for future research. Riccardo Volpi, Pau de Jorge, Diane Larlus, Gabriela Csurka |
CVPR | 4 |
| 2022 | ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity
Ginger Delmas, Rafael S. Rezende, Gabriela Csurka, Diane Larlus |
ICLR | 3 |
| 2022 | CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionabstractMasked Image Modeling (MIM) has recently been established as a potent pre-training paradigm. A pretext task is constructed by masking patches in an input image, and this masked content is then predicted by a neural network using visible patches as sole input. This pre-training leads to state-of-the-art performance when finetuned for high-level semantic tasks, e.g. image classification and object detection. In this paper we instead seek to learn representations that transfer well to a wide variety of 3D vision and lower-level geometric downstream tasks, such as depth prediction or optical flow estimation. Inspired by MIM, we propose an unsupervised representation learning task trained from pairs of images showing the same scene from different viewpoints. More precisely, we propose the pretext task of cross-view completion where the first input image is partially masked, and this masked content has to be reconstructed from the visible content and the second image. In single-view MIM, the masked content often cannot be inferred precisely from the visible portion only, so the model learns to act as a prior influenced by high-level semantics. In contrast, this ambiguity can be resolved with cross-view completion from the second unmasked image, on the condition that the model is able to understand the spatial relationship between the two images. Our experiments show that our pretext task leads to significantly improved performance for monocular 3D vision downstream tasks such as depth estimation. In addition, our model can be directly applied to binocular downstream tasks like optical flow or relative camera pose estimation, for which we obtain competitive results without bells and whistles, i.e., using a generic architecture without any task-specific design. Philippe Weinzaepfel, Vincent Leroy 0003, Thomas Lucas 0002, Romain Brégier, Yohann Cabon, Vaibhav Arora, Leonid Antsfeld, Boris Chidlovskii, Gabriela Csurka, Jérôme Revaud |
NeurIPS | 9 |
| 2022 | Investigating the Role of Image Retrieval for Visual Localization
Martin Humenberger, Yohann Cabon, Noé Pion, Philippe Weinzaepfel, Nicolas Guérin, Torsten Sattler, Gabriela Csurka |
Int. J. Comput. Vis. | 8 |
| 2021 | Large-Scale Localization Datasets in Crowded Indoor SpacesabstractEstimating the precise location of a camera using visual localization enables interesting applications such as augmented reality or robot navigation. This is particularly useful in indoor environments where other localization technologies, such as GNSS, fail. Indoor spaces impose interesting challenges on visual localization algorithms: occlusions due to people, textureless surfaces, large viewpoint changes, low light, repetitive textures, etc. Existing indoor datasets are either comparably small or do only cover a subset of the mentioned challenges. In this paper, we introduce 5 new indoor datasets for visual localization in challenging real-world environments. They were captured in a large shopping mall and a large metro station in Seoul, South Korea, using a dedicated mapping platform consisting of 10 cameras and 2 laser scanners. In order to obtain accurate ground truth camera poses, we developed a robust LiDAR SLAM which provides initial poses that are then refined using a novel structure-from-motion based optimization. We present a benchmark of modern visual localization algorithms on these challenging datasets showing superior performance of structure-based methods using robust image features. The datasets are available at: https://naverlabs.com/datasets Soohyun Ryu, Suyong Yeon, Yonghan Lee 0001, Deokhwa Kim, Cheolho Han, Yohann Cabon, Philippe Weinzaepfel, Nicolas Guérin, Gabriela Csurka, Martin Humenberger |
CVPR | 10 |
| 2021 | Unsupervised Meta-Domain Adaptation for Fashion RetrievalabstractCross-domain fashion item retrieval naturally arises when unconstrained consumer images are used to query for fashion items in a collection of high-quality photographs provided by retailers. To perform this task, approaches typically leverage both consumer and shop domains from a given dataset to learn a domain invariant representation, allowing these images of different nature to be directly compared. When consumer images are not available beforehand, such training is impossible. In this paper, we focus on this challenging and yet practical scenario, and we propose instead to leverage representations learned for cross-domain retrieval from another source dataset and to adapt them to the target dataset for this particular setting. More precisely, we bypass the lack of consumer images and directly target the more challenging meta-domain gap which occurs between consumer images and shop images, independently of their dataset. Assuming that datasets share some similar fashion items, we cluster their shop images and leverage the clusters to automatically generate pseudo-labels. Those are used to associate consumer and shop images across datasets, which in turn allows to learn meta-domain-invariant representations suitable for cross-domain retrieval in the target dataset. The features and code are available at https://github.com/vivoutlaw/UDMA. Vivek Sharma 0001, Naila Murray, Diane Larlus, M. Saquib Sarfraz, Rainer Stiefelhagen, Gabriela Csurka |
WACV | 6 |
| 2020 | Benchmarking Image Retrieval for Visual LocalizationabstractVisual localization, i.e., camera pose estimation in a known scene, is a core component of technologies such as autonomous driving and augmented reality. State-of-the-art localization approaches often rely on image retrieval techniques for one of two tasks: (1) provide an approximate pose estimate or (2) determine which parts of the scene are potentially visible in a given query image. It is common practice to use state-of-the-art image retrieval algorithms for these tasks. These algorithms are often trained for the goal of retrieving the same landmark under a large range of viewpoint changes. However, robustness to viewpoint changes is not necessarily desirable in the context of visual localization. This paper focuses on understanding the role of image retrieval for multiple visual localization tasks. We introduce a benchmark setup and compare state-of-the-art retrieval representations on multiple datasets. We show that retrieval performance on classical landmark retrieval/recognition tasks correlates only for some but not all tasks to localization performance. This indicates a need for retrieval approaches specifically designed for localization tasks. Our benchmark and evaluation protocols are available at https://github.com/naver/kapture-localization. Noé Pion, Martin Humenberger, Gabriela Csurka, Yohann Cabon, Torsten Sattler |
3DV | 3 |
| 2020 | Estimating Low-Rank Region Likelihood MapsabstractLow-rank regions capture geometrically meaningful structures in an image which encompass typical local features such as edges, corners and all kinds of regular, symmetric, often repetitive patterns, that are commonly found in man-made environment. While such patterns are challenging current state-of-the-art feature correspondence methods, the recovered homography of a low-rank texture readily provides 3D structure with respect to a 3D plane, without any prior knowledge of the visual information on that plane. However, the automatic and efficient detection of the broad class of low-rank regions is unsolved. Herein, we propose a novel self-supervised low-rank region detection deep network that predicts a low-rank likelihood map from an image. The evaluation of our method on real-world datasets shows not only that it reliably predicts low-rank regions in the image similarly to our baseline method, but thanks to the data augmentations used in the training phase it generalizes well to difficult cases (e.g. day/night lighting, low contrast, underexposure) where the baseline prediction fails. Gabriela Csurka, Zoltan Kato, Andor Juhasz, Martin Humenberger |
CVPR | 1 |
| 2019 | Visual Localization by Learning Objects-Of-Interest Dense Match RegressionabstractWe introduce a novel CNN-based approach for visual localization from a single RGB image that relies on densely matching a set of Objects-of-Interest (OOIs). In this paper, we focus on planar objects which are highly descriptive in an environment, such as paintings in museums or logos and storefronts in malls or airports. For each OOI, we define a reference image for which 3D world coordinates are available. Given a query image, our CNN model detects the OOIs, segments them and finds a dense set of 2D-2D matches between each detected OOI and its corresponding reference image. Given these 2D-2D matches, together with the 3D world coordinates of each reference image, we obtain a set of 2D-3D matches from which solving a Perspective-n-Point problem gives a pose estimate. We show that 2D-3D matches for reference images, as well as OOI annotations can be obtained for all training images from a single instance annotation per OOI by leveraging Structure-from-Motion reconstruction. We introduce a novel synthetic dataset, VirtualGallery, which targets challenges such as varying lighting conditions and different occlusion levels. Our results show that our method achieves high precision and is robust to these challenges. We also experiment using the Baidu localization dataset captured in a shopping mall. Our approach is the first deep regression-based method to scale to such a larger environment. Philippe Weinzaepfel, Gabriela Csurka, Yohann Cabon, Martin Humenberger |
CVPR | 2 |
| 2019 | Fine-Grained Action Retrieval Through Multiple Parts-of-Speech EmbeddingsabstractWe address the problem of cross-modal fine-grained action retrieval between text and video. Cross-modal retrieval is commonly achieved through learning a shared embedding space, that can indifferently embed modalities. In this paper, we propose to enrich the embedding by disentangling parts-of-speech (PoS) in the accompanying captions. We build a separate multi-modal embedding space for each PoS tag. The outputs of multiple PoS embeddings are then used as input to an integrated multi-modal space, where we perform action retrieval. All embeddings are trained jointly through a combination of PoS-aware and PoS-agnostic losses. Our proposal enables learning specialised embedding spaces that offer multiple views of the same embedded entities. We report the first retrieval results on fine-grained actions for the large-scale EPIC dataset, in a generalised zero-shot setting. Results show the advantage of our approach for both video-to-text and text-to-video action retrieval. We also demonstrate the benefit of disentangling the PoS for the generic task of cross-modal video retrieval on the MSR-VTTdataset. Michael Wray, Gabriela Csurka, Diane Larlus, Dima Damen |
ICCV | 2 |
| 2016 | Domain Adaptation in the Absence of Source Domain DataabstractThe overwhelming majority of existing domain adaptation methods makes an assumption of freely available source domain data. An equal access to both source and target data makes it possible to measure the discrepancy between their distributions and to build representations common to both target and source domains. In reality, such a simplifying assumption rarely holds, since source data are routinely a subject of legal and contractual constraints between data owners and data customers. When source domain data can not be accessed, decision making procedures are often available for adaptation nevertheless. These procedures are often presented in the form of classification, identification, ranking etc. rules trained on source data and made ready for a direct deployment and later reuse. In other cases, the owner of a source data is allowed to share a few representative examples such as class means. In this paper we address the domain adaptation problem in real world applications, where the reuse of source domain data is limited to classification rules or a few representative examples. We extend the recent techniques of feature corruption and their marginalization, both in supervised and unsupervised settings. We test and compare them on private and publicly available source datasets and show that significant performance gains can be achieved despite the absence of source data and shortage of labeled target data. Boris Chidlovskii, Stéphane Clinchant, Gabriela Csurka |
KDD | 3 |
| 2015 | Unsupervised Visual and Textual Information Fusion in CBMIR Using Graph-Based MethodsabstractMultimedia collections are more than ever growing in size and diversity. Effective multimedia retrieval systems are thus critical to access these datasets from the end-user perspective and in a scalable way. We are interested in repositories of image/text multimedia objects and we study multimodal information fusion techniques in the context of content-based multimedia information retrieval. We focus on graph-based methods, which have proven to provide state-of-the-art performances. We particularly examine two such methods: cross-media similarities and random-walk-based scores. From a theoretical viewpoint, we propose a unifying graph-based framework, which encompasses the two aforementioned approaches. Our proposal allows us to highlight the core features one should consider when using a graph-based technique for the combination of visual and textual information. We compare cross-media and random-walk-based results using three different real-world datasets. From a practical standpoint, our extended empirical analyses allow us to provide insights and guidelines about the use of graph-based methods for multimodal information fusion in content-based multimedia information retrieval. Julien Ah-Pine, Gabriela Csurka, Stéphane Clinchant |
ACM Trans. Inf. Syst. | 2 |
| 2013 | What is a good evaluation measure for semantic segmentation?abstractIn this work, we consider the evaluation of the semantic segmentation task. We discuss the strengths and limitations of the few existing measures, and propose new ways to evaluate semantic segmentation. First, we argue that a per-image score instead of one computed over the entire dataset brings a lot more insight. Second, we propose to take contours more carefully into account. Based on the conducted experiments, we suggest best practices for the evaluation. Finally, we present a user study we conducted to better understand how the quality of image segmentations is perceived by humans. Gabriela Csurka, Diane Larlus, Florent Perronnin |
BMVC | 1 |
| 2013 | Tree-Structured CRF Models for Interactive Image LabelingabstractWe propose structured prediction models for image labeling that explicitly take into account dependencies among image labels. In our tree-structured models, image labels are nodes, and edges encode dependency relations. To allow for more complex dependencies, we combine labels in a single node and use mixtures of trees. Our models are more expressive than independent predictors, and lead to more accurate label predictions. The gain becomes more significant in an interactive scenario where a user provides the value of some of the image labels at test time. Such an interactive scenario offers an interesting tradeoff between label accuracy and manual labeling effort. The structured models are used to decide which labels should be set by the user, and transfer the user input to more accurate predictions on other image labels. We also apply our models to attribute-based image classification, where attribute predictions of a test image are mapped to class probabilities by means of a given attribute-class mapping. Experimental results on three publicly available benchmark datasets show that in all scenarios our structured models lead to more accurate predictions, and leverage user input much more effectively than state-of-the-art independent models. Thomas Mensink, Jakob Verbeek, Gabriela Csurka |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Distance-Based Image Classification: Generalizing to New Classes at Near-Zero CostabstractWe study large-scale image classification methods that can incorporate new classes and training images continuously over time at negligible cost. To this end, we consider two distance-based classifiers, the k-nearest neighbor (k-NN) and nearest class mean (NCM) classifiers, and introduce a new metric learning approach for the latter. We also introduce an extension of the NCM classifier to allow for richer class representations. Experiments on the ImageNet 2010 challenge dataset, which contains over 10(6) training images of 1,000 classes, show that, surprisingly, the NCM classifier compares favorably to the more flexible k-NN classifier. Moreover, the NCM performance is comparable to that of linear SVMs which obtain current state-of-the-art performance. Experimentally, we study the generalization performance to classes that were not used to learn the metrics. Using a metric learned on 1,000 classes, we show results for the ImageNet-10K dataset which contains 10,000 classes, and obtain performance that is competitive with the current state-of-the-art while being orders of magnitude faster. Furthermore, we show how a zero-shot class prior based on the ImageNet hierarchy can improve performance when few training images are available. Thomas Mensink, Jakob Verbeek, Florent Perronnin, Gabriela Csurka |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Metric Learning for Large Scale Image Classification: Generalizing to New Classes at Near-Zero Cost
Thomas Mensink, Jakob Verbeek, Florent Perronnin, Gabriela Csurka |
ECCV (2) | 4 |
| 2012 | Images as sets of locally weighted features
Teófilo Emídio de Campos, Gabriela Csurka, Florent Perronnin |
Comput. Vis. Image Underst. | 2 |
| 2011 | Combining Visible and Near-Infrared Cues for image CategorisationabstractStandard digital cameras are sensitive to radiation in the near-infrared domain, but this additional cue is in general discarded. In this paper, we consider the scene categorisation problem in the context of images where both standard visible RGB channels and near infrared information are available. Using efficient local patch-based Fisher Vector image representations, we show based on thorough experimental studies the benefit of using this new type of data. We investigate which image descriptors are relevant, and how to best combine them. In particular, our experiments show that when combining texture and colour information, computed on visible and near-infrared channels, late fusion is the best performing strategy and outperforms the state-of-the-art categorisation methods on RGB-only data. 1 Neda Salamati, Diane Larlus, Gabriela Csurka |
BMVC | 3 |
| 2011 | Learning structured prediction models for interactive image labelingabstractWe propose structured models for image labeling that take into account the dependencies among the image labels explicitly. These models are more expressive than independent label predictors, and lead to more accurate predictions. While the improvement is modest for fully-automatic image labeling, the gain is significant in an interactive scenario where a user provides the value of some of the image labels. Such an interactive scenario offers an interesting trade-off between accuracy and manual labeling effort. The structured models are used to decide which labels should be set by the user, and transfer the user input to more accurate predictions on other image labels. We also apply our models to attribute-based image classification, where attribute predictions of a test image are mapped to class probabilities by means of a given attribute-class mapping. In this case the structured models are built at the attribute level. We also consider an interactive system where the system asks a user to set some of the attribute values in order to maximally improve class prediction performance. Experimental results on three publicly available benchmark data sets show that in all scenarios our structured models lead to more accurate predictions, and leverage user input much more effectively than state-of-the-art independent models. Thomas Mensink, Jakob Verbeek, Gabriela Csurka |
CVPR | 3 |
| 2011 | Assessing the aesthetic quality of photographs using generic image descriptorsabstractIn this paper, we automatically assess the aesthetic properties of images. In the past, this problem has been addressed by hand-crafting features which would correlate with best photographic practices (e.g. “Does this image respect the rule of thirds?”) or with photographic techniques (e.g. “Is this image a macro?”). We depart from this line of research and propose to use generic image descriptors to assess aesthetic quality. We experimentally show that the descriptors we use, which aggregate statistics computed from low-level local features, implicitly encode the aesthetic properties explicitly used by state-of-the-art methods and outperform them by a significant margin. Luca Marchesotti, Florent Perronnin, Diane Larlus, Gabriela Csurka |
ICCV | 4 |
| 2011 | Semantic combination of textual and visual information in multimedia retrievalabstractThe goal of this paper is to introduce a set of techniques we call semantic combination in order to efficiently fuse text and image retrieval systems in the context of multimedia information access. These techniques emerge from the observation that image and textual queries are expressed at different semantic levels and that a single image query is often ambiguous. Overall, the semantic combination techniques overcome a conceptual barrier rather than a technical one: these methods can be seen as a combination of late fusion and image reranking. Albeit simple, this approach has not been used yet. We assess the proposed techniques against late and cross-media fusion using 4 different ImageCLEF datasets. Compared to late fusion, performances significantly increase on two datasets and remain similar on the two other ones. Stéphane Clinchant, Julien Ah-Pine, Gabriela Csurka |
ICMR | 3 |
| 2011 | An Efficient Approach to Semantic Segmentation
Gabriela Csurka, Florent Perronnin |
Int. J. Comput. Vis. | 1 |
| 2011 | Building look & feel concept models from color combinations - With applications in image classification, retrieval, and color transfer
Gabriela Csurka, Sandra Skaff, Luca Marchesotti, Craig Saunders |
Vis. Comput. | 1 |
| 2010 | Trans Media Relevance Feedback for Image AutoannotationabstractAutomatic image annotation is an important tool for keyword-based image retrieval, providing a textual index for non-annotated images. Many image auto annotation methods are based on visual similarity between images to be annotated and images in a training corpus. The annotations of the most similar training images are transferred to the image to be annotated. In this paper we consider using also similarities among the training images, both visual and textual, to derive pseudo relevance models, as well as crossmedia relevance models. We extend a recent state-of-the-art image annotation model to incorporate this information. On two widely used datasets (COREL and IAPR) we show experimentally that the pseudo-relevance models improve the annotation accuracy. Thomas Mensink, Jakob Verbeek, Gabriela Csurka |
BMVC | 3 |
| 2009 | Hierarchical Image-Region Labeling via Structured LearningabstractWe present a graphical model that encodes hierarchical constraints for classifying image regions at multiple scales. We show that inference can be performed efficiently and exactly, rendering it amenable to structured learning. Our model is parametrised using the outputs of a series of first-order classifiers, meaning that it learns which classifiers are useful at different scales, as well as the relationships between classifiers across scales. Example results Correct labeling, using bounding-boxes from VOC2007 (1 − ∆ = 1): Baseline, using no second-order features (1 − ∆ = 0.566): Our model The ‘nodes ’ of our graphical model correspond to overlapping image regions: Second-order features without learning (1 − ∆ = 0.551): Learning of all features (1 − ∆ = 0.770): Colour-code for labels: Edges are formed by connecting nodes at different scales: we connect two nodes precisely when the corresponding image regions overlap at adjacent scales, so that our graphical model forms a quad-tree. First-order (node) features Our image features are based on those from [2], in which image-level, region-level, and patch-level classifiers are proposed. We use all classifiers at all scales (Pr,label is the probability that the region r is labeled label): Φ nodes (r, label) = (0,..., P 1 r,label,..., 0)... (0,..., P features for first classifier n r,label,..., 0) features for nth classifier Thus we learn which classifiers are useful at which scales. Hierarchical constraints We want to ban inconsistent assignments at different scales: aeroplane Julian J. McAuley, Teófilo Emídio de Campos, Gabriela Csurka, Florent Perronnin |
BMVC | 3 |
| 2009 | A framework for visual saliency detection with applications to image thumbnailingabstractWe propose a novel framework for visual saliency detection based on a simple principle: images sharing their global visual appearances are likely to share similar salience. Assuming that an annotated image database is available, we first retrieve the most similar images to the target image; secondly, we build a simple classifier and we use it to generate saliency maps. Finally, we refine the maps and we extract thumbnails. We show that in spite of its simplicity, our framework outperforms state-of-the-art approaches. Another advantage is its ability to deal with visual pop-up and application/task-driven saliency, if appropriately annotated images are available. Luca Marchesotti, Claudio Cifarelli, Gabriela Csurka |
ICCV | 3 |
| 2009 | Crossing textual and visual content in different application scenarios
Julien Ah-Pine, Marco Bressan 0003, Stéphane Clinchant, Gabriela Csurka, Yves Hoppenot, Jean-Michel Renders |
Multim. Tools Appl. | 4 |
| 2009 | Introduction to the special issue on "metadata mining for image understanding"
Gabriela Csurka, Katerina Pastra |
Multim. Tools Appl. | 1 |
| 2008 | A Simple High Performance Approach to Semantic SegmentationabstractWe propose a simple approach to semantic image segmentation. Our system scores low-level patches according to their class relevance, propagates these posterior probabilities to pixels and uses low-level segmentation to guide the semantic segmentation. The two main contributions of this paper are as follows. First, for the patch scoring, we describe each patch with a high-level descriptor based on the Fisher kernel and use a set of linear classifiers. While the Fisher kernel methodology was shown to lead to high accuracy for image classification, it has not been applied to the segmentation problem. Second, we use global image classifiers to take into account the context of the objects to be segmented. If an image as a whole is unlikely to contain an object class, then the corresponding class is not considered in the segmentation pipeline. This increases the classification accuracy and reduces the computational cost. We will show that despite its apparent simplicity, this system provides above state-of-the-art performance on the PASCAL VOC 2007 dataset and state-of-the-art performance on the MSRC 21 dataset. 1 Gabriela Csurka, Florent Perronnin |
BMVC | 1 |
| 2006 | Adapted Vocabularies for Generic Visual Categorization
Florent Perronnin, Christopher R. Dance, Gabriela Csurka, Marco Bressan 0003 |
ECCV (4) | 3 |
| 2006 | Categorization in multiple category systemsabstractWe explore the situation in which documents have to be categorized into more than one category system, a situation we refer to as multiple-view categorization. More particularly, we address the case where two different categorizers have already been built based on non-necessarily identical training sets, each one labeled using one category system. On the top of these categorizers considered as black-boxes, we propose some algorithms able to exploit a third training set containing a few examples annotated in both category systems. Such a situation arises for example in large companies where incoming mails have to be routed to several departments, each one relying on its own category system. We focus here on exploiting possible dependencies between category systems in order to refine the categorization decisions made by categorizers trained independently on different category systems. After a description of the multiple categorization problem, we present several possible solutions, based either on a categorization or reweighting approach, and compare them on real data. Lastly, we show how the multimedia categorization problem can be cast as a multiple categorization problem and assess our methods in this framework. Jean-Michel Renders, Éric Gaussier, Cyril Goutte, François Pacull, Gabriela Csurka |
ICML | 5 |
| 2002 | Toward generic image dewatermarking?abstractDuring the last decade, a significant effort has been put into designing watermarking algorithms. The watermarking community now needs some advanced attacks and fair benchmarks in order to compare the performances of different watermarking technologies. Moreover, attacks permit the weaknesses of an algorithm to be found and consequently trigger further research in order to overcome the problem. These ideas motivated the creation of the European Certimark project. After a short definition of dewatermarking, we present an original attack based on self similarities. This attack is then put to the test with three different publicly available watermarking tools. Finally, we discuss briefly the feasibility of a generic attack i.e. a dewatermarking attack which should succeed in removing whatever watermarks have been inserted by whatever watermarking tools. Christian Rey, Gwenaël J. Doërr, Gabriela Csurka, Jean-Luc Dugelay |
ICIP (3) | 3 |
| 2000 | Stereo Calibration from Rigid MotionsabstractWe describe a method for calibrating a stereo pair of cameras using general or planar motions. The method consists of upgrading a 3D projective representation to affine and to Euclidean without any knowledge, neither about the motion parameters nor about the 3D layout. We investigate the algebraic properties relating projective representation to the plane at infinity and to the intrinsic camera parameters when the camera pair is considered as a moving rigid body. We show that all the computations can be carried out using standard linear resolutions techniques. An error analysis reveals the relative importance of the various steps of the calibration process: projective-to-affine and affine-to-metric upgrades. Extensive experiments performed with calibrated and natural data confirm the error analysis as well as the sensitivity study performed with simulated data. Radu Horaud, Gabriela Csurka, David Demirdjian |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1999 | A Bayesian Approach To Spread Spectrum Watermark Detection and Secure Copyright Protection for Digital Image LibrariesabstractDigital watermarks have been proposed as a method for discouraging illicit copying and distribution of copyrighted material, and to create secure digital image libraries by adding to images copyright and user-right information. Using a robust digital watermark to detect and trace copyright violations has therefore lot of interest. This paper describes an approach to embedding a digital watermark using the Fourier transform. The paper also addresses the difficult problem of oblivious watermark detection. It is shown that, for the CDMA spread spectrum signal described in the paper, it is still possible to positively detect the presence of a watermark without being able to decode it (and even infer the number of bits contained in the watermark) given only the key used to generate it. Finally, through experimental results the usefulness of such measure is shown. Joseph Ó Ruanaidh, Gabriela Csurka |
CVPR | 2 |
| 1999 | Direct Identification of Moving Objects and Background from 2D Motion ModelsabstractThis paper presents the dynamic scene analysis part of an original and consistent framework to video partitioning into shots, camera motion estimation and multiple motion analysis with a view to content-based video indexing. All the information parts required to achieve these different goals result from handling the apparent motion within consecutive image pairs. Within each extracted shot, a binary segmentation of the image is performed into regions whose motion either conforms or not to the 2D estimated dominant motion represented by a quadratic motion model. This paper focuses on a low-cost method based on projective geometry criteria to distinguish non-conforming regions generated by really moving objects from static ones in the scene. The proposed algorithm is validated on a variety of real image sequences. Gabriela Csurka, Patrick Bouthemy |
ICCV | 1 |
| 1999 | Finding the Collineation between Two Projective Reconstructions
Gabriela Csurka, David Demirdjian, Radu Horaud |
Comput. Vis. Image Underst. | 1 |
| 1999 | Algebraic and Geometric Tools to Compute Projective and Permutation InvariantsabstractStudies the computation of projective invariants in pairs of images from uncalibrated cameras and presents a detailed study of the projective and permutation invariants for configurations of points and/or lines. Two basic computational approaches are given, one algebraic and one geometric. In each case, invariants are computed in projective space or directly from image measurements. Finally, we develop combinations of those projective invariants which are insensitive to permutations of the geometric primitives of each of the basic configurations. Gabriela Csurka, Olivier D. Faugeras |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | Autocalibration in the Presence of Critical MotionsabstractAutocalibration is a difficult problem. Not only is its computation very noisesensitive, but there also exist many critical motions that prevent the estimation of some of the camera parameters. When a "stratified" approach is considered, affine and Euclidean calibration are computed in separate steps and it is possible to see that a part of these ambiguities occur during affine-toEuclidean calibration. This paper studies the affine-to-Euclidean step in detail using the real Jordan decomposition of the infinite homography. It gives a new way to compute the autocalibration and analyzes the effects of critical motions on the computation of internal parameters. Finally, it shows that in some cases, it is possible to obtain complete calibration in the presence of critical motions. Keywords : Autocalibration, critical motions, real Jordan decomposition, affine calibration, infinite homography. 1 Introduction This article raises the problem of autocalibration of a camera undergoing rigid mo... David Demirdjian, Gabriela Csurka, Radu Horaud |
BMVC | 2 |
| 1998 | Projective Translations and Affine Stereo CalibrationabstractThis paper investigates the structure of projective translations-rigid translations expressed as homographies in projective space. A seven parameter representation is proposed, which explicitly represents the geometric entities constraining and defining the translation. A practical algebraic method for estimating these parameters is developed. It provides affine calibration of a stereo rig, determines the translation axis, and allows projective translations to be composed. The practical effectiveness of the calibration is evaluated on synthetic and real image data. Andreas Ruf, Gabriela Csurka, Radu Horaud |
CVPR | 2 |
| 1998 | Closed-Form Solutions for the Euclidean Calibration of a Stereo Rig
Gabriela Csurka, David Demirdjian, Andreas Ruf, Radu Horaud |
ECCV (1) | 1 |
| 1998 | Self-Calibration and Euclidean Reconstruction Using Motions of a Stereo RigabstractThis paper describes a method to upgrade projective reconstruction to affine and to metric reconstructions using rigid general motions of a stereo rig. We make clear the algebraic relationships between projective reconstruction, the plane at infinity (affine reconstruction), camera calibration, and metric reconstruction. We show that all the computations can be carried out using standard linear resolution methods and that these methods compare favorably with nonlinear optimization, methods in the presence of Gaussian noise. We carry out a theoretical error analysis which quantify the relative importance of the accuracies of projective-to-affine conversion and affine-to-Euclidean conversion. Experiments with with real data are consistent with the theoretical error analysis and with a sensitivity analysis performed with simulated data. Radu Horaud, Gabriela Csurka |
ICCV | 2 |
| 1998 | 3-D Reconstruction of Urban Scenes from Image Sequences
Olivier D. Faugeras, Luc Robert, Stéphane Laveau, Gabriela Csurka, Cyril Zeller, Cyrille Gauclin, Imad Zoghlami |
Comput. Vis. Image Underst. | 4 |
| 1998 | Computing three dimensional project invariants from a pair of images using the Grassmann-Cayley algebra
Gabriela Csurka, Olivier D. Faugeras |
Image Vis. Comput. | 1 |
| 1997 | Computing Protective and Permutation Invariants of Points and Lines
Gabriela Csurka, Olivier D. Faugeras |
CAIP | 1 |
| 1997 | Characterizing the Uncertainty of the Fundamental Matrix
Gabriela Csurka, Cyril Zeller, Zhengyou Zhang, Olivier D. Faugeras |
Comput. Vis. Image Underst. | 1 |
| 1997 | A Comparison of Projective Reconstruction Methods for Pairs of Views
Charlie Rothwell, Olivier D. Faugeras, Gabriela Csurka |
Comput. Vis. Image Underst. | 3 |
| 1995 | A Comparison of Projective Reconstruction Methods for Pairs of ViewsabstractRecently, different approaches for uncalibrated stereo have been suggested which permit projective reconstruction from multiple views. These use weak calibration which is represented by the epipolar geometry, and so no knowledge of the intrinsic or extrinsic camera parameters is required. We consider projective reconstructions from pairs of views, and compare a number of the available methods. Consequently we conclude which methods are most likely to be of use in applications that are dependent on 3D uncalibrated reconstructions.> Charlie Rothwell, Gabriela Csurka, Olivier D. Faugeras |
ICCV | 2 |