VLDB 2026 Research / reviewers in the wild / expert
Giorgos Tolias
dblp:09/4652
· DBLP profile ↗
57ranked-venue papers
13as first author
18since 2021 · last 2025
0000-0002-9570-3870ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 12 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 7 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ILIAS: Instance-Level Image retrieval At ScaleabstractThis work introduces ILIAS, a new test dataset for Instance-Level Image retrieval At Scale. It is designed to evaluate the ability of current and future foundation models and retrieval techniques to recognize particular objects. The key benefits over existing datasets include large scale, domain diversity, accurate ground truth, and a performance that is far from saturated. ILIAS includes query and positive images for 1,000 object instances, manually collected to capture challenging conditions and diverse domains. Large-scale retrieval is conducted against 100 million distractor images from YFCC100M. To avoid false negatives without extra annotation effort, we include only query objects confirmed to have emerged after 2014, i.e. the compilation date of YFCC100M. An extensive benchmarking is performed with the following observations: i) models fine-tuned on specific domains, such as landmarks or products, excel in that domain but fail on ILIAS ii) learning a linear adaptation layer using multi-domain class supervision results in performance improvements, especially for vision-language models iii) local descriptors in retrieval re-ranking are still a key ingredient, especially in the presence of severe background clutter iv) the text-to-image performance of the vision-language foundation models is surprisingly close to the corresponding image-to-image case. website: https://vrg.fel.cvut.cz/ilias/ Giorgos Kordopatis-Zilos, Vladan Stojnic, Anna Manko, Pavel Suma, Nikolaos-Antonios Ypsilantis, Nikos Efthymiadis, Zakaria Laskar, Jiri Matas, Ondrej Chum, Giorgos Tolias |
CVPR | 10 |
| 2025 | A Dataset for Semantic Segmentation in the Presence of UnknownsabstractBefore deployment in the real-world deep neural networks require thorough evaluation of how they handle both knowns, inputs represented in the training data, and unknowns (anomalies). This is especially important for scene understanding tasks with safety critical applications, such as in autonomous driving. Existing datasets allow evaluation of only knowns or unknowns - but not both, which is required to establish "in the wild" suitability of deep neural network models. To bridge this gap, we propose a novel anomaly segmentation dataset, ISSU, that features a diverse set of anomaly inputs from cluttered real-world environments. The dataset is twice larger than existing anomaly segmentation datasets, and provides a training, validation and test set for controlled in-domain evaluation. The test set consists of a static and temporal part, with the latter comprised of videos. The dataset provides annotations for both closed-set (knowns) and anomalies, enabling closed-set and open-set evaluation. The dataset covers diverse conditions, such as domain and cross-sensor shift, illumination variation and allows ablation of anomaly detection methods with respect to these variations. Evaluation results of current state-of-the-art methods confirm the need for improvements especially in domain-generalization, small and large object segmentation. The code and the dataset are available at https://github.com/vojirt/benchmark_issu. Zakaria Laskar, Tomás Vojír, Matej Grcic, Iaroslav Melekhov, Shankar Gangisetty, Juho Kannala, Jiri Matas, Giorgos Tolias, C. V. Jawahar |
CVPR | 8 |
| 2025 | LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic SegmentationabstractWe propose a training-free method for open-vocabulary semantic segmentation using Vision-and-Language Models (VLMs). Our approach enhances the initial per-patch predictions of VLMs through label propagation, which jointly optimizes predictions by incorporating patch-to-patch relationships. Since VLMs are primarily optimized for cross-modal alignment and not for intra-modal similarity, we use a Vision Model (VM) that is observed to better capture these relationships. We address resolution limitations inherent to patch-based encoders by applying label propagation at the pixel level as a refinement step, significantly improving segmentation accuracy near class boundaries. Our method, called LPOSS+, performs inference over the entire image, avoiding window-based processing and thereby capturing contextual interactions across the full image. LPOSS+ achieves state-of-the-art performance among training-free methods, across a diverse set of datasets. Code: https://github.com/vladan-stojnic/LPOSS Vladan Stojnic, Yannis Kalantidis, Jiri Matas, Giorgos Tolias |
CVPR | 4 |
| 2025 | LOCORE: Image Re-ranking with Long-Context Sequence ModelingabstractWe introduce LoCoRe, Long-Context Re-ranker, a model that takes as input local descriptors corresponding to an image query and a list of gallery images and outputs similarity scores between the query and each gallery image. This model is used for image retrieval, where typically a first ranking is performed with an efficient similarity measure, and then a shortlist of top-ranked images is re-ranked based on a more fine-grained similarity model. Compared to existing methods that perform pair-wise similarity estimation with local descriptors or list-wise re-ranking with global descriptors, LoCoRe is the first method to perform list-wise re-ranking with local descriptors. To achieve this, we leverage efficient long-context sequence models to effectively capture the dependencies between query and gallery images at the local-descriptor level. During testing, we process long shortlists with a sliding window strategy that is tailored to overcome the context size limitations of sequence models. Our approach achieves superior performance compared with other re-rankers on established image retrieval benchmarks of landmarks ($\mathcal{R}{\text{Oxf}}$ and $\mathcal{R}{\text{Par}}$), products (SOP), fashion items (In-Shop), and bird species (CUB-200) while having comparable latency to the pair-wise local descriptor re-rankers. Zilin Xiao, Pavel Suma, Ayush Sachdeva, Hao-Jen Wang, Giorgos Kordopatis-Zilos, Giorgos Tolias, Vicente Ordonez |
CVPR | 6 |
| 2025 | Processing and Acquisition Traces in Visual Encoders: What Does CLIP Know About Your Camera?
Ryan Ramos, Vladan Stojnic, Giorgos Kordopatis-Zilos, Yuta Nakashima, Giorgos Tolias, Noa Garcia |
ICCV | 5 |
| 2025 | Instance-Level Composed Image RetrievalabstractThe progress of composed image retrieval (CIR), a popular research direction in image retrieval, where a combined visual and textual query is used, is held back by the absence of high-quality training and evaluation data. We introduce a new evaluation dataset, i-CIR, which, unlike existing datasets, focuses on an instance-level class definition. The goal is to retrieve images that contain the same particular object as the visual query, presented under a variety of modifications defined by textual queries. Its design and curation process keep the dataset compact to facilitate future research, while maintaining its challenge—comparable to retrieval among more than 40M random distractors—through a semi-automated selection of hard negatives. To overcome the challenge of obtaining clean, diverse, and suitable training data, we leverage pre-trained vision-and-language models (VLMs) in a training-free approach called BASIC. The method separately estimates query-image-to-image and query-text-to-image similarities, performing late fusion to upweight images that satisfy both queries, while down-weighting those that exhibit high similarity with only one of the two. Each individual similarity is further improved by a set of components that are simple and intuitive. BASIC sets a new state of the art on i-CIR but also on existing CIR datasets that follow a semantic-level class definition. Project page: https://vrg.fel.cvut.cz/icir/. Bill Psomas, George Retsinas, Nikos Efthymiadis, Panagiotis Paraskevas Filntisis, Yannis Avrithis, Petros Maragos, Ondrej Chum, Giorgos Tolias |
NeurIPS | 8 |
| 2025 | Composed Image Retrieval for Training-FREE DOMain ConversionabstractThis work addresses composed image retrieval in the context of domain conversion, where the content of a query image is retrieved in the domain specified by the query text. We show that a strong vision-language model provides sufficient descriptive power without additional training. The query image is mapped to the text input space using textual inversion. Unlike common practice that invert in the continuous space of text tokens, we use the discrete word space via a nearest-neighbor search in a text vocabulary. With this inversion, the image is softly mapped across the vocabulary and is made more robust using retrieval-based augmentation. Database images are retrieved by a weighted ensemble of text queries combining mapped words with the domain text. Our method outperforms prior art by a large margin on standard and newly introduced benchmarks. Code: https://github.com/NikosEfth/freedom Nikos Efthymiadis, Bill Psomas, Zakaria Laskar, Konstantinos Karantzalos, Yannis Avrithis, Ondrej Chum, Giorgos Tolias |
WACV | 7 |
| 2025 | Crafting Distribution Shifts for Validation and Training in Single Source Domain GeneralizationabstractSingle-source domain generalization attempts to learn a model on a source domain and deploy it to unseen target domains. Limiting access only to source domain data imposes two key challenges - how to train a model that can generalize and how to verify that it does. The standard practice of validation on the training distribution does not accurately reflect the model's generalization ability, while validation on the test distribution is a malpractice to avoid. In this work, we construct an independent validation set by transforming source domain images with a comprehensive list of augmentations, covering a broad spectrum of potential distribution shifts in target domains. We demonstrate a high correlation between validation and test performance for multiple methods and across various datasets. The proposed validation achieves a relative accuracy improvement over the standard validation equal to 15.4% or 1.6% when used for method selection or learning rate tuning, respectively. Furthermore, we introduce a novel family of methods that increase the shape bias through enhanced edge maps. To benefit from the augmentations during training and preserve the independence of the validation set, a k-fold validation process is designed to separate the augmentation types used in training and validation. The method that achieves the best performance on the augmented validation is selected from the proposed family. It achieves state-of-the-art performance on various standard benchmarks. Code at: https://github.com/NikosEfth/crafting-shifts Nikos Efthymiadis, Giorgos Tolias, Ondrej Chum |
WACV | 2 |
| 2024 | Label Propagation for Zero-shot Classification with Vision-Language ModelsabstractVision-Language Models (VLMs) have demonstrated im-pressive performance on zero-shot classification, i.e. classi-fication when provided merely with a list of class names. In this paper, we tackle the case of zero-shot classification in the presence of unlabeled data. We leverage the graph structure of the unlabeled data and introduce ZLaP, a method based on label propagation (LP) that utilizes geodesic distances for classification. We tailor LP to graphs containing both text and image features and further pro-pose an efficient method for performing inductive infer-ence based on a dual solution and a sparsification step. We perform extensive experiments to evaluate the effectiveness of our method on 14 common datasets and show that ZLaP outperforms the latest related works. Code: https://github.com/vladan-stojnic/ZLaP Vladan Stojnic, Yannis Kalantidis, Giorgos Tolias |
CVPR | 3 |
| 2024 | AMES: Asymmetric and Memory-Efficient Similarity Estimation for Instance-Level Retrieval
Pavel Suma, Giorgos Kordopatis-Zilos, Ahmet Iscen, Giorgos Tolias |
ECCV (59) | 4 |
| 2024 | Composed Image Retrieval for Remote SensingabstractThis work introduces composed image retrieval to remote sensing. It allows to query a large image archive by image examples alternated by a textual description, enriching the descriptive power over unimodal queries, either visual or textual. Various attributes can be modified by the textual part, such as shape, color, or context. A novel method fusing image-to-image and text-to-image similarity is introduced. We demonstrate that a vision-language model possesses sufficient descriptive power and no further learning step or training data are necessary. We present a new evaluation benchmark focused on color, context, density, existence, quantity, and shape modifications. Our work not only sets the state-of-the-art for this task, but also serves as a foundational step in addressing a gap in the field of remote sensing image retrieval. Code at: https://github.com/billpsomas/rscir. Bill Psomas, Ioannis Kakogeorgiou, Nikos Efthymiadis, Giorgos Tolias, Ondrej Chum, Yannis Avrithis, Konstantinos Karantzalos |
IGARSS | 4 |
| 2024 | Training Ensembles with Inliers and Outliers for Semi-supervised Active LearningabstractDeep active learning in the presence of outlier examples poses a realistic yet challenging scenario. Acquiring unlabeled data for annotation requires a delicate balance between avoiding outliers to conserve the annotation budget and prioritizing useful inlier examples for effective training. In this work, we present an approach that leverages three highly synergistic components, which are identified as key ingredients: joint classifier training with inliers and outliers, semi-supervised learning through pseudo-labeling, and model ensembling. Our work demonstrates that ensembling significantly enhances the accuracy of pseudolabeling and improves the quality of data acquisition. By enabling semi-supervision through the joint training process, where outliers are properly handled, we observe a substantial boost in classifier accuracy through the use of all available unlabeled examples. Notably, we reveal that the integration of joint training renders explicit outlier detection unnecessary; a conventional component for acquisition in prior work. The three key components align seamlessly with numerous existing approaches. Through empirical evaluations, we showcase that their combined use leads to a performance increase. Remarkably, despite its simplicity, our proposed approach outperforms all other methods in terms of performance. Code: https://github.com/vladan-stojnic/active-outliers Vladan Stojnic, Zakaria Laskar, Giorgos Tolias |
WACV | 3 |
| 2024 | The 2023 video similarity dataset and challenge
Ed Pizzi, Giorgos Kordopatis-Zilos, Hiral Patel, Gheorghe Postelnicu, Sugosh Nagavara Ravindra, Symeon Papadopoulos, Giorgos Tolias, Matthijs Douze |
Comput. Vis. Image Underst. | 8 |
| 2024 | HSCNet++: Hierarchical Scene Coordinate Classification and Regression for Visual Localization with TransformerabstractAbstract Visual localization is critical to many applications in computer vision and robotics. To address single-image RGB localization, state-of-the-art feature-based methods match local descriptors between a query image and a pre-built 3D model. Recently, deep neural networks have been exploited to regress the mapping between raw pixels and 3D coordinates in the scene, and thus the matching is implicitly performed by the forward pass through the network. However, in a large and ambiguous environment, learning such a regression task directly can be difficult for a single network. In this work, we present a new hierarchical scene coordinate network to predict pixel scene coordinates in a coarse-to-fine manner from a single RGB image. The proposed method, which is an extension of HSCNet, allows us to train compact models which scale robustly to large environments. It sets a new state-of-the-art for single-image localization on the 7-Scenes, 12-Scenes, Cambridge Landmarks datasets, and the combined indoor scenes. Shuzhe Wang, Zakaria Laskar, Iaroslav Melekhov, Yi Zhao 0014, Giorgos Tolias, Juho Kannala |
Int. J. Comput. Vis. | 6 |
| 2023 | Test-time Training for Matching-based Video Object SegmentationabstractThe video object segmentation (VOS) task involves the segmentation of an object over time based on a single initial mask. Current state-of-the-art approaches use a memory of previously processed frames and rely on matching to estimate segmentation masks of subsequent frames. Lacking any adaptation mechanism, such methods are prone to test-time distribution shifts. This work focuses on matching-based VOS under distribution shifts such as video corruptions, stylization, and sim-to-real transfer. We explore test-time training strategies that are agnostic to the specific task as well as strategies that are designed specifically for VOS. This includes a variant based on mask cycle consistency tailored to matching-based VOS methods. The experimental results on common benchmarks demonstrate that the proposed test-time training yields significant improvements in performance. In particular for the sim-to-real scenario and despite using only a single test video, our approach manages to recover a substantial portion of the performance gain achieved through training on real videos. Additionally, we introduce DAVIS-C, an augmented version of the popular DAVIS test set, featuring extreme distribution shifts like image-/video-level corruptions and stylizations. Our results illustrate that test-time training enhances performance even in these challenging cases. Juliette Bertrand, Giorgos Kordopatis-Zilos, Yannis Kalantidis, Giorgos Tolias |
NeurIPS | 4 |
| 2023 | Large-to-small Image Resolution Asymmetry in Deep Metric LearningabstractDeep metric learning for vision is trained by optimizing a representation network to map (non-)matching image pairs to (non-)similar representations. During testing, which typically corresponds to image retrieval, both database and query examples are processed by the same network to obtain the representation used for similarity estimation and ranking. In this work, we explore an asymmetric setup by light-weight processing of the query at a small image resolution to enable fast representation extraction. The goal is to obtain a network for database examples that is trained to operate on large resolution images and benefits from fine-grained image details, and a second network for query examples that operates on small resolution images but preserves a representation space aligned with that of the database network. We achieve this with a distillation approach that transfers knowledge from a fixed teacher network to a student via a loss that operates per image and solely relies on coupled augmentations without the use of any labels. In contrast to prior work that explores such asymmetry from the point of view of different network architectures, this work uses the same architecture but modifies the image resolution. We conclude that resolution asymmetry is a better way to optimize the performance/efficiency trade-off than architecture asymmetry. Evaluation is performed on three standard deep metric learning benchmarks, namely CUB200, Cars196, and SOP. Code: https://github.com/pavelsuma/raml Pavel Suma, Giorgos Tolias |
WACV | 2 |
| 2022 | Recall@k Surrogate Loss with Large Batches and Similarity MixupabstractThis work focuses on learning deep visual representation models for retrieval by exploring the interplay between a new loss function, the batch size, and a new regularization approach. Direct optimization, by gradient descent, of an evaluation metric, is not possible when it is nondifferentiable, which is the case for recall in retrieval. A differentiable surrogate loss for the recall is proposed in this work. Using an implementation that sidesteps the hardware constraints of the GPU memory, the method trains with a very large batch size, which is essential for metrics computed on the entire retrieval database. It is assisted by an efficient mixup regularization approach that operates on pairwise scalar similarities and virtually increases the batch size further. The suggested method achieves state-of-the-art performance in several image retrieval benchmarks when used for deep metric learning. For instance-level recognition, the method outperforms similar approaches that train using an approximation of average precision. Giorgos Tolias, Jiri Matas |
CVPR | 2 |
| 2022 | Edge Augmentation for Large-Scale Sketch Recognition without SketchesabstractThis work addresses scaling up the sketch classification task into a large number of categories. Collecting sketches for training is a slow and tedious process that has so far precluded any attempts to large-scale sketch recognition. We overcome the lack of training sketch data by exploiting labeled collections of natural images that are easier to obtain. To bridge the domain gap we present a novel augmentation technique that is tailored to the task of learning sketch recognition from a training set of natural images. Randomization is introduced in the parameters of edge detection and edge selection. Natural images are translated to a pseudo-novel domain called "randomized Binary Thin Edges" (rBTE), which is used as a training domain instead of natural images. The ability to scale up is demonstrated by training CNN-based sketch recognition of more than 2.5 times larger number of categories than used previously. For this purpose, a dataset of natural images from 874 categories is constructed by combining a number of popular computer vision datasets. The categories are selected to be suitable for sketch recognition. To estimate the performance, a subset of 393 categories with sketches is also collected. Nikos Efthymiadis, Giorgos Tolias, Ondrej Chum |
ICPR | 2 |
| 2020 | Graph Convolutional Networks for Learning with Few Clean and Many Noisy Labels
Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum, Cordelia Schmid |
ECCV (10) | 2 |
| 2020 | Learning and Aggregating Deep Local Descriptors for Instance-Level Recognition
Giorgos Tolias, Tomás Jenícek, Ondrej Chum |
ECCV (1) | 1 |
| 2019 | Label Propagation for Deep Semi-Supervised LearningabstractSemi-supervised learning is becoming increasingly important because it can combine data carefully labeled by humans with abundant unlabeled data to train deep neural networks. Classic methods on semi-supervised learning that have focused on transductive learning have not been fully exploited in the inductive framework followed by modern deep learning. The same holds for the manifold assumption-that similar examples should get the same prediction. In this work, we employ a transductive label propagation method that is based on the manifold assumption to make predictions on the entire dataset and use these predictions to generate pseudo-labels for the unlabeled data and train a deep neural network. At the core of the transductive method lies a nearest neighbor graph of the dataset that we create based on the embeddings of the same network. Therefore our learning process iterates between these two steps. We improve performance on several datasets especially in the few labels regime and show that our work is complementary to current state of the art. Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum |
CVPR | 2 |
| 2019 | Explicit Spatial Encoding for Deep Local DescriptorsabstractWe propose a kernelized deep local-patch descriptor based on efficient match kernels of neural network activations. Response of each receptive field is encoded together with its spatial location using explicit feature maps. Two location parametrizations, Cartesian and polar, are used to provide robustness to a different types of canonical patch misalignment. Additionally, we analyze how the conventional architecture, i.e. a fully connected layer attached after the convolutional part, encodes responses in a spatially variant way. In contrary, explicit spatial encoding is used in our descriptor, whose potential applications are not limited to local-patches. We evaluate the descriptor on standard benchmarks. Both versions, encoding 32 × 32 or 64 × 64 patches, consistently outperform all other methods on all benchmarks. The number ofparameters of the model is independent of the input patch resolution. Arun Mukundan, Giorgos Tolias, Ondrej Chum |
CVPR | 2 |
| 2019 | Targeted Mismatch Adversarial Attack: Query With a Flower to Retrieve the TowerabstractAccess to online visual search engines implies sharing of private user content - the query images. We introduce the concept of targeted mismatch attack for deep learning based retrieval systems to generate an adversarial image to conceal the query image. The generated image looks nothing like the user intended query, but leads to identical or very similar retrieval results. Transferring attacks to fully unseen networks is challenging. We show successful attacks to partially unknown systems, by designing various loss functions for the adversarial image construction. These include loss functions, for example, for unknown global pooling operation or unknown input resolution by the retrieval system. We evaluate the attacks on standard retrieval benchmarks and compare the results retrieved with the original and adversarial image. Giorgos Tolias, Filip Radenovic, Ondrej Chum |
ICCV | 1 |
| 2019 | Understanding and Improving Kernel Local Descriptors
Arun Mukundan, Giorgos Tolias, Andrei Bursuc, Hervé Jégou, Ondrej Chum |
Int. J. Comput. Vis. | 2 |
| 2019 | Graph-based particular object discovery
Oriane Siméoni, Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum |
Mach. Vis. Appl. | 3 |
| 2019 | Fine-Tuning CNN Image Retrieval with No Human AnnotationabstractImage descriptors based on activations of Convolutional Neural Networks (CNNs) have become dominant in image retrieval due to their discriminative power, compactness of representation, and search efficiency. Training of CNNs, either from scratch or fine-tuning, requires a large amount of annotated data, where a high quality of annotation is often crucial. In this work, we propose to fine-tune CNNs for image retrieval on a large collection of unordered images in a fully automated manner. Reconstructed 3D models obtained by the state-of-the-art retrieval and structure-from-motion methods guide the selection of the training data. We show that both hard-positive and hard-negative examples, selected by exploiting the geometry and the camera positions available from the 3D models, enhance the performance of particular-object retrieval. CNN descriptor whitening discriminatively learned from the same training data outperforms commonly used PCA whitening. We propose a novel trainable Generalized-Mean (GeM) pooling layer that generalizes max and average pooling and show that it boosts retrieval performance. Applying the proposed method to the VGG network achieves state-of-the-art performance on the standard benchmarks: Oxford Buildings, Paris, and Holidays datasets. Filip Radenovic, Giorgos Tolias, Ondrej Chum |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Hybrid Diffusion: Spectral-Temporal Graph Filtering for Manifold Ranking
Ahmet Iscen, Yannis Avrithis, Giorgos Tolias, Teddy Furon, Ondrej Chum |
ACCV (2) | 3 |
| 2018 | Fast Spectral Ranking for Similarity SearchabstractDespite the success of deep learning on representing images for particular object retrieval, recent studies show that the learned representations still lie on manifolds in a high dimensional space. This makes the Euclidean nearest neighbor search biased for this task. Exploring the manifolds online remains expensive even if a nearest neighbor graph has been computed offline. This work introduces an explicit embedding reducing manifold search to Euclidean search followed by dot product similarity search. This is equivalent to linear graph filtering of a sparse signal in the frequency domain. To speed up online search, we compute an approximate Fourier basis of the graph offline. We improve the state of art on particular object retrieval datasets including the challenging Instre dataset containing small objects. At a scale of 105 images, the offline cost is only a few hours, while query time is comparable to standard similarity search. Ahmet Iscen, Yannis Avrithis, Giorgos Tolias, Teddy Furon, Ondrej Chum |
CVPR | 3 |
| 2018 | Mining on Manifolds: Metric Learning Without LabelsabstractIn this work we present a novel unsupervised framework for hard training example mining. The only input to the method is a collection of images relevant to the target application and a meaningful initial representation, provided e.g. by pre-trained CNN. Positive examples are distant points on a single manifold, while negative examples are nearby points on different manifolds. Both types of examples are revealed by disagreements between Euclidean and manifold similarities. The discovered examples can be used in training with any discriminative loss. The method is applied to unsupervised fine-tuning of pre-trained networks for fine-grained classification and particular object retrieval. Our models are on par or are outperforming prior models that are fully or partially supervised. Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum |
CVPR | 2 |
| 2018 | Revisiting Oxford and Paris: Large-Scale Image Retrieval BenchmarkingabstractIn this paper we address issues with image retrieval benchmarking on standard and popular Oxford 5k and Paris 6k datasets. In particular, annotation errors, the size of the dataset, and the level of challenge are addressed: new annotation for both datasets is created with an extra attention to the reliability of the ground truth. Three new protocols of varying difficulty are introduced. The protocols allow fair comparison between different methods, including those using a dataset pre-processing stage. For each dataset, 15 new challenging queries are introduced. Finally, a new set of 1M hard, semi-automatically cleaned distractors is selected. An extensive comparison of the state-of-the-art methods is performed on the new benchmark. Different types of methods are evaluated, ranging from local-feature-based to modern CNN based methods. The best results are achieved by taking the best of the two worlds. Most importantly, image retrieval appears far from being solved. Filip Radenovic, Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum |
CVPR | 3 |
| 2018 | Deep Shape Matching
Filip Radenovic, Giorgos Tolias, Ondrej Chum |
ECCV (5) | 2 |
| 2018 | Unsupervised Object Discovery for Instance RecognitionabstractSevere background clutter is challenging in many computer vision tasks, including large-scale image retrieval. Global descriptors, that are popular due to their memory and search efficiency, are especially prone to corruption by such a clutter. Eliminating the impact of the clutter on the image descriptor increases the chance of retrieving relevant images and prevents topic drift due to actually retrieving the clutter in the case of query expansion. In this work, we propose a novel salient region detection method. It captures, in an unsupervised manner, patterns that are both discriminative and common in the dataset. Saliency is based on a centrality measure of a nearest neighbor graph constructed from regional CNN representations of dataset images. The descriptors derived from the salient regions improve particular object retrieval, most noticeably in a large collections containing small objects. Oriane Siméoni, Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum |
WACV | 3 |
| 2018 | Efficient contour match kernel
Giorgos Tolias, Ondrej Chum |
Image Vis. Comput. | 1 |
| 2017 | Multiple-Kernel Local-Patch Descriptor
Arun Mukundan, Giorgos Tolias, Ondrej Chum |
BMVC | 2 |
| 2017 | Efficient Diffusion on Region Manifolds: Recovering Small Objects with Compact CNN RepresentationsabstractQuery expansion is a popular method to improve the quality of image retrieval with both conventional and CNN representations. It has been so far limited to global image similarity. This work focuses on diffusion, a mechanism that captures the image manifold in the feature space. An efficient off-line stage allows optional reduction in the number of stored regions. In the on-line stage, the proposed handling of unseen queries in the indexing stage removes additional computation to adjust the precomputed data. We perform diffusion through a sparse linear system solver, yielding practical query times well below one second. Experimentally, we observe a significant boost in performance of image retrieval with compact CNN descriptors on standard benchmarks, especially when the query object covers only a small part of the image. Small objects have been a common failure case of CNN-based retrieval. Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Teddy Furon, Ondrej Chum |
CVPR | 2 |
| 2017 | Asymmetric Feature Maps with Application to Sketch Based RetrievalabstractWe propose a novel concept of asymmetric feature maps (AFM), which allows to evaluate multiple kernels between a query and database entries without increasing the memory requirements. To demonstrate the advantages of the AFM method, we derive a short vector image representation that, due to asymmetric feature maps, supports efficient scale and translation invariant sketch-based image retrieval. Unlike most of the short-code based retrieval systems, the proposed method provides the query localization in the retrieved image. The efficiency of the search is boosted by approximating a 2D translation search via trigonometric polynomial of scores by 1D projections. The projections are a special case of AFM. An order of magnitude speed-up is achieved compared to traditional trigonometric polynomials. The results are boosted by an image-based average query expansion, exceeding significantly the state of the art on standard benchmarks. Giorgos Tolias, Ondrej Chum |
CVPR | 1 |
| 2017 | Panorama to Panorama Matching for Location RecognitionabstractInternational audience Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Teddy Furon, Ondrej Chum |
ICMR | 2 |
| 2016 | CNN Image Retrieval Learns from BoW: Unsupervised Fine-Tuning with Hard Examples
Filip Radenovic, Giorgos Tolias, Ondrej Chum |
ECCV (1) | 2 |
| 2016 | Image Search with Selective Match Kernels: Aggregation Across Single and Multiple Images
Giorgos Tolias, Yannis Avrithis, Hervé Jégou |
Int. J. Comput. Vis. | 1 |
| 2016 | Erratum to: Image Search with Selective Match Kernels: Aggregation Across Single and Multiple Images
Giorgos Tolias, Yannis Avrithis, Hervé Jégou |
Int. J. Comput. Vis. | 1 |
| 2015 | Kernel Local Descriptors with Implicit Rotation MatchingabstractIn this work we design a kernelized local feature descriptor and propose a matching scheme for aligning patches quickly and automatically. We analyze the SIFT descriptor from a kernel view and identify and reproduce some of its underlying benefits. We overcome the quantization artifacts of SIFT by encoding pixel attributes in a continous manner via explicit feature maps. Experiments performed on the patch dataset of Brown et al. [3] show the superiority of our descriptor over methods based on supervised learning. Andrei Bursuc, Giorgos Tolias, Hervé Jégou |
ICMR | 2 |
| 2015 | Rotation and translation covariant match kernels for image retrieval
Giorgos Tolias, Andrei Bursuc, Teddy Furon, Hervé Jégou |
Comput. Vis. Image Underst. | 1 |
| 2015 | A Comparison of Dense Region Detectors for Image Search and Fine-Grained ClassificationabstractWe consider a pipeline for image classification or search based on coding approaches like bag of words or Fisher vectors. In this context, the most common approach is to extract the image patches regularly in a dense manner on several scales. This paper proposes and evaluates alternative choices to extract patches densely. Beyond simple strategies derived from regular interest region detectors, we propose approaches based on superpixels, edges, and a bank of Zernike filters used as detectors. The different approaches are evaluated on recent image retrieval and fine-grained classification benchmarks. Our results show that the regular dense detector is outperformed by other methods in most situations, leading us to improve the state-of-the-art in comparable setups on standard retrieval and fined-grained benchmarks. As a byproduct of our study, we show that existing methods for blob and superpixel extraction achieve high accuracy if the patches are extracted along the edges and not around the detected regions. Ahmet Iscen, Giorgos Tolias, Philippe Henri Gosselin, Hervé Jégou |
IEEE Trans. Image Process. | 2 |
| 2014 | Orientation Covariant Aggregation of Local Descriptors with Embeddings
Giorgos Tolias, Teddy Furon, Hervé Jégou |
ECCV (6) | 1 |
| 2014 | Towards large-scale geometry indexing by feature selection
Giorgos Tolias, Yannis Kalantidis, Yannis Avrithis, Stefanos D. Kollias |
Comput. Vis. Image Underst. | 1 |
| 2014 | Hough Pyramid Matching: Speeded-Up Geometry Re-ranking for Large Scale Image Retrieval
Yannis Avrithis, Giorgos Tolias |
Int. J. Comput. Vis. | 2 |
| 2014 | Visual query expansion with or without geometry: Refining local descriptors by feature aggregation
Giorgos Tolias, Hervé Jégou |
Pattern Recognit. | 1 |
| 2013 | To Aggregate or Not to aggregate: Selective Match Kernels for Image SearchabstractThis paper considers a family of metrics to compare images based on their local descriptors. It encompasses the VLAD descriptor and matching techniques such as Hamming Embedding. Making the bridge between these approaches leads us to propose a match kernel that takes the best of existing techniques by combining an aggregation procedure with a selective match kernel. Finally, the representation underpinning this kernel is approximated, providing a large scale image search both precise and scalable, as shown by our experiments on several benchmarks. Giorgos Tolias, Yannis Avrithis, Hervé Jégou |
ICCV | 1 |
| 2012 | SymCity: feature selection by symmetry for large scale image retrievalabstractMany problems, including feature selection, vocabulary learning, location and landmark recognition, structure from motion and 3d reconstruction, rely on a learning process that involves wide-baseline matching on multiple views of the same object or scene. In practical large scale image retrieval applications however, most images depict unique views where this idea does not apply. We exploit self-similarities, symmetries and repeating patterns to select features within a single image. We achieve the same performance compared to the full feature set with only a small fraction of its index size on a dataset of unique views of buildings or urban scenes, in the presence of one million distractors of similar nature. Our best solution is linear in the number of correspondences, with practical running times of just a few milliseconds. Giorgos Tolias, Yannis Kalantidis, Yannis Avrithis |
ACM Multimedia | 1 |
| 2011 | Speeded-up, relaxed spatial matchingabstractA wide range of properties and assumptions determine the most appropriate spatial matching model for an application, e.g. recognition, detection, registration, or large scale image retrieval. Most notably, these include discriminative power, geometric invariance, rigidity constraints, mapping constraints, assumptions made on the underlying features or descriptors and, of course, computational complexity. Having image retrieval in mind, we present a very simple model inspired by Hough voting in the transformation space, where votes arise from single feature correspondences. A relaxed matching process allows for multiple matching surfaces or non-rigid objects under one-to-one mapping, yet is linear in the number of correspondences. We apply it to geometry re-ranking in a search engine, yielding superior performance with the same space requirements but a dramatic speed-up compared to the state of the art. Giorgos Tolias, Yannis Avrithis |
ICCV | 1 |
| 2011 | VIRaL: Visual Image Retrieval and Localization
Yannis Kalantidis, Giorgos Tolias, Yannis Avrithis, Marios Phinikettos, Evaggelos Spyrou, Phivos Mylonas, Stefanos D. Kollias |
Multim. Tools Appl. | 2 |
| 2010 | Intelligent content retrieval using a visual vocabulary and geometric constraintsabstractDuring the last decades multimedia processing has emerged as an important technology to retrieve content based on similar data. Moreover, recent developments in the fields of high definition (HD) multimedia content and personal content collections (personal camcorders and digital still image cameras) tend to generate a huge volume of multimedia data everyday. Thus, the need for a meaningful, quick organization and access to generated content is now more than necessary; however, it still remains a rather difficult problem to be tackled both by humans and computers. In this paper we propose an intelligent extension of traditional image analysis methodologies towards more efficient digital content retrieval. The main idea is to extend local feature extraction methodologies by introducing additional geometrical constraints in the process. The proposed approach is tested and evaluated on a number of publicly available image datasets and results are very promising. Evaggelos Spyrou, Yannis Kalantidis, Giorgos Tolias, Phivos Mylonas, Stefanos D. Kollias |
FUZZ-IEEE | 3 |
| 2010 | Image clustering through community detection on hybrid image similarity graphsabstractThe wide adoption of photo sharing applications such as Flickr©and the massive amounts of user-generated content uploaded to them raises an information overload issue for users. An established technique to overcome such an overload is to cluster images into groups based on their similarity and then use the derived clusters to assist navigation and browsing of the collection. In this paper, we present a community detection (i.e. graph-based clustering) approach that makes use of both visual and tagging features of images in order to efficiently extract groups of related images within large image collections. Based on experiments we conducted on a dataset comprising publicly available images from Flickr©, we demonstrate the efficiency of our method, the added value of combining visual and tag features and the utility of the derived clusters for exploring an image collection. Symeon Papadopoulos, Christos Zigkolis, Giorgos Tolias, Yannis Kalantidis, Phivos Mylonas, Ioannis Kompatsiaris, Athena Vakali |
ICIP | 3 |
| 2010 | Retrieving landmark and non-landmark images from community photo collectionsabstractState of the art data mining and image retrieval in community photo collections typically focus on popular subsets, e.g. images containing landmarks or associated to Wikipedia articles. We propose an image clustering scheme that, seen as vector quantization compresses a large corpus of images by grouping visually consistent ones while providing a guaranteed distortion bound. This allows us, for instance, to represent the visual content of all thousands of images depicting the Parthenon in just a few dozens of scene maps and still be able to retrieve any single, isolated, non-landmark image like a house or graffiti on a wall. Starting from a geo-tagged dataset, we first group images geographically and then visually, where each visual cluster is assumed to depict different views of the the same scene. We align all views to one reference image and construct a 2D scene map by preserving details from all images while discarding repeating visual features. Our indexing, retrieval and spatial matching scheme then operates directly on scene maps. We evaluate the precision of the proposed method on a challenging one-million urban image dataset. Yannis Avrithis, Yannis Kalantidis, Giorgos Tolias, Evaggelos Spyrou |
ACM Multimedia | 3 |
| 2010 | Feature map hashing: sub-linear indexing of appearance and global geometryabstractWe present a new approach to image indexing and retrieval, which integrates appearance with global image geometry in the indexing process, while enjoying robustness against viewpoint change, photometric variations, occlusion, and background clutter. We exploit shape parameters of local features to estimate image alignment via a single correspondence. Then, for each feature, we construct a sparse spatial map of all remaining features, encoding their normalized position and appearance, typically vector quantized to visual word. An image is represented by a collection of such feature maps and RANSAC-like matching is reduced to a number of set intersections. Because the induced dissimilarity is still not a metric, we extend min-wise independent permutations to collections of sets and derive a similarity measure for feature map collections. We then exploit sparseness to build an inverted file whereby the retrieval process is sub-linear in the total number of images, ideally linear in the number of relevant ones. We achieve excellent performance on 10^4 images, with a query time in the order of milliseconds. Yannis Avrithis, Giorgos Tolias, Yannis Kalantidis |
ACM Multimedia | 2 |
| 2009 | Large Scale Concept Detection in Video Using a Region Thesaurus
Evaggelos Spyrou, Giorgos Tolias, Yannis Avrithis |
MMM | 2 |
| 2009 | Concept detection and keyframe extraction using a visual thesaurus
Evaggelos Spyrou, Giorgos Tolias, Phivos Mylonas, Yannis Avrithis |
Multim. Tools Appl. | 2 |