EDBT 2026 Demo / reviewers in the wild / expert
Ondrej Chum
dblp:96/63
· DBLP profile ↗
84ranked-venue papers
18as first author
12since 2021 · last 2025
0000-0001-7042-1810ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 77 · 18 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 62 · 14 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ILIAS: Instance-Level Image retrieval At ScaleabstractThis work introduces ILIAS, a new test dataset for Instance-Level Image retrieval At Scale. It is designed to evaluate the ability of current and future foundation models and retrieval techniques to recognize particular objects. The key benefits over existing datasets include large scale, domain diversity, accurate ground truth, and a performance that is far from saturated. ILIAS includes query and positive images for 1,000 object instances, manually collected to capture challenging conditions and diverse domains. Large-scale retrieval is conducted against 100 million distractor images from YFCC100M. To avoid false negatives without extra annotation effort, we include only query objects confirmed to have emerged after 2014, i.e. the compilation date of YFCC100M. An extensive benchmarking is performed with the following observations: i) models fine-tuned on specific domains, such as landmarks or products, excel in that domain but fail on ILIAS ii) learning a linear adaptation layer using multi-domain class supervision results in performance improvements, especially for vision-language models iii) local descriptors in retrieval re-ranking are still a key ingredient, especially in the presence of severe background clutter iv) the text-to-image performance of the vision-language foundation models is surprisingly close to the corresponding image-to-image case. website: https://vrg.fel.cvut.cz/ilias/ Giorgos Kordopatis-Zilos, Vladan Stojnic, Anna Manko, Pavel Suma, Nikolaos-Antonios Ypsilantis, Nikos Efthymiadis, Zakaria Laskar, Jiri Matas, Ondrej Chum, Giorgos Tolias |
CVPR | 9 |
| 2025 | Instance-Level Composed Image RetrievalabstractThe progress of composed image retrieval (CIR), a popular research direction in image retrieval, where a combined visual and textual query is used, is held back by the absence of high-quality training and evaluation data. We introduce a new evaluation dataset, i-CIR, which, unlike existing datasets, focuses on an instance-level class definition. The goal is to retrieve images that contain the same particular object as the visual query, presented under a variety of modifications defined by textual queries. Its design and curation process keep the dataset compact to facilitate future research, while maintaining its challenge—comparable to retrieval among more than 40M random distractors—through a semi-automated selection of hard negatives. To overcome the challenge of obtaining clean, diverse, and suitable training data, we leverage pre-trained vision-and-language models (VLMs) in a training-free approach called BASIC. The method separately estimates query-image-to-image and query-text-to-image similarities, performing late fusion to upweight images that satisfy both queries, while down-weighting those that exhibit high similarity with only one of the two. Each individual similarity is further improved by a set of components that are simple and intuitive. BASIC sets a new state of the art on i-CIR but also on existing CIR datasets that follow a semantic-level class definition. Project page: https://vrg.fel.cvut.cz/icir/. Bill Psomas, George Retsinas, Nikos Efthymiadis, Panagiotis Paraskevas Filntisis, Yannis Avrithis, Petros Maragos, Ondrej Chum, Giorgos Tolias |
NeurIPS | 7 |
| 2025 | Composed Image Retrieval for Training-FREE DOMain ConversionabstractThis work addresses composed image retrieval in the context of domain conversion, where the content of a query image is retrieved in the domain specified by the query text. We show that a strong vision-language model provides sufficient descriptive power without additional training. The query image is mapped to the text input space using textual inversion. Unlike common practice that invert in the continuous space of text tokens, we use the discrete word space via a nearest-neighbor search in a text vocabulary. With this inversion, the image is softly mapped across the vocabulary and is made more robust using retrieval-based augmentation. Database images are retrieved by a weighted ensemble of text queries combining mapped words with the domain text. Our method outperforms prior art by a large margin on standard and newly introduced benchmarks. Code: https://github.com/NikosEfth/freedom Nikos Efthymiadis, Bill Psomas, Zakaria Laskar, Konstantinos Karantzalos, Yannis Avrithis, Ondrej Chum, Giorgos Tolias |
WACV | 6 |
| 2025 | Crafting Distribution Shifts for Validation and Training in Single Source Domain GeneralizationabstractSingle-source domain generalization attempts to learn a model on a source domain and deploy it to unseen target domains. Limiting access only to source domain data imposes two key challenges - how to train a model that can generalize and how to verify that it does. The standard practice of validation on the training distribution does not accurately reflect the model's generalization ability, while validation on the test distribution is a malpractice to avoid. In this work, we construct an independent validation set by transforming source domain images with a comprehensive list of augmentations, covering a broad spectrum of potential distribution shifts in target domains. We demonstrate a high correlation between validation and test performance for multiple methods and across various datasets. The proposed validation achieves a relative accuracy improvement over the standard validation equal to 15.4% or 1.6% when used for method selection or learning rate tuning, respectively. Furthermore, we introduce a novel family of methods that increase the shape bias through enhanced edge maps. To benefit from the augmentations during training and preserve the independence of the validation set, a k-fold validation process is designed to separate the augmentation types used in training and validation. The method that achieves the best performance on the augmented validation is selected from the proposed family. It achieves state-of-the-art performance on various standard benchmarks. Code at: https://github.com/NikosEfth/crafting-shifts Nikos Efthymiadis, Giorgos Tolias, Ondrej Chum |
WACV | 3 |
| 2024 | Co-segmentation Without any Pixel-Level Supervision with Application to Large-Scale Sketch Classification
Nikolaos-Antonios Ypsilantis, Ondrej Chum |
ACCV (10) | 2 |
| 2024 | Composed Image Retrieval for Remote SensingabstractThis work introduces composed image retrieval to remote sensing. It allows to query a large image archive by image examples alternated by a textual description, enriching the descriptive power over unimodal queries, either visual or textual. Various attributes can be modified by the textual part, such as shape, color, or context. A novel method fusing image-to-image and text-to-image similarity is introduced. We demonstrate that a vision-language model possesses sufficient descriptive power and no further learning step or training data are necessary. We present a new evaluation benchmark focused on color, context, density, existence, quantity, and shape modifications. Our work not only sets the state-of-the-art for this task, but also serves as a foundational step in addressing a gap in the field of remote sensing image retrieval. Code at: https://github.com/billpsomas/rscir. Bill Psomas, Ioannis Kakogeorgiou, Nikos Efthymiadis, Giorgos Tolias, Ondrej Chum, Yannis Avrithis, Konstantinos Karantzalos |
IGARSS | 5 |
| 2024 | UDON: Universal Dynamic Online distillatioN for generic image representationsabstractUniversal image representations are critical in enabling real-world fine-grained and instance-level recognition applications, where objects and entities from any domain must be identified at large scale.
Despite recent advances, existing methods fail to capture important domain-specific knowledge, while also ignoring differences in data distribution across different domains.
This leads to a large performance gap between efficient universal solutions and expensive approaches utilising a collection of specialist models, one for each domain.
In this work, we make significant strides towards closing this gap, by introducing a new learning technique, dubbed UDON (Universal Dynamic Online distillatioN).
UDON employs multi-teacher distillation, where each teacher is specialized in one domain, to transfer detailed domain-specific knowledge into the student universal embedding.
UDON's distillation approach is not only effective, but also very efficient, by sharing most model parameters between the student and all teachers, where all models are jointly trained in an online manner.
UDON also comprises a sampling technique which adapts the training process to dynamically allocate batches to domains which are learned slower and require more frequent processing.
This boosts significantly the learning of complex domains which are characterised by a large number of classes and long-tail distributions.
With comprehensive experiments, we validate each component of UDON, and showcase significant improvements over the state of the art in the recent UnED benchmark.
Code: https://github.com/nikosips/UDON. Nikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araújo 0001, Ondrej Chum |
NeurIPS | 4 |
| 2023 | Dark Side Augmentation: Generating Diverse Night Examples for Metric LearningabstractImage retrieval methods based on CNN descriptors rely on metric learning from a large number of diverse examples of positive and negative image pairs. Domains, such as night-time images, with limited availability and variability of training data suffer from poor retrieval performance even with methods performing well on standard benchmarks. We propose to train a GAN-based synthetic-image generator, translating available day-time image examples into night images. Such a generator is used in metric learning as a form of augmentation, supplying training data to the scarce domain. Various types of generators are evaluated and analyzed. We contribute with a novel light-weight GAN architecture that enforces the consistency between the original and translated image through edge consistency. The proposed architecture also allows a simultaneous training of an edge detector that operates on both night and day images. To further increase the variability in the training examples and to maximize the generalization of the trained model, we propose a novel method of diverse anchor mining.The proposed method improves over the state-of-the-art results on a standard Tokyo 24/7 day-night retrieval benchmark while preserving the performance on Oxford and Paris datasets. This is achieved without the need of training image pairs of matching day and night images. The source code is available at https://github.com/mohwald/gandtr. Albert Mohwald, Tomás Jenícek, Ondrej Chum |
ICCV | 3 |
| 2023 | Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image RepresentationsabstractFine-grained and instance-level recognition methods are commonly trained and evaluated on specific domains, in a model per domain scenario. Such an approach, however, is impractical in real large-scale applications. In this work, we address the problem of universal image embedding, where a single universal model is trained and used in multiple domains. First, we leverage existing domain-specific datasets to carefully construct a new large-scale public benchmark for the evaluation of universal image embeddings, with 241k query images, 1.4M index images and 2.8M training images across 8 different domains and 349k classes. We define suitable metrics, training and evaluation protocols to foster future research in this area. Second, we provide a comprehensive experimental evaluation on the new dataset, demonstrating that existing approaches and simplistic extensions lead to worse performance than an assembly of models trained for each domain separately. Finally, we conducted a public research competition on this topic, leveraging industrial datasets, which attracted the participation of more than 1k teams world-wide. This exercise generated many interesting research ideas and findings which we present in detail. Project webpage: https://cmp.felk.cvut.cz/univ_emb/ Nikolaos-Antonios Ypsilantis, Kaifeng Chen, Bingyi Cao, Mário Lipovský, Pelin Dogan-Schönberger, Grzegorz Makosa, Boris Bluntschli, Mojtaba Seyedhosseini, Ondrej Chum, André Araújo 0001 |
ICCV | 9 |
| 2022 | Object-Guided Day-Night Visual localization in Urban ScenesabstractWe introduce Object-Guided localization (OGuL) based on a novel method of local-feature matching. Direct matching of local features is sensitive to significant changes in illumination. In contrast, object detection often survives severe changes in lighting conditions. The proposed method first detects semantic objects and establishes correspondences of those objects between images. Object correspondences provide local coarse alignment of the images in the form of a planar homography. These homographies are consequently used to guide the matching of local features. Experiments on standard urban localization datasets (Aachen, RobotCar-Season) show that OGuL significantly improves localization results with as simple local features as SIFT, and its performance competes with the state-of-the-art CNN-based methods trained for day-to-night localization. Assia Benbihi, Cédric Pradalier, Ondrej Chum |
ICPR | 3 |
| 2022 | Edge Augmentation for Large-Scale Sketch Recognition without SketchesabstractThis work addresses scaling up the sketch classification task into a large number of categories. Collecting sketches for training is a slow and tedious process that has so far precluded any attempts to large-scale sketch recognition. We overcome the lack of training sketch data by exploiting labeled collections of natural images that are easier to obtain. To bridge the domain gap we present a novel augmentation technique that is tailored to the task of learning sketch recognition from a training set of natural images. Randomization is introduced in the parameters of edge detection and edge selection. Natural images are translated to a pseudo-novel domain called "randomized Binary Thin Edges" (rBTE), which is used as a training domain instead of natural images. The ability to scale up is demonstrated by training CNN-based sketch recognition of more than 2.5 times larger number of categories than used previously. For this purpose, a dataset of natural images from 874 categories is constructed by combining a number of popular computer vision datasets. The categories are selected to be suitable for sketch recognition. To estimate the performance, a subset of 393 categories with sketches is also collected. Nikos Efthymiadis, Giorgos Tolias, Ondrej Chum |
ICPR | 3 |
| 2021 | Minimal Solvers for Rectifying From Radially-Distorted Conjugate TranslationsabstractThis paper introduces minimal solvers that jointly solve for radial lens undistortion and affine-rectification using local features extracted from the image of coplanar translated and reflected scene texture, which is common in man-made environments. The proposed solvers accommodate different types of local features and sampling strategies, and three of the proposed variants require just one feature correspondence. State-of-the-art techniques from algebraic geometry are used to simplify the formulation of the solvers. The generated solvers are stable, small and fast. Synthetic and real-image experiments show that the proposed solvers have superior robustness to noise compared to the state of the art. The solvers are integrated with an automated system for rectifying imaged scene planes from coplanar repeated texture. Accurate rectifications on challenging imagery taken with narrow to wide field-of-view lenses demonstrate the applicability of the proposed solvers. James Pritts, Zuzana Kukelova, Viktor Larsson, Yaroslava Lochman, Ondrej Chum |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Graph Convolutional Networks for Learning with Few Clean and Many Noisy Labels
Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum, Cordelia Schmid |
ECCV (10) | 4 |
| 2020 | Learning and Aggregating Deep Local Descriptors for Instance-Level Recognition
Giorgos Tolias, Tomás Jenícek, Ondrej Chum |
ECCV (1) | 3 |
| 2020 | Minimal Solvers for Rectifying from Radially-Distorted Scales and Change of Scales
James Pritts, Zuzana Kukelova, Viktor Larsson, Yaroslava Lochman, Ondrej Chum |
Int. J. Comput. Vis. | 5 |
| 2020 | Saddle: Fast and repeatable features with good coverage
Javier Aldana-Iuit, Dmytro Mishkin, Ondrej Chum, Jiri Matas |
Image Vis. Comput. | 3 |
| 2019 | Label Propagation for Deep Semi-Supervised LearningabstractSemi-supervised learning is becoming increasingly important because it can combine data carefully labeled by humans with abundant unlabeled data to train deep neural networks. Classic methods on semi-supervised learning that have focused on transductive learning have not been fully exploited in the inductive framework followed by modern deep learning. The same holds for the manifold assumption-that similar examples should get the same prediction. In this work, we employ a transductive label propagation method that is based on the manifold assumption to make predictions on the entire dataset and use these predictions to generate pseudo-labels for the unlabeled data and train a deep neural network. At the core of the transductive method lies a nearest neighbor graph of the dataset that we create based on the embeddings of the same network. Therefore our learning process iterates between these two steps. We improve performance on several datasets especially in the few labels regime and show that our work is complementary to current state of the art. Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum |
CVPR | 4 |
| 2019 | Explicit Spatial Encoding for Deep Local DescriptorsabstractWe propose a kernelized deep local-patch descriptor based on efficient match kernels of neural network activations. Response of each receptive field is encoded together with its spatial location using explicit feature maps. Two location parametrizations, Cartesian and polar, are used to provide robustness to a different types of canonical patch misalignment. Additionally, we analyze how the conventional architecture, i.e. a fully connected layer attached after the convolutional part, encodes responses in a spatially variant way. In contrary, explicit spatial encoding is used in our descriptor, whose potential applications are not limited to local-patches. We evaluate the descriptor on standard benchmarks. Both versions, encoding 32 × 32 or 64 × 64 patches, consistently outperform all other methods on all benchmarks. The number ofparameters of the model is independent of the input patch resolution. Arun Mukundan, Giorgos Tolias, Ondrej Chum |
CVPR | 3 |
| 2019 | Local Features and Visual Words Emerge in ActivationsabstractWe propose a novel method of deep spatial matching (DSM) for image retrieval. Initial ranking is based on image descriptors extracted from convolutional neural network activations by global pooling, as in recent state-of-the-art work. However, the same sparse 3D activation tensor is also approximated by a collection of local features. These local features are then robustly matched to approximate the optimal alignment of the tensors. This happens without any network modification, additional layers or training. No local feature detection happens on the original image. No local feature descriptors and no visual vocabulary are needed throughout the whole process. We experimentally show that the proposed method achieves the state-of-the-art performance on standard benchmarks across different network architectures and different global pooling methods. The highest gain in performance is achieved when diffusion on the nearest-neighbor graph of global descriptors is initiated from spatially verified images. Oriane Siméoni, Yannis Avrithis, Ondrej Chum |
CVPR | 3 |
| 2019 | No Fear of the Dark: Image Retrieval Under Varying Illumination Conditions
Tomás Jenícek, Ondrej Chum |
ICCV | 2 |
| 2019 | Targeted Mismatch Adversarial Attack: Query With a Flower to Retrieve the TowerabstractAccess to online visual search engines implies sharing of private user content - the query images. We introduce the concept of targeted mismatch attack for deep learning based retrieval systems to generate an adversarial image to conceal the query image. The generated image looks nothing like the user intended query, but leads to identical or very similar retrieval results. Transferring attacks to fully unseen networks is challenging. We show successful attacks to partially unknown systems, by designing various loss functions for the adversarial image construction. These include loss functions, for example, for unknown global pooling operation or unknown input resolution by the retrieval system. We evaluate the attacks on standard retrieval benchmarks and compare the results retrieved with the original and adversarial image. Giorgos Tolias, Filip Radenovic, Ondrej Chum |
ICCV | 3 |
| 2019 | Linking Art through Human PosesabstractWe address the discovery of composition transfer in artworks based on their visual content. Automated analysis of large art collections, which are growing as a result of art digitization among museums and galleries, is an important tool for art history and assists cultural heritage preservation. Modern image retrieval systems offer good performance on visually similar artworks, but fail in the cases of more abstract composition transfer. The proposed approach links artworks through a pose similarity of human figures depicted in images. Human figures are the subject of a large fraction of visual art from middle ages to modernity and their distinctive poses were often a source of inspiration among artists. The method consists of two steps - fast pose matching and robust spatial verification. We experimentally show that explicit human pose matching is superior to standard content-based image retrieval methods on a manually annotated art composition transfer dataset. Tomás Jenícek, Ondrej Chum |
ICDAR | 2 |
| 2019 | Understanding and Improving Kernel Local Descriptors
Arun Mukundan, Giorgos Tolias, Andrei Bursuc, Hervé Jégou, Ondrej Chum |
Int. J. Comput. Vis. | 5 |
| 2019 | Graph-based particular object discovery
Oriane Siméoni, Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum |
Mach. Vis. Appl. | 5 |
| 2019 | Fine-Tuning CNN Image Retrieval with No Human AnnotationabstractImage descriptors based on activations of Convolutional Neural Networks (CNNs) have become dominant in image retrieval due to their discriminative power, compactness of representation, and search efficiency. Training of CNNs, either from scratch or fine-tuning, requires a large amount of annotated data, where a high quality of annotation is often crucial. In this work, we propose to fine-tune CNNs for image retrieval on a large collection of unordered images in a fully automated manner. Reconstructed 3D models obtained by the state-of-the-art retrieval and structure-from-motion methods guide the selection of the training data. We show that both hard-positive and hard-negative examples, selected by exploiting the geometry and the camera positions available from the 3D models, enhance the performance of particular-object retrieval. CNN descriptor whitening discriminatively learned from the same training data outperforms commonly used PCA whitening. We propose a novel trainable Generalized-Mean (GeM) pooling layer that generalizes max and average pooling and show that it boosts retrieval performance. Applying the proposed method to the VGG network achieves state-of-the-art performance on the standard benchmarks: Oxford Buildings, Paris, and Holidays datasets. Filip Radenovic, Giorgos Tolias, Ondrej Chum |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Hybrid Diffusion: Spectral-Temporal Graph Filtering for Manifold Ranking
Ahmet Iscen, Yannis Avrithis, Giorgos Tolias, Teddy Furon, Ondrej Chum |
ACCV (2) | 5 |
| 2018 | Rectification from Radially-Distorted Scales
James Pritts, Zuzana Kukelova, Viktor Larsson, Ondrej Chum |
ACCV (5) | 4 |
| 2018 | Fast Spectral Ranking for Similarity SearchabstractDespite the success of deep learning on representing images for particular object retrieval, recent studies show that the learned representations still lie on manifolds in a high dimensional space. This makes the Euclidean nearest neighbor search biased for this task. Exploring the manifolds online remains expensive even if a nearest neighbor graph has been computed offline. This work introduces an explicit embedding reducing manifold search to Euclidean search followed by dot product similarity search. This is equivalent to linear graph filtering of a sparse signal in the frequency domain. To speed up online search, we compute an approximate Fourier basis of the graph offline. We improve the state of art on particular object retrieval datasets including the challenging Instre dataset containing small objects. At a scale of 105 images, the offline cost is only a few hours, while query time is comparable to standard similarity search. Ahmet Iscen, Yannis Avrithis, Giorgos Tolias, Teddy Furon, Ondrej Chum |
CVPR | 5 |
| 2018 | Mining on Manifolds: Metric Learning Without LabelsabstractIn this work we present a novel unsupervised framework for hard training example mining. The only input to the method is a collection of images relevant to the target application and a meaningful initial representation, provided e.g. by pre-trained CNN. Positive examples are distant points on a single manifold, while negative examples are nearby points on different manifolds. Both types of examples are revealed by disagreements between Euclidean and manifold similarities. The discovered examples can be used in training with any discriminative loss. The method is applied to unsupervised fine-tuning of pre-trained networks for fine-grained classification and particular object retrieval. Our models are on par or are outperforming prior models that are fully or partially supervised. Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum |
CVPR | 4 |
| 2018 | Radially-Distorted Conjugate TranslationsabstractThis paper introduces the first minimal solvers that jointly solve for affine-rectification and radial lens distortion from coplanar repeated patterns. Even with imagery from moderately distorted lenses, plane rectification using the pinhole camera model is inaccurate or invalid. The proposed solvers incorporate lens distortion into the camera model and extend accurate rectification to wide-angle imagery, which is now common from consumer cameras. The solvers are derived from constraints induced by the conjugate translations of an imaged scene plane, which are integrated with the division model for radial lens distortion. The hidden-variable trick with ideal saturation is used to reformulate the constraints so that the solvers generated by the Gröbner-basis method are stable, small and fast. Rectification and lens distortion are recovered from either one conjugately translated affine-covariant feature or two independently translated similarity-covariant features. The proposed solvers are used in a RANSAC-based estimator, which gives accurate rectifications after few iterations. The proposed solvers are evaluated against the state-of-the-art and demonstrate significantly better rectifcations on noisy measurements. Qualitative results on diverse imagery demonstrate high-accuracy undistortion and rectification. James Pritts, Zuzana Kukelova, Viktor Larsson, Ondrej Chum |
CVPR | 4 |
| 2018 | Revisiting Oxford and Paris: Large-Scale Image Retrieval BenchmarkingabstractIn this paper we address issues with image retrieval benchmarking on standard and popular Oxford 5k and Paris 6k datasets. In particular, annotation errors, the size of the dataset, and the level of challenge are addressed: new annotation for both datasets is created with an extra attention to the reliability of the ground truth. Three new protocols of varying difficulty are introduced. The protocols allow fair comparison between different methods, including those using a dataset pre-processing stage. For each dataset, 15 new challenging queries are introduced. Finally, a new set of 1M hard, semi-automatically cleaned distractors is selected. An extensive comparison of the state-of-the-art methods is performed on the new benchmark. Different types of methods are evaluated, ranging from local-feature-based to modern CNN based methods. The best results are achieved by taking the best of the two worlds. Most importantly, image retrieval appears far from being solved. Filip Radenovic, Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum |
CVPR | 5 |
| 2018 | Local Orthogonal-Group Testing
Ahmet Iscen, Ondrej Chum |
ECCV (2) | 2 |
| 2018 | Deep Shape Matching
Filip Radenovic, Giorgos Tolias, Ondrej Chum |
ECCV (5) | 3 |
| 2018 | Unsupervised Object Discovery for Instance RecognitionabstractSevere background clutter is challenging in many computer vision tasks, including large-scale image retrieval. Global descriptors, that are popular due to their memory and search efficiency, are especially prone to corruption by such a clutter. Eliminating the impact of the clutter on the image descriptor increases the chance of retrieving relevant images and prevents topic drift due to actually retrieving the clutter in the case of query expansion. In this work, we propose a novel salient region detection method. It captures, in an unsupervised manner, patterns that are both discriminative and common in the dataset. Saliency is based on a centrality measure of a nearest neighbor graph constructed from regional CNN representations of dataset images. The descriptors derived from the salient regions improve particular object retrieval, most noticeably in a large collections containing small objects. Oriane Siméoni, Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Ondrej Chum |
WACV | 5 |
| 2018 | Efficient contour match kernel
Giorgos Tolias, Ondrej Chum |
Image Vis. Comput. | 2 |
| 2017 | Multiple-Kernel Local-Patch Descriptor
Arun Mukundan, Giorgos Tolias, Ondrej Chum |
BMVC | 3 |
| 2017 | Efficient Diffusion on Region Manifolds: Recovering Small Objects with Compact CNN RepresentationsabstractQuery expansion is a popular method to improve the quality of image retrieval with both conventional and CNN representations. It has been so far limited to global image similarity. This work focuses on diffusion, a mechanism that captures the image manifold in the feature space. An efficient off-line stage allows optional reduction in the number of stored regions. In the on-line stage, the proposed handling of unseen queries in the indexing stage removes additional computation to adjust the precomputed data. We perform diffusion through a sparse linear system solver, yielding practical query times well below one second. Experimentally, we observe a significant boost in performance of image retrieval with compact CNN descriptors on standard benchmarks, especially when the query object covers only a small part of the image. Small objects have been a common failure case of CNN-based retrieval. Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Teddy Furon, Ondrej Chum |
CVPR | 5 |
| 2017 | Asymmetric Feature Maps with Application to Sketch Based RetrievalabstractWe propose a novel concept of asymmetric feature maps (AFM), which allows to evaluate multiple kernels between a query and database entries without increasing the memory requirements. To demonstrate the advantages of the AFM method, we derive a short vector image representation that, due to asymmetric feature maps, supports efficient scale and translation invariant sketch-based image retrieval. Unlike most of the short-code based retrieval systems, the proposed method provides the query localization in the retrieved image. The efficiency of the search is boosted by approximating a 2D translation search via trigonometric polynomial of scores by 1D projections. The projections are a special case of AFM. An order of magnitude speed-up is achieved compared to traditional trigonometric polynomials. The results are boosted by an image-based average query expansion, exceeding significantly the state of the art on standard benchmarks. Giorgos Tolias, Ondrej Chum |
CVPR | 2 |
| 2017 | Panorama to Panorama Matching for Location RecognitionabstractInternational audience Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Teddy Furon, Ondrej Chum |
ICMR | 5 |
| 2017 | Optimizing explicit feature maps on intervals
Ondrej Chum |
Image Vis. Comput. | 1 |
| 2016 | Coplanar Repeats by Energy Minimization
James Pritts, Denys Rozumnyi, M. Pawan Kumar, Ondrej Chum |
BMVC | 4 |
| 2016 | From Dusk Till Dawn: Modeling in the DarkabstractInternet photo collections naturally contain a large variety of illumination conditions, with the largest difference between day and night images. Current modeling techniques do not embrace the broad illumination range often leading to reconstruction failure or severe artifacts. We present an algorithm that leverages the appearance variety to obtain more complete and accurate scene geometry along with consistent multi-illumination appearance information. The proposed method relies on automatic scene appearance grouping, which is used to obtain separate dense 3D models. Subsequent model fusion combines the separate models into a complete and accurate reconstruction of the scene. In addition, we propose a method to derive the appearance information for the model under the different illumination conditions, even for scene parts that are not observed under one illumination condition. To achieve this, we develop a cross-illumination color transfer technique. We evaluate our method on a large variety of landmarks from across Europe reconstructed from a database of 7.4M images. Filip Radenovic, Johannes L. Schönberger, Dinghuang Ji, Jan-Michael Frahm, Ondrej Chum, Jiri Matas |
CVPR | 5 |
| 2016 | CNN Image Retrieval Learns from BoW: Unsupervised Fine-Tuning with Hard Examples
Filip Radenovic, Giorgos Tolias, Ondrej Chum |
ECCV (1) | 3 |
| 2016 | In the Saddle: Chasing fast and repeatable featuresabstractA novel similarity-covariant feature detector that extracts points whose neighborhoods, when treated as a 3D intensity surface, have a saddle-like intensity profile. The saddle condition is verified efficiently by intensity comparisons on two concentric rings that must have exactly two dark-to-bright and two bright-to-dark transitions satisfying certain geometric constraints. Experiments show that the Saddle features are general, evenly spread and appearing in high density in a range of images. The Saddle detector is among the fastest proposed. In comparison with detector with similar speed, the Saddle features show superior matching performance on number of challenging datasets. Javier Aldana-Iuit, Dmytro Mishkin, Ondrej Chum, Jiri Matas |
ICPR | 3 |
| 2015 | Camera Elevation Estimation from a Single Mountain Landscape PhotographabstractThis work addresses the problem of camera elevation estimation from a single photograph in an outdoor environment. We introduce a new benchmark dataset of one-hundred thousand images with annotated camera elevation called Alps100K. We propose and experimentally evaluate two automatic data-driven approaches to camera elevation estimation: one based on convolutional neural networks, the other on local features. To compare the proposed methods to human performance, an experiment with 100 subjects is conducted. The experimental results show that both proposed approaches outperform humans and that the best result is achieved by their combination. Martin Cadík, Jan Vasícek, Michal Hradis, Filip Radenovic, Ondrej Chum |
BMVC | 5 |
| 2015 | From single image query to detailed 3D reconstructionabstractStructure-from-Motion for unordered image collections has significantly advanced in scale over the last decade. This impressive progress can be in part attributed to the introduction of efficient retrieval methods for those systems. While this boosts scalability, it also limits the amount of detail that the large-scale reconstruction systems are able to produce. In this paper, we propose a joint reconstruction and retrieval system that maintains the scalability of large-scale Structure-from-Motion systems while also recovering the often lost ability of reconstructing fine details of the scene. We demonstrate our proposed method on a large-scale dataset of 7.4 million images downloaded from the Internet. Johannes L. Schönberger, Filip Radenovic, Ondrej Chum, Jan-Michael Frahm |
CVPR | 3 |
| 2015 | Low Dimensional Explicit Feature MapsabstractApproximating non-linear kernels by finite-dimensional feature maps is a popular approach for speeding up training and evaluation of support vector machines or to encode information into efficient match kernels. We propose a novel method of data independent construction of low dimensional feature maps. The problem is cast as a linear program which jointly considers competing objectives: the quality of the approximation and the dimensionality of the feature map. For both shift-invariant and homogeneous kernels the proposed method achieves a better approximations at the same dimensionality or comparable approximations at lower dimensionality of the feature map compared with state-of-the-art methods. Ondrej Chum |
ICCV | 1 |
| 2015 | Towards visual words to wordsabstractWe address the problem of text localization and retrieval in real world images. We are first to study the retrieval of text images, i.e. the selection of images containing text in large collections at high speed. We propose a novel representation, textual visual words, which describe text by generic visual words that geometrically consistently predict bottom and top lines of text. The visual words are discretized SIFT descriptors of Hessian features. The features may correspond to various structures present in the text - character fragments, individual characters or their arrangements. The textual words representation is invariant to affine transformation of the image and local linear change of intensity. Experiments demonstrate that the proposed method outperforms the state-of-the-art on the MS dataset. The proposed method detects blurry, small font, low contrast, noisy text from real world images. Rakesh Mehta, Ondrej Chum, Jiri Matas |
ICDAR | 2 |
| 2015 | Multiple Measurements and Joint Dimensionality Reduction for Large Scale Image Search with Short VectorsabstractThis paper addresses the construction of a short-vector (128D) image representation for large-scale image and particular object retrieval. In particular, the method of joint dimensionality reduction of multiple vocabularies is considered. We study a variety of vocabulary generation techniques: different k-means initializations, different descriptor transformations, different measurement regions for descriptor extraction. Our extensive evaluation shows that different combinations of vocabularies, each partitioning the descriptor space in a different yet complementary manner, results in a significant performance improvement, which exceeds the state-of-the-art. Filip Radenovic, Hervé Jégou, Ondrej Chum |
ICMR | 3 |
| 2014 | Efficient Image Detail Mining
Andrej Mikulík, Filip Radenovic, Ondrej Chum, Jiri Matas |
ACCV (2) | 3 |
| 2014 | Rectification, and Segmentation of Coplanar Repeated PatternsabstractThis paper presents a novel and general method for the detection, rectification and segmentation of imaged coplanar repeated patterns. The only assumption made of the scene geometry is that repeated scene elements are mapped to each other by planar Euclidean transformations. The class of patterns covered is broad and includes nearly all commonly seen, planar, man-made repeated patterns. In addition, novel linear constraints are used to reduce geometric ambiguity between the rectified imaged pattern and the scene pattern. Rectification to within a similarity of the scene plane is achieved from one rotated repeat, or to within a similarity with a scale ambiguity along the axis of symmetry from one reflected repeat. A stratum of constraints is derived that gives the necessary configuration of repeats for each successive level of rectification. A generative model for the imaged pattern is inferred and used to segment the pattern with pixel accuracy. Qualitative results are shown on a broad range of image types on which state-of-the-art methods fail. James Pritts, Ondrej Chum, Jiri Matas |
CVPR | 2 |
| 2014 | Erratum to: Learning Vocabularies over a Fine Quantization
Andrej Mikulík, Michal Perdoch, Ondrej Chum, Jiri Matas |
Int. J. Comput. Vis. | 3 |
| 2013 | Image Retrieval for Online Browsing in Large Image Collections
Andrej Mikulík, Ondrej Chum, Jiri Matas |
SISAP | 2 |
| 2013 | Learning Vocabularies over a Fine Quantization
Andrej Mikulík, Michal Perdoch, Ondrej Chum, Jiri Matas |
Int. J. Comput. Vis. | 3 |
| 2013 | USAC: A Universal Framework for Random Sample ConsensusabstractA computational problem that arises frequently in computer vision is that of estimating the parameters of a model from data that have been contaminated by noise and outliers. More generally, any practical system that seeks to estimate quantities from noisy data measurements must have at its core some means of dealing with data contamination. The random sample consensus (RANSAC) algorithm is one of the most popular tools for robust estimation. Recent years have seen an explosion of activity in this area, leading to the development of a number of techniques that improve upon the efficiency and robustness of the basic RANSAC algorithm. In this paper, we present a comprehensive overview of recent research in RANSAC-based robust estimation by analyzing and comparing various approaches that have been explored over the years. We provide a common context for this analysis by introducing a new framework for robust estimation, which we call Universal RANSAC (USAC). USAC extends the simple hypothesize-and-verify structure of standard RANSAC to incorporate a number of important practical and computational considerations. In addition, we provide a general-purpose C++ software library that implements the USAC framework by leveraging state-of-the-art algorithms for the various modules. This implementation thus addresses many of the limitations of standard RANSAC within a single unified package. We benchmark the performance of the algorithm on a large collection of estimation problems. The implementation we provide can be used by researchers either as a stand-alone tool for robust estimation or as a benchmark for evaluating new techniques. Rahul Raguram, Ondrej Chum, Marc Pollefeys, Jiri Matas, Jan-Michael Frahm |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Fixing the Locally Optimized RANSACabstractThe paper revisits the problem of local optimization for RANSAC. Improvements of the LO-RANSAC procedure are proposed: a use of truncated quadratic cost function, an introduction of a limit on the number of inliers used for the least squares computation and several implementation issues are addressed. The implementation is made publicly available. Extensive experiments demonstrate that the novel algorithm called LO +-RANSAC is (1) very stable (almost non-random in nature), (2) very precise in a broad range of conditions, (3) less sensitive to the choice of inlier-outlier threshold and (4) it offers a significantly better starting point for bundle adjustment than the Gold Standard method advocated in the Hartley-Zisserman book. 1 Karel Lebeda, Jiri Matas, Ondrej Chum |
BMVC | 3 |
| 2012 | Fast computation of min-Hash signatures for image collectionsabstractA new method for highly efficient min-Hash generation for document collections is proposed. It exploits the inverted file structure which is available in many applications based on a bag or a set of words. Fast min-Hash generation is important in applications such as image clustering where good recall and precision requires a large number of min-Hash signatures. Using the set of words represenation, the novel exact min-Hash generation algorithm achieves approximately a 50-fold speed-up on two dataset with 105and 106images respectively. We also propose an approximate min-Hash assignment process which reaches a more than 200-fold speed-up at the cost of missing about 2–3% of matches. We also experimentally show that the method generalizes to other modalities with significantly different statistics. Ondrej Chum, Jiri Matas |
CVPR | 1 |
| 2012 | Negative Evidences and Co-occurences in Image Retrieval: The Benefit of PCA and Whitening
Hervé Jégou, Ondrej Chum |
ECCV (2) | 2 |
| 2012 | Homography estimation from correspondences of local elliptical features
Ondrej Chum, Jiri Matas |
ICPR | 1 |
| 2011 | Total recall II: Query expansion revisitedabstractMost effective particular object and image retrieval approaches are based on the bag-of-words (BoW) model. All state-of-the-art retrieval results have been achieved by methods that include a query expansion that brings a significant boost in performance. We introduce three extensions to automatic query expansion: (i) a method capable of preventing tf-idf failure caused by the presence of sets of correlated features (confusers), (ii) an improved spatial verification and re-ranking step that incrementally builds a statistical model of the query object and (iii) we learn relevant spatial context to boost retrieval performance. The three improvements of query expansion were evaluated on standard Paris and Oxford datasets according to a standard protocol, and state-of-the-art results were achieved. Ondrej Chum, Andrej Mikulík, Michal Perdoch, Jiri Matas |
CVPR | 1 |
| 2010 | Planar Affine Rectification from Change of Scale
Ondrej Chum, Jiri Matas |
ACCV (4) | 1 |
| 2010 | Unsupervised discovery of co-occurrence in sparse high dimensional dataabstractAn efficient min-Hash based algorithm for discovery of dependencies in sparse high-dimensional data is presented. The dependencies are represented by sets of features co-occurring with high probability and are called co-ocsets. Sparse high dimensional descriptors, such as bag of words, have been proven very effective in the domain of image retrieval. To maintain high efficiency even for very large data collection, features are assumed independent. We show experimentally that co-ocsets are not rare, i.e. the independence assumption is often violated, and that they may ruin retrieval performance if present in the query image. Two methods for managing co-ocsets in such cases are proposed. Both methods significantly outperform the state-of-the-art in image retrieval, one is also significantly faster. Ondrej Chum, Jiri Matas |
CVPR | 1 |
| 2010 | Learning a Fine Vocabulary
Andrej Mikulík, Michal Perdoch, Ondrej Chum, Jiri Matas |
ECCV (3) | 3 |
| 2010 | Image Matching and Retrieval by Repetitive PatternsabstractDetection of repetitive patterns in images has been studied for a long time in computer vision. This paper discusses a method for representing a lattice or line pattern by shift-invariant descriptor of the repeating element. The descriptor overcomes shift ambiguity and can be matched between different a views. The pattern matching is then demonstrated in retrieval experiment, where different images of the same buildings are retrieved solely by repetitive patterns. Petr Doubek, Jiri Matas, Michal Perdoch, Ondrej Chum |
ICPR | 4 |
| 2010 | Construction of Precise Local Affine FramesabstractWe propose a novel method for the refinement of Maximally Stable Extremal Region (MSER) boundaries to sub-pixel precision by taking into account the intensity function in the 2 × 2 neighborhood of the contour points. The proposed method improves the repeatability and precision of Local Affine Frames (LAFs) constructed on extremal regions. Additionally, we propose a novel method for detection of local curvature extrema on the refined contour. Experimental evaluation on publicly available datasets shows that matching with the modified LAFs leads to a higher number of correspondences and a higher inlier ratio in more than 80% of the test image pairs. Since the processing time of the contour refinement is negligible, there is no reason not to include the algorithms as a standard part of the MSER detector and LAF constructions. Andrej Mikulík, Jiri Matas, Michal Perdoch, Ondrej Chum |
ICPR | 4 |
| 2010 | Large-Scale Discovery of Spatially Related ImagesabstractWe propose a randomized data mining method that finds clusters of spatially overlapping images. The core of the method relies on the min-Hash algorithm for fast detection of pairs of images with spatial overlap, the so-called cluster seeds. The seeds are then used as visual queries to obtain clusters which are formed as transitive closures of sets of partially overlapping images that include the seed. We show that the probability of finding a seed for an image cluster rapidly increases with the size of the cluster. The properties and performance of the algorithm are demonstrated on data sets with 10(4), 10(5), and 5 x 10(6) images. The speed of the method depends on the size of the database and the number of clusters. The first stage of seed generation is close to linear for databases sizes up to approximately 2(34) approximately 10(10) images. On a single 2.4 GHz PC, the clustering process took only 24 minutes for a standard database of more than 100,000 images, i.e., only 0.014 seconds per image. Ondrej Chum, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Geometric min-Hashing: Finding a (thick) needle in a haystackabstractWe propose a novel hashing scheme for image retrieval, clustering and automatic object discovery. Unlike commonly used bag-of-words approaches, the spatial extent of image features is exploited in our method. The geometric information is used both to construct repeatable hash keys and to increase the discriminability of the description. Each hash key combines visual appearance (visual words) with semi-local geometric information. Compared with the state-of-the-art min-hash, the proposed method has both higher recall (probability of collision for hashes on the same object) and lower false positive rates (random collisions). The advantages of geometric min-hashing approach are most pronounced in the presence of viewpoint and scale change, significant occlusion or small physical overlap of the viewing fields. We demonstrate the power of the proposed method on small object discovery in a large unordered collection of images and on a large scale image clustering problem. Ondrej Chum, Michal Perdoch, Jiri Matas |
CVPR | 1 |
| 2009 | Efficient representation of local geometry for large scale object retrievalabstractState of the art methods for image and object retrieval exploit both appearance (via visual words) and local geometry (spatial extent, relative pose). In large scale problems, memory becomes a limiting factor - local geometry is stored for each feature detected in each image and requires storage larger than the inverted file and term frequency and inverted document frequency weights together. We propose a novel method for learning discretized local geometry representation based on minimization of average reprojection error in the space of ellipses. The representation requires only 24 bits per feature without drop in performance. Additionally, we show that if the gravity vector assumption is used consistently from the feature description to spatial verification, it improves retrieval performance and decreases the memory footprint. The proposed method outperforms state of the art retrieval algorithms in a standard image retrieval benchmark. Michal Perdoch, Ondrej Chum, Jiri Matas |
CVPR | 2 |
| 2008 | Near Duplicate Image Detection: min-Hash and tf-idf Weighting
Ondrej Chum, James Philbin, Andrew Zisserman |
BMVC | 1 |
| 2008 | Lost in quantization: Improving particular object retrieval in large scale image databasesabstractThe state of the art in visual object retrieval from large databases is achieved by systems that are inspired by text retrieval. A key component of these approaches is that local regions of images are characterized using high-dimensional descriptors which are then mapped to ldquovisual wordsrdquo selected from a discrete vocabulary.This paper explores techniques to map each visual region to a weighted set of words, allowing the inclusion of features which were lost in the quantization stage of previous systems. The set of visual words is obtained by selecting words based on proximity in descriptor space. We describe how this representation may be incorporated into a standard tf-idf architecture, and how spatial verification is modified in the case of this soft-assignment. We evaluate our method on the standard Oxford Buildings dataset, and introduce a new dataset for evaluation. Our results exceed the current state of the art retrieval performance on these datasets, particularly on queries with poor initial recall where techniques like query expansion suffer. Overall we show that soft-assignment is always beneficial for retrieval with large vocabularies, at a cost of increased storage requirements for the index. James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, Andrew Zisserman |
CVPR | 2 |
| 2008 | Optimal Randomized RANSACabstractA randomized model verification strategy for RANSAC is presented. The proposed method finds, like RANSAC, a solution that is optimal with user-specified probability. The solution is found in time that is (i) close to the shortest possible and (ii) superior to any deterministic verification strategy. A provably fastest model verification strategy is designed for the (theoretical) situation when the contamination of data by outliers is known. In this case, the algorithm is the fastest possible (on average) of all randomized \\RANSAC algorithms guaranteeing a confidence in the solution. The derivation of the optimality property is based on Wald's theory of sequential decision making, in particular a modified sequential probability ratio test (SPRT). Next, the R-RANSAC with SPRT algorithm is introduced. The algorithm removes the requirement for a priori knowledge of the fraction of outliers and estimates the quantity online. We show experimentally that on standard test data the method has performance close to the theoretically optimal and is 2 to 10 times faster than standard RANSAC and is up to 4 times faster than previously published methods. Ondrej Chum, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | An Exemplar Model for Learning Object ClassesabstractWe introduce an exemplar model that can learn and generate a region of interest around class instances in a training set, given only a set of images containing the visual class. The model is scale and translation invariant. In the training phase, image regions that optimize an objective function are automatically located in the training images, without requiring any user annotation such as bounding boxes. The objective function measures visual similarity between training image pairs, using the spatial distribution of both appearance patches and edges. The optimization is initialized using discriminative features. The model enables the detection (localization) of multiple instances of the object class in test images, and can be used as a precursor to training other visual models that require bounding box annotation. The detection performance of the model is assessed on the PASCAL Visual Object Classes Challenge 2006 test set. For a number of object classes the performance far exceeds the current state of the art of fully supervised methods. Ondrej Chum, Andrew Zisserman |
CVPR | 1 |
| 2007 | Object retrieval with large vocabularies and fast spatial matchingabstractIn this paper, we present a large-scale object retrieval system. The user supplies a query object by selecting a region of a query image, and the system returns a ranked list of images that contain the same object, retrieved from a large corpus. We demonstrate the scalability and performance of our system on a dataset of over 1 million images crawled from the photo-sharing site, Flickr [3], using Oxford landmarks as queries. Building an image-feature vocabulary is a major time and performance bottleneck, due to the size of our dataset. To address this problem we compare different scalable methods for building a vocabulary and introduce a novel quantization method based on randomized trees which we show outperforms the current state-of-the-art on an extensive ground-truth. Our experiments show that the quantization has a major effect on retrieval quality. To further improve query performance, we add an efficient spatial verification stage to re-rank the results returned from our bag-of-words model and show that this consistently improves search quality, though by less of a margin when the visual vocabulary is large. We view this work as a promising step towards much larger, "web-scale" image corpora. James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, Andrew Zisserman |
CVPR | 2 |
| 2007 | Total Recall: Automatic Query Expansion with a Generative Feature Model for Object RetrievalabstractGiven a query image of an object, our objective is to retrieve all instances of that object in a large (1M+) image database. We adopt the bag-of-visual-words architecture which has proven successful in achieving high precision at low recall. Unfortunately, feature detection and quantization are noisy processes and this can result in variation in the particular visual words that appear in different images of the same object, leading to missed results. In the text retrieval literature a standard method for improving performance is query expansion. A number of the highly ranked documents from the original query are reissued as a new query. In this way, additional relevant terms can be added to the query. This is a form of blind relevance feedback and it can fail if `outlier' (false positive) documents are included in the reissued query. In this paper we bring query expansion into the visual domain via two novel contributions. Firstly, strong spatial constraints between the query image and each result allow us to accurately verify each return, suppressing the false positives which typically ruin text-based query expansion. Secondly, the verified images can be used to learn a latent feature model to enable the controlled construction of expanded queries. We illustrate these ideas on the 5000 annotated image Oxford building database together with more than 1M Flickr images. We show that the precision is substantially boosted, achieving total recall in many cases. Ondrej Chum, James Philbin, Josef Sivic, Michael Isard, Andrew Zisserman |
ICCV | 1 |
| 2006 | Geometric Hashing with Local Affine FramesabstractWe propose a novel representation of local image structure and a matching scheme that are insensitive to a wide range of appearance changes. The representation is a collection of local affine frames that are constructed on outer boundaries of maximally stable extremal regions (MSERS) in an affine-covariant way. Each local affine frame is described by a relative location of other local affine frames in its neighborhood. The image is thus represented by quantities that depend only on the location of the boundaries of MSERs. Inter-image correspondences between local affine frames are formed in constant time by geometric hashing. Direct detection of local afine frames removes the requirement of a point-based hashing to establish reference frames in a combinatorial way, which has in the case of affine transform complexily that is cubic in the number of points. Local affine frames, which are also the quantities represented in the hash table, occupy a 6 0 space and hence data collisions are less likely compared with 2 0 point hashing. Experimentally, the robustness of the method and its insensitiviq to photometric changes is demonstrated on images from different spectral bands of satellite sensor; on images of a transparent object and on images of an object taken during day and night. Ondrej Chum, Jiri Matas |
CVPR (1) | 1 |
| 2005 | Matching with PROSAC - Progressive Sample ConsensusabstractA new robust matching method is proposed. The progressive sample consensus (PROSAC) algorithm exploits the linear ordering defined on the set of correspondences by a similarity function used in establishing tentative correspondences. Unlike RANSAC, which treats all correspondences equally and draws random samples uniformly from the full set, PROSAC samples are drawn from progressively larger sets of top-ranked correspondences. Under the mild assumption that the similarity measure predicts correctness of a match better than random guessing, we show that PROSAC achieves large computational savings. Experiments demonstrate it is often significantly faster (up to more than hundred times) than RANSAC. For the derived size of the sampled set of correspondences as a function of the number of samples already drawn, PROSAC converges towards RANSAC in the worst case. The power of the method is demonstrated on wide-baseline matching problems. Ondrej Chum, Jiri Matas |
CVPR (1) | 1 |
| 2005 | Two-View Geometry Estimation Unaffected by a Dominant PlaneabstractA RANSAC-based algorithm for robust estimation of epipolar geometry from point correspondences in the possible presence of a dominant scene plane is presented. The algorithm handles scenes with (i) all points in a single plane, (ii) majority of points in a single plane and the rest off the plane, (iii) no dominant plane. It is not required to know a priori which of the cases (i)-(iii) occurs. The algorithm exploits a theorem we proved, that if five or more of seven correspondences are related by a homography then there is an epipolar geometry consistent with the seven-tuple as well as with all correspondences related by the homography. This means that a seven point sample consisting of two outliers and five inliers lying in a dominant plane produces an epipolar geometry which is wrong and yet consistent with a high number of correspondences. The theorem explains why RANSAC often fails to estimate epipolar geometry in the presence of a dominant plane. Rather surprisingly, the theorem also implies that RANSAC-based homography estimation is faster when drawing nonminimal samples of seven correspondences than minimal samples of four correspondences. Ondrej Chum, Tomás Werner, Jiri Matas |
CVPR (1) | 1 |
| 2005 | Randomized RANSAC with Sequential Probability Ratio TestabstractA randomized model verification strategy for RANSAC is presented. The proposed method finds, like RANSAC, a solution that is optimal with user-controllable probability n. A provably optimal model verification strategy is designed for the situation when the contamination of data by outliers is known, i.e. the algorithm is the fastest possible (on average) of all randomized RANSAC algorithms guaranteeing 1 - n confidence in the solution. The derivation of the optimality property is based on Wald's theory of sequential decision making. The R-RANSAC with SPRT which does not require the a priori knowledge of the fraction of outliers and has results close to the optimal strategy is introduced. We show experimentally that on standard test data the method is 2 to 10 times faster than the standard RANSAC and up to 4 times faster than previously published methods. Jiri Matas, Ondrej Chum |
ICCV | 2 |
| 2005 | The geometric error for homographies
Ondrej Chum, Tomás Pajdla, Peter F. Sturm |
Comput. Vis. Image Underst. | 1 |
| 2004 | Randomized RANSAC with Td, d test
Jiri Matas, Ondrej Chum |
Image Vis. Comput. | 2 |
| 2004 | Robust wide-baseline stereo from maximally stable extremal regions
Jiri Matas, Ondrej Chum, Tomás Pajdla |
Image Vis. Comput. | 2 |
| 2003 | Joint Orientation of EpipolesabstractIt is known that epipolar constraint can be augmented with orientation by formulating it in the oriented projective geometry. This oriented epipolar constraint requires knowing the orientations (signs of overall scales) of epipoles and fundamental matrix. The current belief is that these orientations cannot be obtained from the fundamental matrix only and that additional information is needed, typically, a single correct point correspondence. In contrary to this, we show that fundamental matrix alone encodes orientation of epipoles up to their common scale sign. We present two formulations of this fact. The algebraic formulation gives a closed formula to compute the second epipole from fundamental matrix and the first epipole. The geometric formulation is in terms of the conic formed by intersections of corresponding epipolar lines in the common image plane; we show that the epipoles always lie on different antipodal components of the spherical interpretation of this conic. Further, we show that, under mild assumptions, fundamental matrix can discriminate between two classes of mutual position of a pair of directional cameras. 1 Ondrej Chum, Tomás Werner, Tomás Pajdla |
BMVC | 1 |
| 2002 | Randomized RANSAC with T(d, d) testabstractMany computer vision algorithms include a robust estimation step where model parameters are computed from a data set containing a significant proportion of outliers. The RANSAC algorithm is possibly the most widely used robust estimator in the field of computer vision. In the paper we show that under a broad range of conditions, RANSAC efficiency is significantly improved if its hypothesis evaluation step is randomized. A new Jiri Matas, Ondrej Chum |
BMVC | 2 |
| 2002 | Robust Wide Baseline Stereo from Maximally Stable Extremal RegionsabstractAbstract The wide-baseline stereo problem, i.e. the problem of establishing correspondences between a pair of images taken from different viewpoints is studied. A new set of image elements that are put into correspondence, the so called extremal regions , is introduced. Extremal regions possess highly desirable properties: the set is closed under (1) continuous (and thus projective) transformation of image coordinates and (2) monotonic transformation of image intensities. An efficient (near linear complexity) and practically fast detection algorithm (near frame rate) is presented for an affinely invariant stable subset of extremal regions, the maximally stable extremal regions (MSER). A new robust similarity measure for establishing tentative correspondences is proposed. The robustness ensures that invariants from multiple measurement regions (regions obtained by invariant constructions from extremal regions), some that are significantly larger (and hence discriminative) than the MSERs, may be used to establish tentative correspondences. The high utility of MSERs, multiple measurement regions and the robust metric is demonstrated in wide-baseline experiments on image pairs from both indoor and outdoor scenes. Significant change of scale (3.5×), illumination conditions, out-of-plane rotation, occlusion, locally anisotropic scale change and 3D translation of the viewpoint are all present in the test problems. Good estimates of epipolar geometry (average distance from corresponding points to the epipolar line below 0.09 of the inter-pixel distance) are obtained. Jiri Matas, Ondrej Chum, Tomás Pajdla |
BMVC | 2 |