VLDB 2026 Research / reviewers in the wild / expert
Jiri Matas
dblp:m/JiriMatas · also George Matas, Jirí Matas
· DBLP profile ↗
261ranked-venue papers
22as first author
61since 2021 · last 2026
0000-0003-0863-4844ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 223 · 20 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 167 · 15 first-author · 40 since 2021Databases, data management, data science and information retrieval · 15 · 5 since 2021Systems, architecture and hardware · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identity Verification from Human Scent using Channel Representation of 2D Gas Chromatography-Mass Spectrometry DataabstractThis study examines the feasibility of employing raw two-dimensional gas chromatography/time-of-flight mass spectrometry (GC×GC–ToF-MS) data for the purpose of human scent identity verification. Unlike techniques that require expert-driven identification of compounds, our framework transforms each GC×GC–ToF-MS sample into a multi-channel image. A comprehensive assessment has been conducted on ten feature-embedding methods, and five spatial-alignment strategies.The evaluation is performed on a newly assembled dataset of 252 individuals, comprising 2,528 raw samples and aggregating around 7.5 TB of data. On our dataset, the best method reaches ≈ 53% true positive rate at a 5% false positive rate without explicit spatial registration. Although this performance is below that of well-established biometrics (e.g., iris verification), the results demonstrate the feasibility of raw-odor verification for scenarios where direct line-of-sight or cooperation may be limited and suggest that, for our acquisition/drift regime, explicit registration is not required to reach the best performance.Code and dataset on github.com. Radim Spetlík, Jan Hlavsa, Jana Cechová, Petra Pojmanová, Jiri Matas, Stepán Urban |
WACV | 5 |
| 2026 | Deepfake Detection that Generalizes Across BenchmarksabstractThe generalization of deepfake detectors to unseen manipulation techniques remains a challenge for practical deployment. Although many approaches adapt foundation models by introducing significant architectural complexity, this work demonstrates that robust generalization is achievable through a parameter-efficient adaptation of one of the foundational pre-trained vision encoders. The proposed method, GenD, fine-tunes only the Layer Normalization parameters (0.03% of the total) and enhances generalization by enforcing a hyperspherical feature manifold using L2 normalization and metric learning on it. We conducted an extensive evaluation on 14 benchmark datasets spanning from 2019 to 2025. The proposed method achieves state-of-the-art performance, outperforming more complex, recent approaches in average cross-dataset AUROC. Our analysis yields two primary findings for the field: 1) training on paired real-fake data from the same source video is essential for mitigating shortcut learning and improving generalization, and 2) detection difficulty on academic datasets has not strictly increased over time, with models trained on older, diverse datasets showing strong generalization capabilities. This work delivers a computationally efficient and reproducible method, proving that state-of-the-art generalization is attainable by making targeted, minimal changes to a pre-trained foundational image encoder model. The code is at: https://github.com/yermandy/GenD Andrii Yermakov, Jan Cech, Jiri Matas, Mario Fritz |
WACV | 3 |
| 2026 | Underwater visual tracking with a large scale dataset and image enhancementabstractThis paper presents a new dataset and general tracker enhancement method for Underwater Visual Object Tracking (UVOT). Despite its significance, underwater tracking has remained unexplored due to data inaccessibility. It poses distinct challenges; the underwater environment exhibits non-uniform lighting conditions, low visibility, lack of sharpness, low contrast, camouflage, and reflections from suspended particles. Performance of traditional tracking methods designed primarily for terrestrial or open-air scenarios drops in such conditions. We address the problem by proposing a novel underwater image enhancement algorithm designed specifically to boost tracking quality. The method has resulted in a significant performance improvement, of up to 5.0% AUC, of state-of-the-art (SOTA) visual trackers. To develop robust and accurate UVOT methods, large-scale datasets are required. To this end, we introduce a large-scale UVOT benchmark dataset consisting of 400 video segments and 275,000 manually annotated frames enabling underwater training and evaluation of deep trackers. The videos are labelled with several underwater-specific tracking attributes including watercolor variation, target distractors, camouflage, target relative size, and low visibility conditions. The UVOT400 dataset, tracking results, and the code are publicly available on: https://github.com/BasitAlawode/UVOT400 . Basit Alawode, Sajid Javed, Fayaz Ali Dharejo, Mehnaz Ummar, Arif Mahmoud, Fahad Shahbaz Khan, Jiri Matas |
Neurocomputing | 7 |
| 2026 | PixOOD: Pixel-Level Out-of-Distribution DetectionabstractWe propose a pixel-level out-of-distribution detection algorithm, called PixOOD, which does not require training on samples of anomalous data and is not designed for a specific application which avoids traditional training biases. The PixOOD consists of two main parts - in-distribution data model and decision strategy estimator. In order to model the complex intra-class variability of the in-distribution data at the pixel-level, we propose an online data condensation algorithm which is more robust than standard K-means and is easily trainable through (stochastic) gradient descent techniques. Furthermore, we propose two models for estimating decision strategy, per-class and unified calibration models, each suitable for different applications. We evaluate PixOOD on a wide range of problems. It achieved state-of-the-art results on four out of seven datasets, while being competitive on the rest. Tomás Vojír, Jan Sochman, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | ILIAS: Instance-Level Image retrieval At ScaleabstractThis work introduces ILIAS, a new test dataset for Instance-Level Image retrieval At Scale. It is designed to evaluate the ability of current and future foundation models and retrieval techniques to recognize particular objects. The key benefits over existing datasets include large scale, domain diversity, accurate ground truth, and a performance that is far from saturated. ILIAS includes query and positive images for 1,000 object instances, manually collected to capture challenging conditions and diverse domains. Large-scale retrieval is conducted against 100 million distractor images from YFCC100M. To avoid false negatives without extra annotation effort, we include only query objects confirmed to have emerged after 2014, i.e. the compilation date of YFCC100M. An extensive benchmarking is performed with the following observations: i) models fine-tuned on specific domains, such as landmarks or products, excel in that domain but fail on ILIAS ii) learning a linear adaptation layer using multi-domain class supervision results in performance improvements, especially for vision-language models iii) local descriptors in retrieval re-ranking are still a key ingredient, especially in the presence of severe background clutter iv) the text-to-image performance of the vision-language foundation models is surprisingly close to the corresponding image-to-image case. website: https://vrg.fel.cvut.cz/ilias/ Giorgos Kordopatis-Zilos, Vladan Stojnic, Anna Manko, Pavel Suma, Nikolaos-Antonios Ypsilantis, Nikos Efthymiadis, Zakaria Laskar, Jiri Matas, Ondrej Chum, Giorgos Tolias |
CVPR | 8 |
| 2025 | A Dataset for Semantic Segmentation in the Presence of UnknownsabstractBefore deployment in the real-world deep neural networks require thorough evaluation of how they handle both knowns, inputs represented in the training data, and unknowns (anomalies). This is especially important for scene understanding tasks with safety critical applications, such as in autonomous driving. Existing datasets allow evaluation of only knowns or unknowns - but not both, which is required to establish "in the wild" suitability of deep neural network models. To bridge this gap, we propose a novel anomaly segmentation dataset, ISSU, that features a diverse set of anomaly inputs from cluttered real-world environments. The dataset is twice larger than existing anomaly segmentation datasets, and provides a training, validation and test set for controlled in-domain evaluation. The test set consists of a static and temporal part, with the latter comprised of videos. The dataset provides annotations for both closed-set (knowns) and anomalies, enabling closed-set and open-set evaluation. The dataset covers diverse conditions, such as domain and cross-sensor shift, illumination variation and allows ablation of anomaly detection methods with respect to these variations. Evaluation results of current state-of-the-art methods confirm the need for improvements especially in domain-generalization, small and large object segmentation. The code and the dataset are available at https://github.com/vojirt/benchmark_issu. Zakaria Laskar, Tomás Vojír, Matej Grcic, Iaroslav Melekhov, Shankar Gangisetty, Juho Kannala, Jiri Matas, Giorgos Tolias, C. V. Jawahar |
CVPR | 7 |
| 2025 | ProbPose: A Probabilistic Approach to 2D Human Pose EstimationabstractCurrent state-of-the-art Human Pose Estimation methods ignore out-of-image keypoints in both training and evaluation and use uncalibrated heatmaps as keypoint location representations. We propose ProbPose, which predicts for each keypoint: a calibrated probability of keypoint presence at each location in the activation window, the probability of being outside of it, and its predicted visibility. To address the lack of evaluation protocols for out-of-image keypoints, we introduce the CropCOCO dataset and the Extended OKS (Ex-OKS) metric, which extends OKS to out-of-image points. Tested on COCO, CropCOCO, and OCHuman, ProbPose shows significant gains in out-of-image keypoint localization while also improving in-image localization through data augmentation. Additionally, the model improves robustness along the edges of the bounding box and offers better flexibility in keypoint evaluation. The code and weights are available on the project website1. Miroslav Purkrábek, Jiri Matas |
CVPR | 2 |
| 2025 | LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic SegmentationabstractWe propose a training-free method for open-vocabulary semantic segmentation using Vision-and-Language Models (VLMs). Our approach enhances the initial per-patch predictions of VLMs through label propagation, which jointly optimizes predictions by incorporating patch-to-patch relationships. Since VLMs are primarily optimized for cross-modal alignment and not for intra-modal similarity, we use a Vision Model (VM) that is observed to better capture these relationships. We address resolution limitations inherent to patch-based encoders by applying label propagation at the pixel level as a refinement step, significantly improving segmentation accuracy near class boundaries. Our method, called LPOSS+, performs inference over the entire image, avoiding window-based processing and thereby capturing contextual interactions across the full image. LPOSS+ achieves state-of-the-art performance among training-free methods, across a diverse set of datasets. Code: https://github.com/vladan-stojnic/LPOSS Vladan Stojnic, Yannis Kalantidis, Jiri Matas, Giorgos Tolias |
CVPR | 3 |
| 2025 | LifeCLEF 2025 Teaser: Challenges on Species Presence Prediction and Identification, and Individual Animal Identification
Alexis Joly, Lukás Picek, Stefan Kahl, Hervé Goëau, Lukás Adam, Christophe Botella, Maximilien Servajean, Diego Marcos, César Leblanc, Théo Larcher, Jiri Matas, Klára Janousková, Vojtech Cermák, Kostas Papafitsoros, Robert Planqué, Willem-Pier Vellinga, Holger Klinck, Tom Denton, Pierre Bonnet, Henning Müller |
ECIR (5) | 11 |
| 2025 | Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous CircleabstractHuman pose estimation methods work well on isolated people but struggle with multiple-bodies-in-proximity scenarios. Previous work has addressed this problem by conditioning pose estimation by detected bounding boxes or keypoints, but overlooked instance masks. We propose to iteratively enforce mutual consistency of bounding boxes, instance masks, and poses. The introduced BBox-Mask-Pose (BMP) method uses three specialized models that improve each other's output in a closed loop. All models are adapted for mutual conditioning, which improves robustness in multi-body scenes. MaskPose, a new mask-conditioned pose estimation model, is the best among top-down approaches on OCHuman. BBox-Mask-Pose pushes SOTA on OCHuman dataset in all three tasks - detection, instance segmentation, and pose estimation. It also achieves SOTA performance on COCO pose estimation. The method is especially good in scenes with large instances overlap, where it improves detection by 39% over the baseline detector. With small specialized models and faster runtime, BMP is an effective alternative to large human-centered foundational models. Code and models are available on https://MiraPurkrabek.github.io/BBox-Mask-Pose. Miroslav Purkrábek, Jiri Matas |
ICCV | 2 |
| 2025 | Auto-calibration of Camera Intrinsics and Extrinsics using Lidar and MotionabstractA novel camera autocalibration method is presented. Any camera model can be calibrated, and no calibration targets like checkerboards are used. The method requires the camera to be mounted on a lidar-equipped moving platform travelling through a structured environment along a known path. The primary reason for cross-modal camera calibration is not to solve the sensor fusion problem, but to tap the huge amount of accurate metric data points available from the lidar. The amount of measurements is easily four orders of magnitude higher than in checkerboard based methods. This leads to improved estimation accuracy, especially of higher-order distortion coefficients. In a multi-camera setup, the lidar additionally defines a common reference coordinate system for all cameras.Compared to the majority of published methods on camera-lidar autocalibration, (i) our calibration procedure relies on motion features, (ii) the hard-to-obtain-accurately lidar-lidar and lidar-image feature correspondences are not required, and (iii) both camera extrinsics and intrinsics, including complex distortion models, are autocalibrated. Experiments show that the calibration accuracy reaches or exceeds the accuracy of methods relying on calibration targets. Stepán Obdrzálek, Jiri Matas |
IROS | 2 |
| 2025 | Breaking the Frame: Visual Place Recognition by Overlap PredictionabstractVisual place recognition methods struggle with occlusion and partial visual overlaps. We propose a novel visual place recognition approach based on overlap prediction, called VOP, shifting from traditional reliance on global image similarities and local features to image overlap prediction. VOP proceeds co-visible image sections by obtaining patch-level embeddings using a Vision Transformer backbone and establishing patch-to-patch correspondences without requiring expensive feature detection and matching. Our approach uses a voting mechanism to assess overlap scores for potential database images. It provides a nuanced image retrieval metric in challenging scenarios. Experimental results show that VOP leads to more accurate relative pose estimation and localization results on the retrieved image pairs than state-of-the-art base-lines on a number of large-scale, real-world indoor and outdoor benchmarks. The code is available at https://github.com/weitong8591/vop.git. Tong Wei 0002, Philipp Lindenberger, Jiri Matas, Daniel Barath |
WACV | 3 |
| 2025 | MFTIQ: Multi-Flow Tracker with Independent Matching Quality Estimation
Jonás Serých, Michal Neoral, Jiri Matas |
WACV | 3 |
| 2024 | PixOOD: Pixel-Level Out-of-Distribution Detection
Tomás Vojír, Jan Sochman, Jiri Matas |
ECCV (60) | 3 |
| 2024 | LifeCLEF 2024 Teaser: Challenges on Species Distribution Prediction and Identification
Alexis Joly, Lukás Picek, Stefan Kahl, Hervé Goëau, Vincent Espitalier, Christophe Botella, Benjamin Deneu, Diego Marcos, Joaquim Estopinan, César Leblanc, Théo Larcher, Milan Sulc, Marek Hrúz, Maximilien Servajean, Jiri Matas, Hervé Glotin, Robert Planqué, Willem-Pier Vellinga, Holger Klinck, Tom Denton, Andrew Durso, Ivan Eggel, Pierre Bonnet, Henning Müller |
ECIR (6) | 15 |
| 2024 | Improving 2D Human Pose Estimation in Rare Camera Views with Synthetic DataabstractMethods and datasets for human pose estimation focus predominantly on side- and front-view scenarios. We overcome the limitation by leveraging synthetic data and introduce RePoGen (RarE POses GENerator), an SMPL-based method for generating synthetic humans with comprehensive control over pose and view. Experiments on top-view datasets and a new dataset of real images with diverse poses show that adding the RePoGen data to the COCO dataset outperforms previous approaches to top- and bottom-view pose estimation without harming performance on common views. An ablation study shows that anatomical plausibility, a property prior research focused on, is not a prerequisite for effective performance. The introduced dataset and the corresponding code are available on the project website11https://mirapurkrabek.github.io/RePoGen-paper/. Miroslav Purkrábek, Jiri Matas |
FG | 2 |
| 2024 | MFT: Long-Term Tracking of Every PixelabstractWe propose MFT – Multi-Flow dense Tracker – a novel method for dense, pixel-level, long-term tracking. The approach exploits optical flows estimated not only between consecutive frames, but also for pairs of frames at logarithmically spaced intervals. It selects the most reliable sequence of flows on the basis of estimates of its geometric accuracy and the probability of occlusion, both provided by a pre-trained CNN. We show that MFT achieves competitive performance on the TAP-Vid benchmark, outperforming baselines by a significant margin, and tracking densely orders of magnitude faster than the state-of-the-art point-tracking methods. The method is insensitive to medium-length occlusions and it is robustified by estimating flow with respect to the reference frame, which reduces drift. Michal Neoral, Jonás Serých, Jiri Matas |
WACV | 3 |
| 2024 | Single-Image Deblurring, Trajectory and Shape Recovery of Fast Moving Objects with Denoising Diffusion Probabilistic ModelsabstractBlurry appearance of fast moving objects in video frames was successfully used to reconstruct the object appearance and motion in both 2D and 3D domains. The proposed method addresses the novel, severely ill-posed, task of single-image fast moving object deblurring, shape, and trajectory recovery – previous approaches require at least three consecutive video frames. Given a single image, the method outputs the object 2D appearance and position in a series of sub-frames as if captured by a high-speed camera (i.e. temporal super-resolution). The proposed SI-DDPM-FMO method is trained end-to-end on a synthetic dataset with various moving objects, yet it generalizes well to real-world data from several publicly available datasets. SI-DDPM-FMO performs similarly to or better than recent multi-frame methods and a carefully designed baseline method. Radim Spetlík, Denys Rozumnyi, Jiri Matas |
WACV | 3 |
| 2024 | A New Dataset and a Distractor-Aware Architecture for Transparent Object TrackingabstractAbstract Performance of modern trackers degrades substantially on transparent objects compared to opaque objects. This is largely due to two distinct reasons. Transparent objects are unique in that their appearance is directly affected by the background. Furthermore, transparent object scenes often contain many visually similar objects (distractors), which often lead to tracking failure. However, development of modern tracking architectures requires large training sets, which do not exist in transparent object tracking. We present two contributions addressing the aforementioned issues. We propose the first transparent object trackingtraining datasetTrans2k that consists of over 2k sequences with 104,343 images overall, annotated by bounding boxes and segmentation masks. Standard trackers trained on this dataset consistently improve by up to 16%. Our second contribution is a new distractor-aware transparent object tracker (DiTra) that treats localization accuracy and target identification as separate tasks and implements them by a novel architecture. DiTra sets a new state-of-the-art in transparent object tracking and generalizes well to opaque objects. Alan Lukezic, Ziga Trojer, Jiri Matas, Matej Kristan |
Int. J. Comput. Vis. | 3 |
| 2024 | In Memoriam: Xiaoou Tang
Yasuyuki Matsushita, Svetlana Lazebnik, Jiri Matas |
Int. J. Comput. Vis. | 3 |
| 2024 | Cascaded and Generalizable Neural Radiance Fields for Fast View SynthesisabstractWe present CG-NeRF, a cascade and generalizable neural radiance fields method for view synthesis. Recent generalizing view synthesis methods can render high-quality novel views using a set of nearby input views. However, the rendering speed is still slow due to the nature of uniformly-point sampling of neural radiance fields. Existing scene-specific methods can train and render novel views efficiently but can not generalize to unseen data. Our approach addresses the problems of fast and generalizing view synthesis by proposing two novel modules: a coarse radiance fields predictor and a convolutional-based neural renderer. This architecture infers consistent scene geometry based on the implicit neural fields and renders new views efficiently using a single GPU. We first train CG-NeRF on multiple 3D scenes of the DTU dataset, and the network can produce high-quality and accurate novel views on unseen real and synthetic data using only photometric losses. Moreover, our method can leverage a denser set of reference images of a single scene to produce accurate novel views without relying on additional explicit representations and still maintains the high-speed rendering of the pre-trained model. Experimental results show that CG-NeRF outperforms state-of-the-art generalizable neural rendering methods on various synthetic and real datasets. Phong Nguyen 0001, Lam Huynh, Esa Rahtu, Jiri Matas, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | A Large-Scale Homography BenchmarkabstractWe present a large-scale dataset of Planes in 3D, Pi3D, of roughly 1000 planes observed in 10 000 images from the 1DSfM dataset, and HEB, a large-scale homography estimation benchmark leveraging Pi3D. The applications of the Pi3D dataset are diverse, e.g. training or evaluating monocular depth, surface normal estimation and image matching algorithms. The HEB dataset consists of 226 260 homographies and includes roughly 4M correspondences. The homographies link images that often undergo significant viewpoint and illumination changes. As applications of HEB, we perform a rigorous evaluation of a wide range of robust estimators and deep learning-based correspondence filtering methods, establishing the current state-of-the-art in robust homography estimation. We also evaluate the uncertainty of the SIFT orientations and scales w.r.t. the ground truth coming from the underlying homographies and provide codes for comparing uncertainty of custom detectors. The dataset is available at https://github.com/danini/homography-benchmark. Daniel Barath, Dmytro Mishkin, Michal Polic, Wolfgang Förstner, Jiri Matas |
CVPR | 5 |
| 2023 | Finding Geometric Models by Clustering in the Consensus SpaceabstractWe propose a new algorithm for finding an unknown number of geometric models, e.g., homographies. The problem is formalized as finding dominant model instances progressively without forming crisp point-to-model assignments. Dominant instances are found via a RANSAC-like sampling and a consolidation process driven by a model quality function considering previously proposed instances. New ones are found by clustering in the consensus space. This new formulation leads to a simple iterative algorithm with state-of-the-art accuracy while running in real-time on a number of vision problems - at least two orders of magnitude faster than the competitors on two-view motion estimation. Also, we propose a deterministic sampler reflecting the fact that real-world data tend to form spatially coherent structures. The sampler returns connected components in a progressively densified neighborhood-graph. We present a number of applications where the use of multiple geometric models improves accuracy. These include pose estimation from multiple generalized homographies; trajectory estimation of fast-moving objects; and we also propose a way of using multiple homographies in global SfM algorithms. Source code: https://github.com/danini/clustering-in-consensus-space. Daniel Barath, Denys Rozumnyi, Ivan Eichhardt, Levente Hajder, Jiri Matas |
CVPR | 5 |
| 2023 | Tracking by 3D Model Estimation of Unknown Objects in VideosabstractMost model-free visual object tracking methods formulate the tracking task as object location estimation given by a 2D segmentation or a bounding box in each video frame. We argue that this representation is limited and instead propose to guide and improve 2D tracking with an explicit object representation, namely the textured 3D shape and 6DoF pose in each video frame. Our representation tackles a complex long-term dense correspondence problem between all 3D points on the object for all video frames, including frames where some points are invisible. To achieve that, the estimation is driven by re-rendering the input video frames as well as possible through differentiable rendering, which has not been used for tracking before. The proposed optimization minimizes a novel loss function to estimate the best 3D shape, texture, and 6DoF pose. We improve the state-of-the-art in 2D segmentation tracking on three different datasets with mostly rigid objects. Denys Rozumnyi, Jiri Matas, Marc Pollefeys, Vittorio Ferrari, Martin R. Oswald |
ICCV | 2 |
| 2023 | Adaptive Reordering Sampler with Neurally Guided MAGSACabstractWe propose a new sampler for robust estimators that always selects the sample with the highest probability of consisting only of inliers. After every unsuccessful iteration, the inlier probabilities are updated in a principled way via a Bayesian approach. The probabilities obtained by the deep network are used as prior (so-called neural guidance) inside the sampler. Moreover, we introduce a new loss that exploits, in a geometrically justifiable manner, the orientation and scale that can be estimated for any type of feature, e.g., SIFT or SuperPoint, to estimate two-view geometry. The new loss helps to learn higher-order information about the underlying scene geometry. Benefiting from the new sampler and the proposed loss, we combine the neural guidance with the state-of-the-art MAGSAC++. Adaptive Reordering Sampler with Neurally Guided MAGSAC (ARS-MAGSAC) is superior to the state-of-the-art in terms of accuracy and run-time on the PhotoTourism and KITTI datasets for essential and fundamental matrix estimation. The code and trained models are available at https://github.com/weitong8591/ars_magsac. Tong Wei 0002, Jiri Matas, Daniel Barath |
ICCV | 2 |
| 2023 | Generalized Differentiable RANSACabstractWe propose ▽-RANSAC, a generalized differentiable RANSAC that allows learning the entire randomized robust estimation pipeline. The proposed approach enables the use of relaxation techniques for estimating the gradients in the sampling distribution, which are then propagated through a differentiable solver. The trainable quality function marginalizes over the scores from all the models estimated within ▽-RANSAC to guide the network learning accurate and useful inlier probabilities or to train feature detection and matching networks. Our method directly maximizes the probability of drawing a good hypothesis, allowing us to learn better sampling distributions. We test ▽-RANSAC on various real-world scenarios on fundamental and essential matrix estimation, and 3D point cloud registration, outdoors and indoors, with handcrafted and learning-based features. It is superior to the state-of-the-art in terms of accuracy while running at a similar speed to its less accurate alternatives. The code and trained models are available at https://github.com/weitong8591/differentiable_ransac. Tong Wei 0002, Alexander Shekhovtsov 0001, Jiri Matas, Daniel Barath |
ICCV | 4 |
| 2023 | DocILE Benchmark for Document Information Localization and Extraction
Stepán Simsa, Milan Sulc, Michal Uricár, Ahmed Hamdi, Matej Kocián, Matyás Skalický, Jiri Matas, Antoine Doucet, Mickaël Coustaty, Dimosthenis Karatzas |
ICDAR (2) | 8 |
| 2023 | DoG Accuracy Via Equivariance: Get The Interpolation RightabstractWe study the influence of image interpolation algorithms on local feature detectors operating on a scale pyramid, focusing on the Difference-of-Gaussian, as used in SIFT. We show that commonly used implementations, such as in OpenCV and Kornia, are neither rotational nor scale equivariant. We present a simple solution and demonstrate its positive influence on the downstream image matching tasks. The implementation of the method has been accepted in standard libraries OpenCV [1] and Kornia [2]. Václav Vávra, Dmytro Mishkin, Jiri Matas |
ICIP | 3 |
| 2023 | Efficient Visuo-Haptic Object Shape Completion for Robot ManipulationabstractFor robot manipulation, a complete and accurate object shape is desirable. Here, we present a method that combines visual and haptic reconstruction in a closed-loop pipeline. From an initial viewpoint, the object shape is reconstructed using an implicit surface deep neural network. The location with highest uncertainty is selected for haptic exploration, the object is touched, the new information from touch and a new point cloud from the camera are added, object position is re-estimated and the cycle is repeated. We extend Rustler et al. (2022) by using a new theoretically grounded method to determine the points with highest uncertainty, and we increase the yield of every haptic exploration by adding not only the contact points to the point cloud but also incorporating the empty space established through the robot movement to the object. Additionally, the solution is compact in that the jaws of a closed two-finger gripper are directly used for exploration. The object position is re-estimated after every robot action and multiple objects can be present simultaneously on the table. We achieve a steady improvement with every touch using three different metrics and demonstrate the utility of the better shape reconstruction in grasping experiments on the real robot. On average, grasp success rate increases from 63.3 % to 70.4 % after a single exploratory touch and to 82.7% after five touches. The collected data and code are publicly available (https://osf.io/j6rkd/, https://github.com/ctu-vras/vishac). Lukas Rustler, Jiri Matas, Matej Hoffmann |
IROS | 2 |
| 2023 | Planar Object Tracking via Weighted Optical FlowabstractWe propose WOFT – a novel method for planar object tracking that estimates a full 8 degrees-of-freedom pose, i.e. the homography w.r.t. a reference view. The method uses a novel module that leverages dense optical flow and assigns a weight to each optical flow correspondence, estimating a homography by weighted least squares in a fully differentiable manner. The trained module assigns zero weights to incorrect correspondences (outliers) in most cases, making the method robust and eliminating the need of the typically used non-differentiable robust estimators like RANSAC. The proposed weighted optical flow tracker (WOFT) achieves state-of-the-art performance on two benchmarks, POT-210 [23] and POIC [7], tracking consistently well across a wide range of scenarios. Jonás Serých, Jiri Matas |
WACV | 2 |
| 2023 | Image-Consistent Detection of Road Anomalies as Unpredictable PatchesabstractWe propose a novel method for anomaly detection primarily aiming at autonomous driving. The design of the method, called DaCUP (Detection of anomalies as Consistent Unpredictable Patches), is based on two general properties of anomalous objects: an anomaly is (i) not from a class that could be modelled and (ii) it is not similar (in appearance) to non-anomalous objects in the image. To this end, we propose a novel embedding bottleneck in an auto-encoder like architecture that enables modelling of a diverse, multi-modal known class appearance (e.g. road). Secondly, we introduce novel image-conditioned distance features that allow known class identification in a nearest-neighbour manner on-the-fly, greatly increasing its ability to distinguish true and false positives. Lastly, an inpainting module is utilized to model the uniqueness of detected anomalies and significantly reduce false positives by filtering regions that are similar, thus reconstructable from their neighbourhood. We demonstrate that filtering of regions based on their similarity to neighbour regions, using e.g. an inpainting module, is general and can be used with other methods for reduction of false positives. The proposed method is evaluated on several publicly available datasets for road anomaly detection and on a maritime benchmark for obstacle avoidance. The method achieves state-of-the-art performance in both tasks with the same hyper-parameters with no domain specific design. Tomás Vojír, Jiri Matas |
WACV | 2 |
| 2023 | Binaural SoundNet: Predicting Semantics, Depth and Motion With Binaural SoundsabstractHumans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene understanding purely based on binaural sounds. The considered tasks include predicting the semantic masks of sound-making objects, the motion of sound-making objects, and the depth map of the scene. To this aim, we propose a novel sensor setup and record a new audio-visual dataset of street scenes with eight professional binaural microphones and a 360$\mathrm{^{\circ }}$camera. The co-existence of visual and audio cues is leveraged for supervision transfer. In particular, we employ a cross-modal distillation framework that consists of multiple vision ‘teacher’ methods and a sound ‘student’ method – the student method is trained to generate the same results as the teacher methods do. This way, the auditory system can be trained without using human annotations. To further boost the performance, we propose another novel auxiliary task, coined Spatial Sound Super-Resolution, to increase the directional resolution of sounds. We then formulate the four tasks into one end-to-end trainable multi-tasking network aiming to boost the overall performance. Experimental results show that 1) our method achieves good results for all four tasks, 2) the four tasks are mutually beneficial – training them together achieves the best performance, 3) the number and orientation of microphones are both important, and 4) features learned from the standard spectrogram and features obtained by the classic signal processing pipeline are complementary for auditory perception tasks. The data and code are released on the project page:https://www.trace.ethz.ch/publications/2020/sound_perception/index.html. Dengxin Dai, Arun Balajee Vasudevan, Jiri Matas, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Visual Object Tracking With Discriminative Filters and Siamese Networks: A Survey and OutlookabstractAccurate and robust visual object tracking is one of the most challenging and fundamental computer vision problems. It entails estimating the trajectory of the target in an image sequence, given only its initial location, and segmentation, or its rough approximation in the form of a bounding box. Discriminative Correlation Filters (DCFs) and deep Siamese Networks (SNs) have emerged as dominating tracking paradigms, which have led to significant progress. Following the rapid evolution of visual object tracking in the last decade, this survey presents a systematic and thorough review of more than 90 DCFs and Siamese trackers, based on results in nine tracking benchmarks. First, we present the background theory of both the DCF and Siamese tracking core formulations. Then, we distinguish and comprehensively review the shared as well as specific open research challenges in both these tracking paradigms. Furthermore, we thoroughly analyze the performance of DCF and Siamese trackers on nine benchmarks, covering different experimental aspects of visual tracking: datasets, evaluation metrics, performance, and speed comparisons. We finish the survey by presenting recommendations and suggestions for distinguished open challenges based on our analysis. Sajid Javed, Martin Danelljan, Fahad Shahbaz Khan, Muhammad Haris Khan, Michael Felsberg, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Pose-graph via Adaptive Image Re-ordering
Daniel Barath, Jana Noskova, Ivan Eichhardt, Jiri Matas |
BMVC | 4 |
| 2022 | Trans2k: Unlocking the Power of Deep Models for Transparent Object Tracking
Alan Lukezic, Ziga Trojer, Jiri Matas, Matej Kristan |
BMVC | 3 |
| 2022 | DAD-3DHeads: A Large-scale Dense, Accurate and Diverse Dataset for 3D Head Alignment from a Single ImageabstractWe present DAD-3DHeads, a dense and diverse large-scale dataset, and a robust model for 3D Dense Head Alignment in-the-wild. It contains annotations of over 3.5K land-marks that accurately represent 3D head shape compared to the ground-truth scans. The data-driven model, DAD-3DNet, trained on our dataset, learns shape, expression, and pose parameters, and performs 3D reconstruction of a FLAME mesh. The model also incorporates a landmark prediction branch to take advantage of rich supervision and co-training of multiple related tasks. Experimentally, DAD-3DNet outperforms or is comparable to the state-of-the-art models in (i) 3D Head Pose Estimation on AFLW2000-3D and BIWI, (ii) 3D Face Shape Reconstruction on NoW and Feng, and (iii) 3D Dense Head Alignment and 3D Land-marks Estimation on DAD-3DHeads dataset. Finally, diver-sity of DAD-3DHeads in camera angles, facial expressions, and occlusions enables a benchmark to study in-the-wild generalization and robustness to distribution shifts. The dataset webpage is https://p.farm/research/dad-3dheads. Tetiana Martyniuk, Orest Kupyn, Yana Kurlyak, Igor Krashenyi, Jiri Matas, Viktoriia Sharmanska |
CVPR | 5 |
| 2022 | Recall@k Surrogate Loss with Large Batches and Similarity MixupabstractThis work focuses on learning deep visual representation models for retrieval by exploring the interplay between a new loss function, the batch size, and a new regularization approach. Direct optimization, by gradient descent, of an evaluation metric, is not possible when it is nondifferentiable, which is the case for recall in retrieval. A differentiable surrogate loss for the recall is proposed in this work. Using an implementation that sidesteps the hardware constraints of the GPU memory, the method trains with a very large batch size, which is essential for metrics computed on the entire retrieval database. It is assisted by an efficient mixup regularization approach that operates on pairwise scalar similarities and virtually increases the batch size further. The suggested method achieves state-of-the-art performance in several image retrieval benchmarks when used for deep metric learning. For instance-level recognition, the method outperforms similar approaches that train using an approximation of average precision. Giorgos Tolias, Jiri Matas |
CVPR | 3 |
| 2022 | Point Cloud Color ConstancyabstractIn this paper, we present Point Cloud Color Constancy, in short PCCC, an illumination chromaticity estimation algorithm exploiting a point cloud. We leverage the depth information captured by the time-of-flight (ToF) sensor mounted rigidly with the RGB sensor, and form a 6D cloud where each point contains the coordinates and RGB intensities, noted as (x,y,z, r,g, b). PCCC applies the PointNet architecture to the color constancy problem, deriving the illumination vector point-wise and then making a global decision about the global illumination chromaticity. On two popular RGB-D datasets, which we extend with illumination information, as well as on a novel benchmark, PCCC obtains lower error than the state-of-the-art algorithms. Our method is simple andfast, requiring merely 16 x 16-size input and reaching speed over 140 fps (CPU time), including the cost of building the point cloud and net inference. Xiaoyan Xing, Yanlin Qian, Sibo Feng, Yuhan Dong, Jiri Matas |
CVPR | 5 |
| 2022 | FEAR: Fast, Efficient, Accurate and Robust Visual Tracker
Vasyl Borsuk, Roman Vei, Orest Kupyn, Tetiana Martyniuk, Igor Krashenyi, Jiri Matas |
ECCV (22) | 6 |
| 2022 | Lightweight Monocular Depth with a Novel Neural Architecture Search MethodabstractThis paper presents a novel neural architecture search method, called LiDNAS, for generating lightweight monocular depth estimation models. Unlike previous neural architecture search (NAS) approaches, where finding optimized networks is computationally demanding, the introduced novel Assisted Tabu Search leads to efficient architecture exploration. Moreover, we construct the search space on a pre-defined backbone network to balance layer diversity and search space size. The LiDNAS method outperforms the state-of-the-art NAS approach, proposed for disparity and depth estimation, in terms of search efficiency and output model performance. The LiDNAS optimized models achieve result superior to compact depth estimation state-of-the-art on NYU-Depth-v2, KITTI, and ScanNet, while being 7%-500% more compact in size, i.e the number of model parameters. Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
WACV | 3 |
| 2022 | Danish Fungi 2020 - Not Just Another Image Recognition DatasetabstractWe introduce a novel fine-grained dataset and bench-mark, the Danish Fungi 2020 (DF20). The dataset, constructed from observations submitted to the Atlas of Danish Fungi, is unique in its taxonomy-accurate class labels, small number of errors, highly unbalanced long-tailed class distribution, rich observation metadata, and well-defined class hierarchy. DF20 has zero overlap with ImageNet, al-lowing unbiased comparison of models fine-tuned from publicly available ImageNet checkpoints. The proposed evaluation protocol enables testing the ability to improve classification using metadata – e.g. precise geographic location, habitat, and substrate, facilitates classifier calibration testing, and finally allows to study the impact of the device settings on the classification performance. Experiments using Convolutional Neural Networks (CNN) and the recent Vision Transformers (ViT) show that DF20 presents a challenging task. Interestingly, ViT achieves results su-perior to CNN baselines with 80.45% accuracy and 0.743 macro F1 score, reducing the CNN error by 9% and 12% respectively. A simple procedure for including metadata into the decision process improves the classification accuracy by more than 2.95 percentage points, reducing the error rate by 15%. The source code for all methods and experiments is available at https://sites.google.com/view/danish-fungi-dataset. Lukás Picek, Milan Sulc, Jiri Matas, Thomas S. Jeppesen, Jacob Heilmann-Clausen, Thomas Læssøe, Tobias Frøslev |
WACV | 3 |
| 2022 | The Hitchhiker's Guide to Prior-Shift AdaptationabstractIn many computer vision classification tasks, class priors at test time often differ from priors on the training set. In the case of such prior shift, classifiers must be adapted correspondingly to maintain close to optimal performance. This paper analyzes methods for adaptation of probabilistic classifiers to new priors and for estimating new priors on an unlabeled test set. We propose a novel method to address a known issue of prior estimation methods based on confusion matrices, where inconsistent estimates of decision probabilities and confusion matrices lead to negative values in the estimated priors. Experiments on fine-grained image classification datasets provide insight into the best practice of prior shift estimation and classifier adaptation, and show that the proposed method achieves state-of-the-art results in prior adaptation. Applying the best practice to two tasks with naturally imbalanced priors, learning from web-crawled images and plant species classification, increased the recognition accuracy by 1.1% and 3.4% respectively. Tomás Sipka, Milan Sulc, Jiri Matas |
WACV | 3 |
| 2022 | Graph-Cut RANSAC: Local Optimization on Spatially Coherent StructuresabstractWe propose Graph-Cut RANSAC, GC-RANSAC in short, a new robust geometric model estimation method where the local optimization step is formulated as energy minimization with binary labeling, applying the graph-cut algorithm to select inliers. The minimized energy reflects the assumption that geometric data often form spatially coherent structures - it includes both a unary component representing point-to-model residuals and a binary term promoting spatially coherent inlier-outlier labelling of neighboring points. The proposed local optimization step is conceptually simple, easy to implement, efficient with a globally optimal inlier selection given the model parameters. Graph-Cut RANSAC, equipped with "the bells and whistles" of USAC and MAGSAC++, was tested on a range of problems using a number of publicly available datasets for homography, 6D object pose, fundamental and essential matrix estimation. It is more geometrically accurate than state-of-the-art robust estimators, fails less often and runs faster or with speed similar to less accurate alternatives. The source code is available at https://github.com/danini/graph-cut-ransac. Daniel Barath, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Marginalizing Sample ConsensusabstractA new method for robust estimation, MAGSAC++, is proposed. It introduces a new model quality (scoring) function that does not make inlier-outlier decisions, and a novel marginalization procedure formulated as an M-estimation with a novel class of M-estimators (a robust kernel) solved by an iteratively re-weighted least squares procedure. Instead of the inlier-outlier threshold, it requires only its loose upper bound which can be chosen from a significantly wider range. Also, we propose a new termination criterion and a technique for selecting a set of inliers in a data-driven manner as a post-processing step after the robust estimation finishes. On a number of publicly available real-world datasets for homography, fundamental matrix fitting and relative pose, MAGSAC++ produces results superior to the state-of-the-art robust methods. It is more geometrically accurate, fails fewer times, and it is often faster. It is shown that MAGSAC++ is significantly less sensitive to the setting of the threshold upper bound than the other state-of-the-art algorithms to the inlier-outlier threshold. Therefore, it is easier to be applied to unseen problems and scenes without acquiring information by hand about the setting of the inlier-outlier threshold. The source code and examples both in C++ and Python are available at https://github.com/danini/magsac. Daniel Barath, Jana Noskova, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | A Discriminative Single-Shot Segmentation Network for Visual Object TrackingabstractTemplate-based discriminative trackers are currently the dominant tracking paradigm due to their robustness, but are restricted to bounding box tracking and a limited range of transformation models, which reduces their localization accuracy. We propose a discriminative single-shot segmentation tracker – D3S$_2$, which narrows the gap between visual object tracking and video object segmentation. A single-shot network applies two target models with complementary geometric properties, one invariant to a broad range of transformations, including non-rigid deformations, the other assuming a rigid object to simultaneously achieve robust online target segmentation. The overall tracking reliability is further increased by decoupling the object and feature scale estimation. Without per-dataset finetuning, and trained only for segmentation as the primary output, D3S$_2$outperforms all published trackers on the recent short-term tracking benchmark VOT2020 and performs very close to the state-of-the-art trackers on the GOT-10k, TrackingNet, OTB100 and LaSoT. D3S$_2$outperforms the leading segmentation tracker SiamMask on video object segmentation benchmarks and performs on par with top video object segmentation algorithms. Alan Lukezic, Jiri Matas, Matej Kristan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Monocular Depth Estimation Primed by Salient Point Detection and Normalized Hessian LossabstractDeep neural networks have recently thrived on single image depth estimation. That being said, current developments on this topic highlight an apparent compromise between accuracy and network size. This work proposes an accurate and lightweight framework for monocular depth estimation based on a self-attention mechanism stemming from salient point detection. Specifically, we utilize a sparse set of keypoints to train a FuSaNet model that consists of two major components: Fusion-Net and Saliency-Net. In addition, we introduce a normalized Hessian loss term invariant to scaling and shear along the depth direction, which is shown to substantially improve the accuracy. The proposed method achieves state-of-the-art results on NYU-Depth-v2 and KITTI while using 3.1-38.4 times smaller model in terms of the number of parameters than baseline approaches. Experiments on the SUN-RGBD further demonstrate the generalizability of the proposed method. Lam Huynh, Matteo Pedone, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
3DV | 4 |
| 2021 | RGBD-Net: Predicting Color and Depth Images for Novel Views SynthesisabstractWe propose a new cascaded architecture for novel view synthesis, called RGBD-Net, which consists of two core components: a hierarchical depth regression network and a depth-aware generator network. The former one predicts depth maps of the target views by using adaptive depth scaling, while the latter one leverages the predicted depths and renders spatially and temporally consistent target images. In the experimental evaluation on standard datasets, RGBD-Net not only outperforms the state-of-the-art by a clear margin, but it also generalizes well to new scenes without per-scene optimization. Moreover, we show that RGBD-Net can be optionally trained without depth supervision while still retaining high-quality rendering. Thanks to the depth regression network, RGBD-Net can be also used for creating dense 3D point clouds that are more accurate than those produced by some state-of-the-art multi-view stereo methods. Phong Nguyen 0001, Animesh Karnewar, Lam Huynh, Esa Rahtu, Jiri Matas, Janne Heikkilä |
3DV | 5 |
| 2021 | Monocular Arbitrary Moving Object Discovery and Segmentation
Michal Neoral, Jan Sochman, Jiri Matas |
BMVC | 3 |
| 2021 | Efficient Initial Pose-Graph Generation for Global SfMabstractWe propose ways to speed up the initial pose-graph generation for global Structure-from-Motion algorithms. To avoid forming tentative point correspondences by FLANN and geometric verification by RANSAC, which are the most time-consuming steps of the pose-graph creation, we propose two new methods – built on the fact that image pairs usually are matched consecutively. Thus, candidate relative poses can be recovered from paths in the partly-built pose-graph. We propose a heuristic for the A*traversal, considering global similarity of images and the quality of the pose-graph edges. Given a relative pose from a path, descriptor-based feature matching is made "light-weight" by exploiting the known epipolar geometry. To speed up PROSAC-based sampling when RANSAC is applied, we propose a third method to order the correspondences by their inlier probabilities from previous estimations. The algorithms are tested on 402130 image pairs from the 1DSfM dataset and they speed up the feature matching 17 times and pose estimation 5 times. Source code: https://github.com/danini/pose-graph-initialization Daniel Barath, Dmytro Mishkin, Ivan Eichhardt, Ilia Shipachev, Jiri Matas |
CVPR | 5 |
| 2021 | DeFMO: Deblurring and Shape Recovery of Fast Moving ObjectsabstractObjects moving at high speed appear significantly blurred when captured with cameras. The blurry appearance is especially ambiguous when the object has complex shape or texture. In such cases, classical methods, or even humans, are unable to recover the object’s appearance and motion. We propose a method that, given a single image with its estimated background, outputs the object’s appearance and position in a series of sub-frames as if captured by a high-speed camera (i.e. temporal super-resolution). The proposed generative model embeds an image of the blurred object into a latent space representation, disentangles the background, and renders the sharp appearance. Inspired by the image formation model, we design novel self-supervised loss function terms that boost performance and show good generalization capabilities. The proposed DeFMO method is trained on a complex synthetic dataset, yet it performs well on real-world data from several datasets. DeFMO outperforms the state of the art and generates high-quality temporal super-resolution frames. Denys Rozumnyi, Martin R. Oswald, Vittorio Ferrari, Jiri Matas, Marc Pollefeys |
CVPR | 4 |
| 2021 | Boosting Monocular Depth Estimation with Lightweight 3D Point FusionabstractIn this paper, we propose enhancing monocular depth estimation by adding 3D points as depth guidance. Unlike existing depth completion methods, our approach performs well on extremely sparse and unevenly distributed point clouds, which makes it agnostic to the source of the 3D points. We achieve this by introducing a novel multi-scale 3D point fusion network that is both lightweight and efficient. We demonstrate its versatility on two different depth estimation problems where the 3D points have been acquired with conventional structure-from-motion and Li-DAR. In both cases, our network performs on par with state-of-the-art depth completion methods and achieves significantly higher accuracy when only a small number of points is used while being more compact in terms of the number of parameters. We show that our method outperforms some contemporary deep learning based multi-view stereo and structure-from-motion methods both in accuracy and in compactness. Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
ICCV | 3 |
| 2021 | VSAC: Efficient and Accurate Estimator for H and FabstractWe present VSAC, a RANSAC-type robust estimator with a number of novelties. It benefits from the introduction of the concept of independent inliers that improves significantly the efficacy of the dominant plane handling and, also, allows near error-free rejection of incorrect models, without false positives. The local optimization process and its application is improved so that it is run on average only once. Further technical improvements include adaptive sequential hypothesis verification and efficient model estimation via Gaussian elimination. Experiments on four standard datasets show that VSAC is significantly faster than all its predecessors and runs on average in 1-2 ms, on a CPU. It is two orders of magnitude faster and yet as precise as MAGSAC++, the currently most accurate estimator of two-view geometry. In the repeated runs on EVD, HPatches, PhotoTourism, and Kusvod2 datasets, it never failed. Maksym Ivashechkin, Daniel Barath, Jiri Matas |
ICCV | 3 |
| 2021 | FMODetect: Robust Detection of Fast Moving ObjectsabstractWe propose the first learning-based approach for fast moving objects detection. Such objects are highly blurred and move over large distances within one video frame. Fast moving objects are associated with a deblurring and matting problem, also called deblatting. We show that the separation of deblatting into consecutive matting and deblurring allows achieving real-time performance, i.e. an order of magnitude speed-up, and thus enabling new classes of application. The proposed method detects fast moving objects as a truncated distance function to the trajectory by learning from synthetic data. For the sharp appearance estimation and accurate trajectory estimation, we propose a matting and fitting network that estimates the blurred appearance without background, followed by an energy minimization based deblurring. The state-of-the-art methods are outperformed in terms of recall, precision, trajectory estimation, and sharp appearance reconstruction. Compared to other methods, such as deblatting, the inference is of several orders of magnitude faster and allows applications such as real-time fast moving object detection and retrieval in large video collections. Denys Rozumnyi, Jiri Matas, Filip Sroubek, Marc Pollefeys, Martin R. Oswald |
ICCV | 2 |
| 2021 | Road Anomaly Detection by Partial Image Reconstruction with Segmentation CouplingabstractWe present a novel approach to the detection of unknown objects in the context of autonomous driving. The problem is formulated as anomaly detection, since we assume that the unknown stuff or object appearance cannot be learned. To that end, we propose a reconstruction module that can be used with many existing semantic segmentation networks, and that is trained to recognize and reconstruct road (drivable) surface from a small bottleneck. We postulate that poor reconstruction of the road surface is due to areas that are outside of the training distribution, which is a strong indicator of an anomaly. The road structural similarity error is coupled with the semantic segmentation to incorporate information from known classes and produce final per-pixel anomaly scores. The proposed JSR-Net was evaluated on four datasets, Lost-and-found, Road Anomaly, Road Obstacles, and FishyScapes, achieving state-of-art performance on all, reducing the false positives significantly, while typically having the highest average precision for wide range of operation points. Tomás Vojír, Tomás Sipka, Rahaf Aljundi, Nikolay Chumerin, Daniel Olmeda Reino, Jiri Matas |
ICCV | 6 |
| 2021 | Fast Text vs. Non-text Classification of Images
Jiri Kralicek, Jiri Matas |
ICDAR (4) | 2 |
| 2021 | FEDS - Filtered Edit Distance Surrogate
Jiri Matas |
ICDAR (4) | 2 |
| 2021 | Fast Fourier Intrinsic NetworkabstractWe address the problem of decomposing an image into albedo and shading. We propose the Fast Fourier Intrinsic Network, FFI-Net in short, that operates in the spectral domain, splitting the input into several spectral bands. Weights in FFI-Net are optimized in the spectral domain, allowing faster convergence to a lower error. FFI-Net is lightweight and does not need auxiliary networks for training. The network is trained end-to-end with a novel spectral loss which measures the global distance between the network prediction and corresponding ground truth. FFI-Net achieves state-of-the-art performance on MPI-Sintel, MIT Intrinsic, and IIW datasets. Yanlin Qian, Miaojing Shi, Joni-Kristian Kämäräinen, Jiri Matas |
WACV | 4 |
| 2021 | Guest Editorial: Special Issue on "Computer Vision for All Seasons: Adverse Weather and Lighting Conditions"
Dengxin Dai, Robby T. Tan, Vishal M. Patel, Jiri Matas, Bernt Schiele, Luc Van Gool |
Int. J. Comput. Vis. | 4 |
| 2021 | Image Matching Across Wide Baselines: From Paper to Practice
Yuhe Jin, Dmytro Mishkin, Anastasiia Mishchuk, Jiri Matas, Pascal Fua, Kwang Moo Yi, Eduard Trulls |
Int. J. Comput. Vis. | 4 |
| 2021 | Tracking by DeblattingabstractAbstract Objects moving at high speed along complex trajectories often appear in videos, especially videos of sports. Such objects travel a considerable distance during exposure time of a single frame, and therefore, their position in the frame is not well defined. They appear as semi-transparent streaks due to the motion blur and cannot be reliably tracked by general trackers. We propose a novel approach called Tracking by Deblatting based on the observation that motion blur is directly related to the intra-frame trajectory of an object. Blur is estimated by solving two intertwined inverse problems, blind deblurring and image matting, which we call deblatting. By postprocessing, non-causal Tracking by Deblatting estimates continuous, complete, and accurate object trajectories for the whole sequence. Tracked objects are precisely localized with higher temporal resolution than by conventional trackers. Energy minimization by dynamic programming is used to detect abrupt changes of motion, called bounces. High-order polynomials are then fitted to smooth trajectory segments between bounces. The output is a continuous trajectory function that assigns location for every real-valued time stamp from zero to the number of frames. The proposed algorithm was evaluated on a newly created dataset of videos from a high-speed camera using a novel Trajectory-IoU metric that generalizes the traditional Intersection over Union and measures the accuracy of the intra-frame trajectory. The proposed method outperforms the baselines both in recall and trajectory accuracy. Additionally, we show that from the trajectory function precise physical calculations are possible, such as radius, gravity, and sub-frame object velocity. Velocity estimation is compared to the high-speed camera measurements and radars. Results show high performance of the proposed method in terms of Trajectory-IoU, recall, and velocity estimation. Denys Rozumnyi, Jan Kotera, Filip Sroubek, Jiri Matas |
Int. J. Comput. Vis. | 4 |
| 2021 | Performance Evaluation Methodology for Long-Term Single-Object TrackingabstractA long-term visual object tracking performance evaluation methodology and a benchmark are proposed. Performance measures are designed by following a long-term tracking definition to maximize the analysis probing strength. The new measures outperform existing ones in interpretation potential and in better distinguishing between different tracking behaviors. We show that these measures generalize the short-term performance measures, thus linking the two tracking problems. Furthermore, the new measures are highly robust to temporal annotation sparsity and allow annotation of sequences hundreds of times longer than in the current datasets without increasing manual annotation labor. A new challenging dataset of carefully selected sequences with many target disappearances is proposed. A new tracking taxonomy is proposed to position trackers on the short-term/long-term spectrum. The benchmark contains an extensive evaluation of the largest number of long-term trackers and comparison to state-of-the-art short-term trackers. We analyze the influence of tracking architecture implementations to long-term performance and explore various redetection strategies as well as the influence of visual model update strategies to long-term tracking drift. The methodology is integrated in the VOT toolkit to automate experimental analysis and benchmarking and to facilitate the future development of long-term trackers. Alan Lukezic, Luka Cehovin, Tomás Vojír, Jiri Matas, Matej Kristan |
IEEE Trans. Cybern. | 4 |
| 2020 | LSD_2 - Joint Denoising and Deblurring of Short and Long Exposure Images with CNNs
Janne Mustaniemi, Juho Kannala, Jiri Matas, Simo Särkkä, Janne Heikkilä |
BMVC | 3 |
| 2020 | MAGSAC++, a Fast, Reliable and Accurate Robust EstimatorabstractA new method for robust estimation, MAGSAC++1, is proposed. It introduces a new model quality (scoring) function that does not require the inlier-outlier decision, and a novel marginalization procedure formulated as an M-estimation with a novel class of M-estimators (a robust kernel) solved by an iteratively re-weighted least squares procedure. We also propose a new sampler, Progressive NAPSAC, for RANSAC-like robust estimators. Exploiting the fact that nearby points often originate from the same model in real-world data, it finds local structures earlier than global samplers. The progressive transition from local to global sampling does not suffer from the weaknesses of purely localized samplers. On six publicly available realworld datasets for homography and fundamental matrix fitting, MAGSAC++ produces results superior to the state-of-the-art robust methods. It is faster, more geometrically accurate and fails less often. Daniel Barath, Jana Noskova, Maksym Ivashechkin, Jiri Matas |
CVPR | 4 |
| 2020 | EPOS: Estimating 6D Pose of Objects With SymmetriesabstractWe present a new method for estimating the 6D pose of rigid objects with available 3D models from a single RGB input image. The method is applicable to a broad range of objects, including challenging ones with global or partial symmetries. An object is represented by compact surface fragments which allow handling symmetries in a systematic manner. Correspondences between densely sampled pixels and the fragments are predicted using an encoder-decoder network. At each pixel, the network predicts: (i) the probability of each object's presence, (ii) the probability of the fragments given the object's presence, and (iii) the precise 3D location on each fragment. A data-dependent number of corresponding 3D locations is selected per pixel, and poses of possibly multiple object instances are estimated using a robust and efficient variant of the PnP-RANSAC algorithm. In the BOP Challenge 2019, the method outperforms all RGB and most RGB-D and D methods on the T-LESS and LM-O datasets. On the YCB-V dataset, it is superior to all competitors, with a large margin over the second-best RGB method. Source code is at: cmp.felk.cvut.cz/epos. Tomas Hodan, Daniel Barath, Jiri Matas |
CVPR | 3 |
| 2020 | D3S - A Discriminative Single Shot Segmentation TrackerabstractTemplate-based discriminative trackers are currently the dominant tracking paradigm due to their robustness, but are restricted to bounding box tracking and a limited range of transformation models, which reduces their localization accuracy. We propose a discriminative single-shot segmentation tracker - D3S, which narrows the gap between visual object tracking and video object segmentation. A single-shot network applies two target models with complementary geometric properties, one invariant to a broad range of transformations, including non-rigid deformations, the other assuming a rigid object to simultaneously achieve high robustness and online target segmentation. Without per-dataset finetuning and trained only for segmentation as the primary output, D3S outperforms all trackers on VOT2016, VOT2018 and GOT-10k benchmarks and performs close to the state-of-the-art trackers on the TrackingNet. D3S outperforms the leading segmentation tracker SiamMask on video segmentation benchmark and performs on par with top video object segmentation algorithms, while running an order of magnitude faster, close to real-time. Alan Lukezic, Jiri Matas, Matej Kristan |
CVPR | 2 |
| 2020 | Sub-Frame Appearance and 6D Pose Estimation of Fast Moving ObjectsabstractWe propose a novel method that tracks fast moving objects, mainly non-uniform spherical, in full 6 degrees of freedom, estimating simultaneously their 3D motion trajectory, 3D pose and object appearance changes with a time step that is a fraction of the video frame exposure time. The sub-frame object localization and appearance estimation allows realistic temporal super-resolution and precise shape estimation. The method, called TbD-3D (Tracking by Deblatting in 3D) relies on a novel reconstruction algorithm which solves a piece-wise deblurring and matting problem. The 3D rotation is estimated by minimizing the reprojection error. As a second contribution, we present a new challenging dataset with fast moving objects that change their appearance and distance to the camera. High-speed camera recordings with zero lag between frame exposures were used to generate videos with different frame rates annotated with ground-truth trajectory and pose. Denys Rozumnyi, Jan Kotera, Filip Sroubek, Jiri Matas |
CVPR | 4 |
| 2020 | Guiding Monocular Depth Estimation Using Depth-Attention Volume
Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
ECCV (26) | 3 |
| 2020 | Learning Surrogates via Deep Embedding
Tomas Hodan, Jiri Matas |
ECCV (30) | 3 |
| 2020 | Text Recognition - Real World Data and Where to Find ThemabstractWe present a method for exploiting weakly annotated images to improve text extraction pipelines. The approach uses an arbitrary end-to-end text recognition system to obtain text region proposals and their, possibly erroneous, transcriptions. The method includes matching of imprecise transcriptions to weak annotations and an edit distance guided neighbourhood search. It produces nearly error-free, localised instances of scene text, which we treat as “pseudo ground truth” (PGT). The method is applied to two weakly-annotated datasets. Training with the extracted PGT consistently improves the accuracy of a state of the art recognition model, by 3.7% on average, across different benchmark datasets (image domains) and 24.5% on one of the weakly annotated datasets11Acknowledgements. The authors were supported by Czech Technical University student grant SGS20/171/0HK3/3TJ13, the MEYS VVV project CZ.02.1.01/0.010.0J16 019/0000765 Research Center for Informatics, the Spanish Research project TIN2017-89779-P and the CERCA Programme / Generalitat de Catalunya. Klára Janousková, Jiri Matas, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ICPR | 2 |
| 2020 | Ballroom Dance Recognition from Audio RecordingsabstractWe propose a CNN-based approach to classify ten genres of ballroom dances given audio recordings, five latin and five standard, namely Cha Cha Cha, Jive, Paso Doble, Rumba, Samba, Quickstep, Slow Foxtrot, Slow Waltz, Tango and Viennese Waltz. We utilize a spectrogram of an audio signal and we treat it as an image that is an input of the CNN. The classification is performed independently by 5-seconds spectrogram segments in sliding window fashion and the results are then aggregated. The method was tested on following datasets: Publicly available Extended Ballroom dataset collected by Marchand and Peeters, 2016 and two YouTube datasets collected by us, one in studio quality and the other, more challenging, recorded on mobile phones. The method achieved accuracy 93.9%, 96.7% and 89.8% respectively. The method runs in real-time. We implemented a web application to demonstrate the proposed method. Tomás Pavlín, Jan Cech, Jiri Matas |
ICPR | 3 |
| 2020 | DAL: A Deep Depth-Aware Long-term TrackerabstractThe best RGBD trackers provide high accuracy but are slow to run. On the other hand, the best RGB trackers are fast but clearly inferior on the RGBD datasets. In this work, we propose a deep depth-aware long-term tracker that achieves state-of-the-art RGBD tracking performance and is fast to run. We reformulate deep discriminative correlation filter (DCF) to embed the depth information into deep features. Moreover, the same depth-aware correlation filter is used for target redetection. Comprehensive evaluations show that the proposed tracker achieves state-of-the-art performance on the Princeton RGBD, STC, and the newly-released CDTB benchmarks and runs 20 fps. Yanlin Qian, Alan Lukezic, Matej Kristan, Joni-Kristian Kämäräinen, Jiri Matas |
ICPR | 6 |
| 2020 | Robust Audio-Based Vehicle Counting in Low-to-Moderate Traffic FlowabstractThe paper presents a method for audio-based vehicle counting (VC) in low-to-moderate traffic using one-channel sound. We formulate VC as a regression problem, i.e., we predict the distance between a vehicle and the microphone. Minima of the proposed distance function correspond to vehicles passing by the microphone. VC is carried out via local minima detection in the predicted distance. We propose to set the minima detection threshold at a point where the probabilities of false positives and false negatives coincide so they statistically cancel each other in total vehicle number. The method is trained and tested on a traffic-monitoring dataset comprising 422 short, 20-second one-channel sound files with a total of 1421 vehicles passing by the microphone. Relative VC error in a traffic location not used in the training is below 2% within a wide range of detection threshold values. Experimental results show that the regression accuracy in noisy environments is improved by introducing a novel high-frequency power feature. Slobodan Djukanovic, Jiri Matas, Tuomas Virtanen |
IV | 2 |
| 2020 | Fungi Recognition: A Practical Use CaseabstractThe paper presents a system for visual recognition of 1394 fungi species based on deep convolutional neural networks and its deployment in a citizen-science project. The system allows users to automatically identify observed specimens, while providing valuable data to biologists and computer vision researchers. The underlying classification method scored first in the FGVCx Fungi Classification Kaggle competition organized in connection with the Fine-Grained Visual Categorization (FGVC) workshop at CVPR 2018. We describe our winning submission and evaluate all technicalities that increased the recognition scores, and discuss the issues related to deployment of the system via the web- and mobile- interfaces. Milan Sulc, Lukás Picek, Jiri Matas, Thomas S. Jeppesen, Jacob Heilmann-Clausen |
WACV | 3 |
| 2020 | Saddle: Fast and repeatable features with good coverage
Javier Aldana-Iuit, Dmytro Mishkin, Ondrej Chum, Jiri Matas |
Image Vis. Comput. | 4 |
| 2020 | $\mathbb {H}$H-Patches: A Benchmark and Evaluation of Handcrafted and Learned Local DescriptorsabstractIn this paper, a novel benchmark is introduced for evaluating local image descriptors. We demonstrate limitations of the commonly used datasets and evaluation protocols, that lead to ambiguities and contradictory results in the literature. Furthermore, these benchmarks are nearly saturated due to the recent improvements in local descriptors obtained by learning from large annotated datasets. To address these issues, we introduce a new large dataset suitable for training and testing modern descriptors, together with strictly defined evaluation protocols in several tasks such as matching, retrieval and verification. This allows for more realistic, thus more reliable comparisons in different application scenarios. We evaluate the performance of several state-of-the-art descriptors and analyse their properties. We show that a simple normalisation of traditional hand-crafted descriptors is able to boost their performance to the level of deep learning based descriptors once realistic benchmarks are considered. Additionally we specify a protocol for learning and evaluating using cross validation. We show that when training state-of-the-art descriptors on this dataset, the traditional verification task is almost entirely saturated. Vassileios Balntas, Karel Lenc, Andrea Vedaldi, Tinne Tuytelaars, Jiri Matas, Krystian Mikolajczyk |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Restoration of Fast Moving ObjectsabstractIf an object is photographed at motion in front of a static background, the object will be blurred while the background sharp and partially occluded by the object. The goal is to recover the object appearance from such blurred image. We adopt the image formation model for fast moving objects and consider objects undergoing 2D translation and rotation. For this scenario we formulate the estimation of the object shape, appearance, and motion from a single image and known background as a constrained optimization problem with appropriate regularization terms. Both similarities and differences with blind deconvolution are discussed with the latter caused mainly by the coupling of the object appearance and shape in the acquisition model. Necessary conditions for solution uniqueness are derived and a numerical solution based on the alternating direction method of multipliers is presented. The proposed method is evaluated on a new dataset. Jan Kotera, Jiri Matas, Filip Sroubek |
IEEE Trans. Image Process. | 2 |
| 2019 | MAGSAC: Marginalizing Sample ConsensusabstractA method called, sigma-consensus, is proposed to eliminate the need for a user-defined inlier-outlier threshold in RANSAC. Instead of estimating the noise sigma, it is marginalized over a range of noise scales. The optimized model is obtained by weighted least-squares fitting where the weights come from the marginalization over sigma of the point likelihoods of being inliers. A new quality function is proposed not requiring sigma and, thus, a set of inliers to determine the model quality. Also, a new termination criterion for RANSAC is built on the proposed marginalization approach. Applying sigma-consensus, MAGSAC is proposed with no need for a user-defined sigma and improving the accuracy of robust estimation significantly. It is superior to the state-of-the-art in terms of geometric accuracy on publicly available real-world datasets for epipolar geometry (F and E) and homography estimation. In addition, applying sigma-consensus only once as a post-processing step to the RANSAC output always improved the model quality on a wide range of vision problems without noticeable deterioration in processing time, adding a few milliseconds. Daniel Barath, Jiri Matas, Jana Noskova |
CVPR | 2 |
| 2019 | Object Tracking by Reconstruction With View-Specific Discriminative Correlation FiltersabstractStandard RGB-D trackers treat the target as a 2D structure, which makes modelling appearance changes related even to out-of-plane rotation challenging. This limitation is addressed by the proposed long-term RGB-D tracker called OTR - Object Tracking by Reconstruction. OTR performs online 3D target reconstruction to facilitate robust learning of a set of view-specific discriminative correlation filters (DCFs). The 3D reconstruction supports two performance- enhancing features: (i) generation of an accurate spatial support for constrained DCF learning from its 2D projection and (ii) point-cloud based estimation of 3D pose change for selection and storage of view-specific DCFs which robustly localize the target after out-of-view rotation or heavy occlusion. Extensive evaluation on the Princeton RGB-D tracking and STC Benchmarks shows OTR outperforms the state-of-the-art by a large margin. Ugur Kart, Alan Lukezic, Matej Kristan, Joni-Kristian Kämäräinen, Jiri Matas |
CVPR | 5 |
| 2019 | On Finding Gray PixelsabstractWe propose a novel grayness index for finding gray pixels and demonstrate its effectiveness and efficiency in illumination estimation. The grayness index, GI in short, is derived using the Dichromatic Reflection Model and is learning-free. GI allows to estimate one or multiple illumination sources in color-biased images. On standard single-illumination and multiple-illumination estimation benchmarks, GI outperforms state-of-the-art statistical methods and many recent deep methods. GI is simple and fast, written in a few dozen lines of code, processing a 1080p image in ~0.4 seconds with a non-optimized Matlab code. Yanlin Qian, Joni-Kristian Kämäräinen, Jarno Nikkanen, Jiri Matas |
CVPR | 4 |
| 2019 | Progressive-X: Efficient, Anytime, Multi-Model Fitting AlgorithmabstractThe Progressive-X algorithm, Prog-X in short, is proposed for geometric multi-model fitting. The method interleaves sampling and consolidation of the current data interpretation via repetitive hypothesis proposal, fast rejection, and integration of the new hypothesis into the kept instance set by labeling energy minimization. Due to exploring the data progressively, the method has several beneficial properties compared with the state-of-the-art. First, a clear criterion, adopted from RANSAC, controls the termination and stops the algorithm when the probability of finding a new model with a reasonable number of inliers falls below a threshold. Second, Prog-X is an any-time algorithm. Thus, whenever is interrupted, e.g. due to a time limit, the returned instances cover real and, likely, the most dominant ones. The method is superior to the state-of-the-art in terms of accuracy in both synthetic experiments and on publicly available real-world datasets for homography, two-view motion, and motion segmentation. Daniel Barath, Jiri Matas |
ICCV | 2 |
| 2019 | CDTB: A Color and Depth Visual Object Tracking Dataset and BenchmarkabstractWe propose a new color-and-depth general visual object tracking benchmark (CDTB). CDTB is recorded by several passive and active RGB-D setups and contains indoor as well as outdoor sequences acquired in direct sunlight. The CDTB dataset is the largest and most diverse dataset in RGB-D tracking, with an order of magnitude larger number of frames than related datasets. The sequences have been carefully recorded to contain significant object pose change, clutter, occlusion, and periods of long-term target absence to enable tracker evaluation under realistic conditions. Sequences are per-frame annotated with 13 visual attributes for detailed analysis. Experiments with RGB and RGB-D trackers show that CDTB is more challenging than previous datasets. State-of-the-art RGB trackers outperform the recent RGB-D trackers, indicating a large gap between the two fields, which has not been previously detected by the prior benchmarks. Based on the results of the analysis we point out opportunities for future research in RGB-D tracker design. Alan Lukezic, Ugur Kart, Jani Käpylä, Ahmed Durmush, Joni-Kristian Kämäräinen, Jiri Matas, Matej Kristan |
ICCV | 6 |
| 2019 | Care Label RecognitionabstractThe paper introduces the problem of care label recognition and presents a method addressing it. A care label, also called a care tag, is a small piece of cloth or paper attached to a garment providing instructions for its maintenance and information about e.g. the material and size. The informationand instructions are written as symbols or plain text. Care label recognition is a challenging text and pictogram recognition problem - the often sewn text is small, looking as if printed using a non-standard font; the contrast of the text gradually fades, making OCR progressively more difficult. On the other hand, the information provided is typically redundant and thus it facilitates semi-supervised learning. The presented care label recognition method is based on the recently published End-to-End Method for Multi-LanguageScene Text, E2E-MLT, Busta et al. 2018, exploiting specific constraints, e.g. a care label vocabulary with multi-language equivalences. Experiments conducted on a newly-created dataset of 63 care label images show that even when exploiting problem-specific constraints, a state-of-the-art scene text detection and recognition method achieve precision and recall slightly above 0.6, confirming the challenging nature of the problem. Jiri Kralicek, Jiri Matas, Michal Busta |
ICDAR | 2 |
| 2019 | ICDAR2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition - RRC-MLT-2019abstractWith the growing cosmopolitan culture of modern cities, the need of robust Multi-Lingual scene Text (MLT) detection and recognition systems has never been more immense. With the goal to systematically benchmark and push the state-of-the-art forward, the proposed competition builds on top of the RRC-MLT-2017 with an additional end-to-end task, an additional language in the real images dataset, a large scale multi-lingual synthetic dataset to assist the training, and a baseline End-to-End recognition method. The real dataset consists of 20,000 images containing text from 10 languages. The challenge has 4 tasks covering various aspects of multi-lingual scene text: (a) text detection, (b) cropped word script classification, (c) joint text detection and script classification and (d) end-to-end detection and recognition. In total, the competition received 60 submissions from the research and industrial communities. This paper presents the dataset, the tasks and the findings of the presented RRC-MLT-2019 challenge. Nibal Nayef, Cheng-Lin Liu 0001, Jean-Marc Ogier, Michal Busta, Pinaki Nath Chowdhury, Dimosthenis Karatzas, Wafa Khlif, Jiri Matas, Umapada Pal 0001, Jean-Christophe Burie |
ICDAR | 9 |
| 2019 | Flash Lightens Gray PixeabstractIn the real world, a scene is usually cast by multiple illuminants and herein we address the problem of spatial illumination estimation. Our solution is based on detecting gray pixels with the help of flash photography. We show that flash photography significantly improves the performance of gray pixel detection without illuminant prior, training data or calibration of the flash. We also introduce a novel flash photography dataset generated from the MIT intrinsic dataset. Yanlin Qian, Joni-Kristian Kämäräinen, Jiri Matas |
ICIP | 4 |
| 2019 | Gyroscope-Aided Motion Deblurring with Deep NetworksabstractWe propose a deblurring method that incorporates gyroscope measurements into a convolutional neural network (CNN). With the help of such measurements, it can handle extremely strong and spatially-variant motion blur. At the same time, the image data is used to overcome the limitations of gyro-based blur estimation. To train our network, we also introduce a novel way of generating realistic training data using the gyroscope. The evaluation shows a clear improvement in visual quality over the state-of-the-art while achieving real-time performance. Furthermore, the method is shown to improve the performance of existing feature detectors and descriptors against the motion blur. Janne Mustaniemi, Juho Kannala, Simo Särkkä, Jiri Matas, Janne Heikkilä |
WACV | 4 |
| 2019 | Performance analysis of single-query 6-DoF camera pose estimation in self-driving setups
Junsheng Fu, Said Pertuz, Jiri Matas, Joni-Kristian Kämäräinen |
Comput. Vis. Image Underst. | 3 |
| 2019 | Cumulative attribute space regression for head pose estimation and color constancyabstractTwo-stage Cumulative Attribute (CA) regression has been found effective in regression problems of computer vision such as facial age and crowd density estimation. The first stage regression maps input features to cumulative attributes that encode correlations between target values. The previous works have dealt with single output regression. In this work, we propose cumulative attribute spaces for 2- and 3-output (multivariate) regression. We show how the original CA space can be generalized to multiple output by the Cartesian product (CartCA). However, for target spaces with more than two outputs the CartCA becomes computationally infeasible and therefore we propose an approximate solution - multi-view CA (MvCA) - where CartCA is applied to output pairs. We experimentally verify improved performance of the CartCA and MvCA spaces in 2D and 3D face pose estimation and three-output (RGB) illuminant estimation for color constancy. Ke Chen 0004, Kui Jia, Heikki Huttunen, Jiri Matas, Joni-Kristian Kämäräinen |
Pattern Recognit. | 4 |
| 2018 | FuCoLoT - A Fully-Correlational Long-Term Tracker
Alan Lukezic, Luka Cehovin, Tomás Vojír, Jiri Matas, Matej Kristan |
ACCV (2) | 4 |
| 2018 | Continual Occlusion and Optical Flow Estimation
Michal Neoral, Jan Sochman, Jiri Matas |
ACCV (4) | 3 |
| 2018 | Visual Heart Rate Estimation with Convolutional Neural Network
Radim Spetlík, Vojtech Franc, Jan Cech, Jiri Matas |
BMVC | 4 |
| 2018 | Graph-Cut RANSACabstractA novel method for robust estimation, called Graph-Cut RANSAC1, GC-RANSAC in short, is introduced. To separate inliers and outliers, it runs the graph-cut algorithm in the local optimization (LO) step which is applied when a so-far-the-best model is found. The proposed LO step is conceptually simple, easy to implement, globally optimal and efficient. GC-RANSAC is shown experimentally, both on synthesized tests and real image pairs, to be more geometrically accurate than state-of-the-art methods on a range of problems, e.g. line fitting, homography, affine transformation, fundamental and essential matrix estimation. It runs in real-time for many problems at a speed approximately equal to that of the less accurate alternatives (in milliseconds on standard CPU). Daniel Barath, Jiri Matas |
CVPR | 2 |
| 2018 | DeblurGAN: Blind Motion Deblurring Using Conditional Adversarial NetworksabstractWe present DeblurGAN, an end-to-end learned method for motion deblurring. The learning is based on a conditional GAN and the content loss. DeblurGAN achieves state-of-the art performance both in the structural similarity measure and visual appearance. The quality of the deblurring model is also evaluated in a novel way on a real-world problem - object detection on (de-)blurred images. The method is 5 times faster than the closest competitor - Deep-Deblur [25]. We also introduce a novel method for generating synthetic motion blurred images from sharp ones, allowing realistic dataset augmentation. The model, code and the dataset are available at https://github.com/KupynOrest/DeblurGAN. Orest Kupyn, Volodymyr Budzan, Mykola Mykhailych, Dmytro Mishkin, Jiri Matas |
CVPR | 5 |
| 2018 | Multi-class Model Fitting by Energy Minimization and Mode-Seeking
Daniel Barath, Jiri Matas |
ECCV (16) | 2 |
| 2018 | BOP: Benchmark for 6D Object Pose Estimation
Tomas Hodan, Frank Michel 0002, Eric Brachmann, Wadim Kehl, Anders Glent Buch, Dirk Kraft, Bertram Drost, Joel Vidal, Stephan Ihrke, Xenophon Zabulis, Caner Sahin, Fabian Manhardt, Federico Tombari, Tae-Kyun Kim 0001, Jiri Matas, Carsten Rother |
ECCV (10) | 15 |
| 2018 | Repeatability Is Not Enough: Learning Affine Regions via Discriminability
Dmytro Mishkin, Filip Radenovic, Jiri Matas |
ECCV (9) | 3 |
| 2018 | Detecting Decision Ambiguity from Facial ImagesabstractIn situations when potentially costly decisions are being made, faces of people tend to reflect a level of certainty about the appropriateness of the chosen decision. This fact is known from the psychological literature. In the paper, we propose a method that uses facial images for automatic detection of the decision ambiguity state of a subject. To train and test the method, we collected a large-scale dataset from "Who Wants to Be a Millionaire?" -- a popular TV game show. The videos provide examples of various mental states of contestants, including uncertainty, doubts and hesitation. The annotation of the videos is done automatically from on-screen graphics. The problem of detecting decision ambiguity is formulated as binary classification. Video-clips where a contestant asks for help (audience, friend, 50:50) are considered as positive samples; if he (she) replies directly as negative ones. We propose a baseline method combining a deep convolutional neural network with an SVM. The method has an error rate of 24%. The error of human volunteers on the same dataset is 45%, close to chance. Pavel Jahoda, Antonín Vobecký, Jan Cech, Jiri Matas |
FG | 4 |
| 2018 | Non-Contact Reflectance Photoplethysmography: Progress, Limitations, and MythsabstractPhotoplethysmography (PPG) is a non-invasive method of measuring changes of blood volume in human tissue. The literature on non-contact reflectance PPG related to cardiovascular activity is extensively reviewed. We identify key factors limiting the performance of the PPG methods and reproducibility of the research as: a lack of publicly available datasets and incomplete description of data used in published experiments (missing details on video compression, lighting setup and subject's skin type), use of unreliable pulse oximeter devices for ground-truth reference and missing standard experimental protocols. Two experiments with 5 participants are presented showing that the quality of the reconstructed signal (1) is adversely affected by a reduction of spatial resolution that also amplifies the effects of H.264 video compression and (2) is improved by precise pixel-to-pixel stabilization. Radim Spetlík, Jan Cech, Jiri Matas |
FG | 3 |
| 2018 | Depth Masked Discriminative Correlation FilterabstractDepth information provides a strong cue for occlusion detection and handling, but has been largely omitted in generic object tracking until recently due to lack of suitable benchmark datasets and applications. In this work, we propose a Depth Masked Discriminative Correlation Filter (DM-DCF) which adopts novel depth segmentation based occlusion detection that stops correlation filter updating and depth masking which adaptively adjusts the spatial support for correlation filter. In Princeton RGBD Tracking Benchmark, our DM-DCF is among the state-of-the-art in overall ranking and the winner on multiple categories. Moreover, since it is based on DCF, “DM-DCF” runs an order of magnitude faster than its competitors making it suitable for time constrained applications. Ugur Kart, Joni-Kristian Kämäräinen, Jiri Matas, Lixin Fan, Francesco Cricri |
ICPR | 3 |
| 2018 | Fast Motion Deblurring for Feature Detection and Matching Using Inertial MeasurementsabstractMany computer vision and image processing applications rely on local features. It is well-known that motion blur decreases the performance of traditional feature detectors and descriptors. We propose an inertial-based deblurring method for improving the robustness of existing feature detectors and descriptors against the motion blur. Unlike most deblurring algorithms, the method can handle spatially-variant blur and rolling shutter distortion. Furthermore, it is capable of running in real-time contrary to state-of-the-art algorithms. The limitations of inertial-based blur estimation are taken into account by validating the blur estimates using image data. The evaluation shows that when the method is used with traditional feature detector and descriptor, it increases the number of detected keypoints, provides higher repeatability and improves the localization accuracy. We also demonstrate that such features will lead to more accurate and complete reconstructions when used in the application of 3D visual reconstruction. Janne Mustaniemi, Juho Kannala, Simo Särkkä, Jiri Matas, Janne Heikkilä |
ICPR | 4 |
| 2018 | ALFA: Agglomerative Late Fusion Algorithm for Object DetectionabstractWe propose ALFA - a novel late fusion algorithm for object detection. ALFA is based on agglomerative clustering of object detector predictions taking into consideration both the bounding box locations and the class scores. Each cluster represents a single object hypothesis whose location is a weighted combination of the clustered bounding boxes. ALFA was evaluated using combinations of a pair (SSD and DeNet) and a triplet (SSD, DeNet and Faster R-CNN) of recent object detectors that are close to the state-of-the-art. ALFA achieves state of the art results on PASCAL VOC 2007 and PASCAL VOC 2012, outperforming the individual detectors as well as baseline combination strategies, achieving up to 32% lower error than the best individual detectors and up to 6% lower error than the reference fusion algorithm DBF - Dynamic Belief Fusion. Evgenii Razinkov, Iuliia Saveleva, Jiri Matas |
ICPR | 3 |
| 2018 | Discriminative Correlation Filter Tracker with Channel and Spatial Reliability
Alan Lukezic, Tomás Vojír, Luka Cehovin, Jiri Matas, Matej Kristan |
Int. J. Comput. Vis. | 4 |
| 2017 | Discriminative Correlation Filter with Channel and Spatial ReliabilityabstractShort-term tracking is an open and challenging problem for which discriminative correlation filters (DCF) have shown excellent performance.We introduce the channel and spatial reliability concepts to DCF tracking and provide a novel learning algorithm for its efficient and seamless integration in the filter update and the tracking process. The spatial reliability map adjusts the filter support to the part of the object suitable for tracking. This both allows to enlarge the search region and improves tracking of non-rectangular objects. Reliability scores reflect channel-wise quality of the learned filters and are used as feature weighting coefficients in localization. Experimentally, with only two simple standard features, HoGs and Colornames, the novel CSRDCF method - DCF with Channel and Spatial Reliability - achieves state-of-the-art results on VOT 2016, VOT 2015 and OTB100. The CSR-DCF runs in real-time on a CPU. Alan Lukezic, Tomás Vojír, Luka Cehovin, Jiri Matas, Matej Kristan |
CVPR | 4 |
| 2017 | The World of Fast Moving ObjectsabstractThe notion of a Fast Moving Object (FMO), i.e. an object that moves over a distance exceeding its size within the exposure time, is introduced. FMOs may, and typically do, rotate with high angular speed. FMOs are very common in sports videos, but are not rare elsewhere. In a single frame, such objects are often barely visible and appear as semitransparent streaks. A method for the detection and tracking of FMOs is proposed. The method consists of three distinct algorithms, which form an efficient localization pipeline that operates successfully in a broad range of conditions. We show that it is possible to recover the appearance of the object and its axis of rotation, despite its blurred appearance. The proposed method is evaluated on a new annotated dataset. The results show that existing trackers are inadequate for the problem of FMO localization and a new approach is required. Two applications of localization, temporal superresolution and highlighting, are presented. Denys Rozumnyi, Jan Kotera, Filip Sroubek, Lukás Novotný, Jiri Matas |
CVPR | 5 |
| 2017 | Deep TextSpotter: An End-to-End Trainable Scene Text Localization and Recognition FrameworkabstractA method for scene text localization and recognition is proposed. The novelties include: training of both text detection and recognition in a single end-to-end pass, the structure of the recognition CNN and the geometry of its input layer that preserves the aspect of the text and adapts its resolution to the data.,,The proposed method achieves state-of-the-art accuracy in the end-to-end text recognition on two standard datasets – ICDAR 2013 and ICDAR 2015, whilst being an order of magnitude faster than competing methods - the whole pipeline runs at 10 frames per second on an NVidia K80 GPU. Michal Busta, Lukás Neumann, Jiri Matas |
ICCV | 3 |
| 2017 | Recurrent Color ConstancyabstractWe introduce a novel formulation of temporal color constancy which considers multiple frames preceding the frame for which illumination is estimated. We propose an end-to-end trainable recurrent color constancy network – the RCC-Net – which exploits convolutional LSTMs and a simulated sequence to learn compositional representations in space and time. We use a standard single frame color constancy benchmark, the SFU Gray Ball Dataset, which can be adapted to a temporal setting. Extensive experiments show that the proposed method consistently outperforms single-frame state-of-the-art methods and their temporal variants. Yanlin Qian, Ke Chen 0004, Jarno Nikkanen, Joni-Kristian Kämäräinen, Jiri Matas |
ICCV | 5 |
| 2017 | ICDAR2017 Robust Reading Challenge on COCO-TextabstractThis report presents the final results of the ICDAR 2017 Robust Reading Challenge on COCO-Text. A challenge on scene text detection and recognition based on the largest real scene text dataset currently available: the COCO-Text dataset. The competition is structured around three tasks: Text Localization, Cropped Word Recognition and End-To-End Recognition. The competition received a total of 27 submissions over the different opened tasks. This report describes the datasets and the ground truth, details the performance evaluation protocols used and presents the final results along with a brief summary of the participating methods. Raul Gomez, Baoguang Shi, Lluís Gómez i Bigorda, Lukás Neumann, Andreas Veit, Jiri Matas, Serge J. Belongie, Dimosthenis Karatzas |
ICDAR | 6 |
| 2017 | Inertial-based scale estimation for structure from motion on mobile devicesabstractStructure from motion algorithms have an inherent limitation that the reconstruction can only be determined up to the unknown scale factor. Modern mobile devices are equipped with an inertial measurement unit (IMU), which can be used for estimating the scale of the reconstruction. We propose a method that recovers the metric scale given inertial measurements and camera poses. In the process, we also perform a temporal and spatial alignment of the camera and the IMU. Therefore, our solution can be easily combined with any existing visual reconstruction software. The method can cope with noisy camera pose estimates, typically caused by motion blur or rolling shutter artifacts, via utilizing a Rauch-Tung-Striebel (RTS) smoother. Furthermore, the scale estimation is performed in the frequency domain, which provides more robustness to inaccurate sensor time stamps and noisy IMU samples than the previously used time domain representation. In contrast to previous methods, our approach has no parameters that need to be tuned for achieving a good performance. In the experiments, we show that the algorithm outperforms the state-of-the-art in both accuracy and convergence speed of the scale estimate. The accuracy of the scale is around 1% from the ground truth depending on the recording. We also demonstrate that our method can improve the scale accuracy of the Project Tango's build-in motion tracking. Janne Mustaniemi, Juho Kannala, Simo Särkkä, Jiri Matas, Janne Heikkilä |
IROS | 4 |
| 2017 | Visual Descriptors in Methods for Video HyperlinkingabstractIn this paper, we survey different state-of-the-art visual processing methods and utilize them in hyperlinking. Visual information, calculated using Features Signatures, SIMILE descriptors and convolutional neural networks (CNN), is utilized as similarity between video frames and used to find similar faces, objects and setting. Visual concepts in frames are also automatically recognized and textual output of the recognition is combined with search based on subtitles and transcripts. All presented experiments were performed in the Search and Hyperlinking 2014 MediaEval task and Video Hyperlinking 2015 TRECVid task. Petra Galuscáková, Michal Batko, Jan Cech, Jiri Matas, David Novak, Pavel Pecina |
ICMR | 4 |
| 2017 | Working hard to know your neighbor's margins: Local descriptor learning lossabstractWe introduce a loss for metric learning, which is inspired by the Lowe's matching criterion for SIFT. We show that the proposed loss, that maximizes the distance between the closest positive and closest negative example in the batch, is better than complex regularization methods; it works well for both shallow and deep convolution network architectures. Applying the novel loss to the L2Net CNN architecture results in a compact descriptor named HardNet. It has the same dimensionality as SIFT (128) and shows state-of-art performance in wide baseline stereo, patch verification and instance retrieval benchmarks. Anastasiya Mishchuk, Dmytro Mishkin, Filip Radenovic, Jiri Matas |
NIPS | 4 |
| 2017 | T-LESS: An RGB-D Dataset for 6D Pose Estimation of Texture-Less ObjectsabstractWe introduce T-LESS, a new public dataset for estimating the 6D pose, i.e. translation and rotation, of texture-less rigid objects. The dataset features thirty industry-relevant objects with no significant texture and no discriminative color or reflectance properties. The objects exhibit symmetries and mutual similarities in shape and/or size. Compared to other datasets, a unique property is that some of the objects are parts of others. The dataset includes training and test images that were captured with three synchronized sensors, specifically a structured-light and a time-of-flight RGB-D sensor and a high-resolution RGB camera. There are approximately 39K training and 10K test images from each sensor. Additionally, two types of 3D models are provided for each object, i.e. a manually created CAD model and a semi-automatically reconstructed one. Training images depict individual objects against a black background. Test images originate from twenty test scenes having varying complexity, which increases from simple scenes with several isolated objects to very challenging ones with multiple instances of several objects and with a high amount of clutter and occlusion. The images were captured from a systematically sampled view sphere around the object/scene, and are annotated with accurate ground truth 6D poses of all modeled objects. Initial evaluation results indicate that the state of the art in 6D object pose estimation has ample room for improvement, especially in difficult cases with significant occlusion. The T-LESS dataset is available online at cmp:felk:cvut:cz/t-less. Tomas Hodan, Pavel Haluza, Stepán Obdrzálek, Jiri Matas, Manolis I. A. Lourakis, Xenophon Zabulis |
WACV | 4 |
| 2017 | Systematic evaluation of convolution neural network advances on the Imagenet
Dmytro Mishkin, Nikolay Sergievskiy, Jiri Matas |
Comput. Vis. Image Underst. | 3 |
| 2016 | Accurate Closed-form Estimation of Local Affine Transformations Consistent with the Epipolar Geometry
Daniel Barath, Jiri Matas, Levente Hajder |
BMVC | 2 |
| 2016 | Multi-H: Efficient recovery of tangent planes in stereo images
Daniel Barath, Jiri Matas, Levente Hajder |
BMVC | 2 |
| 2016 | From Dusk Till Dawn: Modeling in the DarkabstractInternet photo collections naturally contain a large variety of illumination conditions, with the largest difference between day and night images. Current modeling techniques do not embrace the broad illumination range often leading to reconstruction failure or severe artifacts. We present an algorithm that leverages the appearance variety to obtain more complete and accurate scene geometry along with consistent multi-illumination appearance information. The proposed method relies on automatic scene appearance grouping, which is used to obtain separate dense 3D models. Subsequent model fusion combines the separate models into a complete and accurate reconstruction of the scene. In addition, we propose a method to derive the appearance information for the model under the different illumination conditions, even for scene parts that are not observed under one illumination condition. To achieve this, we develop a cross-illumination color transfer technique. We evaluate our method on a large variety of landmarks from across Europe reconstructed from a database of 7.4M images. Filip Radenovic, Johannes L. Schönberger, Dinghuang Ji, Jan-Michael Frahm, Ondrej Chum, Jiri Matas |
CVPR | 6 |
| 2016 | In the Saddle: Chasing fast and repeatable featuresabstractA novel similarity-covariant feature detector that extracts points whose neighborhoods, when treated as a 3D intensity surface, have a saddle-like intensity profile. The saddle condition is verified efficiently by intensity comparisons on two concentric rings that must have exactly two dark-to-bright and two bright-to-dark transitions satisfying certain geometric constraints. Experiments show that the Saddle features are general, evenly spread and appearing in high density in a range of images. The Saddle detector is among the fastest proposed. In comparison with detector with similar speed, the Saddle features show superior matching performance on number of challenging datasets. Javier Aldana-Iuit, Dmytro Mishkin, Ondrej Chum, Jiri Matas |
ICPR | 4 |
| 2016 | Deep structured-output regression learning for computational color constancyabstractThe color constancy problem is addressed by structured-output regression on the values of the fully-connected layers of a convolutional neural network. The AlexNet and the VGG are considered and VGG slightly outperformed AlexNet. Best results were obtained with the first fully-connected “fc6” layer and with multi-output support vector regression. Experiments on the SFU Color Checker and Indoor Dataset benchmarks demonstrate that our method achieves competitive performance, outperforming the state of the art on the SFU indoor benchmark. Yanlin Qian, Ke Chen 0004, Joni-Kristian Kämäräinen, Jarno Nikkanen, Jiri Matas |
ICPR | 5 |
| 2016 | Online adaptive hidden Markov model for multi-tracker fusion
Tomás Vojír, Jiri Matas, Jana Noskova |
Comput. Vis. Image Underst. | 2 |
| 2016 | Multi-view facial landmark detection by using a 3D shape model
Jan Cech, Vojtech Franc, Michal Uricár, Jiri Matas |
Image Vis. Comput. | 4 |
| 2016 | Detection of bubbles as concentric circular arrangementsabstractThe paper proposes a method for the detection of bubble-like transparent objects in a liquid. The detection problem is non-trivial since bubble appearance varies considerably due to different lighting conditions causing contrast reversal and multiple interreflections. We formulate the problem as the detection of concentric circular arrangements (CCA). The CCAs are recovered in a hypothesize-optimize-verify framework. The hypothesis generation is based on sampling from the partially linked components of the non-maximum suppressed responses of oriented ridge filters, and is followed by the CCA parameter estimation. Parameter optimization is carried out by minimizing a novel cost-function. The performance was tested on gas dispersion images of pulp suspension and oil dispersion images. The mean error of gas/oil volume estimation was used as a performance criterion due to the fact that the main goal of the applications driving the research was the bubble volume estimation. The method achieved 28 and 13 % of gas and oil volume estimation errors correspondingly outperforming the OpenCV Circular Hough Transform in both cases and the WaldBoost detector in gas volume estimation. Nataliya Strokina, Jiri Matas, Tuomas Eerola, Lasse Lensu, Heikki Kälviäinen |
Mach. Vis. Appl. | 2 |
| 2016 | A Novel Performance Evaluation Methodology for Single-Target TrackersabstractThis paper addresses the problem of single-target tracker performance evaluation. We consider the performance measures, the dataset and the evaluation system to be the most important components of tracker evaluation and propose requirements for each of them. The requirements are the basis of a new evaluation methodology that aims at a simple and easily interpretable tracker comparison. The ranking-based methodology addresses tracker equivalence in terms of statistical significance and practical differences. A fully-annotated dataset with per-frame annotations with several visual attributes is introduced. The diversity of its visual properties is maximized in a novel way by clustering a large number of videos according to their visual attributes. This makes it the most sophistically constructed and annotated dataset to date. A multi-platform evaluation system allowing easy integration of third-party trackers is presented as well. The proposed evaluation methodology was tested on the VOT2014 challenge on the new dataset and 38 trackers, making it the largest benchmark to date. Most of the tested trackers are indeed state-of-the-art since they outperform the standard baselines, resulting in a highly-challenging benchmark. An exhaustive analysis of the dataset from the perspective of tracking difficulty is carried out. To facilitate tracker comparison a new performance visualization technique is proposed. Matej Kristan, Jiri Matas, Ales Leonardis, Tomás Vojír, Roman P. Pflugfelder, Gustavo Fernández, Georg Nebehay, Fatih Porikli, Luka Cehovin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Real-Time Lexicon-Free Scene Text Localization and RecognitionabstractAn end-to-end real-time text localization and recognition method is presented. Its real-time performance is achieved by posing the character detection and segmentation problem as an efficient sequential selection from the set of Extremal Regions. The ER detector is robust against blur, low contrast and illumination, color and texture variation. In the first stage, the probability of each ER being a character is estimated using features calculated by a novel algorithm in constant time and only ERs with locally maximal probability are selected for the second stage, where the classification accuracy is improved using computationally more expensive features. A highly efficient clustering algorithm then groups ERs into text lines and an OCR classifier trained on synthetic fonts is exploited to label character regions. The most probable character sequence is selected in the last stage when the context of each character is known. The method was evaluated on three public datasets. On the ICDAR 2013 dataset the method achieves state-of-the-art results in text localization; on the more challenging SVT dataset, the proposed method significantly outperforms the state-of-the-art methods and demonstrates that the proposed pipeline can incorporate additional prior knowledge about the detected text. The proposed method was exploited as the baseline in the ICDAR 2015 Robust Reading competition, where it compares favourably to the state-of-the art. Lukás Neumann, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Texture-Independent Long-Term Tracking Using Virtual CornersabstractLong-term tracking of an object, given only a single instance in an initial frame, remains an open problem. We propose a visual tracking algorithm, robust to many of the difficulties that often occur in real-world scenes. Correspondences of edge-based features are used, to overcome the reliance on the texture of the tracked object and improve invariance to lighting. Furthermore, we address long-term stability, enabling the tracker to recover from drift and to provide redetection following object disappearance or occlusion. The two-module principle is similar to the successful state-of-the-art long-term TLD tracker; however, our approach offers better performance in benchmarks and extends to cases of low-textured objects. This becomes obvious in cases of plain objects with no texture at all, where the edge-based approach proves the most beneficial. We perform several different experiments to validate the proposed method. First, results on short-term sequences show the performance of tracking challenging (low textured and/or transparent) objects that represent failure cases for competing the state-of-the-art approaches. Second, long sequences are tracked, including one of almost 30 000 frames, which, to the best of our knowledge, is the longest tracking sequence reported to date. This tests the redetection and drift resistance properties of the tracker. Finally, we report the results of the proposed tracker on the VOT Challenge 2013 and 2014 data sets as well as on the VTB1.0 benchmark, and we show relative performance of the tracker compared with its competitors. All the results are comparable with the state of the art on sequences with textured objects and superior on non-textured objects. The new annotated sequences are made publicly available. Karel Lebeda, Simon Hadfield, Jiri Matas, Richard Bowden |
IEEE Trans. Image Process. | 3 |
| 2015 | WxBS: Wide Baseline Stereo GeneralizationsabstractWe have presented a new problem -- the wide multiple baseline stereo (WxBS) -- which considers matching of images that simultaneously differ in more than one image acquisition factor such as viewpoint, illumination, sensor type or where object appearance changes significantly, e.g. over time. A new dataset with the ground truth for evaluation of matching algorithms has been introduced and will be made public. We have extensively tested a large set of popular and recent detectors and descriptors and show than the combination of RootSIFT and HalfRootSIFT as descriptors with MSER and Hessian-Affine detectors works best for many different nuisance factors. We show that simple adaptive thresholding improves Hessian-Affine, DoG, MSER (and possibly other) detectors and allows to use them on infrared and low contrast images. A novel matching algorithm for addressing the WxBS problem has been introduced. We have shown experimentally that the WxBS-M matcher dominantes the state-of-the-art methods both on both the new and existing datasets. Dmytro Mishkin, Jiri Matas, Michal Perdoch, Karel Lenc |
BMVC | 2 |
| 2015 | FASText: Efficient Unconstrained Scene Text DetectorabstractWe propose a novel easy-to-implement stroke detector based on an efficient pixel intensity comparison to surrounding pixels. Stroke-specific keypoints are efficiently detected and text fragments are subsequently extracted by local thresholding guided by keypoint properties. Classification based on effectively calculated features then eliminates non-text regions. The stroke-specific keypoints produce 2 times less region segmentations and still detect 25% more characters than the commonly exploited MSER detector and the process is 4 times faster. After a novel efficient classification step, the number of regions is reduced to 7 times less than the standard method and is still almost 3 times faster. All stages of the proposed pipeline are scale-and rotation-invariant and support a wide variety of scripts (Latin, Hebrew, Chinese, etc.) and fonts. When the proposed detector is plugged into a scene text localization and recognition pipeline, a state-of-the-art text localization accuracy is maintained whilst the processing time is significantly reduced. Michal Busta, Lukás Neumann, Jiri Matas |
ICCV | 3 |
| 2015 | Cascaded Sparse Spatial Bins for Efficient and Effective Generic Object DetectionabstractA novel efficient method for extraction of object proposals is introduced. Its "objectness" function exploits deep spatial pyramid features, a novel fast-to-compute HoG-based edge statistic and the EdgeBoxes score [42]. The efficiency is achieved by the use of spatial bins in a novel combination with sparsity-inducing group normalized SVM. State-of-the-art recall performance is achieved on Pascal VOC07, significantly outperforming methods with comparable speed. Interestingly, when only 100 proposals per image are considered the method attains 78 % recall on VOC07. The method improves mAP of the RCNN class-specific detector, increasing it by 10 points when only 50 proposals are used in each image. The system trained on twenty classes performs well on the two hundred class ILSVRC2013 set confirming generalization capability. David R. Novotny, Jiri Matas |
ICCV | 2 |
| 2015 | ICDAR 2015 competition on Robust ReadingabstractResults of the ICDAR 2015 Robust Reading Competition are presented. A new Challenge 4 on Incidental Scene Text has been added to the Challenges on Born-Digital Images, Focused Scene Images and Video Text. Challenge 4 is run on a newly acquired dataset of 1,670 images evaluating Text Localisation, Word Recognition and End-to-End pipelines. In addition, the dataset for Challenge 3 on Video Text has been substantially updated with more video sequences and more accurate ground truth data. Finally, tasks assessing End-to-End system performance have been introduced to all Challenges. The competition took place in the first quarter of 2015, and received a total of 44 submissions. Only the tasks newly introduced in 2015 are reported on. The datasets, the ground truth specification and the evaluation protocols are presented together with the results and a brief summary of the participating methods. Dimosthenis Karatzas, Lluís Gómez i Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, Masakazu Iwamura, Jiri Matas, Lukás Neumann, Vijay Chandrasekhar 0001, Shijian Lu, Faisal Shafait, Seiichi Uchida, Ernest Valveny |
ICDAR | 7 |
| 2015 | Towards visual words to wordsabstractWe address the problem of text localization and retrieval in real world images. We are first to study the retrieval of text images, i.e. the selection of images containing text in large collections at high speed. We propose a novel representation, textual visual words, which describe text by generic visual words that geometrically consistently predict bottom and top lines of text. The visual words are discretized SIFT descriptors of Hessian features. The features may correspond to various structures present in the text - character fragments, individual characters or their arrangements. The textual words representation is invariant to affine transformation of the image and local linear change of intensity. Experiments demonstrate that the proposed method outperforms the state-of-the-art on the MS dataset. The proposed method detects blurry, small font, low contrast, noisy text from real world images. Rakesh Mehta, Ondrej Chum, Jiri Matas |
ICDAR | 3 |
| 2015 | Efficient Scene text localization and recognition with local character refinementabstractAn unconstrained end-to-end text localization and recognition method is presented. The method detects initial text hypothesis in a single pass by an efficient region-based method and subsequently refines the text hypothesis using a more robust local text model, which deviates from the common assumption of region-based methods that all characters are detected as connected components. Lukás Neumann, Jiri Matas |
ICDAR | 2 |
| 2015 | Detection and fine 3D pose estimation of texture-less objects in RGB-D imagesabstractDespite their ubiquitous presence, texture-less objects present significant challenges to contemporary visual object detection and localization algorithms. This paper proposes a practical method for the detection and accurate 3D localization of multiple texture-less and rigid objects depicted in RGB-D images. The detection procedure adopts the sliding window paradigm, with an efficient cascade-style evaluation of each window location. A simple pre-filtering is performed first, rapidly rejecting most locations. For each remaining location, a set of candidate templates (i.e. trained object views) is identified with a voting procedure based on hashing, which makes the method's computational complexity largely unaffected by the total number of known objects. The candidate templates are then verified by matching feature points in different modalities. Finally, the approximate object pose associated with each detected template is used as a starting point for a stochastic optimization procedure that estimates accurate 3D pose. Experimental evaluation shows that the proposed method yields a recognition rate comparable to the state of the art, while its complexity is sub-linear in the number of templates. Tomas Hodan, Xenophon Zabulis, Manolis I. A. Lourakis, Stepán Obdrzálek, Jiri Matas |
IROS | 5 |
| 2015 | MODS: Fast and robust method for two-view matching
Dmytro Mishkin, Jiri Matas, Michal Perdoch |
Comput. Vis. Image Underst. | 2 |
| 2014 | Efficient Image Detail Mining
Andrej Mikulík, Filip Radenovic, Ondrej Chum, Jiri Matas |
ACCV (2) | 4 |
| 2014 | Rectification, and Segmentation of Coplanar Repeated PatternsabstractThis paper presents a novel and general method for the detection, rectification and segmentation of imaged coplanar repeated patterns. The only assumption made of the scene geometry is that repeated scene elements are mapped to each other by planar Euclidean transformations. The class of patterns covered is broad and includes nearly all commonly seen, planar, man-made repeated patterns. In addition, novel linear constraints are used to reduce geometric ambiguity between the rectified imaged pattern and the scene pattern. Rectification to within a similarity of the scene plane is achieved from one rotated repeat, or to within a similarity with a scale ambiguity along the axis of symmetry from one reflected repeat. A stratum of constraints is derived that gives the necessary configuration of repeats for each successive level of rectification. A generative model for the imaged pattern is inferred and used to segment the pattern with pixel accuracy. Qualitative results are shown on a broad range of image types on which state-of-the-art methods fail. James Pritts, Ondrej Chum, Jiri Matas |
CVPR | 3 |
| 2014 | A 3D Approach to Facial Landmarks: Detection, Refinement, and TrackingabstractA real-time algorithm for accurate localization of facial landmarks in a single monocular image is proposed. The algorithm is formulated as an optimization problem, in which the sum of responses of local classifiers is maximized with respect to the camera pose by fitting a generic (not a person-specific) 3D model. The algorithm simultaneously estimates a head position and orientation and detects the facial landmarks in the image. Despite being local, we show that the basin of attraction is large to the extent it can be initialized by a scanning window face detector. Other experiments on standard datasets demonstrate that the proposed algorithm outperforms a state-of-the-art landmark detector especially for non-frontal face images, and that it is capable of reliable and stable tracking for large set of viewing angles. Jan Cech, Vojtech Franc, Jiri Matas |
ICPR | 3 |
| 2014 | Matching of Images of Non-planar Objects with View Synthesis
Dmytro Mishkin, Jiri Matas |
SOFSEM | 2 |
| 2014 | Erratum to: Learning Vocabularies over a Fine Quantization
Andrej Mikulík, Michal Perdoch, Ondrej Chum, Jiri Matas |
Int. J. Comput. Vis. | 4 |
| 2014 | Robust scale-adaptive mean-shift for tracking
Tomás Vojír, Jana Noskova, Jiri Matas |
Pattern Recognit. Lett. | 3 |
| 2013 | Scene Text Localization and Recognition with Oriented Stroke DetectionabstractAn unconstrained end-to-end text localization and recognition method is presented. The method introduces a novel approach for character detection and recognition which combines the advantages of sliding-window and connected component methods. Characters are detected and recognized as image regions which contain strokes of specific orientations in a specific relative position, where the strokes are efficiently detected by convolving the image gradient field with a set of oriented bar filters. Additionally, a novel character representation efficiently calculated from the values obtained in the stroke detection phase is introduced. The representation is robust to shift at the stroke level, which makes it less sensitive to intra-class variations and the noise induced by normalizing character size and positioning. The effectiveness of the representation is demonstrated by the results achieved in the classification of real-world characters using an euclidian nearest-neighbor classifier trained on synthetic data in a plain form. The method was evaluated on a standard dataset, where it achieves state-of-the-art results in both text localization and recognition. Lukás Neumann, Jiri Matas |
ICCV | 2 |
| 2013 | On Combining Multiple Segmentations in Scene Text RecognitionabstractAn end-to-end real-time scene text localization and recognition method is presented. The three main novel features are: (i) keeping multiple segmentations of each character until the very last stage of the processing when the context of each character in a text line is known, (ii) an efficient algorithm for selection of character segmentations minimizing a global criterion, and (iii) showing that, despite using theoretically scale-invariant methods, operating on a coarse Gaussian scale space pyramid yields improved results as many typographical artifacts are eliminated. The method runs in real time and achieves state-of-the-art text localization results on the ICDAR 2011 Robust Reading dataset. Results are also reported for end-to-end text recognition on the ICDAR 2011 dataset. Lukás Neumann, Jiri Matas |
ICDAR | 2 |
| 2013 | Fast Detection of Multiple Textureless 3-D Objects
Hongping Cai, Tomás Werner, Jiri Matas |
ICVS | 3 |
| 2013 | Image Retrieval for Online Browsing in Large Image Collections
Andrej Mikulík, Ondrej Chum, Jiri Matas |
SISAP | 3 |
| 2013 | Learning Vocabularies over a Fine Quantization
Andrej Mikulík, Michal Perdoch, Ondrej Chum, Jiri Matas |
Int. J. Comput. Vis. | 4 |
| 2013 | USAC: A Universal Framework for Random Sample ConsensusabstractA computational problem that arises frequently in computer vision is that of estimating the parameters of a model from data that have been contaminated by noise and outliers. More generally, any practical system that seeks to estimate quantities from noisy data measurements must have at its core some means of dealing with data contamination. The random sample consensus (RANSAC) algorithm is one of the most popular tools for robust estimation. Recent years have seen an explosion of activity in this area, leading to the development of a number of techniques that improve upon the efficiency and robustness of the basic RANSAC algorithm. In this paper, we present a comprehensive overview of recent research in RANSAC-based robust estimation by analyzing and comparing various approaches that have been explored over the years. We provide a common context for this analysis by introducing a new framework for robust estimation, which we call Universal RANSAC (USAC). USAC extends the simple hypothesize-and-verify structure of standard RANSAC to incorporate a number of important practical and computational considerations. In addition, we provide a general-purpose C++ software library that implements the USAC framework by leveraging state-of-the-art algorithms for the various modules. This implementation thus addresses many of the limitations of standard RANSAC within a single unified package. We benchmark the performance of the algorithm on a large collection of estimation problems. The implementation we provide can be used by researchers either as a stand-alone tool for robust estimation or as a benchmark for evaluating new techniques. Rahul Raguram, Ondrej Chum, Marc Pollefeys, Jiri Matas, Jan-Michael Frahm |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Fixing the Locally Optimized RANSACabstractThe paper revisits the problem of local optimization for RANSAC. Improvements of the LO-RANSAC procedure are proposed: a use of truncated quadratic cost function, an introduction of a limit on the number of inliers used for the least squares computation and several implementation issues are addressed. The implementation is made publicly available. Extensive experiments demonstrate that the novel algorithm called LO +-RANSAC is (1) very stable (almost non-random in nature), (2) very precise in a broad range of conditions, (3) less sensitive to the choice of inlier-outlier threshold and (4) it offers a significantly better starting point for bundle adjustment than the Gold Standard method advocated in the Hartley-Zisserman book. 1 Karel Lebeda, Jiri Matas, Ondrej Chum |
BMVC | 2 |
| 2012 | Visual Tracking in the 21st Century
Jiri Matas |
BMVC | 1 |
| 2012 | Fast computation of min-Hash signatures for image collectionsabstractA new method for highly efficient min-Hash generation for document collections is proposed. It exploits the inverted file structure which is available in many applications based on a bag or a set of words. Fast min-Hash generation is important in applications such as image clustering where good recall and precision requires a large number of min-Hash signatures. Using the set of words represenation, the novel exact min-Hash generation algorithm achieves approximately a 50-fold speed-up on two dataset with 105and 106images respectively. We also propose an approximate min-Hash assignment process which reaches a more than 200-fold speed-up at the cost of missing about 2–3% of matches. We also experimentally show that the method generalizes to other modalities with significantly different statistics. Ondrej Chum, Jiri Matas |
CVPR | 2 |
| 2012 | Real-time scene text localization and recognitionabstractAn end-to-end real-time scene text localization and recognition method is presented. The real-time performance is achieved by posing the character detection problem as an efficient sequential selection from the set of Extremal Regions (ERs). The ER detector is robust to blur, illumination, color and texture variation and handles low-contrast text. In the first classification stage, the probability of each ER being a character is estimated using novel features calculated with O(1) complexity per region tested. Only ERs with locally maximal probability are selected for the second stage, where the classification is improved using more computationally expensive features. A highly efficient exhaustive search with feedback loops is then applied to group ERs into words and to select the most probable character segmentation. Finally, text is recognized in an OCR stage trained using synthetic fonts. The method was evaluated on two public datasets. On the ICDAR 2011 dataset, the method achieves state-of-the-art text localization results amongst published methods and it is the first one to report results for end-to-end text recognition. On the more challenging Street View Text dataset, the method achieves state-of-the-art recall. The robustness of the proposed method against noise and low contrast of characters is demonstrated by “false positives” caused by detected watermark text in the dataset. Lukás Neumann, Jiri Matas |
CVPR | 2 |
| 2012 | Homography estimation from correspondences of local elliptical features
Ondrej Chum, Jiri Matas |
ICPR | 2 |
| 2012 | Detection of bubbles as Concentric Circular Arrangements
Nataliya Strokina, Jiri Matas, Tuomas Eerola, Lasse Lensu, Heikki Kälviäinen |
ICPR | 2 |
| 2012 | Ultra-fast tracking based on zero-shift points
Jan Dupac, Jiri Matas, Filip Naiser |
Image Vis. Comput. | 2 |
| 2012 | Tracking-Learning-DetectionabstractThis paper investigates long-term tracking of unknown objects in a video stream. The object is defined by its location and extent in a single frame. In every frame that follows, the task is to determine the object's location and extent or indicate that the object is not present. We propose a novel tracking framework (TLD) that explicitly decomposes the long-term tracking task into tracking, learning, and detection. The tracker follows the object from frame to frame. The detector localizes all appearances that have been observed so far and corrects the tracker if necessary. The learning estimates the detector's errors and updates it to avoid these errors in the future. We study how to identify the detector's errors and learn from them. We develop a novel learning method (P-N learning) which estimates the errors by a pair of "experts": (1) P-expert estimates missed detections, and (2) N-expert estimates false alarms. The learning process is modeled as a discrete dynamical system and the conditions under which the learning guarantees improvement are found. We describe our real-time implementation of the TLD framework and the P-N learning. We carry out an extensive quantitative evaluation which shows a significant improvement over state-of-the-art approaches. Zdenek Kalal, Krystian Mikolajczyk, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | State of the JournalabstractT year 2012 will mark the end of my term as Editor-in-Chief of the IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). I believe that we have made substantial progress on one of the core challenges that TPAMI faces, namely, the continued growth of machine learning. As I have mentioned, the IEEE does not have a journal whose focus is modern machine learning methods, such as SVMs. Yet many papers in this area are submitted to TPAMI. One factor is that machine learning falls within TPAMI’s scope statement, but perhaps a more important reason is the journal’s excellence in computer vision, an area where machine learning is having a substantial and increasing impact. There is no possibility for TPAMI to ignore this area and continue to thrive, so the journal out of necessity must rise to the challenge of becoming a leading publication in machine learning. My predecessor David Kriegman saw this development clearly, and responded by appointing Zoubin Gharamani, a famous machine learning expert, as an Associate Editor in Chief (AEIC). Zoubin served his complete 4-year term with distinction, and has now moved up to the TPAMI Advisory Board. Over the last few years the number of machine learning submissions has continued to grow substantially, and we have clearly needed additional help. Machine learning is an area where TPAMI faces some distinct challenges. TPAMI has not published a body of truly fundamental papers in machine learning that is comparable to our accomplishments in computer vision, biometrics, or other areas that are closer to the journal’s traditional strengths. As a result, many of the submissions we receive have fallen short of TPAMI’s high standards. This has posed a diffi cult problem because it is challenging to attract top-notch researchers in machine learning as reviewers or AEs when most of the papers they handle must be rejected, and many would, in all honesty, never be submitted to a major machine learning journal. My primary focus throughout my term as EIC has been to address this situation, and I am pleased to report signifi cant progress. As you know, Max Welling joined us as an AEIC. I am happy to announce that Neil Lawrence has also agreed to serve as an AEIC. Neil has served with distinction as an AE for TPAMI, is on the board of JMLR, and will be program chair for AISTATS. (I note with amusement that Max Welling also held these three roles, which suggests a simple automatic classifi er to detect TPAMI AEICs in machine learning!) Neil has published two books in machine learning, and is primarily interested in probabilistic models. With both Max and Neil on board as AEICs we now have suffi cient manpower to address our main challenges. We have raised the bar for machine learning papers to be sent out for review by rejecting papers early on that would have eventually been rejected anyway. Hopefully, the effects of this will be clearly felt by everyone involved in the reviewing process, and the authors of high-quality submissions will benefi t from the increased availability of reviewing resources. Coupled with this effort, Max and Neil are developing a number of high quality special issues on important topics in machine learning (see the call for papers on page 207 of this issue for the fi rst such initiative). The fi eld of machine learning has a major advantage in its commitment to Open Access, which is an issue that the IEEE (along with most publishers) is struggling with. The top journal in machine learning (JMLR) is Open Access, while perhaps the best conference (NIPS) is making its proceedings available in arXiv. This has enormous benefi ts to the machine learning community. I personally believe that TPAMI will, over time, end up moving to an Open Access model, and I will hazard a guess that this will be one of the main challenges that the next EIC will face. On the operational side, the reviewing process on the whole is fairly timely, although exceptions do occur for a variety of reasons, and I want to yet again apologize to the authors whose papers get stalled in the process for one reason or another. To provide some numbers, there were 999 submissions in 2010 (I must confess I was really hoping for one more to come in at the very end). We are on track for a similar number in 2011, with 795 received as I write. The acceptance rate for 2010 submissions so far is 14 percent, though it is important to realize this does not imply 86 percent have been rejected, since a number of such papers are still undergoing revisions. The typical time from submission to final decision is about six months, which is unchanged from last year. Approximately 30 percent of submissions are rejected without review; while this is unpleasant for the authors, it saves them time from having their paper rejected at the end of the full review process and lets them quickly revise their papers for submission to a more appropriate journal. I am happy to report that the issue with the print queue is now under control, and papers now typically appear in print approximately 5.5 months after the fi nal material is uploaded. Short papers are generally published even faster, and authors are urged to consider this option. Of course, papers continue to be published online quite quickly after acceptance. On the topic of online publication, TPAMI is now available in the IEEE Computer Society’s new OnlinePlus format, at a signifi cant discount to the print subscription price. Over time the number of subscribers to the printed journal is falling, and readers who wish to see this format continue should be sure to sign up for print subscriptions. Sing Bing Kang, Jiri Matas, Max Welling, Ramin Zabih |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Editor's Note
Ramin Zabih, Sing Bing Kang, Neil D. Lawrence, Jiri Matas, Max Welling |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Editor's Note
Ramin Zabih, Sing Bing Kang, Neil D. Lawrence, Jiri Matas, Max Welling |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Rotation-Invariant Image and Video Description With Local Binary Pattern FeaturesabstractIn this paper, we propose a novel approach to compute rotation-invariant features from histograms of local noninvariant patterns. We apply this approach to both static and dynamic local binary pattern (LBP) descriptors. For static-texture description, we present LBP histogram Fourier (LBP-HF) features, and for dynamic-texture recognition, we present two rotation-invariant descriptors computed from the LBPs from three orthogonal planes (LBP-TOP) features in the spatiotemporal domain. LBP-HF is a novel rotation-invariant image descriptor computed from discrete Fourier transforms of LBP histograms. The approach can be also generalized to embed any uniform features into this framework, and combining the supplementary information, e.g., sign and magnitude components of the LBP, together can improve the description ability. Moreover, two variants of rotation-invariant descriptors are proposed to the LBP-TOP, which is an effective descriptor for dynamic-texture recognition, as shown by its recent success in different application problems, but it is not rotation invariant. In the experiments, it is shown that the LBP-HF and its extensions outperform noninvariant and earlier versions of the rotation-invariant LBP in the rotation-invariant texture classification. In experiments on two dynamic-texture databases with rotations or view variations, the proposed video features can effectively deal with rotation variations of dynamic textures (DTs). They also are robust with respect to changes in viewpoint, outperforming recent methods proposed for view-invariant recognition of DTs. Guoying Zhao 0001, Timo Ahonen, Jiri Matas, Matti Pietikäinen |
IEEE Trans. Image Process. | 3 |
| 2011 | Total recall II: Query expansion revisitedabstractMost effective particular object and image retrieval approaches are based on the bag-of-words (BoW) model. All state-of-the-art retrieval results have been achieved by methods that include a query expansion that brings a significant boost in performance. We introduce three extensions to automatic query expansion: (i) a method capable of preventing tf-idf failure caused by the presence of sets of correlated features (confusers), (ii) an improved spatial verification and re-ranking step that incrementally builds a statistical model of the query object and (iii) we learn relevant spatial context to boost retrieval performance. The three improvements of query expansion were evaluated on standard Paris and Oxford datasets according to a standard protocol, and state-of-the-art results were achieved. Ondrej Chum, Andrej Mikulík, Michal Perdoch, Jiri Matas |
CVPR | 4 |
| 2011 | Ultra-fast tracking based on zero-shift pointsabstractA novel tracker based on points where the intensity function is locally even is presented. Tracking of these so called zero-shift points (ZSPs) is very efficient, a single point is tracked on average in less than 10 microseconds on a standard notebook. We demonstrate experimentally the robustness of the tracker to image transformations and a relatively long lifetime of ZSPs in real videosequences. Jan Dupac, Jiri Matas |
ICASSP | 2 |
| 2011 | Text Localization in Real-World Images Using Efficiently Pruned Exhaustive SearchabstractAn efficient method for text localization and recognition in real-world images is proposed. Thanks to effective pruning, it is able to exhaustively search the space of all character sequences in real time (200ms on a 640x480 image). The method exploits higher-order properties of text such as word text lines. We demonstrate that the grouping stage plays a key role in the text localization performance and that a robust and precise grouping stage is able to compensate errors of the character detector. The method includes a novel selector of Maximally Stable Extremal Regions (MSER) which exploits region topology. Experimental validation shows that 95.7% characters in the ICDAR dataset are detected using the novel selector of MSERs with a low sensitivity threshold. The proposed method was evaluated on the standard ICDAR 2003 dataset where it achieved state-of-the-art results in both text localization and recognition. Lukás Neumann, Jiri Matas |
ICDAR | 2 |
| 2011 | Linear Regression and Adaptive Appearance Models for Fast Simultaneous Modelling and Tracking
Liam F. Ellis, Nicholas D. H. Dowson, Jiri Matas, Richard Bowden |
Int. J. Comput. Vis. | 3 |
| 2011 | Learning Linear Discriminant Projections for Dimensionality Reduction of Image DescriptorsabstractIn this paper, we present Linear Discriminant Projections (LDP) for reducing dimensionality and improving discriminability of local image descriptors. We place LDP into the context of state-of-the-art discriminant projections and analyze its properties. LDP requires a large set of training data with point-to-point correspondence ground truth. We demonstrate that training data produced by a simulation of image transformations leads to nearly the same results as the real data with correspondence ground truth. This makes it possible to apply LDP as well as other discriminant projection approaches to the problems where the correspondence ground truth is not available, such as image categorization. We perform an extensive experimental evaluation on standard data sets in the context of image matching and categorization. We demonstrate that LDP enables significant dimensionality reduction of local descriptors and performance increases in different applications. The results improve upon the state-of-the-art recognition performance with simultaneous dimensionality reduction from 128 to 30. Hongping Cai, Krystian Mikolajczyk, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Editor's Note
Ramin Zabih, Zoubin Ghahramani, Sing Bing Kang, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Editorial
Ramin Zabih, Zoubin Ghahramani, Sing Bing Kang, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Editor's Note
Ramin Zabih, Sing Bing Kang, Jiri Matas, Max Welling |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Editor's Note
Ramin Zabih, Sing Bing Kang, Jiri Matas, Max Welling |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | State of the Journal
Ramin Zabih, Jiri Matas, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Detection and matching of curvilinear structures
Cédric Lemaitre, Michal Perdoch, Adel Rahmoune, Jiri Matas, Johel Mitéran |
Pattern Recognit. | 4 |
| 2010 | Planar Affine Rectification from Change of Scale
Ondrej Chum, Jiri Matas |
ACCV (4) | 2 |
| 2010 | A Method for Text Localization and Recognition in Real-World Images
Lukás Neumann, Jiri Matas |
ACCV (3) | 2 |
| 2010 | Unsupervised discovery of co-occurrence in sparse high dimensional dataabstractAn efficient min-Hash based algorithm for discovery of dependencies in sparse high-dimensional data is presented. The dependencies are represented by sets of features co-occurring with high probability and are called co-ocsets. Sparse high dimensional descriptors, such as bag of words, have been proven very effective in the domain of image retrieval. To maintain high efficiency even for very large data collection, features are assumed independent. We show experimentally that co-ocsets are not rare, i.e. the independence assumption is often violated, and that they may ruin retrieval performance if present in the query image. Two methods for managing co-ocsets in such cases are proposed. Both methods significantly outperform the state-of-the-art in image retrieval, one is also significantly faster. Ondrej Chum, Jiri Matas |
CVPR | 2 |
| 2010 | Tracking the invisible: Learning where the object might beabstractObjects are usually embedded into context. Visual context has been successfully used in object detection tasks, however, it is often ignored in object tracking. We propose a method to learn supporters which are, be it only temporally, useful for determining the position of the object of interest. Our approach exploits the General Hough Transform strategy. It couples the supporters with the target and naturally distinguishes between strongly and weakly coupled motions. By this, the position of an object can be estimated even when it is not seen directly (e.g., fully occluded or outside of the image region) or when it changes its appearance quickly and significantly. Experiments show substantial improvements in model-free tracking as well as in the tracking of “virtual” points, e.g., in medical applications. Helmut Grabner, Jiri Matas, Luc Van Gool, Philippe C. Cattin |
CVPR | 2 |
| 2010 | P-N learning: Bootstrapping binary classifiers by structural constraintsabstractThis paper shows that the performance of a binary classifier can be significantly improved by the processing of structured unlabeled data, i.e. data are structured if knowing the label of one example restricts the labeling of the others. We propose a novel paradigm for training a binary classifier from labeled and unlabeled examples that we call P-N learning. The learning process is guided by positive (P) and negative (N) constraints which restrict the labeling of the unlabeled set. P-N learning evaluates the classifier on the unlabeled data, identifies examples that have been classified in contradiction with structural constraints and augments the training set with the corrected samples in an iterative process. We propose a theory that formulates the conditions under which P-N learning guarantees improvement of the initial classifier and validate it on synthetic and real data. P-N learning is applied to the problem of on-line learning of object detector during tracking. We show that an accurate object detector can be learned from a single example and an unlabeled video sequence where the object may occur. The algorithm is compared with related approaches and state-of-the-art is achieved on a variety of objects (faces, pedestrians, cars, motorbikes and animals). Zdenek Kalal, Jiri Matas, Krystian Mikolajczyk |
CVPR | 2 |
| 2010 | Learning a Fine Vocabulary
Andrej Mikulík, Michal Perdoch, Ondrej Chum, Jiri Matas |
ECCV (3) | 4 |
| 2010 | Face-TLD: Tracking-Learning-Detection applied to facesabstractA novel system for long-term tracking of a human face in unconstrained videos is built on Tracking-Learning-Detection (TLD) approach. The system extends TLD with the concept of a generic detector and a validator which is designed for real-time face tracking resistent to occlusions and appearance changes. The off-line trained detector localizes frontal faces and the online trained validator decides which faces correspond to the tracked subject. Several strategies for building the validator during tracking are quantitatively evaluated. The system is validated on a sitcom episode (23 min.) and a surveillance (8 min.) video. In both cases the system detects-tracks the face and automatically learns a multi-view model from a single frontal example and an unlabeled video. Zdenek Kalal, Krystian Mikolajczyk, Jiri Matas |
ICIP | 3 |
| 2010 | Image Matching and Retrieval by Repetitive PatternsabstractDetection of repetitive patterns in images has been studied for a long time in computer vision. This paper discusses a method for representing a lattice or line pattern by shift-invariant descriptor of the repeating element. The descriptor overcomes shift ambiguity and can be matched between different a views. The pattern matching is then demonstrated in retrieval experiment, where different images of the same buildings are retrieved solely by repetitive patterns. Petr Doubek, Jiri Matas, Michal Perdoch, Ondrej Chum |
ICPR | 2 |
| 2010 | Forward-Backward Error: Automatic Detection of Tracking FailuresabstractThis paper proposes a novel method for tracking failure detection. The detection is based on the Forward-Backward error, i.e. the tracking is performed forward and backward in time and the discrepancies between these two trajectories are measured. We demonstrate that the proposed error enables reliable detection of tracking failures and selection of reliable trajectories in video sequences. We demonstrate that the approach is complementary to commonly used normalized cross-correlation (NCC). Based on the error, we propose a novel object tracker called Median Flow. State-of-the-art performance is achieved on challenging benchmark video sequences which include non-rigid objects. Zdenek Kalal, Krystian Mikolajczyk, Jiri Matas |
ICPR | 3 |
| 2010 | Construction of Precise Local Affine FramesabstractWe propose a novel method for the refinement of Maximally Stable Extremal Region (MSER) boundaries to sub-pixel precision by taking into account the intensity function in the 2 × 2 neighborhood of the contour points. The proposed method improves the repeatability and precision of Local Affine Frames (LAFs) constructed on extremal regions. Additionally, we propose a novel method for detection of local curvature extrema on the refined contour. Experimental evaluation on publicly available datasets shows that matching with the modified LAFs leads to a higher number of correspondences and a higher inlier ratio in more than 80% of the test image pairs. Since the processing time of the contour refinement is negligible, there is no reason not to include the algorithms as a standard part of the MSER detector and LAF constructions. Andrej Mikulík, Jiri Matas, Michal Perdoch, Ondrej Chum |
ICPR | 2 |
| 2010 | A voting strategy for visual ego-motion from stereoabstractWe present a procedure for egomotion estimation from visual input of a stereo pair of video cameras. The 3D egomotion problem, which has six degrees of freedom in general, is simplified to four dimensions and further decomposed to two two-dimensional subproblems. The decomposition allows us to use a voting strategy to identify the most probable solution, avoiding the random sampling (RANSAC) or other approximation techniques. The input constitutes of image correspondences between consecutive stereo pairs, i.e. feature points do not need to be tracked over time. The experiments show that even if a trajectory is put together as a simple concatenation of frame-to-frame increments, it comes out reliable and precise. Stepán Obdrzálek, Jiri Matas |
Intelligent Vehicles Symposium | 2 |
| 2010 | Efficient Sequential Correspondence Selection by CosegmentationabstractIn many retrieval, object recognition, and wide-baseline stereo methods, correspondences of interest points (distinguished regions) are commonly established by matching compact descriptors such as SIFTs. We show that a subsequent cosegmentation process coupled with a quasi-optimal sequential decision process leads to a correspondence verification procedure that 1) has high precision (is highly discriminative), 2) has good recall, and 3) is fast. The sequential decision on the correctness of a correspondence is based on simple statistics of a modified dense stereo matching algorithm. The statistics are projected on a prominent discriminative direction by SVM. Wald's sequential probability ratio test is performed on the SVM projection computed on progressively larger cosegmented regions. We show experimentally that the proposed sequential correspondence verification (SCV) algorithm significantly outperforms the standard correspondence selection method based on SIFT distance ratios on challenging matching problems. Jan Cech, Jiri Matas, Michal Perdoch |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Large-Scale Discovery of Spatially Related ImagesabstractWe propose a randomized data mining method that finds clusters of spatially overlapping images. The core of the method relies on the min-Hash algorithm for fast detection of pairs of images with spatial overlap, the so-called cluster seeds. The seeds are then used as visual queries to obtain clusters which are formed as transitive closures of sets of partially overlapping images that include the seed. We show that the probability of finding a seed for an image cluster rapidly increases with the size of the cluster. The properties and performance of the algorithm are demonstrated on data sets with 10(4), 10(5), and 5 x 10(6) images. The speed of the method depends on the size of the database and the number of clusters. The first stage of seed generation is close to linear for databases sizes up to approximately 2(34) approximately 10(10) images. On a single 2.4 GHz PC, the clustering process took only 24 minutes for a standard database of more than 100,000 images, i.e., only 0.014 seconds per image. Ondrej Chum, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Editor's Note
Ramin Zabih, Jiri Matas, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Editor's Note
Ramin Zabih, Jiri Matas, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Editor's Note
Ramin Zabih, Jiri Matas, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Geometric min-Hashing: Finding a (thick) needle in a haystackabstractWe propose a novel hashing scheme for image retrieval, clustering and automatic object discovery. Unlike commonly used bag-of-words approaches, the spatial extent of image features is exploited in our method. The geometric information is used both to construct repeatable hash keys and to increase the discriminability of the description. Each hash key combines visual appearance (visual words) with semi-local geometric information. Compared with the state-of-the-art min-hash, the proposed method has both higher recall (probability of collision for hashes on the same object) and lower false positive rates (random collisions). The advantages of geometric min-hashing approach are most pronounced in the presence of viewpoint and scale change, significant occlusion or small physical overlap of the viewing fields. We demonstrate the power of the proposed method on small object discovery in a large unordered collection of images and on a large scale image clustering problem. Ondrej Chum, Michal Perdoch, Jiri Matas |
CVPR | 3 |
| 2009 | Efficient representation of local geometry for large scale object retrievalabstractState of the art methods for image and object retrieval exploit both appearance (via visual words) and local geometry (spatial extent, relative pose). In large scale problems, memory becomes a limiting factor - local geometry is stored for each feature detected in each image and requires storage larger than the inverted file and term frequency and inverted document frequency weights together. We propose a novel method for learning discretized local geometry representation based on minimization of average reprojection error in the space of ellipses. The representation requires only 24 bits per feature without drop in performance. Additionally, we show that if the gravity vector assumption is used consistently from the feature description to spatial verification, it improves retrieval performance and decreases the memory footprint. The proposed method outperforms state of the art retrieval algorithms in a standard image retrieval benchmark. Michal Perdoch, Ondrej Chum, Jiri Matas |
CVPR | 3 |
| 2009 | Integrated vision system for the semantic interpretation of activities where a person handles objects
Markus Vincze, Michael Zillich, Wolfgang Ponweiser, Václav Hlavác, Jiri Matas, Stepán Obdrzálek, Hilary Buxton, A. Jonathan Howell, Kingsley Sage, Antonis A. Argyros, Christof Eberst, Gerald Umgeher |
Comput. Vis. Image Underst. | 5 |
| 2009 | Learning Fast Emulators of Binary Decision Processes
Jan Sochman, Jiri Matas |
Int. J. Comput. Vis. | 2 |
| 2009 | Anytime learning for the NoSLLiP tracker
Karel Zimmermann, Tomás Svoboda, Jiri Matas |
Image Vis. Comput. | 3 |
| 2009 | Introduction of New Associate Editors
Ramin Zabih, Zoubin Ghahramani, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Introduction of New Associate Editors
Ramin Zabih, Jiri Matas, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Tracking by an Optimal Sequence of Linear PredictorsabstractWe propose a learning approach to tracking explicitly minimizing the computational complexity of the tracking process subject to user-defined probability of failure (loss-of-lock) and precision. The tracker is formed by a Number of Sequences of Learned Linear Predictors (NoSLLiP). Robustness of NoSLLiP is achieved by modeling the object as a collection of local motion predictors--object motion is estimated by the outlier-tolerant RANSAC algorithm from local predictions. Efficiency of the NoSLLiP tracker stems from (i) the simplicity of the local predictors and (ii) from the fact that all design decisions--the number of local predictors used by the tracker, their computational complexity (i.e. the number of observations the prediction is based on), locations as well as the number of RANSAC iterations are all subject to the optimization (learning) process. All time-consuming operations are performed during the learning stage--tracking is reduced to only a few hundreds integer multiplications in each step. On PC with 1xK8 3200+, a predictor evaluation requires about 30 microseconds. The proposed approach is verified on publicly-available sequences with approximately 12000 frames with ground-truth. Experiments demonstrates, superiority in frame rates and robustness with respect to the SIFT detector, Lucas-Kanade tracker and other trackers. Karel Zimmermann, Jiri Matas, Tomás Svoboda |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Learning Linear Discriminant Projections for Dimensionality Reduction of Image DescriptorsabstractIn this paper we present Linear Discriminant Projections (LDP) for reducing dimensionality and improving discriminability of local image descriptors. We place LDP into the context of state-of-the-art discriminant projections and analyze its properties. LDP requires large set of training data with point-to-point correspondence ground truth. We demonstrate that a training data produced by a simulation of image transformations leads to nearly the same results as the real data with correspondence ground truth. This makes it possible to apply LDP as well as other discriminant projection approaches to the problems where the correspondence ground truth is not available such as image categorization. We perform an extensive experimental evaluation on standard datasets in the context of image matching and categorization. We demonstrate that LDP enables significant dimensionality reduction of local descriptors and performance increases in different applications. The results improve upon the state-of-the-art recognition performance with simultaneous dimensionality reduction from 128 to 30. Hongping Cai, Krystian Mikolajczyk, Jiri Matas |
BMVC | 3 |
| 2008 | Online Learning and Partitioning of Linear Displacement Predictors for TrackingabstractA novel approach to learning and tracking arbitrary image features is presented. Tracking is tackled by learning the mapping from image inten-sity differences to displacements. Linear regression is used, resulting in low computational cost. An appearance model of the target is built on-the-fly by clustering sub-sampled image templates. The medoidshift algorithm is used to cluster the templates thus identifying various modes or aspects of the target appearance, each mode is associated to the most suitable set of linear predic-tors allowing piecewise linear regression from image intensity differences to warp updates. Despite no hard-coding or offline learning, excellent results are shown on three publicly available video sequences and comparisons with related approaches made. 1 Liam F. Ellis, Jiri Matas, Richard Bowden |
BMVC | 2 |
| 2008 | Weighted Sampling for Large-Scale BoostingabstractThis paper addresses the problem of learning from very large databases where batch learning is impractical or even infeasible. Bootstrap is a popular technique applicable in such situations. We show that sampling strategy used for bootstrapping has a significant impact on the resulting classifier performance. We design a new general sampling strategy ”quasi-random weighted sampling + trimming ” (QWS+) that includes well established strategies as special cases. The QWS+ approach minimizes the variance of hypothesis error estimate and leads to significant improvement in performance compared to standard sampling techniques. The superior performance is demonstrated on several problems including profile and frontal face detection. 1 Zdenek Kalal, Jiri Matas, Krystian Mikolajczyk |
BMVC | 2 |
| 2008 | Efficient sequential correspondence selection by cosegmentationabstractIn many retrieval, object recognition and wide baseline stereo methods, correspondences of interest points are established possibly sublinearly by matching a compact descriptor such as SIFT. We show that a subsequent cosegmentation process coupled with a quasi-optimal sequential decision process leads to a correspondence verification procedure that has (i) high precision (is highly discriminative) (ii) good recall and (iii) is fast. The sequential decision on the correctness of a correspondence is based on trivial attributes of a modified dense stereo matching algorithm. The attributes are projected on a prominent discriminative direction by SVM. Waldpsilas sequential probability ratio test is performed for SVM projection computed on progressively larger co-segmented regions. Experimentally we show that the process significantly outperforms the standard correspondence selection process based on SIFT distance ratios on challenging matching problems. Jan Cech, Jiri Matas, Michal Perdoch |
CVPR | 2 |
| 2008 | Training sequential on-line boosting classifier for visual trackingabstractOn-line boosting allows to adapt a trained classifier to changing environmental conditions or to use sequentially available training data. Yet, two important problems in the on-line boosting training remain unsolved: (i) classifier evaluation speed optimization and, (ii) automatic classifier complexity estimation. In this paper we show how the on-line boosting can be combined with Waldpsilas sequential decision theory to solve both of the problems. The properties of the proposed on-line WaldBoost algorithm are demonstrated on a visual tracking problem. The complexity of the classifier is changing dynamically depending on the difficulty of the problem. On average, a speedup of a factor of 5-10 is achieved compared to the non-sequential on-line boosting. Helmut Grabner, Jan Sochman, Horst Bischof, Jiri Matas |
ICPR | 4 |
| 2008 | Guest Editors' Introduction to the Special Section on CVPR PapersabstractThe four papers in this special section are extended versions of award-winning papers from the 2007 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2007). Simon Baker, Jiri Matas, Ramin Zabih |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Optimal Randomized RANSACabstractA randomized model verification strategy for RANSAC is presented. The proposed method finds, like RANSAC, a solution that is optimal with user-specified probability. The solution is found in time that is (i) close to the shortest possible and (ii) superior to any deterministic verification strategy. A provably fastest model verification strategy is designed for the (theoretical) situation when the contamination of data by outliers is known. In this case, the algorithm is the fastest possible (on average) of all randomized \\RANSAC algorithms guaranteeing a confidence in the solution. The derivation of the optimality property is based on Wald's theory of sequential decision making, in particular a modified sequential probability ratio test (SPRT). Next, the R-RANSAC with SPRT algorithm is introduced. The algorithm removes the requirement for a priori knowledge of the fraction of outliers and estimates the quantity online. We show experimentally that on standard test data the method has performance close to the theoretically optimal and is 2 to 10 times faster than standard RANSAC and is up to 4 times faster than previously published methods. Ondrej Chum, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Learning a Fast Emulator of a Binary Decision Process
Jan Sochman, Jiri Matas |
ACCV (2) | 2 |
| 2007 | Definition of a Model-Based Detector of Curvilinear Regions
Cédric Lemaitre, Johel Mitéran, Jiri Matas |
CAIP | 3 |
| 2007 | Linear Predictors for Fast Simultaneous Modeling and TrackingabstractAn approach for fast tracking of arbitrary image features with no prior model and no offline learning stage is presented. Fast tracking is achieved using banks of linear displacement predictors learnt online. A multi-modal appearance model is also learnt on-the-fly that facilitates the selection of subsets of predictors suitable for prediction in the next frame. The approach is demonstrated in real-time on a number of challenging video sequences and experimentally compared to other simultaneous modeling and tracking approaches with favourable results. Liam F. Ellis, Nicholas D. H. Dowson, Jiri Matas, Richard Bowden |
ICCV | 3 |
| 2007 | Improving Descriptors for Fast Tree Matching by Optimal Linear ProjectionabstractIn this paper we propose to transform an image descriptor so that nearest neighbor (NN) search for correspondences becomes the optimal matching strategy under the assumption that inter-image deviations of corresponding descriptors have Gaussian distribution. The Euclidean NN in the transformed domain corresponds to the NN according to a truncated Mahalanobis metric in the original descriptor space. We provide theoretical justification for the proposed approach and show experimentally that the transformation allows a significant dimensionality reduction and improves matching performance of a state-of-the art SIFT descriptor. We observe consistent improvement in precision-recall and speed of fast matching in tree structures at the expense of little overhead for projecting the descriptors into transformed space. In the context of SIFT vs. transformed M- SIFT comparison, tree search structures are evaluated according to different criteria and query types. All search tree experiments confirm that transformed M-SIFTperforms better than the original SIFT. Krystian Mikolajczyk, Jiri Matas |
ICCV | 2 |
| 2007 | Stable Affine Frames on IsophotesabstractWe propose a new affine-covariant feature, the stable affine frame (SAF). SAFs lie on the boundary of extremal regions, i.e. on isophotes. Instead of requiring the whole isophote to be stable with respect to intensity perturbation as in maximally stable extremal regions (MSERs), stability is required only locally, for the primitives constituting the three-point frames. The primitives are extracted by an affine invariant process that exploits properties of bitangents and algebraic moments. Thus, instead of using closed stable isophotes, i.e. MSERs, and detecting affine frames on them, SAFs are sought even on some unstable extremal regions. We show experimentally on standard datasets that SAFs have repeatability comparable to the best affine covariant detectors tested in the state-of-the-art report (Mikolajczyk et al., 2005) and consistently produce a significantly higher number of features per image. Moreover, the features cover images more evenly than MSERs, which facilitates robustness to occlusion. Michal Perdoch, Jiri Matas, Stepán Obdrzálek |
ICCV | 2 |
| 2007 | Adaptive Parameter Optimization for Real-time TrackingabstractAdaptation of a tracking procedure combined in a common way with a Kalman filter is formulated as an constrained optimization problem, where a trade-off between precision and loss-of-lock probability is explicitly taken into account. While the tracker is learned in order to minimize computational complexity during a learning stage, in a tracking stage the precision is maximized online under a constraint imposed by the loss-of-lock probability resulting in an optimal setting of the tracking procedure. We experimentally show that the proposed method converges to a steady solution in all variables. In contrast to a common Kalman filter based tracking, we achieve a significantly lower state covariance matrix. We also show, that if the covariance matrix is continuously updated, the method is able to adapt to a different situations. If a dynamic model is precise enough the tracker is allowed to spend a longer time with a fine motion estimation, however, if the motion gets saccadic, i.e. unpredictable by the dynamic model, the method automatically gives up the precision in order to avoid loss-of-lock. Karel Zimmermann, Tomás Svoboda, Jiri Matas |
ICCV | 3 |
| 2006 | Geometric Hashing with Local Affine FramesabstractWe propose a novel representation of local image structure and a matching scheme that are insensitive to a wide range of appearance changes. The representation is a collection of local affine frames that are constructed on outer boundaries of maximally stable extremal regions (MSERS) in an affine-covariant way. Each local affine frame is described by a relative location of other local affine frames in its neighborhood. The image is thus represented by quantities that depend only on the location of the boundaries of MSERs. Inter-image correspondences between local affine frames are formed in constant time by geometric hashing. Direct detection of local afine frames removes the requirement of a point-based hashing to establish reference frames in a combinatorial way, which has in the case of affine transform complexily that is cubic in the number of points. Local affine frames, which are also the quantities represented in the hash table, occupy a 6 0 space and hence data collisions are less likely compared with 2 0 point hashing. Experimentally, the robustness of the method and its insensitiviq to photometric changes is demonstrated on images from different spectral bands of satellite sensor; on images of a transparent object and on images of an object taken during day and night. Ondrej Chum, Jiri Matas |
CVPR (1) | 2 |
| 2006 | Editorial: Selection of Papers for the ECCV 2004 Special Issue
Tomás Pajdla, Jiri Matas |
Int. J. Comput. Vis. | 2 |
| 2005 | Sub-linear Indexing for Large Scale Object RecognitionabstractRealistic approaches to large scale object recognition, i.e. for detection and localisation of hundreds or more objects, must support sub-linear time indexing.In the paper, we propose a method capable of recognising one of N objects in log(N) time.The "visual memory" is organised as a binary decision tree that is built to minimise average time to decision.Leaves of the tree represent a few local image areas, and each non-terminal node is associated with a 'weak classifier'.In the recognition phase, a single invariant measurement decides in which subtree a corresponding image area is sought.The method preserves all the strengths of local affine region methodsrobustness to background clutter, occlusion, and large changes of viewpoints.Experimentally we show that it supports near real-time recognition of hundreds of objects with state-of-the-art recognition rates.After the test image is processed (in a second on a current PCs), the recognition via indexing into the visual memory requires milliseconds. Stepán Obdrzálek, Jiri Matas |
BMVC | 2 |
| 2005 | Matching with PROSAC - Progressive Sample ConsensusabstractA new robust matching method is proposed. The progressive sample consensus (PROSAC) algorithm exploits the linear ordering defined on the set of correspondences by a similarity function used in establishing tentative correspondences. Unlike RANSAC, which treats all correspondences equally and draws random samples uniformly from the full set, PROSAC samples are drawn from progressively larger sets of top-ranked correspondences. Under the mild assumption that the similarity measure predicts correctness of a match better than random guessing, we show that PROSAC achieves large computational savings. Experiments demonstrate it is often significantly faster (up to more than hundred times) than RANSAC. For the derived size of the sampled set of correspondences as a function of the number of samples already drawn, PROSAC converges towards RANSAC in the worst case. The power of the method is demonstrated on wide-baseline matching problems. Ondrej Chum, Jiri Matas |
CVPR (1) | 2 |
| 2005 | Two-View Geometry Estimation Unaffected by a Dominant PlaneabstractA RANSAC-based algorithm for robust estimation of epipolar geometry from point correspondences in the possible presence of a dominant scene plane is presented. The algorithm handles scenes with (i) all points in a single plane, (ii) majority of points in a single plane and the rest off the plane, (iii) no dominant plane. It is not required to know a priori which of the cases (i)-(iii) occurs. The algorithm exploits a theorem we proved, that if five or more of seven correspondences are related by a homography then there is an epipolar geometry consistent with the seven-tuple as well as with all correspondences related by the homography. This means that a seven point sample consisting of two outliers and five inliers lying in a dominant plane produces an epipolar geometry which is wrong and yet consistent with a high number of correspondences. The theorem explains why RANSAC often fails to estimate epipolar geometry in the presence of a dominant plane. Rather surprisingly, the theorem also implies that RANSAC-based homography estimation is faster when drawing nonminimal samples of seven correspondences than minimal samples of four correspondences. Ondrej Chum, Tomás Werner, Jiri Matas |
CVPR (1) | 3 |
| 2005 | WaldBoost - Learning for Time Constrained Sequential DetectionabstractIn many computer vision classification problems, both the error and time characterizes the quality of a decision. We show that such problems can be formalized in the framework of sequential decision-making. If the false positive and false negative error rates are given, the optimal strategy in terms of the shortest average time to decision (number of measurements used) is the Wald's sequential probability ratio test (SPRT). We built on the optimal SPRT test and enlarge its capabilities to problems with dependent measurements. We show how to overcome the requirements of SPRT - (i) a priori ordered measurements and (ii) known joint probability density functions. We propose an algorithm with near optimal time and error rate trade-off, called WaldBoost, which integrates the AdaBoost algorithm for measurement selection and ordering and the joint probability density estimation with the optimal SPRT decision strategy. The WaldBoost algorithm is tested on the face detection problem. The results are superior to the state-of-the-art methods in the average evaluation time and comparable in detection rates. Jan Sochman, Jiri Matas |
CVPR (2) | 2 |
| 2005 | Randomized RANSAC with Sequential Probability Ratio TestabstractA randomized model verification strategy for RANSAC is presented. The proposed method finds, like RANSAC, a solution that is optimal with user-controllable probability n. A provably optimal model verification strategy is designed for the situation when the contamination of data by outliers is known, i.e. the algorithm is the fastest possible (on average) of all randomized RANSAC algorithms guaranteeing 1 - n confidence in the solution. The derivation of the optimality property is based on Wald's theory of sequential decision making. The R-RANSAC with SPRT which does not require the a priori knowledge of the fraction of outliers and has results close to the optimal strategy is introduced. We show experimentally that on standard test data the method is 2 to 10 times faster than the standard RANSAC and up to 4 times faster than previously published methods. Jiri Matas, Ondrej Chum |
ICCV | 1 |
| 2005 | A Comparison of Affine Region Detectors
Krystian Mikolajczyk, Tinne Tuytelaars, Cordelia Schmid, Andrew Zisserman, Jiri Matas, Frederik Schaffalitzky, Timor Kadir, Luc Van Gool |
Int. J. Comput. Vis. | 5 |
| 2005 | Feature-Based Affine-Invariant Localization of FacesabstractWe present a novel method for localizing faces in person identification scenarios. Such scenarios involve high resolution images of frontal faces. The proposed algorithm does not require color, copes well in cluttered backgrounds, and accurately localizes faces including eye centers. An extensive analysis and a performance evaluation on the XM2VTS database and on the realistic BioID and BANCA face databases is presented. We show that the algorithm has precision superior to reference methods. Miroslav Hamouz, Josef Kittler, Joni-Kristian Kämäräinen, Pekka Paalanen, Heikki Kälviäinen, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2004 | Randomized RANSAC with Td, d test
Jiri Matas, Ondrej Chum |
Image Vis. Comput. | 1 |
| 2004 | Robust wide-baseline stereo from maximally stable extremal regions
Jiri Matas, Ondrej Chum, Tomás Pajdla |
Image Vis. Comput. | 1 |
| 2003 | Face verification via error correcting output codes
Josef Kittler, Reza Ghaderi, Terry Windeatt, Jiri Matas |
Image Vis. Comput. | 4 |
| 2002 | Randomized RANSAC with T(d, d) testabstractMany computer vision algorithms include a robust estimation step where model parameters are computed from a data set containing a significant proportion of outliers. The RANSAC algorithm is possibly the most widely used robust estimator in the field of computer vision. In the paper we show that under a broad range of conditions, RANSAC efficiency is significantly improved if its hypothesis evaluation step is randomized. A new Jiri Matas, Ondrej Chum |
BMVC | 1 |
| 2002 | Robust Wide Baseline Stereo from Maximally Stable Extremal RegionsabstractAbstract The wide-baseline stereo problem, i.e. the problem of establishing correspondences between a pair of images taken from different viewpoints is studied. A new set of image elements that are put into correspondence, the so called extremal regions , is introduced. Extremal regions possess highly desirable properties: the set is closed under (1) continuous (and thus projective) transformation of image coordinates and (2) monotonic transformation of image intensities. An efficient (near linear complexity) and practically fast detection algorithm (near frame rate) is presented for an affinely invariant stable subset of extremal regions, the maximally stable extremal regions (MSER). A new robust similarity measure for establishing tentative correspondences is proposed. The robustness ensures that invariants from multiple measurement regions (regions obtained by invariant constructions from extremal regions), some that are significantly larger (and hence discriminative) than the MSERs, may be used to establish tentative correspondences. The high utility of MSERs, multiple measurement regions and the robust metric is demonstrated in wide-baseline experiments on image pairs from both indoor and outdoor scenes. Significant change of scale (3.5×), illumination conditions, out-of-plane rotation, occlusion, locally anisotropic scale change and 3D translation of the viewpoint are all present in the test problems. Good estimates of epipolar geometry (average distance from corresponding points to the epipolar line below 0.09 of the inter-pixel distance) are obtained. Jiri Matas, Ondrej Chum, Tomás Pajdla |
BMVC | 1 |
| 2002 | Object Recognition using Local Affine Frames on Distinguished RegionsabstractA novel approach to appearance based object recognition is introduced. The proposed method, based on matching of local image features, reliably recognises objects under very different viewing conditions. First, distinguished regions of data-dependent shape are robustly detected. On these regions, local affine frames are established using several affine invariant constructions. Direct comparison of photometrically normalised colour intensities in local, geometrically aligned frames results in a matching scheme that is invariant to piecewise-affine image deformations, but still remains very discriminative. The potential of the approach is experimentally verified on COIL-100 and SOIL-47 – publicly available image databases. On SOIL-47, 100 % recognition rate is achieved for single training view per object. On COIL-100, 99.9% recognition rate is obtained for 18 training views per object. Robustness to severe occlusions is demonstrated by only a moderate decrease of recognition performance in an experiment where half of each test image is erased. 1 Stepán Obdrzálek, Jiri Matas |
BMVC | 2 |
| 2002 | The Multimodal Neighborhood Signature for Modeling Object Color Appearance and Applications in Object Recognition and Image Retrieval
Jiri Matas, Dimitri Koubaroulis, Josef Kittler |
Comput. Vis. Image Underst. | 1 |
| 2002 | Support vector machines for face authentication
Kenneth Jonsson, Josef Kittler, Yongping Li, Jiri Matas |
Image Vis. Comput. | 4 |
| 2001 | Face Verification via ECOCabstractWe develop a novel approach to face verification based on the Error Correcting Output Coding (ECOC) classifier design concept. In the training phase the client set is repeatedly divided into two ECOC specified sub-sets (superclasses) to train a set of binary classifiers. The output of the classifiers defines the ECOC feature space, in which it is easier to separate transformed patterns representing clients and impostors. The proposed method exhibits superior verification performance on the well known XM2VTS data set as compared with previously reported results. 1 Josef Kittler, Reza Ghaderi, Terry Windeatt, Jiri Matas |
BMVC | 4 |
| 2001 | Face Verification Using Error Correcting Output CodesabstractThe error correcting output coding (ECOC) approach to classifier design decomposes a multi-class problem into a set of complementary two-class problems. We show how to apply the ECOC concept to automatic face verification, which is inherently a two-class problem. The output of the binary classifiers defines the ECOC feature space, in which it is easier to separate transformed patterns representing clients and impostors. We propose two different combining strategies as the matching score for face verification. The first uses the first order Minkowski metric, and requires a threshold to be set. The second is a kernel-based method and has no parameters to set. The proposed method exhibits better performance on the well known XM2VTS data set compared with previous reported results. Josef Kittler, Reza Ghaderi, Terry Windeatt, Jiri Matas |
CVPR (1) | 4 |
| 2001 | Empirical evaluation of a calibration chart detector
Soh Ling Min, Josef Kittler, Jiri Matas |
Mach. Vis. Appl. | 3 |
| 2000 | On Matching Scores for LDA-based Face VerificationabstractWe address the problem of face verification using linear discriminant anal-ysis and investigate the issue of matching score1. We establish the reason behind the success of the normalised correlation. The improved understand-ing about the role of metric then naturally leads to a novel way of measuring the distance between a probe image and a model. In extensive experimen-tal studies on the publicly available XM2VTS database2 using the Lausanne protocol3 we show that the proposed metric is consistently superior to both the Euclidean distance and normalised correlation matching scores. The ef-fect of various photometric normalisations4 on the matching scores is also investigated. 1 Josef Kittler, Yongping Li, Jiri Matas |
BMVC | 3 |
| 2000 | Object Recognition using the Invariant Pixel-Set SignatureabstractA new object recognition method, the Invariant Pixel Set Signature (IPSS), is introduced. Objects are represented with a probability density on the space of invariants computed from measurements (pixel values) inside convex hulls of n-tuples of interest points. Experimentally the method is tested on COIL– 20, a publicly available database of 72 views of 20 natural object rotating on a turntable. With a model built from a single view, recognition performance measured by the average match percentile is above 98 % for 20 degrees and above 96 % for30 degrees. For some object, 100 % first rank is achieved for all 72 views. Robustness to occlusion is shown using images with one half covered. For a small change of viewpoint (10 degrees) recognition of the occluded object is perfect. 1 Jiri Matas, J. Burianek, Josef Kittler |
BMVC | 1 |
| 2000 | Colour Image Retrieval and Object Recognition Using the Multimodal Neighbourhood Signature
Jiri Matas, Dimitri Koubaroulis, Josef Kittler |
ECCV (1) | 1 |
| 2000 | Learning Support Vectors for Face Verification and RecognitionabstractThe paper studies support vector machines (SVM) in the context of face verification and recognition. Our study supports the hypothesis that the SVM approach is able to extract the relevant discriminatory information from the training data and we present results showing superior performance in comparison with benchmark methods. However, when the representation space already captures and emphasises the discriminatory information (e.g., Fisher's linear discriminant), SVM loose their superiority. The results also indicate that the SVM are robust against changes in illumination provided these are adequately represented in the training data. The proposed system is evaluated on a large database of 295 people obtaining highly competitive results: an equal error rate of 1% for verification and a rank-one error rate of 2% for recognition (or 98% correct rank-one recognition). Kenneth Jonsson, Josef Kittler, Yongping Li, Jiri Matas |
FG | 4 |
| 2000 | Wearable face recognition aidabstractThe feasibility of realising a low cost wearable face recognition aid based on a robust correlation algorithm is investigated. The aim of the study is to determine the limiting spatial and grey level resolution of the probe and gallery images that would support successful prompting of the identity of input face images. Low spatial and grey level resolution images are obtained from good quality image data algorithmically. The tests carried out on the XM2VTS database demonstrate that robust correlation is very resilient to degradations of spatial and grey level image resolution. Correct prompts have been generated in 98% cases even for severely degraded images. C. Iordanoglou, Kenneth Jonsson, Josef Kittler, Jiri Matas |
ICASSP | 4 |
| 2000 | Using Gradient Information to Enhance the Progressive Probabilistic Hough TransformabstractWe look at the benefits to be gained in using gradient information to enhance the progressive probabilistic Hough transform (PPHT). It is shown how using the angle information in controlling the voting process and in assigning pixels correctly to a line, PPHT's performance can be significantly improved. The improved algorithm gives results very close to that of the standard Hough transform, but requires significantly less computation. Charles Galambos, Josef Kittler, Jiri Matas |
ICPR | 3 |
| 2000 | The Multimodal Signature Method: An Efficiency and Sensitivity StudyabstractThe multimodal neighbourhood signature (MNS) method has given acceptable results both for the colour-based image retrieval and the object recognition task. Local colour content is concisely represented by invariant features computed from neighbourhoods with multimodal colour density function. In this paper, efficiency related issues regarding the MNS algorithm are investigated. Its performance, speed, sensitivity to internal parameters and storage requirements are tested on a standard colour object recognition experiment. Very good recognition rate (99.9%) was achieved in real time. The MNS signature size is a few hundred bytes on average, an important property for retrieval from large databases. The algorithmic complexity of signature computation and matching are analysed and efficient implementations are proposed. Dimitri Koubaroulis, Jiri Matas, Josef Kittler |
ICPR | 2 |
| 2000 | Comparison of Face Verification Results on the XM2VTS DatabaseabstractPresents results of the face verification contest that was organized in conjunction with International Conference on Pattern Recognition 2000. Participants had to use identical data sets from a large, publicly available multimodal database XM2VTSDB. Training and evaluation was carried out according to an a priori known protocol. Verification results of all tested algorithms have been collected and made public on the XM2VTSDB website, facilitating large scale experiments on classifier combination and fusion. Tested methods included, among others, representatives of the most common approaches to face verification -elastic graph matching, Fisher's linear discriminant and support vector machines. Jiri Matas, Miroslav Hamouz, Kenneth Jonsson, Josef Kittler, Yongping Li, Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas, Teewoon Tan, Hong Yan 0001, Fabrizio Smeraldi, N. Capdevielle, Wulfram Gerstner, Yousri Abdeljaoued, Josef Bigün, Souheil Ben Yacoub, Eddy Mayoraz |
ICPR | 1 |
| 2000 | Robust Detection of Lines Using the Progressive Probabilistic Hough Transform
Jiri Matas, Charles Galambos, Josef Kittler |
Comput. Vis. Image Underst. | 1 |
| 1999 | Support Vector Machines for Face AuthenticationabstractAbstract We present an extensive study of the support vector machine (SVM) sensitivity to various processing steps in the context of face authentication. In particular, we evaluate the impact of the representation space and photometric normalisation technique on the SVM performance. Our study supports the hypothesis that the SVM approach is able to extract the relevant discriminatory information from the training data. We believe that this is the main reason for its superior performance over benchmark methods (e.g. the eigenface technique). However, when the representation space already captures and emphasises the discriminatory information content (e.g. the fisherface method), the SVMs cease to be superior to the benchmark techniques. The SVM performance evaluation is carried out on a large face database containing 295 subjects. Kenneth Jonsson, Josef Kittler, Yongping Li, Jiri Matas |
BMVC | 4 |
| 1999 | Effective Implementation of Linear Discriminant Analysis for Face Recognition and Verification
Yongping Li, Josef Kittler, Jiri Matas |
CAIP | 3 |
| 1999 | Audio-Visual Person VerificationabstractIn this paper we investigate benefits of classifier combination (fusion) for a multimodal system for personal identity verification. The system uses frontal face images and speech. We show that a sophisticated fusion strategy enables the system to outperform its facial and vocal modules when taken seperately. We show that both trained linear weighted schemes and fusion by Support Vector Machine classifier leads to a significant reduction of total error rates. The complete system is tested on data from a publicly available audio-visual database (XM2VTS, 295 subjects) according to a published protocol. Souheil Ben Yacoub, Jürgen Lüttin, Kenneth Jonsson, Jiri Matas, Josef Kittler |
CVPR | 4 |
| 1999 | Progressive Probabilistic Hough Transform for Line DetectionabstractWe present a novel Hough Transform algorithm referred to as Progressive Probabilistic Hough Transform (PPHT). Unlike the Probabilistic HT where Standard HT is performed on a pre-selected fraction of input points, PPHT minimises the amount of computation needed to detect lines by exploiting the difference an the fraction of votes needed to detect reliably lines with different numbers of supporting points. The fraction of points used for voting need not be specified ad hoc or using a priori knowledge, as in the probabilistic HT; it is a function of the inherent complexity of the input data. The algorithm is ideally suited for real-time applications with a fixed amount of available processing time, since voting and line detection is interleaved. The most salient features are likely to be detected first. Experiments show that in many circumstances PPHT has advantages over the Standard HT. Charles Galambos, Josef Kittler, Jiri Matas |
CVPR | 3 |
| 1999 | On Camera Calibration for Scene Model Acquisition and Maintenance Using an Active Vision System
Rupert Young 0001, Jiri Matas, Josef Kittler |
ICVS | 2 |
| 1999 | Fast face localisation and verification
Jiri Matas, Kenneth Jonsson, Josef Kittler |
Image Vis. Comput. | 1 |
| 1998 | Saliency-Based Robust Correlation for Real-Time Face Registration and VerificationabstractWe propose a novel person verification system for real-time face identification. The main features of the system include accurate registration of face images using a robust form of correlation, a framework for global registration of a face database using a minimum spanning tree algorithm and a method for selecting a subset of features optimal for discrimination between clients and impostors. The results indicate that the image registration is of high accuracy and the feature selection is successfully improving on the verification performance. 1 Introduction Verification of person identity based on biometric information is important for many security applications. Examples include access control to buildings, surveillance and intrusion detection. Furthermore, there are many emerging fields that would benefit from developments in person verification technology such as advanced human-computer interfaces and tele-services including tele-shopping and tele-banking. Comparing verific... Kenneth Jonsson, Jiri Matas, Josef Kittler, S. Haberl |
BMVC | 2 |
| 1998 | Progressive Probabilistic Hough TransformabstractIn the paper we present the Progressive Probabilistic Hough Transform (PPHT). Unlike the Probabilistic Hough Transform [4] where Standard Hough Transform is performed on a pre-selected fraction of input points, PPHT minimises the amount of computation needed to detect lines by exploiting the difference in the fraction of votes needed to reliably detect lines with different numbers of supporting points. The fraction of points used for voting need not be specified ad hoc or using a priori knowledge, as in the Probabilistic Hough Transform; it is a function of the inherent complexity of data. The algorithm is ideally suited for real-time applications with a fixed amount of available processing time, since voting and line detection is interleaved. The most salient features are likely to be detected first. Experiments show PPHT has, in many circumstances, advantages over the Standard Hough Transform. 1 Introduction The Hough Transform (HT) is a popular method for the extracti... Jiri Matas, Charles Galambos, Josef Kittler |
BMVC | 1 |
| 1998 | Recognition using labelled objectsabstractGeneral object recognition is a difficult problem. We (1997) proposed a solution for object recognition in an unconstrained environment. We simplified the recognition problem by attaching a special planar pattern on objects of interest. This approach allows us to determine the pose easily. A robust detector for the pattern was developed for the first stage of the solution. This paper investigates the next part of the recognition process, i.e. matching of the models using a modified technique based on the Chamfer matching algorithm. The algorithm is enhanced by augmenting the matching process by an additional step which promotes consistency of image gradient directions of the corresponding points. Soh Ling Min, Jiri Matas, Josef Kittler |
ICPR | 2 |
| 1998 | Selection of speaker independent feature for a speaker verification systemabstractIn this paper we propose an optimisation technique to choose a user independent feature subset from the input feature set for a dynamic time-warping (DTW) based text-dependent speaker verification system. The results indicate that with the optimised feature set the verification error rate of the system can be improved. Medha Pandit, Josef Kittler, Jiri Matas |
ICPR | 3 |
| 1998 | Object-detection with a varying number of eigenspace projectionsabstractWe present a method allowing a significant speed-up of the eigen-detection method (detection based on principle component analysis). We derive a formula for an upper bound on the class-conditional probability (or equivalently a lower bound on the Mahalanobis distance) on which detection is based. Often, the lower bound of Mahalanobis distance (MD) reaches a preset threshold after computation of only a few eigen-projections. In this case the computation of MD can be immediately terminated. Regardless of the precise value of MD, the detection hypothesis (object from class /spl Omega/ is detected) can be rejected. While provably obtaining results identical to the standard technique, we achieved a two- to three-fold speed-up in face detection experiments on images from the CMU database. Michael Reiter, Jiri Matas |
ICPR | 2 |
| 1998 | Hypothesis selection for scene interpretation using grammatical models of scene evolutionabstractA major bottleneck in dynamic scene interpretation is the search that is required through a database to find a model that best matches the observed data. We show that the problem can be alleviated if the object model selection is controlled by a scene evolution model. We adopt a grammatical model to characterise objects and events in a dynamic scene which can be used to generate visual expectations within a particular context. The object hypotheses can be accepted without further search of the database provided a measure of the goodness of fit of the match between the selected model and the visual data falls below a threshold. In this paper we present experiments for determining the necessary thresholds for the model hypotheses testing using the recognition method described by Yang et al. (1994), as well as for assessing the subsequent performance of the scene interpretation system with and without the constraining grammar. Rupert Young 0001, Josef Kittler, Jiri Matas |
ICPR | 3 |
| 1998 | On Combining ClassifiersabstractWe develop a common theoretical framework for combining classifiers which use distinct pattern representations and show that many existing schemes can be considered as special cases of compound classification where all the pattern representations are used jointly to make a decision. An experimental comparison of various classifier combination schemes demonstrates that the combination rule developed under the most restrictive assumptions-the sum rule-outperforms other classifier combinations schemes. A sensitivity analysis of the various schemes to estimation errors is carried out to show that this finding can be justified theoretically. Josef Kittler, Mohamad Hatef, Robert P. W. Duin, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 1997 | Fast Face Localisation and Verification
Jiri Matas, Kenneth Jonsson, Josef Kittler |
BMVC | 1 |
| 1997 | Statistical chromaticity-based lip tracking with B-splinesabstractWe present a statistical, colour-based technique for lip tracking intended to support personal verification. The lips are automatically localised in the original image by exploiting grey-level gradient projections as well as chromaticity models to find the mouth area in an automatically segmented region corresponding to the face area. A B-spline, initially with an elliptic shape is then generated to start up tracking. Tracking proceeds by estimating new lip contour positions according to a statistical chromaticity model for the lips. These measurements are used together with a Lagrangian formulation of contour dynamics to update the new spline control points. The method has been tested on the M2VTS database, where lips were accurately tracked on sequences of speaking subjects consisting of more than hundred frames. The tracker can be used to perform feature extraction from the mouth area as well as for model detection for personal verification applications. M. Ulises Ramos Sánchez, Jiri Matas, Josef Kittler |
ICASSP | 2 |
| 1997 | Object Recognition Using a TagabstractWe propose a method for object recognition in an unconstrained condition. This includes wide range of illumination, unknown view points and complicated background. We simplified the general problem by placing special design patterns (Tsai's (1987) camera calibration chart) on the object that allows us to solve the pose determination problem easily. After establishing the methodological framework, we experiment with this technique on a wide range of real data to show the reliability of the approach. We argue that with this capability we should be able to (a) establish camera position with respect to a landmark, (b) recognise an object which is tagged with the calibration chart and (c) test any camera calibration and 3D pose estimation routines, thus facilitating future research and applications in mobile robots navigation, 3D reconstruction and stereo vision. Jiri Matas, Soh Ling Min, Josef Kittler |
ICIP (1) | 1 |
| 1997 | Combining evidence in personal identity verification systems
Josef Kittler, Jiri Matas, Kenneth Jonsson, M. Ulises Ramos Sánchez |
Pattern Recognit. Lett. | 2 |
| 1996 | Using grammars for scene interpretationabstractA method that employs grammars to direct the inference process of a vision system that does interpretation of dynamic scenes is described. The system uses a set of qualitative image descriptors to drive the interpretation. The result is a 'natural language' description of scene activities. In addition the inference engine generates a set of predictions that can be used to control the interpretation strategy so as to make the processing of new images more efficient. The system has been implemented in an expert system shell to demonstrate the viability of the approach. Results on real images are reported. Henrik I. Christensen, Jiri Matas, Josef Kittler |
ICIP (2) | 2 |
| 1996 | Techniques for the interpretation of thermal paint coated samplesabstractThermal paint, paint that changes colour with maximum temperature, has a variety of uses within the automotive industry. The analysis of samples after exposure to heat is a task currently performed by humans. We present the current findings of our attempts to automate this procedure by means of image analysis. A framework has been developed within which colour is mapped to an RGB space curve which is a function of temperature. This models the context imparted by the physical process, as well as spatial context, and gives us better results than those obtained by using conventional image segmentation techniques. Andrew Griffin, Josef Kittler, Terry Windeatt, Jiri Matas |
ICPR | 4 |
| 1995 | Spatial and Feature Space Clustering: Applications in Image Analysis
Jiri Matas, Josef Kittler |
CAIP | 1 |
| 1995 | On Representation and Matching of Multi-Coloured ObjectsabstractA new representation for objects with multiple colours-the colour adjacency graph (CAG)-is proposed. Each node of the CAG represents a single chromatic component of the image defined as a set of pixels forming a unimodal cluster in the chromatic scattergram. Edges encode information about adjacency of colour components and their reflectance ratio. The CAG is related to both the histogram and region adjacency graph representations. It is shown to be preserving and combining the best features of these two approaches while avoiding their drawbacks. The proposed approach is tested on a range of difficult object recognition and localisation problems involving complex imagery of non-rigid 3D objects under varied viewing conditions with excellent results.> Jiri Matas, Radek Marik, Josef Kittler |
ICCV | 1 |
| 1995 | Colour-based object recognition under spectrally non-uniform illumination
Jiri Matas, Radek Marik, Josef Kittler |
Image Vis. Comput. | 1 |
| 1995 | Intentional control of camera look direction and viewpoint in an active vision system
Paolo Remagnino, John Illingworth, Josef Kittler, Jiri Matas |
Image Vis. Comput. | 4 |
| 1994 | Illumination Invariant Colour RecognitionabstractThe article describes a colour based recognition system with three novel features. Firstly, the proposed system can operate in environments where spectral characteristics of illumination change in both space and time. Secondly, benefits in terms of speed and quality of output are gained by focusing processing to areas of salient colour. Finally, an automatic model acquisition procedure allows rapid creation of the model database. 1 Introduction In the paper we present a colour-based recognition system that aims to demonstrate the advantages of selective processing. We do not attempt to analyse the whole image in the spirit of traditional segmentation methods (eg. [KSK87],[GJT87]); instead we try to find areas where distinctive colour provides least ambiguous information about presence of objects from the model database. Using this approach, standard recognition tasks (eg. What is in the scene?, Where is object X?) can be accomplished without wasting computational resources in parts o... Jiri Matas, Radek Marik, Josef Kittler |
BMVC | 1 |
| 1994 | Recognition of Cylindrical Objects Using Occluding Boundaries Obtained from Colour Based SegmentationabstractThis paper describes a method for model-based recognition of cylindrical ob-jects from occluding boundaries obtained by computationally efficient colour segmentation of a 2D image. The models are invoked by combining geomet-ric and colour features. Occluding boundaries of hypothesized objects are generated using colour segmentation and ground plane constraint. Hypothe-sis verification is achieved by evaluating the fit between occluding boundary generated by the hypothesised object and the edge data. This method dif-fers from existing methods in that it integrates multiple measurements and prior knowledge to achieve robust object recognition. Experiments with real images have been carried out and the results are promising. 1 Dekun Yang, Josef Kittler, Jiri Matas |
BMVC | 3 |
| 1993 | Generation, Verification and Localization of Object Hypotheses based on ColourabstractThis paper presents a model-based method for colour-based recognition of objects from a large database. The algorithm is based on the assumption that surface reflectances of objects in the model database follow the extended dichromatic model proposed by Shafer [Sha84]. Adoption of the dichromatic model allows recovery of body colour - the component, of sensor responses (RGB-values) that is independent of scene geometry and illumination intensity. Both theoretical studies [Hea89b] and experiments [LB90][KSK88] confirm that Shafer's model gives a suitable approximation for reflectances of a wide range of materials. Instead of using traditional techniques (eg. clustering, split-and-merge) to obtain regions of 'similarly' coloured pixels followed by classification a novel approach is argued for. First, for each pixel a list of models with nonzero aposteriori probabilities P(modeU\body colotir) is computed using Bayes formula. Next, regions are formed by grouping pixels with identical most probable hypothesis. Probabilities P(modeli\region) are obtained trough a standard group decision rule [FT80]. We show that the proposed scheme can be used for a number of visual tasks - localization of objects, generation and verification of object hypotheses. Experiments on images of complex indoor scenes confirm that the proposed method can provide reliable information about the surrounding environment. Jiri Matas, Radek Marik, Josef Kittler |
BMVC | 1 |
| 1993 | Junction detection using probabilistic relaxation
Jiri Matas, Josef Kittler |
Image Vis. Comput. | 1 |
| 1992 | Contextual Junction Finder
Jiri Matas, Josef Kittler |
BMVC | 1 |
| 1992 | On Computing the Next Look Camera Parameters in Active Vision
Paolo Remagnino, Josef Kittler, Jiri Matas, John Illingworth |
ECAI | 3 |
| 1991 | Low-level Grouping of Straight Line Segments
A. Etemadi, J.-P. Schmidt, Jiri Matas, John Illingworth, Josef Kittler |
BMVC | 3 |