EDBT 2026 Demo / reviewers in the wild / expert
Krystian Mikolajczyk
dblp:96/433
· DBLP profile ↗
117ranked-venue papers
16as first author
33since 2021 · last 2026
0000-0003-0726-9187ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 92 · 14 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 81 · 12 first-author · 18 since 2021Systems, architecture and hardware · 5 · 3 since 2021Computer networks · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Match Stereo Videos via Bidirectional AlignmentabstractVideo stereo matching is the task of estimating consistent disparity maps from rectified stereo videos. There is considerable scope for improvement in both datasets and methods within this area. Recent learning-based methods often focus on optimizing performance for independent stereo pairs, leading to temporal inconsistencies in videos. Existing video methods typically employ sliding window operation over time dimension, which can result in low-frequency oscillations corresponding to the window size. To address these challenges, we propose a bidirectional alignment mechanism for adjacent frames as a fundamental operation. Building on this, we introduce a novel video processing framework, BiDAStereo, and a plugin stabilizer network, BiDAStabilizer, compatible with general image-based methods. Regarding datasets, current synthetic object-based and indoor datasets are commonly used for training and benchmarking, with a lack of outdoor nature scenarios. To bridge this gap, we present a realistic synthetic dataset and benchmark focused on natural scenes, along with a real-world dataset captured by a stereo camera in diverse urban scenes for qualitative evaluation. Extensive experiments on in-domain, out-of-domain, and robustness evaluation demonstrate the contribution of our methods and datasets, showcasing improvements in prediction quality and achieving state-of-the-art results on various commonly used benchmarks. Junpeng Jing, Ye Mao, Anlan Qiu, Krystian Mikolajczyk |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | MIMO Channel as a Neural Function: Implicit Neural Representations for Extreme CSI CompressionabstractAcquiring and utilizing accurate channel state information (CSI) is crucial for realizing the benefits of massive multiple-input multiple-output (MIMO) technology. Current CSI feedback approaches improve precision by employing advanced deep-learning methods to learn representative CSI features for a subsequent compression process. Diverging from previous works, we treat the CSI compression problem in the context of implicit neural representations. Specifically, each CSI matrix is viewed as a neural function that maps the spatial coordinates (antenna and subchannel) to the corresponding channel gains with physical significance. Rather than transmitting the parameters of the specific neural functions directly, we send low-cost modulations of the CSI matrix, derived through a meta-learning algorithm. These modulations are then applied to a shared base network at the receiver to reconstruct the CSI matrix. Numerical results show that our proposed approach achieves state-of-the-art performance and showcases flexibility in feedback strategies. Maojun Zhang, Yulin Shao, Krystian Mikolajczyk, Deniz Gündüz |
ICASSP | 4 |
| 2025 | No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse ViewsabstractWe introduce SPFSplat, an efficient framework for 3D Gaussian splatting from sparse multi-view images, requiring no ground-truth poses during training or inference. It employs a shared feature extraction backbone, enabling simultaneous prediction of 3D Gaussian primitives and camera poses in a canonical space from unposed inputs within a single feed-forward step. Alongside the rendering loss based on estimated novel-view poses, a reprojection loss is integrated to enforce the learning of pixel-aligned Gaussian primitives for enhanced geometric constraints. This pose-free training paradigm and efficient one-step feed-forward design make SPFSplat well-suited for practical applications. Remarkably, despite the absence of pose supervision, SPFSplat achieves state-of-the-art performance in novel view synthesis even under significant viewpoint changes and limited image overlap. It also surpasses recent methods trained with geometry priors in relative pose estimation. Code and trained models are available on our project page: https://ranrhuang.github.io/spfsplat/. Ranran Huang 0004, Krystian Mikolajczyk |
ICCV | 2 |
| 2025 | Stereo Any Video: Temporally Consistent Stereo Matching
Junpeng Jing, Weixun Luo, Ye Mao, Krystian Mikolajczyk |
ICCV | 4 |
| 2025 | Hypo3D: Exploring Hypothetical Reasoning in 3DabstractThe rise of vision-language foundation models marks an advancement in bridging the gap between human and machine capabilities in 3D scene reasoning. Existing 3D reasoning benchmarks assume real-time scene accessibility, which is impractical due to the high cost of frequent scene updates. To this end, we introduce Hypothetical 3D Reasoning, namely Hypo3D, a benchmark designed to evaluate models’ ability to reason without access to real-time scene data. Models need to imagine the scene state based on a provided change description before reasoning. Hypo3D is formulated as a 3D Visual Question Answering (VQA) benchmark, comprising 7,727 context changes across 700 indoor scenes, resulting in 14,885 question-answer pairs. An anchor-based world frame is established for all scenes, ensuring consistent reference to a global frame for directional terms in context changes and QAs. Extensive experiments show that state-of-the-art foundation models struggle to reason effectively in hypothetically changed scenes. This reveals a substantial performance gap compared to humans, particularly in scenarios involving movement changes and directional reasoning. Even when the change is irrelevant to the question, models often incorrectly adjust their answers. The code and dataset are publicly available at: https://matchlab-imperial.github.io/Hypo3D. Ye Mao, Weixun Luo, Junpeng Jing, Anlan Qiu, Krystian Mikolajczyk |
ICML | 5 |
| 2025 | Closed Loop Interactive Embodied Reasoning for Robot ManipulationabstractEmbodied reasoning systems integrate robotic hardware and cognitive processes to perform complex tasks, typically in response to a natural language query about a specific physical environment. This usually involves changing the belief about the scene or physically interacting and changing the scene (e.g. sort the objects from lightest to heaviest). In order to facilitate the development of such systems we introduce a new modular Closed Loop Interactive Embodied Reasoning (CLIER) approach that takes into account the measurements of non-visual object properties, changes in the scene caused by external disturbances as well as uncertain outcomes of robotic actions. CLIER performs multi-modal reasoning and action planning and generates a sequence of primitive actions that can be executed by a robot manipulator. Our method operates in a closed loop, responding to changes in the environment. Our approach is developed with the use of MuBle simulation environment and tested in$\mathbf{1 0}$interactive benchmark scenarios. We extensively evaluate our reasoning approach in simulation and in real-world manipulation tasks with a success rate above$\mathbf{7 6 \%}$and 64%, respectively. Michal Nazarczuk, Jan Kristof Behrens, Karla Stépánová, Matej Hoffmann, Krystian Mikolajczyk |
ICRA | 5 |
| 2024 | Understanding the Role of the Projector in Knowledge DistillationabstractIn this paper we revisit the efficacy of knowledge distillation as a function matching and metric learning problem. In doing so we verify three important design decisions, namely the normalisation, soft maximum function, and projection layers as key ingredients. We theoretically show that the projector implicitly encodes information on past examples, enabling relational gradients for the student. We then show that the normalisation of representations is tightly coupled with the training dynamics of this projector, which can have a large impact on the students performance. Finally, we show that a simple soft maximum function can be used to address any significant capacity gap problems. Experimental results on various benchmark datasets demonstrate that using these insights can lead to superior or comparable performance to state-of-the-art knowledge distillation techniques, despite being much more computationally efficient. In particular, we obtain these results across image classification (CIFAR100 and ImageNet), object detection (COCO2017), and on more difficult distillation objectives, such as training data efficient transformers, whereby we attain a 77.2% top-1 accuracy with DeiT-Ti on ImageNet. Code and models are publicly available. Roy Miles, Krystian Mikolajczyk |
AAAI | 2 |
| 2024 | Learning to Project for Cross-Task Knowledge Distillation
Dylan Auty, Roy Miles, Benedikt Kolbeinsson, Krystian Mikolajczyk |
BMVC | 4 |
| 2024 | Match-Stereo-Videos: Bidirectional Alignment for Consistent Dynamic Stereo Matching
Junpeng Jing, Ye Mao, Krystian Mikolajczyk |
ECCV (60) | 3 |
| 2024 | AirEyeSeg: Teacher-Student Insights into Robust Fisheye UAV Detection
Zhenyue Gu, Benedikt Kolbeinsson, Krystian Mikolajczyk |
ICPRAM | 3 |
| 2024 | Interactive Learning of Physical Object Properties Through Robot Manipulation and Database of Object MeasurementsabstractThis work presents a framework for automatically extracting physical object properties, such as material composition, mass, volume, and stiffness, through robot manipulation and a database of object measurements. The framework involves exploratory action selection to maximize learning about objects on a table. A Bayesian network models conditional dependencies between object properties, incorporating prior probability distributions and uncertainty associated with measurement actions. The algorithm selects optimal exploratory actions based on expected information gain and updates object properties through Bayesian inference. Experimental evaluation demonstrates effective action selection compared to a baseline and correct termination of the experiments if there is nothing more to be learned. The algorithm proved to behave intelligently when presented with trick objects with material properties in conflict with their appearance. The robot pipeline integrates with a logging module and an online database of objects, containing over 24,000 measurements of 63 objects with different grippers. All code and data are publicly available, facilitating automatic digitization of objects and their physical properties through exploratory manipulations. Andrej Kruzliak, Jiri Hartvich, Shubhan P. Patni, Lukas Rustler, Jan Kristof Behrens, Fares J. Abu-Dakka, Krystian Mikolajczyk, Ville Kyrki, Matej Hoffmann |
IROS | 7 |
| 2024 | OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned ImagesabstractRecent open-world 3D representation learning methods using Vision-Language Models (VLMs) to align 3D point clouds with image-text information have shown superior 3D zero-shot performance. However, CAD-rendered images for this alignment often lack realism and texture variation, compromising alignment robustness. Moreover, the volume discrepancy between 3D and 2D pretraining datasets highlights the need for effective strategies to transfer the representational abilities of VLMs to 3D learning. In this paper, we present OpenDlign, a novel open-world 3D model using depth-aligned images generated from a diffusion model for robust multimodal alignment. These images exhibit greater texture diversity than CAD renderings due to the stochastic nature of the diffusion model. By refining the depth map projection pipeline and designing depth-specific prompts, OpenDlign leverages rich knowledge in pre-trained VLM for 3D representation learning with streamlined fine-tuning. Our experiments show that OpenDlign achieves high zero-shot and few-shot performance on diverse 3D tasks, despite only fine-tuning 6 million parameters on a limited ShapeNet dataset. In zero-shot classification, OpenDlign surpasses previous models by 8.0\% on ModelNet40 and 16.4\% on OmniObject3D. Additionally, using depth-aligned images for multimodal alignment consistently enhances the performance of other state-of-the-art models. Ye Mao, Junpeng Jing, Krystian Mikolajczyk |
NeurIPS | 3 |
| 2024 | Multi-Class Segmentation from Aerial Views using Recursive Noise DiffusionabstractSemantic segmentation from aerial views is a crucial task for autonomous drones, as they rely on precise and accurate segmentation to navigate safely and efficiently. However, aerial images present unique challenges such as diverse viewpoints, extreme scale variations, and high scene complexity. In this paper, we propose an end-to-end multiclass semantic segmentation diffusion model that addresses these challenges. We introduce recursive denoising to allow information to propagate through the denoising process, as well as a hierarchical multi-scale approach that complements the diffusion process. Our method achieves promising results on the UAVid dataset and state-of-the-art performance on the Vaihingen Building segmentation benchmark. Being the first iteration of this method, it shows great promise for future improvements. Our code and models are available at: https://github.com/benediktkol/recursive-noise-diffusion Benedikt Kolbeinsson, Krystian Mikolajczyk |
WACV | 2 |
| 2024 | AirNet: Neural Network Transmission Over the AirabstractState-of-the-art performance for many edge applications is achieved by deep neural networks (DNNs). Often, these DNNs are location- and time-sensitive, and must be delivered over a wireless channel rapidly and efficiently. In this paper, we introduce AirNet, a family of novel methods that allow DNNs to be efficiently delivered over wireless channels under stringent transmit power and latency constraints. It is a part of a new class of joint source-channel coding methods that maximize the accuracy of the DNNs transmitted to the receiver, rather than recover the DNNs with high fidelity. In AirNet, we propose to directly map the DNN parameters to the transmitted channel symbols, while training the network under the channel constraints with robustness to channel noise. AirNet achieves higher accuracy compared to the separation-based alternatives. We further improve its performance by pruning the network below the available bandwidth, and using bandwidth expansion for significant network parameters. We also exploit unequal error protection (UEP) by selectively expanding the important layers. Finally, we propose an ensemble training approach where networks for different channel conditions can be obtained simultaneously, resolving the impractical memory requirements of training distinct networks for different channel conditions. Mikolaj Jankowski, Deniz Gündüz, Krystian Mikolajczyk |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Deep Joint Source-Channel Coding for Adaptive Image Transmission Over MIMO ChannelsabstractWe introduce a vision transformer (ViT)-based deep joint source and channel coding (DeepJSCC) scheme for wireless image transmission over multiple-input multiple-output (MIMO) channels, called DeepJSCC-MIMO. We employ DeepJSCC-MIMO in both open-loop and closed-loop MIMO systems. The novel DeepJSCC-MIMO architecture surpasses the classical separation-based benchmarks, while exhibiting robustness to channel estimation errors and flexibility in adapting to diverse channel conditions and antenna configurations without requiring retraining. Specifically, by harnessing the self-attention mechanism of the ViT, DeepJSCC-MIMO intelligently learns feature mapping and power allocation strategies tailored to the unique characteristics of the source image and prevailing channel conditions. Extensive numerical experiments validate the significant improvements in both distortion quality and perceptual quality achieved by DeepJSCC-MIMO for both open-loop and closed-loop MIMO systems across a wide range of scenarios. Moreover, DeepJSCC-MIMO exhibits robustness to varying channel conditions, channel estimation errors, and different antenna numbers, making it an appealing technology for emerging semantic communication systems. Yulin Shao, Chenghong Bian, Krystian Mikolajczyk, Deniz Gündüz |
IEEE Trans. Wirel. Commun. | 4 |
| 2024 | Transformer-Aided Wireless Image Transmission With Channel FeedbackabstractThis paper presents a novel wireless image transmission paradigm that can exploit feedback from the receiver, called JSCCformer-f. We consider a block feedback channel model, where the transmitter receives noiseless/noisy channel output feedback after each block. The proposed scheme employs a single encoder to facilitate transmission over multiple blocks, refining the receiver’s estimation at each block. Specifically, the unified encoder of JSCCformer-f can leverage the semantic information from the source image, and acquire channel state information and the decoder’s current belief about the source image from the feedback signal to generate coded symbols at each block. Numerical experiments show that our JSCCformer-f scheme achieves state-of-the-art performance with robustness to noise in the feedback link. Additionally, JSCCformer-f can adapt to the channel condition directly through feedback without the need for separate channel estimation. We further extend the scope of the JSCCformer-f approach to include the broadcast channel, which enables the transmitter to generate broadcast codes in accordance with signal semantics and channel feedback from individual receivers. Yulin Shao, Emre Ozfatura, Krystian Mikolajczyk, Deniz Gündüz |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | Vision Transformer for Adaptive Image Transmission over MIMO ChannelsabstractThis paper presents a vision transformer (ViT) based joint source and channel coding (JSCC) scheme for wireless image transmission over multiple-input multiple-output (MIMO) systems, called ViT-MIMO. The proposed ViT-MIMO architecture, in addition to outperforming separation-based benchmarks, can flexibly adapt to different channel conditions without requiring retraining. Specifically, exploiting the self-attention mechanism of the ViT enables the proposed ViT-MIMO model to adaptively learn the feature mapping and power allocation based on the source image and channel conditions. Numerical experiments show that ViT-MIMO can significantly improve the transmission quality across a large variety of scenarios, including varying channel conditions, making it an attractive solution for emerging semantic communication systems. Yulin Shao, Chenghong Bian, Krystian Mikolajczyk, Deniz Gündüz |
ICC | 4 |
| 2023 | Project to Adapt: Domain Adaptation for Depth Completion from Noisy and Sparse Sensor DataabstractAbstract Depth completion aims to predict a dense depth map from a sparse depth input. The acquisition of dense ground-truth annotations for depth completion settings can be difficult and, at the same time, a significant domain gap between real LiDAR measurements and synthetic data has prevented from successful training of models in virtual settings. We propose a domain adaptation approach for sparse-to-dense depth completion that is trained from synthetic data, without annotations in the real domain or additional sensors. Our approach simulates the real sensor noise in an RGB + LiDAR set-up, and consists of three modules: simulating the real LiDAR input in the synthetic domain via projections, filtering the real noisy LiDAR for supervision and adapting the synthetic RGB image using a CycleGAN approach. We extensively evaluate these modules in the KITTI depth completion benchmark. Adrián López Rodríguez, Benjamin Busam, Krystian Mikolajczyk |
Int. J. Comput. Vis. | 3 |
| 2023 | DESC: Domain Adaptation for Depth Estimation via Semantic ConsistencyabstractAbstract Accurate real depth annotations are difficult to acquire, needing the use of special devices such as a LiDAR sensor. Self-supervised methods try to overcome this problem by processing video or stereo sequences, which may not always be available. Instead, in this paper, we propose a domain adaptation approach to train a monocular depth estimation model using a fully-annotated source dataset and a non-annotated target dataset. We bridge the domain gap by leveraging semantic predictions and low-level edge features to provide guidance for the target domain. We enforce consistency between the main model and a second model trained with semantic segmentation and edge maps, and introduce priors in the form of instance heights. Our approach is evaluated on standard domain adaptation benchmarks for monocular depth estimation and show consistent improvement upon the state-of-the-art. Code available at https://github.com/alopezgit/DESC . Adrián López Rodríguez, Krystian Mikolajczyk |
Int. J. Comput. Vis. | 2 |
| 2023 | Key.Net: Keypoint Detection by Handcrafted and Learned CNN Filters RevisitedabstractWe introduce a novel approach for keypoint detection that combines handcrafted and learned CNN filters within a shallow multi-scale architecture. Handcrafted filters provide anchor structures for learned filters, which localize, score, and rank repeatable features. Scale-space representation is used within the network to extract keypoints at different levels. We design a loss function to detect robust features that exist across a range of scales and to maximize the repeatability score. Our Key.Net model is trained on data synthetically created from ImageNet and evaluated on HPatches and other benchmarks. Results show that our approach outperforms state-of-the-art detectors in terms of repeatability, matching performance, and complexity. Key.Net implementations in TensorFlow and PyTorch are available online. Axel Barroso Laguna, Krystian Mikolajczyk |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | OoD-Pose: Camera Pose Regression From Out-of-Distribution Synthetic ViewsabstractIn this paper, we address the problem of camera pose estimation in outdoor and indoor scenarios. We propose a relative pose regression method that can directly regress the camera pose from images with significantly higher accuracy than existing methods of the same class. We first investigate one of the main factors that limits the accuracy of relative pose regression, and then introduce a new approach that significantly improves the performance. Specifically, we propose a method to overcome the biased training data by a novel training technique. It generates poses, guided by a probability distribution of the training set, which are then used to synthesise new views for training. Lastly, we evaluate our approach on widely used benchmarks and show that it achieves significantly lower error compared to prior regression-based methods and retrieval techniques. Tony Ng, Adrián López Rodríguez, Vassileios Balntas, Krystian Mikolajczyk |
3DV | 4 |
| 2022 | Information Theoretic Representation Distillation
Roy Miles, Adrián López Rodríguez, Krystian Mikolajczyk |
BMVC | 3 |
| 2022 | Estimating water turbidity from a smartphone camera
Lina M. Lozano Wilches, Chotiwat Jantarakasem, Laure Sioné, Michael Templeton, Krystian Mikolajczyk |
BMVC | 5 |
| 2022 | ScaleNet: A Shallow Architecture for Scale EstimationabstractIn this paper, we address the problem of estimating scale factors between images. We formulate the scale estimation problem as a prediction of a probability distribution over scale factors. We design a new architecture, SealeNet, that exploits dilated convolutions as well as self- and cross-correlation layers to predict the scale between images. We demonstrate that rectifying images with estimated scales leads to significant performance improvements for various tasks and methods. Specifically, we show how ScaleNet can be combined with sparse local features and dense correspondence networks to improve camera pose estimation, 3D reconstruction, or dense geometric matching in different benchmarks and datasets. We provide an extensive evaluation on several tasks, and analyze the computational overhead of SealeNet. The code, evaluation protocols, and trained models are publicly available at https://github.com/axelBarroso/ScaleNet. Axel Barroso Laguna, Yurun Tian, Krystian Mikolajczyk |
CVPR | 3 |
| 2022 | NinjaDesc: Content-Concealing Visual Descriptors via Adversarial LearningabstractIn the light of recent analyses on privacy-concerning scene revelation from visual descriptors, we develop descriptors that conceal the input image content. In particular, we propose an adversarial learning framework for training visual descriptors that prevent image reconstruction, while maintaining the matching accuracy. We let a feature encoding network and image reconstruction network compete with each other, such that the feature encoder tries to impede the image reconstruction with its generated descriptors, while the reconstructor tries to recover the input image from the descriptors. The experimental results demonstrate that the visual descriptors obtained with our method significantly deteriorate the image reconstruction quality with minimal impact on correspondence matching and camera localization performance. Tony Ng, Hyo Jin Kim 0004, Vincent T. Lee, Daniel DeTone, Tsun-Yi Yang, Tianwei Shen, Eddy Ilg, Vassileios Balntas, Krystian Mikolajczyk, Chris Sweeney |
CVPR | 9 |
| 2022 | Monocular Depth Estimation Using Cues Inspired by Biological Vision SystemsabstractMonocular depth estimation (MDE) aims to transform an RGB image of a scene into a pixelwise depth map from the same camera view. It is fundamentally ill-posed due to missing information: any single image can have been taken from many possible 3D scenes. Part of the MDE task is, therefore, to learn which visual cues in the image can be used for depth estimation, and how. With training data limited by cost of annotation or network capacity limited by computational power, this is challenging.In this work we demonstrate that explicitly injecting visual cue information into the model is beneficial for depth estimation. Following research into biological vision systems, we focus on semantic information and prior knowledge of object sizes and their relations, to emulate the biological cues of relative size, familiar size, and absolute size. We use state-of-the-art semantic and instance segmentation models to provide external information, and exploit language embeddings to encode relational information between classes. We also provide a prior on the average real-world size of objects. This external information overcomes the limitation in data availability, and ensures that the limited capacity of a given network is focused on known-helpful cues, therefore improving performance. We experimentally validate our hypothesis and evaluate the proposed model on the widely used NYUD2 indoor depth estimation benchmark. The results show improvements in depth prediction when the semantic information, size prior and instance size are explicitly provided along with the RGB images, and our method can be easily adapted to any depth estimation system. Dylan Auty, Krystian Mikolajczyk |
ICPR | 2 |
| 2022 | Everything Has Been Done, It Is Our Job to Do It One Better in Image Matching and Localisation
Krystian Mikolajczyk |
ICPRAM | 1 |
| 2022 | PRiDAN: Person Re-identification from Drones with Adaptive Weights and Expanded Neighbourhood
Chatchanan Varojpipath, Krystian Mikolajczyk |
ICPRAM | 2 |
| 2022 | AirNet: Neural Network Transmission over the AirabstractState-of-the-art performance for many emerging edge applications is achieved by deep neural networks (DNNs). Often, the employed DNNs are location- and time-dependent, and the parameters of a specific DNN must be delivered from an edge server to the edge device rapidly and efficiently to carry out time-sensitive inference tasks. This can be considered as a joint source-channel coding (JSCC) problem, in which the goal is not to recover the DNN coefficients with the minimal distortion, but in a manner that provides the highest accuracy in the downstream task. For this purpose we introduce AirNet, a novel training and analog transmission method to deliver DNNs over the air. We first train the DNN with noise injection to counter the wireless channel noise. We also employ pruning to identify the most significant DNN parameters that can be delivered within the available channel bandwidth, knowledge distillation, and nonlinear bandwidth expansion to provide better error protection for the most important network parameters. We show that AirNet achieves significantly higher test accuracy compared to the separation-based alternative, and exhibits graceful degradation with channel quality. Mikolaj Jankowski, Deniz Gündüz, Krystian Mikolajczyk |
ISIT | 3 |
| 2021 | Compressing Local Descriptor Models for Mobile ApplicationsabstractFeature-based image matching has been significantly improved through the use of deep learning and new large datasets. However, there has been little work addressing the computational cost, model size, and matching accuracy tradeoffs for the state of the art models. In this paper, we consider these practical aspects and improve the state-of-the-art HardNet model through the use of depthwise separable layers and an efficient tensor decomposition. We propose the Convolution-Depthwise-Pointwise (CDP) layer, which partitions the weights into a low and full rank decomposition to exploit the naturally emergent structure in the convolutional weights. We can achieve an 8× reduction in the number of parameters on the HardNet model, 13× reduction in the computational complexity, while sacrificing less than 1% on the overall accuracy across the HPatches benchmarks. To further demonstrate the generalisation of this approach, we apply it to other state-of-the-art descriptor models, where we are able to a significant performance improvement. Roy Miles, Krystian Mikolajczyk |
ICASSP | 2 |
| 2021 | Embodied Reasoning for Discovering Object Properties via ManipulationabstractIn this paper, we present an integrated system that includes reasoning from visual and natural language inputs, action and motion planning, executing tasks by a robotic arm, manipulating objects, and discovering their properties. A vision to action module recognises the scene with objects and their attributes and analyses enquiries formulated in natural language. It performs multi-modal reasoning and generates a sequence of simple actions that can be executed by a robot. The scene model and action sequence are sent to a planning and execution module that generates a motion plan with collision avoidance, simulates the actions, and executes them. We use synthetic data to train various components of the system and test on a real robot to show the generalization capabilities. We focus on a tabletop scenario with objects that can be grasped by our embodied agent i.e. a 7DoF manipulator with a two-finger gripper. We evaluate the agent on 60 representative queries repeated 3 times (e.g., ’Check what is on the other side of the soda can’) concerning different objects and tasks in the scene. We perform experiments in a simulated and real environment and report the success rate for various components of the system. Our system achieves up to 80.6% success rate on challenging scenes and queries. We also analyse and discuss the challenges that such an intelligent embodied system faces. Jan Kristof Behrens, Michal Nazarczuk, Karla Stépánová, Matej Hoffmann, Yiannis Demiris, Krystian Mikolajczyk |
ICRA | 6 |
| 2021 | Multi-Attentive Detection of the Spider Monkey Whinny in the (Actual) WildabstractWe study deep bioacoustic event detection through multi-head attention based pooling, exemplified by wildlife monitoring.In the multiple instance learning framework, a core deep neural network learns a projection of the input acoustic signal into a sequence of embeddings, each representing a segment of the input.Sequence pooling is then required to aggregate the information present in the sequence such that we have a single clip-wise representation.We propose an improvement based on Squeeze-and-Excitation mechanisms upon a recently proposed audio tagging ResNet, and show that it performs significantly better than the baseline, as well as a collection of other recent audio models.We then further enhance our model, by performing an extensive comparative study of recent sequence pooling mechanisms, and achieve our best result using multi-head selfattention followed by concatenation of the head-specific pooled embeddings -better than prediction pooling methods, as well as compared to other recent sequence pooling tricks.We perform these experiments on a novel dataset of spider monkey whinny calls we introduce here, recorded in a rainforest in the South-Pacific coast of Costa Rica, with a promising outlook pertaining to minimally invasive wildlife monitoring. Georgios Rizos, Jenna Lawson, Zhuoda Han, Duncan Butler, James Rosindell, Krystian Mikolajczyk, Cristina Banks-Leite, Björn W. Schuller |
Interspeech | 6 |
| 2021 | Wireless Image Retrieval at the EdgeabstractWe study the image retrieval problem at the wireless edge, where an edge device captures an image, which is then used to retrieve similar images from an edge server. These can be images of the same person or a vehicle taken from other cameras at different times and locations. Our goal is to maximize the accuracy of the retrieval task under power and bandwidth constraints over the wireless link. Due to the stringent delay constraint of the underlying application, sending the whole image at a sufficient quality is not possible. We propose two alternative schemes based on digital and analog communications, respectively. In the digital approach, we first propose a deep neural network (DNN) aided retrieval-oriented image compression scheme, whose output bit sequence is transmitted over the channel using conventional channel codes. In the analog joint source and channel coding (JSCC) approach, the feature vectors are directly mapped into channel symbols. We evaluate both schemes on image based re-identification (re-ID) tasks under different channel conditions, including both static and fading channels. We show that the JSCC scheme significantly increases the end-to-end accuracy, speeds up the encoding process, and provides graceful degradation with channel conditions. The proposed architecture is evaluated through extensive simulations on different datasets and channel conditions, as well as through ablation studies. Mikolaj Jankowski, Deniz Gündüz, Krystian Mikolajczyk |
IEEE J. Sel. Areas Commun. | 3 |
| 2020 | HDD-Net: Hybrid Detector Descriptor with Mutual Interactive Learning
Axel Barroso Laguna, Yannick Verdie, Benjamin Busam, Krystian Mikolajczyk |
ACCV (1) | 4 |
| 2020 | V2A - Vision to Action: Learning Robotic Arm Actions Based on Vision and Language
Michal Nazarczuk, Krystian Mikolajczyk |
ACCV (3) | 2 |
| 2020 | Project to Adapt: Domain Adaptation for Depth Completion from Noisy and Sparse Sensor Data
Adrián López Rodríguez, Benjamin Busam, Krystian Mikolajczyk |
ACCV (1) | 3 |
| 2020 | D2D: Keypoint Extraction with Describe to Detect Approach
Yurun Tian, Vassileios Balntas, Tony Ng, Axel Barroso Laguna, Yiannis Demiris, Krystian Mikolajczyk |
ACCV (3) | 6 |
| 2020 | Text Attribute Aggregation and Visual Feature Decomposition for Person Search
Sara Iodice, Krystian Mikolajczyk |
BMVC | 2 |
| 2020 | Cascaded channel pruning using hierarchical self-distillation
Roy Miles, Krystian Mikolajczyk |
BMVC | 2 |
| 2020 | DESC: Domain Adaptation for Depth Estimation via Semantic Consistency
Adrián López Rodríguez, Krystian Mikolajczyk |
BMVC | 2 |
| 2020 | SOLAR: Second-Order Loss and Attention for Image Retrieval
Tony Ng, Vassileios Balntas, Yurun Tian, Krystian Mikolajczyk |
ECCV (25) | 4 |
| 2020 | Deep Joint Source-Channel Coding for Wireless Image RetrievalabstractMotivated by surveillance applications with wireless cameras or drones, we consider the problem of image retrieval over a wireless channel. Conventional systems apply lossy compression on query images to reduce the data that must be transmitted over a bandwidth and power limited wireless link. We first note that reconstructing the original image is not needed for retrieval tasks; hence, we introduce a deep neutral network (DNN) based compression scheme targeting the retrieval task. Then, we completely remove the compression step, and propose another DNN-based communication scheme that directly maps the feature vectors to channel inputs. This joint source-channel coding (JSCC) approach not only improves the end-to-end accuracy, but also simplifies and speeds up the encoding operation which is highly beneficial for power and latency constrained IoT applications. Mikolaj Jankowski, Deniz Gündüz, Krystian Mikolajczyk |
ICASSP | 3 |
| 2020 | SHOP-VRB: A Visual Reasoning Benchmark for Object PerceptionabstractIn this paper we present an approach and a benchmark for visual reasoning in robotics applications, in particular small object grasping and manipulation. The approach and benchmark are focused on inferring object properties from visual and text data. It concerns small household objects with their properties, functionality, natural language descriptions as well as question-answer pairs for visual reasoning queries along with their corresponding scene semantic representations. We also present a method for generating synthetic data which allows to extend the benchmark to other objects or scenes and propose an evaluation protocol that is more challenging than in the existing datasets. We propose a reasoning system based on symbolic program execution. A disentangled representation of the visual and textual inputs is obtained and used to execute symbolic programs that represent a 'reasoning process' of the algorithm. We perform a set of experiments on the proposed benchmark and compare to results from the state of the art methods. These results expose the shortcomings of the existing benchmarks that may lead to misleading conclusions on the actual performance of the visual reasoning systems. Michal Nazarczuk, Krystian Mikolajczyk |
ICRA | 2 |
| 2020 | HyNet: Learning Local Descriptor with Hybrid Similarity Measure and Triplet LossabstractIn this paper, we investigate how L2 normalisation affects the back-propagated descriptor gradients during training. Based on our observations, we propose HyNet, a new local descriptor that leads to state-of-the-art results in matching. HyNet introduces a hybrid similarity measure for triplet margin loss, a regularisation term constraining the descriptor norm, and a new network architecture that performs L2 normalisation of all intermediate feature maps and the output descriptors. HyNet surpasses previous methods by a significant margin on standard benchmarks that include patch matching, verification, and retrieval, as well as outperforming full end-to-end methods on 3D reconstruction tasks. Yurun Tian, Axel Barroso Laguna, Tony Ng, Vassileios Balntas, Krystian Mikolajczyk |
NeurIPS | 5 |
| 2020 | $\mathbb {H}$H-Patches: A Benchmark and Evaluation of Handcrafted and Learned Local DescriptorsabstractIn this paper, a novel benchmark is introduced for evaluating local image descriptors. We demonstrate limitations of the commonly used datasets and evaluation protocols, that lead to ambiguities and contradictory results in the literature. Furthermore, these benchmarks are nearly saturated due to the recent improvements in local descriptors obtained by learning from large annotated datasets. To address these issues, we introduce a new large dataset suitable for training and testing modern descriptors, together with strictly defined evaluation protocols in several tasks such as matching, retrieval and verification. This allows for more realistic, thus more reliable comparisons in different application scenarios. We evaluate the performance of several state-of-the-art descriptors and analyse their properties. We show that a simple normalisation of traditional hand-crafted descriptors is able to boost their performance to the level of deep learning based descriptors once realistic benchmarks are considered. Additionally we specify a protocol for learning and evaluating using cross validation. We show that when training state-of-the-art descriptors on this dataset, the traditional verification task is almost entirely saturated. Vassileios Balntas, Karel Lenc, Andrea Vedaldi, Tinne Tuytelaars, Jiri Matas, Krystian Mikolajczyk |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Domain Adaptation for Object Detection via Style Consistency
Adrián López Rodríguez, Krystian Mikolajczyk |
BMVC | 2 |
| 2019 | Key.Net: Keypoint Detection by Handcrafted and Learned CNN FiltersabstractWe introduce a novel approach for keypoint detection task that combines handcrafted and learned CNN filters within a shallow multi-scale architecture. Handcrafted filters provide anchor structures for learned filters, which localize, score and rank repeatable features. Scale-space representation is used within the network to extract keypoints at different levels. We design a loss function to detect robust features that exist across a range of scales and to maximize the repeatability score. Our Key.Net model is trained on data synthetically created from ImageNet and evaluated on HPatches benchmark. Results show that our approach outperforms state-of-the-art detectors in terms of repeatability, matching performance and complexity. Axel Barroso Laguna, Edgar Riba, Daniel Ponsa, Krystian Mikolajczyk |
ICCV | 4 |
| 2018 | Partial Person Re-identification with Alignment and Hallucination
Sara Iodice, Krystian Mikolajczyk |
ACCV (6) | 2 |
| 2018 | Deep Segmentation and Registration in X-Ray Angiography Video
Athanasios Vlontzos, Krystian Mikolajczyk |
BMVC | 2 |
| 2018 | Person Re-Identification with Vision and LanguageabstractIn this paper we propose a new approach to person re-identification using images and natural language descriptions. We propose a joint vision and language model based on CNN and LSTM architectures to match across the two modalities as well as to enrich visual examples for which there are no language descriptions. We also introduce new annotations in the form of natural language descriptions for two standard Re-ID benchmarks, namely CUHK03 and VIPeR. We perform experiments on these two datasets with techniques based on CNN, hand-crafted features as well as LSTM for analysing visual and natural description data. We investigate and demonstrate the advantages of using natural language descriptions compared to attributes as well as CNN compared to LSTM in the context of Re-ID. We show that the joint use of language and vision can significantly improve the state-of-the-art performance on standard Re-ID benchmarks. Fei Yan 0001, Josef Kittler, Krystian Mikolajczyk |
ICPR | 3 |
| 2018 | Binary Online Learned DescriptorsabstractWe propose a novel approach to generate a binary descriptor optimized for each image patch independently. The approach is inspired by the linear discriminant embedding that simultaneously increases inter and decreases intra class distances. A set of discriminative and uncorrelated binary tests is established from all possible tests in an offline training process. The patch adapted descriptors are then efficiently built online from a subset of features which lead to lower intra-class distances and thus, to a more robust descriptor. We perform experiments on three widely used benchmarks and demonstrate improvements in matching performance, and illustrate that per-patch optimization outperforms global optimization. Vassileios Balntas, Lilian Tang, Krystian Mikolajczyk |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | BreakingNews: Article Annotation by Image and Text ProcessingabstractBuilding upon recent Deep Neural Network architectures, current approaches lying in the intersection of Computer Vision and Natural Language Processing have achieved unprecedented breakthroughs in tasks like automatic captioning or image retrieval. Most of these learning methods, though, rely on large training sets of images associated with human annotations that specifically describe the visual content. In this paper we propose to go a step further and explore the more complex cases where textual descriptions are loosely related to the images. We focus on the particular domain of news articles in which the textual content often expresses connotative and ambiguous relations that are only suggested but not directly inferred from images. We introduce an adaptive CNN architecture that shares most of the structure for multiple tasks including source detection, article illustration and geolocation of articles. Deep Canonical Correlation Analysis is deployed for article illustration, and a new loss function based on Great Circle Distance is proposed for geolocation. Furthermore, we present BreakingNews, a novel dataset with approximately 100K news articles including images, text and captions, and enriched with heterogeneous meta-data (such as GPS coordinates and user comments). We show this dataset to be appropriate to explore all aforementioned problems, for which we provide a baseline performance using various Deep Learning architectures, and different representations of the textual and visual features. We report very promising results and bring to light several limitations of current state-of-the-art in this kind of domain, which we hope will help spur progress in the field. Arnau Ramisa, Fei Yan 0001, Francesc Moreno-Noguer, Krystian Mikolajczyk |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | HPatches: A Benchmark and Evaluation of Handcrafted and Learned Local DescriptorsabstractIn this paper, a novel benchmark is introduced for evaluating local image descriptors. We demonstrate limitations of the commonly used datasets and evaluation protocols, that lead to ambiguities and contradictory results in the literature. Furthermore, these benchmarks are nearly saturated due to the recent improvements in local descriptors obtained by learning from large annotated datasets. To address these issues, we introduce a new large dataset suitable for training and testing modern descriptors, together with strictly defined evaluation protocols in several tasks such as matching, retrieval and verification. This allows for more realistic, thus more reliable comparisons in different application scenarios. We evaluate the performance of several state-of-the-art descriptors and analyse their properties. We show that a simple normalisation of traditional hand-crafted descriptors is able to boost their performance to the level of deep learning based descriptors once realistic benchmarks are considered. Additionally we specify a protocol for learning and evaluating using cross validation. We show that when training state-of-the-art descriptors on this dataset, the traditional verification task is almost entirely saturated. Vassileios Balntas, Karel Lenc, Andrea Vedaldi, Krystian Mikolajczyk |
CVPR | 4 |
| 2017 | Higher-Order Occurrence Pooling for Bags-of-Words: Visual Concept DetectionabstractIn object recognition, the Bag-of-Words model assumes: i) extraction of local descriptors from images, ii) embedding the descriptors by a coder to a given visual vocabulary space which results in mid-level features, iii) extracting statistics from mid-level features with a pooling operator that aggregates occurrences of visual words in images into signatures, which we refer to as First-order Occurrence Pooling. This paper investigates higher-order pooling that aggregates over co-occurrences of visual words. We derive Bag-of-Words with Higher-order Occurrence Pooling based on linearisation of Minor Polynomial Kernel, and extend this model to work with various pooling operators. This approach is then effectively used for fusion of various descriptor types. Moreover, we introduce Higher-order Occurrence Pooling performed directly on local image descriptors as well as a novel pooling operator that reduces the correlation in the image signatures. Finally, First-, Second-, and Third-order Occurrence Pooling are evaluated given various coders and pooling operators on several widely used benchmarks. The proposed methods are compared to other approaches such as Fisher Vector Encoding and demonstrate improved results. Piotr Koniusz, Fei Yan 0001, Philippe Henri Gosselin, Krystian Mikolajczyk |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Learning local feature descriptors with triplets and shallow convolutional neural networks
Vassileios Balntas, Edgar Riba, Daniel Ponsa, Krystian Mikolajczyk |
BMVC | 4 |
| 2016 | Generating commentaries for tennis videosabstractWe present an approach to automatically generating verbal commentaries for tennis games. We introduce a novel application that requires a combination of techniques from computer vision, natural language processing and machine learning. A video sequence is first analysed using state-of-the-art computer vision methods to track the ball, fit the detected edges to the court model, track the players, and recognise their strokes. Based on the recognised visual attributes we formulate the tennis commentary generation problem in the framework of long short-term memory recurrent neural networks as well as structured SVM. In particular, we investigate pre-embedding of descriptive terms and loss function for LSTM. We introduce a new dataset of 633 annotated pairs of tennis videos and corresponding commentary. We perform an automatic as well as human based evaluation, and demonstrate that the proposed pre-embedding and loss function lead to substantially improved accuracy of the generated commentary. Fei Yan 0001, Krystian Mikolajczyk, Josef Kittler |
ICPR | 2 |
| 2016 | Hierarchical online domain adaptation of deformable part-based modelsabstractWe propose an online domain adaptation method for the deformable part-based model (DPM). The online domain adaptation is based on a two-level hierarchical adaptation tree, which consists of instance models in the leaf nodes and a category model at the root node. Moreover, combined with a multiple object tracking procedure (MOT), our proposal neither requires target-domain annotated data nor revisiting the source-domain data for performing the source-to-target domain adaptation of the DPM. From a practical point of view this means that, given a source-domain DPM and new video for training on a new domain without object annotations, our procedure outputs a new DPM adapted to the domain represented by the video. As proof-of-concept we apply our proposal to the challenging task of pedestrian detection. In this case, each instance model is an exemplar classifier trained online with only one pedestrian per frame. The pedestrian instances are collected by MOT and the hierarchical model is constructed dynamically according to the pedestrian trajectories. Our experimental results show that the adapted model achieves the accuracy of recent supervised domain adaptation methods (i.e., requiring manually annotated target-domain data), and improves the source model more than 10 percentage points. Jiaolong Xu, David Vázquez 0001, Krystian Mikolajczyk, Antonio M. López 0001 |
ICRA | 3 |
| 2016 | Deformable part-based tracking by coupled global and local correlation filters
Osman Akin, Erkut Erdem, Aykut Erdem, Krystian Mikolajczyk |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | BOLD - Binary online learned descriptor for efficient image matchingabstractIn this paper we propose a novel approach to generate a binary descriptor optimized for each image patch independently. The approach is inspired by the linear discriminant embedding that simultaneously increases inter and decreases intra class distances. A set of discriminative and uncorrelated binary tests is established from all possible tests in an offline training process. The patch adapted descriptors are then efficiently built online from a subset of tests which lead to lower intra class distances thus a more robust descriptor. A patch descriptor consists of two binary strings where one represents the results of the tests and the other indicates the subset of the patch-related robust tests that are used for calculating a masked Hamming distance. Our experiments on three different benchmarks demonstrate improvements in matching performance, and illustrate that per-patch optimization outperforms global optimization. Vassileios Balntas, Lilian Tang, Krystian Mikolajczyk |
CVPR | 3 |
| 2015 | Deep correlation for matching images and textabstractThis paper addresses the problem of matching images and captions in a joint latent space learnt with deep canonical correlation analysis (DCCA). The image and caption data are represented by the outputs of the vision and text based deep neural networks. The high dimensionality of the features presents a great challenge in terms of memory and speed complexity when used in DCCA framework. We address these problems by a GPU implementation and propose methods to deal with overfitting. This makes it possible to evaluate DCCA approach on popular caption-image matching benchmarks. We compare our approach to other recently proposed techniques and present state of the art results on three datasets. Fei Yan 0001, Krystian Mikolajczyk |
CVPR | 2 |
| 2015 | Full ranking as local descriptor for visual recognition: A comparison of distance metrics on sn
Chi-Ho Chan, Fei Yan 0001, Josef Kittler, Krystian Mikolajczyk |
Pattern Recognit. | 4 |
| 2014 | Leveraging High Level Visual Information for Matching Images and Captions
Fei Yan 0001, Krystian Mikolajczyk |
ACCV (1) | 2 |
| 2014 | Online Learning and Detection with Part-Based, Circulant StructureabstractCirculant Structure Kernel (CSK) has recently been introduced as a simple and extremely efficient tracking method. In this paper, we propose an extension of CSK that explicitly addresses partial occlusion problems which the original CSK suffers from. Our extension is based on a part-based scheme, which improves the robustness and localisation accuracy. Furthermore, we improve the robustness of CSK for long-term tracking by incorporating it into an online learning and detection framework. We provide an extensive comparison to eight recently introduced tracking methods. Our experimental results show that the proposed approach significantly improves the original CSK and provides state-of-the-art results when combined with online learning approach. Osman Akin, Krystian Mikolajczyk |
ICPR | 2 |
| 2014 | Improving Object Tracking with Voting from False Positive DetectionsabstractContext provides additional information in detection and tracking and several works proposed online trained trackers that make use of the context. However, the context is usually considered during tracking as items with motion patterns significantly correlated with the target. We propose a new approach that exploits context in tracking-by-detection and makes use of persistent false positive detections. True detection as well as repeated false positives act as pointers to the location of the target. This is implemented with a generalised Hough voting and incorporated into a state-of-the art online learning framework. The proposed method presents good performance in both speed and accuracy and it improves the current state of the art results in a challenging benchmark. Vassileios Balntas, Lilian Tang, Krystian Mikolajczyk |
ICPR | 3 |
| 2014 | Ranking Images Based on Aesthetic QualitiesabstractWe propose a novel approach for learning image representation based on qualitative assessments of visual aesthetics. It relies on a multi-node multi-state model that represents image attributes and their relations. The model is learnt from pair wise image preferences provided by annotators. To demonstrate the effectiveness we apply our approach to fashion image rating, i.e., comparative assessment of aesthetic qualities. Bag-of-features object recognition is used for the classification of visual attributes such as clothing and body shape in an image. The attributes and their relations are then assigned learnt potentials which are used to rate the images. Evaluation of the representation model has demonstrated a high performance rate in ranking fashion images. Aarushi Gaur, Krystian Mikolajczyk |
ICPR | 2 |
| 2014 | Robust Registration and Filtering for Moving Object Detection in Aerial VideosabstractIn this paper we present a multi-frame motion detection approach for aerial platforms with a two-folded contribution. First, we propose a novel image registration method, which can robustly cope with a large variety of aerial imagery. We show that it can benefit from a hardware accelerated implementation using graphic cards, allowing processing at high frame rate. Second, to handle the inaccuracy of the registration and sensor noise that result in false-alarms, we present an efficient filtering step to reduce incorrect motion hypotheses that arise from background substraction. We show that the proposed filtering significantly improves the precision of the motion detection while maintaining high recall. We introduce a new dataset for evaluating aerial surveillance systems, which will be made available for comparison. We evaluate the registration performance in terms of accuracy and speed as well as the filtering in terms of motion detection performance. Falk Schubert, Krystian Mikolajczyk |
ICPR | 2 |
| 2014 | Guest Editorial: Tracking, Detection and Segmentation
Richard Bowden, John P. Collomosse, Krystian Mikolajczyk |
Int. J. Comput. Vis. | 3 |
| 2014 | Automatic annotation of tennis games: An integration of audio, vision, and learning
Fei Yan 0001, Josef Kittler, David Windridge, William J. Christmas, Krystian Mikolajczyk, Stephen J. Cox, Qiang Huang 0006 |
Image Vis. Comput. | 5 |
| 2013 | A Global-Local Approach to Saliency Detection
Ahmed Boudissa, Joo Kooi Tan, Hyoungseop Kim, Seiji Ishikawa, Takashi Shinomiya, Krystian Mikolajczyk |
CAIP (2) | 6 |
| 2013 | Benchmarking GPU-Based Phase Correlation for Homography-Based Registration of Aerial Imagery
Falk Schubert, Krystian Mikolajczyk |
CAIP (2) | 2 |
| 2013 | Performance Evaluation of Image Filtering for Classification and Retrieval
Falk Schubert, Krystian Mikolajczyk |
ICPRAM | 2 |
| 2013 | Comparison of mid-level feature coding approaches and pooling strategies in visual concept detection
Piotr Koniusz, Fei Yan 0001, Krystian Mikolajczyk |
Comput. Vis. Image Underst. | 3 |
| 2013 | A Robust and Scalable Visual Category and Action Recognition System Using Kernel Discriminant Analysis With Spectral RegressionabstractVisual concept detection and action recognition are one of the most important tasks in content-based multimedia information retrieval (CBMIR) technology. It aims at annotating images using a vocabulary defined by a set of concepts of interest including scenes types (mountains, snow, etc.) or human actions (phoning, playing instrument). This paper describes our system in the ImageCLEF@ICPR10, Pascal VOC 08 Visual Concept Detection and Pascal VOC 10 Action Recognition Challenges. The proposed system ranked first in these large-scale tasks when evaluated independently by the organizers. The proposed system involves state-of-the-art local descriptor computation, vector quantization via clustering, structured scene or object representation via localized histograms of vector codes, similarity measure for kernel construction and classifier learning. The main novelty is the classifier-level and kernel-level fusion using Kernel Discriminant Analysis and Spectral Regression (SR-KDA) with RBF Chi-Squared kernels obtained from various image descriptors. The distinctiveness of the proposed method is also assessed experimentally using a video benchmark: the Mediamill Challenge along with benchmarks from ImageCLEF@ICPR10, Pascal VOC 10 and Pascal VOC 08. From the experimental results, it can be derived that the presented system consistently yields significant performance gains when compared with the state-of-the art methods. The other strong point is the introduction of SR-KDA in the classification stage where the time complexity scales linearly with respect to the number of concepts and the main computational complexity is independent of the number of categories. Muhammad Atif Tahir, Fei Yan 0001, Piotr Koniusz, Muhammad Awais 0001, Mark Barnard, Krystian Mikolajczyk, Ahmed Bouridane, Josef Kittler |
IEEE Trans. Multim. | 6 |
| 2012 | Evaluation of local detectors and descriptors for fast feature matching
Ondrej Miksik, Krystian Mikolajczyk |
ICPR | 2 |
| 2012 | Automatic annotation of court games with structured output learning
Fei Yan 0001, Josef Kittler, Krystian Mikolajczyk, David Windridge |
ICPR | 3 |
| 2012 | Non-Sparse Multiple Kernel Fisher Discriminant Analysis
Fei Yan 0001, Josef Kittler, Krystian Mikolajczyk, Muhammad Atif Tahir |
J. Mach. Learn. Res. | 3 |
| 2012 | Tracking-Learning-DetectionabstractThis paper investigates long-term tracking of unknown objects in a video stream. The object is defined by its location and extent in a single frame. In every frame that follows, the task is to determine the object's location and extent or indicate that the object is not present. We propose a novel tracking framework (TLD) that explicitly decomposes the long-term tracking task into tracking, learning, and detection. The tracker follows the object from frame to frame. The detector localizes all appearances that have been observed so far and corrects the tracker if necessary. The learning estimates the detector's errors and updates it to avoid these errors in the future. We study how to identify the detector's errors and learn from them. We develop a novel learning method (P-N learning) which estimates the errors by a pair of "experts": (1) P-expert estimates missed detections, and (2) N-expert estimates false alarms. The learning process is modeled as a discrete dynamical system and the conditions under which the learning guarantees improvement are found. We describe our real-time implementation of the TLD framework and the P-N learning. We carry out an extensive quantitative evaluation which shows a significant improvement over state-of-the-art approaches. Zdenek Kalal, Krystian Mikolajczyk, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Augmented Kernel Matrix vs Classifier Fusion for Object RecognitionabstractAugmented Kernel Matrix (AKM) has recently been proposed to accommodate for the fact that a single training example may have different importance in different feature spaces, in contrast to Multiple Kernel Learning (MKL) that assigns the same weight to all examples in one feature space.However, the AKM approach is limited to small datasets due to its memory requirements.An alternative way to fuse information from different feature channels is classifier fusion (ensemble methods).There is a significant amount of work on linear programming formulations of classifier fusion (CF) in the case of binary classification.In this paper we derive primal and dual of AKM to draw its correspondence with CF.We propose a multiclass extension of binary ν-LPBoost, which learns the contribution of each class in each feature channel.Existing approaches of CF promote sparse features combinations, due to regularization based on 1 -norm, and lead to a selection of a subset of feature channels, which is not good in case of informative channels.We also generalize existing CF formulations to arbitrary p -norm for binary and multiclass problems which results in more effective use of complementary information.We carry out an extensive comparison and show that the proposed nonlinear CF schemes outperform its sparse counterpart as well as state-of-the-art MKL approaches. Muhammad Awais 0001, Fei Yan 0001, Krystian Mikolajczyk, Josef Kittler |
BMVC | 3 |
| 2011 | Spatial Coordinate Coding to reduce histogram representations, Dominant Angle and Colour Pyramid MatchabstractSpatial Pyramid Match lies at a heart of modern object category recognition systems. Once image descriptors are expressed as histograms of visual words, they are further deployed across spatial pyramid with coarse-to-fine spatial location grids. However, such representation results in extreme histogram vectors of 200K or more elements increasing computational and memory requirements. This paper investigates alternative ways of introducing spatial information during formation of histograms. Specifically, we propose to apply spatial location information at a descriptor level and refer to it as Spatial Coordinate Coding. Alternatively, x, y, radius, or angle is used to perform semi-coding. This is achieved by adding one of the spatial components at the descriptor level whilst applying Pyramid Match to another. Lastly, we demonstrate that Pyramid Match can be applied robustly to other measurements: Dominant Angle and Colour. We demonstrate state-of-the art results on two datasets with means of Soft Assignment and Sparse Coding. Piotr Koniusz, Krystian Mikolajczyk |
ICIP | 2 |
| 2011 | Soft assignment of visual words as Linear Coordinate Coding and optimisation of its reconstruction errorabstractVisual Word Uncertainty also referred to as Soft Assignment is a well established technique for representing images as histograms by flexible assignment of image descriptors to a visual vocabulary. Recently, an attention of the community dealing with the object category recognition has been drawn to Linear Coordinate Coding methods. In this work, we focus on Soft Assignment as it yields good results amidst competitive methods. We show that one can take two views on Soft Assignment: an approach derived from Gaussian Mixture Model or special case of Linear Coordinate Coding. The latter view helps us propose how to optimise smoothing factor of Soft Assignment in a way that minimises descriptor reconstruction error and maximises classification performance. In turns, this renders tedious cross-validation towards establishing this parameter unnecessary and yields it a handy technique. We demonstrate state-of-the-art performance of such optimised assignment on two image datasets and several types of descriptors. Piotr Koniusz, Krystian Mikolajczyk |
ICIP | 2 |
| 2011 | Novel Fusion Methods for Pattern Recognition
Muhammad Awais 0001, Fei Yan 0001, Krystian Mikolajczyk, Josef Kittler |
ECML/PKDD (1) | 3 |
| 2011 | An evaluation of bags-of-words and spatio-temporal shapes for action recognitionabstractBags-of-visual-Words (BoW) and Spatio-Temporal Shapes (STS) are two very popular approaches for action recognition from video. The former (BoW) is an un-structured global representation of videos which is built using a large set of local features. The latter (STS) uses a single feature located on a region of interest (where the actor is) in the video. Despite the popularity of these methods, no comparison between them has been done. Also, given that BoW and STS differ intrinsically in terms of context inclusion and globality/locality of operation, an appropriate evaluation framework has to be designed carefully. This paper compares these two approaches using four different datasets with varied degree of space-time specificity of the actions and varied relevance of the contextual background. We use the same local feature extraction method and the same classifier for both approaches. Further to BoW and STS, we also evaluated novel variations of BoW constrained in time or space. We observe that the STS approach leads to better results in all datasets whose background is of little relevance to action classification. Teófilo Emídio de Campos, Mark Barnard, Krystian Mikolajczyk, Josef Kittler, Fei Yan 0001, William J. Christmas, David Windridge |
WACV | 3 |
| 2011 | Action recognition with appearance-motion features and fast search trees
Krystian Mikolajczyk, Hirofumi Uemura |
Comput. Vis. Image Underst. | 1 |
| 2011 | Learning Linear Discriminant Projections for Dimensionality Reduction of Image DescriptorsabstractIn this paper, we present Linear Discriminant Projections (LDP) for reducing dimensionality and improving discriminability of local image descriptors. We place LDP into the context of state-of-the-art discriminant projections and analyze its properties. LDP requires a large set of training data with point-to-point correspondence ground truth. We demonstrate that training data produced by a simulation of image transformations leads to nearly the same results as the real data with correspondence ground truth. This makes it possible to apply LDP as well as other discriminant projection approaches to the problems where the correspondence ground truth is not available, such as image categorization. We perform an extensive experimental evaluation on standard data sets in the context of image matching and categorization. We demonstrate that LDP enables significant dimensionality reduction of local descriptors and performance increases in different applications. The results improve upon the state-of-the-art recognition performance with simultaneous dimensionality reduction from 128 to 30. Hongping Cai, Krystian Mikolajczyk, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Learning weights for codebook in image classification and retrievalabstractThis paper presents a codebook learning approach for image classification and retrieval. It corresponds to learning a weighted similarity metric to satisfy that the weighted similarity between the same labeled images is larger than that between the differently labeled images with largest margin. We formulate the learning problem as a convex quadratic programming and adopt alternating optimization to solve it efficiently. Experiments on both synthetic and real datasets validate the approach. The codebook learning improves the performance, in particular in the case where the number of training examples is not sufficient for large size codebook. Hongping Cai, Fei Yan 0001, Krystian Mikolajczyk |
CVPR | 3 |
| 2010 | P-N learning: Bootstrapping binary classifiers by structural constraintsabstractThis paper shows that the performance of a binary classifier can be significantly improved by the processing of structured unlabeled data, i.e. data are structured if knowing the label of one example restricts the labeling of the others. We propose a novel paradigm for training a binary classifier from labeled and unlabeled examples that we call P-N learning. The learning process is guided by positive (P) and negative (N) constraints which restrict the labeling of the unlabeled set. P-N learning evaluates the classifier on the unlabeled data, identifies examples that have been classified in contradiction with structural constraints and augments the training set with the corrected samples in an iterative process. We propose a theory that formulates the conditions under which P-N learning guarantees improvement of the initial classifier and validate it on synthetic and real data. P-N learning is applied to the problem of on-line learning of object detector during tracking. We show that an accurate object detector can be learned from a single example and an unlabeled video sequence where the object may occur. The algorithm is compared with related approaches and state-of-the-art is achieved on a variety of objects (faces, pedestrians, cars, motorbikes and animals). Zdenek Kalal, Jiri Matas, Krystian Mikolajczyk |
CVPR | 3 |
| 2010 | lp norm multiple kernel Fisher discriminant analysis for object and image categorisationabstractIn this paper, we generalise multiple kernel Fisher discriminant analysis (MK-FDA) such that the kernel weights can be regularised with an ℓpnorm for any p ≥ 1, in contrast to existing MK-FDA that uses either l1 or l2 norm. We present formulations for both binary and multiclass cases and solve the associated optimisation problems efficiently with semi-infinite programming. We show on three object and image categorisation benchmarks that by learning the intrinsic sparsity of a given set of base kernels using a validation set, the proposed ℓpMK-FDA outperforms its fixed-norm counterparts, and is capable of producing state-of-the-art performance. Moreover, we show that our ℓpMK-FDA outperforms the ℓpmultiple kernel support vector machine (ℓpMK-SVM) which has been recently proposed. Based on this observation and our experience with single kernel FDA and SVM, we argue that the almost century-old FDA is still a strong competitor of the popular SVM. Fei Yan 0001, Krystian Mikolajczyk, Mark Barnard, Hongping Cai, Josef Kittler |
CVPR | 2 |
| 2010 | Face-TLD: Tracking-Learning-Detection applied to facesabstractA novel system for long-term tracking of a human face in unconstrained videos is built on Tracking-Learning-Detection (TLD) approach. The system extends TLD with the concept of a generic detector and a validator which is designed for real-time face tracking resistent to occlusions and appearance changes. The off-line trained detector localizes frontal faces and the online trained validator decides which faces correspond to the tracked subject. Several strategies for building the validator during tracking are quantitatively evaluated. The system is validated on a sitcom episode (23 min.) and a surveillance (8 min.) video. In both cases the system detects-tracks the face and automatically learns a multi-view model from a single frontal example and an unlabeled video. Zdenek Kalal, Krystian Mikolajczyk, Jiri Matas |
ICIP | 2 |
| 2010 | Feature Pairs Connected by Lines for Object RecognitionabstractIn this paper we exploit image edges and segmentation maps to build features for object category recognition. We build a parametric line based image approximation to identify the dominant edge structures. Line ends are used as features described by histograms of gradient orientations. We then form descriptors based on connected line ends to incorporate weak topological constraints which improve their discriminative power. Using point pairs connected by an edge assures higher repeatability than a random pair of points or edges. The results are compared with state-of-the-art, and show significant improvement on challenging recognition benchmark Pascal VOC 2007. Kernel based fusion is performed to emphasize the complementary nature of our descriptors with respect to the state-of-the-art features. Muhammad Awais 0001, Krystian Mikolajczyk |
ICPR | 2 |
| 2010 | Forward-Backward Error: Automatic Detection of Tracking FailuresabstractThis paper proposes a novel method for tracking failure detection. The detection is based on the Forward-Backward error, i.e. the tracking is performed forward and backward in time and the discrepancies between these two trajectories are measured. We demonstrate that the proposed error enables reliable detection of tracking failures and selection of reliable trajectories in video sequences. We demonstrate that the approach is complementary to commonly used normalized cross-correlation (NCC). Based on the error, we propose a novel object tracker called Median Flow. State-of-the-art performance is achieved on challenging benchmark video sequences which include non-rigid objects. Zdenek Kalal, Krystian Mikolajczyk, Jiri Matas |
ICPR | 2 |
| 2010 | On a Quest for Image Descriptors Based on Unsupervised Segmentation MapsabstractThis paper investigates segmentation-based image descriptors for object category recognition. In contrast to commonly used interest points the proposed descriptors are extracted from pairs of adjacent regions given by a segmentation method. In this way we exploit semi-local structural information from the image. We propose to use the segments as spatial bins for descriptors of various image statistics based on gradient, colour and region shape. Proposed descriptors are validated on standard recognition benchmarks. Results show they outperform state-of-the-art reference descriptors with 5.6x less data and achieve comparable results to them with 8.6x less data. The proposed descriptors are complementary to SIFT and achieve state-of-the-art results when combined together within a kernel based classifier. Piotr Koniusz, Krystian Mikolajczyk |
ICPR | 2 |
| 2010 | The University of Surrey Visual Concept Detection System at ImageCLEF@ICPR: Working NotesabstractVisual concept detection is one of the most important tasks in image and video indexing. This paper describes our system in the ImageCLEF@ICPR Visual Concept Detection Task which ranked first for large-scale visual concept detection tasks in terms of Equal Error Rate (EER) and Area under Curve (AUC) and ranked third in terms of hierarchical measure. The presented approach involves state-of-the-art local descriptor computation, vector quantisation via clustering, structured scene or object representation via localised histograms of vector codes, similarity measure for kernel construction and classifier learning. The main novelty is the classifier-level and kernel-level fusion using Kernel Discriminant Analysis with RBF/Power Chi-Squared kernels obtained from various image descriptors. For 32 out of 53 individual concepts, we obtain the best performance of all 12 submissions to this task. Muhammad Atif Tahir, Fei Yan 0001, Mark Barnard, Muhammad Awais 0001, Krystian Mikolajczyk, Josef Kittler |
ICPR | 5 |
| 2009 | Segmentation Based Interest Points and Evaluation of Unsupervised Image Segmentation MethodsabstractThis paper investigates segmentation based interest points for matching and recognition.We propose two simple methods for extracting features from the segmentation maps, which focus on the boundaries and centres of the gravity of the segments.In addition, this can be considered a novel approach for evaluating unsupervised image segmentation algorithms.Former evaluations aim at estimating segmentation quality by how well resulting segments adhere to the contours separating ground-truth foregrounds from backgrounds and therefore explicitly focus on particular objects of interest.In contrast, we propose to measure the robustness of segmentations by the repeatability of features extracted from segments on images related by various geometric and photometric transformations.Further, our evaluation provides a new insight into suitability of the segmentation methods for generating local features for image retrieval or recognition.Several segmentation methods are evaluated and compared to state-of-the art interest point detectors using the repeatability criteria as well as standard matching and recognition benchmarks. Piotr Koniusz, Krystian Mikolajczyk |
BMVC | 2 |
| 2009 | Non-sparse Multiple Kernel Learning for Fisher Discriminant AnalysisabstractWe consider the problem of learning a linear combination of pre-specified kernel matrices in the Fisher discriminant analysis setting. Existing methods for such a task impose an ¿1norm regularisation on the kernel weights, which produces sparse solution but may lead to loss of information. In this paper, we propose to use ¿2norm regularisation instead. The resulting learning problem is formulated as a semi-infinite program and can be solved efficiently. Through experiments on both synthetic data and a very challenging object recognition benchmark, the relative advantages of the proposed method and its ¿1counterpart are demonstrated, and insights are gained as to how the choice of regularisation norm should be made. Fei Yan 0001, Josef Kittler, Krystian Mikolajczyk, Muhammad Atif Tahir |
ICDM | 3 |
| 2009 | A hands-on approach to high-dynamic-range and superresolution fusionabstractThis paper discusses a new framework to enhance image and video quality. Recent advances in high-dynamic-range image fusion and superresolution make it possible to extend the intensity range or to increase the resolution of the image beyond the limitations of the sensor. In this paper, we propose a new way to combine both of these fusion methods in a two-stage scheme. To achieve robust image enhancement in practical application scenarios, we adapt state-of-the-art methods for automatic photometric camera calibration, controlled image acquisition, image fusion and tone mapping. With respect to high-dynamic-range reconstruction, we show that only two input images can sufficiently capture the dynamic range of the scene. The usefulness and performance of this system is demonstrated on images taken with various types of cameras. Falk Schubert, Klaus Schertler, Krystian Mikolajczyk |
WACV | 3 |
| 2008 | Learning Linear Discriminant Projections for Dimensionality Reduction of Image DescriptorsabstractIn this paper we present Linear Discriminant Projections (LDP) for reducing dimensionality and improving discriminability of local image descriptors. We place LDP into the context of state-of-the-art discriminant projections and analyze its properties. LDP requires large set of training data with point-to-point correspondence ground truth. We demonstrate that a training data produced by a simulation of image transformations leads to nearly the same results as the real data with correspondence ground truth. This makes it possible to apply LDP as well as other discriminant projection approaches to the problems where the correspondence ground truth is not available such as image categorization. We perform an extensive experimental evaluation on standard datasets in the context of image matching and categorization. We demonstrate that LDP enables significant dimensionality reduction of local descriptors and performance increases in different applications. The results improve upon the state-of-the-art recognition performance with simultaneous dimensionality reduction from 128 to 30. Hongping Cai, Krystian Mikolajczyk, Jiri Matas |
BMVC | 2 |
| 2008 | Weighted Sampling for Large-Scale BoostingabstractThis paper addresses the problem of learning from very large databases where batch learning is impractical or even infeasible. Bootstrap is a popular technique applicable in such situations. We show that sampling strategy used for bootstrapping has a significant impact on the resulting classifier performance. We design a new general sampling strategy ”quasi-random weighted sampling + trimming ” (QWS+) that includes well established strategies as special cases. The QWS+ approach minimizes the variance of hypothesis error estimate and leads to significant improvement in performance compared to standard sampling techniques. The superior performance is demonstrated on several problems including profile and frontal face detection. 1 Zdenek Kalal, Jiri Matas, Krystian Mikolajczyk |
BMVC | 3 |
| 2008 | Combining High-Resolution Images With Low-Quality VideosabstractRecently a lot of research has been directed towards the question: What can be done with "brute-force vision" using huge amounts of data?Image retrieval methods have been shown to succeed on collections of images with sizes over a million.Various applications such as object recognition, 3D geometrical arrangment of images showing the same scene or inferring missing image regions can benefit from large image databases.Motivated by this research we propose an alternative use of image information stored in large pools like the internet.Given an input video, we can utilize corresponding still images stored at much better quality to improve the overall quality of the video.A hybrid superresolution scheme is applied to smoothly incorporate the high-frequency components.On those areas where hallucination of details fails, a standard MAP-estimation of the high-resolution image is performed.The performance is demonstrated on real data examples. Falk Schubert, Krystian Mikolajczyk |
BMVC | 2 |
| 2008 | Feature Tracking and Motion Compensation for Action RecognitionabstractThis paper discusses an approach to human action recognition via local fea-ture tracking and robust estimation of background motion. The main contri-bution is a robust feature extraction algorithm based on KLT tracker and SIFT as well as a method for estimating dominant planes in the scene. Multiple in-terest point detectors are used to provide large number of features for every frame. The motion vectors for the features are estimated using optical flow and SIFT based matching. The features are combined with image segmenta-tion to estimate dominant homographies, and then separated into static and moving ones regardless the camera motion. The action recognition approach can handle camera motion, zoom, human appearance variations, background clutter and occlusion. The motion compensation shows very good accuracy on a number of test sequences. The recognition system is extensively com-pared to state-of-the art action recognition methods and the results are im-proved. 1 Hirofumi Uemura, Seiji Ishikawa, Krystian Mikolajczyk |
BMVC | 3 |
| 2008 | Action recognition with motion-appearance vocabulary forestabstractIn this paper we propose an approach for action recognition based on a vocabulary forest of local motion-appearance features. Large numbers of features with associated motion vectors are extracted from action data and are represented by many vocabulary trees. Features from a query sequence are matched to the trees and vote for action categories and their locations. Large number of trees make the process efficient and robust. The system is capable of simultaneous categorization and localization of actions using only a few frames per sequence. The approach obtains excellent performance on standard action recognition sequences. We perform large scale experiments on 17 challenging real action categories from Olympic Games1. We demonstrate the robustness of our method to appearance variations, camera motion, scale change, asymmetric actions, background clutter and occlusion. Krystian Mikolajczyk, Hirofumi Uemura |
CVPR | 1 |
| 2007 | Improving Descriptors for Fast Tree Matching by Optimal Linear ProjectionabstractIn this paper we propose to transform an image descriptor so that nearest neighbor (NN) search for correspondences becomes the optimal matching strategy under the assumption that inter-image deviations of corresponding descriptors have Gaussian distribution. The Euclidean NN in the transformed domain corresponds to the NN according to a truncated Mahalanobis metric in the original descriptor space. We provide theoretical justification for the proposed approach and show experimentally that the transformation allows a significant dimensionality reduction and improves matching performance of a state-of-the art SIFT descriptor. We observe consistent improvement in precision-recall and speed of fast matching in tree structures at the expense of little overhead for projecting the descriptors into transformed space. In the context of SIFT vs. transformed M- SIFT comparison, tree search structures are evaluated according to different criteria and query types. All search tree experiments confirm that transformed M-SIFTperforms better than the original SIFT. Krystian Mikolajczyk, Jiri Matas |
ICCV | 1 |
| 2006 | Efficient Clustering and Matching for Object Class RecognitionabstractIn this paper we address the problem of building object class representations based on local features and fast matching in a large database. We propose an efficient algorithm for hierarchical agglomerative clustering. We examine different agglomerative and partitional clustering strategies and compare the quality of obtained clusters. Our combination of partitional-agglomerative clustering gives significant improvement in terms of efficiency while maintaining the same quality of clusters. We also propose a method for building data structures for fast matching in high dimensional feature spaces. These improvements allow to deal with large sets of training data typically used in recognition of multiple object classes. 1 Bastian Leibe, Krystian Mikolajczyk, Bernt Schiele |
BMVC | 2 |
| 2006 | Segmentation Based Multi-Cue Integration for Object DetectionabstractThis paper proposes a novel method for integrating multiple local cues, i.e. local region detectors as well as descriptors, in the context of object detection. Rather than to fuse the outputs of several distinct classifiers in a fixed setup, our approach implements a highly adaptable integration scheme, flexibly recombining the contributions of all individual cues depending on their explanatory power for each new test image. The key idea behind our approach is to integrate the cues over an estimated top-down segmentation, which allows to quantify how much each of them contributed to the object hypothesis. By combining those contributions on a per-pixel level, our approach ensures that each cue is only used for object regions for which it is confident and that potential correlations are effectively factored out. Experimental results on several benchmark data sets show that the proposed multi-cue combination scheme significantly increases detection performance compared to any of its constituent cues alone. Moreover, it provides an interesting evaluation tool to analyze the complementarity of local feature detectors and descriptors. 1 Bastian Leibe, Krystian Mikolajczyk, Bernt Schiele |
BMVC | 2 |
| 2006 | Multiple Object Class Detection with a Generative ModelabstractIn this paper we propose an approach capable of simultaneous recognition and localization of multiple object classes using a generative model. A novel hierarchical representation allows to represent individual images as well as various objects classes in a single, scale and rotation invariant model. The recognition method is based on a codebook representation where appearance clusters built from edge based features are shared among several object classes. A probabilistic model allows for reliable detection of various objects in the same image. The approach is highly efficient due to fast clustering and matching methods capable of dealing with millions of high dimensional features. The system shows excellent performance on several object categories over a wide range of scales, in-plane rotations, background clutter, and partial occlusions. The performance of the proposed multi-object class detection approach is competitive to state of the art approaches dedicated to a single object class recognition problem. Krystian Mikolajczyk, Bastian Leibe, Bernt Schiele |
CVPR (1) | 1 |
| 2005 | An Evaluation of Local Shape-Based Features for Pedestrian DetectionabstractPedestrian detection in real world scenes is a challenging problem. In recent years a variety of approaches have been proposed, and impressive results have been reported on a variety of databases. This paper systematically evaluates (1) various local shape descriptors, namely Shape Context and Local Chamfer descriptor and (2) four different interest point detectors for the detection of pedestrians. Those results are compared to the standard global Chamfer matching approach. A main result of the paper is that Shape Context trained on real edge images rather than on clean pedestrian silhouettes combined with the Hessian-Laplace detector outperforms all other tested approaches. 1 Edgar Seemann, Bastian Leibe, Krystian Mikolajczyk, Bernt Schiele |
BMVC | 3 |
| 2005 | Local Features for Object Class RecognitionabstractIn this paper, we compare the performance of local detectors and descriptors in the context of object class recognition. Recently, many detectors/descriptors have been evaluated in the context of matching as well as invariance to viewpoint changes (Mikolajczyk and Schmid, 2004). However, it is unclear if these results can be generalized to categorization problems, which require different properties of features. We evaluate 5 state-of-the-art scale invariant region detectors and 5 descriptors. Local features are computed for 20 object classes and clustered using hierarchical agglomerative clustering. We measure the quality of appearance clusters and location distributions using entropy as well as precision. We also measure how the clusters generalize from training set to novel test data. Our results indicate that attended SIFT descriptors (Mikolajczyk and Schmid, 2005) computed on Hessian-Laplace regions perform best. Second score is obtained by salient regions (Kadir and Brady, 2001). The results also show that these two detectors provide complementary features. The new detectors/descriptors significantly improve the performance of a state-of-the art recognition approach (Leibe, et al., 2005) in pedestrian detection task Krystian Mikolajczyk, Bastian Leibe, Bernt Schiele |
ICCV | 1 |
| 2005 | A Comparison of Affine Region Detectors
Krystian Mikolajczyk, Tinne Tuytelaars, Cordelia Schmid, Andrew Zisserman, Jiri Matas, Frederik Schaffalitzky, Timor Kadir, Luc Van Gool |
Int. J. Comput. Vis. | 1 |
| 2005 | A Performance Evaluation of Local DescriptorsabstractIn this paper, we compare the performance of descriptors computed for local interest regions, as, for example, extracted by the Harris-Affine detector [Mikolajczyk, K and Schmid, C, 2004]. Many different descriptors have been proposed in the literature. It is unclear which descriptors are more appropriate and how their performance depends on the interest region detector. The descriptors should be distinctive and at the same time robust to changes in viewing conditions as well as to errors of the detector. Our evaluation uses as criterion recall with respect to precision and is carried out for different image transformations. We compare shape context [Belongie, S, et al., April 2002], steerable filters [Freeman, W and Adelson, E, Setp. 1991], PCA-SIFT [Ke, Y and Sukthankar, R, 2004], differential invariants [Koenderink, J and van Doorn, A, 1987], spin images [Lazebnik, S, et al., 2003], SIFT [Lowe, D. G., 1999], complex filters [Schaffalitzky, F and Zisserman, A, 2002], moment invariants [Van Gool, L, et al., 1996], and cross-correlation for different types of interest regions. We also propose an extension of the SIFT descriptor and show that it outperforms the original method. Furthermore, we observe that the ranking of the descriptors is mostly independent of the interest region detector and that the SIFT-based descriptors perform best. Moments and steerable filters show the best performance among the low dimensional descriptors. Krystian Mikolajczyk, Cordelia Schmid |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Human Detection Based on a Probabilistic Assembly of Robust Part Detectors
Krystian Mikolajczyk, Cordelia Schmid, Andrew Zisserman |
ECCV (1) | 1 |
| 2004 | Scale & Affine Invariant Interest Point Detectors
Krystian Mikolajczyk, Cordelia Schmid |
Int. J. Comput. Vis. | 1 |
| 2003 | Shape recognition with edge-based featuresabstractIn this paper we describe an approach to recognizing poorly textured objects, that may contain holes and tubular parts, in cluttered scenes under arbitrary viewing conditions. To this end we develop a number of novel components. First, we introduce a new edge-based local feature detector that is invariant to similarity transformations. The features are localized on edges and a neighbourhood is estimated in a scale invariant manner. Second, the neighbourhood descriptor computed for foreground features is not affected by background clutter, even if the feature is on an object boundary. Third, the descriptor generalizes Lowe’s SIFT method [12] to edges. An object model is learnt from a single training image. The object is then recognized in new images in a series of steps which apply progressively tighter geometric restrictions. A final contribution of this work is to allow sufficient flexibility in the geometric representation that objects in the same visual class can be recognized. Results are demonstrated for various object classes including bikes and rackets. Krystian Mikolajczyk, Andrew Zisserman, Cordelia Schmid |
BMVC | 1 |
| 2003 | A performance evaluation of local descriptorsabstractIn this paper we compare the performance of interest point descriptors. Many different descriptors have been proposed in the literature. However, it is unclear which descriptors are more appropriate and how their performance depends on the interest point detector. The descriptors should be distinctive and at the same time robust to changes in viewing conditions as well as to errors of the point detector. Our evaluation uses as criterion detection rate with respect to false positive rate and is carried out for different image transformations. We compare SIFT descriptors (Lowe, 1999), steerable filters (Freeman and Adelson, 1991), differential invariants (Koenderink ad van Doorn, 1987), complex filters (Schaffalitzky and Zisserman, 2002), moment invariants (Van Gool et al., 1996) and cross-correlation for different types of interest points. In this evaluation, we observe that the ranking of the descriptors does not depend on the point detector and that SIFT descriptors perform best. Steerable filters come second ; they can be considered a good choice given the low dimensionality. Krystian Mikolajczyk, Cordelia Schmid |
CVPR (2) | 1 |
| 2003 | Face Detection and Tracking in a Video by Propagating Detection ProbabilitiesabstractThis paper presents a new probabilistic method for detecting and tracking multiple faces in a video sequence. The proposed method integrates the information of face probabilities provided by the detector and the temporal information provided by the tracker to produce a method superior to the available detection and tracking methods. The three novel contributions of the paper are: 1) Accumulation of probabilities of detection over a sequence. This leads to coherent detection over time and, thus, improves detection results. 2) Prediction of the detection parameters which are position, scale, and pose. This guarantees the accuracy of accumulation as well as a continuous detection. 3) The representation of pose is based on the combination of two detectors, one for frontal views and one for profiles. Face detection is fully automatic and is based on the method developed by Schneiderman and Kanade (2000). It uses local histograms of wavelet coefficients represented with respect to a coordinate frame fixed to the object. A probability of detection is obtained for each image position and at several scales and poses. The probabilities of detection are propagated over time using a Condensation filter and factored sampling. Prediction is based on a zero order model for position, scale, and pose; update uses the probability maps produced by the detection routine. The proposed method can handle multiple faces, appearing/disappearing faces as well as changing scale and pose. Experiments carried out on a large number of sequences taken from commercial movies and the Web show a clear improvement over the results of frame-based detection (in which the detector is applied to each frame of the video sequence). Ragini Choudhury, Cordelia Schmid, Krystian Mikolajczyk |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2002 | An Affine Invariant Interest Point Detector
Krystian Mikolajczyk, Cordelia Schmid |
ECCV (1) | 1 |
| 2001 | Face detection in a video sequence - a temporal approachabstractThis paper presents a new method for detecting faces in a video sequence where detection is not limited to frontal views. The three novel contributions of the paper are : (1) Accumulation of probabilities of detection over a sequence. This allows to obtain a coherent detection over time as well as independence from thresholds. (2) Prediction of the detection parameters which are position, scale and pose. This guarantees the accuracy of accumulation as well as a continuous detection. (3) The way pose is represented. The representation is based on the combination of two detectors, one for frontal views and one for profiles. Face detection is fully automatic and is based on the method developed by Schneiderman [13]. It uses local histograms of wavelet coefficients represented with respect to a coordinate frame fixed to the object. A probability of detection is obtained for each image position, several scales and the two detectors. The probabilities of detection are propagated over time using a Condensation filter and factored sampling. Prediction is based on a zero order model for position, scale and "pose"; update uses the probability maps produced by the detection routine. Experiments show a clear improvement over frame-based detection results. Krystian Mikolajczyk, Ragini Choudhury, Cordelia Schmid |
CVPR (2) | 1 |
| 2001 | Indexing Based on Scale Invariant Interest Points
Krystian Mikolajczyk, Cordelia Schmid |
ICCV | 1 |
| 1993 | A Test-Bed for Computer-Assisted Fusion of Multi-Modality Medical Images
Krystian Mikolajczyk, J. Owczarczyk, W. Recko |
CAIP | 1 |