EDBT 2026 Demo / reviewers in the wild / expert
Miroslaw Bober
dblp:74/623 · also Miroslaw Z. Bober
· DBLP profile ↗
46ranked-venue papers
11as first author
8since 2021 · last 2024
0000-0001-9484-9125ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 21 · 7 first-author · 7 since 2021Systems, architecture and hardware · 2Computer networks · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Segmentation and scene understanding · 57% Probabilistic and Bayesian machine learning · 14% Graph learning · 12% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 100% | |
| Computer graphics and multimedia
5 papers |
Multimedia analysis and retrieval · 70% Image and video processing · 26% Image and video coding · 4% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
scene graph generation |
2.6 | 4 | 2024 | Importance Weighted Structure Learning for Scene Graph Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2024 Constrained Structure Learning for Scene Graph Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Neural Belief Propagation for Scene Graph Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Information retrieval
image retrieval |
0.9 | 2 | 2021 | ACTNET: End-to-End Learning of Feature Activations and Multi-stream Aggregation for Effective Instance Image Retrieval · Int. J. Comput. Vis. 2021 REMAP: Multi-Layer Entropy-Guided Pooling of Dense CNN Features for Image Retrieval · IEEE Trans. Image Process. 2019 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.8 | 1 | 2024 | Importance Weighted Structure Learning for Scene Graph Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Graph learning › high-order interaction
higher-order dependency modeling |
0.7 | 1 | 2023 | Neural Belief Propagation for Scene Graph Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Information retrieval › image retrieval
instance retrieval |
0.5 | 1 | 2021 | ACTNET: End-to-End Learning of Feature Activations and Multi-stream Aggregation for Effective Instance Image Retrieval · Int. J. Comput. Vis. 2021 |
Computer vision › 3D vision › local feature descriptor
binary descriptor |
0.3 | 1 | 2017 | Fast, Compact, and Discriminative: Evaluation of Binary Descriptors for Mobile Applications · IEEE Trans. Multim. 2017 |
Computer vision › 3D vision
local feature descriptor |
0.3 | 1 | 2017 | Fast, Compact, and Discriminative: Evaluation of Binary Descriptors for Mobile Applications · IEEE Trans. Multim. 2017 |
Multimedia analysis and retrieval
image retrieval |
0.3 | 1 | 2017 | Improving Large-Scale Image Retrieval Through Robust Aggregation of Local Descriptors · IEEE Trans. Pattern Anal. Mach. Intell. 2017 |
Multimedia analysis and retrieval › image retrieval
large-scale image retrieval |
0.3 | 1 | 2017 | Improving Large-Scale Image Retrieval Through Robust Aggregation of Local Descriptors · IEEE Trans. Pattern Anal. Mach. Intell. 2017 |
Performance modeling and evaluation
benchmarking |
0.3 | 1 | 2017 | Fast, Compact, and Discriminative: Evaluation of Binary Descriptors for Mobile Applications · IEEE Trans. Multim. 2017 |
Information retrieval › similarity search › metric space similarity search
hamming space search |
0.2 | 1 | 2016 | Improved Hamming Distance Search Using Variable Length Hashing · CVPR 2016 |
Information retrieval
similarity search |
0.2 | 1 | 2016 | Improved Hamming Distance Search Using Variable Length Hashing · CVPR 2016 |
Algorithms and data structures › data structure design › search structures
hashing |
0.2 | 1 | 2016 | Improved Hamming Distance Search Using Variable Length Hashing · CVPR 2016 |
Computer vision › Vision and language
visual relationship detection |
0.1 | 1 | 2021 | Visual Semantic Information Pursuit: A Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2021 |
Machine learning › Representation and self-supervised learning › visual representation › image representation
local feature aggregation |
0.1 | 1 | 2017 | Improving Large-Scale Image Retrieval Through Robust Aggregation of Local Descriptors · IEEE Trans. Pattern Anal. Mach. Intell. 2017 |
Computer vision › Video understanding and tracking › video analytics › video object analysis › object-centric video understanding
moving object recognition |
0.1 | 1 | 2017 | Fast, Compact, and Discriminative: Evaluation of Binary Descriptors for Mobile Applications · IEEE Trans. Multim. 2017 |
Image and video processing › edge detection
edge tracing |
0.1 | 1 | 2008 | Fuzzy chamfer distance and its probabilistic formulation for visual tracking · CVPR 2008 |
Image and video processing
image matching |
0.1 | 1 | 2008 | Fuzzy chamfer distance and its probabilistic formulation for visual tracking · CVPR 2008 |
Multimedia analysis and retrieval
object tracking |
0.1 | 1 | 2008 | Fuzzy chamfer distance and its probabilistic formulation for visual tracking · CVPR 2008 |
Image and video processing
motion estimation |
0.0 | 2 | 1998 | Nonlinear Motion Estimation Using the Supercoupling Approach · IEEE Trans. Pattern Anal. Mach. Intell. 1998 Robust motion analysis · CVPR 1994 |
Image and video processing › video segmentation
motion segmentation |
0.0 | 1 | 1998 | Nonlinear Motion Estimation Using the Supercoupling Approach · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Image and video processing › motion estimation
affine motion estimation |
0.0 | 1 | 1997 | Image Sequence Coding Using Multiple-Level Segmentation and Affine Motion Estimation · IEEE J. Sel. Areas Commun. 1997 |
Image and video coding › video compression
motion compensation |
0.0 | 1 | 1997 | Image Sequence Coding Using Multiple-Level Segmentation and Affine Motion Estimation · IEEE J. Sel. Areas Commun. 1997 |
Image and video coding
video compression |
0.0 | 1 | 1997 | Image Sequence Coding Using Multiple-Level Segmentation and Affine Motion Estimation · IEEE J. Sel. Areas Commun. 1997 |
Image and video processing › motion estimation
optical flow |
0.0 | 1 | 1994 | Robust motion analysis · CVPR 1994 |
Mathematical optimization
global optimization |
0.0 | 1 | 1998 | Nonlinear Motion Estimation Using the Supercoupling Approach · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Computer vision › Video understanding and tracking
motion analysis |
0.0 | 1 | 1994 | Robust motion analysis · CVPR 1994 |
Methods — techniques the papers use, named apart from their topics
message passing neural network · 2.1entropic mirror descent · 1.4multi-stream aggregation · 1.0gumbel-softmax · 0.8mean-field approximation · 0.7mean field variational bayes · 0.7bethe approximation · 0.7fisher vector · 0.6PCA normalization · 0.6variable length hashing · 0.5deep learning · 0.5branch-and-bound · 0.5triplet loss · 0.4convolutional neural network · 0.4KL divergence · 0.4whitening · 0.3matching criteria evaluation · 0.3descriptor extraction · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Importance Weighted Structure Learning for Scene Graph GenerationabstractScene graph generation is a structured prediction task aiming to explicitly model objects and their relationships via constructing a visually-grounded scene graph for an input image. Currently, the message passing neural network based mean field variational Bayesian methodology is the ubiquitous solution for such a task, in which the variational inference objective is often assumed to be the classical evidence lower bound. However, the variational approximation inferred from such loose objective generally underestimates the underlying posterior, which often leads to inferior generation performance. In this paper, we propose a novel importance weighted structure learning method aiming to approximate the underlying log-partition function with a tighter importance weighted lower bound, which is computed from multiple samples drawn from a reparameterizable Gumbel-Softmax sampler. A generic entropic mirror descent algorithm is applied to solve the resulting constrained variational inference task. The proposed method achieves the state-of-the-art performance on various popular scene graph generation benchmarks. Daqi Liu, Miroslaw Bober, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Understanding the Distributions of Aggregation Layers in Deep Neural NetworksabstractThe process of aggregation is ubiquitous in almost all the deep nets' models. It functions as an important mechanism for consolidating deep features into a more compact representation while increasing the robustness to overfitting and providing spatial invariance in deep nets. In particular, the proximity of global aggregation layers to the output layers of DNNs means that aggregated features directly influence the performance of a deep net. A better understanding of this relationship can be obtained using information theoretic methods. However, this requires knowledge of the distributions of the activations of aggregation layers. To achieve this, we propose a novel mathematical formulation for analytically modeling the probability distributions of output values of layers involved with deep feature aggregation. An important outcome is our ability to analytically predict the Kullback-Leibler (KL)-divergence of output nodes in a DNN. We also experimentally verify our theoretical predictions against empirical observations across a broad range of different classification tasks and datasets. Eng-Jon Ong, Syed Sameed Husain, Miroslaw Bober |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Neural Belief Propagation for Scene Graph GenerationabstractScene graph generation aims to interpret an input image by explicitly modelling the objects contained therein and their relationships. In existing methods the problem is predominantly solved by message passing neural network models. Unfortunately, in such models, the variational distributions generally ignore the structural dependencies among the output variables, and most of the scoring functions only consider pairwise dependencies. This can lead to inconsistent interpretations. In this article, we propose a novel neural belief propagation method seeking to replace the traditional mean field approximation with a structural Bethe approximation. To find a better bias-variance trade-off, higher-order dependencies among three or more output variables are also incorporated into the relevant scoring function. The proposed method achieves the state-of-the-art performance on various popular scene graph generation benchmarks. Daqi Liu, Miroslaw Bober, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Constrained Structure Learning for Scene Graph GenerationabstractAs a structured prediction task, scene graph generation aims to build a visually-grounded scene graph to explicitly model objects and their relationships in an input image. Currently, the mean field variational Bayesian framework is the de facto methodology used by the existing methods, in which the unconstrained inference step is often implemented by a message passing neural network. However, such formulation fails to explore other inference strategies, and largely ignores the more general constrained optimization models. In this paper, we present a constrained structure learning method, for which an explicit constrained variational inference objective is proposed. Instead of applying the ubiquitous message-passing strategy, a generic constrained optimization method - entropic mirror descent - is utilized to solve the constrained variational inference step. We validate the proposed generic model on various popular scene graph generation benchmarks and show that it outperforms the state-of-the-art methods. Daqi Liu, Miroslaw Bober, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Efficient Hybrid Network: Inducting Scattering FeaturesabstractRecent work showed that hybrid networks, which combine predefined and learnt filters within a single architecture, are more amenable to theoretical analysis and less prone to overfitting in data-limited scenarios. However, their performance has yet to prove competitive against the conventional counterparts when sufficient amounts of training data are available. In an attempt to address this core limitation of current hybrid networks, we introduce an Efficient Hybrid Network (E-HybridNet). We show that it is the first scattering based approach that consistently outperforms its conventional counterparts on a diverse range of datasets. It is achieved with a novel inductive architecture that embeds scattering features into the network flow using Hybrid Fusion Blocks. We also demonstrate that the proposed design inherits the key property of prior hybrid networks - an effective generalisation in data-limited scenarios. Our approach successfully combines the best of the two worlds: flexibility and power of learnt features and stability and predictability of scattering representations. Dmitry Minskiy, Miroslaw Bober |
ICPR | 2 |
| 2021 | Scattering-Based Hybrid Networks: An Evaluation and Design GuideabstractHybrid networks combine fixed and learnable filters to address the limitations of fully trained CNNs such as poor interpretability, high computational complexity and a need for large training sets. Many hybrid designs were proposed, utilising different filter types, backbone CNNs and different approaches to learning. They were evaluated on different (and often simplistic) datasets, making it difficult to understand their relative performance, their strengths and weaknesses, also there are no design guides on building a hybrid application for the problem at hand. We present and benchmark a collection of 27 networks, some new learnable extensions to existing designs, all within a framework that allows an assessment of a wide range of scattering types and their effects on the system performance. Also, we outline application scenarios most suitable for hybrid networks, identify previously unnoticed trends and provide guidance in building hybrids. Dmitry Minskiy, Miroslaw Bober |
ICIP | 2 |
| 2021 | ACTNET: End-to-End Learning of Feature Activations and Multi-stream Aggregation for Effective Instance Image Retrieval
Syed Sameed Husain, Eng-Jon Ong, Miroslaw Bober |
Int. J. Comput. Vis. | 3 |
| 2021 | Visual Semantic Information Pursuit: A SurveyabstractVisual semantic information comprises two important parts: the meaning of each visual semantic unit and the coherent visual semantic relation conveyed by these visual semantic units. Essentially, the former one is a visual perception task while the latter corresponds to visual context reasoning. Remarkable advances in visual perception have been achieved due to the success of deep learning. In contrast, visual semantic information pursuit, a visual scene semantic interpretation task combining visual perception and visual context reasoning, is still in its early stage. It is the core task of many different computer vision applications, such as object detection, visual semantic segmentation, visual relationship detection, or scene graph generation. Since it helps to enhance the accuracy and the consistency of the resulting interpretation, visual context reasoning is often incorporated with visual perception in current deep end-to-end visual semantic information pursuit methods. Surprisingly, a comprehensive review for this exciting area is still lacking. In this survey, we present a unified theoretical paradigm for all these methods, followed by an overview of the major developments and the future trends in each potential direction. The common benchmark datasets, the evaluation metrics and the comparisons of the corresponding methods are also introduced. Daqi Liu, Miroslaw Bober, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Divergence Based Weighting for Information Channels in Deep Convolutional Neural Networks for Bird Audio DetectionabstractIn this paper, we address the problem of bird audio detection and propose a new convolutional neural network architecture together with a divergence based information channel weighing strategy in order to achieve improved state-of-the-art performance and faster convergence. The effectiveness of the methodology is shown on the Bird Audio Detection Challenge 2018 (Detection and Classification of Acoustic Scenes and Events Challenge, Task 3) development data set. Cemre Zor, Muhammad Awais 0001, Josef Kittler, Miroslaw Bober, Syed Sameed Husain, Qiuqiang Kong, Christian Kroos |
ICASSP | 4 |
| 2019 | Deep Architectures and Ensembles for Semantic Video ClassificationabstractThis paper addresses the problem of accurate semantic labeling of short videos. To this end, a multitude of three different deep nets, ranging from traditional recurrent neural 4 networks (LSTM, GRU), temporal agnostic networks (FV, VLAD, BoW), fully connected neural networks mid-stage AV fusion, and others were considered. Additionally, we also propose a residual architecture-based deep neural network (DNN) for video classification, with state-of-the-art classification performance at significantly reduced complexity. Furthermore, we propose four new approaches to diversity-driven multi-net ensembling, one based on fast correlation measure and three incorporating a DNN-based combiner. We show that significant performance gains can be achieved by ensembling diverse nets and we investigate factors contributing to high diversity. Based on the extensive YouTube8M dataset, we provide an in-depth evaluation and analysis of their behavior. We show that the performance of the ensemble is state-of-the-art achieving the highest accuracy on the YouTube8M Kaggle test data. The performance of the ensemble of classifiers was also evaluated on the HMDB51 and UCF101 datasets, and show that the resulting method achieves comparable accuracy with the state-of-the-art methods using similar input features. Eng-Jon Ong, Syed Sameed Husain, Mikel Bober-Irizar, Miroslaw Bober |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | REMAP: Multi-Layer Entropy-Guided Pooling of Dense CNN Features for Image RetrievalabstractThis paper addresses the problem of very large-scale image retrieval, focusing on improving its accuracy and robustness. We target enhanced robustness of search to factors such as variations in illumination, object appearance and scale, partial occlusions, and cluttered backgrounds -particularly important when search is performed across very large datasets with significant variability. We propose a novel CNN-based global descriptor, called REMAP, which learns and aggregates a hierarchy of deep features from multiple CNN layers, and is trained end-to-end with a triplet loss. REMAP explicitly learns discriminative features which are mutually-supportive and complementary at various semantic levels of visual abstraction. These dense local features are max-pooled spatially at each layer, within multi-scale overlapping regions, before aggregation into a single image-level descriptor. To identify the semantically useful regions and layers for retrieval, we propose to measure the information gain of each region and layer using KL-divergence. Our system effectively learns during training how useful various regions and layers are and weights them accordingly. We show that such relative entropy-guided aggregation outperforms classical CNN-based aggregation controlled by SGD. The entire framework is trained in an end-to-end fashion, outperforming the latest state-of-the-art results. On image retrieval datasets Holidays, Oxford and MPEG, the REMAP descriptor achieves mAP of 95.5%, 91.5% and 80.1% respectively, outperforming any results published to date. REMAP also formed the core of the winning submission to the Google Landmark Retrieval Challenge on Kaggle. Syed Sameed Husain, Miroslaw Bober |
IEEE Trans. Image Process. | 2 |
| 2017 | Movement correction in DCE-MRI through windowed and reconstruction dynamic mode decompositionabstractImages of the kidneys using dynamic contrast-enhanced magnetic resonance renography (DCE-MRR) contains unwanted complex organ motion due to respiration. This gives rise to motion artefacts that hinder the clinical assessment of kidney function. However, due to the rapid change in contrast agent within the DCE-MR image sequence, commonly used intensity-based image registration techniques are likely to fail. While semi-automated approaches involving human experts are a possible alternative, they pose significant drawbacks including inter-observer variability, and the bottleneck introduced through manual inspection of the multiplicity of images produced during a DCE-MRR study. To address this issue, we present a novel automated, registration-free movement correction approach based on windowed and reconstruction variants of dynamic mode decomposition (WR-DMD). Our proposed method is validated on ten different healthy volunteers’ kidney DCE-MRI data sets. The results, using block-matching-block evaluation on the image sequence produced by WR-DMD, show the elimination of $$99\%$$ of mean motion magnitude when compared to the original data sets, thereby demonstrating the viability of automatic movement correction using WR-DMD. Santosh Tirunagari, Norman Poh, Kevin Wells, Miroslaw Bober, Isky Gorden, David Windridge |
Mach. Vis. Appl. | 4 |
| 2017 | Improving Large-Scale Image Retrieval Through Robust Aggregation of Local DescriptorsabstractVisual search and image retrieval underpin numerous applications, however the task is still challenging predominantly due to the variability of object appearance and ever increasing size of the databases, often exceeding billions of images. Prior art methods rely on aggregation of local scale-invariant descriptors, such as SIFT, via mechanisms including Bag of Visual Words (BoW), Vector of Locally Aggregated Descriptors (VLAD) and Fisher Vectors (FV). However, their performance is still short of what is required. This paper presents a novel method for deriving a compact and distinctive representation of image content called Robust Visual Descriptor with Whitening (RVD-W). It significantly advances the state of the art and delivers world-class performance. In our approach local descriptors are rank-assigned to multiple clusters. Residual vectors are then computed in each cluster, normalized using a direction-preserving normalization function and aggregated based on the neighborhood rank. Importantly, the residual vectors are de-correlated and whitened in each cluster before aggregation, leading to a balanced energy distribution in each dimension and significantly improved performance. We also propose a new post-PCA normalization approach which improves separability between the matching and non-matching global descriptors. This new normalization benefits not only our RVD-W descriptor but also improves existing approaches based on FV and VLAD aggregation. Furthermore, we show that the aggregation framework developed using hand-crafted SIFT features also performs exceptionally well with Convolutional Neural Network (CNN) based features. The RVD-W pipeline outperforms state-of-the-art global descriptors on both the Holidays and Oxford datasets. On the large scale datasets, Holidays1M and Oxford1M, SIFT-based RVD-W representation obtains a mAP of 45.1 and 35.1 percent, while CNN-based RVD-W achieve a mAP of 63.5 and 44.8 percent, all yielding superior performance to the state-of-the-art. Syed Sameed Husain, Miroslaw Bober |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Fast, Compact, and Discriminative: Evaluation of Binary Descriptors for Mobile ApplicationsabstractLocal feature descriptors underpin many diverse applications, supporting object recognition, image registration, database search, 3D reconstruction, and more. The recent phenomenal growth in mobile devices and mobile computing in general has created demand for descriptors that are not only discriminative, but also compact in size and fast to extract and match. In response, a large number of binary descriptors have been proposed, each claiming to overcome some limitations of the predecessors. This paper provides a comprehensive evaluation of several promising binary designs. We show that existing evaluation methodologies are not sufficient to fully characterize descriptors' performance and propose a new evaluation protocol and a challenging dataset. In contrast to the previous reviews, we investigate the effects of the matching criteria, operating points, and compaction methods, showing that they all have a major impact on the systems' design and performance. Finally, we provide descriptor extraction times for both general-purpose systems and mobile devices, in order to better understand the real complexity of the extraction task. The objective is to provide a comprehensive reference and a guide that will help in selection and design of the future descriptors. Simone Madeo, Miroslaw Bober |
IEEE Trans. Multim. | 2 |
| 2016 | Improved Hamming Distance Search Using Variable Length HashingabstractThis paper addresses the problem of ultra-large-scale search in Hamming spaces. There has been considerable research on generating compact binary codes in vision, for example for visual search tasks. However the issue of efficient searching through huge sets of binary codes remains largely unsolved. To this end, we propose a novel, unsupervised approach to thresholded search in Hamming space, supporting long codes (e.g. 512-bits) with a wide-range of Hamming distance radii. Our method is capable of working efficiently with billions of codes delivering between one to three orders of magnitude acceleration, as compared to prior art. This is achieved by relaxing the equal-size constraint in the Multi-Index Hashing approach, leading to multiple hash-tables with variable length hash-keys. Based on the theoretical analysis of the retrieval probabilities of multiple hash-tables we propose a novel search algorithm for obtaining a suitable set of hash-key lengths. The resulting retrieval mechanism is shown empirically to improve the efficiency over the state-of-the-art, across a range of datasets, bit-depths and retrieval thresholds. Eng-Jon Ong, Miroslaw Bober |
CVPR | 2 |
| 2014 | Robust and scalable aggregation of local features for ultra large-scale retrievalabstractThis paper is concerned with design of a compact, binary and scalable image representation that is easy to compute, fast to match and delivers beyond state-of-the-art performance in visual recognition of objects, buildings and scenes. A novel descriptor is proposed which combines rank-based multi-assignment with robust aggregation framework and cluster/bit selection mechanisms for size scalability. Extensive performance evaluation is presented, including experiments within the state-of-the art pipeline developed by the MPEG group standardising Compact Descriptors for Visual Search (CVDS). Syed Sameed Husain, Miroslaw Bober |
ICIP | 2 |
| 2014 | Component hashing of variable-length binary aggregated descriptors for fast image searchabstractCompact locally aggregated binary features have shown great advantages in image search. As the exhaustive linear search in Hamming space still entails too much computational complexity for large datasets, recent works proposed to directly use binary codes as hash indices, yielding a dramatic increase in speedup. However, these methods cannot be directly applied to variable-length binary features. In this paper, we propose a Component Hashing (CoHash) algorithm to handle the variable-length binary aggregated descriptors indexing for fast image search. The main idea is to decompose the distance measure between variable-length descriptors into aligned component-to-component matching problems independently, and build multiple hash tables for the visual word components. Given a query, its candidate neighbors are found by using the query binary sub-vectors as indices into their corresponding hash tables. In particular, a bit selection based on conditional mutual information maximization is proposed to reduce the dimensionality of visual word components, which provides a light storage of indices and balances the retrieval accuracy and search cost. Extensive experiments on benchmark datasets show that our approach is 20~25 times faster than linear search, without any noticeable retrieval performance loss. Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001, Wen Gao 0001, Miroslaw Bober |
ICIP | 6 |
| 2013 | Special issue on visual search and augmented reality
Giovanni Cordara, Miroslaw Bober, Yuriy A. Reznik |
Signal Process. Image Commun. | 2 |
| 2012 | The MPEG-7 Video Signature Tools for Content IdentificationabstractThis paper presents the core technologies of the video signature tools recently standardized by ISO/IEC Moving Picture Experts Group (MPEG) as an amendment to the MPEG-7 Standard (ISO/IEC 15938). The video signature is a high-performance content fingerprint that is suitable for desktop scale to web-scale deployment and provides high levels of robustness to common video editing operations and high temporal localization accuracy at extremely low false alarm rates, achieving a detection rate in the order of 96% at a false alarm rate in the order of five false matches per million comparisons. The applications of the video signature are numerous and include rights management and monetization, distribution management, usage monitoring, metadata association, and corporate or personal database management. In this paper, we review the prior work in the field, explain the standardization process and status, and provide details and evaluation results for the video signature tools. Stavros Paschalakis, Kota Iwamoto, Paul Brasnett, Nikola Sprljan, Ryoma Oami, Toshiyuki Nomura, Akio Yamada, Miroslaw Bober |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2010 | Recent developments on standardisation of MPEG-7 Visual Signature ToolsabstractThis paper presents the latest developments and possible new directions for future work in standardisation of Visual Signature Tools within the Moving Picture Experts Group (MPEG). The tools, which include the Image Signature descriptor and the recently completed Video Signature descriptor, form a part of the MPEG-7 specification. They enable fast and robust detection of duplicate or derived visual media content, images and videos. Descriptors of this type are sometimes also referred to as fingerprints or robust hashes. Here we mainly focus on introducing the technology behind the recently completed Video Signature Tools and describe some recent developments and demonstration applications for the Image Signature Tools. Finally, we briefly present MPEG exploratory investigations on requirements of searching for different images containing the same visual objects within the mobile visual search framework. Paul Brasnett, Stavros Paschalakis, Miroslaw Bober |
ICME | 3 |
| 2009 | MPEG-7 visual signature toolsabstractThe MPEG-7 standard offers a comprehensive set of audiovisual content Description Tools to support applications enabling effective and efficient access to multimedia content. MPEG recently identified a need for a set of new, unique descriptors necessary for detection of duplicate or derived visual media content. The MPEG-7 Visual Signature Tools are characterized by efficient detection at low false positive rates, high robustness and enable very fast searching. This paper outlines application scenarios for visual signatures (also known as fingerprints or robust hashes) and discusses the requirements, evaluation process and selection methodology employed by MPEG. The Image Signature technology selected by MPEG and its performance is also described and the latest work on Video Signatures and associated standardisation time-line is summarized. Miroslaw Bober, Paul Brasnett |
ICME | 1 |
| 2008 | Fuzzy chamfer distance and its probabilistic formulation for visual trackingabstractThe paper presents a fuzzy chamfer distance and its probabilistic formulation for edge-based visual tracking. First, connections of the chamfer distance and the Hausdorff distance with fuzzy objective functions for clustering are shown using a reformulation theorem. A fuzzy chamfer distance (FCD) based on fuzzy objective functions and a probabilistic formulation of the fuzzy chamfer distance (PFCD) based on data association methods are then presented for tracking, which can all be regarded as reformulated fuzzy objective functions and minimized with iterative algorithms. Results on challenging sequences demonstrate the performance of the proposed tracking method. Yonggang Jin, Farzin Mokhtarian, Miroslaw Bober, John Illingworth |
CVPR | 3 |
| 2008 | Fast and robust image identificationabstractThis paper presents an image identifier robust to common modifications. A multi-resolution Trace transform is introduced that constructs a set of 1D representations of an image. A binary identifier is extracted from each representation using a Fourier transform. Experimental evaluation of the algorithm and three state-of-the-art methods was carried out on a set of over 60,000 unique images with results demonstrating that the method outperforms the prior art methods in terms of detection, robustness and speed and achieves detection rate over 99% at a false-positive rate below 1 per million, with search speed exceeding 10 million image pairs per second on a desktop PC. Paul Brasnett, Miroslaw Bober |
ICPR | 2 |
| 2005 | Dual LDA - an effective feature space reduction method for face recognitionabstractLinear discriminant analysis (LDA) is a popular feature extraction technique that aims at creating a feature set of enhanced discriminatory power. The authors introduced a novel approach dual LDA (DLDA) and proposed an efficient SVD-based implementation. This paper focuses on feature space reduction aspect of DLDA achieved in course of proper choice of the parameters controlling the DLDA algorithm. The comparative experiments conducted on a collection of five facial databases consisting in total of more than 10000 photos show that DLDA outperforms by a great margin the methods reducing the feature space by means of feature subset selection. Krzysztof Kucharski, Wladyslaw Skarbek, Miroslaw Bober |
AVSS | 3 |
| 2005 | Feature Space Reduction for Face Recognition with Dual Linear Discriminant Analysis
Krzysztof Kucharski, Wladyslaw Skarbek, Miroslaw Bober |
CAIP | 3 |
| 2004 | Dual LDA for Face Recognition
Wladyslaw Skarbek, Krzysztof Kucharski, Miroslaw Bober |
Fundam. Informaticae | 3 |
| 2003 | Face Recognition by Fisher and Scatter Linear Discriminant Analysis
Miroslaw Bober, Krzysztof Kucharski, Wladyslaw Skarbek |
CAIP | 1 |
| 2003 | An FPGA System for the High Speed Extraction, Normalization and Classification of Moment Descriptors
Stavros Paschalakis, Peter Lee 0002, Miroslaw Bober |
FPL | 3 |
| 2003 | A low cost FPGA system for high speed face detection and trackingabstractWe present an FPGA face detection and tracking system for audiovisual communications, with a particular focus on mobile videoconferencing. The advantages of deploying such a technology in a mobile handset are many, including face stabilisation, reduced bit rate, and higher quality video on practical display sizes. Most face detection methods, however, assume at least modest general purpose processing capabilities, making them inappropriate for real-time applications, especially for power-limited devices, as well as modest custom hardware implementations. We present a method which achieves a very high detection and tracking performance and, at the same time, entails a significantly reduced computational complexity, allowing real-time implementations on custom hardware or simple microprocessors. We then propose on FPGA implementation which entails very low logic and memory costs and achieves extremely high processing rates at very low clock speeds. Stavros Paschalakis, Miroslaw Bober |
FPT | 2 |
| 2001 | MPEG-7: Evolution or Revolution?
Miroslaw Bober |
CAIP | 1 |
| 2001 | Segmenting Traffic Scenes from Grey Level and Motion Information
Jorge Badenas, Miroslaw Bober, Filiberto Pla |
Pattern Anal. Appl. | 2 |
| 2001 | MPEG-7 visual shape descriptorsabstractThis paper describes techniques and tools for shape representation and matching, developed in the context of MPEG-7 standardization. The application domains for each descriptor are considered, and the contour-based shape descriptor is presented in some detail. Example applications are also shown. Miroslaw Bober |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | On Accurate and Robust Estimation of Fundamental Matrix
Miroslaw Bober, Nikos Georgis, Josef Kittler |
Comput. Vis. Image Underst. | 1 |
| 1998 | Nonlinear Motion Estimation Using the Supercoupling ApproachabstractThis paper presents the application of a very efficient multiresolution transformation, which is related to the renormalization group approach of physics, to the problem of motion segmentation. The approach proposed is much faster and yields much better results than the full resolution approach. The problem is formulated as one of global optimization where a cost function is constructed to combine the information obtained by various processors as well as the constraints we impose to the problem. The cost function is optimized using the supercoupling multiresolution approach. Miroslaw Bober, Maria Petrou, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | Scalable image coding using Gaussian pyramid vector quantisation with resolution-independent block sizeabstractWe present a new approach to multiresolution vector quantisation. Its main advantage is exploitation of long-range correlations in the image by keeping the vector size constant, independent of the image scale. We also developed a variable block-rate version of the algorithm, which allows better utilisation of the available bit budget by refining only those areas of the image which are not efficiently approximated by lower resolutions of the pyramid. Leszek Cieplinski, Miroslaw Bober |
ICASSP | 2 |
| 1997 | Image Sequence Coding Using Multiple-Level Segmentation and Affine Motion EstimationabstractA very low bit-rate video codec using multiple-level segmentation and affine motion compensation is presented. The translational motion model is adequate to motion compensate small regions even when complex motion is involved; however, it is no longer capable of delivering satisfactory results when applied to large regions or the whole frame. The proposed codec is based on a variable block size algorithm enhanced with global motion compensation, inner block segmentation, and a set of motion models used adaptively in motion compensation. The experimental results show that the proposed method gives better results in terms of the bit rate under the same PSNR constraint for most of the tested sequences as compared with the fixed block size approach and traditional variable block size codec in which only translational motion compensation is utilized. Miroslaw Bober, Josef Kittler |
IEEE J. Sel. Areas Commun. | 2 |
| 1996 | On Accurate and Robust Estimation of Fundamental MatrixabstractThis paper is concerned with the accurate and robust estimation of the fundamental matrix. We show that, given a certain conditions, a basic linear algorithm can yield excellent accuracy, in cases two orders of magnitude better than sophisticated algorithms. The key element of the success is the accuracy and the statistical distribution of the errors of displacement estimates used as input. We propose a low-- level, gradient--based (as opposed to feature--based) algorithm based on the Hough Transform to extract the low--level measurements. We show that it is much more efficient, both in terms of computational expense and accuracy of the final estimate, to remove the errors in the intermediate representation (optic--flow) than to attempt to improve the final estimate by complicated non--linear algorithms. Experimental results are also included. 1 Introduction The correspondence analysis of the images of two views of a scene is a problem which arises in a number of applicat... Miroslaw Bober, Nikos Georgis, Josef Kittler |
BMVC | 1 |
| 1996 | Video coding using affine motion compensated predictionabstractThe aim in variable block size and object based video coding is to motion compensate as large regions as possible. Whereas the translational motion model is adequate to motion compensate small regions even if the image sequence involves complex motion, it is no longer capable of delivering satisfactory results for large regions. A more sophisticated motion model is required in such regions. In this paper, we extend our previous approach by proposing a multiple layer video codec which uses affine model for large regions to improve motion compensation. The translational model is still used for small regions to reduce the bit overhead. A multi-resolution robust Hough transform based motion estimation technique is used to estimate the affine motion parameters. The experimental results show that the proposed method gives better results in terms of the bit rate under the same PSNR constraint for most of the tested sequences as compared with our previous approach and an implementation of the standard H.261. Miroslaw Bober, Josef Kittler |
ICASSP | 2 |
| 1996 | A hybrid codec for very low bit rate video codingabstractIn segmentation based video compression algorithms, the quality of the segmentation and the efficiency of the boundary coding have strong influence on the coding efficiency and image quality. We propose a hybrid codec using motion segmentation and employing a variable size block structure. The image is segmented based on motion and grey level information and represented by variable size blocks combined with an inner block partition. A method of partitioning using fixed patterns is developed to encode the motion boundaries inside a block. This approach improves the boundary coding efficiency by a factor of three as compared with the straight line approximation used in our earlier approach. The experimental results show that the proposed algorithm decreases the bit rate and preserves good image quality. Miroslaw Bober, Josef Kittler |
ICIP (1) | 2 |
| 1995 | Combining the Hough Transform and Multiresolution MRF's for the Robust Motion Estimation
Miroslaw Bober, Josef Kittler |
ACCV | 1 |
| 1995 | Motion based image segmentation for video codingabstractA robust and stable scene segmentation is a prerequisite for the object based coding. Various approaches to this complex task have been proposed, including segmentation of optic flow, grey-level based segmentation or simple division of the scene into moving and stationary regions. In this paper, we propose an algorithm which combines all three approaches in order to get a more robust and accurate segmentation of the moving objects. The experimental results show that the proposed algorithm can significantly reduce over-segmentation and maintain accurate motion boundaries. The use of the proposed approach in video coding can increase the PSNR and reduce the bit rate. Miroslaw Bober, Josef Kittler |
ICIP (3) | 2 |
| 1994 | Robust motion analysisabstractWe develop a new robust algorithm for the estimation of optic flow and extraction of other motion-relevant information. A novel combination of the Hough Transform, and robust statistical methods results in unbiased estimates for multiple motions, parallel segmentation and estimation and increased robustness to noise and changes of illumination. The algorithm is fast, due to application of multiresolution in both image and parameter space. A simple, translational motion model and a complex one coping with rotation and change of scale are applied. Also, an accuracy measure for the derived estimate is introduced. The paper includes experimental tests of this new approach and its comparison with several other widely-cited methods. The experiments were aimed at assessing the effect of noise, change of illumination and multiple motions on the algorithms performance. The results show that our approach is significantly more robust than other methods.> Miroslaw Bober, Josef Kittler |
CVPR | 1 |
| 1994 | Robust Motion Estimation and Multistage Vector Quantisation for Sequence CompressionabstractMotion prediction and spatial coding are the two main techniques used to construct algorithms for image sequence compression. We present an approach which merges a robust motion estimation technique, based on the Hough transform and robust statistics, with multiple stage vector quantization. MSVQ uses global optimization and multipath searching. The algorithm is capable of segmenting multiple motions and uses line segments to code real motion boundaries. This significantly improves the subjective quality of coded sequence on the motion edges. Experimental results show that the proposed algorithm can achieve better PSNR without the increase in bitrate.> Miroslaw Bober, Josef Kittler |
ICIP (2) | 2 |
| 1994 | Multiresolution motion segmentationabstractThis paper shows the application of a very efficient multiresolution transformation which is related to the renormalization group approach of physics, to the problem of motion segmentation. The proposed approach is much faster and yields much better results than the full resolution approach. Maria Petrou, Miroslaw Bober, Josef Kittler |
ICPR (1) | 2 |
| 1994 | Estimation of complex multimodal motion: an approach based on robust statistics and Hough transform
Miroslaw Bober, Josef Kittler |
Image Vis. Comput. | 1 |
| 1993 | Estimation of Complex Multimodal Motion: An Approach based on Robust Statistics and Hough TransformabstractAbstract An application of robust statistics in a Hough transform based motion estimation approach is presented. The algorithm is developed and experiments are performed, proving its superior performance in terms of estimate accuracy, convergence, robustness and better segmentation. Comparative results with standard methods are also included. Miroslaw Bober, Josef Kittler |
BMVC | 1 |