VLDB 2026 Research / reviewers in the wild / expert
Josef Kittler
dblp:k/JosefKittler
· DBLP profile ↗
576ranked-venue papers
58as first author
151since 2021 · last 2026
0000-0002-8110-9205ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 407 · 44 first-author · 100 since 2021Graphics, computer vision, multimedia, augmented reality and games · 309 · 17 first-author · 61 since 2021Databases, data management, data science and information retrieval · 16 · 3 first-author · 6 since 2021Security and privacy · 14 · 1 since 2021Human-computer interaction and ubiquitous computing · 12 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Computer networks · 5 · 3 since 2021Systems, architecture and hardware · 2 · 2 first-authorTheory of computation · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoZSR-VAD: Contextual Zero-Shot Reasoning for Video Anomaly Detection
Mohd Ubaid Wani, Sara Atito Ali Ahmed, Srinivasa Rao Nandam, Josef Kittler, Muhammad Awais 0001 |
ICPR (12) | 4 |
| 2026 | SketchMind: Understanding Abstract Sketches with MLLMs for Fine-Grained Sketch-Based Image Retrieval
Changxing Li, Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001, Josef Kittler |
WWW | 5 |
| 2026 | A Color Information Driven Collaborative Training of Dual Task Parallel Network for Visible and Thermal Infrared Image Fusion and Saliency Object Detection
Zeyang Zhang 0002, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Muhammad Awais 0001, Josef Kittler |
Int. J. Comput. Vis. | 6 |
| 2026 | HACG: Leveraging Hierarchical Alignment and Caption Generation for Text-Video Retrieval
Donglin Zhang 0001, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 4 |
| 2026 | Attack Intensity is Target-Related: Exploration of Sparse Adversarial Attack against Visual Object Trackers
Shao-Chuan Zhao, Tianyang Xu 0001, Zhangyong Tang, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 6 |
| 2026 | Cross-sample Consistency Learning for Semi-supervised Medical Image Segmentation
Tao Zhou 0002, Yunqi Gu, Kaiwen Huang 0002, Huazhu Fu, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 6 |
| 2026 | Interactive image-to-video transfer learning
Cong Wu 0006, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler |
Neural Networks | 5 |
| 2026 | SPD-Updater: Symmetric positive definite manifold geometry based temporal updating for visual object trackingabstractVisual object tracking has witnessed continuous advances in recent years along with the exciting developments in backbone networks. In general, all advanced solutions adhere to the template-based tracking framework, which exhibits powerful representative capacity gained via offline training. However, when the target undergoes appearance changes or occlusion, the tracker, which relies on a fixed template defined in the initial frame, struggles to locate it accurately in such complex situations. To achieve online adaptation, recent studies have introduced dynamic templates. Typically, the adopted solution is to compute reliability scores in the traditional Euclidean space to assess the confidence of the dynamic template. However, the Euclidean metric is unreliable to some extent in high-dimensional feature spaces, potentially resulting in a negative impact by involving incorrect dynamic templates. To overcome this problem, we exploit the compact geometric representation capacity of the Symmetric Positive Definite (SPD) manifold to design a novel score prediction module for the tracker update (SPD-Updater). By switching to an SPD manifold metric, we obtain a more accurate and stable dynamic template, thereby enhancing the model capacity to handle complex situations. To validate the reliability of manifold metric in tracking models, we conduct experiments with trackers using different backbones. The experimental results on LaSOT, GOT-10k, TrackingNet, and UAV123 demonstrate the effectiveness of our approach, reflecting the merit of the SPD metric in online tracking adaptation. Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler |
Neural Networks | 5 |
| 2026 | EvaNet: Toward More Efficient and Consistent Infrared and Visible Image Fusion AssessmentabstractEvaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to capture the true quality of the fusion results but also are computationally demanding. To address these issues, we propose a unified evaluation framework specifically tailored for image fusion. At its core is a lightweight network designed efficiently to approximate widely used metrics, following a divide-and-conquer strategy. Unlike conventional approaches that directly assess similarity between fused and source images, we first decompose the fusion result into infrared and visible components. The evaluation model is then used to measure the degree of information preservation in these separated components, effectively disentangling the fusion evaluation process. During training, we incorporate a contrastive learning strategy and inform our evaluation model by perceptual scene assessment provided by a large language model. Last, we propose the first consistency evaluation framework, which measures the alignment between image fusion metrics and human visual perception, using both independent no-reference scores and downstream tasks performance as objective references. Extensive experiments show that our learning-based evaluation paradigm delivers both superior efficiency (up to 1,000 times faster) and greater consistency across a range of standard image fusion benchmarks. Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Tao Zhou 0002, Hui Li 0037, Zhangyong Tang, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Probabilistically Aligned View-Unaligned Clustering With Adaptive Template SelectionabstractIn most existing multi-view modeling scenarios, cross-view correspondence (CVC) between instances of the same target from different views, like paired image-text data, is a crucial prerequisite for effortlessly deriving a consistent representation. Nevertheless, this premise is frequently compromised in certain applications, where each view is organized and transmitted independently, resulting in the view-unaligned problem (VuP). Restoring CVC of unaligned multi-view data is a challenging and highly demanding task that has received limited attention from the research community. To tackle this practical challenge, we propose to integrate the permutation derivation procedure into the bipartite graph paradigm for view-unaligned clustering, termed Probabilistically Aligned View-unaligned Clustering with Adaptive Template Selection (PAVuC-ATS). Specifically, we learn consistent anchors and view-specific graphs by the bipartite graph, and derive permutations applied to the unaligned graphs by reformulating the alignment between two latent representations as a 2-step transition of a Markov chain with adaptive template selection, thereby achieving the probabilistic alignment. The convergence of the resultant optimization problem is validated both experimentally and theoretically. Extensive experiments on six benchmark datasets demonstrate the superiority of the proposed PAVuC-ATS over the baseline methods. Wenhua Dong, Xiaojun Wu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Celebrating the Life and Research Work of Edwin Hancock
Xiao Bai 0001, Jun Zhou 0001, Richard C. Wilson 0001, Charlotte Davies, Josef Kittler |
Pattern Recognit. | 5 |
| 2026 | Towards data-efficient generalised portrait animation: Landmark-driven latent traversal with fine-grained control mechanism
Yingjie Dai, Josef Kittler |
Pattern Recognit. | 4 |
| 2026 | Unfolded ISTA for deep sparse subspace clustering
Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 4 |
| 2026 | Switcher: Adaptive framework for unified and customized multi-modal object tracking
He Wang 0028, Tianyang Xu 0001, Zhangyong Tang, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 5 |
| 2026 | Single image, any face: Generalisable 3D face generation
Wenqing Wang 0002, Haosen Yang 0003, Josef Kittler, Xiatian Zhu |
Pattern Recognit. | 3 |
| 2026 | RSPH: Robust self-paced hashing for cross-modal retrieval
Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 4 |
| 2026 | Supervised discrete cross-modal hashing with exploiting semantic correlations
Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 4 |
| 2026 | Visual complexity guided diffusion defender for video object tracking and recognition
Shao-Chuan Zhao, Tianyang Xu 0001, Hui Li 0037, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 5 |
| 2026 | Adaptive Continual Learning for Online Visual Object Tracking via Dynamic Grassmannian Appearance ModelingabstractOnline visual object tracking fundamentally constitutes a continual learning challenge, demanding persistent adaptation to target variations within dynamic video streams, while preserving critical features to prevent forgetting. Existing template-based trackers are particularly vulnerable to severe target deformation and drastic background changes in real-world scenarios. Recent explorations enhance adaptability through local update strategies—such as dynamic templates, appearance tokens, and parameter fine-tuning—to address rapid appearance variations. However, these methods inherently propagate target appearance changes without explicit modelling and prioritise short-term adaptation over long-term global representation. Consequently, they fail to balance initial target features with current observations, leading to tracking failure in long-term scenarios. To overcome these limitations, we propose DG-Track, an adaptive continual learning framework leveraging Grassmannian manifold geometry. Specifically, we represent target appearance within a Grassmannian affine subspace and perform continual adaptation via incremental learning. Compared to Euclidean geometry, the Grassmannian manifold captures nonlinear appearance representations, yielding more compact and geometrically consistent temporal dynamics modelling. Furthermore, an adaptive forgetting module dynamically regulates the interplay between the current observed subspace and the initial template, ensuring stable long-term tracking. DG-Track is a plug-and-play solution for online tracking, adding no learnable parameters. Comprehensive experiments with diverse baseline trackers on LaSOT, GOT-10k, TrackingNet, and UAV123 validate the efficacy of our continual learning framework and Grassmannian manifold geometry in enhancing visual object tracking performance. Code is available at https://github.com/xiaoqing0825/DGTrack. Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Probabilistic Embeddings With Evidence Learning and Refinement for Text-Video RetrievalabstractThis paper studies the problem of text-video retrieval, where the goal is to learn accurate cross-modal alignment between videos and text. This problem is challenging because of the matching ambiguity caused by the inherent gap between the heterogeneous video and text modalities. In particular, the differences in the information granularity and abstraction levels between the two modalities hinder a reliable sample-level alignment. Moreover, redundant visual content, sparse textual descriptions, and temporal variability in videos introduce additional uncertainty, resulting in ambiguous matching and suboptimal performance. In this paper, we propose a novel method named Probabilistic Embeddings with Evidence Learning and Refinement (PE2LR), which models video-text pairs as probability distributions and captures uncertainty through the evidence theory. Specifically, we perform distribution-level representation learning to resolve the semantic ambiguity of video-text pairs. To improve the alignment further, we introduce a distribution-based embedding refinement module to ameliorate the semantic consistency across modalities. The proposed PE2LR is able to pull positive sample pairs closer in the embedding space, while pushing the negative pairs apart. Comprehensive experiments on several benchmark datasets (including MSRVTT, DiDeMo, and ActivityNet Captions) demonstrate that our PE2LR achieves state-of-the-art search performance. Donglin Zhang 0001, Zheng Rao, Xing Xu 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Image Process. | 5 |
| 2026 | Beyond Mere Tuning: Harnessing the Full Potential of Prompts for Text-Video RetrievalabstractText-video retrieval has advanced significantly by thanks to the merits of large-scale contrastive languageimage pre-training. Recently, prompt tuning has emerged as a parameter-efficient strategy to solve the substantial training costs associated with these models. However, existing studies are generally confined to simple tuning strategies, without fully exploring the potential of prompts. Unlike its application to singlemodal tasks, prompt tuning in cross-modal learning involves not only the adaptation to specific tasks but also the challenge of dealing with the intricate interaction of semantic information. These aspects remain insufficiently explored, thereby limiting retrieval performance. To this end, we introduce a comprehensive framework for harnessing the full potential of prompts in textvideo retrieval. Specifically, we explore prompts in three respects: i) Task adaptation: we design appropriate prompts to guide the fine-tuning of pre-trained models to perform text-video retrieval tasks. ii) Temporal modelling: to enable temporal modelling, while protecting the original knowledge from being corrupted, we propose a prompt rolling module that enables effective temporal processing without adding extra training parameters. iii) The sampling scope prediction: We introduce a query expansion strategy that exploits the rich semantic information conveyed by multimodal prompt features to define the sampling scope more precisely. Extensive experiments on four mainstream datasets, i.e., MSR-VTT, ActivityNet, DiDeMo, and LSMDC, demonstrate that our method achieves state-of-the-art performance. Donglin Zhang 0001, Xing Xu 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Multim. | 5 |
| 2026 | BusReF: Infrared-Visible Images Registration and Fusion Focus on Reconstructible Area Using One Set of FeaturesabstractIn multi-modal imaging scenarios, the misalignment of images presents a persistent challenge. Conventional image fusion algorithms, aiming to enhance the performance of downstream vision tasks, presuppose strictly registered inputs to achieve satisfactory results. To relax this assumption, a common approach is to register the images first; however, existing multi-modal registration methods are often hindered by complex architectures and a heavy reliance on semantic information. This article proposes BusRef, a unified framework that jointly addresses image registration and fusion, with a specific focus on the Infrared-Visible Image Registration and Fusion (IVRF) task. Within this framework, unaligned image pairs are processed through three sequential stages: coarse registration, fine registration, and fusion. We demonstrate that this integrated approach enables more robust and accurate IVRF. Key to our framework is a novel training and evaluation strategy that employs masks to mitigate the influence of non-reconstructible regions on the loss function, thereby significantly improving the model’s accuracy and robustness. Furthermore, we introduce a gradient-aware fusion network designed to effectively preserve complementary information from both modalities. Comprehensive experiments demonstrate that BusRef achieves superior performance when compared against various state-of-the-art registration and fusion algorithms. Our code is available at https://github.com/Yukarizz/BusReF . Zeyang Zhang 0002, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Congcong Bian, Josef Kittler |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2026 | Resilient Semantic Pseudo-Text Embedding for Zero-Shot Video Moment RetrievalabstractWith the explosive growth of video data, video moment retrieval (VMR) has attracted increasing attention due to its ability to localize semantically relevant moments in untrimmed videos. However, existing VMR approaches usually rely on annotated video-text correspondences or temporal annotations, both of which require significant human effort and are costly to scale. Even worse, the inherent subjectivity in manual labeling often introduces inconsistencies into the training data, further complicating the issue. In this article, we investigate the problem of Zero-Shot Video Moment Retrieval (ZS-VMR) and develop a novel method, Resilient Semantic Pseudo-Text Modeling (RSPT). The core of RSPT is to construct semantically rich pseudo-text embeddings through visually guided perturbations. Specifically, RSPT first generates initial pseudo-texts by injecting random noise into visual features and then learns adaptive noise weights by modeling the correlations between these pseudo-texts and visual features. This enables the generation of diverse and semantically aligned representations from multiple perspectives. To ensure alignment with visual semantics and suppress irrelevant noise, RSPT introduces a quality-aware contrastive loss that regularizes the semantic boundaries of pseudo-texts. Extensive experiments on Charades-STA and ActivityNet-Captions show that RSPT outperforms existing competitive baselines, validating its efficacy. Code is available at https://github.com/dmcsy/RSPT . Donglin Zhang 0001, Weixiang Shi, Xiaojun Wu 0001, Josef Kittler |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | PETS2025: Multi-Authority Multi-Sensor Maritime Surveillance Challenge and EvaluationabstractThis paper presents the outcomes of the PETS2025 challenge, held in conjunction with AVSS 2025 and sponsored by the EU-funded EURMARS project. The challenge introduces a novel maritime surveillance dataset comprising image sequences captured by diverse multi-altitude, multimodal sensors, reflecting the real-world multi-authority environment. The key tasks include: (1) object detection using various sensors across different platforms (ground-based and low-altitude aerial) and spectral ranges (visible, thermal, ultraviolet (UV), and short-wave infrared (SWIR)); (2) long-term tracking of targets in maritime environments spanning both sea and land; and (3) approximating target geolocations by using sensor imagery and telemetry data. Performance evaluations of results submitted by 12 international participants are discussed. The results show the effectiveness of these submissions and highlight ongoing challenges posed by heterogeneous sensors and complex environments. These challenges emphasise the need to further improve detection, tracking, and geolocation approximation for maritime and coastal surveillance. Thanet Markchom, Jonathan N. Boyle, Lulu Chen, James M. Ferryman, Matteo Marturini, Stephan Veigl, Andreas Opitz, Andreas Kriechbaum-Zabini, Romaios Bratskas, Anastasios Gkamaris, Dimitris Papachristos, George Leventakis, Wenjun Fan, Hsiang-Wei Huang, Jeng-Neng Hwang, Pyong-Kun Kim, Kwangju Kim, Chung-I Huang, Kenta Saito, Shunta Kaneko, Kyoko Sudo, Nguyen Thanh Thien, Meng-Yu Kao, Jun-Wei Hsieh, Teepakorn Lilek, Tossapol Pomsuwan, Jinjie Gu, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler, Stephanie Stacy, Alfredo Gabaldon, Peter Tu, Dongyoung Kim, Kyoungoh Lee |
AVSS | 31 |
| 2025 | One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image FusionabstractAdvanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction through pixel-level supervision. This new paradigm provides strong guidance for unsupervised multimodal fusion without relying on abstract semantics, enhancing task-shared feature learning for broader applicability. Owning to the hybrid image features and enhanced universal representations, the proposed GIFNet supports diverse fusion tasks, achieving high performance across both seen and unseen scenarios with a single model. Uniquely, experimental results reveal that our framework also supports single-modality enhancement, offering superior flexibility for practical applications. Our code will be available at https://github.com/AWCXV/GIFNet. Chunyang Cheng, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Zhangyong Tang, Hui Li 0037, Zeyang Zhang 0002, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
CVPR | 10 |
| 2025 | Text Augmented Correlation Transformer For Few-shot Classification & SegmentationabstractFoundation models like CLIP and ALIGN have transformed few-shot and zero-shot vision applications by fusing visual and textual data, yet the integrative few-shot classification and segmentation (FS-CS) task primarily leverages visual cues, overlooking the potential of textual support. In FS-CS scenarios, ambiguous object boundaries and overlapping classes often hinder model performance, as limited visual data struggles to fully capture high-level semantics. To bridge this gap, we present a novel multi-modal FS-CS framework that integrates textual cues into support data, facilitating enhanced semantic disambiguation and fine-grained segmentation. Our approach first investigates the unique contributions of exclusive text-based support, using only class labels to achieve FS-CS. This strategy alone achieves performance competitive with vision-only methods on FS-CS tasks, underscoring the power of textual cues in few-shot learning. Building on this, we introduce a dualmodal prediction mechanism that synthesizes insights from both textual and visual support sets, yielding robust multimodal predictions. This integration significantly elevates FS-CS performance, with classification and segmentation improvements of +3.7/6.6% (1-way 1-shot) and +8.0/6.5% (2-way 1-shot) on COCO-20i, and +2.2/3.8% (1-way 1shot) and +4.3/4.0% (2-way 1-shot) on Pascal-5i. Additionally, in weakly supervised FS-CS settings, our method surpasses visual-only benchmarks using textual support exclusively, further enhanced by our dual-modal predictions. By rethinking the role of text in FS-CS, our work establishes new benchmarks for multi-modal few-shot learning and demonstrates the efficacy of textual cues for improving model generalization and segmentation accuracy. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
CVPR | 4 |
| 2025 | Enhanced Weakly Supervised Few-shot Classification & SegmentationabstractThe emergence of vision-language foundation models has enabled the integration of textual information into vision-based applications. However, in few-shot classification and segmentation (FS-CS), this potential remains underutilised. Commonly, self-supervised vision models have been employed, particularly in weakly-supervised scenarios, to generate pseudo-segmentation masks, as ground truth masks are typically unavailable and only target classification is provided. Despite their success, such models find it difficult to capture accurate semantics when compared to vision-language models. To address this limitation, we propose a novel FS-CS approach that leverages the rich semantic alignment of vision-language models to generate more precise pseudo ground-truth masks. While current vision-language models excel in global visual-text alignment, they struggle with finer, patch-level alignment, which is crucial for detailed segmentation tasks. To overcome this, we introduce a method that enhances patch-level alignment without requiring additional training. In addition, existing FS-CS frameworks typically lacks multi-scale information, limiting their ability to capture fine and coarse features simultaneously. To overcome this, we incorporate a module based on atrous convolutions to inject multi-scale information into the feature maps. Together, these contributions - text enhanced pseudo-mask generation and improved multi-scale feature representation - significantly boost the performance of our model in weakly-supervised settings, surpassing state-of-the-art methods and demonstrating the importance of integrating multi-modal information for robust FS-CS solutions. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
ICASSP | 4 |
| 2025 | Adaptive Hyper-Graph Convolution Network for Skeleton-Based Human Action Recognition with Virtual ConnectionsabstractThe shared topology of human skeletons motivated the recent investigation of graph convolutional network (GCN) solutions for action recognition. However, most of the existing GCNs rely on the binary connection of two neighboring vertices (joints) formed by an edge (bone), overlooking the potential of constructing multi-vertex convolution structures. Although some studies have attempted to utilize hyper-graphs to represent the topology, they rely on a fixed construction strategy, which limits their adaptivity in uncovering the intricate latent relationships within the action. In this paper, we address this oversight and explore the merits of an adaptive hyper-graph convolutional network (Hyper-GCN) to achieve the aggregation of rich semantic information conveyed by skeleton vertices. In particular, our Hyper-GCN adaptively optimises the hyper-graphs during training, revealing the action-driven multi-vertex relations. Besides, virtual connections are often designed to support efficient feature aggregation, implicitly extending the spectrum of dependencies within the skeleton. By injecting virtual connections into hyper-graphs, the semantic clues of diverse action categories can be highlighted. The results of experiments conducted on the NTU-60, NTU-120, and NW-UCLA datasets demonstrate the merits of our Hyper-GCN, compared to the state-of-the-art methods. The code is available at https://github.com/6UOOON9/Hyper-GCN. Youwei Zhou, Tianyang Xu 0001, Cong Wu 0006, Xiaojun Wu 0001, Josef Kittler |
ICCV | 5 |
| 2025 | Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated LearningabstractBatch Normalisation (BN) is widely used in conventional deep neural network training to harmonise the input-output distributions for each batch of data.
However, federated learning, a distributed learning paradigm, faces the challenge of dealing with non-independent and identically distributed data among the client nodes.
Due to the lack of a coherent methodology for updating BN statistical parameters, standard BN degrades the federated learning performance.
To this end, it is urgent to explore an alternative normalisation solution for federated learning.
In this work, we resolve the dilemma of the BN layer in federated learning by developing a customised normalisation approach, Hybrid Batch Normalisation (HBN).
HBN separates the update of statistical parameters (*i.e.*, means and variances used for evaluation) from that of learnable parameters (*i.e.*, parameters that require gradient updates), obtaining unbiased estimates of global statistical parameters in distributed scenarios.
In contrast with the existing solutions, we emphasise the supportive power of global statistics for federated learning.
The HBN layer introduces a learnable hybrid distribution factor, allowing each computing node to adaptively mix the statistical parameters of the current batch with the global statistics.
Our HBN can serve as a powerful plugin to advance federated learning performance.
It reflects promising merits across a wide range of federated learning settings, especially for small batch sizes and heterogeneous data.
Code is available at https://github.com/Hongyao-Chen/HybridBN. Hongyao Chen, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
ICML | 4 |
| 2025 | SymCL: Riemannian Contrastive Learning on the Symmetric Positive Definite Manifold for Visual ClassificationabstractSymmetric Positive Definite (SPD) matric has been proven to be an effective feature descriptor in the realm of artificial intelligence, as it can encode spatiotemporal statistical information of data on a curved Riemannian manifold, i.e., SPD manifold. Although existing Riemannian neural networks have demonstrated superiority in many scientific fields, the inherent reliance on labels within supervised learning renders them susceptible to label errors. Besides, it is insufficient to depend solely on labels to learn effective feature distributions in some complicated data scenarios. Drawing inspiration from the considerable achievements of contrastive learning (CL) across diverse tasks, we extend the conventional CL paradigm to the context of SPD manifolds, which we denote SymCL, paving the way for a novel approach in SPD matrix-based visual classification. Furthermore, we inject a Riemannian triplet loss-based Riemannian metric learning (RML) into the designed SPD manifold CL framework for the sake of improving the discrimination of the learned geometric representations. Extensive experimental results on four datasets verify the effectiveness of the proposed algorithm. Yusheng Bao, Rui Wang 0050, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
IJCNN | 5 |
| 2025 | Learning a Discriminative Grassmannian Neural Network for Visual ClassificationabstractLearning representations on the Grassmannian manifold is popular in quite a few visual classification tasks. With the development of deep learning techniques, several neural networks have recently emerged for processing subspace data. However, the diversely changed appearance of the signal data (video clips and image sets), makes it impossible for the existing Grassmannian networks (GrasNets) that rely on a single cross-entropy loss for end-to-end training to learn effective geometric representations, especially for complicated visual scenarios. To solve this problem, a Riemannian triplet loss-based Riemannian metric learning mechanism is introduced to the original GrasNet, which can explicitly encode and learn the characteristics of the intra- and inter-class data distributions conveyed by the input data during network training. Additionally, given the existence of intra-class diversity and inter-class ambiguity of the input data, we propose a hard sample reward strategy (HSR) to further improve the discriminability of the learned network embedding. Extensive experimental results obtained on four benchmarking datasets demonstrate the effectiveness of the proposed method. Yusheng Bao, Rui Wang 0050, Tianyang Xu 0001, Xiaojun Wu 0001, Umapada Pal 0001, Josef Kittler |
IJCNN | 6 |
| 2025 | Ingredients-Guided and Nutrients-Prompted Network for Food Nutrition EstimationabstractFood plays a vital role in human health, and accurate nutrition estimation is crucial for guiding healthy dietary choices. Traditional biochemical-based assessment methods are often inefficient, costly, and impractical for daily use. With the continuous progress in computer vision, some vision-based nutrition estimation approaches have emerged, typically relying on RGB images alone or in combination with depth images to infer nutritional information. These methods have achieved promising performance and garnered considerable attention. However, these methods often ignore visually imperceptible ingredients such as oil, sugar, and salt, which may significantly influence the estimation of nutritional content. Besides, existing methods lack explicit mechanisms for modeling nutrient-specific information and guiding attention toward nutrition-relevant semantics. To solve the above two issues, we propose a novel ingredients-guided and nutrients-prompted nutrition estimation method. Our method adopts multi-scale feature fusion and integrates RGB and depth modalities to enhance visual representation learning. To account for invisible ingredients, we introduce an ingredients-guided strategy, which enhances the sensitivity to non-visible nutritional factors. Moreover, a nutrient-prompt mechanism is introduced to explicitly guide the focus of the model toward nutrient-relevant attributes during estimation. We validate our method on Nutrition5k, where it consistently outperforms existing state-of-the-art methods, demonstrating its efficacy. Donglin Zhang 0001, Boyuan Ma, Xiaojun Wu 0001, Josef Kittler |
ACM Multimedia | 4 |
| 2025 | Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and BenchmarkingabstractUnifying multiple multi-modal visual object tracking (MMVOT) tasks draws increasing attention due to the complementary nature of different modalities in building robust tracking systems. Existing practices mix all data sensor types in a single training procedure, structuring a parallel paradigm from the data-centric perspective and aiming for a global optimum on the joint distribution of the involved tasks. However, the absence of a unified benchmark where all types of data coexist forces evaluations on separated benchmarks, causing inconsistency between training and testing, thus leading to performance degradation. To address these issues, this work advances in two aspects: A unified benchmark, coined as UniBench300, is introduced to bridge the inconsistency by incorporating multiple task data, reducing inference passes from three to one and cutting time consumption by 27%. The unification process is reformulated in a serial format, progressively integrating new tasks. In this way, the performance degradation can be specified as knowledge forgetting of previous tasks, which naturally aligns with the philosophy of continual learning (CL), motivating further exploration of injecting CL into the unification process. Extensive experiments conducted on two baselines and four benchmarks demonstrate the significance of UniBench300 and the superiority of CL in supporting a stable unification process. Moreover, while conducting dedicated analyses, the performance degradation is found to be negatively correlated with network capacity. Additionally, modality discrepancies contribute to varying degradation levels across tasks (RGBT > RGBD > RGBE in MMVOT), offering valuable insights for future multi-modal vision research. Source codes and the proposed benchmark is available at https://github.com/Zhangyong-Tang/UniBench300. Zhangyong Tang, Tianyang Xu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Tao Zhou 0002, Xiaojun Wu 0001, Josef Kittler |
ACM Multimedia | 7 |
| 2025 | CG-SSL: Concept-Guided Self-Supervised LearningabstractHumans understand visual scenes by first capturing a global impression and then refining this understanding into distinct, object-like components. Inspired by this process, we introduce \textbf{C}oncept-\textbf{G}uided \textbf{S}elf-\textbf{S}upervised \textbf{L}earning (CG-SSL), a novel framework that brings structure and interpretability to representation learning through a curriculum of three training phases: (1) global scene encoding, (2) discovery of visual concepts via tokenised cross-attention, and (3) alignment of these concepts across views.
Unlike traditional SSL methods, which simply enforce similarity between multiple augmented views of the same image, CG-SSL accounts for the fact that these views may highlight different parts of an object or scene. To address this, our method establishes explicit correspondences between views and aligns the representations of meaningful image regions. At its core, CG-SSL augments standard SSL with a lightweight decoder that learns and refines concept tokens via cross-attention with patch features. The concept tokens are trained using masked concept distillation and a feature-space reconstruction objective. A final alignment stage enforces view consistency by geometrically matching concept regions under heavy augmentation, enabling more compact, robust, and disentangled representations of scene regions.
Across multiple backbone sizes, CG-SSL achieves state-of-the-art results on image segmentation benchmarks using $k$-NN and linear probes, substantially outperforming prior methods and approaching, or even surpassing, the performance of leading SSL models trained on over $100\times$ more data. Code and pretrained models will be released. Sara Atito Ali Ahmed, Josef Kittler, Muhammad Imran Razzak, Muhammad Awais 0001 |
NeurIPS | 2 |
| 2025 | Collaborating Vision, Depth, and Thermal Signals for Multi-Modal Tracking: Dataset and AlgorithmabstractExisting multi-modal object tracking approaches primarily focus on dual-modal paradigms, such as RGB-Depth or RGB-Thermal, yet remain challenged in complex scenarios due to limited input modalities. To address this gap, this work introduces a novel multi-modal tracking task that leverages three complementary modalities, including visible RGB, Depth (D), and Thermal Infrared (TIR), aiming to enhance robustness in complex scenarios. To support this task, we construct a new multi-modal tracking dataset, coined RGBDT500, which consists of 500 videos with synchronised frames across the three modalities. Each frame provides spatially aligned RGB, depth, and thermal infrared images with precise object bounding box annotations.Furthermore, we propose a novel multi-modal tracker, dubbed RDTTrack.RDTTrack integrates tri-modal information for robust tracking by leveraging a pretrained RGB-only tracking model and prompt learning techniques.In specific, RDTTrack fuses thermal infrared and depth modalities under a proposed orthogonal projection constraint, then integrates them with RGB signals as prompts for the pre-trained foundation tracking model, effectively harmonising tri-modal complementary cues.The experimental results demonstrate the effectiveness and advantages of the proposed method, showing significant improvements over existing dual-modal approaches in terms of tracking accuracy and robustness in complex scenarios. The dataset and source code are publicly available at https://xuefeng-zhu5.github.io/RGBDT500. Xuefeng Zhu 0003, Tianyang Xu 0001, Yifan Pan, Jinjie Gu, Xi Li 0001, Jiwen Lu, Xiaojun Wu 0001, Josef Kittler |
NeurIPS | 8 |
| 2025 | Learning adaptive detection and tracking collaborations with augmented UAV synthesis for accurate anti-UAV system
Shihan Liu, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler |
Expert Syst. Appl. | 5 |
| 2025 | FusionBooster: A Unified Image Fusion Boosting Paradigm
Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Hui Li 0037, Xi Li 0001, Josef Kittler |
Int. J. Comput. Vis. | 6 |
| 2025 | SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion
Huan Kang, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Chunyang Cheng, Josef Kittler |
Int. J. Comput. Vis. | 7 |
| 2025 | Investigating Self-Supervised Methods for Label-Efficient LearningabstractAbstract Vision transformers combined with self-supervised learning have enabled the development of models which scale across large datasets for several downstream tasks, including classification, segmentation, and detection. However, the potential of these models for low-shot learning across several downstream tasks remains largely under explored. In this work, we conduct a systematic examination of different self-supervised pretext tasks, namely contrastive learning, clustering, and masked image modelling, to assess their low-shot capabilities by comparing different pretrained models. In addition, we explore the impact of various collapse avoidance techniques, such as centring, ME-MAX, and sinkhorn, on these downstream tasks. Based on our detailed analysis, we introduce a framework that combines mask image modelling and clustering as pretext tasks. This framework demonstrates superior performance across all examined low-shot downstream tasks, including multi-class classification, multi-label classification and semantic segmentation. Furthermore, when testing the model on large-scale datasets, we show performance gains in various tasks. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Correction: Investigating Self-Supervised Methods for Label-Efficient Learning
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Pure anomaly detection via self-supervised deep metric learning with adaptive marginabstractWe address the problem of anomaly detection (AD) by a deep network pretrained using self-supervised learning for an auxiliary geometric transformation (GT) classification task. Our key contribution is a novel loss function that augments the standard cross-entropy by an additional term that plays a significant role in the later stages of self-supervised learning. The proposed enabling innovation is a triplet centre loss with an adaptive margin and a learnable metric, which relentlessly drives the GT classes to exhibit continuously improving compactness and inter-class separation. The pretrained network is finetuned for the downstream task using non-anomalous data only, and a GT model for the data is constructed. Anomalies are detected by fusing the output of several decision functions defined using the learnt GT class model. In contrast to the majority of existing methods, our approach strictly adheres to the pure AD design philosophy, which relies on the use of purely non-anomalous data for the design. Extensive experiments on four publicly available AD datasets demonstrate the effectiveness of the proposed contributions and lead to significant performance gains compared to the state-of-the-art (1.8% on F-MNIST, 1.0% on CIFAR-10, 1.2% on CIFAR-100, and 1.7% on CatVsDog). https://github.com/12sf12/Deep-Anomaly-Detection • A three-stage pure anomaly detection framework using self-supervised learning. • A novel loss addressing margin and metric selection issues in triplet-based losses. • We propose a method to compute adaptive margins per sample in each mini-batch. • Experiments show the proposed method outperforms SOTA across all datasets. Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler |
Neurocomputing | 4 |
| 2025 | TENet: Targetness entanglement incorporating with multi-scale pooling and mutually-guided fusion for RGB-E object tracking
Pengcheng Shao, Tianyang Xu 0001, Zhangyong Tang, Linze Li 0002, Xiaojun Wu 0001, Josef Kittler |
Neural Networks | 6 |
| 2025 | Adaptive pooling with dual-stage fusion for skeleton-based action recognition
Cong Wu 0006, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
Neural Networks | 4 |
| 2025 | MMDG-DTI: Drug-target interaction prediction via multimodal feature fusion and domain generalization
Yang Hua 0002, Zhenhua Feng 0001, Xiaoning Song, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 5 |
| 2025 | Which images can be effectively learnt from self-supervised learning?
Michalis Lazarou, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
Pattern Recognit. Lett. | 4 |
| 2025 | Temporal aggregation for real-time RGBT tracking via fast decision-level fusion
Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. Lett. | 4 |
| 2025 | M3Track: Meta-Prompt for Multi-Modal TrackingabstractPrompt-tuning has shown remarkable success in multi-modal visual tracking, which enhances RGB tracking by incorporating an additional modality,e.g., thermal infrared (T), Depth (D), or Event (E), forming an RGB+X tracking paradigm. However, in the testing phase, current approaches utilise a frozen prompt for the entire benchmark, failing to account for the diversity and unique characteristics of individual videos. To address this issue, we inject meta-learning solutions into current prompt-based tracking technique, thereby emphasising sequence-level adaptation. Unlike traditional prompt-based trackers, which keep the parameters of prompt blocks fixed during testing, our approach updates these parameters in the first frame of each video using a meta-learning solution. This allows for enhanced discriminative tracking capabilities tailored to each video. Building on this advancement, we extend our methodology beyond separate implementations for RGBT, RGBD, and RGBE tasks. A unified multi-modal tracker is further derived, resulting in the first unified tracker without any task priors (notification of task type) employed in both training and testing phases. Extensive experimental results on LasHeR, DepthTrack, VisEvent, GTOT, RGBT234, and RGBD1K consistently demonstrate the superiority of the proposed method against the existing prompt-tuning paradigm. Source codes are available athttps://github.com/Zhangyong-Tang/M3Track. Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Signal Process. Lett. | 4 |
| 2025 | Adaptive Colour-Depth Aware Attention for RGB-D Object TrackingabstractRecent advances in RGB-D tracking have been driven by the synergistic combination of high-performing RGB-only trackers and auxiliary depth information. However, most existing methods rely on visual feature descriptors to extract depth features, which are then fused with vision features. This pipeline may lead to performance degradation due to the incongruence between the RGB and depth modalities. In this letter, we propose an efficient and effective transformer-based framework, that explicitly models colour and depth information for RGB-D tracking. Specifically, we first statistically code the colour and depth information of the foreground and background for the template. Then, the spatial attention maps of the search region are obtained using these colour-depth statistical models, enhancing the visual features of the search region for improved object localisation accuracy. The comprehensive experimental results obtained on multiple benchmarks demonstrate the effectiveness and merits of the proposed approach in explicit colour-depth coding for RGB-D tracking. The code and the models are publicly accessible athttps://github.com/xuefeng-zhu5/CDAAT. Xuefeng Zhu 0003, Tianyang Xu 0001, Xiaojun Wu 0001, Zhenhua Feng 0001, Josef Kittler |
IEEE Signal Process. Lett. | 5 |
| 2025 | Deep Discriminative Multi-View ClusteringabstractMulti-view clustering based on deep auto-encoder networks has garnered increasing attention and made significant progress in recent years. However, we argue that most existing methods inadequately explore the discriminability while learning clustering assignments, resulting in models struggling to accurately cluster data, particularly those with ambiguous semantics. To address this problem, we propose a novel framework termed deep discriminative multi-view clustering (DDMvC). This framework is designed to further increase the inter-cluster distances by learning a discriminative projection dictionary with global prior information. To begin with, we enhance the reliability of the dictionary atoms by initializing them with class-specific prototypes derived from concatenated global features across multiple views. Subsequently, we iteratively refine the atoms to guarantee their independence from any specific cluster. Simultaneously, we incorporate contrastive learning for the cluster assignments projected by these atoms, striving for inter-view consistent clustering results. Experimental results on benchmark multi-view datasets demonstrate that our framework achieves the state-of-the-art clustering performance. Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Hui Li 0037, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Revisiting RGBT Tracking Benchmarks From the Perspective of Modality Validity: A New Benchmark, Problem, and SolutionabstractRGBT tracking draws increasing attention because of its robustness in multi-modal warranting (MMW) scenarios, such as nighttime and adverse weather conditions, where relying on a single sensing modality fails to ensure stable tracking results. However, existing benchmarks predominantly contain videos collected in common scenarios where both RGB and thermal infrared (TIR) information are of sufficient quality. This weakens the representativeness of existing benchmarks in severe imaging conditions, leading to tracking failures in MMW scenarios. To bridge this gap, we present a new benchmark considering the modality validity, MV-RGBT, captured specifically from MMW scenarios where either RGB (extreme illumination) or TIR (thermal truncation) modality is invalid. Hence, it is further divided into two subsets according to the valid modality, offering a new compositional perspective for evaluation and providing valuable insights for future designs. Moreover, MV-RGBT is the most diverse benchmark of its kind, featuring 36 different object categories captured across 19 distinct scenes. Furthermore, considering severe imaging conditions in MMW scenarios, a new problem is posed in RGBT tracking, named 'when to fuse', to stimulate the development of fusion strategies for such scenarios. To facilitate its discussion, we propose a new solution with a mixture of experts, named MoETrack, where each expert generates independent tracking results along with a confidence score. Extensive results demonstrate the significant potential of MV-RGBT in advancing RGBT tracking and elicit the conclusion that fusion is not always beneficial, especially in MMW scenarios. Besides, MoETrack achieves state-of-the-art results on several benchmarks, including MV-RGBT, GTOT, and LasHeR. Source codes and benchmarks are available at https://github.com/Zhangyong-Tang/MVRGBT. Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Zhenhua Feng 0001, Josef Kittler |
IEEE Trans. Image Process. | 7 |
| 2025 | DFL-Net: Disentangled Feature Learning Network for Multi-View ClusteringabstractMulti-view clustering aims at partitioning data into their underlying categories by mining shared and complementary information conveyed by different views. Although the integration of deep learning and disentanglement learning has markedly improved clustering performance, our analysis reveals two fundamental limitations in existing approaches: inadequate separation between view-shared and view-exclusive features; and the negative effects of clustering-irrelevant information on feature decoupling. To tackle these issues, we present a novel Disentangled Feature Learning Network (DFL-Net), which utilizes a progressive learning framework to systematically disentangle features. DFL-Net initially establishes view-shared representations through semantic disparity minimization, followed by the construction of orthogonal feature subspaces using cross-view and intra-view independence constraints to isolate view-specific features. Subsequently, DFL-Net enforces clustering consistency across views to adaptively eliminate irrelevant information, thus enhancing the overall effectiveness of disentanglement learning. The framework introduces two significant innovations: a comprehensive feature independence criterion that concurrently reduces intra-view and cross-view feature dependencies, and an irrelevance filtering mechanism that ensures cross-view clustering consistency. Extensive experiments on benchmark datasets demonstrate the superior performance of DFL-Net compared to state-of-the-art methods. Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Learning a Better SPD Network for Signal Classification: A Riemannian Batch Normalization MethodabstractSymmetric positive definite (SPD) matrices have been widely used as Riemannian feature descriptors in various scientific fields, due to their capacity to encode effective manifold-valued representations. Inspired by the architectural principles of Euclidean deep learning, the emerging SPD neural networks have achieved more robust signal classification. Among these advancements, Riemannian batch normalization (RBN) based on the affine-invariant Riemannian metric (AIRM) has emerged as a key technique for enhancing the learning capability of SPD-based networks. Nevertheless, the reliance of singular value decomposition (SVD) makes this metric relatively unstable for the computation of SPD matrices, especially for the ill-conditioned case. To address this limitation, we propose a novel RBN algorithm based on the recently introduced log-Cholesky metric (LCM), which leverages Cholesky decomposition. Unlike AIRM, the LCM offers enhanced numerical stability and allows for more efficient computation. Specifically, the LCM-based Riemannian operators such as Fr $\acute {\mathrm {e}}$ chet mean and parallel transport (PT) are much simpler than those of AIRM, and both have closed forms. Besides, since LCM is the pullback metric from the Cholesky manifold via Cholesky decomposition, the LCM-based RBN on the SPD manifold can be computed in the Cholesky manifold, further boosting the efficiency. Extensive experiments conducted on four benchmarking datasets certify the effectiveness of our proposed algorithm. The source code is now available at: https://github.com/jjscc/CBN.git. Rui Wang 0050, Shaocheng Jin, Zhenyu Cai, Ziheng Chen 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Generative-Based Fusion Mechanism for Multi-Modal TrackingabstractGenerative models (GMs) have received increasing research interest for their remarkable capacity to achieve comprehensive understanding. However, their potential application in the domain of multi-modal tracking has remained unexplored. In this context, we seek to uncover the potential of harnessing generative techniques to address the critical challenge, information fusion, in multi-modal tracking. In this paper, we delve into two prominent GM techniques, namely, Conditional Generative Adversarial Networks (CGANs) and Diffusion Models (DMs). Different from the standard fusion process where the features from each modality are directly fed into the fusion block, we combine these multi-modal features with random noise in the GM framework, effectively transforming the original training samples into harder instances. This design excels at extracting discriminative clues from the features, enhancing the ultimate tracking performance. Based on this, we conduct extensive experiments across two multi-modal tracking tasks, three baseline methods, and four challenging benchmarks. The experimental results demonstrate that the proposed generative-based fusion mechanism achieves state-of-the-art performance by setting new records on GTOT, LasHeR and RGBD1K. Code will be available at https://github.com/Zhangyong-Tang/GMMT. Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003, Josef Kittler |
AAAI | 5 |
| 2024 | SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action RecognitionabstractContrastive learning has achieved great success in skeleton-based action recognition. However, most existing approaches encode the skeleton sequences as entangled spatiotemporal representations and confine the contrasts to the same level of representation. Instead, this paper introduces a novel contrastive learning framework, namely Spatiotemporal Clues Disentanglement Network (SCD-Net). Specifically, we integrate the decoupling module with a feature extractor to derive explicit clues from spatial and temporal domains respectively. As for the training of SCD-Net, with a constructed global anchor, we encourage the interaction between the anchor and extracted clues. Further, we propose a new masking strategy with structural constraints to strengthen the contextual associations, leveraging the latest development from masked image modelling into the proposed SCD-Net. We conduct extensive evaluations on the NTU-RGB+D (60&120) and PKU-MMD (I&II) datasets, covering various downstream tasks such as action recognition, action retrieval, transfer learning, and semi-supervised learning. The experimental results demonstrate the effectiveness of our method, which outperforms the existing state-of-the-art (SOTA) approaches significantly. Our code and supplementary material can be found at https://github.com/cong-wu/SCD-Net. Cong Wu 0006, Xiaojun Wu 0001, Josef Kittler, Tianyang Xu 0001, Muhammad Awais 0001, Zhenhua Feng 0001 |
AAAI | 3 |
| 2024 | Pseudo Labelling for Enhanced Masked Auto Encoders
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
BMVC | 4 |
| 2024 | Enhancing Radiology Report Generation: The Impact of Locally Grounded Vision and Language Training
Sergio Sánchez Santiesteban, Muhammad Awais 0001, Yi-Zhe Song, Josef Kittler |
BMVC | 4 |
| 2024 | C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
Rongchang Li 0001, Zhenhua Feng 0001, Tianyang Xu 0001, Linze Li 0002, Xiaojun Wu 0001, Muhammad Awais 0001, Sara Atito Ali Ahmed, Josef Kittler |
ECCV (38) | 8 |
| 2024 | Efficient Few-Shot Action Recognition via Multi-level Post-reasoning
Cong Wu 0006, Xiaojun Wu 0001, Linze Li 0002, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler |
ECCV (3) | 6 |
| 2024 | Improved Image Captioning Via Knowledge Graph-Augmented ModelsabstractMultimodal foundation models, pre-trained on large-scale data, effectively capture vast amounts of factual and commonsense knowledge. However, these models store all their knowledge within their parameters, requiring increasingly larger models and training data to capture more knowledge. To address this limitation and achieve a more scalable and modular integration of knowledge, we propose a novel knowledge graph-augmented multimodal model. This approach enables a base multimodal model to access pertinent information from an external knowledge graph. Our methodology leverages existing general domain knowledge to facilitate vision-language pre-training using paired images and text descriptions. We conduct comprehensive evaluations demonstrating that our model outperforms state-of-the-art models and yields comparable results to much larger models trained on more extensive datasets. Notably, our model reached a 145 Cider score on MS COCO Captions using only 2.9 million samples, outperforming a 1.4B parameter model by 1.7% despite having 11 times fewer parameters. Sergio Sánchez Santiesteban, Sara Atito Ali Ahmed, Muhammad Awais 0001, Yi-Zhe Song, Josef Kittler |
ICASSP | 5 |
| 2024 | SS-CXR: Self-Supervised Pretraining Using Chest X-Rays Towards A Domain Specific Foundation ModelabstractChest X-rays (CXRs) are widely used imaging modality for the diagnosis and prognosis of lung disease. There is a large body of work where machine learning algorithms are developed for specific tasks. However, the traditional diagnostic tool design methods based on supervised learning are burdened by the need to provide training data annotation, which should be of good quality for better clinical outcomes. Here, we propose an alternative solution, a new self-supervised paradigm, where a general representation from CXRs is learned using a group-masked self-supervised framework. The pre-trained model is then fine-tuned for domain-specific tasks such as covid-19, pneumonia detection, and general health screening. We show that the same pre-training can be used for the lung segmentation task. Our proposed paradigm shows robust performance in multiple downstream tasks which demonstrates the success of the pre-training. Moreover, the performance of the pre-trained models on data with significant drift during test time proves the learning of a better generic representation. The methods are further validated by covid-19 detection in a unique small-scale pediatric data set. The performance gain ($\sim 25 \%$) is significant when compared to a supervised transformer-based method. This adds credence to the strength and reliability of our proposed framework and pre-training strategy. Syed Muhammad Anwar, Abhijeet Parida, Sara Atito Ali Ahmed, Muhammad Awais 0001, Gustavo Nino, Josef Kittler, Marius George Linguraru |
ICIP | 6 |
| 2024 | Investigating Self-Supervised Methods for Label-Efficient LearningabstractVision transformers combined with self-supervised learning have enabled the development of models which scale across large datasets for several downstream tasks like classification, segmentation and detection. The low-shot learning capability of these models, across several low-shot downstream tasks, has been largely under explored. We perform a system level study of different self supervised pretext tasks, namely contrastive learning, clustering, and masked image modelling for their low-shot capabilities by comparing the pretrained models. In addition we also study the effects of collapse avoidance methods, namely centring, ME-MAX, sinkhorn, on these downstream tasks. Based on our detailed analysis, we introduce a framework involving both mask image modelling and clustering as pretext tasks, which performs better across all low-shot downstream tasks, including multi-class classification, multi-label classification and semantic segmentation. Furthermore, when testing the model on full scale datasets, we show performance gains in multi-class classification, multi-label classification and semantic segmentation. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
ICIP | 4 |
| 2024 | Masked Momentum Contrastive Learning for Semantic Understanding by ObservationabstractLarge language models (LLMs) have shown excellent performance in zero-shot learning using natural language prompts. However, in the domain of computer vision (CV), the paradigm of pretraining followed by finetuning remains dominant. The aim of this study is to reduce this gap by utilizing the capability of Self-Supervised Learning (SSL) in semantic understanding for zero-shot segmentation, without relying on human-provided labels or vision-language supervision. We introduce a novel evaluation framework that employs visual prompts, including a threshold and a query patch. This framework evaluates the ability of SSL models to derive concepts from observational data. Through this evaluation, we identify the strengths and limitations of SSL models in understanding semantics. Building on the insights from various SSL methods, we further propose the MMC approach to enhance the representations for objects, which integrates Masked image modeling, Momentum-based self-distillation, and global Contrastive learning. MMC achieves a better balance between the inter-object discriminability and the intra-object compactness of learned features. Our experiments on COCO, DAVIS-2017, PASCAL VOC, and ADE20K demonstrate outstanding performance of MMC’s representations. Jiantao Wu, Shentong Mo, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Syed Sameed Husain, Muhammad Awais 0001 |
ICIP | 5 |
| 2024 | Spatio-Temporal Domain-Aware Network for Skeleton-Based Action Representation Learning
Jiannan Hu, Cong Wu 0006, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
ICPR (29) | 5 |
| 2024 | StableTalk: Advancing Audio-to-Talking Face Generation with Stable Diffusion and Vision Transformer
Fatemeh Nazarieh, Josef Kittler, Muhammad Awais 0001, Diptesh Kanojia, Zhenhua Feng 0001 |
ICPR (6) | 2 |
| 2024 | Learning Explicit Modulation Vectors for Disentangled Transformer Attention-Based RGB-D Visual Tracking
Yifan Pan, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaoqing Luo, Xiaojun Wu 0001, Josef Kittler |
ICPR (16) | 6 |
| 2024 | Attention-Based Patch Matching and Motion-Driven Point Association for Accurate Point Tracking
Han Zang, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaoning Song, Xiaojun Wu 0001, Josef Kittler |
ICPR (16) | 6 |
| 2024 | Multi-frequency Fine-Grained Matching for Audio-Visual Segmentation
Yinhao Zhang, Tianyang Xu 0001, Xiaojun Wu 0001, Shao-Chuan Zhao, Josef Kittler |
ICPR (29) | 5 |
| 2024 | A Riemannian Residual Learning Mechanism for SPD NetworkabstractThe generalization of Euclidean network paradigm to the Riemannian manifolds has attracted much attention for offering useful geometric representations in processing manifold-valued data in recent years. However, the information degradation during data compression mapping hinders Riemannian networks from going deeper, and there are very few solutions specifically designed for this problem. Given the remarkable success of deep Residual learning in Euclidean networks, a novel Riemannian residual learning mechanism (RRLM) is proposed in the context of Symmetric Positive Definite (SPD) manifolds, enabling the characterization of deep spatiotemporal features while preserving the manifold properties. Based on RRLM, a stack of SPD manifold-constrained residual-like blocks is designed on the tail of the original SPDNet(backbone) for the sake of conducting deep Riemannian residual learning. For simplicity, we refer to the network architecture introduced above as Riemannian residual SPD network (ResSPDNet). The experimental results achieved on three types of visual classification tasks, i.e., facial emotion recognition, drone recognition, and action recognition, demonstrate that our method can achieve improved accuracy with a deepened network structure. Zhenyu Cai, Rui Wang 0050, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
IJCNN | 5 |
| 2024 | MVRMLM 2024: Multimodal Video Retrieval and Multimodal Language ModellingabstractAs the proliferation of video content continues, and many video archives lack suitable metadata, therefore, video retrieval, particularly through example-based search, has become increasingly crucial. Existing metadata often fails to meet the needs of specific types of searches, especially when videos contain elements from different modalities, such as visual and audio. Consequently, developing video retrieval methods that can handle multi-modal content is essential. In designing our novel video retrieval framework named Multi-modal Video Search by Examples (MVSE)1, we focused on accuracy (precision and recall), efficiency (retrieval time in seconds), interactivity, and extensibility, with key components including advanced data processing and a user-friendly interface aimed at enhancing search effectiveness and user experience. With the advent of Large Language Models (LLMs), the interaction between multimodal data, including image and audio has been transformed with a significant leap forward towards a bigger goal of artificial general intelligence. This workshop aims to bring together experts from diverse domains to explore the possibilities of developing novel ways of multimodal data search, understanding and interaction. Hui Wang 0001, Josef Kittler, Mark J. F. Gales, Rob Cooper, Maurice D. Mulvenna, Wing W. Y. Ng, Yang Hua 0001, Richard Gault, Abbas Haider, Guanfeng Wu |
ICMR | 2 |
| 2024 | MMDRFuse: Distilled Mini-Model with Dynamic Refresh for Multi-Modality Image FusionabstractIn recent years, Multi-Modality Image Fusion (MMIF) has been applied to many fields, which has attracted many scholars to endeavour to improve the fusion performance. However, the prevailing focus has predominantly been on the architecture design, rather than the training strategies. As a low-level vision task, image fusion is supposed to quickly deliver output images for observation and supporting downstream tasks. Thus, superfluous computational and storage overheads should be avoided. In this work, a lightweight Distilled Mini-Model with a Dynamic Refresh strategy (MMDRFuse) is proposed to achieve this objective. To pursue model parsimony, an extremely small convolutional network with a total of 113 trainable parameters (0.44 KB) is obtained by three carefully designed supervisions. First, digestible distillation is constructed by emphasising external spatial feature consistency, delivering soft supervision with balanced details and saliency for the target network. Second, we develop a comprehensive loss to balance the pixel, gradient, and perception clues from the source images. Third, an innovative dynamic refresh training strategy is used to collaborate history parameters and current supervision during training, together with an adaptive adjust function to optimise the fusion network. Extensive experiments on several public datasets demonstrate that our method exhibits promising advantages in terms of model efficiency and complexity, with superior performance in multiple image fusion tasks and downstream pedestrian detection application. The code of this work is publicly available at https://github.com/yanglinDeng/MMDRFuse. Yanglin Deng, Tianyang Xu 0001, Chunyang Cheng, Xiaojun Wu 0001, Josef Kittler |
ACM Multimedia | 5 |
| 2024 | Dynamic Subframe Splitting and Spatio-Temporal Motion Entangled Sparse Attention for RGB-E Tracking
Pengcheng Shao, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler |
PRCV (13) | 5 |
| 2024 | M-adapter: Multi-level image-to-video adaptation for video action recognition
Rongchang Li 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Linze Li 0002, Josef Kittler |
Comput. Vis. Image Underst. | 7 |
| 2024 | Scene adaptive mechanism for action recognition
Cong Wu 0006, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
Comput. Vis. Image Underst. | 4 |
| 2024 | Multi-modal video search by examples - A video quality impact analysisabstractAbstract As the proliferation of video content continues, and many video archives lack suitable metadata, therefore, video retrieval, particularly through example‐based search, has become increasingly crucial. Existing metadata often fails to meet the needs of specific types of searches, especially when videos contain elements from different modalities, such as visual and audio. Consequently, developing video retrieval methods that can handle multi‐modal content is essential. An innovative Multi‐modal Video Search by Examples (MVSE) framework is introduced, employing state‐of‐the‐art techniques in its various components. In designing MVSE, the authors focused on accuracy, efficiency, interactivity, and extensibility, with key components including advanced data processing and a user‐friendly interface aimed at enhancing search effectiveness and user experience. Furthermore, the framework was comprehensively evaluated, assessing individual components, data quality issues, and overall retrieval performance using high‐quality and low‐quality BBC archive videos. The evaluation reveals that: (1) multi‐modal search yields better results than single‐modal search; (2) the quality of video, both visual and audio, has an impact on the query precision. Compared with image query results, audio quality has a greater impact on the query precision (3) a two‐stage search process (i.e. searching by Hamming distance based on hashing, followed by searching by Cosine similarity based on embedding); is effective but increases time overhead; (4) large‐scale video retrieval is not only feasible but also expected to emerge shortly. Guanfeng Wu, Abbas Haider, Xing Tian, Erfan Loweimi, Chi-Ho Chan, Mengjie Qian 0001, Muhammad Junaid Awan, Ivor T. A. Spence, Rob Cooper, Wing W. Y. Ng, Josef Kittler, Mark J. F. Gales, Hui Wang 0001 |
IET Comput. Vis. | 11 |
| 2024 | Learning Feature Restoration Transformer for Robust Dehazing Visual Object Tracking
Tianyang Xu 0001, Yifan Pan, Zhenhua Feng 0001, Xuefeng Zhu 0003, Chunyang Cheng, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 7 |
| 2024 | A Spatio-Temporal Robust Tracker with Spatial-Channel Transformer and Jitter Suppression
Shao-Chuan Zhao, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 4 |
| 2024 | UniMod1K: Towards a More Universal Large-Scale Dataset and Benchmark for Multi-modal Learning
Xuefeng Zhu 0003, Tianyang Xu 0001, Zongtao Liu, Zhangyong Tang, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 6 |
| 2024 | View-shuffled clustering via the modified Hungarian algorithm
Wenhua Dong, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
Neural Networks | 7 |
| 2024 | RAgE: Robust Age Estimation Through Subject Anchoring With Consistency RegularisationabstractModern facial age estimation systems can achieve high accuracy when training and test datasets are identically distributed and captured under similar conditions. However, domain shifts in data, encountered in practice, lead to a sharp drop in accuracy of most existing age estimation algorithms. In this article, we propose a novel method, namely RAgE, to improve the robustness and reduce the uncertainty of age estimates by leveraging unlabelled data through a subject anchoring strategy and a novel consistency regularisation term. First, we propose an similarity-preserving pseudo-labelling algorithm by which the model generates pseudo-labels for a cohort of unlabelled images belonging to the same subject, while taking into account the similarity among age labels. In order to improve the robustness of the system, a consistency regularisation term is then used to simultaneously encourage the model to produce invariant outputs for the images in the cohort with respect to an anchor image. We propose a novel consistency regularisation term the noise-tolerant property of which effectively mitigates the so-called confirmation bias caused by incorrect pseudo-labels. Experiments on multiple benchmark ageing datasets demonstrate substantial improvements over the state-of-the-art methods and robustness to confounding external factors, including subject's head pose, illumination variation and appearance of expression in the face image. Ali Akbari 0003, Muhammad Awais 0001, Soroush Fatemifar, Syed Safwan Khalid, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Importance Weighted Structure Learning for Scene Graph GenerationabstractScene graph generation is a structured prediction task aiming to explicitly model objects and their relationships via constructing a visually-grounded scene graph for an input image. Currently, the message passing neural network based mean field variational Bayesian methodology is the ubiquitous solution for such a task, in which the variational inference objective is often assumed to be the classical evidence lower bound. However, the variational approximation inferred from such loose objective generally underestimates the underlying posterior, which often leads to inferior generation performance. In this paper, we propose a novel importance weighted structure learning method aiming to approximate the underlying log-partition function with a tighter importance weighted lower bound, which is computed from multiple samples drawn from a reparameterizable Gumbel-Softmax sampler. A generic entropic mirror descent algorithm is applied to solve the resulting constrained variational inference task. The proposed method achieves the state-of-the-art performance on various popular scene graph generation benchmarks. Daqi Liu, Miroslaw Bober, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Edwin Hancock
Josef Kittler, Richard C. Wilson 0001 |
Pattern Recognit. | 1 |
| 2024 | Self-supervised learning for RGB-D object tracking
Xuefeng Zhu 0003, Tianyang Xu 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Xiaojun Wu 0001, Zhenhua Feng 0001, Josef Kittler |
Pattern Recognit. | 7 |
| 2024 | Feature enhancement and coarse-to-fine detection for RGB-D tracking
Xuefeng Zhu 0003, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. Lett. | 4 |
| 2024 | ASiT: Local-Global Audio Spectrogram Vision Transformer for Event ClassificationabstractTransformers, which were originally developed for natural language processing, have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships. Constrained by the data hungry nature of transformers and the limited amount of labelled data, most transformer-based models for audio tasks are finetuned from ImageNet pretrained models, despite the huge gap between the domain of natural images and audio. This has motivated the research in self-supervised pretraining of audio transformers, which reduces the dependency on large amounts of labeled data and focuses on extracting concise representations of audio spectrograms. In this paper, we proposeLocal-GlobalAudioSpectrogram vIsionTransformer, namely ASiT, a novel self-supervised learning framework that captures local and global contextual information by employing group masked model learning and self-distillation. We evaluate our pretrained models on both audio and speech classification tasks, including audio event classification, keyword spotting, and speaker identification. We further conduct comprehensive ablation studies, including evaluations of different pretraining strategies. The proposed ASiT framework significantly boosts the performance on all tasks and sets a new state-of-the-art performance in five audio and speech classification tasks, outperforming recent methods, including the approaches that use additional datasets for pretraining. Sara Atito Ali Ahmed, Muhammad Awais 0001, Wenwu Wang 0001, Mark D. Plumbley, Josef Kittler |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | A Survey of Cross-Modal Visual Content GenerationabstractCross-modal content generation has become very popular in recent years. To generate high-quality and realistic content, a variety of methods have been proposed. Among these approaches, visual content generation has attracted significant attention from academia and industry due to its vast potential in various applications. This survey provides an overview of recent advances in visual content generation conditioned on other modalities, such as text, audio, speech, and music, with a focus on their key contributions to the community. In addition, we summarize the existing publicly available datasets that can be used for training and benchmarking cross-modal visual content generation models. We provide an in-depth exploration of the datasets used for audio-to-visual content generation, filling a gap in the existing literature. Various evaluation metrics are also introduced along with the datasets. Furthermore, we discuss the challenges and limitations encountered in the area, such as modality alignment and semantic coherence. Last, we outline possible future directions for synthesizing visual content from other modalities including the exploration of new modalities, and the development of multi-task multi-modal networks. This survey serves as a resource for researchers interested in quickly gaining insights into this burgeoning field. Fatemeh Nazarieh, Zhenhua Feng 0001, Muhammad Awais 0001, Wenwu Wang 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Deep Metric Learning on the SPD Manifold for Image Set ClassificationabstractThanks to the efficacy of Symmetric Positive Definite (SPD) manifold in characterizing video sequences (image sets), image set-based visual classification has made remarkable progress. However, the issue of large intra-class diversity and inter-class similarity is still an open challenge for the research community. Although several recent studies have alleviated the above issue by constructing Riemannian neural networks for SPD matrix nonlinear processing, the degradation of structural information during multi-stage feature transformation impedes them from going deeper. Besides, a single cross-entropy loss is insufficient for discriminative learning as it neglects the peculiarities of data distribution. To this end, this paper develops a novel framework for image set classification. Specifically, we first choose a mainstream neural network built on the SPD manifold (SPDNet)[25]as the backbone with a stacked SPD manifold autoencoder (SSMAE) built on the tail to enrich the structured representations. Due to the associated reconstruction error terms, the embedding mechanism of both SSMAE and each SPD manifold autoencoder (SMAE) forms an approximate identity mapping, simplifying the training of the suggested deeper network. Then, the ReCov layer is introduced with a nonlinear function for the constructed architecture to narrow the discrepancy of the intra-class distributions from the perspective of regularizing the local statistical information of the SPD data. Afterward, two progressive metric learning stages are coupled with the proposed SSMAE to explicitly capture, encode, and analyze the geometric distributions of the generated deep representations during training. In consequence, not only a more powerful Riemannian network embedding but also effective classifiers can be obtained. Finally, a simple maximum voting strategy is applied to the outputs of the learned multiple classifiers for classification. The proposed model is evaluated on three typical visual classification tasks using widely adopted benchmarking datasets. Extensive experiments show its superiority over the state of the arts. Rui Wang 0050, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Motion Complement and Temporal Multifocusing for Skeleton-Based Action RecognitionabstractModeling sequences with spatial-temporal graph convolutional networks has become a mainstream paradigm in skeleton-based action recognition. However, many existing methods adopt redundant or cluttered structures to mine the key action features, thus making it difficult to achieve a balanced or leading performance in accuracy and efficiency. In this paper, we propose a novel framework, referred to as Motion Complement and Temporal Multifocusing Network (MCTM-Net), to capture the relationships within skeleton sequences by means of an efficient decomposition of the spatiotemporal graph model. Specifically, for spatial modeling, we introduce a motion-related relational descriptor that extends the channel dimension so as to enhance the modeling of motion salient regions as a complement to the conventional physical adjacency relationships. An improved parameterized physical relationship model is also proposed to better fit the data characteristics. As for temporal modeling, we propose an efficient multi-focus temporal information acquisition strategy that aggregates the information from multiple temporal spans and adjacent regions. We conduct extensive experiments on multiple representative datasets, including NTU-RGB+D (60&120), Northwestern-UCLA, and UWA3D Multiview Activity II, to validate our innovations. The experimental results show the effectiveness of our method. The code will be available athttps://github.com/cong-wu/MCMT-Net. Cong Wu 0006, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Distillation, Ensemble and Selection for Building a Better and Faster Siamese Based TrackerabstractVisual object tracking has witnessed continuous improvements in performance, thanks to deep CNN learning that recently emerged. More complex CNN models invariably offer better accuracy. However, there is a conflict between the tracking efficiency and model complexity, which poses a challenge in balancing speed against accuracy. To optimize the trade-off between these two performance criteria, a distillation-ensemble-selection framework is proposed in this paper. Without any modification to the baseline network architecture, the proposed approach enables the construction of a Siamese-based tracker with improved capacity and efficiency. Specifically, multiple student trackers are designed by means of knowledge distillation from a given teacher tracking model. To manage the varying granularity of unknown targets, an ensemble module combines the outputs of the student trackers with the help of a learnable fine-grained attention module. Besides, in the online tracking stage, a selection module adaptively controls the complexity of the tracker by identifying an appropriate subset of the candidate tracker models. We verify the effectiveness of the proposed method in both anchor-based and anchor-free paradigms. The experimental results obtained on standard benchmarking datasets demonstrate the effectiveness of the proposed method, with an outstanding and balanced performance in both accuracy and speed. Shao-Chuan Zhao, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Pluggable Attack for Visual Object TrackingabstractPerforming adversarial attacks on a visual tracker aims to drift the apparent target to the background by adding malicious perturbations to the source images. Demonstrating convincingly their ability to decrease accuracy, existing tracking attackers mislead the target predictions at the decision level, but this is tracker design specific, narrowing their applicability to other tracking approaches. In contrast, we advocate that attacks be performed by corrupting the feature-level clues, i.e., the feature representations extracted by deep networks. The proposed approach provides a general attacking framework for backbone-head tracking architectures. Motivated by the knowledge that the quality of intermediate-level features strongly influences the decision making, four intermediate-level attack methods are proposed to maximise the difference between the feature distributions of natural and adversarial samples, thus decoupling the attack strategies from the form of the output of specific victim trackers. Interestingly, our intermediate-level attacks are compatible with existing decision-level attacks, thus a joint optimisation of these two kinds of adversarial objective functions has the potential to achieve better attacking performance. Hence, the proposed adversarial attack methodology can be used in conjunction with several mainstream tracking paradigms (Discriminative correlation filters, Siamese networks, and Transformer trackers), demonstrating its pluggability. The experimental results on four popular benchmarks, e.g., OTB100, UAV123, LaSOT, and TLP, verify that our method can produce impressive and consistent accuracy degeneration. Shao-Chuan Zhao, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Perceiving Actions via Temporal Video Frame PairsabstractVideo action recognition aims at classifying the action category in given videos. In general, semantic-relevant video frame pairs reflect significant action patterns such as object appearance variation and abstract temporal concepts like speed, rhythm, and so on. However, existing action recognition approaches tend to holistically extract spatiotemporal features. Though effective, there is still a risk of neglecting the crucial action features occurring across frames with a long-term temporal span. Motivated by this, in this article, we propose to perceive actions via frame pairs directly and devise a novel Nest Structure with frame pairs as basic units. Specifically, we decompose a video sequence into all possible frame pairs and hierarchically organize them according to temporal frequency and order, thus transforming the original video sequence into a Nest Structure. Through naturally decomposing actions, the proposed structure can flexibly adapt to diverse action variations such as speed or rhythm changes. Next, we devise a Temporal Pair Analysis module (TPA) to extract discriminative action patterns based on the proposed Nest Structure. The designed TPA module consists of a pair calculation part to calculate the pair features and a pair fusion part to hierarchically fuse the pair features for recognizing actions. The proposed TPA can be flexibly integrated into existing backbones, serving as a side branch to capture various action patterns from multi-level features. Extensive experiments show that the proposed TPA module can achieve consistent improvements over several typical backbones, reaching or updating CNN-based SOTA results on several challenging action recognition benchmarks. Rongchang Li 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2024 | One-pass View-unaligned ClusteringabstractGiven a set of multi-view instances, the prevailing assumption in most existing clustering approaches is that they are complete and exhibit cross-view alignment. However, this assumption is often unrealistic. In such scenarios, it could be satisfied at the cost of data pre-processing, but this would be complex and inconsistent with practical applications. Therefore, developing more effective solutions for the View-unaligned Problem (VuP) is highly desirable. Several pioneering works have tackled the partially VuP, yet handling fully VuP remains a challenge due to the reliance on partially pre-aligned instances. In this paper, we propose One-pass View-unaligned Clustering (OpVuC) that simultaneously aligns and clusters instances in a unified framework. Specifically, we alig shuffled instances with a selected template using an innovative global-local alignment scheme based on the notion of geometric invariance and separate the fully aligned instances using a relaxed$k$-means algorithm. The proposed OpVuC method can handle VuP at any alignment level without requiring any pre-aligned instances. Extensive experiments conducted on several benchmark datasets demonstrate the effectiveness and merits of the proposed OpVuC method. Wenhua Dong, Xiaojun Wu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
IEEE Trans. Multim. | 6 |
| 2024 | Lightweight Multiperson Pose Estimation With Staggered Alignment Self-DistillationabstractAccurate 2D human pose estimation from images is vital for understanding human actions. However, deploying the latest models, e.g., regression-based models, on resource-limited devices remains challenging due to their high computational requirements. In this paper, we address the resolution dilemma in regression-based multiperson pose estimation, where low-resolution inputs cause performance degradation, while high-resolution inputs drastically increase computational costs. To achieve a lightweight regression approach, it becomes crucial to enhance the model's capabilities in low-resolution scenarios. We propose the staggered alignment self-distillation (SASD) method and a corresponding network architecture. Our approach involves training two twin networks with shared weights: a high-resolution network and a low-resolution network. The high-resolution network serves as a teacher, guiding the learning process of the low-resolution network through feature map staggered alignment. The knowledge from the high-resolution network enhances the performance of the low-resolution network during low-resolution inference. Additionally, we employ a normalized skeleton loss to capture the loss of bone-related structure during training. Through extensive experiments on the MS-COCO and CrowdPose datasets, we demonstrate the superiority of our proposed method over state-of-the-art, lightweight multiperson pose estimation techniques, achieving much better performance with lower computational costs. Furthermore, our method achieves comparable performance to recent advanced regression-based pose estimation methods but with only 1/4 of the computational cost. Zhenkun Fan, Zhuoxu Huang, Zhixiang Chen 0003, Tao Xu 0038, Jungong Han, Josef Kittler |
IEEE Trans. Multim. | 6 |
| 2024 | SPD Manifold Deep Metric Learning for Image Set ClassificationabstractBy characterizing each image set as a nonsingular covariance matrix on the symmetric positive definite (SPD) manifold, the approaches of visual content classification with image sets have made impressive progress. However, the key challenge of unhelpfully large intraclass variability and interclass similarity of representations remains open to date. Although, several recent studies have mitigated the two problems by jointly learning the embedding mapping and the similarity metric on the original SPD manifold, their inherent shallow and linear feature transformation mechanism are not powerful enough to capture useful geometric features, especially in complex scenarios. To this end, this article explores a novel approach, termed SPD manifold deep metric learning (SMDML), for image set classification. Specifically, SMDML first selects a prevailing SPD manifold neural network (SPDNet) as the backbone (encoder) to derive an SPD matrix nonlinear representation. To counteract the degradation of structural information during multistage feature embedding, we construct a Riemannian decoder at the end of the encoder, trained by a reconstruction error term (RT), to induce the generated low-dimensional feature manifold of the hidden layer to capture the pivotal information about the visual data describing the imaged scene. We demonstrate through theory and experiments that it is feasible to replace the Riemannian metric with Euclidean distance in RT. Then, the ReCov layer is introduced into the established Riemannian network to regularize the local statistical information within each input feature matrix, which enhances the effectiveness of the learning process. The theoretical analysis of the activation function used in the ReCov layer in terms of continuity and conditions for generating positive definite matrices is beneficial for network design. Inspired by the fact that the single cross-entropy loss used for training is unable to effectively parse the geometric distribution of the deep representations, we finally endow the suggested model with a novel metric learning regularization term. By explicitly incorporating the encoding and processing of the data variations into the network learning process, this term can not only derive a powerful Riemannian representation but also train an effective classifier. The experimental results show the superiority of the proposed approach on three typical visual classification tasks. Rui Wang 0050, Xiaojun Wu 0001, Ziheng Chen 0001, Josef Kittler |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Multi-Level Fusion for Robust RGBT Tracking via Enhanced Thermal RepresentationabstractDue to the limitations of visible (RGB) sensors in challenging scenarios, such as nighttime and foggy environments, the thermal infrared (TIR) modality draws increasing attention as an auxiliary source for robust tracking systems. Currently, the existing methods extract both the RGB and TIR (RGBT) clues in a similar approach, i.e., utilising RGB-pretrained models with or without finetuning, and then aggregate the multi-modal information through a fusion block embedded in a single level. However, the different imaging principles of RGB and TIR data raise questions about the suitability of RGB-pretrained models for thermal data. In this article, it is argued that the modality gap is overlooked, and an alternative training paradigm is proposed for TIR data to ensure consistency between the training and test data, which is achieved by optimising the TIR feature extractor with only TIR data involved. Furthermore, with the goal of making better use of the enhanced thermal representations, a multi-level fusion strategy is inspired by the observation that various fusion strategies at different levels can contribute to a better performance. Specifically, fusion modules at both the feature and decision levels are derived for a comprehensive fusion procedure while the pixel-level fusion strategy is not considered due to the misalignment of multi-modal image pairs. The effectiveness of our method is demonstrated by extensive qualitative and quantitative experiments conducted on several challenging benchmarks. Code will be released at https://github.com/Zhangyong-Tang/MELT . Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Riemannian Local Mechanism for SPD Neural NetworksabstractThe Symmetric Positive Definite (SPD) matrices have received wide attention for data representation in many scientific areas. Although there are many different attempts to develop effective deep architectures for data processing on the Riemannian manifold of SPD matrices, very few solutions explicitly mine the local geometrical information in deep SPD feature representations. Given the great success of local mechanisms in Euclidean methods, we argue that it is of utmost importance to ensure the preservation of local geometric information in the SPD networks. We first analyse the convolution operator commonly used for capturing local information in Euclidean deep networks from the perspective of a higher level of abstraction afforded by category theory. Based on this analysis, we define the local information in the SPD manifold and design a multi-scale submanifold block for mining local geometry. Experiments involving multiple visual tasks validate the effectiveness of our approach. Ziheng Chen 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Zhiwu Huang, Josef Kittler |
AAAI | 6 |
| 2023 | RGBD1K: A Large-Scale Dataset and Benchmark for RGB-D Object TrackingabstractRGB-D object tracking has attracted considerable attention recently, achieving promising performance thanks to the symbiosis between visual and depth channels. However, given a limited amount of annotated RGB-D tracking data, most state-of-the-art RGB-D trackers are simple extensions of high-performance RGB-only trackers, without fully exploiting the underlying potential of the depth channel in the offline training stage. To address the dataset deficiency issue, a new RGB-D dataset named RGBD1K is released in this paper. The RGBD1K contains 1,050 sequences with about 2.5M frames in total. To demonstrate the benefits of training on a larger RGB-D data set in general, and RGBD1K in particular, we develop a transformer-based RGB-D tracker, named SPT, as a baseline for future visual object tracking studies using the new dataset. The results, of extensive experiments using the SPT tracker demonstrate the potential of the RGBD1K dataset to improve the performance of RGB-D tracking, inspiring future developments of effective tracker designs. The dataset and codes will be available on the project homepage: https://github.com/xuefeng-zhu5/RGBD1K. Xuefeng Zhu 0003, Tianyang Xu 0001, Zhangyong Tang, Zucheng Wu, Xiaojun Wu 0001, Josef Kittler |
AAAI | 8 |
| 2023 | Ada2NPT: An Adaptive Nearest Proxies Triplet Loss for Attribute-Aware Face Recognition with Adaptively Compacted Feature Learning
Lei Ju 0005, Zhenhua Feng 0001, Muhammad Awais 0001, Josef Kittler |
ACML | 4 |
| 2023 | Group Masked Model Learning for General Audio RepresentationabstractVision transformers have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships. However, transformers are known to be data hungry which require orders of magnitude more data [1] to train. This has motivated the research in self-supervised pretraining of audio transformers, which reduces the dependency on large amounts of labeled data and focuses on extracting concise representation of the audio spectrograms. In this paper, we propose Audio-GMML, a self-supervised transformer for general audio representations that is based on Group Masked Model Learning (GMML) and a patch aggregation strategy to improve the performance of learned representations and enforce global structure of the given audio. We evaluate our pretrained models on several downstream tasks, setting a new state-of-the-art performance on five audio and speech classification tasks. The code and pretrained weights will be made publicly available for the scientific community. Sara Atito Ali Ahmed, Muhammad Awais 0001, Tony Alex, Josef Kittler |
ICIP | 4 |
| 2023 | GMML is All You NeedabstractVision transformers (ViTs) have generated significant interest in the computer vision community because of their flexibility in exploiting contextual information, whether it is sharply confined local, or long range global. However, they are known to be data hungry and therefore often pretrained on large-scale datasets, e.g. JFT-300M or ImageNet. An ideal learning method would perform best regardless of the size of the dataset, a property lacked by current learning methods, with merely a few existing works studying ViTs with limited data. We propose Group Masked Model Learning (GMML), a self-supervised learning (SSL) method that is able to train ViTs and achieve state-of-the-art (SOTA) performance when pre-trained with limited data. The GMML uses the information conveyed by all concepts in the image. This is achieved by manipulating randomly groups of connected tokens, successively covering different meaningful parts of the image content, and then recovering the hidden information from the visible part of the concept. Unlike most of the existing SSL approaches, GMML does not require momentum encoder, nor relies on careful implementation details such as large batches and gradient stopping. Pretraining, finetuning, and evaluation codes are available under: https://github.com/GMML. Sara Atito Ali Ahmed, Muhammad Awais 0001, Srinivasa Rao Nandam, Josef Kittler |
ICIP | 4 |
| 2023 | ETM-face: effective training sample selection and multi-scale feature learning for face detection
Junyuan He, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
Multim. Tools Appl. | 6 |
| 2023 | U-SPDNet: An SPD manifold learning-based neural network for visual classification
Rui Wang 0050, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
Neural Networks | 5 |
| 2023 | Global Context-Aware Feature Extraction and Visible Feature Enhancement for Occlusion-Invariant Pedestrian Detection in Crowded Scenes
Zhen Liu 0015, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
Neural Process. Lett. | 6 |
| 2023 | Scalable Affine Multi-view Subspace Clustering
Wanrong Yu, Xiaojun Wu 0001, Tianyang Xu 0001, Ziheng Chen 0001, Josef Kittler |
Neural Process. Lett. | 5 |
| 2023 | Deep Order-Preserving Learning With Adaptive Optimal Transport DistanceabstractWe consider a framework for taking into consideration the relative importance (ordinality) of object labels in the process of learning a label predictor function. The commonly used loss functions are not well matched to this problem, as they exhibit deficiencies in capturing natural correlations of the labels and the corresponding data. We propose to incorporate such correlations into our learning algorithm using an optimal transport formulation. Our approach is to learn the ground metric, which is partly involved in forming the optimal transport distance, by leveraging ordinality as a general form of side information in its formulation. Based on this idea, we then develop a novel loss function for training deep neural networks. A highly efficient alternating learning method is then devised to alternatively optimise the ground metric and the deep model in an end-to-end learning manner. This scheme allows us to adaptively adjust the shape of the ground metric, and consequently the shape of the loss function for each application. We back up our approach by theoretical analysis and verify the performance of our proposed scheme by applying it to two learning tasks, i.e. chronological age estimation from the face and image aesthetic assessment. The numerical results on several benchmark datasets demonstrate the superiority of the proposed algorithm. Ali Akbari 0003, Muhammad Awais 0001, Soroush Fatemifar, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | NPT-Loss: Demystifying Face Recognition Losses With Nearest Proxies TripletabstractFace recognition (FR) using deep convolutional neural networks (DCNNs) has seen remarkable success in recent years. One key ingredient of DCNN-based FR is the design of a loss function that ensures discrimination between various identities. The state-of-the-art (SOTA) solutions utilise normalised Softmax loss with additive and/or multiplicative margins. Despite being popular and effective, these losses are justified only intuitively with little theoretical explanations. In this work, we show that under the LogSumExp (LSE) approximation, the SOTA Softmax losses become equivalent to a proxy-triplet loss that focuses on nearest-neighbour negative proxies only. This motivates us to propose a variant of the proxy-triplet loss, entitled Nearest Proxies Triplet (NPT) loss, which unlike SOTA solutions, converges for a wider range of hyper-parameters and offers flexibility in proxy selection and thus outperforms SOTA techniques. We generalise many SOTA losses into a single framework and give theoretical justifications for the assertion that minimising the proposed loss ensures a minimum separability between all identities. We also show that the proposed loss has an implicit mechanism of hard-sample mining. We conduct extensive experiments using various DCNN architectures on a number of FR benchmarks to demonstrate the efficacy of the proposed scheme over SOTA methods. Syed Safwan Khalid, Muhammad Awais 0001, Zhenhua Feng 0001, Chi-Ho Chan, Ammarah Farooq, Ali Akbari 0003, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | LRRNet: A Novel Representation Learning Guided Fusion Network for Infrared and Visible ImagesabstractDeep learning based fusion methods have been achieving promising performance in image fusion tasks. This is attributed to the network architecture that plays a very important role in the fusion process. However, in general, it is hard to specify a good fusion architecture, and consequently, the design of fusion networks is still a black art, rather than science. To address this problem, we formulate the fusion task mathematically, and establish a connection between its optimal solution and the network architecture that can implement it. This approach leads to a novel method proposed in the paper of constructing a lightweight fusion network. It avoids the time-consuming empirical network design by a trial-and-test strategy. In particular we adopt a learnable representation approach to the fusion task, in which the construction of the fusion network architecture is guided by the optimisation algorithm producing the learnable model. The low-rank representation (LRR) objective is the foundation of our learnable model. The matrix multiplications, which are at the heart of the solution are transformed into convolutional operations, and the iterative process of optimisation is replaced by a special feed-forward network. Based on this novel network architecture, an end-to-end lightweight fusion network is constructed to fuse infrared and visible light images. Its successful training is facilitated by a detail-to-semantic information loss function proposed to preserve the image details and to enhance the salient features of the source images. Our experiments show that the proposed fusion network exhibits better fusion performance than the state-of-the-art fusion methods on public datasets. Interestingly, our network requires a fewer training parameters than other existing methods. Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Jiwen Lu, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Neural Belief Propagation for Scene Graph GenerationabstractScene graph generation aims to interpret an input image by explicitly modelling the objects contained therein and their relationships. In existing methods the problem is predominantly solved by message passing neural network models. Unfortunately, in such models, the variational distributions generally ignore the structural dependencies among the output variables, and most of the scoring functions only consider pairwise dependencies. This can lead to inconsistent interpretations. In this article, we propose a novel neural belief propagation method seeking to replace the traditional mean field approximation with a structural Bethe approximation. To find a better bias-variance trade-off, higher-order dependencies among three or more output variables are also incorporated into the relevant scoring function. The proposed method achieves the state-of-the-art performance on various popular scene graph generation benchmarks. Daqi Liu, Miroslaw Bober, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Constrained Structure Learning for Scene Graph GenerationabstractAs a structured prediction task, scene graph generation aims to build a visually-grounded scene graph to explicitly model objects and their relationships in an input image. Currently, the mean field variational Bayesian framework is the de facto methodology used by the existing methods, in which the unconstrained inference step is often implemented by a message passing neural network. However, such formulation fails to explore other inference strategies, and largely ignores the more general constrained optimization models. In this paper, we present a constrained structure learning method, for which an explicit constrained variational inference objective is proposed. Instead of applying the ubiquitous message-passing strategy, a generic constrained optimization method - entropic mirror descent - is utilized to solve the constrained variational inference step. We validate the proposed generic model on various popular scene graph generation benchmarks and show that it outperforms the state-of-the-art methods. Daqi Liu, Miroslaw Bober, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Keep an eye on faces: Robust face detection with heatmap-Assisted spatial attention and scale-Aware layer attentionabstractModern anchor-based face detectors learn discriminative features using large-capacity networks and extensive anchor settings. In spite of their promising results, they are not without problems. First, most anchors extract redundant features from the background. As a consequence, the performance improvements are achieved at the expense of a disproportionate computational complexity. Second, the predicted face boxes are only distinguished by a classifier supervised by pre-defined positive, negative and ignored anchors. This strategy may ignore potential contributions from cohorts of anchors labeled negative/ignored during inference simply because of their inferior initialisation, although they can regress well to a target. In other words, true positives and representative features may get filtered out by unreliable confidence scores. To deal with the first concern and achieve more efficient face detection, we propose a Heatmap-assisted Spatial Attention (HSA) module and a Scale-aware Layer Attention (SLA) module to extract informative features using lower computational costs. To be specific, SLA incorporates the information from all the feature pyramid layers, weighted adaptively to remove redundant layers. HSA predicts a reshaped Gaussian heatmap and employs it to facilitate a spatial feature selection by better highlighting facial areas. For more reliable decision-making, we merge the predicted heatmap scores and classification results by voting. Since our heatmap scores are based on the distance to the face centres, they are able to retain all the well-regressed anchors. The experiments obtained on several well-known benchmarks demonstrate the merits of the proposed method. Lei Ju 0005, Josef Kittler, Muhammad Awais Rana, Wankou Yang, Zhenhua Feng 0001 |
Pattern Recognit. | 2 |
| 2023 | Hybrid Riemannian Graph-Embedding Metric Learning for Image Set ClassificationabstractWith the continuously increasing amount of video data, image set classification has recently received widespread attention in the CV&PR community. However, the intra-class diversity and inter-class ambiguity of representations remain an open challenge. To tackle this issue, several methods have been put forward to perform multiple geometry-aware image set modelling and learning. Although the extracted complementary geometric information is beneficial for decision making, the sophisticated computational paradigm (e.g., scatter matrices computation and iterative optimisation) of such algorithms is counterproductive. As a countermeasure, we propose an effective hybrid Riemannian metric learning framework in this paper. Specifically, we design a multiple graph embedding-guided metric learning framework for the sake of fusing these complementary kernel features, obtained via the explicit RKHS embeddings of the Grassmannian manifold, SPD manifold, and Gaussian embedded Riemannian manifold, into a unified subspace for classification. Furthermore, the involved optimisation problem of the developed model can be solved in terms of a series of sub-problems, achieving improved efficiency theoretically and experimentally. Substantial experiments are carried out to evaluate the efficacy of our approach. The experimental results suggest the superiority of it over the state-of-the-art methods. Ziheng Chen 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Josef Kittler |
IEEE Trans. Big Data | 5 |
| 2023 | CPInformer for Efficient and Robust Compound-Protein Interaction PredictionabstractRecently, deep learning has become the mainstream methodology for Compound-Protein Interaction (CPI) prediction. However, the existing compound-protein feature extraction methods have some issues that limit their performance. First, graph networks are widely used for structural compound feature extraction, but the chemical properties of a compound depend on functional groups rather than graphic structure. Besides, the existing methods lack capabilities in extracting rich and discriminative protein features. Last, the compound-protein features are usually simply combined for CPI prediction, without considering information redundancy and effective feature mining. To address the above issues, we propose a novel CPInformer method. Specifically, we extract heterogeneous compound features, including structural graph features and functional class fingerprints, to reduce prediction errors caused by similar structural compounds. Then, we combine local and global features using dense connections to obtain multi-scale protein features. Last, we apply ProbSparse self-attention to protein features, under the guidance of compound features, to eliminate information redundancy, and to improve the accuracy of CPInformer. More importantly, the proposed method identifies the activated local regions that link a CPI, providing a good visualisation for the CPI state. The results obtained on five benchmarks demonstrate the merits and superiority of CPInformer over the state-of-the-art approaches. Yang Hua 0002, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler, Dongjun Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Fast Self-Guided Multi-View Subspace ClusteringabstractMulti-view subspace clustering is an important topic in cluster analysis. Its aim is to utilize the complementary information conveyed by multiple views of objects to be clustered. Recently, view-shared anchor learning based multi-view clustering methods have been developed to speed up the learning of common data representation. Although widely applied to large-scale scenarios, most of the existing approaches are still faced with two limitations. First, they do not pay sufficient consideration on the negative impact caused by certain noisy views with unclear clustering structures. Second, many of them only focus on the multi-view consistency, yet are incapable of capturing the cross-view diversity. As a result, the learned complementary features may be inaccurate and adversely affect clustering performance. To solve these two challenging issues, we propose a Fast Self-guided Multi-view Subspace Clustering (FSMSC) algorithm which skillfully integrates the view-shared anchor learning and global-guided-local self-guidance learning into a unified model. Such an integration is inspired by the observation that the view with clean clustering structures will play a more crucial role in grouping the clusters when the features of all views are concatenated. Specifically, we first learn a locally-consistent data representation shared by all views in the local learning module, then we learn a globally-discriminative data representation from multi-view concatenated features in the global learning module. Afterwards, a feature selection matrix constrained by the ℓ2,1-norm is designed to construct a guidance from global learning to local learning. In this way, the multi-view consistent and diverse information can be simultaneously utilized and the negative impact caused by noisy views can be overcame to some extent. Extensive experiments on different datasets demonstrate the effectiveness of our proposed fast self-guided learning model, and its promising performance compared to both, the state-of-the-art non-deep and deep multi-view clustering algorithms. The code of this paper is available at https://github.com/chenzhe207/FSMSC. Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Trans. Image Process. | 4 |
| 2023 | Toward Robust Visual Object Tracking With Independent Target-Agnostic Detection and Effective Siamese Cross-Task InteractionabstractAdvanced Siamese visual object tracking architectures are jointly trained using pair-wise input images to perform target classification and bounding box regression. They have achieved promising results in recent benchmarks and competitions. However, the existing methods suffer from two limitations: First, though the Siamese structure can estimate the target state in an instance frame, provided the target appearance does not deviate too much from the template, the detection of the target in an image cannot be guaranteed in the presence of severe appearance variations. Second, despite the classification and regression tasks sharing the same output from the backbone network, their specific modules and loss functions are invariably designed independently, without promoting any interaction. Yet, in a general tracking task, the centre classification and bounding box regression tasks are collaboratively working to estimate the final target location. To address the above issues, it is essential to perform target-agnostic detection so as to promote cross-task interactions in a Siamese-based tracking framework. In this work, we endow a novel network with a target-agnostic object detection module to complement the direct target inference, and to avoid or minimise the misalignment of the key cues of potential template-instance matches. To unify the multi-task learning formulation, we develop a cross-task interaction module to ensure consistent supervision of the classification and regression branches, improving the synergy of different branches. To eliminate potential inconsistencies that may arise within a multi-task architecture, we assign adaptive labels, rather than fixed hard labels, to supervise the network training more effectively. The experimental results obtained on several benchmarks, i.e., OTB100, UAV123, VOT2018, VOT2019, and LaSOT, demonstrate the effectiveness of the advanced target detection module, as well as the cross-task interaction, exhibiting superior tracking performance as compared with the state-of-the-art tracking methods. Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Image Process. | 4 |
| 2023 | WATCH: Two-Stage Discrete Cross-Media HashingabstractDue to the explosive growth of multimedia data in recent years, cross-media hashing (CMH) approaches have recently received increasing attention. To learn the hash codes, most existing supervised CMH algorithms employ the strict binary label information, which has small margins between the incorrect labels (0) and the true labels (1), increasing the risk of classification error. Besides, most existing CMH approaches are one-stage algorithms, in which the hash functions and binary codes can be learned simultaneously, complicating the optimization. To avoid NP-hard optimization, many approaches utilize a relaxation strategy. However, this optimisation trick may cause large quantization errors. To address this, we present a novel tWo-stAge discreTe Cross-media Hashing method based on smooth matrix factorization and label relaxation, named WATCH. The proposed WATCH controls the margins adaptively by the novel label relaxation strategy. This innovation reduces the quantization error significantly. Besides, WATCH is a two-stage model. In stage 1, we employ a discrete smooth matrix factorization model. Then, the hash codes can be generated discretely, reducing the large quantization loss greatly. In stage 2, we adopt an effective hash function learning strategy, which produces more effective hash functions. Comprehensive experiments on several datasets demonstrate that WATCH outperforms some state-of-the-art methods. Donglin Zhang 0001, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | A Theoretical Insight Into the Effect of Loss Function for Deep Semantic-Preserving LearningabstractGood generalization performance is the fundamental goal of any machine learning algorithm. Using the uniform stability concept, this article theoretically proves that the choice of loss function impacts the generalization performance of a trained deep neural network (DNN). The adopted stability-based framework provides an effective tool for comparing the generalization error bound with respect to the utilized loss function. The main result of our analysis is that using an effective loss function makes stochastic gradient descent more stable which consequently leads to the tighter generalization error bound, and so better generalization performance. To validate our analysis, we study learning problems in which the classes are semantically correlated. To capture this semantic similarity of neighboring classes, we adopt the well-known semantics-preserving learning framework, namely label distribution learning (LDL). We propose two novel loss functions for the LDL framework and theoretically show that they provide stronger stability than the other widely used loss functions adopted for training DNNs. The experimental results on three applications with semantically correlated classes, including facial age estimation, head pose estimation, and image esthetic assessment, validate the theoretical insights gained by our analysis and demonstrate the usefulness of the proposed loss functions in practical applications. Ali Akbari 0003, Muhammad Awais 0001, Manijeh Bashar, Josef Kittler |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Discriminative Dictionary Pair Learning With Scale-Constrained Structured Representation for Image ClassificationabstractThe dictionary pair learning (DPL) model aims to design a synthesis dictionary and an analysis dictionary to accomplish the goal of rapid sample encoding. In this article, we propose a novel structured representation learning algorithm based on the DPL for image classification. It is referred to as discriminative DPL with scale-constrained structured representation (DPL-SCSR). The proposed DPL-SCSR utilizes the binary label matrix of dictionary atoms to project the representation into the corresponding label space of the training samples. By imposing a non-negative constraint, the learned representation adaptively approximates a block-diagonal structure. This innovative transformation is also capable of controlling the scale of the block-diagonal representation by enforcing the sum of within-class coefficients of each sample to 1, which means that the dictionary atoms of each class compete to represent the samples from the same class. This implies that the requirement of similarity preservation is considered from the perspective of the constraint on the sum of coefficients. More importantly, the DPL-SCSR does not need to design a classifier in the representation space as the label matrix of the dictionary can also be used as an efficient linear classifier. Finally, the DPL-SCSR imposes the$l_{2,p}$-norm on the analysis dictionary to make the process of feature extraction more interpretable. The DPL-SCSR seamlessly incorporates the scale-constrained structured representation learning, within-class similarity preservation of representation, and the linear classifier into one regularization term, which dramatically reduces the complexity of training and parameter tuning. The experimental results on several popular image classification datasets show that our DPL-SCSR can deliver superior performance compared with the state-of-the-art (SOTA) dictionary learning methods. The MATLAB code of this article is available athttps://github.com/chenzhe207/DPL-SCSR. Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | AXM-Net: Implicit Cross-Modal Feature Alignment for Person Re-identificationabstractCross-modal person re-identification (Re-ID) is critical for modern video surveillance systems. The key challenge is to align cross-modality representations conforming to semantic information present for a person and ignore background information. This work presents a novel convolutional neural network (CNN) based architecture designed to learn semantically aligned cross-modal visual and textual representations. The underlying building block, named AXM-Block, is a unified multi-layer network that dynamically exploits the multi-scale knowledge from both modalities and re-calibrates each modality according to shared semantics. To complement the convolutional design, contextual attention is applied in the text branch to manipulate long-term dependencies. Moreover, we propose a unique design to enhance visual part-based feature coherence and locality information. Our framework is novel in its ability to implicitly learn aligned semantics between modalities during the feature learning stage. The unified feature learning effectively utilizes textual data as a super-annotation signal for visual representation learning and automatically rejects irrelevant information. The entire AXM-Net is trained end-to-end on CUHK-PEDES data. We report results on two tasks, person search and cross-modal Re-ID. The AXM-Net outperforms the current state-of-the-art (SOTA) methods and achieves 64.44% Rank@1 on the CUHK-PEDES test set. It also outperforms by >10% for cross-viewpoint text-to-image Re-ID scenarios on CrossRe-ID and CUHK-SYSU datasets. Ammarah Farooq, Muhammad Awais 0001, Josef Kittler, Syed Safwan Khalid |
AAAI | 3 |
| 2022 | DreamNet: A Deep Riemannian Manifold Network for SPD Matrix Learning
Rui Wang 0050, Xiaojun Wu 0001, Ziheng Chen 0001, Tianyang Xu 0001, Josef Kittler |
ACCV (6) | 5 |
| 2022 | Multi-target regression via non-linear output structure learning
Shervin Rahimzadeh Arashloo, Josef Kittler |
Neurocomputing | 2 |
| 2022 | Subspace clustering via joint ℓ1, 2 and ℓ2, 1 norms
Wenhua Dong, Xiaojun Wu 0001, Josef Kittler |
Inf. Sci. | 3 |
| 2022 | Learning a discriminative SPD manifold neural network for image set classification
Rui Wang 0050, Xiaojun Wu 0001, Ziheng Chen 0001, Tianyang Xu 0001, Josef Kittler |
Neural Networks | 5 |
| 2022 | Distribution Cognisant Loss for Cross-Database Facial Age Estimation With Sensitivity AnalysisabstractExisting facial age estimation studies have mostly focused on intra-database protocols that assume training and test images are captured under similar conditions. This is rarely valid in practical applications, where we typically encounter training and test sets with different characteristics. In this article, we deal with such situations, namely subjective-exclusive cross-database age estimation. We formulate the age estimation problem as the distribution learning framework, where the age labels are encoded as a probability distribution. To improve the cross-database age estimation performance, we propose a new loss function which provides a more robust measure of the difference between ground-truth and predicted distributions. The desirable properties of the proposed loss function are theoretically analysed and compared with the state-of-the-art approaches. In addition, we compile a new balanced large-scale age estimation database. Last, we introduce a novel evaluation protocol, called subject-exclusive cross-database age estimation protocol, which provides meaningful information of a method in terms of the generalisation capability. The experimental results demonstrate that the proposed approach outperforms the state-of-the-art age estimation methods under both intra-database and subject-exclusive cross-database evaluation protocols. In addition, in this article, we provide a comparative sensitivity analysis of various algorithms to identify trends and issues inherent to their performance. This analysis introduces some open problems to the community which might be considered when designing a robust age estimation system. Ali Akbari 0003, Muhammad Awais 0001, Zhenhua Feng 0001, Ammarah Farooq, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Developing a generic framework for anomaly detection
Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler |
Pattern Recognit. | 4 |
| 2022 | Differentiable neural architecture learning for efficient neural networks
Qingbei Guo, Xiaojun Wu 0001, Josef Kittler, Zhiquan Feng |
Pattern Recognit. | 3 |
| 2022 | Face spoofing detection ensemble via multistage optimisation and pruning
Soroush Fatemifar, Shahrokh Asadi, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler |
Pattern Recognit. Lett. | 5 |
| 2022 | Target-Cognisant Siamese Network for Robust Visual Object Tracking
Yingjie Jiang, Xiaoning Song, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. Lett. | 6 |
| 2022 | Multiple Riemannian Manifold-Valued Descriptors Based Image Set Classification With Multi-Kernel Metric LearningabstractThe importance of wild video based image set recognition is monotonically increasing due to the large amount of video data being collected by various devices including surveillance cameras, drive recorders, smart phones, and internet. The content of these videos is often complex, and it raises the question of how to perform image set modeling and feature extraction for image set-based classification. In recent years, image set classification methods have advanced considerably by modeling the image set in terms of a covariance matrix, linear subspace, or Gaussian distribution. Moreover, the distinctive geometry spanned by them include Symmetric Positive Definite (SPD) manifold, Grassmannian manifold, and Gaussian embedded Riemannian manifold, respectively. As a matter of fact, most of the approaches just adopt a single geometric model to describe each given image set, which may lose information useful for classification. To tackle this problem, we propose a novel algorithm to model each image set from a multi-geometric perspective. Specifically, the covariance matrix, linear subspace, and Gaussian distribution are applied to set representation simultaneously. In order to fuse these multiple heterogeneous Riemannian manifold-valued features, the well-equipped Riemannian kernel functions are first employed to map them into high dimensional Hilbert spaces. Then, a multi-kernel metric learning framework is devised to embed the learned hybrid kernels into a common lower dimensional subspace to facilitate classification. We conduct experiments on six widely used datasets each representing a different classification task: video-based face recognition, set-based object categorization, video-based emotion recognition, dynamic scene classification, set-based cell identification, and 3D hand pose estimation, to evaluate the classification performance of the proposed algorithm. The extensive experimental results confirm its superiority over the state-of-the-art methods. Rui Wang 0050, Xiaojun Wu 0001, Kai-Xuan Chen 0001, Josef Kittler |
IEEE Trans. Big Data | 4 |
| 2022 | Graph2Net: Perceptually-Enriched Graph Learning for Skeleton-Based Action RecognitionabstractSkeleton representation has attracted a great deal of attention recently as an extremely robust feature for human action recognition. However, its non-Euclidean structural characteristics raise new challenges for conventional solutions. Recent studies have shown that there is a native superiority in modeling spatiotemporal skeleton information with a Graph Convolutional Network (GCN). Nevertheless, the skeleton graph modeling normally focuses on the physical adjacency of the elements of the human skeleton sequence, which contrasts with the requirement to provide a perceptually meaningful representation. To address this problem, in this paper, we propose a perceptually-enriched graph learning method by introducing innovative features to spatial and temporal skeleton graph modeling. For the spatial information modeling, we incorporate a Local-Global Graph Convolutional Network (LG-GCN) that builds a multifaceted spatial perceptual representation. This helps to overcome the limitations caused by over-reliance on the spatial adjacency relationships in the skeleton. For temporal modeling, we present a Region-Aware Graph Convolutional Network (RA-GCN), which directly embeds the regional relationships conveyed by a skeleton sequence into a temporal graph model. This innovation mitigates the deficiency of the original skeleton graph models. In addition, we strengthened the ability of the proposed channel modeling methods to extract multi-scale representations. These innovations result in a lightweight graph convolutional model, referred to as Graph2Net, that simultaneously extends the spatial and temporal perceptual fields, and thus enhances the capacity of the graph model to represent skeleton sequences. We conduct extensive experiments on NTU-RGB+D 60&120, Northwestern-UCLA, and Kinetics-400 datasets to show that our results surpass the performance of several mainstream methods while limiting the model complexity and computational overhead. Cong Wu 0006, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | A Novel Ground Metric for Optimal Transport-Based Chronological Age EstimationabstractLabel distribution learning (LDL) is the state-of-the-art approach to dealing with a number of real-world applications, such as chronological age estimation from a face image, where there is an inherent similarity among adjacent age labels. LDL takes into account the semantic similarity by assigning a label distribution to each instance. The well-known Kullback-Leibler (KL) divergence is the widely used loss function for the LDL framework. However, the KL divergence does not fully and effectively capture the semantic similarity among age labels, thus leading to suboptimal performance. In this article, we propose a novel loss function based on the optimal transport theory for the LDL-based age estimation. A ground metric function plays an important role in the optimal transport formulation. It should be carefully determined based on the underlying geometric structure of the label space of the application in-hand. The label space in the age estimation problem has a specific geometric structure, that is, closer ages have more inherent semantic relationships. Inspired by this, we devise a novel ground metric function, which enables the loss function to increase the influence of highly correlated ages; thus exploiting the semantic similarity among ages more effectively than the existing loss functions. We then use the proposed loss function, namely, γ -Wasserstein loss, for training a deep neural network (DNN). This leads to a notoriously computationally expensive and nonconvex optimization problem. Following the standard methodology, we formulate the optimization function as a convex problem and then use an efficient iterative algorithm to update the parameters of the DNN. Extensive experiments in age estimation on different benchmark datasets validate the effectiveness of the proposed method, which consistently outperforms state-of-the-art approaches. Ali Akbari 0003, Muhammad Awais 0001, Soroush Fatemifar, Syed Safwan Khalid, Josef Kittler |
IEEE Trans. Cybern. | 5 |
| 2022 | Robust Visual Object Tracking Via Adaptive Attribute-Aware Discriminative Correlation FiltersabstractIn recent years, attention mechanisms have been widely studied in Discriminative Correlation Filter (DCF) based visual object tracking. To realise spatial attention and discriminative feature mining, existing approaches usually apply regularisation terms to the spatial dimension of multi-channel features. However, these spatial regularisation approaches construct a shared spatial attention pattern for all multi-channel features, without considering the diversity across channels. As each feature map (channel) focuses on a specific visual attribute, a shared spatial attention pattern limits the capability for mining important information from different channels. To address this issue, we advocate channel-specific spatial attention for DCF-based trackers. The key ingredient of the proposed method is an Adaptive Attribute-Aware spatial attention mechanism for constructing a novel DCF-based tracker (A$^3$DCF). To highlight the discriminative elements in each feature map, spatial sparsity is imposed in the filter learning stage, moderated by the prior knowledge regarding the expected concentration of signal energy. In addition, we perform a post processing of the identified spatial patterns to alleviate the impact of less significant channels. The net effect is that the irrelevant and inconsistent channels are removed by the proposed method. The results obtained on a number of well-known benchmarking datasets, including OTB2015, DTB70, UAV123, VOT2018, LaSOT, GOT-10 K and TrackingNet, demonstrate the merits of the proposed A$^3$DCF tracker, with improved performance compared to the state-of-the-art methods. Xuefeng Zhu 0003, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler |
IEEE Trans. Multim. | 5 |
| 2022 | Relaxed Block-Diagonal Dictionary Pair Learning With Locality Constraint for Image RecognitionabstractWe propose a novel structured analysis–synthesis dictionary pair learning method for efficient representation and image classification, referred to as relaxed block-diagonal dictionary pair learning with a locality constraint (RBD-DPL). RBD-DPL aims to learn relaxed block-diagonal representations of the input data to enhance the discriminability of both analysis and synthesis dictionaries by dynamically optimizing the block-diagonal components of representation, while the off-block-diagonal counterparts are set to zero. In this way, the learned synthesis subdictionary is allowed to be more flexible in reconstructing the samples from the same class, and the analysis dictionary effectively transforms the original samples into a relaxed coefficient subspace, which is closely associated with the label information. Besides, we incorporate a locality-constraint term as a complement of the relaxation learning to enhance the locality of the analytical encoding so that the learned representation exhibits high intraclass similarity. A linear classifier is trained in the learned relaxed representation space for consistent classification. RBD-DPL is computationally efficient because it avoids both the use of class-specific complementary data matrices to learn discriminative analysis dictionary, as well as the time-consuming$l_{1}/l_{0}$-norm sparse reconstruction process. The experimental results demonstrate that our RBD-DPL achieves at least comparable or better recognition performance than the state-of-the-art algorithms. Moreover, both the training and testing time are significantly reduced, which verifies the efficiency of our method. The MATLAB code of the proposed RBD-DPL is available athttps://github.com/chenzhe207/RBD-DPL. Zhe Chen 0018, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | SymNet: A Simple Symmetric Positive Definite Manifold Deep Learning Method for Image Set ClassificationabstractBy representing each image set as a nonsingular covariance matrix on the symmetric positive definite (SPD) manifold, visual classification with image sets has attracted much attention. Despite the success made so far, the issue of large within-class variability of representations still remains a key challenge. Recently, several SPD matrix learning methods have been proposed to assuage this problem by directly constructing an embedding mapping from the original SPD manifold to a lower dimensional one. The advantage of this type of approach is that it cannot only implement discriminative feature selection but also preserve the Riemannian geometrical structure of the original data manifold. Inspired by this fact, we propose a simple SPD manifold deep learning network (SymNet) for image set classification in this article. Specifically, we first design SPD matrix mapping layers to map the input SPD matrices into new ones with lower dimensionality. Then, rectifying layers are devised to activate the input matrices for the purpose of forming a valid SPD manifold, chiefly to inject nonlinearity for SPD matrix learning with two nonlinear functions. Afterward, we introduce pooling layers to further compress the input SPD matrices, and the log-map layer is finally exploited to embed the resulting SPD matrices into the tangent space via log-Euclidean Riemannian computing, such that the Euclidean learning applies. For SymNet, the (2-D)2principal component analysis (PCA) technique is utilized to learn the multistage connection weights without requiring complicated computations, thus making it be built and trained easier. On the tail of SymNet, the kernel discriminant analysis (KDA) algorithm is coupled with the output vectorized feature representations to perform discriminative subspace learning. Extensive experiments and comparisons with state-of-the-art methods on six typical visual classification tasks demonstrate the feasibility and validity of the proposed SymNet. Rui Wang 0050, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Two-Stage Supervised Discrete Hashing for Cross-Modal RetrievalabstractRecently, hashing-based multimodal learning systems have received increasing attention due to their query efficiency and parsimonious storage costs. However, impeded by the quantization loss caused by numerical optimization, the existing cross-media hashing approaches are unable to capture all the discriminative information present in the original multimodal data. Besides, most cross-modal methods belong to the one-step paradigm, which learn the binary codes and hash function simultaneously, increasing the complexity of optimization. To address these issues, we propose a novel two-stage approach, named the two-stage supervised discrete hashing (TSDH) method. In particular, in the first phase, TSDH generates a latent representation for each modality. These representations are then mapped to a common Hamming space to generate the binary codes. In addition, TSDH directly endows the hash codes with the semantic labels, enhancing the discriminatory power of the learned binary codes. A discrete hash optimization approach is developed to learn the binary codes without relaxation, avoiding the large quantization loss. The proposed hash function learning scheme reuses the semantic information contained by the embeddings, endowing the hash functions with enhanced discriminability. Extensive experiments on several databases demonstrate the effectiveness of the developed TSDH, outperforming several recent competitive cross-media algorithms. Donglin Zhang 0001, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | Separable Batch Normalization for Robust Facial Landmark Localization
Shuangping Jin, Zhenhua Feng 0001, Wankou Yang, Josef Kittler |
BMVC | 4 |
| 2021 | Particle Swarm And Pattern Search Optimisation Of An Ensemble Of Face Anomaly DetectorsabstractWhile the remarkable advances in face matching render face biometric technology more widely applicable, its successful deployment may be compromised by face spoofing. Recent studies have shown that anomaly-based face spoofing detectors offer an interesting alternative to the multiclass counterparts by generalising better to unseen types of attack. In this work, we investigate the merits of fusing multiple anomaly spoofing detectors in the unseen attack scenario via a Weighted Averaging (WA) and client-specific design. We propose to optimise the parameters of WA by a two-stage optimisation method consisting of Particle Swarm Optimisation (PSO) and the Pattern Search (PS) algorithms to avoid the local minimum problem. Besides, we propose a novel scoring normalisation method which could be effectively applied in extreme cases such as heavy-tailed distributions. We evaluate the capability of the proposed system on publicly available face anti-spoofing databases including Replay-Attack, Replay-Mobile and Rose-Youtu. The experimental results demonstrate that the proposed fusion system outperforms the majority of anomaly-based and state-of-the-art multiclass approaches. Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler |
ICIP | 4 |
| 2021 | How Does Loss Function Affect Generalization Performance of Deep Learning? Application to Human Age EstimationabstractGood generalization performance across a wide variety of domains caused by many external and internal factors is the fundamental goal of any machine learning algorithm. This paper theoretically proves that the choice of loss function matters for improving the generalization performance of deep learning-based systems. By deriving the generalization error bound for deep neural models trained by stochastic gradient descent, we pinpoint the characteristics of the loss function that is linked to the generalization error and can therefore be used for guiding the loss function selection process. In summary, our main statement in this paper is: choose a stable loss function, generalize better. Focusing on human age estimation from the face which is a challenging topic in computer vision, we then propose a novel loss function for this learning problem. We theoretically prove that the proposed loss function achieves stronger stability, and consequently a tighter generalization error bound, compared to the other common loss functions for this problem. We have supported our findings theoretically, and demonstrated the merits of the guidance process experimentally, achieving significant improvements. Ali Akbari 0003, Muhammad Awais 0001, Manijeh Bashar, Josef Kittler |
ICML | 4 |
| 2021 | Insight on Attention Modules for Skeleton-Based Action Recognition
Quanyan Jiang, Xiaojun Wu 0001, Josef Kittler |
PRCV (1) | 3 |
| 2021 | Adaptive Channel Selection for Robust Visual Object Tracking with Discriminative Correlation FiltersabstractAbstract Discriminative Correlation Filters (DCF) have been shown to achieve impressive performance in visual object tracking. However, existing DCF-based trackers rely heavily on learning regularised appearance models from invariant image feature representations. To further improve the performance of DCF in accuracy and provide a parsimonious model from the attribute perspective, we propose to gauge the relevance of multi-channel features for the purpose of channel selection. This is achieved by assessing the information conveyed by the features of each channel as a group, using an adaptive group elastic net inducing independent sparsity and temporal smoothness on the DCF solution. The robustness and stability of the learned appearance model are significantly enhanced by the proposed method as the process of channel selection performs implicit spatial regularisation. We use the augmented Lagrangian method to optimise the discriminative filters efficiently. The experimental results obtained on a number of well-known benchmarking datasets demonstrate the effectiveness and stability of the proposed method. A superior performance over the state-of-the-art trackers is achieved using less than $$10\%$$ 10 % deep feature channels. Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 4 |
| 2021 | Dynamic information enhancement for video classification
Rongchang Li 0001, Xiaojun Wu 0001, Cong Wu 0006, Tianyang Xu 0001, Josef Kittler |
Image Vis. Comput. | 5 |
| 2021 | 2D progressive fusion module for action recognition
Xiaojun Wu 0001, Josef Kittler |
Image Vis. Comput. | 3 |
| 2021 | Weak sub-network pruning for strong and efficient neural networks
Qingbei Guo, Xiaojun Wu 0001, Josef Kittler, Zhiquan Feng |
Neural Networks | 3 |
| 2021 | Advanced skeleton-based action recognition via spatial-temporal rotation descriptors
Xiaojun Wu 0001, Josef Kittler |
Pattern Anal. Appl. | 3 |
| 2021 | Visual Semantic Information Pursuit: A SurveyabstractVisual semantic information comprises two important parts: the meaning of each visual semantic unit and the coherent visual semantic relation conveyed by these visual semantic units. Essentially, the former one is a visual perception task while the latter corresponds to visual context reasoning. Remarkable advances in visual perception have been achieved due to the success of deep learning. In contrast, visual semantic information pursuit, a visual scene semantic interpretation task combining visual perception and visual context reasoning, is still in its early stage. It is the core task of many different computer vision applications, such as object detection, visual semantic segmentation, visual relationship detection, or scene graph generation. Since it helps to enhance the accuracy and the consistency of the resulting interpretation, visual context reasoning is often incorporated with visual perception in current deep end-to-end visual semantic information pursuit methods. Surprisingly, a comprehensive review for this exciting area is still lacking. In this survey, we present a unified theoretical paradigm for all these methods, followed by an overview of the major developments and the future trends in each potential direction. The common benchmark datasets, the evaluation metrics and the comparisons of the corresponding methods are also introduced. Daqi Liu, Miroslaw Bober, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Client-specific anomaly detection for face presentation attack detection
Soroush Fatemifar, Shervin Rahimzadeh Arashloo, Muhammad Awais 0001, Josef Kittler |
Pattern Recognit. | 4 |
| 2021 | MOON: Multi-hash codes joint learning for cross-media retrieval
Donglin Zhang 0001, Xiaojun Wu 0001, He-Feng Yin, Josef Kittler |
Pattern Recognit. Lett. | 4 |
| 2021 | Sparse non-negative transition subspace learning for image classification
Zhe Chen 0018, Xiaojun Wu 0001, Yu-Hong Cai, Josef Kittler |
Signal Process. | 4 |
| 2021 | Learning Alternating Deep-Layer Cascaded RepresentationabstractWe propose an alternating deep-layer cascade (A-DLC) architecture for representation learning in the context of image classification. The merits of the proposed model are threefold. First, A-DLC is the first-ever method that alternatively cascades the sparse and collaborative representations using the class-discriminant softmax vector representation at the interface of each cascade section so that the sparsity and collaborativity can simultaneously be considered. Second, A-DLC inherits the hierarchy learning capability that effectively extends the traditional shallow sparse coding to a multi-layer learning model, thus enabling a full exploitation of the inherent latent discriminative information. Third, the simulation results show a significant amelioration in the classification accuracy, compared to earlier one-step single-layer classification algorithms. The Matlab code of this paper is available at https://github.com/chenzhe207/A-DLC. Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Signal Process. Lett. | 4 |
| 2021 | Complementary Discriminative Correlation Filters Based on Collaborative Representation for Visual Object TrackingabstractIn recent years, discriminative correlation filter (DCF) based algorithms have significantly advanced the state of the art in visual object tracking. The key to the success of DCF is an efficient discriminative regression model trained with powerful multi-cue features, including both hand-crafted and deep neural network features. However, the tracking performance is hindered by their inability to respond adequately to abrupt target appearance variations. This issue is posed by the limited representation capability of fixed image features. In this work, we set out to rectify this shortcoming by proposing a complementary representation of a visual content. Specifically, we propose the use of a collaborative representation between successive frames to extract the dynamic appearance information from a target with rapid appearance changes, which results in suppressing the undesirable impact of the background. The resulting collaborative representation coefficients are combined with the original feature maps using a spatially regularised DCF framework for performance boosting. The experimental results on several benchmarking datasets demonstrate the effectiveness and robustness of the proposed method, as compared with a number of state-of-the-art tracking algorithms. Xuefeng Zhu 0003, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Graph Embedding Multi-Kernel Metric Learning for Image Set Classification With Grassmannian Manifold-Valued FeaturesabstractIn the domain of video-based image set classification, a considerable advance has been made by modeling a sequence of video frames (image set) as a linear subspace, which typically resides on a Grassmannian manifold. As a consequence of the large intra-class variations of the video data, there are two open challenges for the modeling task: how to establish appropriate image set models to encode these variations, and how to effectively measure the similarity between any two image sets. As a possible way to tackle these issues, this paper presents a graph embedding multi-kernel metric learning (GEMKML) algorithm for image set classification. The proposed GEMKML implements set modeling, feature extraction, and classification in two steps. Firstly, the proposed framework constructs a novel cascaded feature learning architecture on Grassmannian manifold with the aim of producing more effective Grassmannian manifold-valued feature representations. To make a better use of these learned features, a graph embedding multi-kernel metric learning scheme is then devised to map them into a lower-dimensional Euclidean space, where the inter-class distances are maximized and the intra-class distances are minimized. We evaluate the proposed GEMKML on five different visual classification tasks using widely adopted datasets. The extensive classification results confirm its superiority over the state-of-the-art methods. Rui Wang 0050, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Multim. | 3 |
| 2021 | Robust One-Class Kernel Spectral RegressionabstractThe kernel null-space technique is known to be an effective one-class classification (OCC) technique. Nevertheless, the applicability of this method is limited due to its susceptibility to possible training data corruption and the inability to rank training observations according to their conformity with the model. This article addresses these shortcomings by regularizing the solution of the null-space kernel Fisher methodology in the context of its regression-based formulation. In this respect, first, the effect of the Tikhonov regularization in the Hilbert space is analyzed, where the one-class learning problem in the presence of contamination in the training set is posed as a sensitivity analysis problem. Next, the effect of the sparsity of the solution is studied. For both alternative regularization schemes, iterative algorithms are proposed which recursively update label confidences. Through extensive experiments, the proposed methodology is found to enhance robustness against contamination in the training set compared with the baseline kernel null-space method, as well as other existing approaches in the OCC paradigm, while providing the functionality to rank training samples effectively. Shervin Rahimzadeh Arashloo, Josef Kittler |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Correlation tracking with implicitly extending search region
Xiaojun Wu 0001, Josef Kittler, Tianyang Xu 0001 |
Vis. Comput. | 3 |
| 2020 | Sensitivity of Age Estimation Systems to Demographic Factors and Image Quality: Achievements and ChallengesabstractRecently, impressively growing efforts have been devoted to the challenging task of facial age estimation. The improvements in performance achieved by new algorithms are measured on several benchmarking test databases with different characteristics to check on consistency. While this is a valuable methodology in itself, a significant issue in the most age estimation related studies is that the reported results lack an assessment of intrinsic system uncertainty. Hence, a more in-depth view is required to examine the robustness of age estimation systems in different scenarios. The purpose of this paper is to conduct an evaluative and comparative analysis of different age estimation systems to identify trends, as well as the points of their critical vulnerability. In particular, we investigate four age estimation systems, including the online Microsoft service, two best state-of-the-art approaches advocated in the literature, as well as a novel age estimation algorithm. We analyse the effect of different internal and external factors, including gender, ethnicity, expression, makeup, illumination conditions, quality and resolution of the face images, on the performance of these age estimation systems. The goal of this sensitivity analysis is to provide the biometrics community with the insight and understanding of the critical subject-, camera- and environmental-based factors that affect the overall performance of the age estimation system under study. Ali Akbari 0003, Muhammad Awais 0001, Josef Kittler |
IJCB | 3 |
| 2020 | Cross Modal Person Re-identification with Visual-Textual QueriesabstractClassical person re-identification approaches assume that a person of interest has appeared across different cameras and can be queried by one of the existing images. However, in real-world surveillance scenarios, frequently no visual information will be available about the queried person. In such scenarios, a natural language description of the person by a witness will provide the only source of information for retrieval. In this work, person re-identification using both vision and language information is addressed under all possible gallery and query scenarios. A two stream deep convolutional neural network framework supervised by identity based cross entropy loss is presented. Canonical Correlation Analysis is performed to enhance the correlation between the two modalities in a joint latent embedding space. To investigate the benefits of the proposed approach, a new testing protocol under a multi modal ReID setting is proposed for the test split of the CUHK-PEDES and CUHK-SYSU benchmarks. The experimental results verify that the learnt visual representations are more robust and perform 20% better during retrieval as compared to a single modality system. Ammarah Farooq, Muhammad Awais 0001, Josef Kittler, Ali Akbari 0003, Syed Safwan Khalid |
IJCB | 3 |
| 2020 | A Stacking Ensemble for Anomaly Based Client-Specific Face Spoofing DetectionabstractTo counteract spoofing attacks, the majority of recent approaches to face spoofing attack detection formulate the problem as a binary classification task in which real data and attack-accesses are both used to train spoofing detectors. Although the classical training framework has been demonstrated to deliver satisfactory results, its robustness to unseen attacks is debatable. Inspired by the recent success of anomaly detection models in face spoofing detection, we propose an ensemble of one-class classifiers fused by a Stacking ensemble method to reduce the generalisation error in the more realistic unseen attack scenario. To be consistent with this scenario, anomalous samples are considered neither for training the component anomaly classifiers nor for the design of the Stacking ensemble. To achieve better face-anti spoofing results, we adopt client-specific information to build both constituent classifiers as well as the Stacking combiner. Besides, we propose a novel 2-stage Genetic Algorithm to further improve the generalisation performance of Stacking ensemble. We evaluate the effectiveness of the proposed systems on publicly available face anti-spoofing databases including Replay-Attack, Replay-Mobile and Rose-Youtu. The experimental results following the unseen attack evaluation protocol confirm the merits of the proposed model. Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler |
ICIP | 4 |
| 2020 | A Flatter Loss for Bias Mitigation in Cross-dataset Facial Age EstimationabstractThe most existing studies in the facial age estimation assume training and test images are captured under similar shooting conditions. However, this is rarely valid in real-worlds applications, where training and test sets usually have different characteristics. In this paper, we advocate a cross-dataset protocol for age estimation benchmarking. In order to improve the cross-dataset age estimation performance, we mitigate the inherent bias caused by the learning algorithm itself. To this end, we propose a novel loss function that is more effective for neural network training. The relative smoothness of the proposed loss function is its advantage with regards to the optimisation process performed by stochastic gradient descent (SGD). Compared with existing loss functions, the lower gradient of the proposed loss function leads to the convergence of SGD to a better optimum point, and consequently a better generalisation. The cross-dataset experimental results demonstrate the superiority of the proposed method over the state-of-the-art algorithms in terms of accuracy and generalisation capability. Ali Akbari 0003, Muhammad Awais 0001, Zhenhua Feng 0001, Ammarah Farooq, Josef Kittler |
ICPR | 5 |
| 2020 | Angular Sparsemax for Face RecognitionabstractThe Softmax prediction function is widely used to train Deep Convolutional Neural Networks (DCNNs) for large-scale face recognition and other applications. The limitation of the softmax activation is that the resulting probability distribution always has a full support. This full support leads to larger intraclass variations. In this paper, we formulate a novel loss function, called Angular Sparsemax for face recognition. The proposed loss function promotes sparseness of the hypotheses prediction function similar to Sparsemax [1] with Fenchel-Young regularisation. By introducing an additive angular margin on the score vector, the discriminatory power of the face embedding is further improved. The proposed loss function is experimentally validated on several databases in terms of recognition accuracy. Its performance compares well with the state of the art Arcface loss. Chi-Ho Chan, Josef Kittler |
ICPR | 2 |
| 2020 | Subspace Clustering via Joint Unsupervised Feature SelectionabstractAny high-dimensional data arising from practical applications usually contains irrelevant features that may impact on the performance of existing subspace clustering methods. This paper proposes a novel subspace clustering method which reconstructs the feature matrix by the means of unsupervised feature selection (UFS) to achieve a better dictionary for subspace clustering (SC). Different from most existing clustering methods, the proposed approach uses the reconstructed feature matrix as the dictionary rather than the original data matrix. As the feature matrix reconstructed by representative features is more discriminative and closer to the ground-truth, it results in improved performance. The corresponding non-convex optimization problem is effectively solved using the half-quadratic and augmented Lagrange multiplier methods. Extensive experiments on four real datasets demonstrate the effectiveness of the proposed method. Wenhua Dong, Xiaojun Wu 0001, Hui Li 0037, Zhenhua Feng 0001, Josef Kittler |
ICPR | 5 |
| 2020 | Adaptive Context-Aware Discriminative Correlation Filters for Robust Visual Object TrackingabstractIn recent years, Discriminative Correlation Filters (DCFs) have gained popularity due to their superior performance in visual object tracking. However, existing DCF trackers usually learn filters using fixed attention mechanisms that focus on the centre of an image and suppresses filter amplitudes in surroundings. In this paper, we propose an Adaptive Context-Aware Discriminative Correlation Filter (ACA-DCF) that is able to improve the existing DCF formulation with complementary attention mechanisms. Our ACA-DCF integrates foreground attention and background attention for complementary context-aware filter learning. More importantly, we ameliorate the design using an adaptive weighting strategy that takes complex appearance variations into account. The experimental results obtained on several well-known benchmarks demonstrate the effectiveness and superiority of the proposed method over the state-of-the-art approaches. Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler |
ICPR | 4 |
| 2020 | Fast Discrete Cross-Modal Hashing Based on Label Relaxation and Matrix FactorizationabstractIn recent years, cross-media retrieval has drawn considerable attention due to the exponential growth of multimedia data. Many hashing approaches have been proposed for the cross-media search task. However, there are still open problems that warrant investigation. For example, most existing supervised hashing approaches employ a binary label matrix, which achieves small margins between wrong labels (0) and true labels (1). This may affect the retrieval performance by generating many false negatives and false positives. In addition, some methods adopt a relaxation scheme to solve the binary constraints, which may cause large quantization errors. There are also some discrete hashing methods that have been presented, but most of them are time-consuming. To conquer these problems, we present a label relaxation and discrete matrix factorization method (LRMF) for cross-modal retrieval. It offers a number of innovations. First of all, the proposed approach employs a novel label relaxation scheme to control the margins adaptively, which has the benefit of reducing the quantization error. Second, by virtue of the proposed discrete matrix factorization method designed to learn the binary codes, large quantization errors caused by relaxation can be avoided. The experimental results obtained on two widely-used databases demonstrate that LRMF outperforms state-of-the-art cross- media methods. Donglin Zhang 0001, Xiaojun Wu 0001, Zhen Liu 0015, Jun Yu 0011, Josef Kittler |
ICPR | 5 |
| 2020 | Rectified Wing Loss for Efficient and Robust Facial Landmark Localisation with Convolutional Neural NetworksabstractAbstract Efficient and robust facial landmark localisation is crucial for the deployment of real-time face analysis systems. This paper presents a new loss function, namely Rectified Wing (RWing) loss, for regression-based facial landmark localisation with Convolutional Neural Networks (CNNs). We first systemically analyse different loss functions, including L2, L1 and smooth L1. The analysis suggests that the training of a network should pay more attention to small-medium errors. Motivated by this finding, we design a piece-wise loss that amplifies the impact of the samples with small-medium errors. Besides, we rectify the loss function for very small errors to mitigate the impact of inaccuracy of manual annotation. The use of our RWing loss boosts the performance significantly for regression-based CNNs in facial landmarking, especially for lightweight network architectures. To address the problem of under-representation of samples with large pose variations, we propose a simple but effective boosting strategy, referred to as pose-based data balancing. In particular, we deal with the data imbalance problem by duplicating the minority training samples and perturbing them by injecting random image rotation, bounding box translation and other data augmentation strategies. Last, the proposed approach is extended to create a coarse-to-fine framework for robust and efficient landmark localisation. Moreover, the proposed coarse-to-fine framework is able to deal with the small sample size problem effectively. The experimental results obtained on several well-known benchmarking datasets demonstrate the merits of our RWing loss and prove the superiority of the proposed method over the state-of-the-art approaches. Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001, Xiaojun Wu 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | Exploiting Deep Learning in Limited-Fronthaul Cell-Free Massive MIMO UplinkabstractA cell-free massive multiple-input multiple-output (MIMO) uplink is considered, where quantize-and-forward (QF) refers to the case where both the channel estimates and the received signals are quantized at the access points (APs) and forwarded to a central processing unit (CPU) whereas in combine-quantize-and-forward (CQF), the APs send the quantized version of the combined signal to the CPU. To solve the non-convex sum rate maximization problem, a heuristic sub-optimal scheme is exploited to convert the power allocation problem into a standard geometric programme (GP). We exploit the knowledge of the channel statistics to design the power elements. Employing large-scale-fading (LSF) with a deep convolutional neural network (DCNN) enables us to determine a mapping from the LSF coefficients and the optimal power through solving the sum rate maximization problem using the quantized channel. Four possible power control schemes are studied, which we refer to as i) small-scale fading (SSF)-based QF; ii) LSF-based CQF; iii) LSF use-and-then-forget (UatF)-based QF; and iv) LSF deep learning (DL)-based QF, according to where channel estimation is performed and exploited and how the optimization problem is solved. Numerical results show that for the same fronthaul rate, the throughput significantly increases thanks to the mapping obtained using DCNN. Manijeh Bashar, Ali Akbari 0003, K. Cumanan, Hien Quoc Ngo, Alister Burr, Pei Xiao 0001, Mérouane Debbah, Josef Kittler |
IEEE J. Sel. Areas Commun. | 8 |
| 2020 | Self-grouping convolutional neural networks
Qingbei Guo, Xiaojun Wu 0001, Josef Kittler, Zhiquan Feng |
Neural Networks | 3 |
| 2020 | Learning image features with fewer labels using a semi-supervised deep convolutional networkabstractLearning feature embeddings for pattern recognition is a relevant task for many applications. Deep learning methods such as convolutional neural networks can be employed for this assignment with different training strategies: leveraging pre-trained models as baselines; training from scratch with the target dataset; or fine-tuning from the pre-trained model. Although there are separate systems used for learning features from labelled and unlabelled data, there are few models combining all available information. Therefore, in this paper, we present a novel semi-supervised deep network training strategy that comprises a convolutional network and an autoencoder using a joint classification and reconstruction loss function. We show our network improves the learned feature embedding when including the unlabelled data in the training process. The results using the feature embedding obtained by our network achieve better classification accuracy when compared with competing methods, as well as offering good generalisation in the context of transfer learning. Furthermore, the proposed network ensemble and loss function is highly extensible and applicable in many recognition tasks. Fernando Pereira dos Santos, Cemre Zor, Josef Kittler, Moacir Ponti |
Neural Networks | 3 |
| 2020 | Learning a representation with the block-diagonal structure for pattern classification
He-Feng Yin, Xiaojun Wu 0001, Josef Kittler, Zhenhua Feng 0001 |
Pattern Anal. Appl. | 3 |
| 2020 | Learning discriminative hashing codes for cross-modal retrieval based on multi-view features
Jun Yu 0011, Xiaojun Wu 0001, Josef Kittler |
Pattern Anal. Appl. | 3 |
| 2020 | Covariance descriptors on a Gaussian manifold and their application to image set classification
Kai-Xuan Chen 0001, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 4 |
| 2020 | Noise-robust dictionary learning with slack block-Diagonal structure for face recognition
Zhe Chen 0018, Xiaojun Wu 0001, He-Feng Yin, Josef Kittler |
Pattern Recognit. | 4 |
| 2020 | An accelerated correlation filter tracker
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 4 |
| 2020 | Discriminative block-diagonal covariance descriptors for image set classification
Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. Lett. | 3 |
| 2020 | Low-rank discriminative least squares regression for image classification
Zhe Chen 0018, Xiaojun Wu 0001, Josef Kittler |
Signal Process. | 3 |
| 2020 | Learning Low-Rank and Sparse Discriminative Correlation Filters for Coarse-to-Fine Visual Object TrackingabstractDiscriminative correlation filter (DCF) has achieved advanced performance in visual object tracking with remarkable efficiency guaranteed by its implementation in the frequency domain. However, the effect of the structural relationship of DCF and object features has not been adequately explored in the context of the filter design. To remedy this deficiency, this paper proposes a Low-rank and Sparse DCF (LSDCF) that improves the relevance of features used by discriminative filters. To be more specific, we extend the classical DCF paradigm from ridge regression to lasso regression, and constrain the estimate to be of low-rank across frames, thus identifying and retaining the informative filters distributed on a low-dimensional manifold. To this end, specific temporal-spatial-channel configurations are adaptively learned to achieve enhanced discrimination and interpretability. In addition, we analyse the complementary characteristics between hand-crafted features and deep features, and propose a coarse-to-fine heuristic tracking strategy to further improve the performance of our LSDCF. Last, the augmented Lagrange multiplier optimisation method is used to achieve efficient optimisation. The experimental results obtained on a number of well-known benchmarking datasets, including OTB2013, OTB50, OTB100, TC128, UAV123, VOT2016 and VOT2018, demonstrate the effectiveness and robustness of the proposed method, delivering outstanding performance compared to the state-of-the-art trackers. Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | MDLatLRR: A Novel Decomposition Method for Infrared and Visible Image FusionabstractImage decomposition is crucial for many image processing tasks, as it allows to extract salient features from source images. A good image decomposition method could lead to a better performance, especially in image fusion tasks. We propose a multi-level image decomposition method based on latent low-rank representation(LatLRR), which is called MDLatLRR. This decomposition method is applicable to many image processing fields. In this paper, we focus on the image fusion task. We build a novel image fusion framework based on MDLatLRR which is used to decompose source images into detail parts(salient features) and base parts. A nuclear-norm based fusion strategy is used to fuse the detail parts and the base parts are fused by an averaging strategy. Compared with other state-of-the-art fusion methods, the proposed algorithm exhibits better fusion performance in both subjective and objective evaluation. Hui Li 0037, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Image Process. | 3 |
| 2019 | Spoofing Attack Detection by Anomaly DetectionabstractSpoofing attacks on biometric systems can seriously compromise their practical utility. In this paper we focus on face spoofing detection. The majority of papers on spoofing attack detection formulate the problem as a two or multiclass learning task, attempting to separate normal accesses from samples of different types of spoofing attacks. In this paper we adopt the anomaly detection approach proposed in [1], where the detector is trained on genuine accesses only using one-class classifiers and investigate the merit of subject specific solutions. We show experimentally that subject specific models are superior to the commonly used client independent method. We also demonstrate that the proposed approach is more robust than multiclass formulations to unseen attacks. Soroush Fatemifar, Shervin Rahimzadeh Arashloo, Muhammad Awais 0001, Josef Kittler |
ICASSP | 4 |
| 2019 | Divergence Based Weighting for Information Channels in Deep Convolutional Neural Networks for Bird Audio DetectionabstractIn this paper, we address the problem of bird audio detection and propose a new convolutional neural network architecture together with a divergence based information channel weighing strategy in order to achieve improved state-of-the-art performance and faster convergence. The effectiveness of the methodology is shown on the Bird Audio Detection Challenge 2018 (Detection and Classification of Acoustic Scenes and Events Challenge, Task 3) development data set. Cemre Zor, Muhammad Awais 0001, Josef Kittler, Miroslaw Bober, Syed Sameed Husain, Qiuqiang Kong, Christian Kroos |
ICASSP | 3 |
| 2019 | Joint Group Feature Selection and Discriminative Filter Learning for Robust Visual Object TrackingabstractWe propose a new Group Feature Selection method for Discriminative Correlation Filters (GFS-DCF) based visual object tracking. The key innovation of the proposed method is to perform group feature selection across both channel and spatial dimensions, thus to pinpoint the structural relevance of multi-channel features to the filtering system. In contrast to the widely used spatial regularisation or feature selection methods, to the best of our knowledge, this is the first time that channel selection has been advocated for DCF-based tracking. We demonstrate that our GFS-DCF method is able to significantly improve the performance of a DCF tracker equipped with deep neural network features. In addition, our GFS-DCF enables joint feature selection and filter learning, achieving enhanced discrimination and interpretability of the learned filters. To further improve the performance, we adaptively integrate historical information by constraining filters to be smooth across temporal frames, using an efficient low-rank approximation. By design, specific temporal-spatial-channel configurations are dynamically learned in the tracking process, highlighting the relevant features, and alleviating the performance degrading impact of less discriminative representations and reducing information redundancy. The experimental results obtained on OTB2013, OTB2015, VOT2017, VOT2018 and TrackingNet demonstrate the merits of our GFS-DCF and its superiority over the state-of-the-art trackers. The code is publicly available at \url{https://github.com/XU-TIANYANG/GFS-DCF}. Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler |
ICCV | 4 |
| 2019 | Non-negative Representation Based Discriminative Dictionary Learning for Face Recognition
Zhe Chen 0018, Xiaojun Wu 0001, Josef Kittler |
ICIG (1) | 3 |
| 2019 | Robust 3D Face Alignment with Efficient Fully Convolutional Neural Networks
Xiaojun Wu 0001, Josef Kittler |
ICIG (2) | 3 |
| 2019 | Discriminative Supervised Hashing for Cross-Modal Similarity Search
Jun Yu 0011, Xiaojun Wu 0001, Josef Kittler |
Image Vis. Comput. | 3 |
| 2019 | Sparse subspace clustering via nonconvex approximation
Wenhua Dong, Xiaojun Wu 0001, Josef Kittler, He-Feng Yin |
Pattern Anal. Appl. | 3 |
| 2019 | A sparse regularized nuclear norm based matrix regression for face recognition with contiguous occlusion
Zhe Chen 0018, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. Lett. | 3 |
| 2019 | Sparse subspace clustering via smoothed ℓp minimization
Wenhua Dong, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. Lett. | 3 |
| 2019 | Mining Hard Augmented Samples for Robust Facial Landmark Localization With CNNsabstractEffective data augmentation is crucial for facial landmark localization with convolutional neural networks (CNNs). In this letter, we investigate different data augmentation techniques that can be used to generate sufficient data for training CNN-based facial landmark localization systems. To the best of our knowledge, this is the first study that provides a systematic analysis of different data augmentation techniques in the area. In addition, an online hard augmented example mining (HAEM) strategy is advocated for further performance boosting. We examine the effectiveness of those techniques using a regression-based CNN architecture. The experimental results obtained on the AFLW and COFW datasets demonstrate the importance of data augmentation and the effectiveness of HAEM. The performance achieved using these techniques is superior to the state-of-the-art algorithms. Zhenhua Feng 0001, Josef Kittler, Xiaojun Wu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2019 | Delta Divergence: A Novel Decision Cognizant Measure of Classifier IncongruenceabstractIn pattern recognition, disagreement between two classifiers regarding the predicted class membership of an observation can be indicative of an anomaly and its nuance. Since, in general, classifiers base their decisions on class a posteriori probabilities, the most natural approach to detecting classifier incongruence is to use divergence. However, existing divergences are not particularly suitable to gauge classifier incongruence. In this paper, we postulate the properties that a divergence measure should satisfy and propose a novel divergence measure, referred to as delta divergence. In contrast to existing measures, it focuses on the dominant (most probable) hypotheses and, thus, reduces the effect of the probability mass distributed over the non dominant hypotheses (clutter). The proposed measure satisfies other important properties, such as symmetry, and independence of classifier confidence. The relationship of the proposed divergence to some baseline measures, and its superiority, is shown experimentally. Josef Kittler, Cemre Zor |
IEEE Trans. Cybern. | 1 |
| 2019 | Learning Adaptive Discriminative Correlation Filters via Temporal Consistency Preserving Spatial Feature Selection for Robust Visual Object TrackingabstractWith efficient appearance learning models, discriminative correlation filter (DCF) has been proven to be very successful in recent video object tracking benchmarks and competitions. However, the existing DCF paradigm suffers from two major issues, i.e., spatial boundary effect and temporal filter degradation. To mitigate these challenges, we propose a new DCF-based tracking method. The key innovations of the proposed method include adaptive spatial feature selection and temporal consistent constraints, with which the new tracker enables joint spatial-temporal filter learning in a lower dimensional discriminative manifold. More specifically, we apply structured spatial sparsity constraints to multi-channel filters. Consequently, the process of learning spatial filters can be approximated by the lasso regularization. To encourage temporal consistency, the filter model is restricted to lie around its historical value and updated locally to preserve the global structure in the manifold. Last, a unified optimization framework is proposed to jointly select temporal consistency preserving spatial features and learn discriminative filters with the augmented Lagrangian method. Qualitative and quantitative evaluations have been conducted on a number of well-known benchmarking datasets such as OTB2013, OTB50, OTB100, Temple-Colour, UAV123, and VOT2018. The experimental results demonstrate the superiority of the proposed method over the state-of-the-art approaches. Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Image Process. | 4 |
| 2018 | Wing Loss for Robust Facial Landmark Localisation With Convolutional Neural NetworksabstractWe present a new loss function, namely Wing loss, for robust facial landmark localisation with Convolutional Neural Networks (CNNs). We first compare and analyse different loss functions including L2, L1 and smooth L1. The analysis of these loss functions suggests that, for the training of a CNN-based localisation model, more attention should be paid to small and medium range errors. To this end, we design a piece-wise loss function. The new loss amplifies the impact of errors from the interval (-w, w) by switching from L1 loss to a modified logarithm function. To address the problem of under-representation of samples with large out-of-plane head rotations in the training set, we propose a simple but effective boosting strategy, referred to as pose-based data balancing. In particular, we deal with the data imbalance problem by duplicating the minority training samples and perturbing them by injecting random image rotation, bounding box translation and other data augmentation approaches. Last, the proposed approach is extended to create a two-stage framework for robust facial landmark localisation. The experimental results obtained on AFLW and 300W demonstrate the merits of the Wing loss function, and prove the superiority of the proposed method over the state-of-the-art approaches. Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001, Patrik Huber 0001, Xiaojun Wu 0001 |
CVPR | 2 |
| 2018 | Semi-supervised Adversarial Learning to Generate Photorealistic Face Images of New Identities from 3D Morphable Model
Baris Gecer, Binod Bhattarai, Josef Kittler, Tae-Kyun Kim 0001 |
ECCV (11) | 3 |
| 2018 | Evaluation of Dense 3D Reconstruction from 2D Face Images in the WildabstractThis paper investigates the evaluation of dense 3D face reconstruction from a single 2D image in the wild. To this end, we organise a competition that provides a new benchmark dataset that contains 2000 2D facial images of 135 subjects as well as their 3D ground truth face scans. In contrast to previous competitions or challenges, the aim of this new benchmark dataset is to evaluate the accuracy of a 3D dense face reconstruction algorithm using real, accurate and high-resolution 3D ground truth face scans. In addition to the dataset, we provide a standard protocol as well as a Python script for the evaluation. Last, we report the results obtained by three state-of-the-art 3D face reconstruction systems on the new benchmark dataset. The competition is organised along with the 2018 13th IEEE Conference on Automatic Face & Gesture Recognition. Zhenhua Feng 0001, Patrik Huber 0001, Josef Kittler, Peter J. B. Hancock, Xiaojun Wu 0001, Qijun Zhao, Willem P. Koppen, Matthias Rätsch |
FG | 3 |
| 2018 | Intelligent Signal Processing Mechanisms for Nuanced Anomaly Detection in Action Audio-Visual Data StreamsabstractWe consider the problem of anomaly detection in an audiovisual analysis system designed to interpret sequences of actions from visual and audio cues. The scene activity recognition is based on a generative framework, with a high-level inference model for contextual recognition of sequences of actions. The system is endowed with anomaly detection mechanisms, which facilitate differentiation of various types of anomalies. This is accomplished using intelligence provided by a classifier incongruence detector, classifier confidence module and data quality assessment system, in addition to the classical outlier detection module. The paper focuses on one of the mechanisms, the classifier incongruence detector, the purpose of which is to flag situations when the video and audio modalities disagree in action interpretation. We demonstrate the merit of using the Delta divergence measure for this purpose. We show that this measure significantly enhances the incongruence detection rate in the Human Action Manipulation complex activity recognition data set. Josef Kittler, Ioannis Kaloskampis, Cemre Zor, Yulia Hicks, Wenwu Wang 0001 |
ICASSP | 1 |
| 2018 | Riemannian kernel based Nyström method for approximate infinite-dimensional covariance descriptors with application to image set classificationabstractIn the domain of pattern recognition, using the CovDs (Covariance Descriptors) to represent data and taking the metrics of the resulting Riemannian manifold into account have been widely adopted for the task of image set classification. Recently, it has been proven that infinite-dimensional CovDs are more discriminative than their low-dimensional counterparts. However, the form of infinite-dimensional CovDs is implicit and the computational load is high. We propose a novel framework for representing image sets by approximating infinite-dimensional CovDs in the paradigm of the Nyström method based on a Riemannian kernel. We start by modeling the images via CovDs, which lie on the Riemannian manifold spanned by SPD (Symmetric Positive Definite) matrices. We then extend the Nyström method to the SPD manifold and obtain the approximations of CovDs in RKHS (Reproducing Kernel Hilbert Space). Finally, we approximate infinite-dimensional CovDs via these approximations. Empirically, we apply our framework to the task of image set classification. The experimental results obtained on three benchmark datasets show that our proposed approximate infinite-dimensional CovDs outperform the original CovDs. Kai-Xuan Chen 0001, Xiaojun Wu 0001, Rui Wang 0050, Josef Kittler |
ICPR | 4 |
| 2018 | Infrared and Visible Image Fusion using a Deep Learning FrameworkabstractIn recent years, deep learning has become a very active research tool which is used in many image processing fields. In this paper, we propose an effective image fusion method using a deep learning framework to generate a single image which contains all the features from infrared and visible images. First, the source images are decomposed into base parts and detail content. Then the base parts are fused by weighted-averaging. For the detail content, we use a deep learning network to extract multi-layer features. Using these features, we use$l_{1}$-norm and weighted-average strategy to generate several candidates of the fused detail content. Once we get these candidates, the max selection strategy is used to get the final fused detail content. Finally, the fused image will be reconstructed by combining the fused base part and the detail content. The experimental results demonstrate that our proposed method achieves state-of-the-art performance in both objective assessment and visual quality. The Code of our fusion method is available at https://github.com/exceptionLi/imagefusion_deeplearning. Hui Li 0037, Xiaojun Wu 0001, Josef Kittler |
ICPR | 3 |
| 2018 | Multiple Manifolds Metric Learning with Application to Image Set ClassificationabstractIn image set classification, a considerable advance has been made by modeling the original image sets by second order statistics or linear subspace, which typically lie on the Riemannian manifold. Specifically, they are Symmetric Positive Definite (SPD) manifold and Grassmann manifold respectively, and some algorithms have been developed on them for classification tasks. Motivated by the inability of existing methods to extract discriminatory features for data on Riemannian manifolds, we propose a novel algorithm which combines multiple manifolds as the features of the original image sets. In order to fuse these manifolds, the well-studied Riemannian kernels have been utilized to map the original Riemannian spaces into high dimensional Hilbert spaces. A metric Learning method has been devised to embed these kernel spaces into a lower dimensional common subspace for classification. The state-of-the-art results achieved on three datasets corresponding to two different classification tasks, namely face recognition and object categorization, demonstrate the effectiveness of the proposed method. Rui Wang 0050, Xiaojun Wu 0001, Kai-Xuan Chen 0001, Josef Kittler |
ICPR | 4 |
| 2018 | Non-negative Subspace Representation Learning Scheme for Correlation Filter Based TrackingabstractDiscriminative correlation filter (DCF) based tracking methods have achieved great success recently. However, the temporal learning scheme in the current paradigm is of a linear recursion form determined by a fixed learning rate which can not adaptively feedback appearance variations. In this paper, we propose a unified non-negative subspace representation constrained leaning scheme for DCF. The subspace is constructed by several templates with auxiliary memory mechanisms. Then the current template is projected onto the subspace to find the non-negative representation and to determine the corresponding template weights. Our learning scheme enables efficient combination of correlation filter and subspace structure. The experimental results on OTB50 demonstrate the effectiveness of our learning formulation. Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
ICPR | 3 |
| 2018 | Person Re-Identification with Vision and LanguageabstractIn this paper we propose a new approach to person re-identification using images and natural language descriptions. We propose a joint vision and language model based on CNN and LSTM architectures to match across the two modalities as well as to enrich visual examples for which there are no language descriptions. We also introduce new annotations in the form of natural language descriptions for two standard Re-ID benchmarks, namely CUHK03 and VIPeR. We perform experiments on these two datasets with techniques based on CNN, hand-crafted features as well as LSTM for analysing visual and natural description data. We investigate and demonstrate the advantages of using natural language descriptions compared to attributes as well as CNN compared to LSTM in the context of Re-ID. We show that the joint use of language and vision can significantly improve the state-of-the-art performance on standard Re-ID benchmarks. Fei Yan 0001, Josef Kittler, Krystian Mikolajczyk |
ICPR | 2 |
| 2018 | Semi-supervised Hashing for Semi-Paired Cross-View RetrievalabstractRecently, hashing techniques have gained importance in large-scale retrieval tasks because of their retrieval speed. Most of the existing cross-view frameworks assume that data are well paired. However, the fully-paired multiview situation is not universal in real applications. The aim of the method proposed in this paper is to learn the hashing function for semi-paired cross-view retrieval tasks. To utilize the label information of partial data, we propose a semi-supervised hashing learning framework which jointly performs feature extraction and classifier learning. The experimental results on two datasets show that our method outperforms several state-of-the-art methods in terms of retrieval accuracy. Jun Yu 0011, Xiaojun Wu 0001, Josef Kittler |
ICPR | 3 |
| 2018 | Improve the Spoofing Resistance of Multimodal Verification with Representation-Based Measures
Zengxi Huang, Zhenhua Feng 0001, Josef Kittler, Yiguang Liu |
PRCV (3) | 3 |
| 2018 | Robust Low-Rank Recovery with a Distance-Measure Structure for Face Recognition
Zhe Chen 0018, Xiaojun Wu 0001, He-Feng Yin, Josef Kittler |
PRICAI | 4 |
| 2018 | Error sensitivity analysis of Delta divergence - a novel measure for classifier incongruence detectionabstractThe state of classifier incongruence in decision making systems incorporating multiple classifiers is often an indicator of anomaly caused by an unexpected observation or an unusual situation. Its assessment is important as one of the key mechanisms for domain anomaly detection. In this paper, we investigate the sensitivity of Delta divergence, a novel measure of classifier incongruence, to estimation errors. Statistical properties of Delta divergence are analysed both theoretically and experimentally. The results of the analysis provide guidelines on the selection of threshold for classifier incongruence detection based on this measure. Josef Kittler, Cemre Zor, Ioannis Kaloskampis, Yulia Hicks, Wenwu Wang 0001 |
Pattern Recognit. | 1 |
| 2018 | Gaussian mixture 3D morphable face modelabstract3D Morphable Face Models (3DMM) have been used in pattern recognition for some time now. They have been applied as a basis for 3D face recognition, as well as in an assistive role for 2D face recognition to perform geometric and photometric normalisation of the input image, or in 2D face recognition system training. The statistical distribution underlying 3DMM is Gaussian. However, the single-Gaussian model seems at odds with reality when we consider different cohorts of data, e.g. Black and Chinese faces. Their means are clearly different. This paper introduces the Gaussian Mixture 3DMM (GM-3DMM) which models the global population as a mixture of Gaussian subpopulations, each with its own mean. The proposed GM-3DMM extends the traditional 3DMM naturally, by adopting a shared covariance structure to mitigate small sample estimation problems associated with data in high dimensional spaces. We construct a GM-3DMM, the training of which involves a multiple cohort dataset, SURREY-JNU, comprising 942 3D face scans of people with mixed backgrounds. Experiments in fitting the GM-3DMM to 2D face images to facilitate their geometric and photometric normalisation for pose and illumination invariant face recognition demonstrate the merits of the proposed mixture of Gaussians 3D face model. Willem P. Koppen, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001, William J. Christmas, Xiaojun Wu 0001, He-Feng Yin |
Pattern Recognit. | 3 |
| 2018 | Articulated motion and deformable objects
Jun Wan 0001, Sergio Escalera, Francisco José Perales López, Josef Kittler |
Pattern Recognit. | 4 |
| 2018 | Dictionary Integration Using 3D Morphable Face Models for Pose-Invariant Collaborative-Representation-Based ClassificationabstractThe paper presents a dictionary integration algorithm using 3D morphable face models (3DMM) for pose-invariant collaborative-representation-based face classification. To this end, we first fit a 3DMM to the 2D face images of a dictionary to reconstruct the 3D shape and texture of each image. The 3D faces are used to render a number of virtual 2D face images with arbitrary pose variations to augment the training data, by merging the original and rendered virtual samples to create an extended dictionary. Second, to reduce the information redundancy of the extended dictionary and improve the sparsity of reconstruction coefficient vectors using collaborative-representation-based classification (CRC), we exploit an on-line class elimination scheme to optimise the extended dictionary by identifying the training samples of the most representative classes for a given query. The final goal is to perform pose-invariant face classification using the proposed dictionary integration method and the on-line pruning strategy under the CRC framework. Experimental results obtained for a set of well-known face data sets demonstrate the merits of the proposed method, especially its robustness to pose variations. Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Josef Kittler, Xiaojun Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2017 | Dynamic Attention-Controlled Cascaded Shape Regression Exploiting Training Data Augmentation and Fuzzy-Set Sample WeightingabstractWe present a new Cascaded Shape Regression (CSR) architecture, namely Dynamic Attention-Controlled CSR (DAC-CSR), for robust facial landmark detection on unconstrained faces. Our DAC-CSR divides facial landmark detection into three cascaded sub-tasks: face bounding box refinement, general CSR and attention-controlled CSR. The first two stages refine initial face bounding boxes and output intermediate facial landmarks. Then, an online dynamic model selection method is used to choose appropriate domain-specific CSRs for further landmark refinement. The key innovation of our DAC-CSR is the fault-tolerant mechanism, using fuzzy set sample weighting, for attention-controlled domain-specific model training. Moreover, we advocate data augmentation with a simple but effective 2D profile face generator, and context-aware feature extraction for better facial feature representation. Experimental results obtained on challenging datasets demonstrate the merits of our DAC-CSR over the state-of-the-art methods. Zhenhua Feng 0001, Josef Kittler, William J. Christmas, Patrik Huber 0001, Xiaojun Wu 0001 |
CVPR | 2 |
| 2017 | Optical-flow features empirical mode decomposition for motion anomaly detectionabstractIn video data analysis of dynamic scenes, temporal characteristics of moving objects play an important role in decision-making. However, the temporal consistency of typical features used for video interpretation is low due to the overlap of the spectra of informative video signal component and the stochastic variations perturbing it. We propose a novel method for object motion anomaly detection in video designed to overcome this problem. It is based on empirical mode decomposition. We show in experiments on a benchmarking dataset that the deterministic component of an optical flow feature obtained using the proposed method is able to isolate the periodic behaviour of the motion from the stochastic values, facilitating much simpler analysis of the motion patterns and achieving impressive anomaly detection performance. Moacir Ponti, Tiago S. Nazaré, Josef Kittler |
ICASSP | 3 |
| 2017 | Maritime anomaly detection in ferry tracksabstractThis paper proposes a methodology for the automatic detection of anomalous shipping tracks traced by ferries. The approach comprises a set of models as a basis for outlier detection: A Gaussian process (GP) model regresses displacement information collected over time, and a Markov chain based detector makes use of the direction (heading) information. GP regression is performed together with Median Absolute Deviation to account for contaminated training data. The methodology utilizes the coordinates of a given ferry recorded on a second by second basis via Automatic Identification System. Its effectiveness is demonstrated on a dataset collected in the Solent area. Cemre Zor, Josef Kittler |
ICASSP | 2 |
| 2017 | An anomaly detection approach to face spoofing detection: A new formulation and evaluation protocolabstractFace anti-spoofing problem can be quite challenging due to various factors including diversity of face spoofing attacks, any new means of spoofing, the problem of imaging sensor interoperability and other environmental factors in addition to the small sample size. Taking into account these observations, in this work, first, a new evaluation protocol called “innovative attack evaluation protocol” to study the effect of occurrence of unseen attack types is proposed which better reflects the realistic conditions in spoofing attacks. Second, a new formulation of the problem based on the anomaly detection concept is proposed where the training data comes from the positive class only. The test data, of course, may come from the positive or negative class. Finally, a thorough evaluation and comparison of 20 different one-class and two-class systems is performed and demonstrated that the anomaly-based formulation is not inferior as compared with the conventional two-class approach. Shervin Rahimzadeh Arashloo, Josef Kittler |
IJCB | 2 |
| 2017 | Unconstrained Face Detection and Open-Set Face Recognition ChallengeabstractFace detection and recognition benchmarks have shifted toward more difficult environments. The challenge presented in this paper addresses the next step in the direction of automatic detection and identification of people from outdoor surveillance cameras. While face detection has shown remarkable success in images collected from the web, surveillance cameras include more diverse occlusions, poses, weather conditions and image blur. Although face verification or closed-set face identification have surpassed human capabilities on some datasets, open-set identification is much more complex as it needs to reject both unknown identities and false accepts from the face detector. We show that unconstrained face detection can approach high detection rates albeit with moderate false accept rates. By contrast, open-set face recognition is currently weak and requires much more attention. Manuel Günther, Peiyun Hu, Christian Herrmann 0001, Chi-Ho Chan, Min Jiang 0003, Shufan Yang, Akshay Raj Dhamija, Deva Ramanan, Jürgen Beyerer, Josef Kittler, Mohamad Al Jazaery, Mohammad Iqbal Nouyed, Guodong Guo, Cezary Stankiewicz, Terrance E. Boult |
IJCB | 10 |
| 2017 | Efficient 3D morphable face model fitting
Guosheng Hu, Fei Yan 0001, Josef Kittler, William J. Christmas, Chi-Ho Chan, Zhenhua Feng 0001, Patrik Huber 0001 |
Pattern Recognit. | 3 |
| 2017 | A decision cognizant Kullback-Leibler divergenceabstractIn decision making systems involving multiple classifiers there is the need to assess classifier (in)congruence, that is to gauge the degree of agreement between their outputs. A commonly used measure for this purpose is the Kullback–Leibler (KL) divergence. We propose a variant of the KL divergence, named decision cognizant Kullback–Leibler divergence (DC-KL), to reduce the contribution of the minority classes, which obscure the true degree of classifier incongruence. We investigate the properties of the novel divergence measure analytically and by simulation studies. The proposed measure is demonstrated to be more robust to minority class clutter. Its sensitivity to estimation noise is also shown to be considerably lower than that of the classical KL divergence. These properties render the DC-KL divergence a much better statistic for discriminating between classifier congruence and incongruence in pattern recognition systems. Moacir Ponti, Josef Kittler, Mateus Riva, Teófilo Emídio de Campos, Cemre Zor |
Pattern Recognit. | 2 |
| 2017 | Special section: CIARP 2015
Alvaro Pardo, Josef Kittler |
Pattern Recognit. Lett. | 2 |
| 2017 | Real-Time 3D Face Fitting and Texture Fusion on In-the-Wild VideosabstractWe present a fully automatic approach to real-time 3D face reconstruction from monocular in-the-wild videos. With the use of a cascaded-regressor-based face tracking and a 3D morphable face model shape fitting, we obtain a semidense 3D face shape. We further use the texture information from multiple frames to build a holistic 3D face representation from the video footage. Our system is able to capture facial expressions and does not require any person-specific training. We demonstrate the robustness of our approach on the challenging 300 Videos in the Wild (300-VW) dataset. Our real-time fitting framework is available as an open-source library at http://4dface.org. Patrik Huber 0001, Philipp Kopp, William J. Christmas, Matthias Rätsch, Josef Kittler |
IEEE Signal Process. Lett. | 5 |
| 2016 | Face Recognition Using a Unified 3D Morphable Model
Guosheng Hu, Fei Yan 0001, Chi-Ho Chan, Weihong Deng, William J. Christmas, Josef Kittler, Neil Robertson 0002 |
ECCV (8) | 6 |
| 2016 | Generating commentaries for tennis videosabstractWe present an approach to automatically generating verbal commentaries for tennis games. We introduce a novel application that requires a combination of techniques from computer vision, natural language processing and machine learning. A video sequence is first analysed using state-of-the-art computer vision methods to track the ball, fit the detected edges to the court model, track the players, and recognise their strokes. Based on the recognised visual attributes we formulate the tennis commentary generation problem in the framework of long short-term memory recurrent neural networks as well as structured SVM. In particular, we investigate pre-embedding of descriptive terms and loss function for LSTM. We introduce a new dataset of 633 annotated pairs of tennis videos and corresponding commentary. We perform an automatic as well as human based evaluation, and demonstrate that the proposed pre-embedding and loss function lead to substantially improved accuracy of the generated commentary. Fei Yan 0001, Krystian Mikolajczyk, Josef Kittler |
ICPR | 3 |
| 2016 | BeamECOC: A local search for the optimization of the ECOC matrixabstractError Correcting Output Coding (ECOC) is a multiclass classification technique in which multiple binary classifiers are trained according to a preset code matrix such that each one learns a separate dichotomy of the classes. While ECOC is one of the best solutions for multi-class problems, one issue which makes it suboptimal is that the training of the base classifiers is done independently of the generation of the code matrix. In this paper, we propose to modify a given ECOC matrix to improve its performance by reducing this decoupling. The proposed algorithm uses beam search to iteratively modify the original matrix, using validation accuracy as a guide. It does not involve further training of the classifiers and can be applied to any ECOC matrix. We evaluate the accuracy of the proposed algorithm (BeamECOC) using 10-fold cross-validation experiments on 6 UCI datasets, using random code matrices of different sizes, and base classifiers of different strengths. Compared to the random ECOC approach, BeamECOC increases the average cross-validation accuracy in 83.3% of the experimental settings involving all datasets, and gives better results than the state-of-the-art in 75% of the scenarios. By employing BeamECOC, it is also possible to reduce the number of columns of a random matrix down to 13% and still obtain comparable or even better results at times. Cemre Zor, Berrin A. Yanikoglu, Erinc Merdivan, Terry Windeatt, Josef Kittler, Ethem Alpaydin |
ICPR | 5 |
| 2016 | Multi-label classification using stacked spectral kernel discriminant analysis
Muhammad Atif Tahir, Josef Kittler, Ahmed Bouridane |
Neurocomputing | 2 |
| 2016 | SALIC: Social Active Learning for Image ClassificationabstractIn this paper, we present SALIC, an active learning method for selecting the most appropriate user tagged images to expand the training set of a binary classifier. The process of active learning can be fully automated in this social context by replacing the human oracle with the images' tags. However, their noisy nature adds further complexity to the sample selection process since, apart from the images' informativeness (i.e., how much they are expected to inform the classifier if we knew their label), our confidence about their actual label should also be maximized (i.e., how certain the oracle is on the images' true contents). The main contribution of this work is in proposing a probabilistic approach for jointly maximizing the two aforementioned quantities. In the examined noisy context, the oracle's confidence is necessary to provide a contextual-based indication of the images' true contents, while the samples' informativeness is required to reduce the computational complexity and minimize the mistakes of the unreliable oracle. To prove this, first, we show that SALIC allows us to select training data as effectively as typical active learning, without the cost of manual annotation. Finally, we argue that the speed-up achieved when learning actively in this social context (where labels can be obtained without the cost of human annotation) is necessary to cope with the continuously growing requirements of large-scale applications. In this respect, we demonstrate that SALIC requires ten times less training data in order to reach the same performance as a straightforward informativeness-agnostic learning approach. Elisavet Chatzilari, Spiros Nikolopoulos, Ioannis Kompatsiaris, Josef Kittler |
IEEE Trans. Multim. | 4 |
| 2016 | Mean-Shift and Sparse Sampling-Based SMC-PHD Filtering for Audio Informed Visual Speaker TrackingabstractThe probability hypothesis density (PHD) filter based on sequential Monte Carlo (SMC) approximation (also known as SMC-PHD filter) has proven to be a promising algorithm for multispeaker tracking. However, it has a heavy computational cost as surviving, spawned, and born particles need to be distributed in each frame to model the state of the speakers and to estimate jointly the variable number of speakers with their states. In particular, the computational cost is mostly caused by the born particles as they need to be propagated over the entire image in every frame to detect the new speaker presence in the view of the visual tracker. In this paper, we propose to use the audio data to improve the visual SMC-PHD (V-SMC-PHD) filter by using the direction of arrival angles of the audio sources to determine when to propagate the born particles and reallocate the surviving and spawned particles. The tracking accuracy of the audio-visual SMC-PHD (AV-SMC-PHD) algorithm is further improved by using a modified mean-shift algorithm to search and climb density gradients iteratively to find the peak of the probability distribution, and the extra computational complexity introduced by mean-shift is controlled with a sparse sampling technique. These improved algorithms, named as AVMS-SMC-PHD and sparse-AVMS-SMC-PHD, respectively, are compared systematically with AV-SMC-PHD and V-SMC-PHD based on the AV16.3, AMI, and CLEAR datasets. Volkan Kilic, Mark Barnard, Wenwu Wang 0001, Adrian Hilton 0001, Josef Kittler |
IEEE Trans. Multim. | 5 |
| 2015 | Fitting 3D Morphable Face Models using local featuresabstractIn this paper, we propose a novel fitting method that uses local image features to fit a 3D Morphable Face Model to 2D images. To overcome the obstacle of optimising a cost function that contains a non-differentiable feature extraction operator, we use a learning-based cascaded regression method that learns the gradient direction from data. The method allows to simultaneously solve for shape and pose parameters. Our method is thoroughly evaluated on Morphable Model generated data and first results on real data are presented. Compared to traditional fitting methods, which use simple raw features like pixel colour or edge maps, local features have been shown to be much more robust against variations in imaging conditions. Our approach is unique in that we are the first to use local features to fit a 3D Morphable Model. Because of the speed of our method, it is applicable for realtime applications. Our cascaded regression framework is available as an open source library at github.com/patrikhuber/superviseddescent. Patrik Huber 0001, Zhenhua Feng 0001, William J. Christmas, Josef Kittler, Matthias Rätsch |
ICIP | 4 |
| 2015 | Audio informed visual speaker tracking with SMC-PHD filterabstractSequential Monte Carlo probability hypothesis density (SMC-PHD) filter has received much interest in the field of nonlinear non-Gaussian visual tracking due to its ability to handle a variable number of speakers. The SMC-PHD filter employs surviving, spawned and born particles to model the state of the speakers and jointly estimates the variable number of speakers with their states. The born particles play a critical role in the detection of new speakers, which makes it necessary to propagate them in each frame. However, this increases the computational cost of the visual tracker. Here, we propose to use audio data to determine when to propagate the born particles and re-allocate the surviving and spawned particles. In our framework, we employ audio data as an aid to visual SMC-PHD (V-SMC-PHD) filter by using the direction of arrival (DOA) angles of the audio sources to reshape the distribution of the particles. Experimental results on the AV16:3 dataset with multi-speaker sequences show that our proposed audio-visual SMC-PHD (AV-SMC-PHD) filter improves the tracking performance in terms of estimation accuracy and computational efficiency. Volkan Kilic, Mark Barnard, Wenwu Wang 0001, Adrian Hilton 0001, Josef Kittler |
ICME | 5 |
| 2015 | Assessment of algorithms for mitosis detection in breast cancer histopathology images
Mitko Veta, Paul J. van Diest, Stefan M. Willems, Anant Madabhushi, Angel Cruz-Roa, Fabio A. González 0001, Anders Boesen Lindbo Larsen, Jacob S. Vestergaard, Anders Bjorholm Dahl, Dan C. Ciresan, Jürgen Schmidhuber, Alessandro Giusti, Luca Maria Gambardella, Faik Boray Tek, Thomas Walter 0003, Ching-Wei Wang, Satoshi Kondo, Bogdan J. Matuszewski, Frédéric Precioso, Violet Snell, Josef Kittler, Teófilo Emídio de Campos, Adnan Mujahid Khan, Nasir M. Rajpoot, Evdokia Arkoumani, Miangela M. Lacle, Max A. Viergever, Josien P. W. Pluim |
Medical Image Anal. | 22 |
| 2015 | Full ranking as local descriptor for visual recognition: A comparison of distance metrics on sn
Chi-Ho Chan, Fei Yan 0001, Josef Kittler, Krystian Mikolajczyk |
Pattern Recognit. | 3 |
| 2015 | Random Cascaded-Regression Copse for Robust Facial Landmark DetectionabstractIn this letter, we present a random cascaded-regression copse (R-CR-C) for robust facial landmark detection. Its key innovations include a new parallel cascade structure design, and an adaptive scheme for scale-invariant shape update and local feature extraction. Evaluation on two challenging benchmarks shows the superiority of the proposed algorithm to state-of-the-art methods. Zhenhua Feng 0001, Patrik Huber 0001, Josef Kittler, William J. Christmas, Xiaojun Wu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2015 | Face Spoofing Detection Based on Multiple Descriptor Fusion Using Multiscale Dynamic Binarized Statistical Image FeaturesabstractFace recognition has been the focus of attention for the past couple of decades and, as a result, a significant progress has been made in this area. However, the problem of spoofing attacks can challenge face biometric systems in practical applications. In this paper, an effective countermeasure against face spoofing attacks based on a kernel discriminant analysis approach is presented. Its success derives from different innovations. First, it is shown that the recently proposed multiscale dynamic texture descriptor based on binarized statistical image features on three orthogonal planes (MBSIF-TOP) is effective in detecting spoofing attacks, showing promising performance compared with existing alternatives. Next, by combining MBSIF-TOP with a blur-tolerant descriptor, namely, the dynamic multiscale local phase quantization (MLPQ-TOP) representation, the robustness of the spoofing attack detector can be further improved. The fusion of the information provided by MBSIF-TOP and MLPQ-TOP is realized via a kernel fusion approach based on a fast kernel discriminant analysis (KDA) technique. It avoids the costly eigen-analysis computations by solving the KDA problem via spectral regression. The experimental evaluation of the proposed system on different databases demonstrates its advantages in detecting spoofing attacks in various imaging conditions, compared with the existing methods. Shervin Rahimzadeh Arashloo, Josef Kittler, William J. Christmas |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2015 | Cascaded Collaborative Regression for Robust Facial Landmark Detection Trained Using a Mixture of Synthetic and Real Images With Dynamic WeightingabstractA large amount of training data is usually crucial for successful supervised learning. However, the task of providing training samples is often time-consuming, involving a considerable amount of tedious manual work. In addition, the amount of training data available is often limited. As an alternative, in this paper, we discuss how best to augment the available data for the application of automatic facial landmark detection. We propose the use of a 3D morphable face model to generate synthesized faces for a regression-based detector training. Benefiting from the large synthetic training data, the learned detector is shown to exhibit a better capability to detect the landmarks of a face with pose variations. Furthermore, the synthesized training data set provides accurate and consistent landmarks automatically as compared to the landmarks annotated manually, especially for occluded facial parts. The synthetic data and real data are from different domains; hence the detector trained using only synthesized faces does not generalize well to real faces. To deal with this problem, we propose a cascaded collaborative regression algorithm, which generates a cascaded shape updater that has the ability to overcome the difficulties caused by pose variations, as well as achieving better accuracy when applied to real faces. The training is based on a mix of synthetic and real image data with the mixing controlled by a dynamic mixture weighting schedule. Initially, the training uses heavily the synthetic data, as this can model the gross variations between the various poses. As the training proceeds, progressively more of the natural images are incorporated, as these can model finer detail. To improve the performance of the proposed algorithm further, we designed a dynamic multi-scale local feature extraction method, which captures more informative local features for detector training. An extensive evaluation on both controlled and uncontrolled face data sets demonstrates the merit of the proposed algorithm. Zhenhua Feng 0001, Guosheng Hu, Josef Kittler, William J. Christmas, Xiaojun Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Audio Assisted Robust Visual Tracking With Adaptive Particle FilteringabstractThe problem of tracking multiple moving speakers in indoor environments has received much attention. Earlier techniques were based purely on a single modality, e.g., vision. Recently, the fusion of multi-modal information has been shown to be instrumental in improving tracking performance, as well as robustness in the case of challenging situations like occlusions (by the limited field of view of cameras or by other speakers). However, data fusion algorithms often suffer from noise corrupting the sensor measurements which cause non-negligible detection errors. Here, a novel approach to combining audio and visual data is proposed. We employ the direction of arrival angles of the audio sources to reshape the typical Gaussian noise distribution of particles in the propagation step and to weight the observation model in the measurement step. This approach is further improved by solving a typical problem associated with the PF, whose efficiency and accuracy usually depend on the number of particles and noise variance used in state estimation and particle propagation. Both parameters are specified beforehand and kept fixed in the regular PF implementation which makes the tracker unstable in practice. To address these problems, we design an algorithm which adapts both the number of particles and noise variance based on tracking error and the area occupied by the particles in the image. Experiments on the AV16.3 dataset show the advantage of our proposed methods over the baseline PF method and an existing adaptive PF algorithm for tracking occluded speakers with a significantly reduced number of particles. Volkan Kilic, Mark Barnard, Wenwu Wang 0001, Josef Kittler |
IEEE Trans. Multim. | 4 |
| 2014 | Transductive Transfer Machine
Nazli FarajiDavar, Teófilo Emídio de Campos, Josef Kittler |
ACCV (3) | 3 |
| 2014 | Adaptive Transductive Transfer Machine
Nazli FarajiDavar, Teófilo Emídio de Campos, Josef Kittler |
BMVC | 3 |
| 2014 | Audio-visual tracking of a variable number of speakers with a random finite set approach
Volkan Kilic, Xionghu Zhong, Mark Barnard, Wenwu Wang 0001, Josef Kittler |
FUSION | 5 |
| 2014 | Robust face recognition by an albedo based 3D morphable modelabstractLarge pose and illumination variations are very challenging for face recognition. The 3D Morphable Model (3DMM) approach is one of the effective methods for pose and illumination invariant face recognition. However, it is very difficult for the 3DMM to recover the illumination of the 2D input image because the ratio of the albedo and illumination contributions in a pixel intensity is ambiguous. Unlike the traditional idea of separating the albedo and illumination contributions using a 3DMM, we propose a novel Albedo Based 3D Morphable Model (AB3DMM), which removes the illumination component from the images using illumination normalisation in a preprocessing step. A comparative study of different illumination normalisation methods for this step is conducted on PIE and Multi-PIE databases. The results show that overall performance of our method outperforms state-of-the-art methods. Guosheng Hu, Chi-Ho Chan, Fei Yan 0001, William J. Christmas, Josef Kittler |
IJCB | 5 |
| 2014 | How many more images do we need? Performance prediction of bootstrapping for image classificationabstractMotivated by the recently introduced scalable concept detection challenge that requires classifiers for hundreds or even thousands of concepts, the objective of this work is to predict the cases where the enhancement of an initial classifier with additional training images is not expected to provide significant improvements. To facilitate this objective, we need a model for predicting the performance gain of a bootstrapping process prior to actually applying it. In order to train this model, we propose two features; the initial classifier's maturity (i.e. how close is the current hyperplane to the optimal) and the oracle's reliability (i.e. how reliable is the oracle in providing the correct labels of new training data). Thus, the contribution of our work is on proposing a method that is able to exploit the correlation between the expected performance boost and these two indicators. As a result, we can considerably improve the scalability properties of such bootstrapping processes by concentrating on the most prominent models and thus reducing the overall processing load. Elisavet Chatzilari, Spiros Nikolopoulos, Ioannis Kompatsiaris, Josef Kittler |
ICIP | 4 |
| 2014 | Classifier Incongruence Detection for Anomaly Flagging in Machine Perception
Josef Kittler |
ICPRAM | 1 |
| 2014 | On detection of novel categories and subcategories of images using incongruenceabstractNovelty detection is a crucial task in the development of autonomous vision systems. It aims at detecting if samples do not conform with the learnt models. In this paper, we consider the problem of detecting novelty in object recognition problems in which the set of object classes are grouped to form a semantic hierarchy. We follow the idea that, within a semantic hierarchy, novel samples can be defined as samples whose categorization at a specific level contrasts with the categorization at a more general level. This measure indicates if a sample is novel and, in that case, if it is likely to belong to a novel broad category or to a novel sub-category. We present an evaluation of this approach on two hierarchical subsets of the Caltech256 objects dataset and on the SUN scenes dataset, with different classification schemes. We obtain an improvement over Weinshall et al. and show that it is possible to bypass their normalisation heuristic. We demonstrate that this approach achieves good novelty detection rates as far as the conceptual taxonomy is congruent with the visual hierarchy, but tends to fail if this assumption is not satisfied. Dalia Coppi, Teófilo Emídio de Campos, Fei Yan 0001, Josef Kittler, Rita Cucchiara |
ICMR | 4 |
| 2014 | Automatic annotation of tennis games: An integration of audio, vision, and learning
Fei Yan 0001, Josef Kittler, David Windridge, William J. Christmas, Krystian Mikolajczyk, Stephen J. Cox, Qiang Huang 0006 |
Image Vis. Comput. | 2 |
| 2014 | Domain Anomaly Detection in Machine Perception: A System Architecture and TaxonomyabstractWe address the problem of anomaly detection in machine perception. The concept of domain anomaly is introduced as distinct from the conventional notion of anomaly used in the literature. We propose a unified framework for anomaly detection which exposes the multifaceted nature of anomalies and suggest effective mechanisms for identifying and distinguishing each facet as instruments for domain anomaly detection. The framework draws on the Bayesian probabilistic reasoning apparatus which clearly defines concepts such as outlier, noise, distribution drift, novelty detection (object, object primitive), rare events, and unexpected events. Based on these concepts we provide a taxonomy of domain anomaly events. One of the mechanisms helping to pinpoint the nature of anomaly is based on detecting incongruence between contextual and noncontextual sensor(y) data interpretation. The proposed methodology has wide applicability. It underpins in a unified way the anomaly detection applications found in the literature. To illustrate some of its distinguishing features, in here the domain anomaly detection methodology is applied to the problem of anomaly detection for a video annotation system. Josef Kittler, William J. Christmas, Teófilo Emídio de Campos, David Windridge, Fei Yan 0001, John Illingworth, Magda Osman |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Corrigendum to "A user-specific and selective multimodal biometric fusion strategy by ranking subjects" [Pattern Recognition 46 (2013) 3341-3357]
Norman Poh, Arun Ross, Weifeng Li 0001, Josef Kittler |
Pattern Recognit. | 4 |
| 2014 | HEp-2 fluorescence pattern classification
Violet Snell, William J. Christmas, Josef Kittler |
Pattern Recognit. | 3 |
| 2014 | Fast pose invariant face recognition using super coupled multiresolution Markov Random Fields on a GPU
Shervin Rahimzadeh Arashloo, Josef Kittler |
Pattern Recognit. Lett. | 2 |
| 2014 | Celebrating the life and work of Maria Petrou
Edwin R. Hancock, Josef Kittler |
Pattern Recognit. Lett. | 2 |
| 2014 | Multilevel Chinese Takeaway Process and Label-Based Processes for Rule Induction in the Context of Automated Sports Video AnnotationabstractWe propose four variants of a novel hierarchical hidden Markov models strategy for rule induction in the context of automated sports video annotation including a multilevel Chinese takeaway process (MLCTP) based on the Chinese restaurant process and a novel Cartesian product label-based hierarchical bottom-up clustering (CLHBC) method that employs prior information contained within label structures. Our results show significant improvement by comparison against the flat Markov model: optimal performance is obtained using a hybrid method, which combines the MLCTP generated hierarchical topological structures with CLHBC generated event labels. We also show that the methods proposed are generalizable to other rule-based environments including human driving behavior and human actions. Aftab Khan 0001, David Windridge, Josef Kittler |
IEEE Trans. Cybern. | 3 |
| 2014 | Class-Specific Kernel Fusion of Multiple Descriptors for Face Verification Using Multiscale Binarised Statistical Image FeaturesabstractThis paper addresses face verification in unconstrained settings. For this purpose, first, a nonlinear binary class-specific kernel discriminant analysis classifier (CS-KDA) based on spectral regression kernel discriminant analysis is proposed. By virtue of the two-class formulation, the proposed CS-KDA approach offers a number of desirable properties such as specificity of the transformation for each subject, computational efficiency, simplicity of training, isolation of the enrolment of each client from others and increased speed in probe testing. Using the proposed CS-KDA approach, a regional discriminative face image representation based on a multiscale variant of the binarized statistical image features is proposed next. The proposed component-based representation when coupled with the dense pixel-wise alignments provided by a symmetric MRF matching model reduces the sensitivity to misalignments and pose variations, gauging the similarity more effectively. Finally, the discriminative representation is combined with two other effective image descriptors, namely the multiscale local binary patterns and the multiscale local phase quantization histograms via a kernel fusion approach to further enhance system accuracy. The experimental evaluation of the proposed methodology on challenging databases demonstrates its advantage over other methods. Shervin Rahimzadeh Arashloo, Josef Kittler |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Dynamic Texture Recognition Using Multiscale Binarized Statistical Image FeaturesabstractA spatio-temporal descriptor for representation and recognition of time-varying textures is proposed [binarized statistical image features on three orthogonal planes (BSIF-TOP)] in this paper. The descriptor, similar in spirit to the well known local binary patterns on three orthogonal planes approach, estimates histograms of binary coded image sequences on three orthogonal planes corresponding to spatial/spatio-temporal dimensions. However, unlike some other methods which generate the code in a heuristic fashion, binary code generation in the BSIF-TOP approach is realized by filtering operations on different regions of spatial/spatio-temporal support and by binarizing the filter responses. The filters are learnt via independent component analysis on each of three planes after preprocessing using a whitening transformation. By extending the BSIF-TOP descriptor to a multiresolution scheme, the descriptor is able to capture the spatio-temporal content of an image sequence at multiple scales, improving its representation capacity. In the evaluations on the UCLA, Dyntex, and Dyntex++ dynamic texture databases, the proposed method achieves very good performance compared to existing approaches. Shervin Rahimzadeh Arashloo, Josef Kittler |
IEEE Trans. Multim. | 2 |
| 2014 | Robust Multi-Speaker Tracking via Dictionary Learning and Identity ModelingabstractWe investigate the problem of visual tracking of multiple human speakers in an office environment. In particular, we propose novel solutions to the following challenges: (1) robust and computationally efficient modeling and classification of the changing appearance of the speakers in a variety of different lighting conditions and camera resolutions; (2) dealing with full or partial occlusions when multiple speakers cross or come into very close proximity; (3) automatic initialization of the trackers, or re-initialization when the trackers have lost lock caused by e.g. the limited camera views. First, we develop new algorithms for appearance modeling of the moving speakers based on dictionary learning (DL), using an off-line training process. In the tracking phase, the histograms (coding coefficients) of the image patches derived from the learned dictionaries are used to generate the likelihood functions based on Support Vector Machine (SVM) classification. This likelihood function is then used in the measurement step of the classical particle filtering (PF) algorithm. To improve the computational efficiency of generating the histograms, a soft voting technique based on approximate Locality-constrained Soft Assignment (LcSA) is proposed to reduce the number of dictionary atoms (codewords) used for histogram encoding. Second, an adaptive identity model is proposed to track multiple speakers whilst dealing with occlusions. This model is updated online using Maximum a Posteriori (MAP) adaptation, where we control the adaptation rate using the spatial relationship between the subjects. Third, to enable automatic initialization of the visual trackers, we exploit audio information, the Direction of Arrival (DOA) angle, derived from microphone array recordings. Such information provides, a priori, the number of speakers and constrains the search space for the speaker's faces. The proposed system is tested on a number of sequences from three publicly available and challenging data corpora (AV16.3, EPFL pedestrian data set and CLEAR) with up to five moving subjects. Mark Barnard, Piotr Koniusz, Wenwu Wang 0001, Josef Kittler, Syed M. Naqvi, Jonathon A. Chambers |
IEEE Trans. Multim. | 4 |
| 2013 | Audio-visual face detection for tracking in a meeting room environment
Mark Barnard, Wenwu Wang 0001, Josef Kittler, Syed M. Naqvi, Jonathon A. Chambers |
FUSION | 3 |
| 2013 | Audio head pose estimation using the direct to reverberant speech ratioabstractHead pose is an important cue in many applications such as, speech recognition and face recognition. Most approaches to head pose estimation to date have used visual information to model and recognise a subject's head in different configurations. These approaches have a number of limitations such as, inability to cope with occlusions, changes in the appearance of the head, and low resolution images. We present here a novel method for determining coarse head pose orientation purely from audio information, exploiting the direct to reverberant speech energy ratio (DRR) within a highly reverberant meeting room environment. Our hypothesis is that a speaker facing towards a microphone will have a higher DRR and a speaker facing away from the microphone will have a lower DRR. This hypothesis is confirmed by experiments conducted on the publicly available AV16.3 database. Mark Barnard, Wenwu Wang 0001, Josef Kittler |
ICASSP | 3 |
| 2013 | Audio constrained particle filter based visual trackingabstractWe present a robust and efficient audio-visual (AV) approach to speaker tracking in a room environment. A challenging problem with visual tracking is to deal with occlusions (caused by the limited field of view of cameras or by other speakers). Another challenge is associated with the particle filtering (PF) algorithm, commonly used for visual tracking, which requires a large number of particles to ensure the distribution is well modelled. In this paper, we propose a new method of fusing audio into the PF based visual tracking. We use the direction of arrival angles (DOAs) of the audio sources to reshape the typical Gaussian noise distribution of particles in the propagation step and to weight the observation model in the measurement step. Experiments on AV16.3 datasets show the advantage of our proposed method over the baseline PF method for tracking occluded speakers with a significantly reduced number of particles. Volkan Kilic, Mark Barnard, Wenwu Wang 0001, Josef Kittler |
ICASSP | 4 |
| 2013 | Photometric Normalization for Face Recognition using Local discrete cosine TransformabstractVariations in illumination is one of major limiting factors of face recognition system performance. The effect of changes in the incident light on face images is analyzed, as well as its influence on the low frequency components of the image. Starting from this analysis, a new photometric normalization method for illumination invariant face recognition is presented. Low-frequency Discrete Cosine Transform coefficients in the logarithmic domain are used in a local way to reconstruct a slowly varying component of the face image which is caused by illumination. After smoothing, this component is subtracted from the original logarithmic image to compensate for illumination variations. Compared to other preprocessing algorithms, our method achieved a very good performance with a total error rate very similar to that produced by the best performing state-of-the-art algorithm. An in-depth analysis of the two preprocessing methods revealed notable differences in their behavior, which is exploited in a multiple classifier fusion framework to achieve further performance improvement. The superiority of the proposal is demonstrated in both face verification and identification experiments. Heydi Mendez Vazquez, Josef Kittler, Chi-Ho Chan, Edel B. García Reyes |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2013 | Multiscale Local Phase Quantization for Robust Component-Based Face Recognition Using Kernel Fusion of Multiple DescriptorsabstractFace recognition subject to uncontrolled illumination and blur is challenging. Interestingly, image degradation caused by blurring, often present in real-world imagery, has mostly been overlooked by the face recognition community. Such degradation corrupts face information and affects image alignment, which together negatively impact recognition accuracy. We propose a number of countermeasures designed to achieve system robustness to blurring. First, we propose a novel blur-robust face image descriptor based on Local Phase Quantization (LPQ) and extend it to a multiscale framework (MLPQ) to increase its effectiveness. To maximize the insensitivity to misalignment, the MLPQ descriptor is computed regionally by adopting a component-based framework. Second, the regional features are combined using kernel fusion. Third, the proposed MLPQ representation is combined with the Multiscale Local Binary Pattern (MLBP) descriptor using kernel fusion to increase insensitivity to illumination. Kernel Discriminant Analysis (KDA) of the combined features extracts discriminative information for face recognition. Last, two geometric normalizations are used to generate and combine multiple scores from different face image scales to further enhance the accuracy. The proposed approach has been comprehensively evaluated using the combined Yale and Extended Yale database B (degraded by artificially induced linear motion blur) as well as the FERET, FRGC 2.0, and LFW databases. The combined system is comparable to state-of-the-art approaches using similar system configurations. The reported work provides a new insight into the merits of various face representation and fusion methods, as well as their role in dealing with variable lighting and blur degradation. Chi-Ho Chan, Muhammad Atif Tahir, Josef Kittler, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | A user-specific and selective multimodal biometric fusion strategy by ranking subjects
Norman Poh, Arun Ross, Weifeng Lee, Josef Kittler |
Pattern Recognit. | 4 |
| 2013 | A Robust and Scalable Visual Category and Action Recognition System Using Kernel Discriminant Analysis With Spectral RegressionabstractVisual concept detection and action recognition are one of the most important tasks in content-based multimedia information retrieval (CBMIR) technology. It aims at annotating images using a vocabulary defined by a set of concepts of interest including scenes types (mountains, snow, etc.) or human actions (phoning, playing instrument). This paper describes our system in the ImageCLEF@ICPR10, Pascal VOC 08 Visual Concept Detection and Pascal VOC 10 Action Recognition Challenges. The proposed system ranked first in these large-scale tasks when evaluated independently by the organizers. The proposed system involves state-of-the-art local descriptor computation, vector quantization via clustering, structured scene or object representation via localized histograms of vector codes, similarity measure for kernel construction and classifier learning. The main novelty is the classifier-level and kernel-level fusion using Kernel Discriminant Analysis and Spectral Regression (SR-KDA) with RBF Chi-Squared kernels obtained from various image descriptors. The distinctiveness of the proposed method is also assessed experimentally using a video benchmark: the Mediamill Challenge along with benchmarks from ImageCLEF@ICPR10, Pascal VOC 10 and Pascal VOC 08. From the experimental results, it can be derived that the presented system consistently yields significant performance gains when compared with the state-of-the art methods. The other strong point is the introduction of SR-KDA in the classification stage where the time complexity scales linearly with respect to the number of concepts and the main computational complexity is independent of the number of categories. Muhammad Atif Tahir, Fei Yan 0001, Piotr Koniusz, Muhammad Awais 0001, Mark Barnard, Krystian Mikolajczyk, Ahmed Bouridane, Josef Kittler |
IEEE Trans. Multim. | 8 |
| 2012 | Resolution-Aware 3D Morphable ModelabstractThe 3D Morphable Model (3DMM) is currently receiving considerable attention for \nhuman face analysis. Most existing work focuses on fitting a 3DMM to high resolution \nimages. However, in many applications, fitting a 3DMM to low-resolution images \nis also important. In this paper, we propose a Resolution-Aware 3DMM (RA- \n3DMM), which consists of 3 different resolution 3DMMs: High-Resolution 3DMM \n(HR- 3DMM), Medium-Resolution 3DMM (MR-3DMM) and Low-Resolution 3DMM \n(LR-3DMM). RA-3DMM can automatically select the best model to fit the input images \nof different resolutions. The multi-resolution model was evaluated in experiments \nconducted on PIE and XM2VTS databases. The experimental results verified that HR- \n3DMM achieves the best performance for input image of high resolution, and MR- \n3DMM and LR-3DMM worked best for medium and low resolution input images, respectively. \nA model selection strategy incorporated in the RA-3DMM is proposed based \non these results. The RA-3DMM model has been applied to pose correction of face images \nranging from high to low resolution. The face verification results obtained with \nthe pose-corrected images show considerable performance improvement over the result \nwithout pose correction in all resolutions Guosheng Hu, Chi-Ho Chan, Josef Kittler, William J. Christmas |
BMVC | 3 |
| 2012 | A dictionary learning approach to trackingabstractThe problem of tracking people using multiple cameras is of much current interest as a means of providing cues for audio-visual blind source separation in dynamic environments. Here we investigate the use of one of the current state-of-the-art techniques in object recognition combined with one of the most popular methods of modelling object motion, particle filters, for tracking people. The dictionary learning or Bag-of-Words approach to object recognition has proved to be very effective in recent years, as shown in a number of large comparisons such as the PASCAL Visual Object recognition Challenge (VOC). In this paper we use this proven object recognition method within the framework of a particle filter. This provides a more accurate and robust tracking of people in a multiple camera environment. We also demonstrate that the dictionary learning approach can provide a principled method for the fusion of multiple features. Mark Barnard, Wenwu Wang 0001, Josef Kittler, Syed M. Naqvi, Jonathon A. Chambers |
ICASSP | 3 |
| 2012 | Blur kernel estimation to improve recognition of blurred facesabstractThis paper proposes an efficient blind deconvolution method to deblur face images for face recognition. The method involves a salient edge map construction, blur kernel estimation and face image deconvolution. The combined Yale and Extended Yale face database B containing different illumination changes and blur conditions are used to evaluated the face identification system. The results show that the accuracy of the face recognition systems implemented with the proposed method improves the accuracy when the faces are degraded by blur in general and motion blur in particular. Chi-Ho Chan, Josef Kittler |
ICIP | 2 |
| 2012 | Automatic face annotation by multilinear AAM with Missing Values
Zhenhua Feng 0001, Josef Kittler, William J. Christmas, Xiaojun Wu 0001, Sebastian Pfeiffer |
ICPR | 2 |
| 2012 | An intrinsic coordinate system for 3D face registration
Willem P. Koppen, Chi-Ho Chan, William J. Christmas, Josef Kittler |
ICPR | 4 |
| 2012 | A discriminative parametric approach to video-based score-level fusion for biometric authentication
Norman Poh, Josef Kittler, Fuad M. Alkoot |
ICPR | 2 |
| 2012 | Texture and shape in fluorescence pattern identification for auto-immune disease diagnosis
Violet Snell, William J. Christmas, Josef Kittler |
ICPR | 3 |
| 2012 | Automatic annotation of court games with structured output learning
Fei Yan 0001, Josef Kittler, Krystian Mikolajczyk, David Windridge |
ICPR | 2 |
| 2012 | Multi-modal region selection approach for training object detectorsabstractOur purpose in this work is to boost the performance of object classifiers learned using the self-training paradigm. We exploit the multi-modal nature of tagged images found in social networks, to optimize the process of region selection when retraining the initial model. More specifically, the proposed approach uses a small number of manually labelled regions to train the initial object detection classifiers. Then, a large number of loosely tagged images, pre-segmented by an automatic segmentation algorithm, is used to enhance the initial training set with additional image regions. However, in contrast to the typical case of self-training where the image regions are selected based solely on how well they fit to the original classification model, our approach aims at optimizing this selection by making combined use of both visual and textual information. The experimental results show that the object detection classifiers generated using the proposed approach outperform the classifiers generated using the typical self-training paradigm. Elisavet Chatzilari, Spiros Nikolopoulos, Ioannis Kompatsiaris, Josef Kittler |
ICMR | 4 |
| 2012 | Non-Sparse Multiple Kernel Fisher Discriminant Analysis
Fei Yan 0001, Josef Kittler, Krystian Mikolajczyk, Muhammad Atif Tahir |
J. Mach. Learn. Res. | 2 |
| 2012 | A Unified Framework for Biometric Expert Fusion Incorporating Quality MeasuresabstractThis paper proposes a unified framework for quality-based fusion of multimodal biometrics. Quality-dependent fusion algorithms aim to dynamically combine several classifier (biometric expert) outputs as a function of automatically derived (biometric) sample quality. Quality measures used for this purpose quantify the degree of conformance of biometric samples to some predefined criteria known to influence the system performance. Designing a fusion classifier to take quality into consideration is difficult because quality measures cannot be used to distinguish genuine users from impostors, i.e., they are nondiscriminative yet still useful for classification. We propose a general Bayesian framework that can utilize the quality information effectively. We show that this framework encompasses several recently proposed quality-based fusion algorithms in the literature--Nandakumar et al., 2006; Poh et al., 2007; Kryszczuk and Drygajo, 2007; Kittler et al., 2007; Alonso-Fernandez, 2008; Maurer and Baker, 2007; Poh et al., 2010. Furthermore, thanks to the systematic study concluded herein, we also develop two alternative formulations of the problem, leading to more efficient implementation (with fewer parameters) and achieving performance comparable to, or better than, the state of the art. Last but not least, the framework also improves the understanding of the role of quality in multiple classifier combination. Norman Poh, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Inverse random under sampling for class imbalance problem and its application to multi-label classification
Muhammad Atif Tahir, Josef Kittler, Fei Yan 0001 |
Pattern Recognit. | 2 |
| 2012 | Multilabel classification using heterogeneous ensemble of multi-label classifiers
Muhammad Atif Tahir, Josef Kittler, Ahmed Bouridane |
Pattern Recognit. Lett. | 2 |
| 2012 | Differential Edit Distance: A Metric for Scene Segmentation EvaluationabstractIn this paper, a novel approach to evaluating video temporal decomposition algorithms is presented. The evaluation measures typically used to this end are nonlinear combinations of precision-recall or coverage-overflow, which are not metrics and additionally possess undesirable properties, such as nonsymmetricity. To alleviate these drawbacks, we introduce a novel unidimensional measure that is proven to be metric and satisfies a number of qualitative prerequisites that previous measures do not. This measure is named differential edit distance (DED), since it can be seen as a variation of the well-known edit distance. After defining DED, we further introduce an algorithm that computes it in less than cubic time. Finally, DED is extensively compared with state-of-the-art measures, namely, the harmonic means (F-score) of precision-recall and coverage-overflow. The experiments include comparisons of qualitative properties, the time required for optimizing the parameters of scene segmentation algorithms with the help of these measures, and a user study gauging the agreement of these measures with the users' assessment of the segmentation results. The results confirm that the proposed measure is a unidimensional metric that is effective in evaluating scene segmentation techniques and in helping to optimize their parameters. Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | Local Ordinal Contrast Pattern Histograms for Spatiotemporal, Lip-Based Speaker AuthenticationabstractLip region deformation during speech contains biometric information and is termed visual speech. This biometric information can be interpreted as being genetic or behavioral depending on whether static or dynamic features are extracted. In this paper, we use a texture descriptor called local ordinal contrast pattern (LOCP) with a dynamic texture representation called three orthogonal planes to represent both the appearance and dynamics features observed in visual speech. This feature representation, when used in standard speaker verification engines, is shown to improve the performance of the lip-biometric trait compared to the state-of-the-art. The best baseline state-of-the-art performance was a half total error rate (HTER) of 13.35% for the XM2VTS database. We obtained HTER of less than 1%. The resilience of the LOCP texture descriptor to random image noise is also investigated. Finally, the effect of the amount of video information on speaker verification performance suggests that with the proposed approach, speaker identity can be verified with a much shorter biometric trait record than the length normally required for voice-based biometrics. In summary, the performance obtained is remarkable and suggests that there is enough discriminative information in the mouth-region to enable its use as a primary biometric trait. Chi-Ho Chan, Budhaditya Goswami, Josef Kittler, William J. Christmas |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2012 | User-Specific Cohort Selection and Score Normalization for Biometric SystemsabstractAn increasing body of evidence suggests that cohort-based score normalization can improve the performance of biometric authentication. This approach relies on the use of N cohort biometric templates, which can be computationally expensive. We contribute to the advancement of cohort score normalization in two ways. First, we show both theoretically and empirically that the most similar and the most dissimilar cohort templates to a target user contain discriminative information. We then investigate the extraction of this information using polynomial regression. Extensive evaluation on the face and fingerprint modalities in the Biosecure DS2 dataset indicates that the proposed method outperforms the state-of-the-art cohort score normalization methods, while reducing the computation cost by as much as half. Amin Merati, Norman Poh, Josef Kittler |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2011 | Augmented Kernel Matrix vs Classifier Fusion for Object RecognitionabstractAugmented Kernel Matrix (AKM) has recently been proposed to accommodate for the fact that a single training example may have different importance in different feature spaces, in contrast to Multiple Kernel Learning (MKL) that assigns the same weight to all examples in one feature space.However, the AKM approach is limited to small datasets due to its memory requirements.An alternative way to fuse information from different feature channels is classifier fusion (ensemble methods).There is a significant amount of work on linear programming formulations of classifier fusion (CF) in the case of binary classification.In this paper we derive primal and dual of AKM to draw its correspondence with CF.We propose a multiclass extension of binary ν-LPBoost, which learns the contribution of each class in each feature channel.Existing approaches of CF promote sparse features combinations, due to regularization based on 1 -norm, and lead to a selection of a subset of feature channels, which is not good in case of informative channels.We also generalize existing CF formulations to arbitrary p -norm for binary and multiclass problems which results in more effective use of complementary information.We carry out an extensive comparison and show that the proposed nonlinear CF schemes outperform its sparse counterpart as well as state-of-the-art MKL approaches. Muhammad Awais 0001, Fei Yan 0001, Krystian Mikolajczyk, Josef Kittler |
BMVC | 4 |
| 2011 | Speaker authentication using video-based lip informationabstractThe lip-region can be interpreted as either a genetic or behavioural biometric trait depending on whether static or dynamic information is used. In this paper, we use a texture descriptor called Local Ordinal Contrast Pattern (LOCP) in conjunction with a novel spatiotemporal sampling method called Windowed Three Orthogonal Planes (WTOP) to represent both appearance and dynamics features ob served in visual speech. This representation, with standard speaker verification engines, is shown to improve the performance of the lip biometric trait compared to the state-of-the-art. The improvement obtained suggests that there is enough discriminative information in the mouth-region to enable its use as a primary biometric as opposed to a "soft" biometric trait. Budhaditya Goswami, Chi-Ho Chan, Josef Kittler, William J. Christmas |
ICASSP | 3 |
| 2011 | Heterogeneous information fusion: A novel fusion paradigm for biometric systemsabstractOne of the most promising ways to improve biometric person recognition is indisputably via information fusion, that is, to combine different sources of information. This pa per proposes a novel fusion paradigm that combines heterogeneous sources of information such as user-specific, cohort and quality information. Two formulations of this problem are proposed, differing in the assumption on the independence of the information sources. Unlike the more common multimodal/multi-algorithmic fusion, the novel paradigm has to deal with information that is not necessarily discriminative but still it is relevant. The methodology can be applied to any biometric system. Furthermore, extensive experiments based on 30 face and fingerprint experiments indicate that the performance gain with respect to the baseline system is about 30%. In contrast, solving this problem using conventional fusion paradigm leads to degraded results. Norman Poh, Amin Merati, Josef Kittler |
IJCB | 3 |
| 2011 | Face recognition using multi-scale local phase quantisation and Linear Regression ClassifierabstractLinear Regression Classifier (LRC) is state-of-the-art face recognition method that represent a probe image as a linear combination of class specific models. However, this method views the image as a point in a feature space, and thus LRC cannot accommodate severe luminance alterations. Histogram-based features, such as Multiscale Local Phase Quantisation histogram (MLPQH) have gained reputation as powerful and attractive texture descriptors showing excellent results in terms of accuracy and computational complexity in face recognition. In this paper, MLPQH features are integrated with "face" features to confront the illumination problem in LRC. The main novelty is the fusion of histogram and face features using z-score normalisation and LRC classifier. The proposed system is evaluated on two benchmarks: ORL and Extended Yale B. The results indicate a significant increase in the performance when compared with state-of the-art face recognition methods. Muhammad Atif Tahir, Chi-Ho Chan, Josef Kittler, Ahmed Bouridane |
ICIP | 3 |
| 2011 | Novel Fusion Methods for Pattern Recognition
Muhammad Awais 0001, Fei Yan 0001, Krystian Mikolajczyk, Josef Kittler |
ECML/PKDD (1) | 4 |
| 2011 | An evaluation of bags-of-words and spatio-temporal shapes for action recognitionabstractBags-of-visual-Words (BoW) and Spatio-Temporal Shapes (STS) are two very popular approaches for action recognition from video. The former (BoW) is an un-structured global representation of videos which is built using a large set of local features. The latter (STS) uses a single feature located on a region of interest (where the actor is) in the video. Despite the popularity of these methods, no comparison between them has been done. Also, given that BoW and STS differ intrinsically in terms of context inclusion and globality/locality of operation, an appropriate evaluation framework has to be designed carefully. This paper compares these two approaches using four different datasets with varied degree of space-time specificity of the actions and varied relevance of the contextual background. We use the same local feature extraction method and the same classifier for both approaches. Further to BoW and STS, we also evaluated novel variations of BoW constrained in time or space. We observe that the STS approach leads to better results in all datasets whose background is of little relevance to action classification. Teófilo Emídio de Campos, Mark Barnard, Krystian Mikolajczyk, Josef Kittler, Fei Yan 0001, William J. Christmas, David Windridge |
WACV | 4 |
| 2011 | Pose-invariant face recognition by matching on multi-resolution MRFs linked by supercoupling transform
Shervin Rahimzadeh Arashloo, Josef Kittler, William J. Christmas |
Comput. Vis. Image Underst. | 2 |
| 2011 | Incremental Linear Discriminant Analysis Using Sufficient Spanning Sets and Its Applications
Tae-Kyun Kim 0001, Björn Stenger, Josef Kittler, Roberto Cipolla |
Int. J. Comput. Vis. | 3 |
| 2011 | Energy Normalization for Pose-Invariant Face Recognition Based on MRF Model Image MatchingabstractA pose-invariant face recognition system based on an image matching method formulated on MRFs is presented. The method uses the energy of the established match between a pair of images as a measure of goodness-of-match. The method can tolerate moderate global spatial transformations between the gallery and the test images and alleviate the need for geometric preprocessing of facial images by encapsulating a registration step as part of the system. It requires no training on non-frontal face images. A number of innovations, such as a dynamic block size and block shape adaptation, as well as label pruning and error pre-whitening measures have been introduced to increase the effectiveness of the approach. The experimental evaluation of the method is performed on two publicly available databases. First, the method is tested on the rotation shots of the XM2VTS data set in a verification scenario. Next, the evaluation is conducted in an identification scenario on the CMU-PIE database. The method compares favorably with the existing 2D or 3D generative model-based methods on both databases in both identification and verification scenarios. Shervin Rahimzadeh Arashloo, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | On Combining Local DCT with Preprocessing Sequence for Face Recognition under Varying Lighting Conditions
Heydi Mendez Vazquez, Josef Kittler, Chi-Ho Chan, Edel B. García Reyes |
CIARP | 2 |
| 2010 | lp norm multiple kernel Fisher discriminant analysis for object and image categorisationabstractIn this paper, we generalise multiple kernel Fisher discriminant analysis (MK-FDA) such that the kernel weights can be regularised with an ℓpnorm for any p ≥ 1, in contrast to existing MK-FDA that uses either l1 or l2 norm. We present formulations for both binary and multiclass cases and solve the associated optimisation problems efficiently with semi-infinite programming. We show on three object and image categorisation benchmarks that by learning the intrinsic sparsity of a given set of base kernels using a validation set, the proposed ℓpMK-FDA outperforms its fixed-norm counterparts, and is capable of producing state-of-the-art performance. Moreover, we show that our ℓpMK-FDA outperforms the ℓpmultiple kernel support vector machine (ℓpMK-SVM) which has been recently proposed. Based on this observation and our experience with single kernel FDA and SVM, we argue that the almost century-old FDA is still a strong competitor of the popular SVM. Fei Yan 0001, Krystian Mikolajczyk, Mark Barnard, Hongping Cai, Josef Kittler |
CVPR | 5 |
| 2010 | Ball event recognition using hmm for automatic tennis annotationabstractA key prerequisite of automatic video indexing and summarisation is the description of events and actions. In the context of many sports, the motion of the ball and agents plays an essential role in describing events. However, the only existing solution for the tennis event recognition problem in the literature is the work in which relies on a set of heuristic rules such as proximity between ball and players or court lines to classify ball event candidates. We present hidden Markov models (HMMs) paradigm to automatically learn to identify events from ball trajectories and demonstrate that its ability to capture the dynamics of the ball movement lead to a much higher performance. Ibrahim Almajai, Josef Kittler, Teófilo Emídio de Campos, William J. Christmas, Fei Yan 0001, David Windridge, Aftab Khan 0001 |
ICIP | 2 |
| 2010 | Sparse representation of (Multiscale) histograms for face recognition robust to registration and illumination problemsabstractWe combine sparse representation with a multiresolution histogram face descriptor to create a powerful representation method for face recognition. The multi resolution histogram descriptor is based on local binary patterns or local phase coding to achieve invariance to various types of image degradation phenomena. By its nature, the histogram descriptor is also robust to geometric misalignment of the gallery and query faces. The proposed face recognition method is evaluated on Yale Face Database B and the extended Yale Face Database B, yielding very impressive results. Chi-Ho Chan, Josef Kittler |
ICIP | 2 |
| 2010 | Fusion of visible and synthesised near infrared information for face authenticationabstractChanges in illumination conditions can cause drastic variations in face appearance and affect the performance of a face authentication system. Near infrared (NIR) face imaging systems have been proposed as a promising way towards illumination invariant face verification. We show that when NIR face images cannot be observed, learning the relationship between NIR information and the corresponding visible images can provide useful complementary information about visible light image data. In particular, we use Canonical Correlation Analysis (CCA) to synthesise the NIR eigenfaces from their corresponding visible ones. In this paper, the verification performance of a CCA-based synthesising algorithm is developed first. Although, synthesised NIR images do not perform as well as the real NIR, it is shown that by fusing the visible and synthesised near infrared information at the score level, the performance of the authentication system considerably improves. Seyed Mohammad Mavadati, Mohammad Sadeghi 0001, Josef Kittler |
ICIP | 3 |
| 2010 | Lattice-Based Anomaly Rectification for Sport Video AnnotationabstractAnomaly detection has received much attention within the literature as a means of determining, in an unsupervised manner, whether a learning domain has changed in a fundamental way. This may require continuous adaptive learning to be abandoned and a new learning process initiated in the new domain. A related problem is that of anomaly rectification; the adaptation of the existing learning mechanism to the change of domain. As a concrete instantiation of this notion, the current paper investigates a novel lattice-based HMM induction strategy for arbitrary court-game environments. We test (in real and simulated domains) the ability of the method to adapt to a change of rule structures going from tennis singles to tennis doubles. Our long term aim is to build a generic system for transferring game-rule inferences. Aftab Khan 0001, David Windridge, Teófilo Emídio de Campos, Josef Kittler, William J. Christmas |
ICPR | 4 |
| 2010 | Model and Score Adaptation for Biometric Systems: Coping With Device Interoperability and Changing Acquisition ConditionsabstractThe performance of biometric systems can be significantly affected by changes in signal quality. In this paper, two types of changes are considered: change in acquisition environment and in sensing devices. We investigated three solutions: (i) model-level adaptation, (ii) score-level adaptation (normalisation), and (iii) the combination of the two, called “compound” adaptation. In order to cope with the above changing conditions, the model-level adaptation attempts to update the parameters of the expert systems (classifiers). This approach requires the authenticity of the candidate samples used for adaptation be known (corresponding to supervised adaptation), or can be estimated (unsupervised adaptation). In comparison, the score-level adaptation merely involves post processing the expert output, with the objective of rendering the associated decision threshold to be dependent only on the class priors despite the changing acquisition conditions. Since the above adaptation strategies treat the underlying biometric experts/classifiers as a black-box, they can be applied to any unimodal or multimodal biometric system, thus facilitating system-level integration and performance optimisation. Our contributions are: (i) proposal of compound adaptation; (ii) investigation and comparison of two different quality-dependent score normalisation strategies; and, (iii) empirical comparison of the merit of the above three solutions on the BANCA face (video) and speech database. Norman Poh, Josef Kittler, Sébastien Marcel, Driss Matrouf, Jean-François Bonastre |
ICPR | 2 |
| 2010 | The University of Surrey Visual Concept Detection System at ImageCLEF@ICPR: Working NotesabstractVisual concept detection is one of the most important tasks in image and video indexing. This paper describes our system in the ImageCLEF@ICPR Visual Concept Detection Task which ranked first for large-scale visual concept detection tasks in terms of Equal Error Rate (EER) and Area under Curve (AUC) and ranked third in terms of hierarchical measure. The presented approach involves state-of-the-art local descriptor computation, vector quantisation via clustering, structured scene or object representation via localised histograms of vector codes, similarity measure for kernel construction and classifier learning. The main novelty is the classifier-level and kernel-level fusion using Kernel Discriminant Analysis with RBF/Power Chi-Squared kernels obtained from various image descriptors. For 32 out of 53 individual concepts, we obtain the best performance of all 12 submissions to this task. Muhammad Atif Tahir, Fei Yan 0001, Mark Barnard, Muhammad Awais 0001, Krystian Mikolajczyk, Josef Kittler |
ICPR | 6 |
| 2010 | On design and optimization of face verification systems that are smart-card based
Thirimachos Bourlai, Josef Kittler, Kieron Messer |
Mach. Vis. Appl. | 2 |
| 2010 | The Multiscenario Multienvironment BioSecure Multimodal Database (BMDB)abstractA new multimodal biometric database designed and acquired within the framework of the European BioSecure Network of Excellence is presented. It is comprised of more than 600 individuals acquired simultaneously in three scenarios: 1) over the Internet, 2) in an office environment with desktop PC, and 3) in indoor/outdoor environments with mobile portable hardware. The three scenarios include a common part of audio/video data. Also, signature and fingerprint data have been acquired both with desktop PC and mobile portable hardware. Additionally, hand and iris data were acquired in the second scenario using desktop PC. Acquisition has been conducted by 11 European institutions. Additional features of the BioSecure Multimodal Database (BMDB) are: two acquisition sessions, several sensors in certain modalities, balanced gender and age distributions, multimodal realistic scenarios with simple and quick tasks per modality, cross-European diversity, availability of demographic data, and compatibility with other multimodal databases. The novel acquisition conditions of the BMDB allow us to perform new challenging research and evaluation of either monomodal or multimodal biometric systems, as in the recent BioSecure Multimodal Evaluation campaign. A description of this campaign including baseline results of individual modalities from the new database is also given. The database is expected to be available for research purposes through the BioSecure Association during 2008. Javier Ortega-Garcia, Julian Fierrez, Fernando Alonso-Fernandez, Javier Galbally, Manuel R. Freire, Joaquín González-Rodríguez, Carmen García-Mateo, José Luis Alba-Castro, Elisardo González-Agulla, Enrique Otero Muras, Sonia Garcia-Salicetti, Lorène Allano, Van-Bao Ly, Bernadette Dorizzi, Josef Kittler, Thirimachos Bourlai, Norman Poh, Farzin Deravi, Ming W. R. Ng, Michael C. Fairhurst, Jean Hennebert, Andreas Humm, Massimo Tistarelli, Linda Brodo, Jonas Richiardi, Andrzej Drygajlo, Harald Ganster, Federico Sukno, Sri-Kaushik Pavani, Alejandro F. Frangi, Lale Akarun, Arman Savran |
IEEE Trans. Pattern Anal. Mach. Intell. | 15 |
| 2010 | A multimodal biometric test bed for quality-dependent, cost-sensitive and client-specific score-level fusion algorithms
Norman Poh, Thirimachos Bourlai, Josef Kittler |
Pattern Recognit. | 3 |
| 2010 | An Evaluation of Video-to-Video Face VerificationabstractPerson recognition using facial features, e.g., mug-shot images, has long been used in identity documents. However, due to the widespread use of web-cams and mobile devices embedded with a camera, it is now possible to realize facial video recognition, rather than resorting to just still images. In fact, facial video recognition offers many advantages over still image recognition; these include the potential of boosting the system accuracy and deterring spoof attacks. This paper presents an evaluation of person identity verification using facial video data, organized in conjunction with the International Conference on Biometrics (ICB 2009). It involves 18 systems submitted by seven academic institutes. These systems provide for a diverse set of assumptions, including feature representation and preprocessing variations, allowing us to assess the effect of adverse conditions, usage of quality information, query selection, and template construction for video-to-video face authentication. Norman Poh, Chi-Ho Chan, Josef Kittler, Sébastien Marcel, Chris McCool, Enrique Argones-Rúa, José Luis Alba-Castro, Mauricio Villegas, Roberto Paredes, Vitomir Struc, Nikola Pavesic, Albert Ali Salah, Hui Fang 0003, Nicholas Costen |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2010 | On-line Learning of Mutually Orthogonal Subspaces for Face Recognition by Image SetsabstractWe address the problem of face recognition by matching image sets. Each set of face images is represented by a subspace (or linear manifold) and recognition is carried out by subspace-to-subspace matching. In this paper, 1) a new discriminative method that maximises orthogonality between subspaces is proposed. The method improves the discrimination power of the subspace angle based face recognition method by maximizing the angles between different classes. 2) We propose a method for on-line updating the discriminative subspaces as a mechanism for continuously improving recognition accuracy. 3) A further enhancement called locally orthogonal subspace method is presented to maximise the orthogonality between competing classes. Experiments using 700 face image sets have shown that the proposed method outperforms relevant prior art and effectively boosts its accuracy by online learning. It is shown that the method for online learning delivers the same solution as the batch computation at far lower computational cost and the locally orthogonal method exhibits improved accuracy. We also demonstrate the merit of the proposed face recognition method on portal scenarios of multiple biometric grand challenge. Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla |
IEEE Trans. Image Process. | 2 |
| 2010 | Quality-Based Score Normalization With Device Qualitative Information for Multimodal Biometric FusionabstractAs biometric technology is rolled out on a larger scale, it will be a common scenario (known as cross-device matching) to have a template acquired by one biometric device used by another during testing. This requires a biometric system to work with different acquisition devices, an issue known as device interoperability. We further distinguish two subproblems, depending on whether the device identity is known or unknown. In the latter case, we show that the device information can be probabilistically inferred given quality measures (e.g., image resolution) derived from the raw biometric data. By keeping the template unchanged, cross-device matching can result in significant degradation in performance. We propose to minimize this degradation by using device-specific quality-dependent score normalization. In the context of fusion, after having normalized each device output independently, these outputs can be combined using the naive Bayes principal. We have compared and categorized several state-of-the-art quality-based score normalization procedures, depending on how the relationship between quality measures and score is modeled, as follows: 1) direct modeling; 2) modeling via the cluster index of quality measures; and 3) extending 2) to further include the device information (device-specific cluster index). Experimental results carried out on the Biosecure DS2 data set show that the last approach can reduce both false acceptance and false rejection rates simultaneously. Furthermore, the compounded effect of normalizing each system individually in multimodal fusion is a significant improvement in performance over the baseline fusion (without using any quality information) when the device information is given. Norman Poh, Josef Kittler, Thirimachos Bourlai |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2009 | Hierarchical Image Matching for Pose-invariant Face RecognitionabstractThe paper addresses the problem of face recognition under arbitrary pose. A hi-erarchical MRF-based image matching method for finding pixel-wise correspondences between facial images viewed from different angles is proposed and used to densely reg-ister a pair of facial images. The goodness-of-match between two faces is then measured in terms of the normalized energy of the match which is a combination of both structural differences between faces as well as their texture distinctiveness. The method needs no training on non-frontal images and circumvents the need for geometrical normalization of facial images. It is also robust to moderate scale changes between images. The proposed approach is evaluated on the CMU PIE database and promising results are obtained. 1 Shervin Rahimzadeh Arashloo, Josef Kittler |
BMVC | 2 |
| 2009 | 3D-assisted Facial Texture Super-ResolutionabstractIn this paper we propose a new framework for super-resolving facial images under arbitrary pose. While example-based super-resolution methods have demonstrated impressive results for face super-resolution under given pose and imaging conditions, they have limitations dealing with different poses and illuminations. Due to these limitations their application to face super-resolution in generalized situations is either impractical or sub-optimal. The proposed framework utilizes a 3D morphable face model in order to address the problem of face super-resolution under arbitrary pose. This framework does not assume any pre-defined pose for the subject thus it can be readily applied to any pose. The main contribution of this work is defining a framework in which a 3D morphable model can be used in conjunction with any of the example-based super-resolution methods. Experimental results prove the potential power of this method in face superresolution and its application to face recognition. Pouria Mortazavian, Josef Kittler, William J. Christmas |
BMVC | 2 |
| 2009 | Non-sparse Multiple Kernel Learning for Fisher Discriminant AnalysisabstractWe consider the problem of learning a linear combination of pre-specified kernel matrices in the Fisher discriminant analysis setting. Existing methods for such a task impose an ¿1norm regularisation on the kernel weights, which produces sparse solution but may lead to loss of information. In this paper, we propose to use ¿2norm regularisation instead. The resulting learning problem is formulated as a semi-infinite program and can be solved efficiently. Through experiments on both synthetic data and a very challenging object recognition benchmark, the relative advantages of the proposed method and its ¿1counterpart are demonstrated, and insights are gained as to how the choice of regularisation norm should be made. Fei Yan 0001, Josef Kittler, Krystian Mikolajczyk, Muhammad Atif Tahir |
ICDM | 2 |
| 2009 | Prototype selection based on sequential searchabstractIn this paper, we propose and explore the use of the sequential search for solving the prototype selection problem since this kind of search has shown good performance for solving selection problems. We propose three prototype selection methods based José Arturo Olvera-López, José Fco. Martínez-Trinidad, Jesús Ariel Carrasco-Ochoa, Josef Kittler |
Intell. Data Anal. | 4 |
| 2009 | A linear-complexity reparameterisation strategy for the hierarchical bootstrapping of capabilities within perception-action architectures
Mikhail Shevchenko, David Windridge, Josef Kittler |
Image Vis. Comput. | 3 |
| 2009 | Designing a smart-card-based face verification system: empirical investigation
Thirimachos Bourlai, Josef Kittler, Kieron Messer |
Mach. Vis. Appl. | 2 |
| 2009 | Influence of compression on 3D face recognition
Lorenzo Granai, Jose Rafael Tena, Miroslav Hamouz, Josef Kittler |
Pattern Recognit. Lett. | 4 |
| 2009 | Benchmarking quality-dependent and cost-sensitive score-level multimodal biometric fusion algorithmsabstractAutomatically verifying the identity of a person by means of biometrics (e.g., face and fingerprint) is an important application in our day-to-day activities such as accessing banking services and security control in airports. To increase the system reliability, several biometric devices are often used. Such a combined system is known as a multimodal biometric system. This paper reports a benchmarking study carried out within the framework of the BioSecure DS2 (Access Control) evaluation campaign organized by the University of Surrey, involving face, fingerprint, and iris biometrics for person authentication, targeting the application of physical access control in a medium-size establishment with some 500 persons. While multimodal biometrics is a well-investigated subject in the literature, there exists no benchmark for a fusion algorithm comparison. Working towards this goal, we designed two sets of experiments: quality-dependent and cost-sensitive evaluation. The quality-dependent evaluation aims at assessing how well fusion algorithms can perform under changing quality of raw biometric images principally due to change of devices. The cost-sensitive evaluation, on the other hand, investigates how well a fusion algorithm can perform given restricted computation and in the presence of software and hardware failures, resulting in errors such as failure-to-acquire and failure-to-match. Since multiple capturing devices are available, a fusion algorithm should be able to handle this nonideal but nevertheless realistic scenario. In both evaluations, each fusion algorithm is provided with scores from each biometric comparison subsystem as well as the quality measures of both the template and the query data. The response to the call of the evaluation campaign proved very encouraging, with the submission of 22 fusion systems. To the best of our knowledge, this campaign is the first attempt to benchmark quality-based multimodal fusion algorithms. In the presence of changing image quality which may be due to a change of acquisition devices and/or device capturing configurations, we observe that the top performing fusion algorithms are those that exploit automatically derived quality measurements. Our evaluation also suggests that while using all the available biometric sensors can definitely increase the fusion performance, this comes at the expense of increased cost in terms of acquisition time, computation time, the physical cost of hardware, and its maintenance cost. As demonstrated in our experiments, a promising solution which minimizes the composite cost is sequential fusion, where a fusion algorithm sequentially uses match scores until a desired confidence is reached, or until all the match scores are exhausted, before outputting the final combined score. Norman Poh, Thirimachos Bourlai, Josef Kittler, Lorène Allano, Fernando Alonso-Fernandez, Onkar Ambekar, John P. Baker, Bernadette Dorizzi, Omolara Fatukasi, Julian Fierrez, Harald Ganster, Javier Ortega-Garcia, Donald E. Maurer, Albert Ali Salah, Tobias Scheidat, Claus Vielhauer |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2008 | Subsurface scattering deconvolution for improved NIR-visible facial image correlationabstractSignificant improvements in face-recognition performance have recently been achieved by obtaining near infrared (NIR) probe images. We demonstrate that by taking into account the differential effects of sub-surface scattering, correlation between facial images in the visible (VIS) and NIR wavelengths can be significantly improved. Hence, by using Fourier analysis and Gaussian deconvolution with variable thresholds for the scattering deconvolution radius and frequency, sub-surface scattering effects are largely eliminated from perpendicular isomap transformations of the facial images. (Isomap images are obtained via scanning reconstruction, as in our case, or else, more generically, via model fitting). Thus, small-scale features visible in both the VIS and NIR, such as skin-pores and certain classes of skin-mottling, can be equally weighted within the correlation analysis. The method can consequently serves as the basis for more detailed forms of facial comparison. Josef Kittler, David Windridge, Debaditya Goswami |
FG | 1 |
| 2008 | A family of methods for quality-based multimodal biometric fusion using generative classifiersabstractAutomatically verifying the identity of a person by means of biometrics (e.g., face and fingerprint) is an important application in our day-to-day activities such as accessing banking services and security control in airports. To increase the system reliability, several biometric devices are often used. This paper considers how auxiliary information such as the quality associated with a biometric sample and the device information can be used when combining the output of several biometric devices. Since both these sources of information are not discriminative in distinguishing genuine users from impostors, combining them is indeed a challenging problem. We advance the state of the art of multimodal biometric fusion in two ways: first, we unify several existing generative classifiers using Bayesian networks. Second, we propose a novel fusion classifier incorporating both the quality and device information simultaneously. Our experiments based on the Biosecure DS2 dataset suggests that the proposed classifier can systematically achieve the best generalization performance compared to currently available state-of-the-art classifiers. Norman Poh, Josef Kittler |
ICARCV | 2 |
| 2008 | Feature condensing algorithm for feature selectionabstractA new unsupervised filter-based feature selection method is introduced. Its principle consists in merging similar features into clusters using a distance measure derived from the correlation coefficient. Subsequently, only one representative feature is selected from each cluster. In experiments with real-world data, we show that the proposed method is benefical as a pre-filtering step for more sophisticated feature selection techniques. Pavel Krízek, Josef Kittler, Václav Hlavác |
ICPR | 2 |
| 2008 | On using error bounds to optimize cost-sensitive multimodal biometric authenticationabstractWhile using more biometric traits in multimodal biometric fusion can effectively increase the system robustness, often, the cost associated to adding additional systems is not considered. In this paper, we propose an algorithm that can efficiently bound the biometric system error. This helps not only to speed up the search for the optimal system configuration by an order of magnitude but also unexpectedly to enhance the robustness to population mismatch. This suggests that bounding the error of biometric system from above can possibly be better than directly estimating it from the data. The latter strategy can be susceptible to spurious biometric samples and the particular choice of users. The efficiency of the proposal is achieved thanks to the use of Chernoff bound in estimating the authentication error. Unfortunately, such a bound assumes that the match scores are normally distributed, which is not necessarily the correct distribution model. We propose to transform simultaneously the class conditional match scores (genuine user or impostor scores) into ones that are more conforming to normal distributions using a modified criterion of the Box-Cox transform. Norman Poh, Josef Kittler |
ICPR | 2 |
| 2008 | A note on an extreme case of the generalized optimal discriminant transformation
Marco Loog, Xiaojun Wu 0001, Jieping Lu, Jing-Yu Yang 0001, Shitong Wang 0001, Josef Kittler |
Neurocomputing | 6 |
| 2008 | Layered Data Association Using Graph-Theoretic Formulation with Application to Tennis Ball Tracking in Monocular SequencesabstractIn this paper, we propose a multilayered data association scheme with graph-theoretic formulation for tracking multiple objects that undergo switching dynamics in clutter. The proposed scheme takes as input object candidates detected in each frame. At the object candidate level, "tracklets'' are "grown'' from sets of candidates that have high probabilities of containing only true positives. At the tracklet level, a directed and weighted graph is constructed, where each node is a tracklet, and the edge weight between two nodes is defined according to the "compatibility'' of the two tracklets. The association problem is then formulated as an all-pairs shortest path (APSP) problem in this graph. Finally, at the path level, by analyzing the APSPs, all object trajectories are identified, and track initiation and track termination are automatically dealt with. By exploiting a special topological property of the graph, we have also developed a more efficient APSP algorithm than the general-purpose ones. The proposed data association scheme is applied to tennis sequences to track tennis balls. Experiments show that it works well on sequences where other data association methods perform poorly or fail completely. Fei Yan 0001, William J. Christmas, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Gesture spotting for low-resolution sports video annotation
Myung-Cheol Roh, William J. Christmas, Josef Kittler, Seong-Whan Lee |
Pattern Recognit. | 3 |
| 2008 | Incorporating Model-Specific Score Distribution in Speaker Verification SystemsabstractIt has been shown that the authentication performance of a biometric system is dependent on the models/templates specific to a user. As a result, some users may be more easily recognized or impersonated than others. The various categories of users have been characterized by Doddington(1988). We refer to this unbalanced performance across users as the Doddington's zoo effect. In the context of fusion, we argue that this effect is system-dependent, i.e., a user model that is easily impersonated (a lamb) in one system may be easily recognized in another system (a sheep). While in principle, a fusion system could be trained to cope with the changing animal behavior of users from system to system, the lack of training data makes it impossible. We believe that one major cause of the Doddington's zoo effect is the variation of class conditional scores from one speaker model to another. We propose a two-level fusion framework that effectively realizes a fusion classifier adapted to each user. First, one applies aclient-specific(or model-specific) score normalization procedure to each of the system outputs to be combined. Then, one feeds the resulting normalized outputs to a fusion classifier (common to all users) as input to obtain a final combined score. Two existing model-specific score normalization procedures are considered in this framework, i.e., F- and Z-norms. In addition to them, a novel score normalization method called model-specific log-likelihood ratio (MS-LLR) is also proposed. While Z-norm is impostor-centric, i.e., it makes use of only the impostor score statistics, F-norm and the proposed MS-LLR are client-impostor centric, i.e., they consider both the client and impostor score statistics simultaneously. Our findings based on the XM2VTS and the NIST2005 databases show that when client-impostor centric normalization procedures are used to implement the proposed two-level fusion framework, the res Norman Poh, Josef Kittler |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Image Feature Localization by Multiple Hypothesis Testing of Gabor FeaturesabstractSeveral novel and particularly successful object and object category detection and recognition methods based on image features, local descriptions of object appearance, have recently been proposed. The methods are based on a localization of image features and a spatial constellation search over the localized features. The accuracy and reliability of the methods depend on the success of both tasks: image feature localization and spatial constellation model search. In this paper, we present an improved algorithm for image feature localization. The method is based on complex-valued multi resolution Gabor features and their ranking using multiple hypothesis testing. The algorithm provides very accurate local image features over arbitrary scale and rotation. We discuss in detail issues such as selection of filter parameters, confidence measure, and the magnitude versus complex representation, and show on a large test sample how these influence the performance. The versatility and accuracy of the method is demonstrated on two profoundly different challenging problems (faces and license plates). Jarmo Ilonen, Joni-Kristian Kämäräinen, Pekka Paalanen, Miroslav Hamouz, Josef Kittler, Heikki Kälviäinen |
IEEE Trans. Image Process. | 5 |
| 2007 | 2D face pose normalisation using a 3D morphable modelabstractThe ever growing need for improved security, surveillance and identity protection, calls for the creation of evermore reliable and robust face recognition technology that is scalable and can be deployed in all kinds of environments without compromising its effectiveness. In this paper we study the impact that pose correction has on the performance of 2D face recognition. To measure the effect, we use a state of the art 2D recognition algorithm. The pose correction is performed by means of 3D morphable model. Our results on the non frontal XM2VTS database showed that pose correction can improve recognition rates up to 30%. Jose Rafael Tena, Raymond S. Smith, Miroslav Hamouz, Josef Kittler, Adrian Hilton 0001, John Illingworth |
AVSS | 4 |
| 2007 | All Pairs Shortest Path Formulation for Multiple Object Tracking with Application to Tennis Video AnalysisabstractIn previous work, we developed a novel data association algorithm with graph-theoretic formulation, and used it to track a tennis ball in broadcast tennis video. However, the track initiation/termination was not automatic, and it could not deal with situations in which more than one ball appeared in the scene. In this paper, we extend our previous work to track multiple tennis balls fully automatically. The algorithm presented in this paper requires the set of all-pairs shortest paths in a directed and edge-weighted graph. We also propose an efficient All-Pairs Shortest Path algorithm by exploiting a special topological property of the graph. Comparative experiments show that the proposed data association algorithm performs well both in terms of efficiency and tracking accuracy. 1 Fei Yan 0001, William J. Christmas, Josef Kittler |
BMVC | 3 |
| 2007 | Improving Stability of Feature Selection Methods
Pavel Krízek, Josef Kittler, Václav Hlavác |
CAIP | 2 |
| 2007 | Quality Controlled Multimodal Fusion of Biometric Experts
Omolara Fatukasi, Josef Kittler, Norman Poh |
CIARP | 2 |
| 2007 | A Method for Estimating Authentication Performance over Time, with Applications to Face Biometrics
Norman Poh, Josef Kittler, Raymond S. Smith, Jose Rafael Tena |
CIARP | 2 |
| 2007 | Incremental Linear Discriminant Analysis Using Sufficient Spanning Set ApproximationsabstractThis paper presents a new incremental learning solution for linear discriminant analysis (LDA). We apply the concept of the sufficient spanning set approximation in each update step, i.e. for the between-class scatter matrix, the projected data matrix as well as the total scatter matrix. The algorithm yields a more general and efficient solution to incremental LDA than previous methods. It also significantly reduces the computational complexity while providing a solution which closely agrees with the batch LDA result. The proposed algorithm has a time complexity of O(Nd2) and requires O(Nd) space, where d is the reduced subspace dimension and N the data dimension. We show two applications of incremental LDA: First, the method is applied to semi-supervised learning by integrating it into an EM framework. Secondly, we apply it to the task of merging large databases which were collected during MPEG standardization for face image retrieval. Tae-Kyun Kim 0001, Shu-Fai Wong, Björn Stenger, Josef Kittler, Roberto Cipolla |
CVPR | 4 |
| 2007 | Object Localisation Using Generative Probability Model for Spatial Constellation and Local Image FeaturesabstractIn this paper we apply state-of-the-art approach to object detection and localisation by incorporating local descriptors and their spatial configuration into a generative probability model. In contrast to the recent semi- supervised methods we do not utilise interest point detectors, but apply a supervised approach where local image features (landmarks) are annotated in a training set and therefore their appearance and spatial variation can be learnt. Our method enables working in purely probabilistic search spaces providing a MAP estimate of object location, and in contrast to the recent methods, no background class needs to be formed. Using the training set we can estimate pdfs for both spatial constellation and local feature appearance. By applying an inference bias that the largest pdf mode has probability one, we are able to combine prior information (spatial configuration of the features) and observations (image feature appearance) into posterior distribution which can be generatively sampled, e.g. using MCMC techniques. The MCMC methods are sensitive to initialisation, but as a solution, we also propose a very efficient and accurate RANSAC-based method for finding good initial hypotheses of object poses. The complete method can robustly and accurately detect and localise objects under any homography. Joni-Kristian Kämäräinen, Miroslav Hamouz, Josef Kittler, Pekka Paalanen, Jarmo Ilonen, Alexander Drobchenko |
ICCV | 3 |
| 2007 | An extreme case of the generalized optimal discriminant transformation and its application to face recognition
Xiaojun Wu 0001, Jieping Lu, Jing-Yu Yang 0001, Shitong Wang 0001, Josef Kittler |
Neurocomputing | 5 |
| 2007 | Discriminative Learning and Recognition of Image Set Classes Using Canonical CorrelationsabstractWe address the problem of comparing sets of images for object recognition, where the sets may represent variations in an object's appearance due to changing camera pose and lighting conditions. Canonical Correlations (also known as principal or canonical angles), which can be thought of as the angles between two d-dimensional subspaces, have recently attracted attention for image set matching. Canonical correlations offer many benefits in accuracy, efficiency, and robustness compared to the two main classical methods: parametric distribution-based and nonparametric sample-based matching of sets. Here, this is first demonstrated experimentally for reasonably sized data sets using existing methods exploiting canonical correlations. Motivated by their proven effectiveness, a novel discriminative learning method over sets is proposed for set classification. Specifically, inspired by classical Linear Discriminant Analysis (LDA), we develop a linear discriminant function that maximizes the canonical correlations of within-class sets and minimizes the canonical correlations of between-class sets. Image sets transformed by the discriminant function are then compared by the canonical correlations. Classical orthogonal subspace method (OSM) is also investigated for the similar purpose and compared with the proposed method. The proposed method is evaluated on various object recognition problems using face image sets with arbitrary motion captured under different illuminations and image sets of 500 general objects taken at different views. The method is also applied to object category recognition using ETH-80 database. The proposed method is shown to outperform the state-of-the-art methods in terms of accuracy and efficiency. Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Introduction to the Special Issue on Biometrics: Progress and DirectionsabstractThe guest editors provide an overview of the articles selected for this special issue. The issue's goal is to document the current state-of-the-art, acknowledge the latest breakthroughs achieved by scientists working in the area of biometric recognition, and identify future promising research areas. It is thought the selection of papers discussed should give readers a good idea of where researchers have been focusing, both on long- studied problems still needing more work and on newer challenges. A fundamental of the field of biometrics is an ever-increasing need for better recognition and stronger security. But, as public and commercial biometric deployments increase in number, there is also more need to understand privacy issues and to provide greater ease-of-use. The volume and quality of papers in this special issue indicate that much progress has been made in many aspects of the biometrics field and that there are challenging and promising future directions still to follow. Salil Prabhakar, Josef Kittler, Davide Maltoni, Lawrence O'Gorman, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Visual Bootstrapping for Unsupervised Symbol Grounding
Josef Kittler, Mikhail Shevchenko, David Windridge |
ACIVS | 1 |
| 2006 | On Optimisation of Smart Card Face Verification SystemsabstractThe optimisation of a smart card face verification system (SCFVS) design is a complex task. As the parameters involved are not independent, the search space is of exponential complexity. We investigate simplified optimisation strategies and demonstrate that both system performance and speed of access can be improved by jointly optimised parameter setting and level of probe compression. Experimental results suggest that the choice of one strategy over another is a matter of the amount of time available for the system design, system performance and response time. Thirimachos Bourlai, Josef Kittler, Kieron Messer |
AVSS | 2 |
| 2006 | Incremental Learning of Locally Orthogonal Subspaces for Set-based Object RecognitionabstractOrthogonal subspaces are effective models to represent object image sets (generally any high-dimensional vector sets). Canonical correlation analysis of the orthogonal subspaces provides a good solution to discriminate objects with sets of images. In such a recognition task involving image sets, an efficient learning over a large volume of image sets, which may be increasing over time, is important. In this paper, an incremental learning method of orthogonal subspaces is proposed by updating the principal components of the class correlation and total correlation matrices separately, yielding the same solution as the batch computation with far lower computational cost. A novel concept of local orthogonality is further proposed to cope with non-linear manifolds of data vectors and find a more optimal solution of orthogonal subspaces for a certain neighbouring object image sets. In the experiments using 700 face image sets, the locally orthogonal subspaces outperformed the orthogonal subspaces as well as relevant state-of-the-art methods in accuracy. Note that the locally orthogonal subspaces are also amenable to incremental updating due to their linear property. 1 Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla |
BMVC | 2 |
| 2006 | General Pose Face Recognition Using Frontal Face Model
Jean-Yves Guillemaut, Josef Kittler, Mohammad Sadeghi 0001, William J. Christmas |
CIARP | 2 |
| 2006 | A Comparative Study of Face Representations in the Frequency Domain
Eduardo Garea Llano, Josef Kittler, Kieron Messer, Heydi Mendez Vazquez |
CIARP | 2 |
| 2006 | A Novel Data Association Algorithm for Object Tracking in Clutter with Application to Tennis Video AnalysisabstractIt is well recognised that data association is critically important for object tracking. However, in the presence of successive misdetections, a large number of false candidates and an unknown number of abrupt model switchings that happen unpredictably, the data association problem can be very difficult. We tackle these difficulties by using a layered data association scheme. At the object level, trajectories are "grown" from sets of object candidates that have high probabilities of containing only true positives; by this means the otherwise combinatorial complexity is significantly reduced. Dijkstra’s shortest path algorithm is then used to perform data association at the trajectory level. The algorithm is applied to low-quality tennis video sequences to track a tennis ball. Experiments show that the algorithm is robust to abrupt model switchings, and performs well in heavily cluttered environments. Fei Yan 0001, Alexey Kostin, William J. Christmas, Josef Kittler |
CVPR (1) | 4 |
| 2006 | Learning Discriminative Canonical Correlations for Object Recognition with Image Sets
Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla |
ECCV (3) | 2 |
| 2006 | Robust Player Gesture Spotting and Recognition in Low-Resolution Sports Video
Myung-Cheol Roh, William J. Christmas, Josef Kittler, Seong-Whan Lee |
ECCV (4) | 3 |
| 2006 | Fusion of Talking Face Biometric Modalities for Personal Identity VerificationabstractWe describe a personal identity verification system based on lip dynamics biometric. The lip shape is represented in terms of a B-spline model, tracked over time. The coordinates of the 11 control points of the B-spline model are used as features for each frame. An utterance consisting of N frames produces a sequence of 22 dimensional feature vectors that is matched to the template using dynamic time warping. The verification error rate achived by the systems on the XM2VTS database is about 14%. By fusing the system with face and voice biometrics the error rate is reduced to a fractiopn of one percent. M. Ulises Ramos Sánchez, Josef Kittler |
ICASSP (5) | 2 |
| 2006 | Design and Fusion of Pose-Invariant Face-Identification ExpertsabstractWe address the problem of pose-invariant face recognition based on a single model image. To cope with novel view face images, a model of the effect of pose changes on face appearance must be available. Face images at an arbitrary pose can be mapped to a reference pose by the model yielding view-invariant representation. Such a model typically relies on dense correspondences of different view face images, which are difficult to establish in practice. Errors in the correspondences seriously degrade the accuracy of any recognizer. Therefore, we assume only the minimal possible set of correspondences, given by the corresponding eye positions. We investigate a number of approaches to pose-invariant face recognition exploiting such a minimal set of facial features correspondences. Four different methods are proposed as pose-invariant face recognition "experts" and combined in a single framework of expert fusion. Each expert explicitly or implicitly realizes the three sequential functions jointly required to capture the nonlinear manifolds of face pose changes: representation, view transformation, and class discriminative feature extraction. Within this structure, the experts are designed for diversity. We compare a design in which the three stages are sequentially optimized with two methods which employ an overall single nonlinear function learnt from different view face images. We also propose an approach exploiting a three-dimensional face data. A lookup table storing facial feature correspondences between different pose images, found by 3-D face models, is constructed. The designed experts are different in their nature owing to different sources of information and architectures used. The proposed fusion architecture of the pose-invariant face experts achieves an impressive accuracy gain by virtue of the individual experts diversity. It is experimentally shown that the individual experts outperform the classical linear discriminant analysis (LDA) method on the XM2VTS face data set consisting of about 300 face classes. Further impressive performance gains are obtained by combining the outputs of the experts using different fusion strategies Tae-Kyun Kim 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | A Tennis Ball Tracking Algorithm for Automatic Annotation of Tennis MatchabstractSeveral tennis ball tracking algorithms have been reported in the literature. However, most of them use high quality video and multiple cameras, and the emphasis has been on coordinating the cameras, or visualising the tracking results. In this paper, we propose a tennis ball tracking algorithm for low quality off-air video recorded with a single camera. Multiple visual cues are exploited for tennis candidate detection. A particle filter with improved sampling efficiency is used to track the tennis candidates. Experimental results show that our algorithm is robust and has a tracking accuracy that is sufficiently high for automatic annotation of tennis matches. 1 Fei Yan 0001, William J. Christmas, Josef Kittler |
BMVC | 3 |
| 2005 | Face Recognition Using Active Near-IR IlluminationabstractA new approach to overcome the problem caused by illumination variation in face recognition is proposed in this paper. Active Near-Infrared (Near-IR) illumination projected by a Light Emitting Diode (LED) light source is used to provide a constant illumination. The difference between two face images captured when the LED light is on and off respectively, is the image of a face under just the LED illumination, and is independent of ambient illumination. In preliminary experiments with various ambient illuminations, significantly better results are achieved for both automatic and semi-automatic face recognition experiments on LED illuminated faces than on face images under ambient illuminations. 1 Xuan Zou, Josef Kittler, Kieron Messer |
BMVC | 2 |
| 2005 | 3D Assisted 2D Face Recognition: Methodology
Josef Kittler, Miroslav Hamouz, Jose Rafael Tena, Adrian Hilton 0001, John Illingworth, M. Ruiz |
CIARP | 1 |
| 2005 | Performance measures of the tomographic classifier fusion methodologyabstractWe seek to quantify both the classification performance and estimation error robustness of the authors' tomographic classifier fusion methodology by contrasting it in field tests and model scenarios with the sum and product classifier fusion methodologies. In particular, we seek to confirm that the tomographic methodology represents a generally optimal strategy across the entire range of problem dimensionalities, and at a sufficient margin to justify the general advocation of its use. Final results indicate, in particular, a near 25% improvement on the next nearest performing combination scheme at the extremity of the tested dimensional range. David Windridge, Josef Kittler |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2005 | Component-based LDA face description for image retrieval and MPEG-7 standardisation
Tae-Kyun Kim 0001, Wonjun Hwang, Josef Kittler |
Image Vis. Comput. | 4 |
| 2005 | Feature-Based Affine-Invariant Localization of FacesabstractWe present a novel method for localizing faces in person identification scenarios. Such scenarios involve high resolution images of frontal faces. The proposed algorithm does not require color, copes well in cluttered backgrounds, and accurately localizes faces including eye centers. An extensive analysis and a performance evaluation on the XM2VTS database and on the realistic BioID and BANCA face databases is presented. We show that the algorithm has precision superior to reference methods. Miroslav Hamouz, Josef Kittler, Joni-Kristian Kämäräinen, Pekka Paalanen, Heikki Kälviäinen, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Locally Linear Discriminant Analysis for Multimodally Distributed Classes for Face Recognition with a Single Model ImageabstractWe present a novel method of nonlinear discriminant analysis involving a set of locally linear transformations called "Locally Linear Discriminant Analysis (LLDA)." The underlying idea is that global nonlinear data structures are locally linear and local structures can be linearly aligned. Input vectors are projected into each local feature space by linear transformations found to yield locally linearly transformed classes that maximize the between-class covariance while minimizing the within-class covariance. In face recognition, linear discriminant analysis (LDA) has been widely adopted owing to its efficiency, but it does not capture nonlinear manifolds of faces which exhibit pose variations. Conventional nonlinear classification methods based on kernels such as generalized discriminant analysis (GDA) and support vector machine (SVM) have been developed to overcome the shortcomings of the linear method, but they have the drawback of high computational cost of classification and overfitting. Our method is for multiclass nonlinear discrimination and it is computationally highly efficient as compared to GDA. The method does not suffer from overfitting by virtue of the linear base structure of the solution. A novel gradient-based learning algorithm is proposed for finding the optimal set of local linear bases. The optimization does not exhibit a local-maxima problem. The transformation functions facilitate robust face recognition in a low-dimensional subspace, under pose variations, using a single model image. The classification results are given for both synthetic and real face data. Tae-Kyun Kim 0001, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Object recognition by symmetrised graph matching using relaxation labelling with an inhibitory mechanism
Alexey Kostin, Josef Kittler, William J. Christmas |
Pattern Recognit. Lett. | 2 |
| 2005 | A unified approach to the generation of semantic cues for sports video annotation
Kieron Messer, William J. Christmas, Edward Jaser, Josef Kittler, Barbara Levienaise-Obadia, Dimitri Koubaroulis |
Signal Process. | 4 |
| 2005 | Fast robust correlationabstractA new, fast, statistically robust, exhaustive, translational image-matching technique is presented: fast robust correlation. Existing methods are either slow or non-robust, or rely on optimization. Fast robust correlation works by expressing a robust matching surface as a series of correlations. Speed is obtained by computing correlations in the frequency domain. Computational cost is analyzed and the method is shown to be fast. Speed is comparable to conventional correlation and, for large images, thousands of times faster than direct robust matching. Three experiments demonstrate the advantage of the technique over standard correlation. Alistair John Fitch, Alexander Kadyrov, William J. Christmas, Josef Kittler |
IEEE Trans. Image Process. | 4 |
| 2004 | A New Kernel Direct Discriminant Analysis (KDDA) Algorithm for Face RecognitionabstractWe propose a new kernel direct discriminant analysis (KDDA) algorithm in this paper. First, a recently advocated direct linear discriminant analysis (DLDA) algorithm is overviewed. Then the new KDDA algorithm is developed which can be considered as a kernel version of the DLDA algorithm. The design of the minimum distance classifier in the new kernel subspace is then discussed. The results of experiments on two well-known facial databases show the effectiveness of the proposed method in face recognition. The results of experiments also confirm that DLDA can be viewed as a special case of the proposed KDDA algorithm. 1. Xiaojun Wu 0001, Josef Kittler, Jing-Yu Yang 0001, Kieron Messer, Shitong Wang 0001 |
BMVC | 2 |
| 2004 | Use of Context in Automatic Annotation of Sports Videos
Ilias Kolonias, William J. Christmas, Josef Kittler |
CIARP | 3 |
| 2004 | Hierarchical Decision Making Scheme for Sports Video Categorisation with Temporal Post-Processing
Edward Jaser, Josef Kittler, William J. Christmas |
CVPR (2) | 2 |
| 2004 | Multiple Classifier System Approach to Model Pruning in Object Recognition
Josef Kittler, Alireza Ahmadyfard |
ECCV (4) | 1 |
| 2004 | The Role Of Relational Constraints In Region MatchingabstractWe propose a graph-based representation for the elliptic region shape descriptors introduced by Tuytelaars et al.13 In this representation we use image profiles to describe the relation between a pair of image regions. This new representation and a graph matching technique proposed in Ref. 1 are the basis of an object recognition method. An experimental comparative study between the original method and the new graph-based method is carried out. The results show that the graph-based method is more robust to scaling than the original method. Moreover, the misclassification rate using the graph-based method is considerably lower than that yielded by the original method. Alireza Ahmadyfard, Josef Kittler |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2004 | Editorial
Fabio Roli, Josef Kittler |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2004 | Fast Branch & Bound Algorithms for Optimal Feature SelectionabstractA novel search principle for optimal feature subset selection using the Branch & Bound method is introduced. Thanks to a simple mechanism for predicting criterion values, a considerable amount of time can be saved by avoiding many slow criterion evaluations. We propose two implementations of the proposed prediction mechanism that are suitable for use with nonrecursive and recursive criterion forms, respectively. Both algorithms find the optimum usually several times faster than any other known Branch & Bound algorithm. As the algorithm computational efficiency is crucial, due to the exponential nature of the search problem, we also investigate other factors that affect the search performance of all Branch & Bound algorithms. Using a set of synthetic criteria, we show that the speed of the Branch & Bound algorithms strongly depends on the diversity among features, feature stability with respect to different subsets, and criterion function dependence on feature set size. We identify the scenarios where the search is accelerated the most dramatically (finish in linear time), as well as the worst conditions. We verify our conclusions experimentally on three real data sets using traditional probabilistic distance criteria. Petr Somol, Pavel Pudil, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Multiple classifier combination for face-based identity verification
Jacek Czyz, Josef Kittler, Luc Vandendorpe |
Pattern Recognit. | 2 |
| 2004 | Independent component analysis in a local facial residue space for face recognition
Tae-Kyun Kim 0001, Wonjun Hwang, Josef Kittler |
Pattern Recognit. | 4 |
| 2004 | An analytical algorithm for determining the generalized optimal set of discriminant vectors
Xiaojun Wu 0001, Josef Kittler, Jing-Yu Yang 0001, Shitong Wang 0001 |
Pattern Recognit. | 2 |
| 2003 | Discriminant Analysis by Locally Linear TransformationsabstractWe present a novel discriminant analysis learning method which is applicable to non-linear data structures. The method can deal with pattern classification problems which have a multi-modal distribution for each class and samples of other classes may be closer to a class than those of the class itself. Conventional linear discriminant analysis (LDA) and LDA mixture model can not solve this linearly non-separable problem. Several local linear transformations are considered to yield locally transformed classes that maximize the between-class covariance and minimize the within-class covariance. The method invloves a novel gradient based algorithm for finding the optimal set of local linear bases. It does not have a local-maxima problem and stably converges to the global maximum point. The method is computationally efficienct as compared to the previous non-linear discriminant analysis based on the kernel approach. The method does not suffer from an overfitting problem by virtue of the linear base structure of the solution. The classification results are given for both simulated data and real face data. 1 Tae-Kyun Kim 0001, Josef Kittler, Seok-Cheol Kee |
BMVC | 2 |
| 2003 | Application of Characteristic Function Method in Target DetectionabstractTarget detection is one of the important elements of Automatic Target Recognition (ATR) systems. In this paper, we propose a new approach to detect outliers in radar returns, based on modelling the background using an empirical distribution rather than a parametric distribution. The key innovation lies in the use of the Characteristic Function (CF) to describe the distribution. The experimental results show a promising performance improvement in terms of detection rate and lower false alarm rate, compared with the conventional Gaussian model which employs the Mahalanobis metric as a distance function. 1 Mohammad Hamiruce Marhaban, Josef Kittler |
BMVC | 2 |
| 2003 | Independent Component Analysis in a Facial Local Residue SpaceabstractIn this paper, we propose an ICA (Independent Component Analysis) based face recognition algorithm, which is robust to illumination and pose variation. Generally, it is well known that the first few eigenfaces represent illumination variation rather than identity. Most PCA (Principal Component Analysis)-based methods have overcome illumination variation by discarding the projection to a few leading eigenfaces. The space spanned after removing a few leading eigenfaces is called the "residual face space". We found that ICA in the residual face space provides more efficient encoding in terms of redundancy reduction and robustness to pose variation as well as illumination variation, owing to its ability to represent non-Gaussian statistics. Moreover, a face image is separated into several facial components, local spaces, and each local space is represented by the ICA bases (independent components) of its corresponding residual space. The statistical models of face images in local spaces are relatively simple and facilitate classification by a linear encoding. Various experimental results show that the accuracy of face recognition is significantly improved by the proposed method under large illumination and pose variations. Tae-Kyun Kim 0001, Wonjun Hwang, Seok-Cheol Kee, Josef Kittler |
CVPR (1) | 5 |
| 2003 | Using a pictorial dictionary as a high level user interface for visual information retrievalabstractThe need for efficient retrieval of visual information is now widely accepted across many research domains. While much progress has been made in the area of low level representation and matching, visual information retrieval systems are often limited by the users ability to express a given query. Current retrieval technology does not allow human operators to formulate queries by means of high level semantics. In this paper we propose a 'pictorial dictionary' scheme to address these problems. Lee Gregory, Josef Kittler |
ICIP (2) | 2 |
| 2003 | Face description based on decomposition and combining of a facial space with LDAabstractWe propose a method of efficient face description for facial image retrieval from a large data set. The novel descriptor is obtained by decomposing the face image into several components and then combining the component features. The decomposition combined with LDA (linear discriminant analysis) provides discriminative facial features that are less sensitive to light and pose changes. Each component is represented in its Fisher space and another LDA is then applied to compactly combine the features of the components. To enhance retrieval accuracy further, a simple pose classification and transformation technique is performed, followed by recursive matching. The experimental results obtained on the MPEG-7 data set show an impressive accuracy of our algorithm as compared with the conventional PCA/ICA/LDA methods. Tae-Kyun Kim 0001, Wonjun Hwang, Seok-Cheol Kee, Josef Kittler |
ICIP (3) | 5 |
| 2003 | A Multiple Classifier System Approach to Affine Invariant Object Recognition
Alireza Ahmadyfard, Josef Kittler |
ICVS | 2 |
| 2003 | A Multimedia System Architecture for Automatic Annotation of Sports Videos
William J. Christmas, Edward Jaser, Kieron Messer, Josef Kittler |
ICVS | 4 |
| 2003 | Face verification via error correcting output codes
Josef Kittler, Reza Ghaderi, Terry Windeatt, Jiri Matas |
Image Vis. Comput. | 1 |
| 2003 | Sum Versus Vote Fusion in Multiple Classifier SystemsabstractAmidst the conflicting experimental evidence of superiority of one over the other, we investigate the Sum and majority Vote combining rules in a two class case, under the assumption of experts being of equal strength and estimation errors conditionally independent and identically distributed. We show, analytically, that, for Gaussian estimation error distributions, Sum always outperforms Vote. For heavy tail distributions, we demonstrate by simulation that Vote may outperform Sum. Results on synthetic data confirm the theoretical predictions. Experiments on real data support the general findings, but also show the effect of the usual assumptions of conditional independence, identical error distributions, and common target outputs of the experts not being fully satisfied. Josef Kittler, Fuad M. Alkoot |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | A Morphologically Optimal Strategy for Classifier Combination: Multiple Expert Fusion as a Tomographic ProcessabstractWe specify an analogy in which the various classifier combination methodologies are interpreted as the implicit reconstruction, by tomographic means, of the composite probability density function spanning the entirety of the pattern space, the process of feature selection in this scenario amounting to an extremely bandwidth-limited Radon transformation of the training data. This metaphor, once elaborated, immediately suggests techniques for improving the process, ultimately defining, in reconstructive terms, an optimal performance criterion for such combinatorial approaches. David Windridge, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | A comparative study of two object recognition methodsabstractAn experimental comparative study between two representation methods for the recognition of 3D objects from a 2D view is carried out. The two methods compared are our ARG region-based representation [1] and the elliptic region-based method of Tuytelaars et al[9]. The results of the experiments conducted show that the former method outperforms the latter particularly under sever scaling and also when applied to objects with curved surfaces. Alireza Ahmadyfard, Josef Kittler |
BMVC | 2 |
| 2002 | Orientation CorrelationabstractA new method of translational image registration is presented: 'orientation correlation'. The method is fast, exhaustive, statistically robust, and illumination invariant. No existing method has all of these properties. A modification that is particularly well suited to matching images of differing modalities, 'squared orientation correlation', is also given. Orientation correlation works by correlating 'orientation images'. Each pixel in a orientation image is a complex number that represents the orientation of intensity gradient. This representation is invariant to illumination change. Angles of gradient orientation are matched. Andrews robust kernel function is applied to angle differences. Through the use of correlation the method is exhaustive. The method is fast as the correlation can be computed using Fast Fourier Transforms. Alistair John Fitch, Alexander Kadyrov, William J. Christmas, Josef Kittler |
BMVC | 4 |
| 2002 | Enhancing the performance of personal identity authentication systems by fusion of face verification expertsabstractWe investigate the behavior knowledge space method (see Xu, L. et al., IEEE Transactions SMC, vol.22, no.3, p.418-35, 1992) and decision templates method (see Kuncheva, L. et al., Pattern Recognition, vol.34, p.299-314, 2001) of classifier fusion in the context of face verification. The study involves six experts which are not only correlated, but also their performance levels differ by as much as a factor of three. Through extensive experiments on the XM2VTS database using the Lausanne protocol, we found that the behavior knowledge space fusion strategy achieved consistently better results than the decision templates method. Most importantly, it exhibited quasi monotonic behavior as the number of experts combined increased. This is a very important conclusion, as it means that the performance of the multimodal system is not degraded by adding experts. Josef Kittler, Marco Ballette, Jacek Czyz, Fabio Roli, Luc Vandendorpe |
ICME (2) | 1 |
| 2002 | The Multimodal Neighborhood Signature for Modeling Object Color Appearance and Applications in Object Recognition and Image Retrieval
Jiri Matas, Dimitri Koubaroulis, Josef Kittler |
Comput. Vis. Image Underst. | 3 |
| 2002 | Using relaxation technique for region-based object recognition
Alireza Ahmadyfard, Josef Kittler |
Image Vis. Comput. | 2 |
| 2002 | Support vector machines for face authentication
Kenneth Jonsson, Josef Kittler, Yongping Li, Jiri Matas |
Image Vis. Comput. | 2 |
| 2002 | Multiple Classifier Fusion in Probabilistic Neural Networks
Jirí Grim, Josef Kittler, Pavel Pudil, Petr Somol |
Pattern Anal. Appl. | 2 |
| 2002 | Moderating k-NN Classifiers
Josef Kittler, Fuad M. Alkoot |
Pattern Anal. Appl. | 1 |
| 2002 | Model Selection by Predictive Validation
Josef Kittler, Kieron Messer, Mohammad Sadeghi 0001 |
Pattern Anal. Appl. | 1 |
| 2002 | Modified product fusion
Fuad M. Alkoot, Josef Kittler |
Pattern Recognit. Lett. | 2 |
| 2002 | An active mesh based tracker for improved feature correspondences
Andrew Griffin, Josef Kittler |
Pattern Recognit. Lett. | 2 |
| 2001 | Face Verification via ECOCabstractWe develop a novel approach to face verification based on the Error Correcting Output Coding (ECOC) classifier design concept. In the training phase the client set is repeatedly divided into two ECOC specified sub-sets (superclasses) to train a set of binary classifiers. The output of the classifiers defines the ECOC feature space, in which it is easier to separate transformed patterns representing clients and impostors. The proposed method exhibits superior verification performance on the well known XM2VTS data set as compared with previously reported results. 1 Josef Kittler, Reza Ghaderi, Terry Windeatt, Jiri Matas |
BMVC | 1 |
| 2001 | Real Time Segmentation of Lip Pixels for Lip Tracker Initialization
Mohammad Sadeghi 0001, Josef Kittler, Kieron Messer |
CAIP | 2 |
| 2001 | Face Verification Using Error Correcting Output CodesabstractThe error correcting output coding (ECOC) approach to classifier design decomposes a multi-class problem into a set of complementary two-class problems. We show how to apply the ECOC concept to automatic face verification, which is inherently a two-class problem. The output of the binary classifiers defines the ECOC feature space, in which it is easier to separate transformed patterns representing clients and impostors. We propose two different combining strategies as the matching score for face verification. The first uses the first order Minkowski metric, and requires a threshold to be set. The second is a kernel-based method and has no parameters to set. The proposed method exhibits better performance on the well known XM2VTS data set compared with previous reported results. Josef Kittler, Reza Ghaderi, Terry Windeatt, Jiri Matas |
CVPR (1) | 1 |
| 2001 | Segmentation of lip pixels for lip tracker initialisationabstractWe propose a novel image segmentation method for lip tracker initialisation which is based on a Gaussian mixture model of the pixel RGB values. The model is built using the predictive validation technique advocated by Kittler, Messer and Sadeghi (see Second International Conference on Advances in Pattern Recognition, Brazil, March 2001) which has been modified to allow modelling with full covariance matrices. A subsequent grouping of the mixture components provides the basis for a Bayesian rule labelling of the pixels as lip or non-lip. We test the proposed method on a database of 145 images and demonstrate that its accuracy is significantly better than the segmentation obtained by k-means clustering. Moreover, the proposed method does not require the number of segments to be specified a priori. Josef Kittler, Kieron Messer, Mohammad Sadeghi 0001 |
ICIP (1) | 1 |
| 2001 | Generation of semantic cues for sports video annotationabstractThe use of video and audio features for automated annotation of audio-visual data is becoming widespread. A major limitation of many of the current methods is that the stored indexing features are too low-level-they relate directly to properties of the data. We apply a further stage of processing that associates the feature measurements with real-world objects or events. The outputs, which we call "cues", denote the probability of the object being present in the scene. An additional advantage of this approach is that the cues from different types of features are presented in a homogeneous way. Kieron Messer, Josef Kittler, Barbara Levienaise-Obadia, William J. Christmas, Dimitri Koubaroulis |
ICIP (3) | 2 |
| 2001 | Empirical evaluation of a calibration chart detector
Soh Ling Min, Josef Kittler, Jiri Matas |
Mach. Vis. Appl. | 2 |
| 2001 | Recognition of polyhedral objects using triplets of projected spatial edges based on a single perspective image
Kok Cheong Wong, Josef Kittler |
Pattern Recognit. | 2 |
| 2000 | Region-Based Object Recognition: Pruning Multiple Representations and HypothesesabstractWe address the problem of object recognition in computer vision. We rep-resent each model and the scene in the form of Attributed Relational Graph. A multiple region representation is provided at each node of the scene ARG to increase the representation reliability. The process of matching the scene ARG against the stored models is facilitated by a novel method for identi-fying the most probable representation from among the multiple candidates. The scene and model graph matching is accomplished using probabilistic relaxation which has been modified to minimise the label clutter. The exper-imental results obtained on real data demonstrate promising performance of the proposed recognition system. 1 Alireza Ahmadyfard, Josef Kittler |
BMVC | 2 |
| 2000 | On Matching Scores for LDA-based Face VerificationabstractWe address the problem of face verification using linear discriminant anal-ysis and investigate the issue of matching score1. We establish the reason behind the success of the normalised correlation. The improved understand-ing about the role of metric then naturally leads to a novel way of measuring the distance between a probe image and a model. In extensive experimen-tal studies on the publicly available XM2VTS database2 using the Lausanne protocol3 we show that the proposed metric is consistently superior to both the Euclidean distance and normalised correlation matching scores. The ef-fect of various photometric normalisations4 on the matching scores is also investigated. 1 Josef Kittler, Yongping Li, Jiri Matas |
BMVC | 1 |
| 2000 | Object Recognition using the Invariant Pixel-Set SignatureabstractA new object recognition method, the Invariant Pixel Set Signature (IPSS), is introduced. Objects are represented with a probability density on the space of invariants computed from measurements (pixel values) inside convex hulls of n-tuples of interest points. Experimentally the method is tested on COIL– 20, a publicly available database of 72 views of 20 natural object rotating on a turntable. With a model built from a single view, recognition performance measured by the average match percentile is above 98 % for 20 degrees and above 96 % for30 degrees. For some object, 100 % first rank is achieved for all 72 views. Robustness to occlusion is shown using images with one half covered. For a small change of viewpoint (10 degrees) recognition of the occluded object is perfect. 1 Jiri Matas, J. Burianek, Josef Kittler |
BMVC | 3 |
| 2000 | Data and Decision Level Fusion of Temporal Information for Automatic Target RecognitionabstractAutomatic Target Recognition (ATR) is a demanding application that requires separation of targets from a noisy background in a sequence of images. In our previous work [5] the background was adaptively described using twodimensional filters designed by Principle Component Analysis on sampled two-dimensional image patches. Significant improvements in performance have been obtained by decision level fusion over time. In this paper we extend this idea and utilise the temporal nature of the data further to design a set of three-dimensional texture filters based on randomly sampled threedimensional image patches . We show that by virtue of data level fusion, using these new filters the true-positive rate can be increased further whilst reducing the number of false-positives. Kieron Messer, Josef Kittler |
BMVC | 2 |
| 2000 | Probabilistic PCA and ICA Subspace Mixture Models for Image SegmentationabstractHigh-dimensional data, such as images represented as points in the space spanned by their pixel values, can often be described in a significantly smaller number of dimensions than the original. One of the ways of finding lowdimensional representations is to train a mixture model of principal component analysers (PCA) on the data. However, some types of data do not fulfill the assumptions of PCA, calling for application of different subspace methods. One such a method is ICA, which has been shown in recent years to be able to find interesting basis vectors (features) in signal and image data. In this paper, a mixture model of ICA subspaces is developed similar to a mixture model of PCA subspaces proposed by others. The new algorithm is applied to a natural texture segmentation problem and is shown to give encouraging results. Dick de Ridder, Josef Kittler, Robert P. W. Duin |
BMVC | 2 |
| 2000 | Video Shot Cut Detection using Adaptive ThresholdingabstractThe performance of shot detection methods in video sequences can be improved by the use of a threshold that adapts itself to the sequence statistics. In this paper we present some new techniques for adapting the threshold. We then compare the new techniques with an existing one, leading to an improved shot detection method. Yusseri Yusoff, William J. Christmas, Josef Kittler |
BMVC | 3 |
| 2000 | Colour Image Retrieval and Object Recognition Using the Multimodal Neighbourhood Signature
Jiri Matas, Dimitri Koubaroulis, Josef Kittler |
ECCV (1) | 3 |
| 2000 | Learning Support Vectors for Face Verification and RecognitionabstractThe paper studies support vector machines (SVM) in the context of face verification and recognition. Our study supports the hypothesis that the SVM approach is able to extract the relevant discriminatory information from the training data and we present results showing superior performance in comparison with benchmark methods. However, when the representation space already captures and emphasises the discriminatory information (e.g., Fisher's linear discriminant), SVM loose their superiority. The results also indicate that the SVM are robust against changes in illumination provided these are adequately represented in the training data. The proposed system is evaluated on a large database of 295 people obtaining highly competitive results: an equal error rate of 1% for verification and a rank-one error rate of 2% for recognition (or 98% correct rank-one recognition). Kenneth Jonsson, Josef Kittler, Yongping Li, Jiri Matas |
FG | 2 |
| 2000 | Wearable face recognition aidabstractThe feasibility of realising a low cost wearable face recognition aid based on a robust correlation algorithm is investigated. The aim of the study is to determine the limiting spatial and grey level resolution of the probe and gallery images that would support successful prompting of the identity of input face images. Low spatial and grey level resolution images are obtained from good quality image data algorithmically. The tests carried out on the XM2VTS database demonstrate that robust correlation is very resilient to degradations of spatial and grey level image resolution. Correct prompts have been generated in 98% cases even for severely degraded images. C. Iordanoglou, Kenneth Jonsson, Josef Kittler, Jiri Matas |
ICASSP | 3 |
| 2000 | Improving the Performance of the Product Fusion StrategyabstractAmong existing classifier combination rules the most widely used are sum, product and vote. Although product is more directly related to the compound class posterior probability, it does not perform well. Sum, which is derived under restricting assumptions, outperforms product, especially if the class aposteriori probability estimates are subject to high levels of noise. We establish the cause of product's degraded performance and propose a method to improve it. Tests on real and synthetic data demonstrate that the modified product has a number of advantages in relation to other rules that we experiment with. Fuad M. Alkoot, Josef Kittler |
ICPR | 2 |
| 2000 | Using Gradient Information to Enhance the Progressive Probabilistic Hough TransformabstractWe look at the benefits to be gained in using gradient information to enhance the progressive probabilistic Hough transform (PPHT). It is shown how using the angle information in controlling the voting process and in assigning pixels correctly to a line, PPHT's performance can be significantly improved. The improved algorithm gives results very close to that of the standard Hough transform, but requires significantly less computation. Charles Galambos, Josef Kittler, Jiri Matas |
ICPR | 2 |
| 2000 | Coping with 3D Artifacts in Video SequencesabstractSeveral video processing techniques with applications in video compression, shot-cut detection, mosaicing and so on are based around the mapping of the positions of feature points from video frame to video frame. This in turn is often based on the assumption that for the scene under observation, a two dimensional approximation of the structure is adequate - both in terms of motion estimation and in the subsequent use of the feature correspondences obtained from the motion estimation process. This assumption is in general false. We propose a method for motion estimation based on an active mesh. In doing so we attempt to model the motion estimation not as an unconstrained motion of feature points moving independently, but as the estimation of a set of planar patches each undergoing a 3D perspective motion. In doing so we note that our method allows us to fit to 3D metrics with a high degree of accuracy. Andrew Griffin, Josef Kittler |
ICPR | 2 |
| 2000 | The Multimodal Signature Method: An Efficiency and Sensitivity StudyabstractThe multimodal neighbourhood signature (MNS) method has given acceptable results both for the colour-based image retrieval and the object recognition task. Local colour content is concisely represented by invariant features computed from neighbourhoods with multimodal colour density function. In this paper, efficiency related issues regarding the MNS algorithm are investigated. Its performance, speed, sensitivity to internal parameters and storage requirements are tested on a standard colour object recognition experiment. Very good recognition rate (99.9%) was achieved in real time. The MNS signature size is a few hundred bytes on average, an important property for retrieval from large databases. The algorithmic complexity of signature computation and matching are analysed and efficient implementations are proposed. Dimitri Koubaroulis, Jiri Matas, Josef Kittler |
ICPR | 3 |
| 2000 | Defining Quantization Strategies and a Perceptual Similarity Measure for Texture-Based Annotation and RetrievalabstractWe introduce an approach for texture-based annotation and retrieval. Given the outputs of 12 Gabor filters, we derive a texture feature space where the sensitivity of the features to illumination changes is attenuated by a suitable normalisation. We then annotate images by defining and selecting codes representing the quantised levels of the texture features appearing in each image. The annotations are stored in a hash table for retrieval efficiency. Ranking schemes are proposed to order the images retrieved at query time. In particular, we use results from psychological studies on the human perception of similarity to formulate a similarity measure. The choice of quantisation of the texture feature space can influence the accuracy of the retrieval. We compared several quantisation schemes in retrieval experiments involving texture images. We found that a uniform quantisation and a quantisation heuristically taking the variance of the texture features into account lead to the best retrieval performance. Barbara Levienaise-Obadia, Josef Kittler, William J. Christmas |
ICPR | 2 |
| 2000 | Comparison of Face Verification Results on the XM2VTS DatabaseabstractPresents results of the face verification contest that was organized in conjunction with International Conference on Pattern Recognition 2000. Participants had to use identical data sets from a large, publicly available multimodal database XM2VTSDB. Training and evaluation was carried out according to an a priori known protocol. Verification results of all tested algorithms have been collected and made public on the XM2VTSDB website, facilitating large scale experiments on classifier combination and fusion. Tested methods included, among others, representatives of the most common approaches to face verification -elastic graph matching, Fisher's linear discriminant and support vector machines. Jiri Matas, Miroslav Hamouz, Kenneth Jonsson, Josef Kittler, Yongping Li, Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas, Teewoon Tan, Hong Yan 0001, Fabrizio Smeraldi, N. Capdevielle, Wulfram Gerstner, Yousri Abdeljaoued, Josef Bigün, Souheil Ben Yacoub, Eddy Mayoraz |
ICPR | 4 |
| 2000 | Fast Unit Selection Algorithm for Neural Network DesignabstractIn this paper a fast neural network pruning algorithm is presented which is based on an analysis of the weights in a trained network. We demonstrate that this technique selects a lean architecture whilst experiencing no corresponding degradation in performance. Our unit selection algorithm is compared to a state of the art network pruning algorithm taken from the literature and is found to offer several advantages, i.e., its simplicity, its speed and the ability to select the leanest architecture. Kieron Messer, Josef Kittler |
ICPR | 2 |
| 2000 | A Comparative Study of Different Segmentation Approaches for Audio Track IndexingabstractThis paper investigate methods for test-independent speaker-based segmentation of audio track. The focus is on online segmentation of the data. We compare the different criteria based on distance measures and linear discriminant analysis (LDA) for automatic segmentation of audio track. Experiments were performed on a TV panel show using various sets of input features. This series includes speech, music and speech of multiple speakers talking simultaneously. The experimental results show that Mahalonabis distance criterion gives better performance than other distance measures and LDA. Medha Pandit, Josef Kittler, Yongping Li, Edward H. S. Chilton |
ICPR | 2 |
| 2000 | The Adaptive Subspace Map for Texture SegmentationabstractA nonlinear mixture-of-subspaces model is proposed to describe images. Images or image patches, when translated, rotated or scaled, lie in low-dimensional subspaces of the high-dimensional space spanned by the grey values. These manifolds can locally be approximated by a linear subspace. The adaptive subspace map is a method to learn such a mixture-of-subspaces from the data. Due to its general nature, various clustering and subspace-finding algorithms can be used. In the paper, two clustering algorithms are compared in an application to some texture segmentation problems. It is shown to compare well to a standard Gabor filter bank approach. Dick de Ridder, Josef Kittler, Olaf Lemmers, Robert P. W. Duin |
ICPR | 2 |
| 2000 | Robust Detection of Lines Using the Progressive Probabilistic Hough Transform
Jiri Matas, Charles Galambos, Josef Kittler |
Comput. Vis. Image Underst. | 3 |
| 2000 | Enhancing CSS-based shape retrieval for objects with shallow concavities
Sadegh Abbasi, Farzin Mokhtarian, Josef Kittler |
Image Vis. Comput. | 3 |
| 2000 | On the generalised stock-cutting problem
Nikos Georgis, Maria Petrou, Josef Kittler |
Mach. Vis. Appl. | 3 |
| 2000 | Probabilistic relaxation and the Hough transform
Josef Kittler |
Pattern Recognit. | 1 |
| 2000 | Combining multiple classifiers by averaging or by multiplying?
David M. J. Tax, Martijn van Breukelen, Robert P. W. Duin, Josef Kittler |
Pattern Recognit. | 4 |
| 2000 | Edge postprocessing using probabilistic relaxationabstractIn this paper, we develop the theory of probabilistic relaxation when the objects to be labeled are arranged in a rectangular grid with known adjacency relations. In this case a dictionary of permissible label configurations is available. The novelty of this work lies in the inclusion of measurements concerning binary relations between the objects to be labeled. These are compared with the corresponding binary relations between the nodes of the dictionary. This way, one of the major objections to probabilistic relaxation, namely, the disregard of the data after the initial assignment of probabilities, is removed. The theory we develop is demonstrated by applying it to the problem of edge relaxation labeling. We show that the inclusion of binary relations greatly improves the performance of algorithms of this kind and compare our approach with previously developed dictionary based approaches, both theoretically and experimentally. Also, a comparison with other edge-postprocessing strategies is provided. Petros Papachristou, Maria Petrou, Josef Kittler |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 1999 | Support Vector Machines for Face AuthenticationabstractAbstract We present an extensive study of the support vector machine (SVM) sensitivity to various processing steps in the context of face authentication. In particular, we evaluate the impact of the representation space and photometric normalisation technique on the SVM performance. Our study supports the hypothesis that the SVM approach is able to extract the relevant discriminatory information from the training data. We believe that this is the main reason for its superior performance over benchmark methods (e.g. the eigenface technique). However, when the representation space already captures and emphasises the discriminatory information content (e.g. the fisherface method), the SVMs cease to be superior to the benchmark techniques. The SVM performance evaluation is carried out on a large face database containing 295 subjects. Kenneth Jonsson, Josef Kittler, Yongping Li, Jiri Matas |
BMVC | 2 |
| 1999 | Adaptive Texture Representation Methods for Automatic Target RecognitionabstractAutomatic Target Recognition (ATR) is a demanding application that re-quires separation of targets from a noisy background in a sequence of im-ages. In this paper, two adaptive methods for describing such a background are proposed which are based on Principal and Independent Component Ana-lysis of sampled image patches. Coupled together with feature selection and outlier detection techniques they enable the ATR system to adapt to certain backgrounds and identify non-standard elements in the images as targets. The methods proposed are compared with a standard wavelet-based approach and are shown to perform somewhat better on a difficult image sequence. 1 Kieron Messer, Dick de Ridder, Josef Kittler |
BMVC | 3 |
| 1999 | Effective Implementation of Linear Discriminant Analysis for Face Recognition and Verification
Yongping Li, Josef Kittler, Jiri Matas |
CAIP | 2 |
| 1999 | Audio-Visual Person VerificationabstractIn this paper we investigate benefits of classifier combination (fusion) for a multimodal system for personal identity verification. The system uses frontal face images and speech. We show that a sophisticated fusion strategy enables the system to outperform its facial and vocal modules when taken seperately. We show that both trained linear weighted schemes and fusion by Support Vector Machine classifier leads to a significant reduction of total error rates. The complete system is tested on data from a publicly available audio-visual database (XM2VTS, 295 subjects) according to a published protocol. Souheil Ben Yacoub, Jürgen Lüttin, Kenneth Jonsson, Jiri Matas, Josef Kittler |
CVPR | 5 |
| 1999 | Progressive Probabilistic Hough Transform for Line DetectionabstractWe present a novel Hough Transform algorithm referred to as Progressive Probabilistic Hough Transform (PPHT). Unlike the Probabilistic HT where Standard HT is performed on a pre-selected fraction of input points, PPHT minimises the amount of computation needed to detect lines by exploiting the difference an the fraction of votes needed to detect reliably lines with different numbers of supporting points. The fraction of points used for voting need not be specified ad hoc or using a priori knowledge, as in the probabilistic HT; it is a function of the inherent complexity of the input data. The algorithm is ideally suited for real-time applications with a fixed amount of available processing time, since voting and line detection is interleaved. The most salient features are likely to be detected first. Experiments show that in many circumstances PPHT has advantages over the Standard HT. Charles Galambos, Josef Kittler, Jiri Matas |
CVPR | 2 |