VLDB 2026 Research / reviewers in the wild / expert
Muhammad Awais 0001
dblp:80/3639-1
· DBLP profile ↗
65ranked-venue papers
3as first author
50since 2021 · last 2026
0000-0002-1122-0709ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 3 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 2 first-author · 26 since 2021Security and privacy · 2Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DAT3: Dual-Teacher Topology Adversarial Training for Defending Against Adversarial Attacks
Huaxing Feng, He-Feng Yin, Sara Atito Ali Ahmed, Muhammad Awais 0001 |
ICPR (11) | 6 |
| 2026 | Channel-Aware Probing for Multi-channel Imaging
Umar Marikkar, Syed Sameed Husain, Muhammad Awais 0001, Sara Atito Ali Ahmed |
ICPR (7) | 3 |
| 2026 | Gaze-Guided Multimodal LLMs for Social Scene Understanding
Shayan Nasiriboukani, Muhammad Awais 0001, Sara Atito Ali Ahmed |
ICPR (11) | 2 |
| 2026 | CoZSR-VAD: Contextual Zero-Shot Reasoning for Video Anomaly Detection
Mohd Ubaid Wani, Sara Atito Ali Ahmed, Srinivasa Rao Nandam, Josef Kittler, Muhammad Awais 0001 |
ICPR (12) | 5 |
| 2026 | A Color Information Driven Collaborative Training of Dual Task Parallel Network for Visible and Thermal Infrared Image Fusion and Saliency Object Detection
Zeyang Zhang 0002, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Muhammad Awais 0001, Josef Kittler |
Int. J. Comput. Vis. | 5 |
| 2026 | Probabilistically Aligned View-Unaligned Clustering With Adaptive Template SelectionabstractIn most existing multi-view modeling scenarios, cross-view correspondence (CVC) between instances of the same target from different views, like paired image-text data, is a crucial prerequisite for effortlessly deriving a consistent representation. Nevertheless, this premise is frequently compromised in certain applications, where each view is organized and transmitted independently, resulting in the view-unaligned problem (VuP). Restoring CVC of unaligned multi-view data is a challenging and highly demanding task that has received limited attention from the research community. To tackle this practical challenge, we propose to integrate the permutation derivation procedure into the bipartite graph paradigm for view-unaligned clustering, termed Probabilistically Aligned View-unaligned Clustering with Adaptive Template Selection (PAVuC-ATS). Specifically, we learn consistent anchors and view-specific graphs by the bipartite graph, and derive permutations applied to the unaligned graphs by reformulating the alignment between two latent representations as a 2-step transition of a Markov chain with adaptive template selection, thereby achieving the probabilistic alignment. The convergence of the resultant optimization problem is validated both experimentally and theoretically. Extensive experiments on six benchmark datasets demonstrate the superiority of the proposed PAVuC-ATS over the baseline methods. Wenhua Dong, Xiaojun Wu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | How Effectively Can Large Language Models Connect SNP Variants and ECG Phenotypes for Cardiovascular Risk Prediction?abstractCardiovascular disease (CVD) prediction remains a tremendous challenge due to its multifactorial etiology and global burden of morbidity and mortality. Despite the growing availability of genomic and electrophysiological data, extracting biologically meaningful insights from such high-dimensional, noisy, and sparsely annotated datasets remains a non-trivial task. Recently, LLMs has been applied effectively to predict structural variations in biological sequences. In this work, we explore the potential of fine-tuned LLMs to predict cardiac diseases and SNPs potentially leading to CVD risk using genetic markers derived from high-throughput genomic profiling. We investigate the effect of genetic patterns associated with cardiac conditions and evaluate how LLMs can learn latent biological relationships from structured and semi-structured genomic data obtained by mapping genetic aspects that are inherited from the family tree. By framing the problem as a Chain of Thought (CoT) reasoning task, the models are prompted to generate disease labels and articulate informed clinical deductions across diverse patient profiles and phenotypes. The findings highlight the promise of LLMs in contributing to early detection, risk assessment, and ultimately, the advancement of personalized medicine in cardiac care. Niranjana Arun Menon, Iqra Farooq, Yulong Li 0002, Yutong Xie 0001, Muhammad Awais 0001, Muhammad Imran Razzak |
BIBM | 6 |
| 2025 | One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image FusionabstractAdvanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction through pixel-level supervision. This new paradigm provides strong guidance for unsupervised multimodal fusion without relying on abstract semantics, enhancing task-shared feature learning for broader applicability. Owning to the hybrid image features and enhanced universal representations, the proposed GIFNet supports diverse fusion tasks, achieving high performance across both seen and unseen scenarios with a single model. Uniquely, experimental results reveal that our framework also supports single-modality enhancement, offering superior flexibility for practical applications. Our code will be available at https://github.com/AWCXV/GIFNet. Chunyang Cheng, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Zhangyong Tang, Hui Li 0037, Zeyang Zhang 0002, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
CVPR | 9 |
| 2025 | Text Augmented Correlation Transformer For Few-shot Classification & SegmentationabstractFoundation models like CLIP and ALIGN have transformed few-shot and zero-shot vision applications by fusing visual and textual data, yet the integrative few-shot classification and segmentation (FS-CS) task primarily leverages visual cues, overlooking the potential of textual support. In FS-CS scenarios, ambiguous object boundaries and overlapping classes often hinder model performance, as limited visual data struggles to fully capture high-level semantics. To bridge this gap, we present a novel multi-modal FS-CS framework that integrates textual cues into support data, facilitating enhanced semantic disambiguation and fine-grained segmentation. Our approach first investigates the unique contributions of exclusive text-based support, using only class labels to achieve FS-CS. This strategy alone achieves performance competitive with vision-only methods on FS-CS tasks, underscoring the power of textual cues in few-shot learning. Building on this, we introduce a dualmodal prediction mechanism that synthesizes insights from both textual and visual support sets, yielding robust multimodal predictions. This integration significantly elevates FS-CS performance, with classification and segmentation improvements of +3.7/6.6% (1-way 1-shot) and +8.0/6.5% (2-way 1-shot) on COCO-20i, and +2.2/3.8% (1-way 1shot) and +4.3/4.0% (2-way 1-shot) on Pascal-5i. Additionally, in weakly supervised FS-CS settings, our method surpasses visual-only benchmarks using textual support exclusively, further enhanced by our dual-modal predictions. By rethinking the role of text in FS-CS, our work establishes new benchmarks for multi-modal few-shot learning and demonstrates the efficacy of textual cues for improving model generalization and segmentation accuracy. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
CVPR | 5 |
| 2025 | Enhanced Weakly Supervised Few-shot Classification & SegmentationabstractThe emergence of vision-language foundation models has enabled the integration of textual information into vision-based applications. However, in few-shot classification and segmentation (FS-CS), this potential remains underutilised. Commonly, self-supervised vision models have been employed, particularly in weakly-supervised scenarios, to generate pseudo-segmentation masks, as ground truth masks are typically unavailable and only target classification is provided. Despite their success, such models find it difficult to capture accurate semantics when compared to vision-language models. To address this limitation, we propose a novel FS-CS approach that leverages the rich semantic alignment of vision-language models to generate more precise pseudo ground-truth masks. While current vision-language models excel in global visual-text alignment, they struggle with finer, patch-level alignment, which is crucial for detailed segmentation tasks. To overcome this, we introduce a method that enhances patch-level alignment without requiring additional training. In addition, existing FS-CS frameworks typically lacks multi-scale information, limiting their ability to capture fine and coarse features simultaneously. To overcome this, we incorporate a module based on atrous convolutions to inject multi-scale information into the feature maps. Together, these contributions - text enhanced pseudo-mask generation and improved multi-scale feature representation - significantly boost the performance of our model in weakly-supervised settings, surpassing state-of-the-art methods and demonstrating the importance of integrating multi-modal information for robust FS-CS solutions. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
ICASSP | 5 |
| 2025 | SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic SoundscapesabstractSelf-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the self-supervised pre-training has sufficiently equipped them to handle real-world audio. However, a critical question remains: how well do these models actually perform in real-world conditions, where audio is typically polyphonic and complex, involving multiple overlapping sound sources? Current audio self-supervised learning (SSL) methods are often benchmarked on datasets predominantly featuring monophonic audio, such as environmental sounds, and speech. As a result, the ability of SSL models to generalize to polyphonic audio, a common characteristic in natural scenarios, remains underexplored. This limitation raises concerns about the practical robustness of SSL models in more realistic audio settings. To address this gap, we introduce Self-Supervised Learning from Audio Mixtures (SSLAM), a novel direction in audio SSL research, designed to improve the model’s ability to learn from polyphonic data while maintaining strong performance on monophonic data. We thoroughly evaluate SSLAM on standard audio SSL benchmark datasets which are predominantly monophonic and conduct a comprehensive comparative analysis against state-of-the-art (SOTA) methods using a range of high-quality, publicly available polyphonic datasets. SSLAM not only improves model performance on polyphonic audio, but also maintains or exceeds performance on standard audio SSL benchmarks. Notably, it achieves up to a 3.9% improvement on the AudioSet-2M(AS-2M), reaching a mean average precision (mAP) of 50.2. For polyphonic datasets, SSLAM sets new SOTA in both linear evaluation and fine-tuning regimes with performance improvements of up to 9.1%(mAP). These results demonstrate SSLAM's effectiveness in both polyphonic and monophonic soundscapes, significantly enhancing the performance of audio SSL models. Code and pre-trained models are available at https://github.com/ta012/SSLAM. Tony Alex, Sara Atito Ali Ahmed, Armin Mustafa, Muhammad Awais 0001, Philip J. B. Jackson |
ICLR | 4 |
| 2025 | CG-SSL: Concept-Guided Self-Supervised LearningabstractHumans understand visual scenes by first capturing a global impression and then refining this understanding into distinct, object-like components. Inspired by this process, we introduce \textbf{C}oncept-\textbf{G}uided \textbf{S}elf-\textbf{S}upervised \textbf{L}earning (CG-SSL), a novel framework that brings structure and interpretability to representation learning through a curriculum of three training phases: (1) global scene encoding, (2) discovery of visual concepts via tokenised cross-attention, and (3) alignment of these concepts across views.
Unlike traditional SSL methods, which simply enforce similarity between multiple augmented views of the same image, CG-SSL accounts for the fact that these views may highlight different parts of an object or scene. To address this, our method establishes explicit correspondences between views and aligns the representations of meaningful image regions. At its core, CG-SSL augments standard SSL with a lightweight decoder that learns and refines concept tokens via cross-attention with patch features. The concept tokens are trained using masked concept distillation and a feature-space reconstruction objective. A final alignment stage enforces view consistency by geometrically matching concept regions under heavy augmentation, enabling more compact, robust, and disentangled representations of scene regions.
Across multiple backbone sizes, CG-SSL achieves state-of-the-art results on image segmentation benchmarks using $k$-NN and linear probes, substantially outperforming prior methods and approaching, or even surpassing, the performance of leading SSL models trained on over $100\times$ more data. Code and pretrained models will be released. Sara Atito Ali Ahmed, Josef Kittler, Muhammad Imran Razzak, Muhammad Awais 0001 |
NeurIPS | 4 |
| 2025 | Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM CaptioningabstractDespite significant advances in inference-time search for vision–language models (VLMs), existing approaches remain both computationally expensive and prone to unpenalized, low-confidence generations which often lead to persistent hallucinations. We introduce \textbf{Value-guided Inference with Margin-based Reward (ViMaR)}, a two-stage inference framework that improves both efficiency and output fidelity by combining a temporal-difference value model with a margin-aware reward adjustment. In the first stage, we perform a single pass to identify the highest-value caption among diverse candidates. In the second stage, we selectively refine only those segments that were overlooked or exhibit weak visual grounding, thereby eliminating frequently rewarded evaluations. A calibrated margin-based penalty discourages low-confidence continuations while preserving descriptive richness. Extensive experiments across multiple VLM architectures demonstrate that ViMaR generates captions that are significantly more reliable, factually accurate, detailed, and explanatory, while achieving over 4$\times$ speedup compared to existing value-guided methods. Specifically, we show that ViMaR trained solely on LLaVA Mistral-7B \textit{generalizes effectively to guide decoding in stronger unseen models}. To further validate this, we adapt ViMaR to steer generation in both LLaVA-OneVision-Qwen2-7B and Qwen2.5-VL-3B, leading to consistent improvements in caption quality and demonstrating robust cross-model guidance. This cross-model generalization highlights ViMaR's flexibility and modularity, positioning it as a scalable and transferable inference-time decoding strategy. Furthermore, when ViMaR-generated captions are used for self-training, the underlying models achieve substantial gains across a broad suite of visual comprehension benchmarks, underscoring the potential of fast, accurate, and self-improving VLM pipelines.
Code: https://github.com/ankan8145/ViMaR Ankan Deria, Adinath Madhavrao Dukre, Sara Atito Ali Ahmed, Sudipta Roy 0002, Muhammad Awais 0001, Muhammad Haris Khan, Muhammad Imran Razzak |
NeurIPS | 6 |
| 2025 | Investigating Self-Supervised Methods for Label-Efficient LearningabstractAbstract Vision transformers combined with self-supervised learning have enabled the development of models which scale across large datasets for several downstream tasks, including classification, segmentation, and detection. However, the potential of these models for low-shot learning across several downstream tasks remains largely under explored. In this work, we conduct a systematic examination of different self-supervised pretext tasks, namely contrastive learning, clustering, and masked image modelling, to assess their low-shot capabilities by comparing different pretrained models. In addition, we explore the impact of various collapse avoidance techniques, such as centring, ME-MAX, and sinkhorn, on these downstream tasks. Based on our detailed analysis, we introduce a framework that combines mask image modelling and clustering as pretext tasks. This framework demonstrates superior performance across all examined low-shot downstream tasks, including multi-class classification, multi-label classification and semantic segmentation. Furthermore, when testing the model on large-scale datasets, we show performance gains in various tasks. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | Correction: Investigating Self-Supervised Methods for Label-Efficient Learning
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | Pure anomaly detection via self-supervised deep metric learning with adaptive marginabstractWe address the problem of anomaly detection (AD) by a deep network pretrained using self-supervised learning for an auxiliary geometric transformation (GT) classification task. Our key contribution is a novel loss function that augments the standard cross-entropy by an additional term that plays a significant role in the later stages of self-supervised learning. The proposed enabling innovation is a triplet centre loss with an adaptive margin and a learnable metric, which relentlessly drives the GT classes to exhibit continuously improving compactness and inter-class separation. The pretrained network is finetuned for the downstream task using non-anomalous data only, and a GT model for the data is constructed. Anomalies are detected by fusing the output of several decision functions defined using the learnt GT class model. In contrast to the majority of existing methods, our approach strictly adheres to the pure AD design philosophy, which relies on the use of purely non-anomalous data for the design. Extensive experiments on four publicly available AD datasets demonstrate the effectiveness of the proposed contributions and lead to significant performance gains compared to the state-of-the-art (1.8% on F-MNIST, 1.0% on CIFAR-10, 1.2% on CIFAR-100, and 1.7% on CatVsDog). https://github.com/12sf12/Deep-Anomaly-Detection • A three-stage pure anomaly detection framework using self-supervised learning. • A novel loss addressing margin and metric selection issues in triplet-based losses. • We propose a method to compute adaptive margins per sample in each mini-batch. • Experiments show the proposed method outperforms SOTA across all datasets. Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler |
Neurocomputing | 2 |
| 2025 | Which images can be effectively learnt from self-supervised learning?
Michalis Lazarou, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
Pattern Recognit. Lett. | 3 |
| 2024 | DTF-AT: Decoupled Time-Frequency Audio Transformer for Event ClassificationabstractConvolutional neural networks (CNNs) and Transformer-based networks have recently enjoyed significant attention for various audio classification and tagging tasks following their wide adoption in the computer vision domain. Despite the difference in information distribution between audio spectrograms and natural images, there has been limited exploration of effective information retrieval from spectrograms using domain-specific layers tailored for the audio domain. In this paper, we leverage the power of the Multi-Axis Vision Transformer (MaxViT) to create DTF-AT (Decoupled Time-Frequency Audio Transformer) that facilitates interactions across time, frequency, spatial, and channel dimensions. The proposed DTF-AT architecture is rigorously evaluated across diverse audio and speech classification tasks, consistently establishing new benchmarks for state-of-the-art (SOTA) performance. Notably, on the challenging AudioSet 2M classification task, our approach demonstrates a substantial improvement of 4.4% when the model is trained from scratch and 3.2% when the model is initialised from ImageNet-1K pretrained weights. In addition, we present comprehensive ablation studies to investigate the impact and efficacy of our proposed approach. The codebase and pretrained weights are available on https://github.com/ta012/DTFAT.git Tony Alex, Armin Mustafa, Muhammad Awais 0001, Philip J. B. Jackson |
AAAI | 4 |
| 2024 | SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action RecognitionabstractContrastive learning has achieved great success in skeleton-based action recognition. However, most existing approaches encode the skeleton sequences as entangled spatiotemporal representations and confine the contrasts to the same level of representation. Instead, this paper introduces a novel contrastive learning framework, namely Spatiotemporal Clues Disentanglement Network (SCD-Net). Specifically, we integrate the decoupling module with a feature extractor to derive explicit clues from spatial and temporal domains respectively. As for the training of SCD-Net, with a constructed global anchor, we encourage the interaction between the anchor and extracted clues. Further, we propose a new masking strategy with structural constraints to strengthen the contextual associations, leveraging the latest development from masked image modelling into the proposed SCD-Net. We conduct extensive evaluations on the NTU-RGB+D (60&120) and PKU-MMD (I&II) datasets, covering various downstream tasks such as action recognition, action retrieval, transfer learning, and semi-supervised learning. The experimental results demonstrate the effectiveness of our method, which outperforms the existing state-of-the-art (SOTA) approaches significantly. Our code and supplementary material can be found at https://github.com/cong-wu/SCD-Net. Cong Wu 0006, Xiaojun Wu 0001, Josef Kittler, Tianyang Xu 0001, Muhammad Awais 0001, Zhenhua Feng 0001 |
AAAI | 6 |
| 2024 | Pseudo Labelling for Enhanced Masked Auto Encoders
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
BMVC | 5 |
| 2024 | Enhancing Radiology Report Generation: The Impact of Locally Grounded Vision and Language Training
Sergio Sánchez Santiesteban, Muhammad Awais 0001, Yi-Zhe Song, Josef Kittler |
BMVC | 2 |
| 2024 | C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
Rongchang Li 0001, Zhenhua Feng 0001, Tianyang Xu 0001, Linze Li 0002, Xiaojun Wu 0001, Muhammad Awais 0001, Sara Atito Ali Ahmed, Josef Kittler |
ECCV (38) | 6 |
| 2024 | Max-AST: Combining Convolution, Local and Global Self-Attentions for Audio Event ClassificationabstractIn the domain of audio transformer architectures, prior research has extensively investigated isotropic architectures that capture the global context through full self-attention and hierarchical architectures that progressively transition from local to global context utilising hierarchical structures with convolutions or window-based attention. However, the idea of imbuing each individual block with both local and global contexts, thereby creating a hybrid transformer block, remains relatively under-explored in the field.To facilitate this exploration, we introduce Multi Axis Audio Spectrogram Transformer (Max-AST), an adaptation of MaxViT to the audio domain. Our approach leverages convolution, local window-attention, and global grid-attention in all the transformer blocks. The proposed model excels in efficiency compared to prior methods and consistently outperforms state-of-the-art techniques, achieving significant gains of up to 2.6% on the AudioSet full set. Further, we performed detailed ablations to analyse the impact of each of these components on audio feature learning. The source code is available at https://github.com/ta012/MaxAST.git Tony Alex, Armin Mustafa, Muhammad Awais 0001, Philip J. B. Jackson |
ICASSP | 4 |
| 2024 | Improved Image Captioning Via Knowledge Graph-Augmented ModelsabstractMultimodal foundation models, pre-trained on large-scale data, effectively capture vast amounts of factual and commonsense knowledge. However, these models store all their knowledge within their parameters, requiring increasingly larger models and training data to capture more knowledge. To address this limitation and achieve a more scalable and modular integration of knowledge, we propose a novel knowledge graph-augmented multimodal model. This approach enables a base multimodal model to access pertinent information from an external knowledge graph. Our methodology leverages existing general domain knowledge to facilitate vision-language pre-training using paired images and text descriptions. We conduct comprehensive evaluations demonstrating that our model outperforms state-of-the-art models and yields comparable results to much larger models trained on more extensive datasets. Notably, our model reached a 145 Cider score on MS COCO Captions using only 2.9 million samples, outperforming a 1.4B parameter model by 1.7% despite having 11 times fewer parameters. Sergio Sánchez Santiesteban, Sara Atito Ali Ahmed, Muhammad Awais 0001, Yi-Zhe Song, Josef Kittler |
ICASSP | 3 |
| 2024 | SS-CXR: Self-Supervised Pretraining Using Chest X-Rays Towards A Domain Specific Foundation ModelabstractChest X-rays (CXRs) are widely used imaging modality for the diagnosis and prognosis of lung disease. There is a large body of work where machine learning algorithms are developed for specific tasks. However, the traditional diagnostic tool design methods based on supervised learning are burdened by the need to provide training data annotation, which should be of good quality for better clinical outcomes. Here, we propose an alternative solution, a new self-supervised paradigm, where a general representation from CXRs is learned using a group-masked self-supervised framework. The pre-trained model is then fine-tuned for domain-specific tasks such as covid-19, pneumonia detection, and general health screening. We show that the same pre-training can be used for the lung segmentation task. Our proposed paradigm shows robust performance in multiple downstream tasks which demonstrates the success of the pre-training. Moreover, the performance of the pre-trained models on data with significant drift during test time proves the learning of a better generic representation. The methods are further validated by covid-19 detection in a unique small-scale pediatric data set. The performance gain ($\sim 25 \%$) is significant when compared to a supervised transformer-based method. This adds credence to the strength and reliability of our proposed framework and pre-training strategy. Syed Muhammad Anwar, Abhijeet Parida, Sara Atito Ali Ahmed, Muhammad Awais 0001, Gustavo Nino, Josef Kittler, Marius George Linguraru |
ICIP | 4 |
| 2024 | Investigating Self-Supervised Methods for Label-Efficient LearningabstractVision transformers combined with self-supervised learning have enabled the development of models which scale across large datasets for several downstream tasks like classification, segmentation and detection. The low-shot learning capability of these models, across several low-shot downstream tasks, has been largely under explored. We perform a system level study of different self supervised pretext tasks, namely contrastive learning, clustering, and masked image modelling for their low-shot capabilities by comparing the pretrained models. In addition we also study the effects of collapse avoidance methods, namely centring, ME-MAX, sinkhorn, on these downstream tasks. Based on our detailed analysis, we introduce a framework involving both mask image modelling and clustering as pretext tasks, which performs better across all low-shot downstream tasks, including multi-class classification, multi-label classification and semantic segmentation. Furthermore, when testing the model on full scale datasets, we show performance gains in multi-class classification, multi-label classification and semantic segmentation. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
ICIP | 5 |
| 2024 | Masked Momentum Contrastive Learning for Semantic Understanding by ObservationabstractLarge language models (LLMs) have shown excellent performance in zero-shot learning using natural language prompts. However, in the domain of computer vision (CV), the paradigm of pretraining followed by finetuning remains dominant. The aim of this study is to reduce this gap by utilizing the capability of Self-Supervised Learning (SSL) in semantic understanding for zero-shot segmentation, without relying on human-provided labels or vision-language supervision. We introduce a novel evaluation framework that employs visual prompts, including a threshold and a query patch. This framework evaluates the ability of SSL models to derive concepts from observational data. Through this evaluation, we identify the strengths and limitations of SSL models in understanding semantics. Building on the insights from various SSL methods, we further propose the MMC approach to enhance the representations for objects, which integrates Masked image modeling, Momentum-based self-distillation, and global Contrastive learning. MMC achieves a better balance between the inter-object discriminability and the intra-object compactness of learned features. Our experiments on COCO, DAVIS-2017, PASCAL VOC, and ADE20K demonstrate outstanding performance of MMC’s representations. Jiantao Wu, Shentong Mo, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Syed Sameed Husain, Muhammad Awais 0001 |
ICIP | 7 |
| 2024 | StableTalk: Advancing Audio-to-Talking Face Generation with Stable Diffusion and Vision Transformer
Fatemeh Nazarieh, Josef Kittler, Muhammad Awais 0001, Diptesh Kanojia, Zhenhua Feng 0001 |
ICPR (6) | 3 |
| 2024 | View-shuffled clustering via the modified Hungarian algorithm
Wenhua Dong, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
Neural Networks | 6 |
| 2024 | RAgE: Robust Age Estimation Through Subject Anchoring With Consistency RegularisationabstractModern facial age estimation systems can achieve high accuracy when training and test datasets are identically distributed and captured under similar conditions. However, domain shifts in data, encountered in practice, lead to a sharp drop in accuracy of most existing age estimation algorithms. In this article, we propose a novel method, namely RAgE, to improve the robustness and reduce the uncertainty of age estimates by leveraging unlabelled data through a subject anchoring strategy and a novel consistency regularisation term. First, we propose an similarity-preserving pseudo-labelling algorithm by which the model generates pseudo-labels for a cohort of unlabelled images belonging to the same subject, while taking into account the similarity among age labels. In order to improve the robustness of the system, a consistency regularisation term is then used to simultaneously encourage the model to produce invariant outputs for the images in the cohort with respect to an anchor image. We propose a novel consistency regularisation term the noise-tolerant property of which effectively mitigates the so-called confirmation bias caused by incorrect pseudo-labels. Experiments on multiple benchmark ageing datasets demonstrate substantial improvements over the state-of-the-art methods and robustness to confounding external factors, including subject's head pose, illumination variation and appearance of expression in the face image. Ali Akbari 0003, Muhammad Awais 0001, Soroush Fatemifar, Syed Safwan Khalid, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Self-supervised learning for RGB-D object tracking
Xuefeng Zhu 0003, Tianyang Xu 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Xiaojun Wu 0001, Zhenhua Feng 0001, Josef Kittler |
Pattern Recognit. | 4 |
| 2024 | ASiT: Local-Global Audio Spectrogram Vision Transformer for Event ClassificationabstractTransformers, which were originally developed for natural language processing, have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships. Constrained by the data hungry nature of transformers and the limited amount of labelled data, most transformer-based models for audio tasks are finetuned from ImageNet pretrained models, despite the huge gap between the domain of natural images and audio. This has motivated the research in self-supervised pretraining of audio transformers, which reduces the dependency on large amounts of labeled data and focuses on extracting concise representations of audio spectrograms. In this paper, we proposeLocal-GlobalAudioSpectrogram vIsionTransformer, namely ASiT, a novel self-supervised learning framework that captures local and global contextual information by employing group masked model learning and self-distillation. We evaluate our pretrained models on both audio and speech classification tasks, including audio event classification, keyword spotting, and speaker identification. We further conduct comprehensive ablation studies, including evaluations of different pretraining strategies. The proposed ASiT framework significantly boosts the performance on all tasks and sets a new state-of-the-art performance in five audio and speech classification tasks, outperforming recent methods, including the approaches that use additional datasets for pretraining. Sara Atito Ali Ahmed, Muhammad Awais 0001, Wenwu Wang 0001, Mark D. Plumbley, Josef Kittler |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | A Survey of Cross-Modal Visual Content GenerationabstractCross-modal content generation has become very popular in recent years. To generate high-quality and realistic content, a variety of methods have been proposed. Among these approaches, visual content generation has attracted significant attention from academia and industry due to its vast potential in various applications. This survey provides an overview of recent advances in visual content generation conditioned on other modalities, such as text, audio, speech, and music, with a focus on their key contributions to the community. In addition, we summarize the existing publicly available datasets that can be used for training and benchmarking cross-modal visual content generation models. We provide an in-depth exploration of the datasets used for audio-to-visual content generation, filling a gap in the existing literature. Various evaluation metrics are also introduced along with the datasets. Furthermore, we discuss the challenges and limitations encountered in the area, such as modality alignment and semantic coherence. Last, we outline possible future directions for synthesizing visual content from other modalities including the exploration of new modalities, and the development of multi-task multi-modal networks. This survey serves as a resource for researchers interested in quickly gaining insights into this burgeoning field. Fatemeh Nazarieh, Zhenhua Feng 0001, Muhammad Awais 0001, Wenwu Wang 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | One-pass View-unaligned ClusteringabstractGiven a set of multi-view instances, the prevailing assumption in most existing clustering approaches is that they are complete and exhibit cross-view alignment. However, this assumption is often unrealistic. In such scenarios, it could be satisfied at the cost of data pre-processing, but this would be complex and inconsistent with practical applications. Therefore, developing more effective solutions for the View-unaligned Problem (VuP) is highly desirable. Several pioneering works have tackled the partially VuP, yet handling fully VuP remains a challenge due to the reliance on partially pre-aligned instances. In this paper, we propose One-pass View-unaligned Clustering (OpVuC) that simultaneously aligns and clusters instances in a unified framework. Specifically, we alig shuffled instances with a selected template using an innovative global-local alignment scheme based on the notion of geometric invariance and separate the fully aligned instances using a relaxed$k$-means algorithm. The proposed OpVuC method can handle VuP at any alignment level without requiring any pre-aligned instances. Extensive experiments conducted on several benchmark datasets demonstrate the effectiveness and merits of the proposed OpVuC method. Wenhua Dong, Xiaojun Wu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
IEEE Trans. Multim. | 5 |
| 2023 | Ada2NPT: An Adaptive Nearest Proxies Triplet Loss for Attribute-Aware Face Recognition with Adaptively Compacted Feature Learning
Lei Ju 0005, Zhenhua Feng 0001, Muhammad Awais 0001, Josef Kittler |
ACML | 3 |
| 2023 | Variational Autoencoders with Decremental Information Bottleneck for Disentanglement
Jiantao Wu, Shentong Mo, Xingshen Zhang, Muhammad Awais 0001, Zhenhua Feng 0001, Lin Wang 0004 |
BMVC | 4 |
| 2023 | Group Masked Model Learning for General Audio RepresentationabstractVision transformers have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships. However, transformers are known to be data hungry which require orders of magnitude more data [1] to train. This has motivated the research in self-supervised pretraining of audio transformers, which reduces the dependency on large amounts of labeled data and focuses on extracting concise representation of the audio spectrograms. In this paper, we propose Audio-GMML, a self-supervised transformer for general audio representations that is based on Group Masked Model Learning (GMML) and a patch aggregation strategy to improve the performance of learned representations and enforce global structure of the given audio. We evaluate our pretrained models on several downstream tasks, setting a new state-of-the-art performance on five audio and speech classification tasks. The code and pretrained weights will be made publicly available for the scientific community. Sara Atito Ali Ahmed, Muhammad Awais 0001, Tony Alex, Josef Kittler |
ICIP | 2 |
| 2023 | GMML is All You NeedabstractVision transformers (ViTs) have generated significant interest in the computer vision community because of their flexibility in exploiting contextual information, whether it is sharply confined local, or long range global. However, they are known to be data hungry and therefore often pretrained on large-scale datasets, e.g. JFT-300M or ImageNet. An ideal learning method would perform best regardless of the size of the dataset, a property lacked by current learning methods, with merely a few existing works studying ViTs with limited data. We propose Group Masked Model Learning (GMML), a self-supervised learning (SSL) method that is able to train ViTs and achieve state-of-the-art (SOTA) performance when pre-trained with limited data. The GMML uses the information conveyed by all concepts in the image. This is achieved by manipulating randomly groups of connected tokens, successively covering different meaningful parts of the image content, and then recovering the hidden information from the visible part of the concept. Unlike most of the existing SSL approaches, GMML does not require momentum encoder, nor relies on careful implementation details such as large batches and gradient stopping. Pretraining, finetuning, and evaluation codes are available under: https://github.com/GMML. Sara Atito Ali Ahmed, Muhammad Awais 0001, Srinivasa Rao Nandam, Josef Kittler |
ICIP | 2 |
| 2023 | LT-ViT: A Vision Transformer for Multi-Label Chest X-Ray ClassificationabstractVision Transformers (ViTs) are widely adopted in medical imaging tasks, and some existing efforts have been directed towards vision-language training for Chest X-rays (CXRs). However, we envision that there still exists a potential for improvement in vision-only training for CXRs using ViTs, by aggregating information from multiple scales, which has been proven beneficial for non-transformer networks. Hence, we have developed LT-ViT, a transformer that utilizes combined attention between image tokens and randomly initialized auxiliary tokens that represent labels. Our experiments demonstrate that LT-ViT (1) surpasses the state-of-the-art performance using pure ViTs on two publicly available CXR datasets, (2) is generalizable to other pre-training methods and therefore is agnostic to model initialization, and (3) enables model interpretability without grad-cam and its variants. Umar Marikkar, Sara Atito Ali Ahmed, Muhammad Awais 0001, Adam Mahdi |
ICIP | 3 |
| 2023 | Deep Order-Preserving Learning With Adaptive Optimal Transport DistanceabstractWe consider a framework for taking into consideration the relative importance (ordinality) of object labels in the process of learning a label predictor function. The commonly used loss functions are not well matched to this problem, as they exhibit deficiencies in capturing natural correlations of the labels and the corresponding data. We propose to incorporate such correlations into our learning algorithm using an optimal transport formulation. Our approach is to learn the ground metric, which is partly involved in forming the optimal transport distance, by leveraging ordinality as a general form of side information in its formulation. Based on this idea, we then develop a novel loss function for training deep neural networks. A highly efficient alternating learning method is then devised to alternatively optimise the ground metric and the deep model in an end-to-end learning manner. This scheme allows us to adaptively adjust the shape of the ground metric, and consequently the shape of the loss function for each application. We back up our approach by theoretical analysis and verify the performance of our proposed scheme by applying it to two learning tasks, i.e. chronological age estimation from the face and image aesthetic assessment. The numerical results on several benchmark datasets demonstrate the superiority of the proposed algorithm. Ali Akbari 0003, Muhammad Awais 0001, Soroush Fatemifar, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | NPT-Loss: Demystifying Face Recognition Losses With Nearest Proxies TripletabstractFace recognition (FR) using deep convolutional neural networks (DCNNs) has seen remarkable success in recent years. One key ingredient of DCNN-based FR is the design of a loss function that ensures discrimination between various identities. The state-of-the-art (SOTA) solutions utilise normalised Softmax loss with additive and/or multiplicative margins. Despite being popular and effective, these losses are justified only intuitively with little theoretical explanations. In this work, we show that under the LogSumExp (LSE) approximation, the SOTA Softmax losses become equivalent to a proxy-triplet loss that focuses on nearest-neighbour negative proxies only. This motivates us to propose a variant of the proxy-triplet loss, entitled Nearest Proxies Triplet (NPT) loss, which unlike SOTA solutions, converges for a wider range of hyper-parameters and offers flexibility in proxy selection and thus outperforms SOTA techniques. We generalise many SOTA losses into a single framework and give theoretical justifications for the assertion that minimising the proposed loss ensures a minimum separability between all identities. We also show that the proposed loss has an implicit mechanism of hard-sample mining. We conduct extensive experiments using various DCNN architectures on a number of FR benchmarks to demonstrate the efficacy of the proposed scheme over SOTA methods. Syed Safwan Khalid, Muhammad Awais 0001, Zhenhua Feng 0001, Chi-Ho Chan, Ammarah Farooq, Ali Akbari 0003, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | A Theoretical Insight Into the Effect of Loss Function for Deep Semantic-Preserving LearningabstractGood generalization performance is the fundamental goal of any machine learning algorithm. Using the uniform stability concept, this article theoretically proves that the choice of loss function impacts the generalization performance of a trained deep neural network (DNN). The adopted stability-based framework provides an effective tool for comparing the generalization error bound with respect to the utilized loss function. The main result of our analysis is that using an effective loss function makes stochastic gradient descent more stable which consequently leads to the tighter generalization error bound, and so better generalization performance. To validate our analysis, we study learning problems in which the classes are semantically correlated. To capture this semantic similarity of neighboring classes, we adopt the well-known semantics-preserving learning framework, namely label distribution learning (LDL). We propose two novel loss functions for the LDL framework and theoretically show that they provide stronger stability than the other widely used loss functions adopted for training DNNs. The experimental results on three applications with semantically correlated classes, including facial age estimation, head pose estimation, and image esthetic assessment, validate the theoretical insights gained by our analysis and demonstrate the usefulness of the proposed loss functions in practical applications. Ali Akbari 0003, Muhammad Awais 0001, Manijeh Bashar, Josef Kittler |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | AXM-Net: Implicit Cross-Modal Feature Alignment for Person Re-identificationabstractCross-modal person re-identification (Re-ID) is critical for modern video surveillance systems. The key challenge is to align cross-modality representations conforming to semantic information present for a person and ignore background information. This work presents a novel convolutional neural network (CNN) based architecture designed to learn semantically aligned cross-modal visual and textual representations. The underlying building block, named AXM-Block, is a unified multi-layer network that dynamically exploits the multi-scale knowledge from both modalities and re-calibrates each modality according to shared semantics. To complement the convolutional design, contextual attention is applied in the text branch to manipulate long-term dependencies. Moreover, we propose a unique design to enhance visual part-based feature coherence and locality information. Our framework is novel in its ability to implicitly learn aligned semantics between modalities during the feature learning stage. The unified feature learning effectively utilizes textual data as a super-annotation signal for visual representation learning and automatically rejects irrelevant information. The entire AXM-Net is trained end-to-end on CUHK-PEDES data. We report results on two tasks, person search and cross-modal Re-ID. The AXM-Net outperforms the current state-of-the-art (SOTA) methods and achieves 64.44% Rank@1 on the CUHK-PEDES test set. It also outperforms by >10% for cross-viewpoint text-to-image Re-ID scenarios on CrossRe-ID and CUHK-SYSU datasets. Ammarah Farooq, Muhammad Awais 0001, Josef Kittler, Syed Safwan Khalid |
AAAI | 2 |
| 2022 | Distribution Cognisant Loss for Cross-Database Facial Age Estimation With Sensitivity AnalysisabstractExisting facial age estimation studies have mostly focused on intra-database protocols that assume training and test images are captured under similar conditions. This is rarely valid in practical applications, where we typically encounter training and test sets with different characteristics. In this article, we deal with such situations, namely subjective-exclusive cross-database age estimation. We formulate the age estimation problem as the distribution learning framework, where the age labels are encoded as a probability distribution. To improve the cross-database age estimation performance, we propose a new loss function which provides a more robust measure of the difference between ground-truth and predicted distributions. The desirable properties of the proposed loss function are theoretically analysed and compared with the state-of-the-art approaches. In addition, we compile a new balanced large-scale age estimation database. Last, we introduce a novel evaluation protocol, called subject-exclusive cross-database age estimation protocol, which provides meaningful information of a method in terms of the generalisation capability. The experimental results demonstrate that the proposed approach outperforms the state-of-the-art age estimation methods under both intra-database and subject-exclusive cross-database evaluation protocols. In addition, in this article, we provide a comparative sensitivity analysis of various algorithms to identify trends and issues inherent to their performance. This analysis introduces some open problems to the community which might be considered when designing a robust age estimation system. Ali Akbari 0003, Muhammad Awais 0001, Zhenhua Feng 0001, Ammarah Farooq, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Developing a generic framework for anomaly detection
Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler |
Pattern Recognit. | 2 |
| 2022 | Face spoofing detection ensemble via multistage optimisation and pruning
Soroush Fatemifar, Shahrokh Asadi, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler |
Pattern Recognit. Lett. | 3 |
| 2022 | A Novel Ground Metric for Optimal Transport-Based Chronological Age EstimationabstractLabel distribution learning (LDL) is the state-of-the-art approach to dealing with a number of real-world applications, such as chronological age estimation from a face image, where there is an inherent similarity among adjacent age labels. LDL takes into account the semantic similarity by assigning a label distribution to each instance. The well-known Kullback-Leibler (KL) divergence is the widely used loss function for the LDL framework. However, the KL divergence does not fully and effectively capture the semantic similarity among age labels, thus leading to suboptimal performance. In this article, we propose a novel loss function based on the optimal transport theory for the LDL-based age estimation. A ground metric function plays an important role in the optimal transport formulation. It should be carefully determined based on the underlying geometric structure of the label space of the application in-hand. The label space in the age estimation problem has a specific geometric structure, that is, closer ages have more inherent semantic relationships. Inspired by this, we devise a novel ground metric function, which enables the loss function to increase the influence of highly correlated ages; thus exploiting the semantic similarity among ages more effectively than the existing loss functions. We then use the proposed loss function, namely, γ -Wasserstein loss, for training a deep neural network (DNN). This leads to a notoriously computationally expensive and nonconvex optimization problem. Following the standard methodology, we formulate the optimization function as a convex problem and then use an efficient iterative algorithm to update the parameters of the DNN. Extensive experiments in age estimation on different benchmark datasets validate the effectiveness of the proposed method, which consistently outperforms state-of-the-art approaches. Ali Akbari 0003, Muhammad Awais 0001, Soroush Fatemifar, Syed Safwan Khalid, Josef Kittler |
IEEE Trans. Cybern. | 2 |
| 2021 | Particle Swarm And Pattern Search Optimisation Of An Ensemble Of Face Anomaly DetectorsabstractWhile the remarkable advances in face matching render face biometric technology more widely applicable, its successful deployment may be compromised by face spoofing. Recent studies have shown that anomaly-based face spoofing detectors offer an interesting alternative to the multiclass counterparts by generalising better to unseen types of attack. In this work, we investigate the merits of fusing multiple anomaly spoofing detectors in the unseen attack scenario via a Weighted Averaging (WA) and client-specific design. We propose to optimise the parameters of WA by a two-stage optimisation method consisting of Particle Swarm Optimisation (PSO) and the Pattern Search (PS) algorithms to avoid the local minimum problem. Besides, we propose a novel scoring normalisation method which could be effectively applied in extreme cases such as heavy-tailed distributions. We evaluate the capability of the proposed system on publicly available face anti-spoofing databases including Replay-Attack, Replay-Mobile and Rose-Youtu. The experimental results demonstrate that the proposed fusion system outperforms the majority of anomaly-based and state-of-the-art multiclass approaches. Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler |
ICIP | 2 |
| 2021 | How Does Loss Function Affect Generalization Performance of Deep Learning? Application to Human Age EstimationabstractGood generalization performance across a wide variety of domains caused by many external and internal factors is the fundamental goal of any machine learning algorithm. This paper theoretically proves that the choice of loss function matters for improving the generalization performance of deep learning-based systems. By deriving the generalization error bound for deep neural models trained by stochastic gradient descent, we pinpoint the characteristics of the loss function that is linked to the generalization error and can therefore be used for guiding the loss function selection process. In summary, our main statement in this paper is: choose a stable loss function, generalize better. Focusing on human age estimation from the face which is a challenging topic in computer vision, we then propose a novel loss function for this learning problem. We theoretically prove that the proposed loss function achieves stronger stability, and consequently a tighter generalization error bound, compared to the other common loss functions for this problem. We have supported our findings theoretically, and demonstrated the merits of the guidance process experimentally, achieving significant improvements. Ali Akbari 0003, Muhammad Awais 0001, Manijeh Bashar, Josef Kittler |
ICML | 2 |
| 2021 | Client-specific anomaly detection for face presentation attack detection
Soroush Fatemifar, Shervin Rahimzadeh Arashloo, Muhammad Awais 0001, Josef Kittler |
Pattern Recognit. | 3 |
| 2020 | Sensitivity of Age Estimation Systems to Demographic Factors and Image Quality: Achievements and ChallengesabstractRecently, impressively growing efforts have been devoted to the challenging task of facial age estimation. The improvements in performance achieved by new algorithms are measured on several benchmarking test databases with different characteristics to check on consistency. While this is a valuable methodology in itself, a significant issue in the most age estimation related studies is that the reported results lack an assessment of intrinsic system uncertainty. Hence, a more in-depth view is required to examine the robustness of age estimation systems in different scenarios. The purpose of this paper is to conduct an evaluative and comparative analysis of different age estimation systems to identify trends, as well as the points of their critical vulnerability. In particular, we investigate four age estimation systems, including the online Microsoft service, two best state-of-the-art approaches advocated in the literature, as well as a novel age estimation algorithm. We analyse the effect of different internal and external factors, including gender, ethnicity, expression, makeup, illumination conditions, quality and resolution of the face images, on the performance of these age estimation systems. The goal of this sensitivity analysis is to provide the biometrics community with the insight and understanding of the critical subject-, camera- and environmental-based factors that affect the overall performance of the age estimation system under study. Ali Akbari 0003, Muhammad Awais 0001, Josef Kittler |
IJCB | 2 |
| 2020 | Cross Modal Person Re-identification with Visual-Textual QueriesabstractClassical person re-identification approaches assume that a person of interest has appeared across different cameras and can be queried by one of the existing images. However, in real-world surveillance scenarios, frequently no visual information will be available about the queried person. In such scenarios, a natural language description of the person by a witness will provide the only source of information for retrieval. In this work, person re-identification using both vision and language information is addressed under all possible gallery and query scenarios. A two stream deep convolutional neural network framework supervised by identity based cross entropy loss is presented. Canonical Correlation Analysis is performed to enhance the correlation between the two modalities in a joint latent embedding space. To investigate the benefits of the proposed approach, a new testing protocol under a multi modal ReID setting is proposed for the test split of the CUHK-PEDES and CUHK-SYSU benchmarks. The experimental results verify that the learnt visual representations are more robust and perform 20% better during retrieval as compared to a single modality system. Ammarah Farooq, Muhammad Awais 0001, Josef Kittler, Ali Akbari 0003, Syed Safwan Khalid |
IJCB | 2 |
| 2020 | A Stacking Ensemble for Anomaly Based Client-Specific Face Spoofing DetectionabstractTo counteract spoofing attacks, the majority of recent approaches to face spoofing attack detection formulate the problem as a binary classification task in which real data and attack-accesses are both used to train spoofing detectors. Although the classical training framework has been demonstrated to deliver satisfactory results, its robustness to unseen attacks is debatable. Inspired by the recent success of anomaly detection models in face spoofing detection, we propose an ensemble of one-class classifiers fused by a Stacking ensemble method to reduce the generalisation error in the more realistic unseen attack scenario. To be consistent with this scenario, anomalous samples are considered neither for training the component anomaly classifiers nor for the design of the Stacking ensemble. To achieve better face-anti spoofing results, we adopt client-specific information to build both constituent classifiers as well as the Stacking combiner. Besides, we propose a novel 2-stage Genetic Algorithm to further improve the generalisation performance of Stacking ensemble. We evaluate the effectiveness of the proposed systems on publicly available face anti-spoofing databases including Replay-Attack, Replay-Mobile and Rose-Youtu. The experimental results following the unseen attack evaluation protocol confirm the merits of the proposed model. Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler |
ICIP | 2 |
| 2020 | A Flatter Loss for Bias Mitigation in Cross-dataset Facial Age EstimationabstractThe most existing studies in the facial age estimation assume training and test images are captured under similar shooting conditions. However, this is rarely valid in real-worlds applications, where training and test sets usually have different characteristics. In this paper, we advocate a cross-dataset protocol for age estimation benchmarking. In order to improve the cross-dataset age estimation performance, we mitigate the inherent bias caused by the learning algorithm itself. To this end, we propose a novel loss function that is more effective for neural network training. The relative smoothness of the proposed loss function is its advantage with regards to the optimisation process performed by stochastic gradient descent (SGD). Compared with existing loss functions, the lower gradient of the proposed loss function leads to the convergence of SGD to a better optimum point, and consequently a better generalisation. The cross-dataset experimental results demonstrate the superiority of the proposed method over the state-of-the-art algorithms in terms of accuracy and generalisation capability. Ali Akbari 0003, Muhammad Awais 0001, Zhenhua Feng 0001, Ammarah Farooq, Josef Kittler |
ICPR | 2 |
| 2020 | Rectified Wing Loss for Efficient and Robust Facial Landmark Localisation with Convolutional Neural NetworksabstractAbstract Efficient and robust facial landmark localisation is crucial for the deployment of real-time face analysis systems. This paper presents a new loss function, namely Rectified Wing (RWing) loss, for regression-based facial landmark localisation with Convolutional Neural Networks (CNNs). We first systemically analyse different loss functions, including L2, L1 and smooth L1. The analysis suggests that the training of a network should pay more attention to small-medium errors. Motivated by this finding, we design a piece-wise loss that amplifies the impact of the samples with small-medium errors. Besides, we rectify the loss function for very small errors to mitigate the impact of inaccuracy of manual annotation. The use of our RWing loss boosts the performance significantly for regression-based CNNs in facial landmarking, especially for lightweight network architectures. To address the problem of under-representation of samples with large pose variations, we propose a simple but effective boosting strategy, referred to as pose-based data balancing. In particular, we deal with the data imbalance problem by duplicating the minority training samples and perturbing them by injecting random image rotation, bounding box translation and other data augmentation strategies. Last, the proposed approach is extended to create a coarse-to-fine framework for robust and efficient landmark localisation. Moreover, the proposed coarse-to-fine framework is able to deal with the small sample size problem effectively. The experimental results obtained on several well-known benchmarking datasets demonstrate the merits of our RWing loss and prove the superiority of the proposed method over the state-of-the-art approaches. Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001, Xiaojun Wu 0001 |
Int. J. Comput. Vis. | 3 |
| 2019 | Spoofing Attack Detection by Anomaly DetectionabstractSpoofing attacks on biometric systems can seriously compromise their practical utility. In this paper we focus on face spoofing detection. The majority of papers on spoofing attack detection formulate the problem as a two or multiclass learning task, attempting to separate normal accesses from samples of different types of spoofing attacks. In this paper we adopt the anomaly detection approach proposed in [1], where the detector is trained on genuine accesses only using one-class classifiers and investigate the merit of subject specific solutions. We show experimentally that subject specific models are superior to the commonly used client independent method. We also demonstrate that the proposed approach is more robust than multiclass formulations to unseen attacks. Soroush Fatemifar, Shervin Rahimzadeh Arashloo, Muhammad Awais 0001, Josef Kittler |
ICASSP | 3 |
| 2019 | Divergence Based Weighting for Information Channels in Deep Convolutional Neural Networks for Bird Audio DetectionabstractIn this paper, we address the problem of bird audio detection and propose a new convolutional neural network architecture together with a divergence based information channel weighing strategy in order to achieve improved state-of-the-art performance and faster convergence. The effectiveness of the methodology is shown on the Bird Audio Detection Challenge 2018 (Detection and Classification of Acoustic Scenes and Events Challenge, Task 3) development data set. Cemre Zor, Muhammad Awais 0001, Josef Kittler, Miroslaw Bober, Syed Sameed Husain, Qiuqiang Kong, Christian Kroos |
ICASSP | 2 |
| 2018 | Wing Loss for Robust Facial Landmark Localisation With Convolutional Neural NetworksabstractWe present a new loss function, namely Wing loss, for robust facial landmark localisation with Convolutional Neural Networks (CNNs). We first compare and analyse different loss functions including L2, L1 and smooth L1. The analysis of these loss functions suggests that, for the training of a CNN-based localisation model, more attention should be paid to small and medium range errors. To this end, we design a piece-wise loss function. The new loss amplifies the impact of errors from the interval (-w, w) by switching from L1 loss to a modified logarithm function. To address the problem of under-representation of samples with large out-of-plane head rotations in the training set, we propose a simple but effective boosting strategy, referred to as pose-based data balancing. In particular, we deal with the data imbalance problem by duplicating the minority training samples and perturbing them by injecting random image rotation, bounding box translation and other data augmentation approaches. Last, the proposed approach is extended to create a two-stage framework for robust facial landmark localisation. The experimental results obtained on AFLW and 300W demonstrate the merits of the Wing loss function, and prove the superiority of the proposed method over the state-of-the-art approaches. Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001, Patrik Huber 0001, Xiaojun Wu 0001 |
CVPR | 3 |
| 2018 | Gaussian mixture 3D morphable face modelabstract3D Morphable Face Models (3DMM) have been used in pattern recognition for some time now. They have been applied as a basis for 3D face recognition, as well as in an assistive role for 2D face recognition to perform geometric and photometric normalisation of the input image, or in 2D face recognition system training. The statistical distribution underlying 3DMM is Gaussian. However, the single-Gaussian model seems at odds with reality when we consider different cohorts of data, e.g. Black and Chinese faces. Their means are clearly different. This paper introduces the Gaussian Mixture 3DMM (GM-3DMM) which models the global population as a mixture of Gaussian subpopulations, each with its own mean. The proposed GM-3DMM extends the traditional 3DMM naturally, by adopting a shared covariance structure to mitigate small sample estimation problems associated with data in high dimensional spaces. We construct a GM-3DMM, the training of which involves a multiple cohort dataset, SURREY-JNU, comprising 942 3D face scans of people with mixed backgrounds. Experiments in fitting the GM-3DMM to 2D face images to facilitate their geometric and photometric normalisation for pose and illumination invariant face recognition demonstrate the merits of the proposed mixture of Gaussians 3D face model. Willem P. Koppen, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001, William J. Christmas, Xiaojun Wu 0001, He-Feng Yin |
Pattern Recognit. | 4 |
| 2017 | Medical image retrieval using deep convolutional neural network
Adnan Qayyum, Syed Muhammad Anwar, Muhammad Awais 0001, Muhammad Majid |
Neurocomputing | 3 |
| 2013 | A Robust and Scalable Visual Category and Action Recognition System Using Kernel Discriminant Analysis With Spectral RegressionabstractVisual concept detection and action recognition are one of the most important tasks in content-based multimedia information retrieval (CBMIR) technology. It aims at annotating images using a vocabulary defined by a set of concepts of interest including scenes types (mountains, snow, etc.) or human actions (phoning, playing instrument). This paper describes our system in the ImageCLEF@ICPR10, Pascal VOC 08 Visual Concept Detection and Pascal VOC 10 Action Recognition Challenges. The proposed system ranked first in these large-scale tasks when evaluated independently by the organizers. The proposed system involves state-of-the-art local descriptor computation, vector quantization via clustering, structured scene or object representation via localized histograms of vector codes, similarity measure for kernel construction and classifier learning. The main novelty is the classifier-level and kernel-level fusion using Kernel Discriminant Analysis and Spectral Regression (SR-KDA) with RBF Chi-Squared kernels obtained from various image descriptors. The distinctiveness of the proposed method is also assessed experimentally using a video benchmark: the Mediamill Challenge along with benchmarks from ImageCLEF@ICPR10, Pascal VOC 10 and Pascal VOC 08. From the experimental results, it can be derived that the presented system consistently yields significant performance gains when compared with the state-of-the art methods. The other strong point is the introduction of SR-KDA in the classification stage where the time complexity scales linearly with respect to the number of concepts and the main computational complexity is independent of the number of categories. Muhammad Atif Tahir, Fei Yan 0001, Piotr Koniusz, Muhammad Awais 0001, Mark Barnard, Krystian Mikolajczyk, Ahmed Bouridane, Josef Kittler |
IEEE Trans. Multim. | 4 |
| 2011 | Augmented Kernel Matrix vs Classifier Fusion for Object RecognitionabstractAugmented Kernel Matrix (AKM) has recently been proposed to accommodate for the fact that a single training example may have different importance in different feature spaces, in contrast to Multiple Kernel Learning (MKL) that assigns the same weight to all examples in one feature space.However, the AKM approach is limited to small datasets due to its memory requirements.An alternative way to fuse information from different feature channels is classifier fusion (ensemble methods).There is a significant amount of work on linear programming formulations of classifier fusion (CF) in the case of binary classification.In this paper we derive primal and dual of AKM to draw its correspondence with CF.We propose a multiclass extension of binary ν-LPBoost, which learns the contribution of each class in each feature channel.Existing approaches of CF promote sparse features combinations, due to regularization based on 1 -norm, and lead to a selection of a subset of feature channels, which is not good in case of informative channels.We also generalize existing CF formulations to arbitrary p -norm for binary and multiclass problems which results in more effective use of complementary information.We carry out an extensive comparison and show that the proposed nonlinear CF schemes outperform its sparse counterpart as well as state-of-the-art MKL approaches. Muhammad Awais 0001, Fei Yan 0001, Krystian Mikolajczyk, Josef Kittler |
BMVC | 1 |
| 2011 | Novel Fusion Methods for Pattern Recognition
Muhammad Awais 0001, Fei Yan 0001, Krystian Mikolajczyk, Josef Kittler |
ECML/PKDD (1) | 1 |
| 2010 | Feature Pairs Connected by Lines for Object RecognitionabstractIn this paper we exploit image edges and segmentation maps to build features for object category recognition. We build a parametric line based image approximation to identify the dominant edge structures. Line ends are used as features described by histograms of gradient orientations. We then form descriptors based on connected line ends to incorporate weak topological constraints which improve their discriminative power. Using point pairs connected by an edge assures higher repeatability than a random pair of points or edges. The results are compared with state-of-the-art, and show significant improvement on challenging recognition benchmark Pascal VOC 2007. Kernel based fusion is performed to emphasize the complementary nature of our descriptors with respect to the state-of-the-art features. Muhammad Awais 0001, Krystian Mikolajczyk |
ICPR | 1 |
| 2010 | The University of Surrey Visual Concept Detection System at ImageCLEF@ICPR: Working NotesabstractVisual concept detection is one of the most important tasks in image and video indexing. This paper describes our system in the ImageCLEF@ICPR Visual Concept Detection Task which ranked first for large-scale visual concept detection tasks in terms of Equal Error Rate (EER) and Area under Curve (AUC) and ranked third in terms of hierarchical measure. The presented approach involves state-of-the-art local descriptor computation, vector quantisation via clustering, structured scene or object representation via localised histograms of vector codes, similarity measure for kernel construction and classifier learning. The main novelty is the classifier-level and kernel-level fusion using Kernel Discriminant Analysis with RBF/Power Chi-Squared kernels obtained from various image descriptors. For 32 out of 53 individual concepts, we obtain the best performance of all 12 submissions to this task. Muhammad Atif Tahir, Fei Yan 0001, Mark Barnard, Muhammad Awais 0001, Krystian Mikolajczyk, Josef Kittler |
ICPR | 4 |