VLDB 2026 Research / reviewers in the wild / expert
Arif Mahmood
dblp:18/4138
· DBLP profile ↗
91ranked-venue papers
14as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 8 first-author · 21 since 2021Artificial intelligence and machine learning · 44 · 4 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorComputer networks · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenMix: Effective data augmentation with generative diffusion model image editingabstractData augmentation is widely used to enhance generalization in visual classification tasks. However, traditional methods struggle when source and target domains differ, as in domain adaptation, due to their inability to address domain gaps. This paper introduces GenMix, a generalizable prompt-guided generative data augmentation approach that enhances both in-domain and cross-domain image classification. Our technique leverages image editing to generate augmented images based on custom conditional prompts, designed specifically for each problem type. By blending portions of the input image with its edited generative counterpart and incorporating fractal patterns, our approach mitigates unrealistic images and label ambiguity, improving the performance and adversarial robustness of the resulting models. Efficacy of our method is established with extensive experiments on eight public datasets for general and fine-grained classification, in both in-domain and cross-domain settings. Additionally, we demonstrate performance improvements for self-supervised learning, learning with data scarcity, and adversarial robustness. As compared to the existing state-of-the-art methods, our technique achieves stronger performance across the board. Khawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood, Karthik Nandakumar, Naveed Akhtar |
Expert Syst. Appl. | 3 |
| 2026 | A Comprehensive Benchmark for Evaluating Night-time Visual Object Tracking
Arif Mahmood, Muhammad Haris Khan |
Int. J. Comput. Vis. | 2 |
| 2026 | Correction: A Comprehensive Benchmark for Evaluating Night-time Visual Object Tracking
Arif Mahmood, Muhammad Haris Khan |
Int. J. Comput. Vis. | 2 |
| 2026 | Improving generative adversarial network generalization for facialexpression synthesis
Arbish Akram, Nazar Khan, Arif Mahmood |
Multim. Tools Appl. | 3 |
| 2026 | Enhancing GNN learning with node augmentation
Maria Marrium, Arif Mahmood, Muhammad Haris Khan, M. Saad Shakeel, Wenxiong Kang |
Neural Networks | 2 |
| 2026 | AquaticCLIP: A Vision-Language Foundation Model and Dataset for Underwater Scene AnalysisabstractThe preservation of aquatic biodiversity is critical in mitigating the effects of climate change. Aquatic scene understanding plays a pivotal role in aiding marine scientists in their decision-making processes. In this article, we introduce AquaticCLIP, a novel contrastive language-image pretraining (CLIP) model tailored for aquatic scene understanding. AquaticCLIP presents an underwater domain-specific learning framework that aligns images and texts in aquatic environments, enabling tasks such as segmentation, classification, detection, and object counting. By leveraging our large-scale underwater image-text paired dataset without the need for ground-truth (GT) annotations, our model enriches existing vision-language models (VLMs) in the aquatic domain. For this purpose, we construct a 2-million underwater image-text paired dataset using heterogeneous resources, including YouTube, Netflix, National Geographic (NatGeo), etc. To fine-tune AquaticCLIP, we propose a prompt-guided vision encoder (PGVE) that progressively aggregates patch features via learnable prompts, while a vision-guided mechanism enhances the language encoder by incorporating visual context. The model is optimized through a contrastive pretraining loss to align visual and textual modalities. AquaticCLIP achieves notable performance improvements in zero-shot settings across multiple underwater computer vision tasks, outperforming existing methods in both accuracy and robustness. Our model sets a new benchmark for vision-language applications in underwater environments. The code and dataset for AquaticCLIP are publicly available on GitHub at: https://github.com/BasitAlawode/AquaticCLIP. Basit Alawode, Iyyakutti Iyappan Ganapathi, Sajid Javed, Mohammed Bennamoun, Arif Mahmood |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual RepresentationabstractIn Computational Pathology (CPath), the introduction of Vision-Language Models (VLMs) has opened new avenues for research, focusing primarily on aligning image-text pairs at a single magnification level. However, this approach might not be sufficient for tasks like cancer subtype classification, tissue phenotyping, and survival analysis due to the limited level of detail that a single-resolution image can provide. Addressing this, we propose a novel multi-resolution paradigm leveraging Whole Slide Images (WSIs) to extract histology patches at multiple resolutions and generate corresponding textual descriptions through advanced CPath VLM. We introduce visual-textual alignment at multiple resolutions as well as cross-resolution alignment to establish more effective text-guided visual representations. Cross-resolution alignment using a multi-modal encoder enhances the model’s ability to capture context from multiple resolutions in histology images. Our model aims to capture a broader range of information, supported by novel loss functions, enriches feature representation, improves discriminative ability, and enhances generalization across different resolutions. Pre-trained on a comprehensive TCGA dataset with 34 million image-language pairs at various resolutions, our fine-tuned model outperforms State-Of-The-Art (SOTA) counterparts across multiple datasets and tasks, demonstrating its effectiveness in CPath. The code is available on GitHub at: https://github.com/BasitAlawode/MR-PLIP. Shahad Albastaki, Anabia Sohail, Iyyakutti Iyappan Ganapathi, Basit Alawode, Asim Khan, Sajid Javed, Naoufel Werghi, Mohammed Bennamoun, Arif Mahmood |
CVPR | 9 |
| 2025 | Unsupervised Discovery of Facial Landmarks and Head PoseabstractUnsupervised landmark and head pose estimation is fundamental in fields like biometrics, augmented reality, and emotion recognition, offering accurate spatial data without relying on labeled datasets. It enhances scalability, adaptability, and generalization across diverse settings, where manual labeling is costly. In this work we exploit Stable Diffusion to approach the challenging problem of unsupervised landmarks and head pose estimation and make following contributions. (a) We propose a semantic-aware landmark localization algorithm including a consistent landmarks selection technique. (b) To encode landmarks and their holistic configuration, we propose learning image-aware textual embedding. (c) A novel algorithm for landmarks-guided 3D head pose estimation is also proposed. (d) We refine the landmarks using head pose by innovating a 3D rendering based augmentation and pose-based batching technique while the refined landmarks, consequently improving the head pose. (e) We report a new state-of-the-art in unsupervised facial landmark estimation across five challenging datasets including AFLW2000, MAFL, Cat-Heads, LS3D and a facial landmark tracking benchmark 300VW. In unsupervised head pose estimation, we outperform existing methods on BIWI and AFLW2000 by visible margins. Moreover, our method provides a significant training speedup over the existing best unsupervised landmark detection method.1 Satyajit Tourani, Siddharth Tourani, Arif Mahmood, Muhammad Haris Khan |
CVPR | 3 |
| 2025 | Localization Lens for Improving Medical Vision-Language Models
Hasan Farooq, Murtaza Taj, Mehwish Nasim, Arif Mahmood |
MICCAI (9) | 4 |
| 2025 | Diffusion-Guided Graph Data AugmentationabstractGraph Neural Networks (GNNs) have achieved remarkable success in a wide range of applications. However, when trained on limited or low-diversity datasets, GNNs are prone to overfitting and memorization, which impacts their generalization. To address this, graph data augmentation (GDA) has become a crucial task to enhance the performance and generalization of GNNs.
Traditional GDA methods employ simple transformations that result in limited performance gains. Although recent diffusion-based augmentation methods offer improved results, they are sparse, task-specific, and constrained by class labels. In this work, we propose a more general and effective diffusion-based GDA framework that is task-agnostic and label-free.
For better training stability and reduced computational cost, we employ a graph variational auto-encoder (GVAE) to learn a compact latent graph representation. A diffusion model is used in the learned latent space to generate both consistent and diverse augmentations.
For a fixed augmentation budget, our algorithm selects a subset of samples that would benefit the most from the augmentation.
To further improve performance, we also perform test-time augmentation, leveraged by the label-free nature of our method.
Thanks to the efficient utilization of GVAE and latent diffusion, our algorithm significantly enhances machine learning safety measures, including calibration, robustness to corruptions, and prediction consistency. Moreover, our method has shown improved robustness against four types of adversarial attacks and achieves better generalization performance.
To demonstrate the effectiveness of the proposed method, we compare it with 30 existing methods on 12 benchmark datasets across node classification, link prediction, and graph classification in various learning settings, including semi-supervised, supervised, and long-tailed data distributions.
The code will soon be made publicly available. Maria Marrium, Arif Mahmood, Muhammad Haris Khan, M. Saad Shakeel, Wenxiong Kang |
NeurIPS | 2 |
| 2025 | Enhancing 3D Human Pose Estimation Amidst Severe Occlusion With Dual Transformer FusionabstractIn the field of 3D Human Pose Estimation from monocular videos, the presence of diverse occlusion types presents a formidable challenge. Prior research has made progress by harnessing spatial and temporal cues to infer 3D poses from 2D joint observations. This paper introduces a Dual Transformer Fusion (DTF) algorithm, a novel approach to obtain a holistic 3D pose estimation, even in the presence of severe occlusions. Confronting the issue of occlusion-induced missing joint data, we propose a temporal interpolation-based occlusion guidance mechanism. To enable precise 3D Human Pose Estimation, our approach leverages the innovative DTF architecture, which first generates a pair of intermediate views. Each intermediate-view undergoes spatial refinement through a self-refinement schema. Subsequently, these intermediate-views are fused to yield the final 3D human pose estimation. The entire system is end-to-end trainable. Through extensive experiments conducted on the Human3.6 M and MPI-INF-3DHP datasets, our method's performance is rigorously evaluated. Notably, our approach outperforms existing state-of-the-art methods on both datasets, yielding substantial improvements. Mehwish Ghafoor, Arif Mahmood, Muhammad Bilal 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Depth Attention for Robust RGB Tracking
Arif Mahmood, Muhammad Haris Khan |
ACCV (2) | 2 |
| 2024 | NT-VOT211: A Large-Scale Benchmark for Night-Time Visual Object Tracking
Arif Mahmood, Muhammad Haris Khan |
ACCV (2) | 2 |
| 2024 | Diffusemix: Label-Preserving Data Augmentation with Diffusion ModelsabstractRecently, a number of image-mixing-based augmentation techniques have been introduced to improve the gen-eralization of deep neural networks. In these techniques, two or more randomly selected natural images are mixed together to generate an augmented image. Such methods may not only omit important portions of the input images but also introduce label ambiguities by mixing images across labels resulting in misleading supervisory signals. To address these limitations, we propose Diffusemix, a novel data augmentation technique that leverages a diffusion model to reshape training images, supervised by our bespoke conditional prompts. First, concatenation of a partial natural image and its generated counterpart is ob-tained which helps in avoiding the generation of unrealistic images or label ambiguities. Then, to enhance resilience against adversarial attacks and improves safety measures, a randomly selected structural pattern from a set of frac-tal images is blended into the concatenated image to form the final augmented image for training. Our empirical results on seven different datasets reveal that Diffusemix achieves superior performance compared to existing state-of-the-art methods on tasks including general classification, fine- grained classification, fine-tuning, data scarcity, and adversarial robustness. Augmented datasets and codes are available here: https://diffusemix.github.io/ Khawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood, Karthik Nandakumar |
CVPR | 3 |
| 2024 | CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language AlignmentabstractThis paper proposes Comprehensive Pathology Language Image Pretraining (CPLIP), a new unsupervised technique designed to enhance the alignment of images and text in histopathology for tasks such as classification and segmentation. This methodology enriches vision-language models by leveraging extensive data without needing ground truth annotations. CPLIP involves constructing a pathology-specific dictionary, generating textual descriptions for images using language models, and retrieving relevant images for each text snippet via a pretrained model. The model is then fine-tuned using a many-to-many contrastive learning method to align complex interrelated concepts across both modalities. Evaluated across multiple histopathology tasks, CPLIP shows notable improvements in zero-shot learning scenarios, outperforming existing methods in both interpretability and robustness and setting a higher benchmark for the application of vision-language models in the field. To encourage further research and replication, the code for CPLIP is available on GitHub at https://cplip.github.io/ Sajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Naoufel Werghi, Mohammed Bennamoun |
CVPR | 2 |
| 2024 | Pose-Guided Self-Training with Two-Stage Clustering for Unsupervised Landmark DiscoveryabstractUnsupervised landmarks discovery (ULD) for an object category is a challenging computer vision problem. In pursuit of developing a robust ULD framework, we explore the potential of a recent paradigm of self-supervised learning algorithms, known as diffusion models. Some recent works have shown that these models implicitly contain important correspondence cues. Towards harnessing the potential of diffusion models for the ULD task, we make the following core contributions. First, we propose a ZeroShot ULD baseline based on simple clustering of random pixel locations with nearest neighbour matching. It delivers better results than existing ULD methods. Second, motivated by the ZeroShot performance, we develop a ULD algorithm based on diffusion features using self-training and clustering which also outperforms prior methods by notable margins. Third, we introduce a new proxy task based on generating latent pose codes and also propose a two-stage clustering to facilitate effective pseudo-labeling, resulting in a significant performance improvement. Overall, our approach consistently outperforms state-of-the-art methods on four challenging benchmarks AFLW, MAFL, CatHeads and LS3D by significant margins. Code and models are available at: https://github.com/skt9/pose-proxy-uld/. Siddharth Tourani, Ahmed Alwheibi, Arif Mahmood, Muhammad Haris Khan |
CVPR | 3 |
| 2024 | Unsupervised mutual transformer learning for multi-gigapixel Whole Slide Image classification
Sajid Javed, Arif Mahmood, Talha Qaiser, Naoufel Werghi, Nasir M. Rajpoot |
Medical Image Anal. | 2 |
| 2024 | Detection and Localization of Firearm Carriers in Complex Scenes for Improved Safety MeasuresabstractDetecting firearms and accurately localizing individuals carrying them in images or videos is of paramount importance in security, surveillance, and content customization. However, this task presents significant challenges in complex environments due to clutter and the diverse shapes of firearms. To address this problem, we propose a novel approach that leverages human–firearm interaction information, which provides valuable clues for localizing firearm carriers. Our approach incorporates an attention mechanism that effectively distinguishes humans and firearms from the background by focusing on relevant areas. Additionally, we introduce a saliency-driven locality-preserving constraint to learn essential features while preserving foreground information in the input image. By combining these components, our approach achieves exceptional results on a newly proposed dataset. To handle inputs of varying sizes, we pass paired human–firearm instances with attention masks as channels through a deep network for feature computation, utilizing an adaptive average pooling (AAP) layer. We extensively evaluate our approach against existing methods in human–object interaction (HOI) detection and achieve significant results (AP = 77.8%) compared to the baseline approach (AP = 63.1%). This demonstrates the effectiveness of leveraging attention mechanisms and saliency-driven locality preservation for accurate human–firearm interaction detection. Our findings contribute to advancing the fields of security and surveillance, enabling more efficient firearm localization and identification in diverse scenarios. Arif Mahmood, Abdul Basit 0019, Muhammad Akhtar Munir, Mohsen Ali |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Clustering Aided Weakly Supervised Training to Detect Anomalous Events in Surveillance VideosabstractFormulating learning systems for the detection of real-world anomalous events using only video-level labels is a challenging task mainly due to the presence of noisy labels as well as the rare occurrence of anomalous events in the training data. We propose a weakly supervised anomaly detection system that has multiple contributions including a random batch selection mechanism to reduce interbatch correlation and a normalcy suppression block (NSB) which learns to minimize anomaly scores over normal regions of a video by utilizing the overall information available in a training batch. In addition, a clustering loss block (CLB) is proposed to mitigate the label noise and to improve the representation learning for the anomalous and normal regions. This block encourages the backbone network to produce two distinct feature clusters representing normal and anomalous events. An extensive analysis of the proposed approach is provided using three popular anomaly detection datasets including UCF-Crime, ShanghaiTech, and UCSD Ped2. The experiments demonstrate the superior anomaly detection capability of our approach. Muhammad Zaigham Zaheer, Arif Mahmood, Marcella Astrid, Seung-Ik Lee |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Unsupervised Landmark Discovery Using Consistency-Guided Bottleneck
Mamona Awan, Muhammad Haris Khan, Sanoojan Baliah, Muhammad Ahmad Waseem, Salman Khan 0001, Fahad Shahbaz Khan, Arif Mahmood |
BMVC | 7 |
| 2023 | Higher-Order Sparse Convolutions in Graph Neural NetworksabstractGraph Neural Networks (GNNs) have been applied to many problems in computer sciences. Capturing higher-order relationships between nodes is crucial to increase the expressive power of GNNs. However, existing methods to capture these relationships could be infeasible for large-scale graphs. In this work, we introduce a new higher-order sparse convolution based on the Sobolev norm of graph signals. Our Sparse Sobolev GNN (S-SobGNN) computes a cascade of filters on each layer with increasing Hadamard powers to get a more diverse set of functions, and then a linear combination layer weights the embeddings of each filter. We evaluate S-SobGNN in several applications of semi-supervised learning. S-SobGNN shows competitive performance in all applications as compared to several state-of-the-art methods. Jhony-Heriberto Giraldo-Zuluaga, Sajid Javed, Arif Mahmood, Fragkiskos D. Malliaros, Thierry Bouwmans |
ICASSP | 3 |
| 2023 | Single-branch Network for Multimodal TrainingabstractWith the rapid growth of social media platforms, users are sharing billions of multimedia posts containing audio, images, and text. Researchers have focused on building autonomous systems capable of processing such multimedia data to solve challenging multimodal tasks including cross-modal retrieval, matching, and verification. Existing works use separate networks to extract embeddings of each modality to bridge the gap between them. The modular structure of their branched networks is fundamental in creating numerous multimodal applications and has become a defacto standard to handle multiple modalities. In contrast, we propose a novel single-branch network capable of learning discriminative representation of unimodal as well as multimodal tasks without changing the network. An important feature of our single-branch network is that it can be trained either using single or multiple modalities without sacrificing performance. We evaluated our proposed single-branch network on the challenging multimodal problem (face-voice association) for cross-modal verification and matching tasks with various loss formulations. Experimental results demonstrate the superiority of our proposed single-branch network over the existing methods in a wide range of experiments. Code: https://github.com/msaadsaeed/SBNet Muhammad Saad Saeed, Shah Nawaz, Muhammad Haris Khan, Muhammad Zaigham Zaheer, Karthik Nandakumar, Muhammad Haroon Yousaf, Arif Mahmood |
ICASSP | 7 |
| 2023 | MACC Net: Multi-task attention crowd counting network
Sahar Aldhaheri, Reem Alotaibi, Bandar Ahmed Alzahrani, Anas Hadi, Arif Mahmood, Areej Alhothali, Ahmed Barnawi |
Appl. Intell. | 5 |
| 2023 | Population affinity propagation approach for points of dispensing location allocation
Nusaybah Alghanmi, Reem Alotaibi, Sultanah Alshammari, Arif Mahmood |
Appl. Intell. | 4 |
| 2023 | Learning Structure Aware Deep Spectral EmbeddingabstractSpectral Embedding (SE) has often been used to map data points from non-linear manifolds to linear subspaces for the purpose of classification and clustering. Despite significant advantages, the subspace structure of data in the original space is not preserved in the embedding space. To address this issue subspace clustering has been proposed by replacing the SE graph affinity with a self-expression matrix. It works well if the data lies in a union of linear subspaces however, the performance may degrade in real-world applications where data often spans non-linear manifolds. To address this problem we propose a novel structure-aware deep spectral embedding by combining a spectral embedding loss and a structure preservation loss. To this end, a deep neural network architecture is proposed that simultaneously encodes both types of information and aims to generate structure-aware spectral embedding. The subspace structure of the input data is encoded by using attention-based self-expression learning. The proposed algorithm is evaluated on six publicly available real-world datasets. The results demonstrate the excellent clustering performance of the proposed algorithm compared to the existing state-of-the-art methods. The proposed algorithm has also exhibited better generalization to unseen data points and it is scalable to larger datasets without requiring significant computational resources. Hira Yaseen, Arif Mahmood |
IEEE Trans. Image Process. | 2 |
| 2023 | Knowledge Distillation in Histology Landscape by Multi-Layer Features SupervisionabstractAutomatic tissue classification is a fundamental task in computational pathology for profiling tumor micro-environments. Deep learning has advanced tissue classification performance at the cost of significant computational power. Shallow networks have also been end-to-end trained using direct supervision however their performance degrades because of the lack of capturing robust tissue heterogeneity. Knowledge distillation has recently been employed to improve the performance of the shallow networks used as student networks by using additional supervision from deep neural networks used as teacher networks. In the current work, we propose a novel knowledge distillation algorithm to improve the performance of shallow networks for tissue phenotyping in histology images. For this purpose, we propose multi-layer feature distillation such that a single layer in the student network gets supervision from multiple teacher layers. In the proposed algorithm, the size of the feature map of two layers is matched by using a learnable multi-layer perceptron. The distance between the feature maps of the two layers is then minimized during the training of the student network. The overall objective function is computed by summation of the loss over multiple layers combination weighted with a learnable attention-based parameter. The proposed algorithm is named as Knowledge Distillation for Tissue Phenotyping (KDTP). Experiments are performed on five different publicly available histology image classification datasets using several teacher-student network combinations within the KDTP algorithm. Our results demonstrate a significant performance increase in the student networks by using the proposed KDTP algorithm compared to direct supervision-based training methods. Sajid Javed, Arif Mahmood, Talha Qaiser, Naoufel Werghi |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Quantification of Occlusion Handling Capability of a 3D Human Pose Estimation Frameworkabstract3D human pose estimation using monocular images is an important yet challenging task. Existing 3D pose detection methods exhibit excellent performance under normal conditions however their performance may degrade due to occlusion. Recently some occlusion aware methods have also been proposed, however, the occlusion handling capability of these networks has not yet been thoroughly investigated. In the current work, we propose an occlusion-guided 3D human pose estimation framework and quantify its occlusion handling capability by using different protocols. The proposed method estimates more accurate 3D human poses using 2D skeletons with missing joints as input. Missing joints are handled by introducing occlusion guidance that provides extra information about the absence or presence of a joint. Temporal information has also been exploited to better estimate the missing joints. A large number of experiments are performed for the quantification of occlusion handling capability of the proposed method on three publicly available datasets in various settings including random missing joints, fixed body parts missing, and complete frames missing, using mean per joint position error criterion. In addition to that, the quality of the predicted 3D poses is also evaluated using action classification performance as a criterion. 3D poses estimated by the proposed method achieved significantly improved action recognition performance in the presence of missing joints. Our experiments demonstrate the effectiveness of the proposed framework for handling the missing joints as well as quantification of the occlusion handling capability of the deep neural networks. Mehwish Ghafoor, Arif Mahmood |
IEEE Trans. Multim. | 2 |
| 2022 | Face Pyramid Vision Transformer
Khawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood |
BMVC | 3 |
| 2022 | Generative Cooperative Learning for Unsupervised Video Anomaly DetectionabstractVideo anomaly detection is well investigated in weakly-supervised and one-class classification (OCC) settings. However, unsupervised video anomaly detection methods are quite sparse, likely because anomalies are less frequent in occurrence and usually not well-defined, which when coupled with the absence of ground truth supervision, could adversely affect the performance of the learning algorithms. This problem is challenging yet rewarding as it can completely eradicate the costs of obtaining laborious annotations and enable such systems to be deployed without human intervention. To this end, we propose a novel unsupervised Generative Cooperative Learning (GCL) approach for video anomaly detection that exploits the low frequency of anomalies towards building a cross-supervision between a generator and a discriminator. In essence, both networks get trained in a cooperative fashion, thereby allowing unsupervised learning. We conduct extensive experiments on two large-scale video anomaly detection datasets, UCF crime and ShanghaiTech. Consistent improvement over the existing state-of-the-art unsupervised and OCC methods corroborate the effectiveness of our approach. Muhammad Zaigham Zaheer, Arif Mahmood, Muhammad Haris Khan, Mattia Segù, Fisher Yu 0001, Seung-Ik Lee |
CVPR | 2 |
| 2022 | Learning to localize image forgery using end-to-end attention network
Iyyakutti Iyappan Ganapathi, Sajid Javed, Syed Sadaf Ali, Arif Mahmood, Ngoc-Son Vu, Naoufel Werghi |
Neurocomputing | 4 |
| 2022 | Moving objects segmentation using generative adversarial modeling
Maryam Sultana, Arif Mahmood, Thierry Bouwmans, Muhammad Haris Khan, Soon Ki Jung |
Neurocomputing | 2 |
| 2022 | AAC: Automatic Augmentation for Crowd Counting
Rui Wang 0077, Reem Alotaibi, Bander A. Alzahrani, Arif Mahmood, Gaoxiang Wu, Abeer Alshehri, Sahar Aldhaheri |
Neurocomputing | 4 |
| 2022 | Multi-scale attention guided network for end-to-end face alignment and recognition
M. Saad Shakeel, Wenxiong Kang, Arif Mahmood |
J. Vis. Commun. Image Represent. | 5 |
| 2022 | Nucleus classification in histology images using message passing network
Taimur Hassan, Sajid Javed, Arif Mahmood, Talha Qaiser, Naoufel Werghi, Nasir M. Rajpoot |
Medical Image Anal. | 3 |
| 2022 | A boosting framework for human posture recognition using spatio-temporal features along with radon transform
Salma Aftab, Syed Farooq Ali, Arif Mahmood, Umar Suleman |
Multim. Tools Appl. | 3 |
| 2022 | Fake visual content detection using two-stream convolutional neural networks
Bilal Yousaf, Waqas Sultani, Arif Mahmood, Junaid Qadir 0001 |
Neural Comput. Appl. | 4 |
| 2022 | Unsupervised moving object segmentation using background subtraction and optimal adversarial noise sample search
Maryam Sultana, Arif Mahmood, Soon Ki Jung |
Pattern Recognit. | 2 |
| 2022 | Predictive Auto-Scaling of Multi-Tier Applications Using Performance Varying Cloud ResourcesabstractThe performance of the same type of cloud resources, such as virtual machines (VMs), varies over time mainly due to hardware heterogeneity, resource contention among co-located VMs, and virtualization overhead. The performance variation can be significant, introducing challenges to learn workload-specific resource provisioning policies to automatically scale the cloud-hosted applications to maintain the desired response time. Moreover, auto-scaling multi-tier applications using minimal resources is even more challenging because bottlenecks may occur on multiple tiers concurrently. In this paper, we address the problem of using performance varying VMs for gracefully auto-scaling a multi-tier application using minimal resources to handle dynamically increasing workloads and satisfy the response time requirements. The proposed system uses a supervised learning method to identify the appropriate resources provisioning for multi-tier applications based on the prediction of the application response time and the request arrival rate. The supervised learning method learns a state transition configuration map which encodes a resource allocation states invariant to the underlying VMs performance variations. This configuration map helps to use performance varying resources in predictive autoscaling method. Our experimental evaluation using a real-world multi-tier web application hosted on a public cloud shows an improved application performance with minimal resources compared to conventional predictive auto-scaling methods. Waheed Iqbal, Abdelkarim Erradi, Muhammad Abdullah 0004, Arif Mahmood |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | Hierarchical Spatiotemporal Graph Regularized Discriminative Correlation Filter for Visual Object TrackingabstractVisual object tracking is a fundamental and challenging task in many high-level vision and robotics applications. It is typically formulated by estimating the target appearance model between consecutive frames. Discriminative correlation filters (DCFs) and their variants have achieved promising speed and accuracy for visual tracking in many challenging scenarios. However, because of the unwanted boundary effects and lack of geometric constraints, these methods suffer from performance degradation. In the current work, we propose hierarchical spatiotemporal graph-regularized correlation filters for robust object tracking. The target sample is decomposed into a large number of deep channels, which are then used to construct a spatial graph such that each graph node corresponds to a particular target location across all channels. Such a graph effectively captures the spatial structure of the target object. In order to capture the temporal structure of the target object, the information in the deep channels obtained from a temporal window is compressed using the principal component analysis, and then, a temporal graph is constructed such that each graph node corresponds to a particular target location in the temporal dimension. Both spatial and temporal graphs span different subspaces such that the target and the background become linearly separable. The learned correlation filter is constrained to act as an eigenvector of the Laplacian of these spatiotemporal graphs. We propose a novel objective function that incorporates these spatiotemporal constraints into the DCFs framework. We solve the objective function using alternating direction methods of multipliers such that each subproblem has a closed-form solution. We evaluate our proposed algorithm on six challenging benchmark datasets and compare it with 33 existing state-of-the art trackers. Our results demonstrate an excellent performance of the proposed algorithm compared to the existing trackers. Sajid Javed, Arif Mahmood, Jorge Dias 0001, Lakmal D. Seneviratne, Naoufel Werghi |
IEEE Trans. Cybern. | 2 |
| 2022 | Stabilizing Adversarially Learned One-Class Novelty Detection Using Pseudo AnomaliesabstractRecently, anomaly scores have been formulated using reconstruction loss of the adversarially learned generators and/or classification loss of discriminators. Unavailability of anomaly examples in the training data makes optimization of such networks challenging. Attributed to the adversarial training, performance of such models fluctuates drastically with each training step, making it difficult to halt the training at an optimal point. In the current study, we propose a robust anomaly detection framework that overcomes such instability by transforming the fundamental role of the discriminator from identifying real vs. fake data to distinguishing good vs. bad quality reconstructions. For this purpose, we propose a method that utilizes the current state as well as an old state of the same generator to create good and bad quality reconstruction examples. The discriminator is trained on these examples to detect the subtle distortions that are often present in the reconstructions of anomalous data. In addition, we propose an efficient generic criterion to stop the training of our model, ensuring elevated performance. Extensive experiments performed on six datasets across multiple domains including image and video based anomaly detection, medical diagnosis, and network security, have demonstrated excellent performance of our approach. Muhammad Zaigham Zaheer, Jin Ha Lee 0002, Arif Mahmood, Marcella Astrid, Seung-Ik Lee |
IEEE Trans. Image Process. | 3 |
| 2022 | An End-to-End Human Abnormal Behavior Recognition Framework for Crowds With Mentally Disordered IndividualsabstractAbnormal or violent behavior by people with mental disorders is common. When individuals with mental disorders exhibit abnormal behavior in public places, they may cause physical and mental harm to others as well as to themselves. Thus, it is necessary to monitor their behavior using visual surveillance systems. However, it is challenging to automatically detect human abnormal behavior (especially for individuals with mental disorders) based on motion recognition technologies. To address these issues, in the current work, we propose an end-to-end abnormal behaviour detection framework from a new perspective in conjunction with the Graph Convolutional Network (GCN) and a 3D Convolutional Neural Network (3DCNN). Specifically, we first train a one-class classifier to extract features and estimate abnormality scores. To improve the performance of abnormal behavior detection, GCN is used to model the similarity between video clips for the correction of noisy labels. Then, based on this framework, GCN recognizes the normal behavior clips in the abnormal video and removes them, while the clips identified as abnormal behavior are retained. Finally, a 3D CNN is used to extract spatiotemporal features to classify different abnormal behaviors. In order to better detect the violent behavior of individuals with mental disorders, the paper focuses on the UCF-Crime dataset with various types of violent behaviors. By experimenting with this dataset, the classification accuracy reaches 37.9%, which is significantly better than that of the current state-of-the-art approaches. Yixue Hao, Zaiyang Tang, Bander A. Alzahrani, Reem Alotaibi, Reem Alharthi, Miaomiao Zhao, Arif Mahmood |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Leveraging orientation for weakly supervised object detection with application to firearm localization
Javed Iqbal 0007, Muhammad Akhtar Munir, Arif Mahmood, Afsheen Rafaqat Ali, Mohsen Ali |
Neurocomputing | 3 |
| 2021 | 4G-VOS: Video Object Segmentation using guided context embedding
Mustansar Fiaz, Muhammad Zaigham Zaheer, Arif Mahmood, Seung-Ik Lee, Soon Ki Jung |
Knowl. Based Syst. | 3 |
| 2021 | Spatially Constrained Context-Aware Hierarchical Deep Correlation Filters for Nucleus Detection in Histology Images
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi, Nasir M. Rajpoot |
Medical Image Anal. | 2 |
| 2021 | Statistically correlated multi-task learning for autonomous driving
Waseem Abbas 0002, Murtaza Taj, Arif Mahmood |
Neural Comput. Appl. | 4 |
| 2021 | Human face super-resolution on poor quality surveillance video footage
Matthew N. Dailey, Arif Mahmood, Jednipat Moonrinta, Mongkol Ekpanyapong |
Neural Comput. Appl. | 3 |
| 2021 | Unsupervised Moving Object Detection in Complex Scenes Using Adversarial RegularizationsabstractMoving object detection (MOD) is a fundamental step in many high-level vision-based applications, such as human activity analysis, visual object tracking, autonomous vehicles, surveillance, and security. Most of the existing MOD algorithms observe performance degradation in the presence of complex scenes containing camouflage objects, shadows, dynamic backgrounds, and varying illumination conditions, and captured by static cameras. To appropriately handle these challenges, we propose a Generative Adversarial Network (GAN) based on a moving object detection algorithm, called MOD_GAN. In the proposed algorithm, scene-specific GANs are trained in an unsupervised MOD setting, thereby enabling the algorithm to learn generating background sequences using input from uniformly distributed random noise samples. In addition to adversarial loss, during training, norm-based loss in the image space and discriminator feature-space is also minimized between the generated images and the training data. The additional losses enable the generator to learn subtle background details, resulting in a more realistic complex scene generation. During testing, a novel back-propagation based algorithm is used to generate images with statistics similar to the test images. More appropriate random noise samples are searched by directly minimizing the loss function between the test and generated images both in the image and discriminator feature-spaces. The network is not updated in this step; only the input noise samples are iteratively modified to minimize the loss function. Moreover, motion information is used to ensure that this loss is only computed on small-motion pixels. A novel dataset containing outdoor time-lapsed images from dawn to dusk with a full illumination variation cycle is also proposed to better compare the MOD algorithms in outdoor scenes. Accordingly, extensive experiments on five benchmark datasets and comparison with 30 existing methods demonstrate the strength of the proposed algorithm. Maryam Sultana, Arif Mahmood, Soon Ki Jung |
IEEE Trans. Multim. | 2 |
| 2021 | Web Application Resource Requirements Estimation Based on the Workload Latent FeaturesabstractMost cloud computing platforms offer reactive resource auto-scaling mechanisms for dealing with variable traffic patterns to deliver the desired QoS properties while keeping low provisioning costs. However, a range of scenarios have not been fully addressed by the current auto-scaling solutions, particularly dealing with a rapid increase in workload and the risk of thrashing due to frequent workload variations. A reactive system is vulnerable in such conditions. Realizing the full potential of auto-scaling still remains challenging particularly due to the need of accurately estimating the application resource requirements for time-varying workload patterns. In this work, we propose and evaluate a novel method using only application access logs to estimate more accurately the hardware resource demands and application response time. In particular, we propose novel workload latent features which we compute by applying unsupervised learning on the access logs. We use these latent features to estimate the application hardware resource requirements and response time for various workload patterns. We evaluate the proposed method using multiple benchmark web applications and compare it with current state-of-the-art. Extensive experimental evaluations show an excellent performance of our proposed workload latent features in estimating response time, CPU, memory, and bandwidth utilization. Abdelkarim Erradi, Waheed Iqbal, Arif Mahmood, Athman Bouguettaya |
IEEE Trans. Serv. Comput. | 3 |
| 2020 | CLAWS: Clustering Assisted Weakly Supervised Learning with Normalcy Suppression for Anomalous Event Detection
Muhammad Zaigham Zaheer, Arif Mahmood, Marcella Astrid, Seung-Ik Lee |
ECCV (22) | 2 |
| 2020 | Localizing Firearm Carriers By Identifying Human-Object PairsabstractVisual identification of gunmen in a crowd is a challenging problem, that requires resolving the association of a person with an object (firearm). We present a novel approach to address this problem, by defining human-object interaction (and non-interaction) bounding boxes. In a given image, human and firearms are separately detected. Each detected human is paired with each detected firearm, allowing us to create a paired bounding box that contains both object and the human. A network is trained to classify these paired-bounding-boxes into human carrying the identified firearm or not. Extensive experiments were performed to evaluate the effectiveness of the algorithm, including exploiting full pose of the human, hand-keypoints, and their association with the firearm. The knowledge of spatially localized features is key to the success of our method by using multi-size proposals with adaptive average pooling. We have also extended a previously existing firearm detection dataset, by adding more images and tagging in the extended dataset the human-firearm pairs (including bounding boxes for firearms and gunmen). The experimental results $({78.5 AP}_{hold})$ demonstrate effectiveness of the proposed method. Abdul Basit 0019, Muhammad Akhtar Munir, Mohsen Ali, Naoufel Werghi, Arif Mahmood |
ICIP | 5 |
| 2020 | CS-RPCA: Clustered Sparse RPCA for Moving Object DetectionabstractMoving object detection (MOD) is an important step for many computer vision applications. In the last decade, it is evident that RPCA has shown to be a potential solution for MOD and achieved a promising performance under various challenging background scenes. However, because of the lack of different types of features, RPCA still shows degraded performance in many complicated background scenes such as dynamic backgrounds, cluttered foreground objects, and camouflage. To address these problems, this paper presents a Clustered Sparse RPCA (CS-RPCA) for MOD under challenging environments. The proposed algorithm extracts multiple features from video sequences and then employs RPCA to get the low-rank and sparse component from each representation. The sparse subspaces are then emerged into a common sparse component using Grassmann manifold. We proposed a novel objective function which computes the composite sparse component from multiple representations and it is solved using non-negative matrix factorization method. The proposed algorithm is evaluated on two challenging datasets for MOD. Results demonstrate excellent performance of the proposed algorithm as compared to existing state-of-the-art methods. Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi |
ICIP | 2 |
| 2020 | Dynamic Background Subtraction Using Least Square Adversarial LearningabstractDynamic Background Subtraction (BS) is a fundamental problem in many vision-based applications. BS in real complex environments has several challenging conditions like illumination variations, shadows, camera jitters, and bad weather. In this study, we aim to address the challenges of BS in complex scenes by exploiting conditional least squares adversarial networks. During training, a scene-specific conditional least squares adversarial network with two additional regularizations including L1-Loss and Perceptual-Loss is employed to learn the dynamic background variations. The given input to the model is video frames conditioned on corresponding ground truth to learn the dynamic changes in complex scenes. Afterwards, testing is performed on unseen test video frames so that the generator would conduct dynamic background subtraction. The proposed method consisting of three loss-terms including least squares adversarial loss, L1-Loss and Perceptual-Loss is evaluated on two benchmark datasets CDnet2014 and BMC. The results of our proposed method show improved performance on both datasets compared with 10 existing state-of-the-art methods. Maryam Sultana, Arif Mahmood, Thierry Bouwmans, Soon Ki Jung |
ICIP | 2 |
| 2020 | Ocean Color Net (OCN) for the Barents SeaabstractOver recent years, rapid environmental changes in the Arctic and subarctic regions have caused significant alterations in the ecosystem structure and seasonality, including the primary productivity of the Barents Sea. This work aims at improving methodology for studying these features, by estimating chlorophyll-a (chl-a) concentrations in the transitional Barents Sea by remotely sensing its optical properties, in order to better understand the large-scale algal bloom dynamics in the region. The in-situ measurements of chl-a are collected from the year 2016 to 2018 over a wide area of the Barents Sea to cover the spatial and temporal variations in chl-a concentration. Optical images of the Barents Sea are captured by the Multi-Spectral Imager Instrument on Sentinel-2. Using these remotely sensed optical images and the in-situ measurements, we propose a match-up dataset creation method based on the distribution of the remotely sensed reflectance spectra. Different Machine Learning (ML) techniques are assessed to estimate concentration of chl-a using the match-up dataset. Most of these techniques have not been investigated before in the subarctic region such as the Barents Sea. The Ocean Color Net (OCN) regression model proposed in this study has outperformed other ML-based techniques including Support Vector Regression, Gaussian Process Regression, and the globally trained Case-2 Regional/Coast Colour (C2RCC) processing chain model C2RCC-Nets, as well as empirical methods based on spectral band ratios. A wide range of experiments has demonstrated the effectiveness of the proposed OCN for ocean color remote sensing in the subarctic region. The performance of the OCN is also presented spatially by computing chl-a maps in the Barents Sea. Muhammad Asim 0003, Camilla Brekke, Arif Mahmood, Torbjørn Eltoft, Marit Reigstad |
IGARSS | 3 |
| 2020 | Masked Linear Regression for Learning Local Receptive Fields for Facial Expression Synthesis
Nazar Khan, Arbish Akram, Arif Mahmood, Sania Ashraf, Kashif Murtaza |
Int. J. Comput. Vis. | 3 |
| 2020 | Cellular community detection for tissue phenotyping in colorectal cancer histology images
Sajid Javed, Arif Mahmood, Muhammad Moazam Fraz, Navid Alemi Koohbanani, Ksenija Benes, Yee-Wah Tsang, Katherine Hewitt, David B. A. Epstein, David R. J. Snead, Nasir M. Rajpoot |
Medical Image Anal. | 2 |
| 2020 | A Self-Reasoning Framework for Anomaly Detection Using Video-Level LabelsabstractAnomalous event detection in surveillance videos is a challenging and practical research problem among image and video processing community. Compared to the frame-level annotations of anomalous events, obtaining video-level annotations is quite fast and cheap though such high-level labels may contain significant noise. More specifically, an anomalous labeled video may actually contain anomaly only in a short duration while the rest of the video frames may be normal. In the current work, we propose a weakly supervised anomaly detection framework based on deep neural networks which is trained in a self-reasoning fashion using only video-level labels. To carry out the self-reasoning based training, we generate pseudo labels by using binary clustering of spatio-temporal video features which helps in mitigating the noise present in the labels of anomalous videos. Our proposed formulation encourages both the main network and the clustering to complement each other in achieving the goal of more accurate anomaly detection. The proposed framework has been evaluated on publicly available real-world anomaly detection datasets including UCF-crime, ShanghaiTech and UCSD Ped2. The experiments demonstrate superiority of our proposed framework over the current state-of-the-art methods. Muhammad Zaigham Zaheer, Arif Mahmood, Hochul Shin, Seung-Ik Lee |
IEEE Signal Process. Lett. | 2 |
| 2020 | Robust Structural Low-Rank TrackingabstractVisual object tracking is an essential task for many computer vision applications. It becomes very challenging when the target appearance changes especially in the presence of occlusion, background clutter, and sudden illumination variations. Methods, that incorporate sparse representation and low-rank assumptions on the target particles have achieved promising results. However, because of the lack of structural constraints, these methods show performance degradation when facing the aforementioned challenges. To alleviate these limitations, we propose a new structural low-rank modeling algorithm for robust object tracking in complex scenarios. In the proposed algorithm, we consider spatial and temporal appearance consistency constraints, among the particles in the low-rank subspace, embedded in four different graphs. The resulting objective function encoding these constraints is novel and it is solved using linearized alternating direction method with adaptive penalty both in batch fashion as well as in online fashion. Our proposed objective function jointly learns the spatial and temporal structure of the target particles in consecutive frames and makes the proposed tracker consistent against many complex tracking scenarios. Results on four challenging datasets demonstrate excellent performance of the proposed algorithm as compared to current state-of-the-art methods. Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi |
IEEE Trans. Image Process. | 2 |
| 2020 | Multiplex Cellular Communities in Multi-Gigapixel Colorectal Cancer Histology Images for Tissue PhenotypingabstractIn computational pathology, automated tissue phenotyping in cancer histology images is a fundamental tool for profiling tumor microenvironments. Current tissue phenotyping methods use features derived from image patches which may not carry biological significance. In this work, we propose a novel multiplex cellular community-based algorithm for tissue phenotyping integrating cell-level features within a graph-based hierarchical framework. We demonstrate that such integration offers better performance compared to prior deep learning and texture-based methods as well as to cellular community based methods using uniplex networks. To this end, we construct celllevel graphs using texture, alpha diversity and multi-resolution deep features. Using these graphs, we compute cellular connectivity features which are then employed for the construction of a patch-level multiplex network. Over this network, we compute multiplex cellular communities using a novel objective function. The proposed objective function computes a low-dimensional subspace from each cellular network and subsequently seeks a common low-dimensional subspace using the Grassmann manifold. We evaluate our proposed algorithm on three publicly available datasets for tissue phenotyping, demonstrating a significant improvement over existing state-of-the-art methods. Sajid Javed, Arif Mahmood, Naoufel Werghi, Ksenija Benes, Nasir M. Rajpoot |
IEEE Trans. Image Process. | 2 |
| 2019 | Structural Low-Rank TrackingabstractVisual object tracking is an important step for many computer vision applications. The task becomes very challenging when the target undergoes heavy occlusion, background clutters, and sudden illumination variations. Methods that incorporate sparse representation and low-rank assumptions on the target particles have achieved promising results. However, because of the lack of structural constraints, these methods show performance degradation when an object faces the aforementioned challenges. To alleviate these limitations, we propose a new structural low-rank modeling algorithm for robust object tracking. In the proposed algorithm, we enforce local spatial, global spatial and temporal appearance consistency among the particles in the low-rank subspace by constructing three graphs. The Laplacian matrices of these graphs are incorporated into the novel low-rank objective function which is solved using linearized alternating direction method with an adaptive penalty. Our proposed objective function jointly learns the spatial, global, and temporal structure of the target particles in consecutive frames and makes the proposed tracker consistent against many complex tracking scenarios. Results on two challenging benchmark datasets show the superiority of the proposed algorithm as compared to current state-of-the-art methods. Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi |
AVSS | 2 |
| 2019 | Action recognition in poor-quality spectator crowd videos using head distribution-based person segmentation
Arif Mahmood, Somaya Al-Máadeed |
Mach. Vis. Appl. | 1 |
| 2019 | Unsupervised deep context prediction for background estimation and foreground segmentation
Maryam Sultana, Arif Mahmood, Sajid Javed, Soon Ki Jung |
Mach. Vis. Appl. | 2 |
| 2019 | Moving Object Detection in Complex Scene Using Spatiotemporal Structured-Sparse RPCAabstractMoving object detection is a fundamental step in various computer vision applications. Robust Principal Component Analysis (RPCA) based methods have often been employed for this task. However, the performance of these methods deteriorates in the presence of dynamic background scenes, camera jitter, camouflaged moving objects, and/or variations in illumination. It is because of an underlying assumption that the elements in the sparse component are mutually independent, and thus the spatiotemporal structure of the moving objects is lost. To address this issue, we propose a spatiotemporal structured sparse RPCA algorithm for moving objects detection, where we impose spatial and temporal regularization on the sparse component in the form of graph Laplacians. Each Laplacian corresponds to a multi-feature graph constructed over superpixels in the input matrix. We enforce the sparse component to act as eigenvectors of the spatial and temporal graph Laplacians while minimizing the RPCA objective function. These constraints incorporate a spatiotemporal subspace structure within the sparse component. Thus, we obtain a novel objective function for separating moving objects in the presence of complex backgrounds. The proposed objective function is solved using a linearized alternating direction method of multipliers based batch optimization. Moreover, we also propose an online optimization algorithm for real-time applications. We evaluated both the batch and online solutions using six publicly available datasets that included most of the aforementioned challenges. Our experiments demonstrated the superior performance of the proposed algorithms compared with the current state-of-the-art methods. Sajid Javed, Arif Mahmood, Somaya Al-Máadeed, Thierry Bouwmans, Soon Ki Jung |
IEEE Trans. Image Process. | 2 |
| 2018 | Dynamic workload patterns prediction for proactive auto-scaling of web applications
Waheed Iqbal, Abdelkarim Erradi, Arif Mahmood |
J. Netw. Comput. Appl. | 3 |
| 2018 | Spatiotemporal Low-Rank Modeling for Complex Scene Background InitializationabstractBackground modeling constitutes the building block of many computer-vision tasks. Traditional schemes model the background as a low rank matrix with corrupted entries. These schemes operate in batch mode and do not scale well with the data size. Moreover, without enforcing spatiotemporal information in the low-rank component, and because of occlusions by foreground objects and redundancy in video data, the design of a background initialization method robust against outliers is very challenging. To overcome these limitations, this paper presents a spatiotemporal low-rank modeling method on dynamic video clips for estimating the robust background model. The proposed method encodes spatiotemporal constraints by regularizing spectral graphs. Initially, a motion-compensated binary matrix is generated using optical flow information to remove redundant data and to create a set of dynamic frames from the input video sequence. Then two graphs are constructed, one between frames for temporal consistency and the other between features for spatial consistency, to encode the local structure for continuously promoting the intrinsic behavior of the low-rank model against outliers. These two terms are then incorporated in the iterative Matrix Completion framework for improved segmentation of background. Rigorous evaluation on severely occluded and dynamic background sequences demonstrates the superior performance of the proposed method over state-of-the-art approaches. Sajid Javed, Arif Mahmood, Thierry Bouwmans, Soon Ki Jung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Improving security surveillance by hidden cameras
Hadia Tazeem, Muhammad Shahid Farid, Arif Mahmood |
Multim. Tools Appl. | 3 |
| 2017 | Background-Foreground Modeling Based on Spatiotemporal Sparse Subspace ClusteringabstractBackground estimation and foreground segmentation are important steps in many high-level vision tasks. Many existing methods estimate background as a low-rank component and foreground as a sparse matrix without incorporating the structural information. Therefore, these algorithms exhibit degraded performance in the presence of dynamic backgrounds, photometric variations, jitter, shadows, and large occlusions. We observe that these backgrounds often span multiple manifolds. Therefore, constraints that ensure continuity on those manifolds will result in better background estimation. Hence, we propose to incorporate the spatial and temporal sparse subspace clustering into the robust principal component analysis (RPCA) framework. To that end, we compute a spatial and temporal graph for a given sequence using motion-aware correlation coefficient. The information captured by both graphs is utilized by estimating the proximity matrices using both the normalized Euclidean and geodesic distances. The low-rank component must be able to efficiently partition the spatiotemporal graphs using these Laplacian matrices. Embedded with the RPCA objective function, these Laplacian matrices constrain the background model to be spatially and temporally consistent, both on linear and nonlinear manifolds. The solution of the proposed objective function is computed by using the linearized alternating direction method with adaptive penalty optimization scheme. Experiments are performed on challenging sequences from five publicly available datasets and are compared with the 23 existing state-of-the-art methods. The results demonstrate excellent performance of the proposed algorithm for both the background estimation and foreground segmentation. Sajid Javed, Arif Mahmood, Thierry Bouwmans, Soon Ki Jung |
IEEE Trans. Image Process. | 2 |
| 2017 | Using Geodesic Space Density Gradients for Network Community DetectionabstractMany real world complex systems naturally map to network data structures instead of geometric spaces because the only available information is the presence or absence of a link between two entities in the system. To enable data mining techniques to solve problems in the network domain, the nodes need to be mapped to a geometric space. We propose this mapping by representing each network node with its geodesic distances from all other nodes. The space spanned by the geodesic distance vectors is the geodesic space of that network. The position of different nodes in the geodesic space encode the network structure. In this space, considering a continuous density field induced by each node, density at a specific point is the summation of density fields induced by all nodes. We drift each node in the direction of positive density gradient using an iterative algorithm till each node reaches a local maximum. Due to the network structure captured by this space, the nodes that drift to the same region of space belong to the same communities in the original network. We use the direction of movement and final position of each node as important clues for community membership assignment. The proposed algorithm is compared with more than 10 state-of-the-art community detection techniques on two benchmark networks with known communities using Normalized Mutual Information criterion. The proposed algorithm outperformed these methods by a significant margin. Moreover, the proposed algorithm has also shown excellent performance on many real-world networks. Arif Mahmood, Michael Small, Somaya Al-Máadeed, Nasir M. Rajpoot |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Subspace based network community detection using sparse linear codingabstractInformation mining from networks by identifying communities is an important problem across a number of research fields including social science, biology, physics, and medicine. Most existing community detection algorithms are graph theoretic and lack the ability to detect accurate community boundaries if the ratio of intra-community to inter-community links is low. Also, algorithms based on modularity maximization may fail to resolve communities smaller than a specific size if the community size varies significantly. We propose a fundamentally different community detection algorithm based on the fact that each network community spans a different subspace in the geodesic space. Therefore, each node can only be efficiently represented as a linear combination of nodes spanning the same subspace (Fig. 1). To make the process of community detection more robust, we use sparse linear coding with ł1 norm constraint. In order to find a community label for each node, sparse spectral clustering algorithm is used. The proposed community detection technique is compared with more than ten state of the art methods on two benchmark networks (with known clusters) using normalized mutual information criterion. Our proposed algorithm outperformed existing methods with a significant margin on both benchmark networks. Arif Mahmood, Michael Small |
ICDE | 1 |
| 2016 | Motion-Aware Graph Regularized RPCA for background modeling of complex scenesabstractComputing a background model from a given sequence of video frames is a prerequisite for many computer vision applications. Recently, this problem has been posed as learning a low-dimensional subspace from high dimensional data. Many contemporary subspace segmentation methods have been proposed to overcome the limitations of the methods developed for simple background scenes. Unfortunately, because of the absence of motion information and without preserving intrinsic geometric structure of video data, most existing algorithms do not provide promising nature of the low-rank component for complex scenes. Such as largely occluded background by foreground objects, superfluity in video frames in order to cope with intermittent motion of foreground objects, sudden lighting condition variation, and camera jitter sequences. To overcome these difficulties, we propose a motion-aware regularization of graphs on low-rank component for video background modeling. We compute optical flow and use this information to make a motion-aware matrix. In order to learn the locality and similarity information within a video we compute inter-frame and intra-frame graphs which we use to preserve geometric information in the low-rank component. Finally, we use linearized alternating direction method with parallel splitting and adaptive penalty to incorporate the preceding steps to recover the model of the background. Experimental evaluations on challenging sequences demonstrate promising results over state-of-the-art methods. Sajid Javed, Soon Ki Jung, Arif Mahmood, Thierry Bouwmans |
ICPR | 3 |
| 2016 | Histogram of Oriented Principal Components for Cross-View Action RecognitionabstractExisting techniques for 3D action recognition are sensitive to viewpoint variations because they extract features from depth images which are viewpoint dependent. In contrast, we directly process pointclouds for cross-view action recognition from unknown and unseen views. We propose the histogram of oriented principal components (HOPC) descriptor that is robust to noise, viewpoint, scale and action speed variations. At a 3D point, HOPC is computed by projecting the three scaled eigenvectors of the pointcloud within its local spatio-temporal support volume onto the vertices of a regular dodecahedron. HOPC is also used for the detection of spatio-temporal keypoints (STK) in 3D pointcloud sequences so that view-invariant STK descriptors (or Local HOPC descriptors) at these key locations only are used for action recognition. We also propose a global descriptor computed from the normalized spatio-temporal distribution of STKs in 4-D, which we refer to as STK-D. We have evaluated the performance of our proposed descriptors against nine existing techniques on two cross-view and three single-view human action recognition datasets. The experimental results show that our techniques provide significant improvement over state-of-the-art methods. Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Discriminative human action classification using locality-constrained linear coding
Hossein Rahmani 0001, Du Q. Huynh, Arif Mahmood, Ajmal Mian |
Pattern Recognit. Lett. | 3 |
| 2016 | Constrained Metric Learning by Permutation Inducing IsometriesabstractThe choice of metric critically affects the performance of classification and clustering algorithms. Metric learning algorithms attempt to improve performance, by learning a more appropriate metric. Unfortunately, most of the current algorithms learn a distance function which is not invariant to rigid transformations of images. Therefore, the distances between two images and their rigidly transformed pair may differ, leading to inconsistent classification or clustering results. We propose to constrain the learned metric to be invariant to the geometry preserving transformations of images that induce permutations in the feature space. The constraint that these transformations are isometries of the metric ensures consistent results and improves accuracy. Our second contribution is a dimension reduction technique that is consistent with the isometry constraints. Our third contribution is the formulation of the isometry constrained logistic discriminant metric learning (IC-LDML) algorithm, by incorporating the isometry constraints within the objective function of the LDML algorithm. The proposed algorithm is compared with the existing techniques on the publicly available labeled faces in the wild, viewpoint-invariant pedestrian recognition, and Toy Cars data sets. The IC-LDML algorithm has outperformed existing techniques for the tasks of face recognition, person identification, and object classification by a significant margin. Joel Bosveld, Arif Mahmood, Du Q. Huynh, Lyle Noakes |
IEEE Trans. Image Process. | 2 |
| 2016 | Subspace Based Network Community Detection Using Sparse Linear CodingabstractInformation mining from networks by identifying communities is an important problem across a number of research fields including social science, biology, physics, and medicine. Most existing community detection algorithms are graph theoretic and lack the ability to detect accurate community boundaries if the ratio of intra-community to inter-community links is low. Also, algorithms based on modularity maximization may fail to resolve communities smaller than a specific size if the community size varies significantly. In this paper, we present a fundamentally different community detection algorithm based on the fact that each network community spans a different subspace in the geodesic space. Therefore, each node can only be efficiently represented as a linear combination of nodes spanning the same subspace. To make the process of community detection more robust, we use sparse linear coding with l1norm constraint. In order to find a community label for each node, sparse spectral clustering algorithm is used. The proposed community detection technique is compared with more than 10 state of the art methods on two benchmark networks (with known clusters) using normalized mutual information criterion. Our proposed algorithm outperformed existing algorithms with a significant margin on both benchmark networks. The proposed algorithm has also shown excellent performance on three real-world networks. Arif Mahmood, Michael Small |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Periocular region-based person identification in the visible, infrared and hyperspectral imagery
Arif Mahmood, Ajmal Mian, Chris McDonald |
Neurocomputing | 2 |
| 2015 | Hyperspectral Face Recognition With Spatiospectral Information Fusion and PLS RegressionabstractHyperspectral imaging offers new opportunities for face recognition via improved discrimination along the spectral dimension. However, it poses new challenges, including low signal-to-noise ratio, interband misalignment, and high data dimensionality. Due to these challenges, the literature on hyperspectral face recognition is not only sparse but is limited to ad hoc dimensionality reduction techniques and lacks comprehensive evaluation. We propose a hyperspectral face recognition algorithm using a spatiospectral covariance for band fusion and partial least square regression for classification. Moreover, we extend 13 existing face recognition techniques, for the first time, to perform hyperspectral face recognition.We formulate hyperspectral face recognition as an image-set classification problem and evaluate the performance of seven state-of-the-art image-set classification techniques. We also test six state-of-the-art grayscale and RGB (color) face recognition algorithms after applying fusion techniques on hyperspectral images. Comparison with the 13 extended and five existing hyperspectral face recognition techniques on three standard data sets show that the proposed algorithm outperforms all by a significant margin. Finally, we perform band selection experiments to find the most discriminative bands in the visible and near infrared response spectrum. Arif Mahmood, Ajmal Mian |
IEEE Trans. Image Process. | 2 |
| 2014 | Sparse Kernel Learning for Image Set Classification
Arif Mahmood, Ajmal Mian |
ACCV (2) | 2 |
| 2014 | Semi-supervised Spectral Clustering for Image Set ClassificationabstractWe present an image set classification algorithm based on unsupervised clustering of labeled training and unlabeled test data where labels are only used in the stopping criterion. The probability distribution of each class over the set of clusters is used to define a true set based similarity measure. To this end, we propose an iterative sparse spectral clustering algorithm. In each iteration, a proximity matrix is efficiently recomputed to better represent the local subspace structure. Initial clusters capture the global data structure and finer clusters at the later stages capture the subtle class differences not visible at the global scale. Image sets are compactly represented with multiple Grassmannian manifolds which are subsequently embedded in Euclidean space with the proposed spectral clustering algorithm. We also propose an efficient eigenvector solver which not only reduces the computational cost of spectral clustering by many folds but also improves the clustering quality and final classification results. Experiments on five standard datasets and comparison with seven existing techniques show the efficacy of our algorithm. Arif Mahmood, Ajmal Mian, Robyn A. Owens |
CVPR | 1 |
| 2014 | HOPC: Histogram of Oriented Principal Components of 3D Pointclouds for Action Recognition
Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
ECCV (2) | 2 |
| 2014 | Action Classification with Locality-Constrained Linear CodingabstractWe propose an action classification algorithm which uses Locality-constrained Linear Coding (LLC) to capture discriminative information of human body variations in each spatio-temporal subsequence of a video sequence. Our proposed method divides the input video into equally spaced overlapping spatio-temporal sub sequences, each of which is decomposed into blocks and then cells. We use the Histogram of Oriented Gradient (HOG3D) feature to encode the information in each cell. We justify the use of LLC for encoding the block descriptor by demonstrating its superiority over Sparse Coding (SC). Our sequence descriptor is obtained via a logistic regression classifier with L2 regularization. We evaluate and compare our algorithm with ten state-of-the-art algorithms on five benchmark datasets. Experimental results show that, on average, our algorithm gives better accuracy than these ten algorithms. Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
ICPR | 2 |
| 2014 | Real time action recognition using histograms of depth gradients and random decision forestsabstractWe propose an algorithm which combines the discriminative information from depth images as well as from 3D joint positions to achieve high action recognition accuracy. To avoid the suppression of subtle discriminative information and also to handle local occlusions, we compute a vector of many independent local features. Each feature encodes spatiotemporal variations of depth and depth gradients at a specific space-time location in the action volume. Moreover, we encode the dominant skeleton movements by computing a local 3D joint position difference histogram. For each joint, we compute a 3D space-time motion volume which we use as an importance indicator and incorporate in the feature vector for improved action discrimination. To retain only the discriminant features, we train a random decision forest (RDF). The proposed algorithm is evaluated on three standard datasets and compared with nine state-of-the-art algorithms. Experimental results show that, on the average, the proposed algorithm outperform all other algorithms in accuracy and have a processing speed of over 112 frames/second. Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
WACV | 2 |
| 2014 | An image composition algorithm for handling global visual effects
Muhammad Shahid Farid, Arif Mahmood |
Multim. Tools Appl. | 2 |
| 2013 | Hyperspectral Face Recognition using 3D-DCT and Partial Least SquaresabstractHyperspectral imaging offers new opportunities for inter-person facial discrimina-tion. However, compact and discriminative feature extraction from high dimensional hyperspectral image cubes is a challenging task. We propose a spatio-spectral feature extraction method based on the 3D Discrete Cosine Transform (3D-DCT). The 3D-DCT optimally compacts information in the low frequency coefficients. Therefore, we rep-resent each hyperspectral facial cube by a small number of low frequency DCT coef-ficients and formulate Partial Least Square (PLS) regression for accurate classification. The proposed algorithm is evaluated on three standard hyperspectral face databases. Ex-perimental results show that the proposed algorithm outperforms five current state of the art hyperspectral face recognition algorithms by a significant margin. 1 Arif Mahmood, Ajmal Mian |
BMVC | 2 |
| 2013 | Periocular biometric recognition using image setsabstractHuman identification based on iris biometrics requires high resolution iris images of a cooperative subject. Such images cannot be obtained in non-intrusive applications such as surveillance. However, the full region around the eye, known as the periocular region, can be acquired non-intrusively and used as a biometric. In this paper we investigate the use of periocular region for person identification. Current techniques have focused on choosing a single best frame, mostly manually, for matching. In contrast, we formulate, for the first time, person identification based on periocular regions as an image set classification problem. We generate periocular region image sets from the Multi Bio-metric Grand Challenge (MBGC) NIR videos. Periocular regions of the right eyes are mirrored and combined with those of the left eyes to form an image set. Each image set contains periocular regions of a single subject. For imageset classification, we use six state-of-the-art techniques and report their comparative recognition and verification performances. Our results show that image sets of periocular regions achieve significantly higher recognition rates than currently reported in the literature for the same database. Arif Mahmood, Ajmal Mian, Chris McDonald |
WACV | 2 |
| 2012 | Hierarchical Sparse Spectral Clustering For Image Set ClassificationabstractWe present a structural matching technique for robust classification based on image sets. In set based classification, a probe set is matched with a number of gallery sets and assigned the label of the most similar set. We represent each image set by a sparse dictionary and compute a similarity matrix by matching all the dictionary atoms of the gallery and probe sets. The similarity matrix comprises the sparse coding coefficients and forms a fully connected directed graph. The nodes of the graph are the dictionary atoms and the edges are the sparse coefficients. The graph is converted to an undirected graph with positive edge weights and spectral clustering is used to cut the graph into two balanced partitions using the normalized cut algorithm. This process is repeated until the graph reduces to critical and non-critical partitions. A critical partition contains atoms with the same gallery label along with one or more probe atoms whereas a non-critical partition either consists of only probe atoms or atoms with multiple gallery labels with no probe atom. Using the critical partitions, we define a novel set based similarity measure and assign the probe set the label of the gallery set with maximum similarity. The proposed algorithm is applied to image set based face recognition using two standard databases. Comparison with existing techniques shows the validity and robustness of our algorithm in the presence of outlier images. Arif Mahmood, Ajmal Mian |
BMVC | 1 |
| 2012 | Correlation-Coefficient-Based Fast Template Matching Through Partial EliminationabstractPartial computation elimination techniques are often used for fast template matching. At a particular search location, computations are prematurely terminated as soon as it is found that this location cannot compete with an already known best match location. Due to the nonmonotonic growth pattern of the correlation-based similarity measures, partial computation elimination techniques have been traditionally considered inapplicable to speed up these measures. In this paper, we show that partial elimination techniques may be applied to a correlation coefficient by using a monotonic formulation, and we propose basic-mode and extended-mode partial correlation elimination algorithms for fast template matching. The basic-mode algorithm is more efficient on small template sizes, whereas the extended mode is faster on medium and larger templates. We also propose a strategy to decide which algorithm to use for a given data set. To achieve a high speedup, elimination algorithms require an initial guess of the peak correlation value. We propose two initialization schemes including a coarse-to-fine scheme for larger templates and a two-stage technique for small- and medium-sized templates. Our proposed algorithms are exact, i.e., having exhaustive equivalent accuracy, and are compared with the existing fast techniques using real image data sets on a wide variety of template sizes. While the actual speedups are data dependent, in most cases, our proposed algorithms have been found to be significantly faster than the other algorithms. Arif Mahmood, Sohaib Khan |
IEEE Trans. Image Process. | 1 |
| 2010 | Exploiting Transitivity of Correlation for Fast Template MatchingabstractElimination Algorithms are often used in template matching to provide a significant speed-up by skipping portions of the computation while guaranteeing the same best-match location as exhaustive search. In this work, we develop elimination algorithms for correlation-based match measures by exploiting the transitivity of correlation. We show that transitive bounds can result in a high computational speed-up if strong autocorrelation is present in the dataset. Generally strong intrareference local autocorrelation is found in natural images, strong inter-reference autocorrelation is found if objects are to be tracked across consecutive video frames and strong intertemplate autocorrelation is found if consecutive video frames are to be matched with a reference image. For each of these cases, the transitive bounds can be adapted to result in an efficient elimination algorithm. The proposed elimination algorithms are exact, that is, they guarantee to yield the same peak location as exhaustive search over the entire solution space. While the speed-up obtained is data dependent, we show empirical results of up to an order of magnitude faster computation as compared to the currently used efficient algorithms on a variety of datasets. Arif Mahmood, Sohaib Khan |
IEEE Trans. Image Process. | 1 |
| 2009 | Early terminating algorithms for Adaboost based detectorsabstractIn this paper we propose an early termination algorithm for speeding up the detection phase of the Adaboost based detectors. In the basic algorithm, at a specific search location, the AdaBoost ensemble response is computed as monotonic decreasing function of weak learners. As more weak learners are evaluated, the response either decreases or remains the same. As soon as the current response becomes lower than the AdaBoost global threshold, remaining computations may be skipped without any loss of accuracy. We further extend the basic algorithm by integrating it with the Non Maxima Suppression (NMS) process. Any candidate location may be discarded, as soon as its current response becomes lower than another candidate location, within the same non-maxima suppression window. In our experiments, our proposed algorithm has been found to be an order of magnitude faster than the traditionally used AdaBoost detector, for the application of edge-corner detection. Speedup comparisons are also done with other three well known edge corner detectors. The early terminated AdaBoost detector has been found to be significantly faster than all three of these detectors. Arif Mahmood, Sohaib Khan |
ICIP | 1 |
| 2008 | Exploiting local auto-correlation function for fast video to reference image alignmentabstractDigital images of natural scenes are usually characterized by strong spatial correlation between adjacent pixels which has been successfully exploited in the coding of still and moving pictures. In this work we show that the strong spatial correlation of natural images can also be used to speedup the video to reference image alignment algorithms. To this end, we divide the search locations in the reference image into groups. The target frames are matched with only one location in each group, while on the remaining locations we evaluate exact theoretic upper bounds on the correlation coefficient. These bounds are used to eliminate majority of the search locations and thus result in significant speedup without effecting the value or location of the global maxima. In our experiments, up to 83.3% search locations are found to be eliminated and the speedup is up to 5.3 times the FFT based implementation and up to 7.9 times the spatial domain techniques. Arif Mahmood, Sohaib Khan |
ICIP | 1 |
| 2007 | Exploiting Inter-frame Correlation for Fast Video to Reference Image Alignment
Arif Mahmood, Sohaib Khan |
ACCV (1) | 1 |
| 2007 | Video Coding With Linear Compensation (VCLC)abstractBlock based motion compensation techniques are commonly used in video encoding to reduce the temporal redundancy of the signal. In these techniques, each block in a video frame is matched with another block in a previous frame. The match criteria normally used is the minimization of the sum of absolute differences (SAD). Traditional encoders take the difference of the current block and its best matching block, and this differential signal is used for further processing. Rather than directly encoding the difference between the two blocks, we propose that the difference between the current block and its first order linear estimate from the best matching block should be used. This choice of using linear compensated differential signal is motivated by observing frequent brightness and contrast changes in real videos. We show two important theoretical results: (1) The variance of the linear compensated differential signal is always less than or equal to the variance of differential signal in traditional encoders. (2) The optimal criteria for finding the best matching block, in our proposed scheme, is the maximization of the magnitude of correlation coefficient. The theoretical results are verified through experimentation on a large dataset taken from several commercial videos. For the same number of bits per pixel, our proposed scheme exhibits an improvement in peak signal to noise ratio (PSNR) of up to 5 dB when compared to the traditional encoding scheme. Arif Mahmood, Zartash Afzal Uzmi, Sohaib Khan |
ICC | 1 |
| 2007 | Early Termination Algorithms for Correlation Coefficient Based Block MatchingabstractBlock based motion compensation techniques make frequent use of early termination algorithms (ETA) to reduce the computational cost of block matching process. ETAs have been well studied in the context of sum of absolute differences (SAD) match measure and are effective in eliminating a large percentage of computations. As compared to SAD, the correlation coefficient (rho) is a more robust measure but has high computational cost because no ETAs for rho have been reported in literature. In this paper, we propose two types of ETAs for correlation coefficient: growth based and the bound based. In growth based ETA, rho is computed as a monotonically decreasing measure. At a specific search location, when the partial value of rho falls below the yet known maxima, remaining calculations are discarded. In bound based ETA, a new upper-bound on rho is derived which is tighter than the currently used Cauchy-Schwartz inequality. The search locations where the proposed bound falls shorter than the yet known maxima are eliminated from the search space. Both types of algorithms are implemented in a cascade and tested on a commercial video dataset. In our experiments, up to 88% computations are found to be eliminated. In terms of execution time, our algorithm is up to 13.7 times faster than the FFTW based implementation and up to 4.6 times faster than the current best known spatial domain technique. Arif Mahmood, Sohaib Khan |
ICIP (2) | 1 |