Muhammad Zaigham Zaheer

dblp:260/4206 · also Muhamamad Zaigham Zaheer · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0001-8272-1351ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Multi-Agent Diffusion Approach for MRI Anomaly Segmentation via Modality-Specific LoRA Specialization
Wafa Al Ghallabi, Muhammad Zaigham Zaheer, Ritesh Thawkar, Omkar Thawakar, Salman Khan 0001, Fahad Shahbaz Khan
WACV2
2026 GenMix: Effective data augmentation with generative diffusion model image editing
abstract
Data augmentation is widely used to enhance generalization in visual classification tasks. However, traditional methods struggle when source and target domains differ, as in domain adaptation, due to their inability to address domain gaps. This paper introduces GenMix, a generalizable prompt-guided generative data augmentation approach that enhances both in-domain and cross-domain image classification. Our technique leverages image editing to generate augmented images based on custom conditional prompts, designed specifically for each problem type. By blending portions of the input image with its edited generative counterpart and incorporating fractal patterns, our approach mitigates unrealistic images and label ambiguity, improving the performance and adversarial robustness of the resulting models. Efficacy of our method is established with extensive experiments on eight public datasets for general and fine-grained classification, in both in-domain and cross-domain settings. Additionally, we demonstrate performance improvements for self-supervised learning, learning with data scarcity, and adversarial robustness. As compared to the existing state-of-the-art methods, our technique achieves stronger performance across the board.
Khawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood, Karthik Nandakumar, Naveed Akhtar
Expert Syst. Appl.2
2024 Collaborative Learning of Anomalies with Privacy (CLAP) for Unsupervised Video Anomaly Detection: A New Baseline
abstract
Unsupervised (US) video anomaly detection (VAD) in surveillance applications is gaining more popularity recently due to its practical real-world applications. As surveillance videos are privacy sensitive and the availability of large-scale video data may enable better US- VAD systems, collaborative learning can be highly rewarding in this setting. However, due to the extremely challenging nature of the US- VAD task, where learning is carried out without any annotations, privacy-preserving collaborative learning of us- VAD systems has not been studied yet. In this paper, we propose a new baseline for anomaly detection capable of localizing anomalous events in complex surveil-lance videos in a fully unsupervised fashion without any labels on a privacy-preserving participant-based distributed training configuration. Additionally, we propose three new evaluation protocols to benchmark anomaly detection approaches on various scenarios of collaborations and data availability. Based on these protocols, we modify existing VAD datasets to extensively evaluate our approach as well as existing US SOTA methods on two large-scale datasets including UCF-Crime and XD- Violence. All proposed evaluation protocols, dataset splits, and codes are available here: https://github.com/AnasEmadllICLAP.
Anas Al-lahham, Muhammad Zaigham Zaheer, Nurbek Tastan, Karthik Nandakumar
CVPR2
2024 Diffusemix: Label-Preserving Data Augmentation with Diffusion Models
abstract
Recently, a number of image-mixing-based augmentation techniques have been introduced to improve the gen-eralization of deep neural networks. In these techniques, two or more randomly selected natural images are mixed together to generate an augmented image. Such methods may not only omit important portions of the input images but also introduce label ambiguities by mixing images across labels resulting in misleading supervisory signals. To address these limitations, we propose Diffusemix, a novel data augmentation technique that leverages a diffusion model to reshape training images, supervised by our bespoke conditional prompts. First, concatenation of a partial natural image and its generated counterpart is ob-tained which helps in avoiding the generation of unrealistic images or label ambiguities. Then, to enhance resilience against adversarial attacks and improves safety measures, a randomly selected structural pattern from a set of frac-tal images is blended into the concatenated image to form the final augmented image for training. Our empirical results on seven different datasets reveal that Diffusemix achieves superior performance compared to existing state-of-the-art methods on tasks including general classification, fine- grained classification, fine-tuning, data scarcity, and adversarial robustness. Augmented datasets and codes are available here: https://diffusemix.github.io/
Khawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood, Karthik Nandakumar
CVPR2
2024 Feature Map Purification for Enhancing Adversarial Robustness of Deep Timeseries Classifiers
abstract
Deep-learning based timeseries classifiers are known to be susceptible to adversarial attacks, where the adversary adds imperceptible perturbations to the input sample to cause mis-classification. While principled adversarial defense mechanisms such as adversarial training and certified robustness have been proposed for image classifiers, they have seldom been studied in the timeseries domain. Existing defenses for timeseries classifiers are primarily centered around adversarial sample detection, but have mixed performance and fail to generalize well across attacks. This work proposes an alternative approach based on purifying the intermediate representations within a deep convolutional timeseries (DCT) classifier. We design a learnable sub-network with residual connections that filters the feature maps in multiple wavelet basis spaces to suppress the adversarial perturbations. Given any pretrained non-robust DCT classifier, the proposed feature map purification (FeMPure) module can be trained in isolation without affecting the given classifier and can be seamlessly plugged back in to enhance the adversarial robustness of the original classifier. Experiments based on 2 well-known architectures for DCT classifiers, 6 adversarial attacks, and 80 public-domain datasets demonstrate that the proposed FeMPure approach can provide good adversarial robustness, irrespective of whether the adversary is unaware or has full knowledge of the defense mechanism. With minor modifications, the FeMPure approach can also be employed for adversarial sample detection or for enhancing certified robustness.
Mubarak G. Abdu-Aguye, Muhammad Zaigham Zaheer, Karthik Nandakumar
ICDM2
2024 A Synopsis of FAME 2024 Challenge: Associating Faces with Voices in Multilingual Environments
Muhammad Saad Saeed, Shah Nawaz, Marta Moscati, Rohan Kumar Das, Muhammad Salman Tahir, Muhammad Zaigham Zaheer, Muhammad Irzam Liaqat, Muhammad Haris Khan, Karthik Nandakumar, Muhammad Haroon Yousaf, Markus Schedl
ACM Multimedia6
2024 A Coarse-to-Fine Pseudo-Labeling (C2FPL) Framework for Unsupervised Video Anomaly Detection
abstract
Detection of anomalous events in videos is an important problem in applications such as surveillance. Video anomaly detection (VAD) is well-studied in the one-class classification (OCC) and weakly supervised (WS) settings. However, fully unsupervised (US) video anomaly detection methods, which learn a complete system without any annotation or human supervision, have not been explored in depth. This is because the lack of any ground truth annotations significantly increases the magnitude of the VAD challenge. To address this challenge, we propose a simple-but-effective two-stage pseudo-label generation framework that produces segment-level (normal/anomaly) pseudo-labels, which can be further used to train a segment-level anomaly detector in a supervised manner. The proposed coarse-to-fine pseudo-label (C2FPL) generator employs carefully-designed hierarchical divisive clustering and statistical hypothesis testing to identify anomalous video segments from a set of completely unlabeled videos. The trained anomaly detector can be directly applied on segments of an unseen test video to obtain segment-level, and subsequently, frame-level anomaly predictions. Extensive studies on two large-scale public-domain datasets, UCF-Crime and XD-Violence, demonstrate that the proposed unsupervised approach achieves superior performance compared to all existing OCC and US methods, while yielding comparable performance to the state-of-the-art WS methods.
Anas Al-lahham, Nurbek Tastan, Muhammad Zaigham Zaheer, Karthik Nandakumar
WACV3
2024 Exploiting autoencoder's weakness to generate pseudo anomalies
Marcella Astrid, Muhammad Zaigham Zaheer, Djamila Aouada, Seung-Ik Lee
Neural Comput. Appl.2
2024 Clustering Aided Weakly Supervised Training to Detect Anomalous Events in Surveillance Videos
abstract
Formulating learning systems for the detection of real-world anomalous events using only video-level labels is a challenging task mainly due to the presence of noisy labels as well as the rare occurrence of anomalous events in the training data. We propose a weakly supervised anomaly detection system that has multiple contributions including a random batch selection mechanism to reduce interbatch correlation and a normalcy suppression block (NSB) which learns to minimize anomaly scores over normal regions of a video by utilizing the overall information available in a training batch. In addition, a clustering loss block (CLB) is proposed to mitigate the label noise and to improve the representation learning for the anomalous and normal regions. This block encourages the backbone network to produce two distinct feature clusters representing normal and anomalous events. An extensive analysis of the proposed approach is provided using three popular anomaly detection datasets including UCF-Crime, ShanghaiTech, and UCSD Ped2. The experiments demonstrate the superior anomaly detection capability of our approach.
Muhammad Zaigham Zaheer, Arif Mahmood, Marcella Astrid, Seung-Ik Lee
IEEE Trans. Neural Networks Learn. Syst.1
2023 Single-branch Network for Multimodal Training
abstract
With the rapid growth of social media platforms, users are sharing billions of multimedia posts containing audio, images, and text. Researchers have focused on building autonomous systems capable of processing such multimedia data to solve challenging multimodal tasks including cross-modal retrieval, matching, and verification. Existing works use separate networks to extract embeddings of each modality to bridge the gap between them. The modular structure of their branched networks is fundamental in creating numerous multimodal applications and has become a defacto standard to handle multiple modalities. In contrast, we propose a novel single-branch network capable of learning discriminative representation of unimodal as well as multimodal tasks without changing the network. An important feature of our single-branch network is that it can be trained either using single or multiple modalities without sacrificing performance. We evaluated our proposed single-branch network on the challenging multimodal problem (face-voice association) for cross-modal verification and matching tasks with various loss formulations. Experimental results demonstrate the superiority of our proposed single-branch network over the existing methods in a wide range of experiments. Code: https://github.com/msaadsaeed/SBNet
Muhammad Saad Saeed, Shah Nawaz, Muhammad Haris Khan, Muhammad Zaigham Zaheer, Karthik Nandakumar, Muhammad Haroon Yousaf, Arif Mahmood
ICASSP4
2023 DCTM: Dilated Convolutional Transformer Model for Multimodal Engagement Estimation in Conversation
abstract
Conversational engagement estimation is posed as a regression problem, entailing the identification of the favorable attention and involvement of the participants in the conversation. This task arises as a crucial pursuit to gain insights into human's interaction dynamics and behavior patterns within a conversation. In this research, we introduce a dilated convolutional Transformer for modeling and estimating human engagement in the MULTIMEDIATE 2023 competition. Our proposed system surpasses the baseline models, exhibiting a noteworthy 7% improvement on test set and 4% on validation set. Moreover, we employ different modality fusion mechanism and show that for this type of data, a simple concatenated method with self-attention fusion gains the best performance.
Vu Ngoc Tu, Van Thong Huynh, Soo-Hyung Kim, Shah Nawaz, Karthik Nandakumar, Muhammad Zaigham Zaheer
ACM Multimedia7
2023 PseudoBound: Limiting the anomaly reconstruction capability of one-class classifiers using pseudo anomalies
abstract
Due to the rarity of anomalous events, video anomaly detection is typically approached as one-class classification (OCC) problem. Typically in OCC, an autoencoder (AE) is trained to reconstruct the normal only training data with the expectation that, in test time, it can poorly reconstruct the anomalous data. However, previous studies have shown that, even trained with only normal data, AEs can often reconstruct anomalous data as well, resulting in a decreased performance. To mitigate this problem, we propose to limit the anomaly reconstruction capability of AEs by incorporating pseudo anomalies during the training of an AE. Extensive experiments using five types of pseudo anomalies show the robustness of our training mechanism towards any kind of pseudo anomaly. Moreover, we demonstrate the effectiveness of our proposed pseudo anomaly based training approach against several existing state-of-the-art (SOTA) methods on three benchmark video anomaly datasets, outperforming all the other reconstruction-based approaches in two datasets and showing the second best performance in the other dataset.
Marcella Astrid, Muhammad Zaigham Zaheer, Seung-Ik Lee
Neurocomputing2
2023 Fingerprint pattern classification using deep transfer learning and data augmentation
Divine Senanu Ametefe, Suzi Seroja Sarnin, Darmawaty Mohd Ali, Muhammad Zaigham Zaheer
Vis. Comput.4
2022 Face Pyramid Vision Transformer
Khawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood
BMVC2
2022 Generative Cooperative Learning for Unsupervised Video Anomaly Detection
abstract
Video anomaly detection is well investigated in weakly-supervised and one-class classification (OCC) settings. However, unsupervised video anomaly detection methods are quite sparse, likely because anomalies are less frequent in occurrence and usually not well-defined, which when coupled with the absence of ground truth supervision, could adversely affect the performance of the learning algorithms. This problem is challenging yet rewarding as it can completely eradicate the costs of obtaining laborious annotations and enable such systems to be deployed without human intervention. To this end, we propose a novel unsupervised Generative Cooperative Learning (GCL) approach for video anomaly detection that exploits the low frequency of anomalies towards building a cross-supervision between a generator and a discriminator. In essence, both networks get trained in a cooperative fashion, thereby allowing unsupervised learning. We conduct extensive experiments on two large-scale video anomaly detection datasets, UCF crime and ShanghaiTech. Consistent improvement over the existing state-of-the-art unsupervised and OCC methods corroborate the effectiveness of our approach.
Muhammad Zaigham Zaheer, Arif Mahmood, Muhammad Haris Khan, Mattia Segù, Fisher Yu 0001, Seung-Ik Lee
CVPR1
2022 Stabilizing Adversarially Learned One-Class Novelty Detection Using Pseudo Anomalies
abstract
Recently, anomaly scores have been formulated using reconstruction loss of the adversarially learned generators and/or classification loss of discriminators. Unavailability of anomaly examples in the training data makes optimization of such networks challenging. Attributed to the adversarial training, performance of such models fluctuates drastically with each training step, making it difficult to halt the training at an optimal point. In the current study, we propose a robust anomaly detection framework that overcomes such instability by transforming the fundamental role of the discriminator from identifying real vs. fake data to distinguishing good vs. bad quality reconstructions. For this purpose, we propose a method that utilizes the current state as well as an old state of the same generator to create good and bad quality reconstruction examples. The discriminator is trained on these examples to detect the subtle distortions that are often present in the reconstructions of anomalous data. In addition, we propose an efficient generic criterion to stop the training of our model, ensuring elevated performance. Extensive experiments performed on six datasets across multiple domains including image and video based anomaly detection, medical diagnosis, and network security, have demonstrated excellent performance of our approach.
Muhammad Zaigham Zaheer, Jin Ha Lee 0002, Arif Mahmood, Marcella Astrid, Seung-Ik Lee
IEEE Trans. Image Process.1
2021 Learning Not to Reconstruct Anomalies
Marcella Astrid, Muhammad Zaigham Zaheer, Jae-Yeong Lee, Seung-Ik Lee
BMVC2
2021 4G-VOS: Video Object Segmentation using guided context embedding
Mustansar Fiaz, Muhammad Zaigham Zaheer, Arif Mahmood, Seung-Ik Lee, Soon Ki Jung
Knowl. Based Syst.2
2020 Old Is Gold: Redefining the Adversarially Learned One-Class Classifier Training Paradigm
abstract
A popular method for anomaly detection is to use the generator of an adversarial network to formulate anomaly score over reconstruction loss of input. Due to the rare occurrence of anomalies, optimizing such networks can be a cumbersome task. Another possible approach is to use both generator and discriminator for anomaly detection. However, attributed to the involvement of adversarial training, this model is often unstable in a way that the performance fluctuates drastically with each training step. In this study, we propose a framework that effectively generates stable results across a wide range of training steps and allows us to use both the generator and the discriminator of an adversarial model for efficient and robust anomaly detection. Our approach transforms the fundamental role of a discriminator from identifying real and fake data to distinguishing between good and bad quality reconstructions. To this end, we prepare training examples for the good quality reconstruction by employing the current generator, whereas poor quality examples are obtained by utilizing an old state of the same generator. This way, the discriminator learns to detect subtle distortions that often appear in reconstructions of the anomaly inputs. Extensive experiments performed on Caltech-256 and MNIST image datasets for novelty detection show superior results. Furthermore, on UCSD Ped2 video dataset for anomaly detection, our model achieves a frame-level AUC of 98.1%, surpassing recent state-of-the-art methods.
Muhammad Zaigham Zaheer, Jin Ha Lee 0002, Marcella Astrid, Seung-Ik Lee
CVPR1
2020 CLAWS: Clustering Assisted Weakly Supervised Learning with Normalcy Suppression for Anomalous Event Detection
Muhammad Zaigham Zaheer, Arif Mahmood, Marcella Astrid, Seung-Ik Lee
ECCV (22)1
2020 A Self-Reasoning Framework for Anomaly Detection Using Video-Level Labels
abstract
Anomalous event detection in surveillance videos is a challenging and practical research problem among image and video processing community. Compared to the frame-level annotations of anomalous events, obtaining video-level annotations is quite fast and cheap though such high-level labels may contain significant noise. More specifically, an anomalous labeled video may actually contain anomaly only in a short duration while the rest of the video frames may be normal. In the current work, we propose a weakly supervised anomaly detection framework based on deep neural networks which is trained in a self-reasoning fashion using only video-level labels. To carry out the self-reasoning based training, we generate pseudo labels by using binary clustering of spatio-temporal video features which helps in mitigating the noise present in the labels of anomalous videos. Our proposed formulation encourages both the main network and the clustering to complement each other in achieving the goal of more accurate anomaly detection. The proposed framework has been evaluated on publicly available real-world anomaly detection datasets including UCF-crime, ShanghaiTech and UCSD Ped2. The experiments demonstrate superiority of our proposed framework over the current state-of-the-art methods.
Muhammad Zaigham Zaheer, Arif Mahmood, Hochul Shin, Seung-Ik Lee
IEEE Signal Process. Lett.1