VLDB 2026 Research / reviewers in the wild / expert
M. Saad Shakeel
dblp:168/9432 · also Muhammad Saad Shakeel
· DBLP profile ↗
17ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0003-3729-2409ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing GNN learning with node augmentation
Maria Marrium, Arif Mahmood, Muhammad Haris Khan, M. Saad Shakeel, Wenxiong Kang |
Neural Networks | 4 |
| 2026 | Interactive feature learning framework for mask-occluded face recognition
M. Saad Shakeel |
Pattern Anal. Appl. | 1 |
| 2025 | Diffusion-Guided Graph Data AugmentationabstractGraph Neural Networks (GNNs) have achieved remarkable success in a wide range of applications. However, when trained on limited or low-diversity datasets, GNNs are prone to overfitting and memorization, which impacts their generalization. To address this, graph data augmentation (GDA) has become a crucial task to enhance the performance and generalization of GNNs.
Traditional GDA methods employ simple transformations that result in limited performance gains. Although recent diffusion-based augmentation methods offer improved results, they are sparse, task-specific, and constrained by class labels. In this work, we propose a more general and effective diffusion-based GDA framework that is task-agnostic and label-free.
For better training stability and reduced computational cost, we employ a graph variational auto-encoder (GVAE) to learn a compact latent graph representation. A diffusion model is used in the learned latent space to generate both consistent and diverse augmentations.
For a fixed augmentation budget, our algorithm selects a subset of samples that would benefit the most from the augmentation.
To further improve performance, we also perform test-time augmentation, leveraged by the label-free nature of our method.
Thanks to the efficient utilization of GVAE and latent diffusion, our algorithm significantly enhances machine learning safety measures, including calibration, robustness to corruptions, and prediction consistency. Moreover, our method has shown improved robustness against four types of adversarial attacks and achieves better generalization performance.
To demonstrate the effectiveness of the proposed method, we compare it with 30 existing methods on 12 benchmark datasets across node classification, link prediction, and graph classification in various learning settings, including semi-supervised, supervised, and long-tailed data distributions.
The code will soon be made publicly available. Maria Marrium, Arif Mahmood, Muhammad Haris Khan, M. Saad Shakeel, Wenxiong Kang |
NeurIPS | 4 |
| 2025 | RSNet: Region-Specific Network for Contactless Palm Vein AuthenticationabstractMore palm features, such as veins and shapes obtained from an enlarged contactless palm vein region of interest (ROI), have been shown to improve recognition performance. However, a few efforts have been made to adequately utilize these features for mining identity information. To address this issue, we propose a Region-Specific Network (RSNet) for contactless palm vein authentication. Our RSNet is a dual-branch structure for global and local feature extraction. Firstly, a Region-based Local feature Enhancement Block (RLEB) is proposed at the local branch to extract region-specific features. In the RLEB, the intermediate feature maps are divided into three asymmetrical patches based on the physiological characteristics of palm vein and palm shape for extracting diversified features, enhancing the local feature representation. Then, a Multi-scale Aggregation Block (MAB) is proposed that efficiently aggregates multi-scale features at a more granular level. Furthermore, to guide the global and local branches in learning complementary feature aspects, a difference loss is introduced to apply a soft subspace orthogonality constraint between the global and local vectors during training. The global branch is designed to assist the learning process of local features, without being adopted for inference. Extensive experiments have demonstrated the effectiveness and superiority of our method, and the RSNet achieves new State-Of-The-Art (SOTA) authentication performance on seven public contactless palm vein databases in the open-set scenario. Dacan Luo, Junduan Huang, Weili Yang, M. Saad Shakeel, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | MRGait: A Multi-range feature learning framework for Cross-View Gait Recognition
M. Saad Shakeel, Kun Liu 0029, Xiaochuan Liao, Wenxiong Kang |
MMAsia | 1 |
| 2024 | CAAM: A calibrated augmented attention module for masked face recognition
M. Saad Shakeel |
J. Vis. Commun. Image Represent. | 1 |
| 2024 | Adaptive Positive Sample Selection and Dynamic Soft Label Assignment for Keypoint DetectionabstractPose estimation plays a crucial role in human-centered vision applications. Some recent efforts achieved pose estimation by keypoints detection. Drawing inspiration from object detection, they treated keypoints as objects and achieved unbiased estimation through implementation of classification and regression heads. However, they still failed to achieve satisfactory performance for detecting heavily occluded keypoints and required elaborate and unavoidable post-processing steps. With a thorough exploration of keypoints’ characteristics, we have developed a novel Adaptive positive Sample selection and dynamic soft Label Assignment (ASLA) scheme tailored for keypoint detection. Specifically, we select positive samples for each keypoint according to the summation distance from the sample coordinates and their predicted coordinates to their corresponding ground truth (GT) in the training phase. For occluded keypoints, the positive samples defined by our method may fall in the semantically relevant regions of pedestrians, rather than the spatially adjacent regions of obstructions, significantly improving their localization performance. Meanwhile, we dynamically assign classification labels to these positive samples based on the distance between their predicted coordinates and their corresponding GT, which ensures that high quality positive samples are assigned with high classification labels. Benefiting from the practical design of our ASLA, the post-processing step is not essential; however, the simple vector-level post-processing would be the icing on the cake. Finally, we extensively evaluate our ASLA performance on two popular human pose estimation benchmarks, COCO and MPII, and comprehensive experiments show that our ASLA significantly outperforms state-of-the-art algorithms. Our code and models will be available athttps://github.com/SCUT-BIP-Lab/ASLA. Wenxiao Tang, M. Saad Shakeel, Wenxiong Kang, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | FVFSNet: Frequency-Spatial Coupling Network for Finger Vein AuthenticationabstractFinger vein biometrics is becoming an important source of human authentication due to its advantages in terms of liveness detection, high security, and user convenience. Although there exist a lot of deep learning-based methods for finger vein authentication, they only extract features from finger vein images in the spatial domain and may lose some important information that is present in other domains, such as the frequency domain. Motivated by this conjecture and the remarkable performance of image feature extraction in the frequency domain, this work explores a method capable of extracting finger vein features in both the spatial and frequency domains. Therefore, the features extracted from different domains can complement each other. In addition, we propose a novel frequency-spatial coupling network (FVFSNet) for finger vein authentication. FVFSNet is mainly composed of three parts: (1) the frequency domain processing module (FDPM), (2) the spatial domain processing module (SDPM), and (3) the frequency-spatial coupling module (FSCM). The FDPM is used to extract the finger vein features present in the frequency domain, which is mainly composed of the frequency-spatial domain transformation and the frequency domain convolution layer. The SDPM is used to extract the finger vein features present in the spatial domain, which is mainly composed of convolution layers with an efficient design. The FSCM is used to couple the features extracted from the FDPM and SDPM, which is mainly composed of the channel and spatial attention mechanisms. To validate our conjecture and the performances of FVFSNet, extensive experiments are conducted on nine commonly used publicly available finger vein datasets. Experimental results show that the frequency domain constitutional neural network has a surprising effect on finger vein authentication, and the proposed FVFSNet achieves the state-of-the-art performance with the advantages of lightweight and low computational cost. Junduan Huang, An Zheng, M. Saad Shakeel, Weili Yang, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Target Category Agnostic Knowledge Distillation With Frequency-Domain SupervisionabstractExisting knowledge distillation approaches require task-related data to train portable student networks for satisfactory performance. Nevertheless, in real-world applications, a majority of the data is unavailable over concerns about personal privacy, commercial confidentiality, etc. To mitigate the difficulty of acquiring the target dataset, the mainstream knowledge distillation methods generate training samples from the teacher network. However, the generated images still differ from the authentic ones, which limit the student network performance. To solve this issue, we propose a convenient and cost-negligible method to build datasets for training student networks. Specifically, we crawl the data from the web without knowing any category of the target dataset, named irrelevant category crawler data (ICCD). To prevent the performance collapse due to the data distribution gap between our ICCD and the target dataset, we propose a pseudo-classification strategy and frequency-domain supervision (PCFS) for target category agnostic knowledge distillation with ICCD. The pseudo-classification strategy classifies ICCD into different pseudo-categories by the teacher network, and uniformly but randomly samples images from each pseudo-category to construct the pseudo-target dataset. Furthermore, we transform the feature maps from the spatial domain to the frequency domain, and utilize the high- and low-frequency signals of the teacher network to impose strong constraints on the student network. Extensive experiments conducted on various test sets demonstrate the effectiveness of our proposed PCFS, which outperforms existing data-free methods and achieves comparable performance to those using the target training set. Code is available athttps://github.com/SCUT-BIP-Lab/PCFS-DFKD. Wenxiao Tang, M. Saad Shakeel, Zisheng Chen, Wenxiong Kang |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | DDAD: Detachable Crowd Density Estimation Assisted Pedestrian DetectionabstractDetecting pedestrians is a challenging computer vision task, especially in the intelligent transportation system. Mainstream pedestrian detection methods purely utilize information of bounding boxes, which overlooks the role of other valuable attributes (e.g., head, head-shoulders, and keypoints) of pedestrians and leads to sub-optimal solutions. Some works leveraged these valuable attributes with a minor performance improvement at the expense of increased computational complexity during the inference phase. To alleviate this dilemma, we propose a simple yet effective method, namely Detachable crowd Density estimation Assisted pedestrian Detection (DDAD), which leverages the crowd density attributes to assist pedestrian detection in the real-world scenes (e.g., crowded scenes and small-scale pedestrian scenes). The advantage of the crowd density estimation is that it allows the network to focus more on the human head and the small-scale pedestrians, which improves the features representation of pedestrians heavily occluded or far from cameras. Our DDAD works on a principle of multi-task learning and can be seamlessly applied to both one-stage and two-stage pedestrian detectors by equipping them with an extra detachable branch of crowd density estimation. The equipped crowd density estimation branch is trained with the annotations derived from the existing pedestrian bounding box annotations, occurring no extra annotation cost. Moreover, it can be removed during the inference phase without sacrificing the inference speed. Extensive experiments conducted on two challenging datasets, i.e., CrowdHuman and CityPersons, demonstrate that our proposed DDAD achieves a significant improvement upon the state-of-the-art methods. Code is available at https://github.com/SCUT-BIP-Lab/ DDAD. Wenxiao Tang, Kun Liu 0029, M. Saad Shakeel, Wenxiong Kang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | BAM: A Bidirectional Attention Module for Masked Face RecognitionabstractMasked Face Recognition (MFR) is a recent addition to the directory of existing challenges in facial biometrics. Due to the limited exposure of facial regions due to mask-occlusion, it is essential to exploit the available non-occluded regions as much as possible for identity feature learning. Aiming to address this issue, we propose a dual-branch bidirectional attention module (BAM), which consists of a spatial attention block (SAB) and a channel attention block (CAB) in each branch. In the first stage, the SAB performs bidirectional interactions between the original feature map and its augmented version to highlight informative spatial locations for feature learning. The learned bidirectional spatial attention maps are then passed through a channel attention block (CAB) to assign high weights to only informative feature channels. Finally, the channel-wise calibrated feature responses are fused to generate a final attention-aware feature representation for MFR. Extensive experiments indicate that our proposed BAM is superior to various state-of-the-art methods in terms of recognizing mask-occluded face images under complex facial variations. M. Saad Shakeel |
VCIP | 1 |
| 2022 | Validity of facial features' geometric measurements for real-time assessment of mental fatigue in construction equipment operators
Imran Mehmood, Heng Li 0001, Waleed Umer, Aamir Arsalan, M. Saad Shakeel, Shahnawaz Anwer |
Adv. Eng. Informatics | 5 |
| 2022 | Deep low-rank feature learning and encoding for cross-age face recognition
M. Saad Shakeel, Kin-Man Lam 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Multi-scale attention guided network for end-to-end face alignment and recognition
M. Saad Shakeel, Wenxiong Kang, Arif Mahmood |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Learning upper patch attention using dual-branch training strategy for masked face recognition
M. Saad Shakeel, Wenxiong Kang |
Pattern Recognit. | 3 |
| 2019 | Learning sparse discriminant low-rank features for low-resolution face recognition
M. Saad Shakeel, Kin-Man Lam 0001, Shun-Cheung Lai |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Deep-feature encoding-based discriminative model for age-invariant face recognition
M. Saad Shakeel, Kin-Man Lam 0001 |
Pattern Recognit. | 1 |