VLDB 2026 Research / reviewers in the wild / expert
Abdelfatah Hassan Ahmed
dblp:336/7541
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-8455-6123ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PRISM-X: Progressive semi-supervised threat detection in X-ray scans with self-guided multimodal refinement
Abdelfatah Hassan Ahmed, Mohammad Irshaid, Mohamad Alansari, Divya Velayudhan, Mohammed Tarnini, Mohammed El-Amine Azz, Naser A. Abou-Elheggag, Taimur Hassan, Ernesto Damiani, Naoufel Werghi |
Inf. Process. Manag. | 1 |
| 2026 | X-SSL: Self-supervised X-ray threat detection with zero-shot and multi-modal learningabstractAutomated X-ray threat detection is challenged by cluttered baggage scans, severe object occlusions, and the scarcity of annotated datasets. Traditional supervised approaches are impractical, as they require large amounts of labeled data, which is difficult to obtain given the rarity of threat objects. Unsupervised methods, on the other hand, often fail to differentiate between threat and nonthreat items due to the complex grayscale nature and high object overlap inherent in X-ray imagery. To overcome these limitations, we propose X-SSL a novel self-supervised learning framework that eliminates manual annotations designed to perform threat localization. Our approach integrates spatial region extraction using MaskCut for zero-shot object proposal generation, contrastive multi-modal clustering that leverages both image and text encoders to cluster and label proposals into threat and nonthreat categories, and self-supervised knowledge distillation where a teacher–student model refines multiscale features from global and local image crops for improved representation learning. We evaluated X-SSL on two benchmark datasets: PIDray (39,000+ images) and CLCXray (14,000+ images), demonstrating significant improvements over previous state-of-the-art (SOTA) methods. In the hidden PIDray subset, X-SSL improves detection AP to 18.16 (+4.41 AP) and segmentation AP to 13.76 (+1.01 AP) over previous methods. On CLCXray, it achieves 39.26 detection AP (+5.93 AP) and 38.67 segmentation AP (+10.20 AP), significantly surpassing previous approaches. For classification, X-SSL achieves an accuracy of 43% on PIDray Hidden and 65% on CLCXray, further highlighting its superior performance compared to existing weakly supervised and unsupervised methods. Code will be available here: https://github.com/yonathan-kiflom/X-SSL . Yonathan Michael, Mohamad Alansari, Abdelfatah Hassan Ahmed, Naoufel Werghi, Andreas Henschel |
Inf. Process. Manag. | 3 |
| 2025 | STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security InspectionabstractAdvancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated threats and concealment tactics, and existing approaches are constrained by a closed-set paradigm with predefined labels. To address these challenges, we introduce STCray, the first multimodal X-ray baggage security dataset, comprising 46,642 image-caption paired scans across 21 threat categories, generated using an X-ray scanner for airport security. STCray is meticulously developed with our specialized protocol that ensures domain-aware, coherent captions, that lead to the multi-modal instruction following data in X-ray baggage security. This allows us to train a domain-aware visual AI assistant named STING-BEE that supports a range of vision-language tasks, including scene comprehension, referring threat localization, visual grounding, and visual question answering (VQA), establishing novel baselines for multi-modal learning in X-ray baggage security. Further, STING-BEE shows state-of-the-art generalization in cross-domain settings. Code, data, and models are available at https://divs1159.github.io/STING-BEE/. Divya Velayudhan, Abdelfatah Hassan Ahmed, Mohamad Alansari, Neha Gour, Abderaouf Behouch, Taimur Hassan, Syed Talal Wasim, Nabil Maalej, Muzammal Naseer, Juergen Gall, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi |
CVPR | 2 |
| 2024 | Feature Fusion for Human Activity Recognition using Parameter-Optimized Multi-Stage Graph Convolutional Network and Transformer ModelsabstractHuman activity recognition is a crucial area of research that involves understanding human movements using computer and machine vision technology. Deep learning has emerged as a powerful tool for this task, with models such as Convolutional Neural Networks (CNNs) and Transformers being employed to capture various aspects of human motion. One of the key contributions of this work is the demonstration of the effectiveness of feature fusion in improving human activity recognition accuracy, which has important implications for the development of more accurate and robust activity recognition systems. This approach addresses a limitation in the field, where the performance of existing models is often limited by their inability to capture both spatial and temporal features effectively. This work presents an approach for human activity recognition using sensory data extracted from four distinct datasets: HuGaDB, PKU-MMD, LARa, and TUG. Two models, the Parameter-Optimized Multi-Stage Graph Convolutional Network (PO-MS-GCN) and a Transformer, were trained and evaluated on each dataset to calculate accuracy and F1-score. Subsequently, the features from the last layer of each model were combined and fed into a classifier. The findings prove that PO MS-GCN outperforms state-of-the-art models in human activity recognition. Specifically, HuGaDB achieved an accuracy of 92.7% and f1-score of 95.2%, TUG achieved an accuracy of 93.2% and f1-score of 98.3%, while LARa and PKU-MMD achieved lower accuracies of 64.31% and 69%, respectively, with corresponding f1-scores of 40.63% and 48.16%. Moreover, feature fusion exceeded the PO-MS-GCN’s results in PKU-MMD, LARa, and TUG datasets. Mohammad Belal, Taimur Hassan, Abdelfatah Hassan Ahmed, Ahmad Aljarah, Nael Alsheikh, Irfan Hussain |
AVSS | 3 |
| 2024 | CLIFS: Clip-Driven Few-Shot Learning for Baggage Threat ClassificationabstractBaggage screening in airports is a cornerstone in airport security measures. The advent of computer vision technologies in recent years has led to the development of several automated systems for identifying security threats in baggage scans. However, existing methods struggle to adapt to new threat categories when faced with a scarcity of data samples, and the rapid emergence of new threats. Hence, in this paper, we propose a novel CLIP-driven few-shot framework (CLIFS) to explore the potential of multi-modality using text-image fusion through contrastive learning to learn relevant contextual features for recognizing security threats with limited samples. By integrating features from GPT-4 generated captions with image features, CLIFS leverages both visual and textual data to significantly improve threat classification performance with limited samples in a few-shot learning context. Our proposed CLIFS was rigorously tested on the SIXray public available baggage X-ray dataset, where it outperformed state-of-the-art by 31.3% in accuracy and 28.40% in F1-score for the challenging 5-shots scenario, demonstrating its robustness and effectiveness in classifying threats from limited data samples. Abdelfatah Hassan Ahmed, Divya Velayudhan, Mahmoud Elmezain, Muaz Al Radi, Abderrahmene Boudiaf, Taimur Hassan, Mohamed Deriche 0001, Mohammed Bennamoun, Naoufel Werghi |
ICIP | 1 |
| 2024 | SMO-CLIP: Enhancing Anomalous Smoke Density Assessment Using A Hybrid LLM-VLM ApproachabstractFlare stacks are among the crucial components in the safety and emission control of petrochemical plants. However, due to the imperceptibility of smoke and contaminants, analyzing these released particles during flare stack operation is one of the top challenges. To stress the problem, our work presents a novel solution called SMO-CLIP that can hybridize knowledge from Vision-Language Models (VLMs), specifically the Contrastive Language Image Pretraining (CLIP) model, with extra insights derived from GPT-4 Large Language Model (LLM). Furthermore, two new tasks, Finegrained Smoke Density Recognition (FSDR) and Coarsegrained Smoke Density Recognition (CSDR) are investigated in this paper to accurately detect and evaluate varying smoke intensities. Notable advancements over current approaches are observed through extensive experiments, demonstrating the superior performance of the proposed approach against state-of-the-art models. Muaz Al Radi, Mahmoud Said Elmezain, Abdelfatah Hassan Ahmed, Abderrahmene Boudiaf, Said Boumaraf, Jorge Dias 0001, Hamad Karki, Sajid Javed, Khalid Yousef Al Awadhi, Naoufel Werghi |
ICIP | 4 |
| 2024 | Enhancing security in X-ray baggage scans: A contour-driven learning approach for abnormality classification and instance segmentation
Abdelfatah Hassan Ahmed, Divya Velayudhan, Taimur Hassan, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Programmable broad learning system for baggage threat recognition
Muhammad Shafay, Abdelfatah Hassan Ahmed, Taimur Hassan, Jorge Dias 0001, Naoufel Werghi |
Multim. Tools Appl. | 2 |
| 2024 | ViT-LSTM synergy: a multi-feature approach for speaker identification and mask detection
Ali Bou Nassif, Ismail Shahin, Mohamed Bader, Abdelfatah Hassan Ahmed, Naoufel Werghi |
Neural Comput. Appl. | 4 |
| 2024 | Autonomous Localization of X-Ray Baggage Threats via Weakly Supervised LearningabstractAutonomous X-ray baggage security screening has shown significant strides recently, proving itself a viable solution to the flaws in manual screening, thanks to advancements in deep learning. However, these data-hungry techniques feed on extensively annotated data involving strenuous labor, impeding their advances in baggage screening. Consequently, we present a context-aware transformer for weakly supervised localization to relieve the annotation burden and provide visual interpretability that aids screeners in threat recognition and researchers in identifying the pitfalls of existing systems. The proposed approach can generalize and localize different types of contraband with only cost-effective binary labels without explicit training on item detection. Context extraction block, integrated into the dual-token framework, generates threat-aware context maps, while the token scoring block focuses on minimizing partial activations. Experimental results surpass state of the art (SOTA) methods in terms of classification and localization accuracies. Furthermore, we analyze failures to determine current vulnerabilities and provide new insights for future research. Divya Velayudhan, Abdelfatah Hassan Ahmed, Taimur Hassan, Neha Gour, Muhammad Owais, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Context-Aware Transformers for Weakly Supervised Baggage Threat LocalizationabstractRecent advances in deep learning have facilitated significant progress in the autonomous detection of concealed security threats from baggage X-ray scans, a plausible solution to overcome the pitfalls of manual screening. However, these data-hungry schemes rely on extensive instance-level annotations that involve strenuous skilled labor. Hence, this paper proposes a context-aware transformer for weakly supervised baggage threat localization, exploiting their inherent capacity to learn long-range semantic relations to capture the object-level context of the illegal items. Unlike the conventional single-class token transformers, the proposed dual-token architecture can generalize well to different threat categories by learning the threat-specific semantics from the token-wise attention to generate context maps. The framework has been evaluated on two public datasets, Compass-XP and SIXray, and surpassed other SOTA approaches. Divya Velayudhan, Abdelfatah Hassan Ahmed, Taimur Hassan, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi |
ICIP | 2 |
| 2022 | Balanced Affinity Loss for Highly Imbalanced Baggage Threat Contour-Driven Instance SegmentationabstractAutonomous detection of threat items from baggage X-ray imagery is one of the most vital and challenging tasks. Manual detection of these items is a cumbersome, slow, and error-ridden process which is also limited by the examination capacity of the security inspector. To overcome these limitations, many researchers have proposed deep learning-driven approaches to recognize suspicious objects from the baggage X-ray scans. However, threat items are rarely seen in the real world compared to innocuous baggage content. Therefore, when trained with imbalanced data, the performance of the conventional threat detection models drastically decreases. This paper addresses these issues with a contour-driven instance segmentation model optimized with a novel combined loss function, dubbed balanced affinity loss function. In addition to mitigating the class imbalance, this function best handles the fine-grained classification aspect inferred by contours and the instance segmentation. We validated the proposed system on three public baggage X-ray datasets, where it outperformed state-of-the-art methods by 7.76%, 25.81%, and 8.78% in terms of intersection-over-union score. Abdelfatah Hassan Ahmed, Ahmad Obeid 0001, Divya Velayudhan, Taimur Hassan, Ernesto Damiani, Naoufel Werghi |
ICIP | 1 |