VLDB 2026 Research / reviewers in the wild / expert
Waseem Ullah
dblp:272/8485
· DBLP profile ↗
29ranked-venue papers
10as first author
29since 2021 · last 2026
0000-0001-5191-9023ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Practical Study of Retrieval-Derived Knowledge Injection in Video Captioning
Karina Abubakirova, Waseem Ullah, Mohsen Guizani |
IWCMC | 2 |
| 2026 | Spiking Transformer-Based Framework for Modeling and Forecasting Energy Generation and Consumption
Ayesha Alhammadi, Waseem Ullah, Mohsen Guizani |
IWCMC | 2 |
| 2026 | Efficient Attention Based UNet for Rooftop Photovoltaic Panel Segmentation
Khawla Almarzooqi, Waseem Ullah, Mohsen Guizani |
IWCMC | 2 |
| 2026 | Adaptive and Reliable Quality Enhancement in Metaverse: A BSUM Approach
Maher Guizani, Latif U. Khan, Waseem Ullah, Mohammad A. Islam 0001, Fakhri Karray |
IWCMC | 3 |
| 2026 | FedVLP: Federated Vision-Language Prompting for Disaster Recognition
Waseem Ullah, Latif U. Khan, Maher Guizani |
IWCMC | 1 |
| 2026 | Graph-Based Temporal Attention Network for Anomaly Recognition in Internet of Things Video SurveillanceabstractAccurate anomaly recognition in video surveillance is critical to ensuring public safety, infrastructure protection, and intelligent security operations. However, existing methods often fail to generalize across diverse and dynamic environments due to their limited capacity to model complex spatio-temporal patterns. In this paper, we propose a novel Graph-Based Temporal Attention Memory Network (G-TAMNet) that significantly advances video anomaly detection by capturing intricate temporal dependencies and spatial relationships. By integrating temporal self-attention into a graph convolutional network (GCN) framework, our model enhances the learning of salient and contextually relevant features that traditional approaches tend to overlook. Furthermore, to enable deployment in resource-constrained settings, such as IoTbased surveillance systems, we incorporate model quantization strategies that reduce computational overhead without compromising detection accuracy. Extensive evaluations of three widely used benchmarks, UCFCrime, LAD-2000, and RWF-2000 the effectiveness of the proposed G-TAMNet model, achieving accuracy, precision, recall, and F1-scores of 54.3% / 61.1% / 59.01% / 61.1%,79.3% / 69.1% / 70.0% / 69.1%,and 94.0% / 94.0% / 94.2% / 94.1%, respectively. These results reflect consistent performance gains of 3.3%, 9.9%, and 0.74% in accuracy over existing state-ofthe-art methods on the corresponding datasets. Such improvements underscore the robustness, scalability, and practical viability of GTAMNet for real-time anomaly detection in resource-constrained IoT surveillance systems. In addition, its reliable generalization across diverse environments highlights the potential of the model to advance intelligent security analytics and transform real-world applications in the domain of information forensics. Waseem Ullah, Latif U. Khan, Mohsen Guizani, Chang-Dong Wang 0001, Di Wu 0001 |
IEEE Internet Things J. | 1 |
| 2026 | Class-incremental learning network for real-time anomaly recognition in surveillance environments
Adnan Hussain, Waseem Ullah, Noman Khan, Zulfiqar Ahmad Khan 0002, Hikmat Yar, Sung Wook Baik |
Pattern Recognit. | 2 |
| 2025 | Joint Communication and Sensing in Metaverse over UAVs: A Deep Reinforcement Learning ApproachabstractThe development of wireless systems based on metaverse to effectively enable novel applications (e.g., surveillance and remote healthcare) gained significant interest from the research community. A metaverse combines many technologies with a virtual representation of the real world system for enabling applications. Although a metaverse enables many applications, there are challenges associated with its deployment. Therefore, in this paper, we consider sensing in metaverse-empowered unmanned aerial vehicle (UAV) networks. A massive number of sensors are considered to effectively sense the physical world states and share them with the metaverse. A problem is formulated to minimize the cost of sensing in terms of energy and latency. Our problem is based on optimizing the resource allocation and association for sensing units. We present a solution based on double deep Q-network (DDQN) for joint association and resource allocation in sensing for metaverse over UAVs network because the defined problem is non-convex and combinatorial. In the end, we validate our idea with comprehensive simulation results. Sabina Jangirova, Latif U. Khan, Branislava Jankovic, Waseem Ullah, Mohsen Guizani |
ISCC | 4 |
| 2025 | UAV-Assisted Real-Time Disaster Detection Using Optimized Transformer ModelabstractDangerous surroundings and difficult-to-reach landscapes introduce significant complications for adequate disaster management and recuperation. These problems can be solved by engaging unmanned aerial vehicles (UAVs) provided with embedded platforms and optical sensors. In this work, we focus on enabling onboard aerial image processing to ensure proper and real-time disaster detection. Such a setting usually causes challenges due to the limited hardware resources of UAVs. However, privacy, connectivity, and latency issues can be avoided. We suggest a UAV-assisted edge framework for disaster detection, leveraging our proposed model optimized for onboard real-time aerial image classification. The optimization of the model is achieved using post-training quantization techniques. To address the limited number of disaster cases in existing benchmark datasets and therefore ensure real-world adoption of our model, we construct a novel dataset, DisasterEye, featuring disaster scenes captured by UAVs and individuals on-site. Experimental results reveal the efficacy of our model, reaching high accuracy with lowered inference latency and memory use on both traditional machines and resource-limited devices. This shows that the scalability and adaptability of our method make it a powerful solution for real-time disaster management on resource-constrained UAV platforms.The code and DisasterEye dataset are available at: https://github.com/Branislava98/TensorRT. Branislava Jankovic, Sabina Jangirova, Waseem Ullah, Latif U. Khan, Mohsen Guizani |
ISCC | 3 |
| 2025 | Resource Optimized Split Federated Learning: A Reinforcement Learning and Optimization ApproachabstractFederated learning (FL) offers many benefits, such as better privacy preservation and less communication overhead for scenarios with frequent data generation. In FL, local models are trained on end-devices and then migrated to the network edge or cloud for global aggregation. This aggregated model is shared back with end-devices to further improve their local models. This iterative process continues until convergence is achieved. Although FL has many merits, it has many challenges. The prominent one is computing resource constraints. End-devices typically have fewer computing resources and are unable to learn well the local models. Therefore, split FL (SFL) was introduced to address this problem. However, enabling SFL is also challenging due to wireless resource constraints and uncertainties. We formulate a joint end-devices computing resources optimization, task-offloading, and resource allocation problem for SFL at the network edge. Our problem formulation has a mixed-integer non-linear programming problem nature and hard to solve due to the presence of both binary and continuous variables. We propose a double deep Q-network (DDDQN) and optimization-based solution. Finally, we validate the proposed method using extensive simulation results. Maher Guizani, Latif U. Khan, Waseem Ullah, Mohammad A. Islam 0001 |
IWCMC | 3 |
| 2025 | Efficient Aerial Fire Detection on Resource-Constrained Devices Using Cross-Architecture Knowledge DistillationabstractFire poses significant ecological danger and can lead to human losses if not mitigated early. Fire detection systems on remote sensing devices have become critical to disaster prevention and management. In this paper, we propose a lightweight and effective fire detection model based on MobileViT-XS, compressed through the distillation of knowledge from a stronger teacher model. Our model combines the local feature extraction capabilities of convolutional networks with the global context capabilities of vision transforms. This allows our proposed model to achieve high detection accuracy and real-time inference, making it suitable for deployment on UAVs and other resource-constrained devices. Through rigorous experiments on the ADSF dataset, the proposed method achieves an accuracy, F1-score, and recall of 95.50%, along with a precision of 95.52%, thereby outperforming contemporary methods. Extensive evaluations on ADSF confirm that our method not only surpasses existing state-of-the-art techniques in accuracy but also provides significantly faster processing speeds. Sabina Jangirova, Branislava Jankovic, Waseem Ullah, Latif U. Khan, Mohsen Guizani |
IWCMC | 3 |
| 2025 | Vision-Language Models for Edge Networks: A Comprehensive SurveyabstractVision Large Language Models (VLMs) combine visual understanding with natural language processing, enabling tasks like image captioning, visual question answering, and video analysis. While VLMs show impressive capabilities across domains such as autonomous vehicles, smart surveillance, and healthcare, their deployment on resource-constrained edge devices remains challenging due to processing power, memory, and energy limitations. This survey explores recent advancements in optimizing VLMs for edge environments, focusing on model compression techniques, including pruning, quantization, knowledge distillation, and specialized hardware solutions that enhance efficiency. We provide a detailed discussion of efficient training and fine-tuning methods, edge deployment challenges, and privacy considerations. Additionally, we discuss the diverse applications of lightweight VLMs across healthcare, environmental monitoring, and autonomous systems, illustrating their growing impact. By highlighting key design strategies, current challenges, and offering recommendations for future directions, this survey aims to inspire further research into the practical deployment of VLMs, ultimately making advanced AI accessible in resource-limited settings. Ahmed Sharshar, Latif U. Khan, Waseem Ullah, Mohsen Guizani |
IEEE Internet Things J. | 3 |
| 2025 | Early progression detection from MCI to AD using multi-view MRI for enhanced assisted living
Nasir Rahim, Waseem Ullah, Jatin Bedi, Younhyun Jung |
Image Vis. Comput. | 3 |
| 2025 | Optimized cross-module attention network and medium-scale dataset for effective fire detection
Zulfiqar Ahmad Khan 0002, Fath U Min Ullah, Hikmat Yar, Waseem Ullah, Noman Khan, Min Je Kim, Sung Wook Baik |
Pattern Recognit. | 4 |
| 2025 | Big Data Analysis for Industrial Activity Recognition Using Attention-Inspired Sequential Temporal Convolution NetworkabstractDeep-learning-based human activity recognition (HAR) methods have significantly transformed a wide range of domains over recent years. However, the adoption of Big Data techniques in industrial applications remains challenging due to issues such as generalized weight optimization, diverse viewpoints, and the complex spatiotemporal features of videos. To address these challenges, this work presents an industrial HAR framework consisting of two main phases. First, a squeeze bottleneck attention block (SBAB) is introduced to enhance the learning capabilities of the backbone model for contextual learning, which allows for the selection and refinement of an optimal feature vector. In the second phase, we propose an effective sequential temporal convolutional network (STCN), which is designed in parallel fashion to mitigate the issues of exploding and vanishing gradients associated with sequence learning. The high-dimensional spatiotemporal feature vectors from the STCN undergo further refinement through our proposed SBAB in a sequential manner, to optimize the features for HAR and enhance the overall performance. The efficacy of the proposed framework is validated through extensive experiments on six datasets, including data from industrial and general activities. Altaf Hussain 0002, Tanveer Hussain 0001, Waseem Ullah, Samee Ullah Khan, Min Je Kim, Khan Muhammad 0001, Javier Del Ser, Sung Wook Baik |
IEEE Trans. Big Data | 3 |
| 2025 | Block Successive Upper-Bound Minimization for Resource Scheduling in Wireless MetaverseabstractIn recent years, there has been a rising trend towards emerging applications (e.g., brain-computer interaction and haptics-based autonomous cars) with diverse requirements. To effectively enable these applications via autonomous operation and intelligent analytics, one can use a metaverseFor more details on how a metaverse can enable emerging applications and architecture, please refer to khan2024ametaverse. In a metaverse, we have two spaces: (a) a meta space based on a virtual model that performs analysis and resource management and (b) a physical space comprised of real world entities. A metaverse effectively enables emerging applications by performing three main tasks: (a) distributed learning of metaverse models; (b) instantly serving the end-users; and (c) sensing of the physical environment and sharing it with the meta space for synchronized operation. To perform these tasks, efficient wireless resource management is needed. Therefore, a novel resource scheduling framework for the wireless metaverse to enable various applications is proposed. Our aim is to minimize the cost of learning and sensing in metaverse. Subsequently, we formulate a problem. Meanwhile, the reliability as well as latency constraints of the service-requesting devices/users will be fulfilled. We assign multiple resource blocks to learning and sensing devices/units, whereas we use a concept of puncturing for service-requesting devices/users upon arrival. We use a scheme that is based on block successive upper-bound minimization and convex optimization for solving our formulated problem. At the end, we use empirical cumulative distribution function vs. cost and cost vs. metaverse entities for numerical evaluations. Latif U. Khan, Waseem Ullah, Sami Muhaidat, Mohsen Guizani, Bechir Hamdaoui |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | TDS-Net: Transformer enhanced dual-stream network for video Anomaly Detection
Adnan Hussain, Waseem Ullah, Noman Khan, Zulfiqar Ahmad Khan 0002, Min Je Kim, Sung Wook Baik |
Expert Syst. Appl. | 2 |
| 2024 | An efficient deep learning architecture for effective fire detection in smart surveillance
Hikmat Yar, Zulfiqar Ahmad Khan 0002, Imad Rida, Waseem Ullah, Min Je Kim, Sung Wook Baik |
Image Vis. Comput. | 4 |
| 2024 | A Trapezoid Attention Mechanism for Power Generation and Consumption ForecastingabstractEffective operation of smart grids relies on accurate forecasting models for renewable power generation (RPG) and power consumption. The intermittent and unpredictable nature of RPG, coupled with diverse consumption patterns, underscores the importance of robust forecasting approaches. Existing models often employ stacked layers, integrating direct features into fully connected layers, yielding suboptimal results with limited generalization capabilities. Addressing these limitations, we propose a novel two-stream architecture for RPG and power consumption forecasting. The first stream leverages dilated causal convolutional layers to capture intricate patterns, while the second stream focuses on temporal information extraction. Importantly, we fine-tune the hyperparameters of both streams using Bayesian algorithms to optimize the learning process. The outputs from these two streams are then intelligently combined and channeled into our innovative trapezoid attention module (TAM) for feature refinement, resulting in superior pattern representation. The TAM incorporates three distinct dimensions (spatial, temporal, and spatiotemporal) and enriches feature maps by integrating a skip connection from the pre-TAM features. The output post-TAM features are then employed for final forecasting. Our approach showcases remarkable performance in short-term forecasting across a spectrum of datasets, including RPG, regional, residential, and industrial power consumption. By addressing the shortcomings of existing forecasting models, our research contributes to the advancement of smart grid technologies, ensuring more reliable and efficient energy management. Zulfiqar Ahmad Khan 0002, Tanveer Hussain 0001, Waseem Ullah, Sung Wook Baik |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | TransCNN: Hybrid CNN and transformer mechanism for surveillance anomaly detection
Waseem Ullah, Tanveer Hussain 0001, Fath U Min Ullah, Mi Young Lee, Sung Wook Baik |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | Sequential attention mechanism for weakly supervised video anomaly detection
Waseem Ullah, Fath U Min Ullah, Zulfiqar Ahmad Khan 0002, Sung Wook Baik |
Expert Syst. Appl. | 1 |
| 2023 | A modified YOLOv5 architecture for efficient fire detection in smart cities
Hikmat Yar, Zulfiqar Ahmad Khan 0002, Fath U Min Ullah, Waseem Ullah, Sung Wook Baik |
Expert Syst. Appl. | 4 |
| 2023 | AD-Graph: Weakly Supervised Anomaly Detection Graph Neural NetworkabstractThe main challenge faced by video‐based real‐world anomaly detection systems is the accurate learning of unusual events that are irregular, complicated, diverse, and heterogeneous in nature. Several techniques utilizing deep learning have been created to detect anomalies, yet their effectiveness on real‐world data is often limited due to the insufficient incorporation of motion patterns. To address these problems and enhance the traditional functionality of anomaly detection systems for surveillance video data, we propose a weakly supervised graph neural‐network‐assisted video anomaly detection framework called AD‐Graph. To identify temporal information from a series of frames, we extract 3D visual and motion features and represent these in a language‐based knowledge graph format. Next, a robust clustering strategy is applied to group together meaningful neighbourhoods of the graph with similar vertices. Furthermore, spectral filters are applied to these graphs, and spectral graph theory is used to generate graph signals and detect anomalous events. Extensive experimental results over two challenging datasets, UCF‐Crime and ShanghaiTech, show improvements of 0.35% and 0.78% against a state‐of‐the‐art model. Waseem Ullah, Tanveer Hussain 0001, Fath U Min Ullah, Khan Muhammad 0001, Mahmoud Hassaballah, Joel J. P. C. Rodrigues, Sung Wook Baik, Victor Hugo C. de Albuquerque |
Int. J. Intell. Syst. | 1 |
| 2023 | Vision transformer attention with multi-reservoir echo state network for anomaly recognition
Waseem Ullah, Tanveer Hussain 0001, Sung Wook Baik |
Inf. Process. Manag. | 1 |
| 2022 | Artificial Intelligence of Things-assisted two-stream neural network for anomaly detection in surveillance Big Video Data
Waseem Ullah, Amin Ullah, Tanveer Hussain 0001, Khan Muhammad 0001, Ali Asghar Heidari, Javier Del Ser, Sung Wook Baik, Victor Hugo C. de Albuquerque |
Future Gener. Comput. Syst. | 1 |
| 2022 | Intelligent dual stream CNN and echo state network for anomaly detection
Waseem Ullah, Tanveer Hussain 0001, Zulfiqar Ahmad Khan 0002, Umair Haroon, Sung Wook Baik |
Knowl. Based Syst. | 1 |
| 2022 | A Multi-Stream Sequence Learning Framework for Human Interaction RecognitionabstractHuman interaction recognition (HIR) is challenging due to multiple humans’ involvement and their mutual interaction in a single frame, generated from their movements. Mainstream literature is based on three-dimensional (3-D) convolutional neural networks (CNNs), processing only visual frames, where human joints data play a vital role in accurate interaction recognition. Therefore, this article proposes a multistream network for HIR that intelligently learns from skeletons’ key points and spatiotemporal visual representations. The first stream localises the joints of the human body using a pose estimation model and transmits them to a 1-D CNN and bidirectional long short-term memory to efficiently extract the features of the dynamic movements of each human skeleton. The second stream feeds the series of visual frames to a 3-D convolutional neural network to extract the discriminative spatiotemporal features. Finally, the outputs of both streams are integrated via fully connected layers that precisely classify the ongoing interactions between humans. To validate the performance of the proposed network, we conducted a comprehensive set of experiments on two benchmark datasets, UT-interaction and TV human interaction, and found 1.15% and 10.0% improvement in the accuracy. Umair Haroon, Amin Ullah, Tanveer Hussain 0001, Waseem Ullah, Khan Muhammad 0001, Mi Young Lee, Sung Wook Baik |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2021 | Splicing sites prediction of human genome using machine learning techniques
Waseem Ullah, Khan Muhammad 0001, Ijaz Ul Haq, Amin Ullah, Saeed Ullah Khattak |
Multim. Tools Appl. | 1 |
| 2021 | CNN features with bi-directional LSTM for real-time anomaly detection in surveillance networks
Waseem Ullah, Amin Ullah, Ijaz Ul Haq, Khan Muhammad 0001, Sung Wook Baik |
Multim. Tools Appl. | 1 |