Amin Ullah

dblp:214/5936 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Computer networks · 6 · 4 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Energy-Efficient Path Planning and Task Scheduling for UAV-Assisted Wireless Networks via Double Dueling Deep Q-Networks with LSTM
abstract
Efficient path planning is crucial to address mission delays, excessive energy consumption, and suboptimal coverage in dynamic Unmanned Aerial Vehicle (UAV) operations. Traditional approaches often overlook temporal patterns in environmental changes and communication demands, leading to inefficient routes and reduced mission performance. To overcome these limitations, this research presents an innovative D3QN-LSTM framework that integrates Double Dueling Deep Q-Networks (D3QN) with Long Short-Term Memory (LSTM) modules for intelligent, temporally-aware UAV path planning. The path planning problem is formulated as a Markov Decision Process (MDP), where the D3QN agent, enhanced with LSTM-based temporal modeling, learns optimal flight trajectories by predicting future changes in mission demands and energy states. Three configurations were evaluated: a baseline D3QN architecture, a D3QN with LSTM integration, and an enhanced D3QN with a deep LSTM architecture. Simulation results demonstrate that incorporating temporal learning reduces mission completion time by up to 14.2%, decreases energy consumption by 12.7%, and improves area coverage efficiency compared to conventional methods. The proposed architecture enables UAVs to perform energy-aware, proactive routing while maintaining operational responsiveness, with promising applications in surveillance, disaster response, and environmental monitoring.
Amin Ullah, Qazi Mazhar ul Haq
CCNC1
2026 Biobanking for intelligent medicine: assessment and evaluation with the SHARE principle
abstract
BACKGROUND: Biobanks are essential for intelligent medicine but face fragmentation and heterogeneity. No standardized framework exists for assessing biobank data value using public information; this study addresses this gap. MATERIALS AND METHODS: We systematically evaluated 94 global biobanks (2010-2024) through a literature review and structured data extraction. Based on 12 international standards, we developed the 5-dimensional SHARE principle (Standardization, Hierarchical structuring, Analytical compatibility, Regulatory compliance, Evolutionary adaptability), operationalized into 10 indicators with a 4-tier scoring system. GPT-4o provided AI-supported prescoring, which was validated by 10 experts and through case studies, including RARPKB. RESULTS: The SHARE principle and a classification map of 94 biobanks were generated. AI and expert scoring showed substantial consistency (κ = 0.62; 95% CI, 0.54-0.70). Biobanks were categorized into 4 tiers: Traditional (60-69), Data-Driven (70-79), Knowledge-Guided (80-89), and Generative and Reasoning-oriented Biobank (90-100). Case validation confirmed utility for disease-specific biobanks. DISCUSSION: We highlight the principal findings, critically examine the reliance on public documentation, propose mitigation strategies, and discuss indicator weighting and implications for translational informatics. CONCLUSIONS: The SHARE principle provides a scalable, standardized method for assessing biobank data value, supporting biobank development, resource discovery, and the development of AI-driven biomedical ecosystems for intelligent medicine.
Amin Ullah, Yingbo Zhang, Hui Zong, Xingyun Liu, Bairong Shen
J. Am. Medical Informatics Assoc.2
2025 Convex Potential Mirror Langevin Algorithm for Efficient Sampling of Energy-Based Models
abstract
This paper introduces the Convex Potential Mirror Langevin Algorithm (CPMLA), a novel method to improve sampling efficiency for Energy-Based Models (EBMs). CPMLA uses mirror Langevin dynamics with a convex potential flow as a dynamic mirror map for EBM sampling. This dynamic mirror map enables targeted geometric exploration on the data manifold, accelerating convergence to the target distribution. Theoretical analysis proves that CPMLA achieves exponential convergence with vanishing bias under relaxed log-concave conditions, supporting its efficiency in adapting to complex data distributions. Experiments on benchmarks like CIFAR-10, SVHN, and CelebA demonstrate CPMLA's improved sampling quality and inference efficiency over existing techniques.
Zitao Yang, Amin Ullah, Fuxin Li
NeurIPS2
2024 CVAE-SM: A Conditional Variational Autoencoder with Style Modulation for Efficient Uncertainty Quantification
abstract
Deep learning has brought transformative advancements to object segmentation, especially in marine robotics contexts such as waste management and subaquatic infrastructure oversight. However, a central challenge persists: calibrating the prediction confidence of the model to ensure robust and reliable outcomes, especially within the demanding underwater environment. Existing solutions for estimating uncertainty are often computationally intensive and have largely centered around Bayesian neural networks or ensemble methods. In this paper, we present a Conditional Variational Autoencoder-based framework (CVAE-SM), which is capable of generating diverse latent codes for improved uncertainty quantification in image segmentation. Our method, enhanced by a style modulator, merges content features, and latent codes more effectively, leading to refined prediction of uncertainty levels. We further introduce a dataset of perturbed underwater images to benchmark uncertainty quantification in this domain. The proposed model not only surpasses peers in segmentation metrics but also matches ensemble models in uncertainty predictions, all while being 2.5 times faster.
Amin Ullah, Taiqing Yan, Fuxin Li
ICRA1
2024 Deep-ReID: deep features and autoencoder assisted image patching strategy for person re-identification in smart cities surveillance
Samee Ullah Khan, Tanveer Hussain 0001, Amin Ullah, Sung Wook Baik
Multim. Tools Appl.3
2023 LocateUAV: Unmanned Aerial Vehicle Location Estimation via Contextual Analysis in an IoT Environment
abstract
Object detection supported by unmanned aerial vehicles (UAVs) has generated significant interest in recent years including applications, such as surveillance, search for missing persons, traffic, and disaster management. Location awareness is a challenging task, particularly, the deployment of UAVs in a global positioning system (GPS) restricted environment or GPS sensor failure. To mitigate this problem, we propose LocateUAV, a novel location awareness framework, to detect UAV’s location by processing the data from the visual sensor in real time using a lightweight convolutional neural network (CNN). Assuming that the drone is in an IoT environment, first, the object detection technique is applied to detect the object of interest (OOI), namely, signboard. Subsequently, optical character recognition (OCR) is applied to extract useful contextual information. In the final step, the extracted information is forwarded to the map application programming interface (API) to locate the UAV. We also present a newly created data set for LocateUAV, which comprises challenging scenarios for context analysis. Moreover, we also compress an existing lightweight model up to 45 MB for efficient processing over UAV, which is 19.5% when compared with the size of the original model. Finally, an in-depth comparison of various trained and efficient object detection and OCR techniques is presented to facilitate future research on the development of flex drone that can extract information from the surroundings of a location in a GPS-restricted environment.
Naqqash Dilshad, Amin Ullah, Jeong-Wook Seo
IEEE Internet Things J.2
2023 Efficient Person Reidentification for IoT-Assisted Cyber-Physical Systems
abstract
The main objective of this study is to propose a cyber–physical system (CPS)-based person reidentification (P-ReID) framework for smart surveillance. The Internet of Things (IoT)-based interconnected vision sensors in smart cities are considered essential elements of a CPS, and contribute significantly to urban security. However, the reidentification of targeted persons using emerging edge AI techniques still faces certain challenges. To improve efficiency at the edge and overcome the traditional sensing of video cameras, we employed an AI-based P-ReID framework for CPS that is functional in IoT environments. In addition, we present dual attention dilated network (DADNet), which integrates an energy-efficient convolutional neural network (CNN) with a self-attention module to substantially improve the person matching probability. Furthermore, we applied dual feature fusion to intelligently integrate discriminative and robust features using early and late fusion strategies that allow DADNet to significantly consider the foreground and marginally utilize the background information. Furthermore, we impose diversity orthogonality regularization over several CNN layers, which boosts the performance of DADNet, resulting in an appropriate usage over IoT networks. A comprehensive set of ablation studies, comparison with other state-of-the-art approaches, and a time complexity analysis confirm the strength of our DADNet for reidentification tasks in AI-enabled IoT settings that are well suited for a CPS.
Samee Ullah Khan, Ijaz Ul Haq, Noman Khan, Amin Ullah, Khan Muhammad 0001, Huiling Chen 0001, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Internet Things J.4
2023 AI-Driven Salient Soccer Events Recognition Framework for Next-Generation IoT-Enabled Environments
abstract
The salient event recognition of soccer matches in the next-generation Internet of Things (Nx-IoT) environment aims to analyze the performance of players/teams by the sports analytics and managerial staff. The embedded Nx-IoT devices carried by the soccer players during the match capture and transmit data to an artificial intelligence (AI)-assisted computing platform. The interconnectivity of data acquisition devices with an AI-assisted computing platform in the Nx-IoT environment will not only allow the spectators to track the formation of their favorite players during a soccer match but will also enable the managerial staff to evaluate the players’ performance in the soccer match as well as in practice sessions. This Nx-IoT-enabled salient event detection feature can be provided to spectators and sports’ managerial staff as a financial technology (FinTech) service. In this article, we propose an efficient deep-learning-based framework for multiperson salient soccer event recognition in IoT-enabled FinTech. The proposed framework performs event recognition in three steps: 1) frames preprocessing; 2) frame-level discriminative features extraction; and 3) high-level events recognition in soccer videos. Moreover, we introduce a new soccer video events (SVE) data set containing videos of six salient events of soccer games. To provide a strong baseline, we evaluate our newly created SVE data set using different traditional machine learning and deep learning algorithms. We also perform event recognition on untrimmed soccer videos using our proposed framework and compare the results with state-of-the-art methods. The obtained results validate the suitability of our proposed framework for salient event recognition in Nx-IoT environments.
Khan Muhammad 0001, Hayat Ullah, Mohammad S. Obaidat, Amin Ullah, Arslan Munir, Victor Hugo C. de Albuquerque
IEEE Internet Things J.4
2023 Deep Learning Assists Surveillance Experts: Toward Video Data Prioritization
abstract
Video summarization (VS) suppresses high-dimensional (HD) video data by only extracting the important information. However, prior research has not focused on the need for surveillance VS, that is used for many applications to assist video surveillance experts, including video retrieval and data storage. In addition, mainstream techniques commonly use two-dimensional (2-D) deep models for VS, ignoring event occurrences. Accordingly, we present a two-fold 3-D deep learning-assisted VS framework. First, we employ an inflated 3-D ConvNet model to extract temporal features; these features are optimized using a proposed encoder mechanism. The input video is temporally segmented using a feature comparison technique for selecting a single frame from each video segment. The segmented shots are evaluated using our novel shot segmentation evaluation scheme and are input into a saliency computation mechanism for keyframe selection in a second fold. Qualitative and quantitative analyses over VS benchmarks and surveillance videos demonstrate the superior performance of our framework, with 0.3- and 4.2-unit increases in the F1 scores for YouTube and title-based video summarization datasets, respectively. Along with accurate VS, a key contribution of our study is the novel shot segmentation criterion prior to VS, which can be used as a benchmark in future research to effectively prioritize HD visual data.
Tanveer Hussain 0001, Fath U Min Ullah, Samee Ullah Khan, Amin Ullah, Umair Haroon, Khan Muhammad 0001, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics4
2022 Artificial Intelligence of Things-assisted two-stream neural network for anomaly detection in surveillance Big Video Data
Waseem Ullah, Amin Ullah, Tanveer Hussain 0001, Khan Muhammad 0001, Ali Asghar Heidari, Javier Del Ser, Sung Wook Baik, Victor Hugo C. de Albuquerque
Future Gener. Comput. Syst.2
2022 An intelligent system for complex violence pattern analysis and detection
abstract
Video surveillance has shown encouraging outcomes to monitor human activities and prevent crimes in real time. To this extent, violence detection (VD) has received substantial attention from the research community due to its vast applications, such as ensuring security over public areas and industrial settings through smart machine intelligence. However, because of changing illumination, complex background and low resolution, the analysis of violence patterns remains challenging in the industrial video surveillance domain. In this paper, we propose a computationally intelligent VD approach to precisely detect violent scenes through deep analysis of surveillance video sequential patterns. First, the video stream acquired through the vision sensor is processed by a lightweight convolutional neural network (CNN) for the segmentation of important shots. Next, temporal optical flow features are extracted from the informative shots via a residential optical flow CNN. These are concatenated with appearance-invariant features extracted from a Darknet CNN model. Finally, a multilayer long short-term memory network is plugged to generate the final feature map for learning the violence patterns in a sequence of frames. In addition, we contribute to the existing surveillance VD data set by considering its indoor and outdoor scenarios separately for the proposed method's evaluation, achieving a 2% increase in accuracy over surveillance fight data set. Experiments also show encouraging results over the state of the art on other challenging benchmark data sets.
Fath U Min Ullah, Mohammad S. Obaidat, Khan Muhammad 0001, Amin Ullah, Sung Wook Baik, Fabio Cuzzolin, Joel J. P. C. Rodrigues, Victor Hugo C. de Albuquerque
Int. J. Intell. Syst.4
2022 Enabling automation and edge intelligence over resource constraint IoT devices for smart home
Mansoor Nasir, Khan Muhammad 0001, Amin Ullah, Jamil Ahmad 0003, Sung Wook Baik
Neurocomputing3
2022 A Multi-Stream Sequence Learning Framework for Human Interaction Recognition
abstract
Human interaction recognition (HIR) is challenging due to multiple humans’ involvement and their mutual interaction in a single frame, generated from their movements. Mainstream literature is based on three-dimensional (3-D) convolutional neural networks (CNNs), processing only visual frames, where human joints data play a vital role in accurate interaction recognition. Therefore, this article proposes a multistream network for HIR that intelligently learns from skeletons’ key points and spatiotemporal visual representations. The first stream localises the joints of the human body using a pose estimation model and transmits them to a 1-D CNN and bidirectional long short-term memory to efficiently extract the features of the dynamic movements of each human skeleton. The second stream feeds the series of visual frames to a 3-D convolutional neural network to extract the discriminative spatiotemporal features. Finally, the outputs of both streams are integrated via fully connected layers that precisely classify the ongoing interactions between humans. To validate the performance of the proposed network, we conducted a comprehensive set of experiments on two benchmark datasets, UT-interaction and TV human interaction, and found 1.15% and 10.0% improvement in the accuracy.
Umair Haroon, Amin Ullah, Tanveer Hussain 0001, Waseem Ullah, Khan Muhammad 0001, Mi Young Lee, Sung Wook Baik
IEEE Trans. Hum. Mach. Syst.2
2021 Human action recognition using attention based LSTM network with dilated CNN features
Khan Muhammad 0001, Mustaqeem Khan 0001, Amin Ullah, Ali Shariq Imran, Mustafa Servet Kiran, Giovanna Sannino, Victor Hugo C. de Albuquerque
Future Gener. Comput. Syst.3
2021 Conflux LSTMs Network: A Novel Approach for Multi-View Action Recognition
abstract
Multi-view action recognition (MVAR) is an optimal technique to acquire numerous clues from different views data for effective action recognition, however, it is not well explored yet. There exist several challenges to MVAR domain such as divergence in viewpoints, invisible regions, and different scales of appearance in each view require better solutions for real world applications. In this paper, we present a conflux long short-term memory (LSTMs) network to recognize actions from multi-view cameras. The proposed framework has four major steps; 1) frame level feature extraction, 2) its propagation through conflux LSTMs network for view self-reliant patterns learning, 3) view inter-reliant patterns learning and correlation computation, and 4) action classification. First, we extract deep features from a sequence of frames using a pre-trained VGG19 CNN model for each view. Second, we forward the extracted features to conflux LSTMs network to learn the view self-reliant patterns. In the next step, we compute the inter-view correlations using the pairwise dot product from output of the LSTMs network corresponding to different views to learn the view inter-reliant patterns. In the final step, we use flatten layers followed by SoftMax classifier for action recognition. Experimental results over benchmark datasets compared to state-of-the-art report an increase of 3% and 2% on northwestern-UCLA and MCAD datasets, respectively.
Amin Ullah, Khan Muhammad 0001, Tanveer Hussain 0001, Sung Wook Baik
Neurocomputing1
2021 Multiview Summarization and Activity Recognition Meet Edge Computing in IoT Environments
abstract
Multiview video summarization (MVS) has not received much attention from the research community due to inter-view correlations and views' overlapping, etc. The majority of previous MVS works are offline, relying on only summary, and require additional communication bandwidth and transmission time, with no focus on foggy environments. We propose an edge intelligence-based MVS and activity recognition framework that combines artificial intelligence with Internet of Things (IoT) devices. In our framework, resource-constrained devices with cameras use a lightweight CNN-based object detection model to segment multiview videos into shots, followed by mutual information computation that helps in a summary generation. Our system does not rely solely on a summary, but encodes and transmits it to a master device using a neural computing stick for inter-view correlations computation and efficient activity recognition, an approach which saves computation resources, communication bandwidth, and transmission time. Experiments show an increase of 0.4 unit in F-measure on an MVS Office dateset and 0.2% and 2% improved accuracy for UCF-50 and YouTube 11 datesets, respectively, with lower storage and transmission times. The processing time is reduced from 1.23 to 0.45 s for a single frame and optimally 0.75 seconds faster MVS. A new dateset is constructed by synthetically adding fog to an MVS dateset to show the adaptability of our system for both certain and uncertain IoT surveillance environments.
Tanveer Hussain 0001, Khan Muhammad 0001, Amin Ullah, Javier Del Ser, Amir Hossein Gandomi, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Internet Things J.3
2021 Splicing sites prediction of human genome using machine learning techniques
Waseem Ullah, Khan Muhammad 0001, Ijaz Ul Haq, Amin Ullah, Saeed Ullah Khattak
Multim. Tools Appl.4
2021 CNN features with bi-directional LSTM for real-time anomaly detection in surveillance networks
Waseem Ullah, Amin Ullah, Ijaz Ul Haq, Khan Muhammad 0001, Sung Wook Baik
Multim. Tools Appl.2
2021 Deep Learning for Safe Autonomous Driving: Current Challenges and Future Directions
abstract
Advances in information and signal processing technologies have a significant impact on autonomous driving (AD), improving driving safety while minimizing the efforts of human drivers with the help of advanced artificial intelligence (AI) techniques. Recently, deep learning (DL) approaches have solved several real-world problems of complex nature. However, their strengths in terms of control processes for AD have not been deeply investigated and highlighted yet. This survey highlights the power of DL architectures in terms of reliability and efficient real-time performance and overviews state-of-the-art strategies for safe AD, with their major achievements and limitations. Furthermore, it covers major embodiments of DL along the AD pipeline including measurement, analysis, and execution, with a focus on road, lane, vehicle, pedestrian, drowsiness detection, collision avoidance, and traffic sign detection through sensing and vision-based DL methods. In addition, we discuss on the performance of several reviewed methods by using different evaluation metrics, with critics on their pros and cons. Finally, this survey highlights the current issues of safe DL-based AD with a prospect of recommendations for future research, rounding up a reference material for newcomers and researchers willing to join this vibrant area of Intelligent Transportation Systems.
Khan Muhammad 0001, Amin Ullah, Jaime Lloret Mauri, Javier Del Ser, Victor Hugo C. de Albuquerque
IEEE Trans. Intell. Transp. Syst.2
2020 One-Shot Learning for Surveillance Anomaly Recognition using Siamese 3D CNN
abstract
One-shot image recognition has been explored for many applications in computer vision community. However, its applications in video analytics is not deeply investigated yet. For instance, surveillance anomaly recognition is an open challenging problem and one of its hurdles is the lack of accurate temporally annotated data. This paper addresses the lack of data issue using one-shot learning strategy and proposes an anomaly recognition framework which exploits a 3D CNN siamese network that yields the similarity between two anomaly sequences. This paper also investigates the existing 3D CNNs for this task and then proposes a lightweight 3D CNN model that efficiently handles one-shot anomaly recognition. Once our network is trained, then we can use the powerful discriminative 3D CNN features to predict anomalies not only for the new data but also for entirely new classes. The proposed model is trained using temporally annotated test set of UCF Crime dataset. Finally, the trained model is used to recognize the anomalies and produce temporal automatic labels for the video level weakly annotated training set of the dataset.
Amin Ullah, Khan Muhammad 0001, Kilichbek Haydarov, Ijaz Ul Haq, Mi Young Lee, Sung Wook Baik
IJCNN1
2020 Human Behavior Understanding in Big Multimedia Data Using CNN based Facial Expression Recognition
Sana Zahir, Amin Ullah, Zahid Akhtar, Khan Muhammad 0001
Mob. Networks Appl.3
2020 Cloud-Assisted Multiview Video Summarization Using CNN and Bidirectional LSTM
abstract
The massive amount of video data produced by surveillance networks in industries instigate various challenges in exploring these videos for many applications, such as video summarization (VS), analysis, indexing, and retrieval. The task of multiview video summarization (MVS) is very challenging due to the gigantic size of data, redundancy, overlapping in views, light variations, and interview correlations. To address these challenges, various low-level features and clustering-based soft computing techniques are proposed that cannot fully exploit MVS. In this article, we achieve MVS by integrating deep neural network based soft computing techniques in a two-tier framework. The first online tier performs target-appearance-based shots segmentation and stores them in a lookup table that is transmitted to cloud for further processing. The second tier extracts deep features from each frame of a sequence in the lookup table and pass them to deep bidirectional long short-term memory (DB-LSTM) to acquire probabilities of informativeness and generates a summary. Experimental evaluation on benchmark dataset and industrial surveillance data from YouTube confirms the better performance of our system compared to the state-of-the-art MVS methods.
Tanveer Hussain 0001, Khan Muhammad 0001, Amin Ullah, Zehong Cao, Sung Wook Baik, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics3
2019 Action recognition using optimized deep autoencoder and CNN for surveillance data streams of non-stationary environments
Amin Ullah, Khan Muhammad 0001, Ijaz Ul Haq, Sung Wook Baik
Future Gener. Comput. Syst.1
2019 Efficient Image Recognition and Retrieval on IoT-Assisted Energy-Constrained Platforms From Big Data Repositories
abstract
The advanced computational capabilities of many resource constrained devices, such as smartphones have enabled various research areas including image retrieval from big data repositories for numerous Internet of Things (IoT) applications. The major challenges for image retrieval using smartphones in an IoT environment are the computational complexity and storage. To deal with big data in IoT environment for image retrieval, this paper proposes a light-weighted deep learning-based system for energy-constrained devices. The system first detects and crops face regions from an image using Viola-Jones algorithm with additional face and nonface classifier to eliminate the miss-detection problem. Second, the system uses convolutional layers of a cost effective pretrained CNN model with defined features to represent faces. Next, features of the big data repository are indexed to achieve a faster matching process for real-time retrieval. Finally, Euclidean distance is used to find similarity between query and repository images. For experimental evaluation, we created a local facial images dataset, including both single and group facial images. This dataset can be used by other researchers as a benchmark for comparison with other real-time facial image retrieval systems. The experimental results show that our proposed system outperforms other state-of-the-art feature extraction methods in terms of efficiency and retrieval for IoT-assisted energy-constrained platforms.
Irfan Mehmood, Amin Ullah, Khan Muhammad 0001, Der-Jiunn Deng, Weizhi Meng 0001, Fadi M. Al-Turjman, Victor Hugo C. de Albuquerque
IEEE Internet Things J.2
2018 Integrating salient colors with rotational invariant texture features for image representation in retrieval systems
Amin Ullah, Jamil Ahmad 0003, Naveed Abbas, Seungmin Rho, Sung Wook Baik
Multim. Tools Appl.2