VLDB 2026 Research / reviewers in the wild / expert
Muhammad Ghulam
dblp:49/10239 · also Ghulam Muhammad
· DBLP profile ↗
135ranked-venue papers
24as first author
63since 2021 · last 2026
0000-0002-9781-3969ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 44 · 4 first-author · 29 since 2021Artificial intelligence and machine learning · 35 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 12 first-author · 9 since 2021Systems, architecture and hardware · 18 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Personalized Federated Transformer Architecture With Digital Twin for Enhanced Environmental Perception in Intelligent IoV SystemsabstractWith 6G-enabled Intelligent Internet of Vehicles (IIoV) generating massive amounts of sensory data, traditional deep learning models struggle to capture long-range relationships across different sensor types while preserving privacy. This paper proposes DT-Trans, a privacy-preserving federated learning framework that combines Digital Twin technology with Vision Transformers. Our framework first trains a global perception model on synthetic digital twin data, then fine-tunes it efficiently for real-world vehicles. By grouping vehicles with similar driving patterns and allowing them to collaboratively train personalized model components, DT-Trans achieves significant accuracy improvements while maintaining data privacy. The Twin-Enhanced Vision Transformer (TE-ViT) is introduced as the global perception backbone; it is pre-trained on massive synthetic DT data and then fine-tuned via parameter-efficient LoRA adapters to bridge the domain gap between virtual and physical worlds. The Cluster-Enhanced Decoupled PFL (CD-PFL-Trans) algorithm splits each TE-ViT into (i) a shared Transformer encoder (base layer) and (ii) client-specific Transformer decoder heads (personalized layer). Hierarchical clustering on decoder parameters groups clients with similar traffic patterns, enabling group-wise aggregation without exchanging raw sensory data. DT-Trans outperforms CNN-based FedAvg/FedPer by 9.3%-16.2% mAP on V&PKITTI perception tasks and up to 42.8% accuracy improvement on CINIC-10 classification under severe heterogeneity, while reducing on-device FLOPs by 34 % via Transformer sparsity techniques. Our work advances Transformer architectures for scalable, privacy-preserving perception in IIoV. Xuewei Chao, Jiachen Jiang, Wenyan Ma, Yang Li 0111, Jing Nie 0002, Sezai Ercisli, Muhammad Ghulam |
IEEE Internet Things J. | 7 |
| 2026 | A Blockchain-Enabled Image Encryption Protocol Based on Quantum Walk for Securing Industrial Internet of Things EnvironmentsabstractThe swift, secure transmission of data, especially sensitive data, such as that acquired from high-resolution image sensors, is a major focus of the Industrial Internet of Things (IIoT). Current strategies do not strike a balance between security and efficiency, causing messages to be lost and/or processing to be delayed to an unacceptable level. We present a new quantum-inspired quantum walk (Q-IQW) encryption protocol. The Q-IQW protocol is used to implement the first blockchain for the safe transfer of data between IIoT devices. For the first time, quantum hash functions based on Q-IQW are used to link blocks in a chain instead of traditional hash functions. The main contributions of the protocol are the efficient data transfer between IIoT nodes and the ability to fully control their data. This work focuses on simulating the protocol to understand the theoretical and empirical performance of the protocol. The results show that in terms of entropy of 7.99 and unified average changing intensity (UACI) of 33.56%, the protocol is very robust as evidenced by 99.58% number of pixel change rate (NPCR), and low correlation, i.e., the protocol exhibits very high security. Further, the protocol is reliable in that it can perfectly recover images under high distortion, for example, a 50% data occlusion in the encrypted images. These outcomes underscore the dependability, utility, and efficiency of the protocol in protecting IIoT information to make it a viable remedy for maintaining the integrity of the information during transfer and storage. Sunil Prajapat, Pankaj Kumar 0006, Muhammad Ghulam, Ashok Kumar Das |
IEEE Internet Things J. | 4 |
| 2026 | QHSA-ViT: A Quantum Discrete-Fourier-Transform-Based Hierarchical Self-Attention Fusion Vision Transformer for Traffic Sign Recognition in Intelligent Vehicular NetworksabstractWith the rapid advancement of the intelligent Internet of Vehicles (IoV), accurate traffic sign classification is essential to ensure driving safety and improve environmental perception. However, conventional image classification models often rely on local features and spatial domain processing, lacking global context modeling and facing computational limitations. To address these challenges, this paper proposes a quantum discrete Fourier transform-based hierarchical self-attention Vision Transformer (QHSA-ViT). Using the parallelism and high-dimensional feature extraction capabilities of quantum computing, the proposed model enhances the quality and efficiency of representation. Specifically, a quantum frequency domain feature representation (QFDFR) module based on a quantum discrete Fourier transform (QDFT) is introduced to capture rich spectral features, while a quantum self-attention fusion (QSAF) module built on a linear combination of unitaries (LCU) and generalized quantum singular value transformation (GQSVT) integrates multilevel attention. The experimental results on five benchmark datasets, including GTSRB, show that QHSA-ViT outperforms baseline models with an average improvement of 9.01% in accuracy and 8.48% in the F1 score. These results validate the effectiveness of the proposed model and highlight its practical applicability and scalability for understanding traffic scenes in intelligent IoV. Zhiguo Qu, Mengqing Zhou, Le Sun 0003, Yimin Yu, Muhammad Ghulam |
IEEE Internet Things J. | 5 |
| 2026 | PCNA-IDS: An integrated lightweight intrusion detection system in internet of vehicles with federated contrastive learning and differential privacy
Zhiguo Qu, Zihong Cai, Le Sun 0003, Muhammad Ghulam |
Knowl. Based Syst. | 4 |
| 2026 | Quantum-driven attention and relation-aware distillation for multi-modal emotion recognition in conversations
Muhammad Ghulam |
Pattern Recognit. | 2 |
| 2026 | Transfer Learning-Enabled System for Drone Medicine Delivery Based on Spatio-Temporal Remote Sensing Data in Edge Cloud NetworksabstractThese days, satellite remote sensing data is employed for different drone applications. The main goal is to provide imaginary information about electromagnetic locations and patterns of geolocations insight into Earth. The Internet of Drone Things (IoDT) exploits remote sensing data to deliver medicine from source to destination. However, many existing medicine delivery systems based on drones need longer execution times and more efficiency in delivering medicine to the right destinations. This paper presents transfer learning, which empowers a spatiotemporal remote sensing data training system for medicine delivery in edge cloud networks based on IoDT applications. The objective is to deliver the medicine to the original destination with the highest score and process all drone tasks based on their given deadlines. We present the offloading spatiotemporal training and scheduling (OSPTS) algorithm methodology that completes the data collection process and medicine delivery in different locations. Therefore, we solve the problem as a combinatorial problem and find the optimal solution based on searching and convolutional neural networks (CNN). Transfer learning and convolutional neural networks are sub-schemes of the OSPTS that train the remote sensing data on edge nodes and point clouds for optimal medicine delivery. Simulation results show that the OSPTS obtained the highest score for medicine delivery in the correct position with less processing time than existing systems. Abdullah Lakhan, Tor-Morten Grønli, Ahmet Soylu, Muhammad Ghulam, Qurat-Ul-Ain Mastoi, Huaming Wu |
IEEE Trans. Cloud Comput. | 4 |
| 2026 | EEG-based Multimodal Emotion Recognition: Recent Progress, Challenges, and Future DirectionsabstractEmotion recognition is a crucial part of cognitive computing. Traditional emotion recognition systems include audio-visual modality. However, a recent trend in recognizing emotions is to use physiological signals such as the Electroencephalogram (EEG). EEG signals, together with audio-visual and other physiological signals, improve the performance of emotion recognition systems. This article presents a systematic literature review on EEG-based multimodal (multimedia) emotion recognition systems for the last 5 years. Three major research questions are addressed: (1) What kind of learning models are used in EEG-based multimedia emotion recognition? (2) What are the publicly available related datasets? (3) What are the challenges and future directions of this topic? The answers to the research questions are provided in different subsections. Muhammad Ghulam, Sumayah A. Almuntasheri, Fadia Alenezi, Nwraan Alhadi, Victor C. M. Leung |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Deep Learning-Based Privacy Preserving Multimodal Biometrics Recognition for Cross-Silo DatasetsabstractABSTRACT Different biometric modalities, such as fingerprints and left and right eye irises, contain physiological characteristics that offer high accuracy in identification processes. These modalities complement each other; for example, fingerprints provide intricate ridge patterns, while irises exhibit stable, precise features that perform well in challenging environments. A new proposed framework based on federated learning with optimised features, pre‐trained deep learning models, linear discriminant analysis and dense neural networks ensures privacy protection for multi‐modal biometric recognition across diverse biometric datasets. The system obtains better accuracy levels alongside increased robustness through the combination of fingerprint and iris scan technology that functions across independent and identically distributed (IID) and non‐independent and non‐identically distributed (non‐IID) conditions. Privacy protection functions as a key asset of federated learning because it allows distributed training operations through non‐raw data sharing, supporting high classification results. The system's performance is enhanced by implementing feature fusion alongside dimensionality reduction methods, which enhance both the efficiency and resistance to noise and variabilities. The system establishes an essential reference point for distributed and heterogeneous real‐world biometric recognition because it implements accurate computation with enhanced efficiency together with privacy protection. The IID data experiments demonstrated 98.86% training accuracy while achieving precision and recall at precise levels of 98.86% and 96.59%. All metrics achieved 100% on the validation data set while keeping loss at zero. The system's performance slightly decreased under non‐IID training data conditions, which resulted in 95.01% training accuracy and 0.18 training loss. The reported precision levels matched recall values since both measurements reached 97.99% and 95.01%. The system maintained perfect validation results through all metrics, which demonstrated a strong ability to generalise beyond data distribution impediments. The integration of multimodal biometric systems with federated learning enables the optimisation of large‐scale solutions because it establishes efficient but accurate and secure applications across domains that include surveillance and security together with healthcare. Isha Kansal, Vikas Khuallar, Gifty Gupta, Deepali Gupta, Sapna Juneja, Ali Nauman, Muhammad Ghulam |
Expert Syst. J. Knowl. Eng. | 7 |
| 2025 | Contextual embedded text summarizer system: A hybrid approachabstractAbstract Selecting crucial sentences from a document is a pivotal task in automatic text summarization systems. Abstractive summarization involves rephrasing key content through advanced natural language techniques, generating a concise, new text conveying critical information. Conversely, extractive summarization reproduces important material from the original text. In the proposed method, a hybrid ensemble approach combines BERTsum for extractive summarization and Longformer2Roberta for abstractive summarization for generating a contextual semantic rich summary for a huge collection of text. These proposed system‐generated summaries were evaluated against reference summaries using the ROUGE package at three rouge levels (Rouge‐1, Rouge‐2, and Rouge‐L). The Proposed contextual embedded hybrid text summarization model has shown significant performance improvement in multiple levels of Rouge score and word mover distance (WMD) of generated summary with a reference summary. The proposed hybrid model demonstrates superior performance over existing state‐of‐the‐art summarizing models on three distinct datasets CNN dataset, WikiSum, and Gigaword dataset. The proposed hybrid model as a text summarizer involves leveraging its capabilities to process longer sequences of text with domain‐specific contextual summaries. This transformers‐based text summarization model has great potential in developing expert systems in various research domains such as health decision support systems, the education sector, customer support chatbots, financial analysis investment recommendations, and financial assistance. Pooja Kherwa, Jyoti Arora, Deepali Gupta, Sapna Juneja, Muhammad Ghulam, Ali Nauman |
Expert Syst. J. Knowl. Eng. | 6 |
| 2025 | From AI to the Era of Explainable AI in Healthcare 5.0: Current State and Future OutlookabstractABSTRACT Artificial intelligence (AI) and explainable artificial intelligence (XAI) are advancing rapidly, with the potential to deliver significant benefits to modern society. The healthcare sector, in particular, has experienced transformative changes; overall, these technologies are helping to address numerous challenges, such as cancer cell detection, tumour zone identification in animal bodies, predictions of major and minor diseases, diagnosis, and more. This article provides an in‐depth and detailed overview of AI and XAI, focusing on recent trends and their implications for advancing Healthcare 5.0 applications. Initially, the study examines the key concepts and exceptional features of AI, XAI, and Healthcare 5.0. Additional emphasis is placed on state‐of‐the‐art practices currently being implemented in healthcare, particularly those involving AI and XAI. Subsequently, it establishes a coherent link between AI and XAI in Healthcare 5.0, grounded in contemporary advancements. Based on the findings, algorithms are recommended to address initial obstacles to integrating AI into the Healthcare 5.0 framework. Proposals for further enhancing Healthcare 5.0 performance through the integration of XAI and its unique features are discussed in detail. The work also provides in‐depth implementation strategies and highlights model‐specific trends within AI and XAI frameworks in Healthcare 5.0. Particular attention is given to AI model predictions in healthcare settings, emphasising their contributions to improved patient feedback and the delivery of more sophisticated care. Most importantly, this research highlights the potential for AI and XAI to support sustainable advancements in Healthcare 5.0 applications. Finally, significant issues are analysed, and an open discussion is presented on future guidelines for the blending of AI with XAI, and Healthcare 5.0 applications. Anichur Rahman, Dipanjali Kundu, Tanoy Debnath, Muaz Rahman, Utpol Kanti Das, Abu Saleh Musa Miah, Muhammad Ghulam |
Expert Syst. J. Knowl. Eng. | 7 |
| 2025 | Smart Internet of Everything Model for Knowledge-Graph-Based Reliable RecommendationabstractUser attention, doubt and anxiety about reliability in intelligent decisions is continuously increasing with increase in the data overload on the Internet of Things (IoT) frameworks. Recommender system (RS) faces noisy inputs and provides vague recommendations with target disparity, explanation ambiguity, and performance biasness as a consequence. The inactive or less interactive users suffer more from these issues due to the lack of sufficient information about their browsing history on the system’s end. In this work, therefore, we introduce smart Internet of Everything model for knowledge graph-based reliable recommendation (KGR) to overcome irrelevant feeds-in from the IoT networks and ensure pertinence-based data quality at the knowledge base to address the highlighted research challenges in the current IoT-based RSs. Particularly, we verify relevance of the incoming contents with the concerned application scenarios, translate the received data to the embedding space, and apply data quality inspection check on the underlying data. We independently encapsulate user-to-item interactions, and provide independent streams of low-level representations of users and items to the prediction module. It uses deep nonnegative matrix factorization technique to process user-item representations and acquire the required preferences. In experiments on four real world datasets, KGR outperforms the-state-of-the-art methods by successfully meeting the aforementioned challenges. Nasrullah Khan, Muhammad Ghulam, Xiaoyuan Jing |
IEEE Internet Things J. | 2 |
| 2025 | Generative AI-Enabled Quantum Encryption Algorithm for Securing IoT-Based Healthcare Application Using BlockchainabstractThe integration of artificial intelligence (AI) with the Internet of Things (IoT) has transformed numerous domains through the AI of Things (AIoT). Nonetheless, AIoT encounters issues related to energy usage and carbon emissions as mobile technology continues to progress. Generative AI (GAI) possesses significant potential to mitigate carbon emissions associated with AIoT, owing to its higher reasoning and generative powers. Conventional security protocols frequently encounter issues with computational efficiency, latency, and overall security comprehensiveness. Blockchain technology, characterized by its decentralized and immutable properties, is a viable approach for improving electronic healthcare data transmission and node authentication in IoT networks. This research examines secure data transmission and node encryption in IoT systems, with a particular emphasis on data management. Conventional approaches encounter constraints in computational efficiency, latency, and comprehensive security. This study presents a novel protocol that combines GAI and blockchain technology with quantum encryption to enhance authentication and ensure secure data transmission. The algorithm comprises multiple consecutive processes, including the encoding and transmission of node requests, followed by the authentication process utilizing hash functions and digital signatures. The authentication approach utilizes a challenge-response technique, guaranteeing that only nodes with authentic credentials can advance. Thereafter, a dynamic key exchange protocol and quantum encryption method provide secure data delivery. The results indicate the procedure’s effectiveness in ensuring secure and regulated access to patient data, underscoring its significance in medical facilities. The system’s functionalities are augmented by a thorough evaluation employing machine learning. The findings indicate that the system exhibits an accuracy of 99.4%, precision of 99.10%, recall of 98.66%, F1-score of 98.50%, and security of 99.2%. An extensive analysis and comparison with the state-of-the-art methods demonstrate the significant advancements of the suggested method in tackling cryptographic security challenges. The algorithm offers a thorough approach to protecting IoT applications, especially in managing healthcare data. Sunil Prajapat, Pankaj Kumar 0006, Ashok Kumar Das, Muhammad Ghulam |
IEEE Internet Things J. | 4 |
| 2025 | SCS-QBCT: A Supply Chain System-Driven Efficient Quantum Blockchain Cross-Chain Transaction SchemeabstractThe development of supply chain systems demands optimization of various technologies in terms of efficiency and resource conservation. As an emerging technology, cross-chain technology in blockchain aims to achieve interoperability and resource sharing between different blockchain networks, enhancing data liquidity and system efficiency. However, relay chains in cross-chain interactions require storing a large number of transaction records, leading to excessive communication and storage loads, which can cause network performance degradation, storage resource exhaustion, and low transaction processing efficiency. To address these issues, this paper proposes a supply chain system-driven efficient quantum blockchain cross-chain transaction scheme (SCS-QBCT). Firstly, SCS-QBCT uses the quantum Fourier transform (QFT) to convert transaction records on relay chains from the time domain to the frequency domain, reducing data redundancy and significantly lowering storage space consumption. Secondly, a multifunctional smart contract, designed to include conventional functions, enables value transfer, transaction withdrawal, transaction query, node identity management, and transaction type identification. Furthermore, inverse quantum Fourier transform (IQFT) is used to restore quantum state transaction records in blocks to classical records, supporting transaction query requests. Finally, the experimental results and theoretical analysis show that SCS-QBCT performs excellently in reducing storage consumption, improving system efficiency, practicality, and security, and meeting the optimization goals of supply chain systems. Zhiguo Qu, Le Sun 0003, Yimin Yu, Muhammad Ghulam |
IEEE Internet Things J. | 5 |
| 2025 | QCACNN: A Quantum Convolutional Neural Network Algorithm for Traffic Sign Recognition in Carbon-Intelligent Electric VehiclesabstractTraffic sign recognition is essential for autonomous driving, enhancing driving efficiency and ensuring road safety. Traffic signs utilize colors and patterns to relay critical information, with color accuracy being especially significant. Traditional models, reliant on extensive data and computing resources, struggle to meet the efficiency demands of electric vehicles with carbon-intelligent computing. Leveraging the physical properties of quantum superposition states, quantum computing offers a solution with its unique parallel computing capabilities, potentially enhancing recognition efficiency and facilitating real-time processing. Quantum convolutional neural networks (QCNNs) show promise in processing large-scale image data with improved efficiency and accuracy. However, most of QCNN researches focus on grayscale image classification, with limited studies on multichannel data. This article introduces the quantum channel attention convolutional neural network (QCACNN), which incorporates a quantum channel attention layer (QCAL) to enhance multichannel data classification. The experimental results demonstrate that QCACNNs surpasses traditional QCNNs in traffic sign recognition, performing comparably to conventional CNNs and SENet models with fewer computing resources. Detailed performance analysis and ablation studies validate the effectiveness of each component within this architecture. Quantum noise resistance tests confirm the robustness of QCACNN, making it a viable solution for electric vehicles with enhanced scalability. Zhiguo Qu, Yichen Xia, Le Sun 0003, Muhammad Ghulam |
IEEE Internet Things J. | 5 |
| 2025 | DAQFL: Dynamic Aggregation Quantum Federated Learning Algorithm for Intelligent Diagnosis in Internet of Medical ThingsabstractFederated learning (FL) is a privacy-preserving alternative to centralized machine learning, where model training is performed on local devices and only global model updates are shared, effectively addressing challenges, such as data silos and privacy protection. Recently, quantum FL (QFL), an emerging FL branch, has garnered significant attention in many industry applications. However, existing QFL algorithms predominantly employ average weighting for global model training, which shows poor performance on heterogeneous healthcare data. To address this challenge, this study proposes a dynamic aggregation QFL algorithm (DAQFL) for intelligent diagnosis. Specifically, it utilizes quantum neural networks (QNNs) as local training models and designs corresponding variational quantum circuits (VQC). To mitigate performance degradation caused by the heterogeneity of medical industrial data, a dynamic aggregation method based on accuracy is proposed to enhance global model performance effectively. Extensive experiments with three distribution settings, including independent and identically distributed (IID), non-independent and identically distributed (Non-IID), and long-tail datasets, show that DAQFL outperforms baseline algorithms in accuracy and training speed. It also performs well in privacy protection and robustness of anti-noise, improving its suitability for real-world medical applications. Zhiguo Qu, Xuemeng Zhao, Le Sun 0003, Muhammad Ghulam |
IEEE Internet Things J. | 4 |
| 2025 | Gating Memory Network Multilayer Perceptron for Traffic Forecasting in Internet of Vehicles SystemsabstractWith the increasing number of distributed edge intelligence (DEI) sensors in cities, Internet of Vehicles (IoV) systems can obtain more fine-grained information. The information plays a crucial role in tasks, such as road monitoring and traffic forecasting. However, graph-based methods used in IoV systems mainly have two limitations: i) They struggle to achieve both efficiency and high prediction performance, and ii) Most lack user-friendly interfaces to intuitively display data collected by DEI sensors and prediction results. To alleviate the first limitation, we propose a novel deep learning model, the gating memory network multilayer perceptron (GMMLP). In the model, the traditional time-consuming graph-based operations are replaced with multilayer perceptrons (MLPs). A gating mechanism is employed to help reduce redundant information from raw input. An integrated memory network is utilized to memorize common patterns implicitly. To alleviate the second limitation and make GMMLP more accessible, we design a DEI-based human-computer interaction and visualization system called the traffic easy access system (TEAS). Traffic departments and ordinary users can access roads in the state recorded by DEI sensors on multiple terminals. Professionals can easily train and test traffic forecasting models by using this system. We validate the performance of our proposed model on multiple datasets. The experimental results demonstrate that our model achieves excellent performance in both inference efficiency and prediction accuracy. Le Sun 0003, Wenzhang Dai, Muhammad Ghulam |
IEEE Internet Things J. | 3 |
| 2025 | Privacy Preservation in AI-Driven IoT for Vehicles via Hierarchical Sharding BlockchainabstractThe AI-driven Internet of Things (AIoT) has been widely applied in the field of Internet of Vehicles (IoV) for vehicular cooperation. Federated learning (FL), due to its ability to protect users’ data privacy, reduce communication overhead, and facilitate real-time decision making, is widely applied in the augmented intelligence of things for vehicles (AIoV). However, integrating FL with AIoV poses challenges, including the absence of fine-grained access control, insufficient safeguards for FL tasks and vehicle identities, inadequate security for data transmission, and shortcomings in protecting data storage. These vulnerabilities may lead to risks such as vehicle tracking, model information theft, and data tampering. To address these challenges, we propose a privacy preservation mechanism for AIoV via cloud–edge–vehicle hierarchical sharding blockchain. First, we propose a hierarchical anonymous authentication scheme for IoV devices with stronger scalability and higher fault tolerance. Vehicles only know the attributes of each other or which shard they belong to. Second, we present a secure FL task assignment scheme for AIoV. Edge nodes utilize attribute-based encryption to deploy fine-grained FL tasks based on vehicle attributes. Only users who meet the attributes can decrypt the content, protecting FL tasks content and participant identities. Third, we present a secure data transmission scheme between AIoV devices to protect the identity and data privacy of both parties, while also achieving noninteractive key agreement. Additionally, we propose a scalable secure data sharing and storage scheme based on hierarchical sharding blockchain, aiming to reduce storage overhead and minimize trust costs. Mingzhe Zhai, Qianhong Wu, Yizhong Liu, Yang Yang 0062, Muhammad Ghulam, Prayag Tiwari |
IEEE Internet Things J. | 6 |
| 2025 | SGUVE-Net: Semantic-Guided Underwater Video Enhancement Network for Real-Time IoT-Based Marine MonitoringabstractUnderwater video enhancement is crucial for marine research and monitoring applications, particularly in the context of the Internet of Things (IoT), where autonomous underwater vehicles (AUVs) and sensor networks are deployed for environmental monitoring and tracking marine life. However, the scarcity of undistorted underwater video data and distortions, such as motion blur and water turbidity, limit the effectiveness of enhancement models. Existing methods typically focus on frame-by-frame enhancement and overlook temporal coherence and computational efficiency. To address these issues, we propose SGUVE-Net (Semantic-Guided Underwater Video Enhancement Network), which combines a multi-scale feature-aware network with a semantic branch for localized enhancement. The main branch employs an encoder-decoder architecture, combining spatial group shifting and dual attention mechanisms to fully exploit contextual information for precise alignment. In contrast, the semantic branch focuses on enhancing key regions of the video frames by incorporating high-level semantic cues, which improves motion tracking accuracy and mitigates motion blur, thereby enhancing video quality for real-time IoT applications. They complement each other to achieve differentiated modeling of static background details and dynamic target features. Experimental results show that SGUVE-Net outperforms state-of-the-art methods across several metrics, providing an effective solution for underwater video enhancement in IoT systems. Jingchun Zhou, Wenyu Fan, Bing Long, Dehuan Zhang, Zongxin He, Qiuping Jiang, Muhammad Ghulam |
IEEE Internet Things J. | 7 |
| 2025 | Degradation-Decoupling Vision Enhancement for Intelligent Underwater Robot Vision Perception SystemabstractUnderwater robots rely on high-quality visual data for precise monitoring and manipulation, yet complex underwater environments often degrade image quality through color distortion, texture blurring, and detail loss. Existing enhancement methods partially address these issues, but fail to effectively decouple nonlinear relationships among degradation factors, leading to inconsistent performance. To address these challenges, we propose a degradation-content decoupling-based underwater image enhancement network (DCDN). The framework integrates a super-fusion cascade module for dynamic feature weighting, reducing artifacts, and combines multichannel color space transformation with texture-guided correction to decouple and optimize degradation factors. This approach improves color fidelity and texture detail restoration by refining color information and adapting local textures. Experiments on public datasets demonstrate that DCDN outperforms existing methods in various underwater scenarios. This work enhances the visual capabilities of underwater robots, supporting intelligent transportation applications, such as marine logistics and underwater inspections. Jingchun Zhou, Chunjiang Liu, Bing Long, Dehuan Zhang, Qiuping Jiang, Muhammad Ghulam |
IEEE Internet Things J. | 6 |
| 2025 | Analysis of multimodal fusion strategies in deep learning for ischemic stroke lesion segmentation on computed tomography perfusion data
Chintha Sri Pothu Raju, Bala Chakravarthy Neelapu, Rabul Hussain Laskar, Muhammad Ghulam |
Multim. Tools Appl. | 4 |
| 2025 | Attention-based Fusion for Stroke Lesion Segmentation on Computed Tomography Perfusion DataabstractIn recent times, stroke has emerged as a significant threat to humans, transforming affected brain tissue into core and penumbra regions. As the penumbra becomes irreversible over time, early core region segmentation is crucial. Automatic segmentation systems offer an efficient alternative to manual segmentation and aid radiologists in stroke lesion segmentation using computed tomography and Computed Tomography Perfusion (CTP) maps that comprise four parameter maps. This automatic segmentation is increasingly used in interactive, multimedia-based systems for diagnostic tools and AI-driven health applications. Top-performing models that follow the patch-processing approach suffer from high inference times. To incorporate effective feature extraction at image-level inferences, which reduces the inference time, we present a hybrid fusion technique that combines early and bottleneck fusion, leveraging two separate encoders for effective feature extraction. Moreover, fusing the information from various fusion methods arbitrarily may not yield optimal results. Consequently, we have introduced two attention modules, i.e., cross-modal attention and cross-fusion attention modules, designed for the effective integration of features derived from diverse modalities and multiple fusion strategies, respectively. The findings highlight a considerable reduction in computational time alongside achieving a comparable Dice score. Additionally, the incorporation of hybrid fusion and attention modules in the baseline notably increased the Dice score from 0.482 to 0.521 in the validation dataset and achieved 0.48 in the test dataset of ISLES 2018. It also demonstrates competitive performance compared to existing models while maintaining efficient prediction times. Chintha Sri Pothu Raju, Rabul Hussain Laskar, Zulfiqar Ali 0001, Muhammad Ghulam |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | A novel integrated quantum-resistant cryptography for secure scientific data exchange in ad hoc networks
Kranthi Kumar Singamaneni, Muhammad Ghulam |
Ad Hoc Networks | 2 |
| 2024 | Fuzzy fractional generalized Bagley-Torvik equation with fuzzy Caputo gH-differentiability
Muhammad Ghulam, Muhammad Akram 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Randomized attention and dual-path system for electrocardiogram identity recognition
Le Sun 0003, Huiyun Li, Muhammad Ghulam |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | A cascaded deep learning framework for iris centre localization in facial imageabstractAbstract Accurate iris centre localization is crucial in many computer vision and facial biometric applications such as gaze estimation, human–computer interaction, iris recognition, and liveness detection. However, it is challenging in an uncontrolled environment due to variations like pose, scale, rotation, specular reflection, and image quality. Therefore, a cascaded deep learning framework for iris centre localization in facial images is proposed that is robust to the abovementioned variations. The proposed approach consists of (i) YOLOv3 for eye detection, (ii) UNet for iris segmentation, and (iii) statistical modelling for iris centre localization. The eyes are first detected using the YOLOv3, and subsequently, iris segmentation is performed within the detected eyes using the UNet. Following iris segmentation, statistical modelling is employed to enhance the localization accuracy of the iris centre. Experiments were performed on benchmark databases, resulting in a standardized error measure SED of 3.405 pixels for BioID and 3.259 pixels for GI4E databases. In addition, the robustness of the proposed eye detection model was further evaluated on the Yale B for illumination variations and the CAS‐PEAL for pose variations. Naseem Ahmad, Muhammad Ghulam, Kuldeep Singh Yadav, Rabul Hussain Laskar, Ashraf Hossain, Zulfiqar Ali 0001 |
Expert Syst. J. Knowl. Eng. | 2 |
| 2024 | Multi-focal channel attention for medical image segmentationabstractAbstract Medical image segmentation through the use of deep learning is becoming a trend targeting to automate disease detection and provide effective treatment. Traditionally huge manual efforts are associated with medical image processing by medical staff. The wide use of neural networks driven by the progressive advancement in computing power and the availability of training data provides instrumental means to automate medical image processing. A polyp refers to an abnormal growth that can occur in various parts of the body. While the majority of polyps are noncancerous or benign, there are instances where certain types can develop into cancer. Detecting and segmenting polyps is highly valuable for identifying early signs and potentially enabling more effective treatment of colon cancer. In this paper, we introduce a novel method for segmenting polyp images. The proposed method utilizes convolutional neural networks (CNNs) and employs an enhanced attention mechanism. To cater to a more granular attention view, this paper enhances the Convolutional Block Attention Module by introducing the multi‐focal channel attention (MFCA) concept, which we call MFCA. In the proposed MFCA channel attention, focal attention spots allow us to consider scattered areas of interest more effectively. To evaluate the effectiveness of the proposed method, multiple experiments are conducted on five well‐known benchmark datasets for polyp image segmentation: Kvasir, CVC‐Clinic DB, CVC‐Colon DB, CVC‐T, and ETIS‐Larib. Various testing scenarios are examined, including the impact of two focal attention spots and four focal attention spots. The experimental results demonstrated that the proposed method achieved superior performance compared to previous state‐of‐the‐art techniques on two of the benchmark datasets, and ranked second on a third dataset. Hamdan Al Jowair, Mansour Alsulaiman, Muhammad Ghulam |
Expert Syst. J. Knowl. Eng. | 3 |
| 2024 | QB-IMD: A Secure Medical Data Processing System With Privacy Protection Based on Quantum Blockchain for IoMTabstractSecurity and privacy are issues that cannot be ignored when collecting and processing medical data in the Internet of Medical Things (IoMT). The blockchain technology is a decentralized ledger system that has diverse application scenarios in the medical field. The blockchain technology relies on traditional cryptography to ensure data integrity and verifiability, but the creation of quantum computing has made it possible to break traditional encryption and signature methods. Therefore, quantum blockchain can provide a higher level of security for handling medical data. This article innovatively designs a new medical data processing system based on quantum blockchain (QB-IMD). In QB-IMD, a quantum blockchain structure and a novel electronic medical record algorithm (QEMR) are proposed to ensure that the processed data is legitimate and tamper-proof. QEMR combines quantum signature and quantum identity authentication to avoid the potential security risks of digital signatures. In addition, through delegated computing by quantum cloud, medical diagnostic data can be computed without leaking to quantum cloud servers, thus protecting user privacy. Through mathematical proof, theoretical analysis, and simulation, it is demonstrated that our scheme can resist six attacks and is feasible to protect user privacy. Zhiguo Qu, Yunyi Meng, Muhammad Ghulam, Prayag Tiwari |
IEEE Internet Things J. | 4 |
| 2024 | Fuzzy fractional epidemiological model for Middle East respiratory syndrome coronavirus on complex heterogeneous network using Caputo derivative
Muhammad Ghulam, Muhammad Akram 0001 |
Inf. Sci. | 1 |
| 2024 | Fuzzy Langevin fractional delay differential equations under granular derivative
Muhammad Ghulam, Muhammad Akram 0001, Nawab Hussain, Tofigh Allahviranloo |
Inf. Sci. | 1 |
| 2024 | Biometric identity recognition based on contrastive positive-unlabeled learning
Le Sun 0003, Yiwen Hua, Muhammad Ghulam |
J. Inf. Secur. Appl. | 3 |
| 2024 | Mitigating human fall injuries: A novel system utilizing 3D 4-stream convolutional neural networks and image fusion
Thamer Alanazi, Khalid Babutain, Muhammad Ghulam |
Image Vis. Comput. | 3 |
| 2024 | CABnet: A channel attention dual adversarial balancing network for multimodal image fusion
Le Sun 0003, Mengqi Tang, Muhammad Ghulam |
Image Vis. Comput. | 3 |
| 2024 | Tuberculosis detection in chest radiograph using convolutional neural network architecture and explainable artificial intelligence
Saad I. Nafisah, Muhammad Ghulam |
Neural Comput. Appl. | 2 |
| 2024 | Guest Editorial Insights of Machine Learning into Medical Decision Making Systems: From Research to Practice
Muhammad Ghulam, Farook Sattar, Zulfiqar Ali 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Ischemic Stroke Segmentation by Transformer and Convolutional Neural Network Using Few-Shot LearningabstractStroke is a major factor in causing disability and fatalities. Doctors use computerized tomography (CT) and magnetic resonance imaging (MRI) scans to assess the severity of a stroke. Automatic image segmentation can help doctors diagnose strokes more quickly and accurately, but it is challenging due to the variability of stroke lesions and the limited availability of labeled data. Deep learning is the cutting-edge technique of machine learning and artificial intelligence, which needs an extensive labeled dataset for effective training. Unfortunately, in the medical domain, the availability of labeled data is severely limited, posing a challenge for conventional deep- learning approaches. In this article, we introduce a system that utilizes deep learning in the form of fusing transformer-based and convolutional neural network (CNN)-based features and few-shot learning techniques to segment ischemic strokes in multimedia MRIs. To accomplish this, we employ two different methods. The first method involves parallel fusion, where we combine CNN-based and transformer-based features. The second method utilizes serial fusion, combining CNN-based and transformer models using few-shot learning. Through the integration of transformer and CNN models, we can extract both global and local features and enhance the system's performance. Moreover, we tackle the issue of limited labeled data by integrating few-shot learning techniques. Additionally, our system optimizes efficiency by selecting only the slices with lesions, disregarding unlesioned slices. The system under consideration is trained with the BraTS2020 dataset, evaluated on the ISLES 2015 dataset, and contrasted the performance with cutting-edge systems. The suggested system attains a dice coefficient score of 0.76, surpassing the scores of previous cutting-edge systems by a substantial margin. Fatima Alshehri, Muhammad Ghulam |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Computer-based Blind Diagnostic System for Classification of Healthy and Disordered VoicesabstractA large population around the world is suffering from voice-related complications. Computer-based voice disorder detection systems can play a substantial role in the early detection of voice disorders by providing complementary information to early-career otolaryngologists and general practitioners. However, various studies have concluded that the recording environment of voice samples affects disorder detection. This influence of the recording environment is a major obstacle in developing such systems when a local voice disorder database is not available. In addition, sometimes the number of samples is not sufficient for training the system. To overcome these issues, a blind detection system for voice disorders is designed and implemented in this study. Hence, without any prior knowledge of voice disorders, the proposed system has the ability to detect those disorders. The developed system relies only on healthy voice samples which can be recorded locally in the desired environment. The generation of a reference model for healthy subjects and decision criteria to detect voice disorders are two major tasks in the proposed systems. These tasks are implemented with two different types of speech features. Moreover, the unsupervised reference model is created by using DBSCAN and k-means algorithms. The overall performance of the system is 74.9%in terms of the geometric mean of sensitivity and specificity. The results of the proposed system are encouraging and better than the performance of Multidimensional Voice Program (MDVP) parameters which are widely used for disorder assessment by otolaryngologists in clinics. Zulfiqar Ali 0001, Alba Garcia Seco de Herrera, Tamer A. Mesallam, Muhammad Ghulam |
CBMS | 4 |
| 2023 | Incommensurate non-homogeneous system of fuzzy linear fractional differential equations using the fuzzy bunch of real functions
Muhammad Akram 0001, Muhammad Ghulam, Tofigh Allahviranloo, Witold Pedrycz |
Fuzzy Sets Syst. | 2 |
| 2023 | Dynamic Convolution With Multilevel Attention for EEG-Based Motor Imagery DecodingabstractBrain–computer interface (BCI) is an innovative technology that utilizes artificial intelligence (AI) and wearable electroencephalography (EEG) sensors to decode brain signals and enhance the quality of life. EEG-based motor imagery (MI) brain signal is used in many BCI applications, including smart healthcare, smart homes, and robotics control. However, the restricted ability to decode brain signals is a major factor preventing BCI technology from expanding significantly. In this study, we introduce a dynamic attention temporal convolutional network (D-ATCNet) for decoding EEG-based MI signals. The D-ATCNet model uses dynamic convolution (Dy-conv) and multilevel attention to enhance the performance of MI classification with a relatively small number of parameters. D-ATCNet has two main blocks: 1) dynamic and 2) temporal convolution. Dy-conv uses multilevel attention to encode low-level MI-EEG information and temporal convolution uses shifted window with self-attention to extract high-level temporal information from the encoded signal. The proposed model performs better than the existing methods with an accuracy of 71.3% for subject independent and 87.08% for subject dependent using the BCI competition IV-2a data set. Hamdi Altaheri, Muhammad Ghulam, Mansour Alsulaiman |
IEEE Internet Things J. | 2 |
| 2023 | Internet of Things: Device Capabilities, Architectures, Protocols, and Smart Applications in Healthcare DomainabstractNowadays, the Internet has spread to practically every country around the world and is having unprecedented effects on people’s lives. The Internet of Things (IoT) is getting more popular and has a high level of interest in both practitioners and academicians in the age of wireless communication due to its diverse applications. The IoT is a technology that enables everyday things to become savvier, everyday computation toward becoming intellectual, and everyday communication to become a little more insightful. In this article, the most common and popular IoT device capabilities, architectures, and protocols are demonstrated in brief to provide a clear overview of the IoT technology to the researchers in this area. The common IoT device capabilities, including hardware (Raspberry Pi, Arduino, and ESP8266) and software (operating systems (OSs), and built-in tools) platforms are described in detail. The widely used architectures that have recently evolved and used are the three-layer architecture, service-oriented architecture, and middleware-based architecture. The popular protocols for IoT are demonstrated which include constrained application protocol, message queue telemetry transport, extensible messaging and presence protocol, advanced message queuing protocol, data distribution service, low power wireless personal area network, Bluetooth low energy, and ZigBee that are frequently utilized to develop smart IoT applications. Additionally, this research provides an in-depth overview of the potential healthcare applications based on IoT technologies in the context of addressing various healthcare concerns. Finally, this article summarizes state-of-the-art knowledge, highlights open issues and shortcomings, and provides recommendations for further studies which would be quite beneficial to anyone with a desire to work in this field and make breakthroughs to get expertise in this area. Md. Milon Islam, Sheikh Nooruddin, Fakhri Karray, Muhammad Ghulam |
IEEE Internet Things J. | 4 |
| 2023 | Stacked Autoencoder-Based Intrusion Detection System to Combat Financial FraudulentabstractWith the rapid progress of wireless communication technologies along with their digital revolutions, the quantity of the Internet of Things (IoT) has been increased by manifolds, resulting in a huge increase in data volume and network traffic. It became easier for an intruder to pretend as a valid service provider, and generate different types of network attacks. This becomes even more severe when the service involves digital financial transactions for possible urbanization. This article proposes an intrusion detection system (IDS) based on a stacked autoencoder (AE) and a deep neural network (DNN). The stacked AE learns the features of the input network record in an unsupervised manner to decrease the feature width. Then, the DNN is trained in a supervised manner to extract deep-learned features for the classifier. In the proposed system, the stacked AE has two latent layers and the DNN has two or three layers, where each layer has a fully connected layer, a batch normalization, and a dropout. The system was evaluated on three publicly available data sets: 1) KDDCup99; 2) NSL-KDD; and 3) aegean Wi-Fi intrusion data sets. Experimental results exhibited that the proposed IDS achieved 94.2%, 99.7%, and 99.9% accuracy, respectively, for multiclass classification. Muhammad Ghulam, M. Shamim Hossain, Sahil Garg |
IEEE Internet Things J. | 1 |
| 2023 | Explicit analytical solutions of an incommensurate system of fractional differential equations in a fuzzy environment
Muhammad Akram 0001, Muhammad Ghulam, Tofigh Allahviranloo |
Inf. Sci. | 2 |
| 2023 | A few-shot learning-based ischemic stroke segmentation system using weighted MRI fusion
Fatima Alshehri, Muhammad Ghulam |
Image Vis. Comput. | 2 |
| 2023 | Multi parallel U-net encoder network for effective polyp image segmentation
Hamdan Al Jowair, Mansour Alsulaiman, Muhammad Ghulam |
Image Vis. Comput. | 3 |
| 2023 | Medical image-based detection of COVID-19 using Deep Convolution Neural Networks
Loveleen Gaur, Ujwal Bhatia, N. Z. Jhanjhi, Muhammad Ghulam, Mehedi Masud |
Multim. Syst. | 4 |
| 2023 | Deep learning techniques for classification of electroencephalogram (EEG) motor imagery (MI) signals: a review
Hamdi Altaheri, Muhammad Ghulam, Mansour Alsulaiman, Syed Umar Amin, Ghadir Ali Altuwaijri, Wadood Abdul, Mohamed Abdelkader Bencherif, Mohammed Faisal |
Neural Comput. Appl. | 2 |
| 2023 | Physics-Informed Attention Temporal Convolutional Network for EEG-Based Motor Imagery ClassificationabstractThe brain-computer interface (BCI) is a cutting-edge technology that has the potential to change the world. Electroencephalogram (EEG) motor imagery (MI) signal has been used extensively in many BCI applications to assist disabled people, control devices or environments, and even augment human capabilities. However, the limited performance of brain signal decoding is restricting the broad growth of the BCI industry. In this article, we propose an attention-based temporal convolutional network (ATCNet) for EEG-based motor imagery classification. The ATCNet model utilizes multiple techniques to boost the performance of MI classification with a relatively small number of parameters. ATCNet employs scientific machine learning to design a domain-specific deep learning model with interpretable and explainable features, multihead self-attention to highlight the most valuable features in MI-EEG data, temporal convolutional network to extract high-level temporal features, and convolutional-based sliding window to augment the MI-EEG data efficiently. The proposed model outperforms the current state-of-the-art techniques in the BCI Competition IV-2a dataset with an accuracy of 85.38% and 70.97% for the subject-dependent and subject-independent modes, respectively. Hamdi Altaheri, Muhammad Ghulam, Mansour Alsulaiman |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Light Deep Models for Cognitive Computing in Intelligent Transportation SystemsabstractThe paper proposes light convolutional neural network (CNN) models for the use of cognitive networking in an intelligent transportation system (ITS). There are two CNN models, one with 1D convolution and connectors, and the other with a tree-like structure. The 1D CNN model is deployed to process 1D temporal data such as the driver’s body temperature and electrocardiogram (ECG) data to measure emotion, while the deep tree CNN model is used to process image data obtained from car camera sensors. As the driver’s cognitive state can frequently change depending on the situation and location of the car, different edge controllers should handle the car sensors’ data within a short period of time. The tree-based deep learning model that can be branched and processed independently in the edge devices can be executed with less computation. This reduces the load and the time of the execution of the model. The light 1D CNN model has less learnable parameters, and hence can be executed in real-time. The cognitive state of a driver is measured by the facial emotion, body temperature, and ECG signal of the driver. The proposed is tested using a publicly available facial emotion database, and the accuracy and the information density are around 94-96% and 4.4, respectively. Muhammad Ghulam, M. Shamim Hossain |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Lightweight and Anonymity-Preserving User Authentication Scheme for IoT-Based HealthcareabstractInternet of Things (IoT) produces massive heterogeneous data from various applications, including digital health, smart hospitals, automated pathology labs, and so forth. IoT sensor nodes are integrated with the medical equipment to enable the health workers to monitor the patients’ health condition and appliances in real time. However, due to security vulnerabilities, an unauthorized user can access health-related information or control the IoT nodes attached to the patient’s body resulting in unprecedented outcomes. Due to wireless channels as a medium of communication, IoT poses several threats such as a denial of service attack, man-in-the-middle attack, and modification attack to the IoT networks’ security and privacy. The proposed research presents a lightweight and anonymity-preserving user authentication protocol to counter these security threats. The given scheme establishes a secure session for the legitimate user and prohibits unauthorized users from gaining access to the IoT sensor nodes. The proposed protocol uses only lightweight cryptography primitives (hash) to alleviate the node’s tiny processor burden. The proposed protocol is efficient and superior because it has low computational and communication costs than conventional protocols. The proposed scheme uses password protection to let only the legitimate user access the IoT sensor nodes to obtain the patient’s real-time health report. Mehedi Masud, Gurjot Singh Gaba, Karanjeet Choudhary, M. Shamim Hossain, Mohammed F. Alhamid, Muhammad Ghulam |
IEEE Internet Things J. | 6 |
| 2022 | Special issue deep learning for multimedia healthcareabstractText, radiological pictures, audio notes, video, and other types of multimedia healthcare data are all generated by today's smart healthcare system [1].The evolution of COVID-19 has resulted in an incremental rise in current healthcare data.The study of multimodal healthcare data on such a big scale has revealed both obstacles and potential.Thanks to artificial intelligence (AI) and, more specifically, deep learning (DL) algorithms, which have been widely used by researchers for handling massive amounts of epidemic data, predicting live epidemic crises, and initiating new research directions in the analysis of healthcare multimedia data [2].As a result, deep learning for multimedia healthcare data analysis is becoming a hot topic in multimedia and computer vision research.The call for papers attracted 54 submissions and after a rigorous review, 20 papers have been accepted for this special issue.A brief summary of papers in this special issue is presented in the following:The paper titled "A Novel Study for Automatic Twoclass Covid-19 Diagnosis (between Covid-19 and Healthy, Pneumonia) on X-ray Images using Texture Analysis and 2-D/3-D Convolutional Neural Networks" aims to diagnose COVID-19 early using X-ray images, automatic two-class classification was carried out in four different titles: COVID-19/Healthy, COVID-19 Pneumonia/Bacterial Pneumonia, COVID-19 Pneumonia/Viral Pneumonia, and COVID-19 Pneumonia/Other Pneumonia.In the study, besides using M. Shamim Hossain, Josu Bilbao, Diana P. Tobón, Muhammad Ghulam, Abdulmotaleb El Saddik |
Multim. Syst. | 4 |
| 2022 | Deep learning in multimedia healthcare applications: a review
Diana P. Tobón, M. Shamim Hossain, Muhammad Ghulam, Josu Bilbao, Abdulmotaleb El Saddik |
Multim. Syst. | 3 |
| 2022 | Attention-Inception and Long- Short-Term Memory-Based Electroencephalography Classification for Motor Imagery Tasks in RehabilitationabstractIn recent years, the contributions of deep learning have had a phenomenal impact on electroencephalography-based brain-computer interfaces. While the decoding accuracy of electroencephalography signals has continued to increase, the process has caused deep learning models to continuously expand in terms of size and computational resource requirements. However, due to their increased size and computational requirements, it has become difficult to embed, store, and execute deep learning models for artificial intelligence of things, cloud-based, or edge devices used in rehabilitation. Hence, this article proposes a novel deep learning-based lightweight model based on attention-inception convolutional neural network and long- short-term memory. The proposed model achieves excellent accuracy on public competition datasets while requiring few parameters and low computational time. Using the BCI competition IV 2a dataset and the high gamma dataset, the proposed model achieved 82.8% and 97.1% accuracies, respectively. Syed Umar Amin, Hamdi Altaheri, Muhammad Ghulam, Wadood Abdul, Mansour Alsulaiman |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Privacy-Preserving Serverless Computing Using Federated Learning for Smart GridsabstractThe smart power grid is a critical energy infrastructure where real-time electricity usage data is collected to predict future energy requirements. The existing prediction models focus on the centralized frameworks, where the collected data from various home area networks (HANs) are forwarded to a central server. This process leads to cybersecurity threats. This article proposes a federated learning based model with privacy preservation of smart grids data using serverless cloud computing. The model considers the blockchain-enabled dew servers in each HAN for local data storage and local model training. Advanced perturbation and normalization techniques are used to reduce the inverse impact of irregular workload on the training results. The experiment conducted on benchmarks datasets demonstrates that the proposed model minimizes the computation and communication costs, attacking probability, and improves the test accuracy. Overall, the proposed model enables smart grids with robust privacy preservation and high accuracy. Mehedi Masud, M. Shamim Hossain, Avinash Kaur, Muhammad Ghulam, Ahmed Ghoneim |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | Special Section on AI-empowered Multimedia Data Analytics for Smart HealthcareabstractNo abstract available. M. Shamim Hossain, Rita Cucchiara, Muhammad Ghulam, Diana P. Tobón, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Transfer reinforcement learning-based road object detection in next generation IoT domain
Ke Wang 0068, Chien-Ming Chen 0001, M. Shamim Hossain, Muhammad Ghulam, Sachin Kumar 0002, Saru Kumari |
Comput. Networks | 4 |
| 2021 | Incentive mechanism for collaborative distributed learning in Artificial Intelligence of Things
Jiali Yin, M. Shamim Hossain, Muhammad Ghulam |
Future Gener. Comput. Syst. | 4 |
| 2021 | A Lightweight and Robust Secure Key Establishment Protocol for Internet of Medical Things in COVID-19 Patients CareabstractDue to the outbreak of COVID-19, the Internet of Medical Things (IoMT) has enabled the doctors to remotely diagnose the patients, control the medical equipment, and monitor the quarantined patients through their digital devices. Security is a major concern in IoMT because the Internet of Things (IoT) nodes exchange sensitive information between virtual medical facilities over the vulnerable wireless medium. Hence, the virtual facilities must be protected from adversarial threats through secure sessions. This article proposes a lightweight and physically secure mutual authentication and secret key establishment protocol that uses physical unclonable functions (PUFs) to enable the network devices to verify the doctor's legitimacy (user) and sensor node before establishing a session key. PUF also protects the sensor nodes deployed in an unattended and hostile environment from tampering, cloning, and side-channel attacks. The proposed protocol exhibits all the necessary security properties required to protect the IoMT networks, like authentication, confidentiality, integrity, and anonymity. The formal AVISPA and informal security analysis demonstrate its robustness against attacks like impersonation, replay, a man in the middle, etc. The proposed protocol also consumes fewer resources to operate and is safe from physical attacks, making it more suitable for IoT-enabled medical network applications. Mehedi Masud, Gurjot Singh Gaba, Salman AlQahtani, Muhammad Ghulam, Brij B. Gupta, Pardeep Kumar 0001, Ahmed Ghoneim |
IEEE Internet Things J. | 4 |
| 2021 | Emotion Recognition for Cognitive Edge Computing Using Deep LearningabstractThe growing use of the Internet of Things (IoT) has increased the volume of data to be processed by manifolds. Edge computing can lessen the load of transmitting a massive volume of data to the cloud. It can also provide reduced latency and real-time experience to the users. This article proposes an emotion recognition system from facial images based on edge computing. A convolutional neural network (CNN) model is proposed to recognize emotion. The model is trained in a cloud during off time and downloaded to an edge server. During the testing, an end device such as a smartphone captures a face image and does some preprocessing, which includes face detection, face cropping, contrast enhancement, and image resizing. The preprocessed image is then sent to the edge server. The edge server runs the CNN model and infers a decision on emotion. The decision is then transmitted back to the smartphone. Two data sets, JAFFE and extended Cohn–Kanade (CK+), are used for the evaluation. Experimental results show that the proposed system is energy efficient, has less learnable parameters, and good recognition accuracy. The accuracies using the JAFFE and CK+ data sets are 93.5% and 96.6%, respectively. Muhammad Ghulam, M. Shamim Hossain |
IEEE Internet Things J. | 1 |
| 2021 | Blockchain for Secure-GaS: Blockchain-Powered Secure Natural Gas IoT System With AI-Enabled Gas Prediction and Transaction in Smart CityabstractThe traditional natural gas Internet-of-Things (IoT) system has many problems, such as centralized management of resources, noncirculation of data between stations, insecurity of transaction information or account books, and lack of contract consensus. In order to ensure data security and reliable transaction, this article introduces artificial intelligence (AI) and blockchain technology and constructs an AI-enabled and blockchain-powered natural gas IoT system in a smart city. In this article, the natural gas output prediction model based on temporal pattern attention-based LSTMs (TPA-LSTMs) is used to enable the system to sense the change of natural gas deliverability. In addition, we establish a blockchain-based secure natural gas transaction scheme, which dynamically matches the purchase contract and sale contract to maximize the interests of the buyer and the seller and obtain a transaction contract. The experimental results show that our model can predict the output value of natural gas in real time and select the appropriate transaction matching scheme according to the dynamic demand for sales. Wenjing Xiao, Haoquan Wang, M. Shamim Hossain, Mubarak Alrashoud, Muhammad Ghulam |
IEEE Internet Things J. | 7 |
| 2021 | EEG-Based Pathology Detection for Home Health MonitoringabstractAn electroencephalogram (EEG)-based remote pathology detection system is proposed in this study. The system uses a deep convolutional network consisting of 1D and 2D convolutions. Features from different convolutional layers are fused using a fusion network. Various types of networks are investigated; the types include a multilayer perceptron (MLP) with a varying number of hidden layers, and an autoencoder. Experiments are done using a publicly available EEG signal database that contains two classes: normal and abnormal. The experimental results demonstrate that the proposed system achieves greater than 89% accuracy using the convolutional network followed by the MLP with two hidden layers. The proposed system is also evaluated in a cloud-based framework, and its performance is found to be comparable with the performance obtained using only a local server. Muhammad Ghulam, M. Shamim Hossain, Neeraj Kumar 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2021 | Efficient Flow Processing in 5G-Envisioned SDN-Based Internet of Vehicles Using GPUsabstractIn the 5G-envisioned Internet of vehicles (IoV), a significant volume of data is exchanged through networks between intelligent transport systems (ITS) and clouds or fogs. With the introduction of Software-Defined Networking (SDN), the problems mentioned above are resolved by high-speed flow-based processing of data in network systems. To classify flows of packets in the SDN network, high throughput packet classification systems are needed. Although software packet classifiers are cheaper and more flexible than hardware classifiers, they could only deliver limited performance. A key idea to resolve this problem is parallelizing packet classification on graphical processing units (GPUs). In this paper, we study parallel forms of Tuple Space Search and Pruned Tuple Space Search algorithms for the flow classification suitable for GPUs using CUDA (Compute Unified Device Architecture). The key idea behind the offered methodology is to transfer the stream of packets from host memory to the global memory of the CUDA device, then assigning each of them to a classifier thread. To evaluate the proposed method, the GPU-based versions of the algorithms were implemented on two different CUDA devices, and two different CPU-based implementations of the algorithms were used as references. Experimental results showed that GPU computing enhances the performance of Pruned Tuple Space Search remarkably more than Tuple Space Search. Moreover, results evinced the computational efficiency of the proposed method for parallelizing packet classification algorithms. Mahdi Abbasi, Ali Najafi, Milad Rafiee, Mohammad Reza Khosravi, Varun G. Menon, Muhammad Ghulam |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | LACCVoV: Linear Adaptive Congestion Control With Optimization of Data Dissemination Model in Vehicle-to-Vehicle CommunicationabstractVehicle-to-vehicle communication assists road-side information exchange granting ease of access and sharing between users. The communication between the vehicles is short-lived due to interference and data congestion in the resource constraint medium. This manuscript introduces a linear adaptive congestion control (LACC) augmenting the benefits of greedy routing and data dissemination model (DDM). LACC focuses on selecting beneficiary vehicle by assessing its end-to-end service capacity and link stability preference. Different from the conventional greedy approach, routing is aided by a linear integer programming module for smart decisions on neighbor selection. The interrupts in data transmission and forwarding due to non-localized vehicles, congested routing paths and paused transmissions are addressed using LACC as a series of linear optimization. This helps to improve the performance of vehicular communication estimated using delay, message delivery, outage, and beacon messages. Arun Kumar Sangaiah, Jaya Subalakshmi Ramamoorthi, Joel J. P. C. Rodrigues, Mohamed Abdur Rahman 0001, Muhammad Ghulam, Mubarak Alrashoud |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Pre-Trained Convolutional Neural Networks for Breast Cancer Detection Using Ultrasound ImagesabstractVolunteer computing based data processing is a new trend in healthcare applications. Researchers are now leveraging volunteer computing power to train deep learning networks consisting of billions of parameters. Breast cancer is the second most common cause of death in women among cancers. The early detection of cancer may diminish the death risk of patients. Since the diagnosis of breast cancer manually takes lengthy time and there is a scarcity of detection systems, development of an automatic diagnosis system is needed for early detection of cancer. Machine learning models are now widely used for cancer detection and prediction research for improving the successive therapy of patients. Considering this need, this study implements pre-trained convolutional neural network based models for detecting breast cancer using ultrasound images. In particular, we tuned the pre-trained models for extracting key features from ultrasound images and included a classifier on the top layer. We measured accuracy of seven popular state-of-the-art pre-trained models using different optimizers and hyper-parameters through fivefold cross validation. Moreover, we consider Grad-CAM and occlusion mapping techniques to examine how well the models extract key features from the ultrasound images to detect cancers. We observe that after fine tuning, DenseNet201 and ResNet50 show 100% accuracy with Adam and RMSprop optimizers. VGG16 shows 100% accuracy using the Stochastic Gradient Descent optimizer. We also develop a custom convolutional neural network model with a smaller number of layers compared to large layers in the pre-trained models. The model also shows 100% accuracy using the Adam optimizer in classifying healthy and breast cancer patients. It is our belief that the model will assist healthcare experts with improved and faster patient screening and pave a way to further breast cancer research. Mehedi Masud, M. Shamim Hossain, Hesham Alhumyani, Sultan S. Alshamrani, Omar Cheikhrouhou, Saleh Ibrahim, Muhammad Ghulam, Amr Ezz El-Din Rashed, Brij B. Gupta |
ACM Trans. Internet Techn. | 7 |
| 2021 | eDiaPredict: An Ensemble-based Framework for Diabetes PredictionabstractMedical systems incorporate modern computational intelligence in healthcare. Machine learning techniques are applied to predict the onset and reoccurrence of the disease, identify biomarkers for survivability analysis depending upon certain health conditions of the patient. Early prediction of diseases like diabetes is essential as the number of diabetic patients of all age groups is increasing rapidly. To identify underlying reasons for the onset of diabetes in its early stage has become a challenging task for medical practitioners. Continuously increasing diabetic patient data has necessitated for the applications of efficient machine learning algorithms, which learns from the trends of the underlying data and recognizes the critical conditions in patients. In this article, an ensemble-based framework named e DiaPredict is proposed. It uses ensemble modeling, which includes an ensemble of different machine learning algorithms comprising XGBoost, Random Forest, Support Vector Machine, Neural Network, and Decision tree to predict diabetes status among patients. The performance of eDiaPredict has been evaluated using various performance parameters like accuracy, sensitivity, specificity, Gini Index, precision, area under curve, area under convex hull, minimum error rate, and minimum weighted coefficient. The effectiveness of the proposed approach is shown by its application on the PIMA Indian diabetes dataset wherein an accuracy of 95% is achieved. Ashima Singh, Arwinder Dhillon, Neeraj Kumar 0001, M. Shamim Hossain, Muhammad Ghulam, Manoj Kumar 0008 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | Multiple contents offloading mechanism in AI-enabled opportunistic networks
Wei-Che Chien, Shih-Yun Huang, Chin-Feng Lai, Han-Chieh Chao, M. Shamim Hossain, Muhammad Ghulam |
Comput. Commun. | 6 |
| 2020 | Joint power and time allocation in energy harvesting of UAV operating system
Qiang Liu 0020, Jun Yang 0014, Jing Lv, Kai Hwang 0001, M. Shamim Hossain, Muhammad Ghulam |
Comput. Commun. | 7 |
| 2020 | Deep learning-based intelligent face recognition in IoT-cloud environment
Mehedi Masud, Muhammad Ghulam, Hesham Alhumyani, Sultan S. Alshamrani, Omar Cheikhrouhou, Saleh Ibrahim, M. Shamim Hossain |
Comput. Commun. | 2 |
| 2020 | Cervical cancer classification using convolutional neural networks and extreme learning machines
Ahmed Ghoneim, Muhammad Ghulam, M. Shamim Hossain |
Future Gener. Comput. Syst. | 2 |
| 2020 | Towards energy-aware cloud-oriented cyber-physical therapy system
M. Shamim Hossain, Mohamed Abdur Rahman 0001, Muhammad Ghulam |
Future Gener. Comput. Syst. | 3 |
| 2020 | Follow me Robot-Mind: Cloud brain based personalized robot service with migration
Long Hu, Yinging Jiang, Fangxin Wang 0001, Kai Hwang 0001, M. Shamim Hossain, Muhammad Ghulam |
Future Gener. Comput. Syst. | 6 |
| 2020 | A knowledge-driven approach for activity recognition in smart homes based on activity profiling
Majdi Rawashdeh, Mohammed G. H. al Zamil, Samer Samarah, M. Shamim Hossain, Muhammad Ghulam |
Future Gener. Comput. Syst. | 5 |
| 2020 | Attention-based sentiment analysis using convolutional and recurrent neural network
Mohd Usama, Belal Ahmad, Enmin Song, M. Shamim Hossain, Mubarak Alrashoud, Muhammad Ghulam |
Future Gener. Comput. Syst. | 6 |
| 2020 | Blockchain-Enabled Distributed Security Framework for Next-Generation IoT: An Edge Cloud and Software-Defined Network-Integrated ApproachabstractThe Internet of Things (IoT) plays a vital role in the real world by providing autonomous support for communications and operations, thus enabling and promoting novel services that are commonly used in day-to-day life. It is important to do research on security frameworks for next-generation IoT and develop state-of-the-art confidentiality protection schemes to deal with various attacks on IoT networks. In order to offer prominent features like continuous confidentiality, authentication, and robustness, the blockchain technology comes out as a sustainable solution. A blockchain-enabled distributed security framework using edge cloud and software-defined networking (SDN) is presented in this article. The security attack detection is achieved at the cloud layer, and security attacks are consequently reduced at the edge layer of the IoT network. The SDN-enabled gateway offers dynamic network traffic flow management, which contributes to the security attack recognition through determining doubtful network traffic flows and diminishes security attacks through hindering doubtful flows. The results obtained show that the proposed security framework can efficiently and effectively meet the data confidentiality challenges introduced by the integration of blockchain, edge cloud, and SDN paradigm. Darshan Vishwasrao Medhane, Arun Kumar Sangaiah, M. Shamim Hossain, Muhammad Ghulam, Jin Wang 0001 |
IEEE Internet Things J. | 4 |
| 2020 | Privacy-preserving based task allocation with mobile edge clouds
Yongfeng Qian, M. Shamim Hossain, Long Hu, Muhammad Ghulam, Syed Umar Amin |
Inf. Sci. | 5 |
| 2020 | Toward cognitive support for automated defect detection
Ehab Essa, M. Shamim Hossain, Ahmad S. Tolba 0001, Hazem M. Raafat, Samir Elmougy, Muhammad Ghulam |
Neural Comput. Appl. | 6 |
| 2020 | Tree-Based Deep Networks for Edge DevicesabstractThis article proposes a tree-based deep model for effective load distribution to edge devices without much loss of accuracy. The input image is divided into groups of volumes, and each volume is passed through a tree structure. The tree structure has many branches and levels, each of which is represented by a convolutional layer. The layers are independent of each other. Therefore, various edge devices can update the parameters of the layers in parallel independently. Experiments are performed using a benchmark dataset and a publicly available date fruits database. Experimental results show that the proposed model has a high information density by reducing the number of parameters without much loss of accuracy. Muhammad Ghulam, M. Shamim Hossain, Abdulsalam Yassine |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Leveraging Deep Learning Techniques for Malaria Parasite Detection Using Mobile ApplicationabstractMalaria is a contagious disease that affects millions of lives every year. Traditional diagnosis of malaria in laboratory requires an experienced person and careful inspection to discriminate healthy and infected red blood cells (RBCs). It is also very time-consuming and may produce inaccurate reports due to human errors. Cognitive computing and deep learning algorithms simulate human intelligence to make better human decisions in applications like sentiment analysis, speech recognition, face detection, disease detection, and prediction. Due to the advancement of cognitive computing and machine learning techniques, they are now widely used to detect and predict early disease symptoms in healthcare field. With the early prediction results, healthcare professionals can provide better decisions for patient diagnosis and treatment. Machine learning algorithms also aid the humans to process huge and complex medical datasets and then analyze them into clinical insights. This paper looks for leveraging deep learning algorithms for detecting a deadly disease, malaria, for mobile healthcare solution of patients building an effective mobile system. The objective of this paper is to show how deep learning architecture such as convolutional neural network (CNN) which can be useful in real-time malaria detection effectively and accurately from input images and to reduce manual labor with a mobile application. To this end, we evaluate the performance of a custom CNN model using a cyclical stochastic gradient descent (SGD) optimizer with an automatic learning rate finder and obtain an accuracy of 97.30% in classifying healthy and infected cell images with a high degree of precision and sensitivity. This outcome of the paper will facilitate microscopy diagnosis of malaria to a mobile application so that reliability of the treatment and lack of medical expertise can be solved. Mehedi Masud, Hesham Alhumyani, Sultan S. Alshamrani, Omar Cheikhrouhou, Saleh Ibrahim, Muhammad Ghulam, M. Shamim Hossain, Mohammad Shorfuzzaman |
Wirel. Commun. Mob. Comput. | 6 |
| 2020 | Light Deep Model for Pulmonary Nodule Detection from CT Scan Images for Mobile DevicesabstractThe emergence of cognitive computing and big data analytics revolutionize the healthcare domain, more specifically in detecting cancer. Lung cancer is one of the major reasons for death worldwide. The pulmonary nodules in the lung can be cancerous after development. Early detection of the pulmonary nodules can lead to early treatment and a significant reduction of death. In this paper, we proposed an end-to-end convolutional neural network- (CNN-) based automatic pulmonary nodule detection and classification system. The proposed CNN architecture has only four convolutional layers and is, therefore, light in nature. Each convolutional layer consists of two consecutive convolutional blocks, a connector convolutional block, nonlinear activation functions after each block, and a pooling block. The experiments are carried out using the Lung Image Database Consortium (LIDC) database. From the LIDC database, 1279 sample images are selected of which 569 are noncancerous, 278 are benign, and the rest are malignant. The proposed system achieved 97.9% accuracy. Compared to other famous CNN architecture, the proposed architecture has much lesser flops and parameters and is thereby suitable for real-time medical image analysis. Mehedi Masud, Muhammad Ghulam, M. Shamim Hossain, Hesham Alhumyani, Sultan S. Alshamrani, Omar Cheikhrouhou, Saleh Ibrahim |
Wirel. Commun. Mob. Comput. | 2 |
| 2019 | Deep Learning for EEG motor imagery classification based on multi-layer CNNs feature fusion
Syed Umar Amin, Mansour Alsulaiman, Muhammad Ghulam, Mohamed Amine Mekhtiche, M. Shamim Hossain |
Future Gener. Comput. Syst. | 3 |
| 2019 | Deep convolutional tree networks
Abduljawad A. Amory, Muhammad Ghulam, Hassan Mathkour |
Future Gener. Comput. Syst. | 2 |
| 2019 | IoT big data analytics for smart homes with fog and cloud computing
Abdulsalam Yassine, Shailendra Singh 0007, M. Shamim Hossain, Muhammad Ghulam |
Future Gener. Comput. Syst. | 4 |
| 2019 | Emotion recognition using secure edge and cloud computing
M. Shamim Hossain, Muhammad Ghulam |
Inf. Sci. | 2 |
| 2019 | Smart healthcare monitoring: a voice pathology detection paradigm for smart cities
M. Shamim Hossain, Muhammad Ghulam, Atif Alamri |
Multim. Syst. | 2 |
| 2019 | Automatic Fruit Classification Using Deep Learning for Industrial ApplicationsabstractFruit classification is an important task in many industrial applications. A fruit classification system may be used to help a supermarket cashier identify the fruit species and prices. It may also be used to help people decide whether specific fruit species meet their dietary requirements. In this paper, we propose an efficient framework for fruit classification using deep learning. More specifically, the framework is based on two different deep learning architectures. The first is a proposed light model of six convolutional neural network layers, whereas the second is a fine-tuned visual geometry group-16 pretrained deep learning model. Two color image datasets, one of which is publicly available, are used to evaluate the proposed framework. The first dataset (dataset 1) consists of clear fruit images, whereas the second dataset (dataset 2) contains fruit images that are challenging to classify. Classification accuracies of 99.49% and 99.75% were achieved on dataset 1 for the first and second models, respectively. On dataset 2, the first and second models obtained accuracies of 85.43% and 96.75%, respectively. M. Shamim Hossain, Muneer H. Al-Hammadi, Muhammad Ghulam |
IEEE Trans. Ind. Informatics | 3 |
| 2019 | Enforcing Position-Based Confidentiality With Machine Learning Paradigm Through Mobile Edge Computing in Real-Time Industrial InformaticsabstractPosition-based services (PBSs) that deliver networked amenities based on roaming user's positions have become progressively popular with the propagation of smart mobile devices. Position is one of the important circumstances in PBSs. For effective PBSs, extraction and recognition of meaningful positions and estimating the subsequent position are fundamental procedures. Several researchers and practitioners have tried to recognize and predict positions using various techniques; however, only few deliberate the progress of position-based real-time applications considering significant tasks of PBSs. In this paper, a method for conserving position confidentiality of roaming PBSs users using machine learning techniques is proposed. We recommend a three-phase procedure for roaming PBS users. It identifies user position by merging decision trees and k-nearest neighbor and estimates user destination along with the position track sequence using hidden Markov models. Moreover, a mobile edge computing service policy is followed in the proposed paradigm, which will ensure the timely delivery of PBSs. The benefits of mobile edge service policy offer position confidentiality and low latency by means of networking and computing services at the vicinity of roaming users. Thorough experiments are conducted, and it is confirmed that the proposed method achieved above 90% of the position confidentiality in PBSs. Arun Kumar Sangaiah, Darshan Vishwasrao Medhane, Tao Han 0004, M. Shamim Hossain, Muhammad Ghulam |
IEEE Trans. Ind. Informatics | 5 |
| 2019 | Applying Deep Learning for Epilepsy Seizure Detection and Brain Mapping VisualizationabstractDeep Convolutional Neural Network (CNN) has achieved remarkable results in computer vision tasks for end-to-end learning. We evaluate here the power of a deep CNN to learn robust features from raw Electroencephalogram (EEG) data to detect seizures. Seizures are hard to detect, as they vary both inter- and intra-patient. In this article, we use a deep CNN model for seizure detection task on an open-access EEG epilepsy dataset collected at the Boston Children's Hospital. Our deep learning model is able to extract spectral, temporal features from EEG epilepsy data and use them to learn the general structure of a seizure that is less sensitive to variations. For cross-patient EEG data, our method produced an overall sensitivity of 90.00%, specificity of 91.65%, and overall accuracy of 98.05% for the whole dataset of 23 patients. The system can detect seizures with an accuracy of 99.46%. Thus, it can be used as an excellent cross-patient seizure classifier. The results show that our model performs better than the previous state-of-the-art models for patient-specific and cross-patient seizure detection task. The method gave an overall accuracy of 99.65% for patient-specific data. The system can also visualize the special orientation of band power features. We use correlation maps to relate spectral amplitude features to the output in the form of images. By using the results from our deep learning model, this visualization method can be used as an effective multimedia tool for producing quick and relevant brain mapping images that can be used by medical experts for further investigation. M. Shamim Hossain, Syed Umar Amin, Mansour Alsulaiman, Muhammad Ghulam |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2018 | An intelligent healthcare system for detection and classification to discriminate vocal fold disorders
Zulfiqar Ali 0001, M. Shamim Hossain, Muhammad Ghulam, Arun Kumar Sangaiah |
Future Gener. Comput. Syst. | 3 |
| 2018 | Edge-centric multimodal authentication system using encrypted biometric templates
Zulfiqar Ali 0001, M. Shamim Hossain, Muhammad Ghulam, Ihsan Ullah 0002, Hamid R. Abachi, Atif Alamri |
Future Gener. Comput. Syst. | 3 |
| 2018 | Iris Recognition Using Multi-Algorithmic Approaches for Cognitive Internet of things (CIoT) Framework
Ramadan Gad, Muhammad Talha 0001, Ahmed A. Abd El-Latif 0001, Mohamed Zorkany, Ayman El-Sayed, Nawal A. El-Fishawy, Muhammad Ghulam |
Future Gener. Comput. Syst. | 7 |
| 2018 | Collaborative analysis model for trending images on social networks
M. Shamim Hossain, Mohammed F. Alhamid, Muhammad Ghulam |
Future Gener. Comput. Syst. | 3 |
| 2018 | Improving consumer satisfaction in smart cities using edge computing and caching: A case study of date fruits classification
M. Shamim Hossain, Muhammad Ghulam, Syed Umar Amin |
Future Gener. Comput. Syst. | 2 |
| 2018 | Cloud-assisted secure video transmission and sharing framework for smart cities
M. Shamim Hossain, Muhammad Ghulam, Wadood Abdul, Biao Song, Brij B. Gupta |
Future Gener. Comput. Syst. | 2 |
| 2018 | Emotion-Aware Connected Healthcare Big Data Towards 5GabstractThe recent development of big data-oriented wireless technologies in terms of emerging 5G, edge computing, interconnected devices of the Internet of Things (IoT), and data analytics, as well as techniques, have enabled connected healthcare services for a happier and healthier life. Although, the quality of the healthcare services can be enhanced through big data-oriented wireless technologies, however, the challenges remain for not considering emotional care, especially for children, elderly, and mentally ill people. In this paper, we propose an emotion-aware connected healthcare system using a powerful emotion detection module. Different IoT devices are used to capture speech and image signals of a patient in a smart home scenario. These signals are used as the input to the emotion detection module. Speech and image signals are processed separately, and classification scores using these signals are fused to produce a final score to take a decision about the emotion. If the emotion is detected as pain, caregivers can visit the patient. Several experiments were performed to validate the proposed system, and good accuracies, up to 99.87%, were achieved for emotion detection. The proposed framework would greatly contribute personalized and seamless emotion-aware healthcare services toward 5G. M. Shamim Hossain, Muhammad Ghulam |
IEEE Internet Things J. | 2 |
| 2018 | Reliable service delivery in Tele-health care systems
Majdi Rawashdeh, Mohammed G. H. al Zamil, M. Shamim Hossain, Samer Samarah, Syed Umar Amin, Muhammad Ghulam |
J. Netw. Comput. Appl. | 6 |
| 2018 | Transferring activity recognition models in FOG computing architecture
Samer Samarah, Mohammed G. H. al Zamil, Majdi Rawashdeh, M. Shamim Hossain, Muhammad Ghulam, Atif Alamri |
J. Parallel Distributed Comput. | 5 |
| 2018 | Cognitive IoT-Cloud Integration for Smart Healthcare: Case Study for Epileptic Seizure Detection and Monitoring
Musaed Alhussein, Muhammad Ghulam, M. Shamim Hossain, Syed Umar Amin |
Mob. Networks Appl. | 2 |
| 2018 | Verifying the Images Authenticity in Cognitive Internet of Things (CIoT)-Oriented Cyber Physical System
M. Shamim Hossain, Muhammad Ghulam, Muhammad Al-Qurishi |
Mob. Networks Appl. | 2 |
| 2018 | Telesurgery Robot Based on 5G Tactile Internet
Yiming Miao, Limei Peng, M. Shamim Hossain, Muhammad Ghulam |
Mob. Networks Appl. | 5 |
| 2018 | Cloud-oriented emotion feedback-based Exergames framework
M. Shamim Hossain, Muhammad Ghulam, Muhammad Al-Qurishi, Mehedi Masud, Ahmad S. Al-Mogren, Wadood Abdul, Atif Alamri |
Multim. Tools Appl. | 2 |
| 2017 | Cyber-physical cloud-oriented multi-sensory smart home framework for elderly people: An energy efficiency perspective
M. Shamim Hossain, Mohamed Abdur Rahman 0001, Muhammad Ghulam |
J. Parallel Distributed Comput. | 3 |
| 2017 | User emotion recognition from a larger pool of social network data using active learning
Muhammad Ghulam, Mohammed F. Alhamid |
Multim. Tools Appl. | 1 |
| 2017 | Speaker recognition based on Arabic phonemes
Mansour Alsulaiman, Awais Mahmood, Muhammad Ghulam |
Speech Commun. | 3 |
| 2016 | Short-term and long-term memory analysis of learning using 2D and 3D educational contentsabstractThe effect of 2D and 3D educational content learning on memory has been studied using electroencephalography (EEG) brain signal. A hypothesis is set that the 3D materials are better than the 2D materials for learning and memory recall. To test the hypothesis, we proposed a classification system that will predict true or false recall for short-term memory (STM) and long-term memory (LTM) after learning by either 2D or 3D educational contents. For this purpose, EEG brain signals are recorded during learning and testing; the signals are then analysed in the time domain using different types of features in various frequency bands. The features are then fed into a support vector machine (SVM)-based classifier. The experimental results indicate that the learning and memory recall using 2D and 3D contents do not have significant differences for both the STM and the LTM. Muhammad Ghulam, Muhammad Hussain 0001, Muneer H. Al-Hammadi, Hatim A. Aboalsamh, Hassan Mathkour, Amir Malik Saeed |
Behav. Inf. Technol. | 1 |
| 2016 | Cloud-assisted Industrial Internet of Things (IIoT) - Enabled framework for health monitoring
M. Shamim Hossain, Muhammad Ghulam |
Comput. Networks | 2 |
| 2016 | Artificially intelligent recognition of Arabic speaker using voice print-based local featuresabstractLocal features for any pattern recognition system are based on the information extracted locally. In this paper, a local feature extraction technique was developed. This feature was extracted in the time–frequency plain by taking the moving average on the diagonal directions of the time–frequency plane. This feature captured the time–frequency events producing a unique pattern for each speaker that can be viewed as a voice print of the speaker. Hence, we referred to this technique as voice print-based local feature. The proposed feature was compared to other features including mel-frequency cepstral coefficient (MFCC) for speaker recognition using two different databases. One of the databases used in the comparison is a subset of an LDC database that consisted of two short sentences uttered by 182 speakers. The proposed feature attained 98.35% recognition rate compared to 96.7% for MFCC using the LDC subset. Awais Mahmood, Mansour Alsulaiman, Muhammad Ghulam, Sheeraz Akram |
J. Exp. Theor. Artif. Intell. | 3 |
| 2016 | Audio-Visual Emotion Recognition Using Big Data Towards 5G
M. Shamim Hossain, Muhammad Ghulam, Mohammed F. Alhamid, Biao Song, Khalid Al Mutib |
Mob. Networks Appl. | 2 |
| 2016 | STCAPLRS: A Spatial-Temporal Context-Aware Personalized Location Recommendation SystemabstractNewly emerging location-based social media network services (LBSMNS) provide valuable resources to understand users’ behaviors based on their location histories. The location-based behaviors of a user are generally influenced by both user intrinsic interest and the location preference, and moreover are spatial-temporal context dependent. In this article, we propose a spatial-temporal context-aware personalized location recommendation system (STCAPLRS), which offers a particular user a set of location items such as points of interest or venues (e.g., restaurants and shopping malls) within a geospatial range by considering personal interest, local preference, and spatial-temporal context influence. STCAPLRS can make accurate recommendation and facilitate people’s local visiting and new location exploration by exploiting the context information of user behavior, associations between users and location items, and the location and content information of location items. Specifically, STCAPLRS consists of two components: offline modeling and online recommendation. The core module of the offline modeling part is a context-aware regression mixture model that is designed to model the location-based user behaviors in LBSMNS to learn the interest of each individual user, the local preference of each individual location, and the context-aware influence factors. The online recommendation part takes a querying user along with the corresponding querying spatial-temporal context as input and automatically combines the learned interest of the querying user, the local preference of the querying location, and the context-aware influence factor to produce the top- k recommendations. We evaluate the performance of STCAPLRS on two real-world datasets: Dianping and Foursquare. The results demonstrate the superiority of STCAPLRS in recommending location items for users in terms of both effectiveness and efficiency. Moreover, the experimental analysis results also illustrate the excellent interpretability of STCAPLRS. Quan Fang, Changsheng Xu, M. Shamim Hossain, Muhammad Ghulam |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2016 | Big Data-Driven Service Composition Using Parallel Clustered Particle Swarm Optimization in Mobile EnvironmentabstractThe proliferation of mobile computing and smartphone technologies has resulted in an increasing number and range of services from myriad service providers. These mobile service providers support numerous emerging services with differing quality metrics but similar functionality. Facilitating an automated service workflow requires fast selection and composition of services from the services pool. The mobile environment is ambient and dynamic in nature, requiring more efficient techniques to deliver the required service composition promptly to users. Selecting the optimum required services in a minimal time from the numerous sets of dynamic services is a challenge. This work addresses the challenge as an optimization problem. An algorithm is developed by combining particle swarm optimization and k-means clustering. It runs in parallel using MapReduce in the Hadoop platform. By using parallel processing, the optimum service composition is obtained in significantly less time than alternative algorithms. This is essential for handling large amounts of heterogeneous data and services from various sources in the mobile environment. The suitability of this proposed approach for big data-driven service composition is validated through modeling and simulation. M. Shamim Hossain, Mohammad Moniruzzaman, Muhammad Ghulam, Ahmed Ghoneim, Atif Alamri |
IEEE Trans. Serv. Comput. | 3 |
| 2015 | An investigation of MDVP parameters for voice pathology detection on three different databases
Ahmed Y. Al-nasheri, Zulfiqar Ali 0001, Muhammad Ghulam, Mansour Alsulaiman |
INTERSPEECH | 3 |
| 2015 | Date fruits classification using texture descriptors and shape-size features
Muhammad Ghulam |
Eng. Appl. Artif. Intell. | 1 |
| 2015 | Cloud-Assisted Speech and Face Recognition Framework for Health Monitoring
M. Shamim Hossain, Muhammad Ghulam |
Mob. Networks Appl. | 2 |
| 2015 | Spectro-temporal directional derivative based automatic speech recognition for a serious game scenario
Muhammad Ghulam, Mehedi Masud, Abdulhameed Alelaiwi, Mohamed Abdur Rahman 0001, Ali Karime, Atif Alamri, M. Shamim Hossain |
Multim. Tools Appl. | 1 |
| 2015 | Audio-Visual Emotion-Aware Cloud Gaming FrameworkabstractThe promising potential and emerging applications of cloud gaming have drawn increasing interest from academia, industry, and the general public. However, providing a high-quality gaming experience in the cloud gaming framework is a challenging task because of the tradeoff between resource consumption and player emotion, which is affected by the game screen. We tackle this problem by leveraging emotion-aware screen effects in the cloud gaming framework and combining them with remote display technology. The first stage in the framework is the learning or training stage, which establishes a relationship between screen features and emotions using Gaussian mixture model-based classifiers. In the operating stage, a linear programming model provides appropriate screen changes based on the real-time user emotion obtained in the first stage. Our experiments demonstrate the effectiveness of the proposed framework. The results show that our proposed framework can provide a high quality gaming experience while generating an acceptable amount of workload for the cloud server in terms of resource consumption. M. Shamim Hossain, Muhammad Ghulam, Biao Song, Mohammad Mehedi Hassan, Abdulhameed Alelaiwi, Atif Alamri |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Word-of-Mouth Understanding: Entity-Centric Multimodal Aspect-Opinion Mining in Social MediaabstractMost existing approaches on aspect-opinion mining focus on the text domain and cannot be applied to social media where the aspects are essentially multimodal and the opinions depend on the specific aspects. To address the problem of multimodal aspect-opinion mining for entities by leveraging multiple cross-collection sources in social media, in this paper we propose a multimodal aspect-opinion model (mmAOM) considering both user-generated photos and textual documents to simultaneously capture correlations between textual and visual modalities, as well as associations between aspects and opinions . By identifying the aspects and the corresponding opinions related to entities, we apply the mmAOM to entity association visualization and multimodal aspect-opinion retrieval. We have conducted extensive experiments on real-world datasets of entities including Flickr photos, Tripadvisor reviews, and news articles. Qualitative and quantitative evaluation results have validated the effectiveness of the multimodal aspect-opinion mining model, and demonstrated the utility of the derived aspects and opinions from mmAOM in applications of entity association visualization and aspect-opinion retrieval. Quan Fang, Changsheng Xu, Jitao Sang 0001, M. Shamim Hossain, Muhammad Ghulam |
IEEE Trans. Multim. | 5 |
| 2014 | Voice pathology detection using auto-correlation of different filters bankabstractThis paper investigates the contribution of frequency bands for automatic voice pathology detection. First, the input voice signal is passed through a number of time-domain band-pass filters. The center frequencies are spaced on an octave scale. Each filter output is then divided into overlapping frames. Auto-correlation function is applied to each block to find the first largest peak, in areas other than near the dc value, and its corresponding lag. Therefore, each frame is having only these two features (peak value and lag). As classifier, we use Gaussian mixture models (GMM) and support vector machine (SVM), separately. Two well-known available databases, one in English (MEEI) and the other one in German (SVD), are used in the investigation. The results demonstrate that the most significant frequency range to detect voice pathology is between 1500 Hz and 3500 Hz. Using this filter band and with only two features, the accuracy is above 97% in case of the MEEI database. Ahmed Y. Al-nasheri, Zulfiqar Ali 0001, Muhammad Ghulam, Mansour Alsulaiman |
AICCSA | 3 |
| 2014 | Automatic pronunciation error detection of nonnative Arabic SpeechabstractComputer assisted language learning (CALL) and, more specifically, computer assisted pronunciation training (CAPT) have received considerable attention in recent years. CAPT allows continuous feedback to the learner without requiring the sole attention of the teacher; it facilitates self study and encourages interactive use of the language in preference to rote learning. One of the important processes in CAPT system is error detection, which locates the errors in the utterance. Although Arabic is currently one of the most widely spoken languages in the world, there has been relatively little research about detection of the pronunciation error by nonnative speakers compared to the other languages. This research is concerned with detecting pronunciation errors of nonnative Arabic speakers from Pakistan and India. All the sounds in this study were taken from King Saud University (KSU) Arabic Speech Database. By analyzing the speech of the Pakistani and Indian speakers in KSU database we found that five phonemes were often mispronounced by nonnative speakers, hence this research will concentrate on pronunciation errors in these five phonemes. The system was built with native and nonnative speakers, and tested with nonnative only. For each phoneme, the Goodness of Pronunciation (GOP) was calculated and compared with a threshold to decide if the phoneme was pronounced correctly or not. The result showed that GOP gave high accuracy, where the scoring accuracy was very good to excellent from 87% to 100%, and the false rejection was zero to less than 10%. This machine judgment is compared with human judgment and the comparison shows excellent agreement between them. Afnan Al Hindi, Mansour Alsulaiman, Muhammad Ghulam, Saad Alkahtani |
AICCSA | 3 |
| 2014 | Detection and classification of voice pathology using feature selectionabstractThe aim of this study is to apply automatic speech recognition (ASR) mechanism to improve the amount of information extracted from the voice and to increase the accuracy of the system by using selective highly discriminative features among different types of acoustic features. For feature extraction, we applied three techniques which are Mel Frequency Cepstral Coefficient (MFCC), Linear Prediction Cepstral Coefficients (LPCC), and RelAtive SpecTrA - Perceptual Linear Predictive (RASTA-PLP) with a number of selected coefficients from each technique by using t-test, Kruskal-Wallis test, or genetic algorithm (GA). Then for classification, either support vector machine (SVM) or Gaussian Mixture Model (GMM) is used. The experimental results on a selected MEEI subset database show that the proposed method gives high accuracies compared with some recent related methods both in detection and classification tasks. The highest accuracy of 99.9875 % with a standard deviation of 0.0263 is achieved in case of detection, and 99.8578 % with a standard deviation of 0.1657 in case of multi-class pathology classification. Malak Al Mojaly, Muhammad Ghulam, Mansour Alsulaiman |
AICCSA | 2 |
| 2014 | Comparison between WLD and LBP descriptors for non-intrusive image forgery detectionabstractDue to the availability of easy-to-use and powerful image editing tools, the authentication of digital images cannot be taken for granted and it gives rise to non-intrusive forgery detection problem because all imaging devices do not embed watermark. We investigated the detection of copy-move and splicing, the two harmful types of image forgery, using textural properties of images. Tampering distorts the texture micro-patterns in an image and texture descriptors can be employed to detect tampering. We did comparative study to examine the effect of two state-of-the-art best texture descriptors: Multiscale Local Binary Pattern (Multi-LBP) and Multiscale Weber Law Descriptor (Multi-WLD). Multiscale texture descriptors extracted from the chrominance components of an image are passed to Support Vector Machine (SVM) to identify it as authentic or forged. The performance comparison reveals that Multi-WLD performs better than Multi-LBP in detecting copy-move and splicing forgeries. Multi-WLD also outperforms state-of-the-art passive forgery detection techniques. Muhammad Hussain 0001, Sahar Q. Saleh, Hatim A. Aboalsamh, Muhammad Ghulam, George Bebis |
INISTA | 4 |
| 2014 | Accurate and robust localization of duplicated region in copy-move image forgery
Maryam Jaberi, George Bebis, Muhammad Hussain 0001, Muhammad Ghulam |
Mach. Vis. Appl. | 4 |
| 2014 | Image forgery detection using steerable pyramid transform and local binary pattern
Muhammad Ghulam, Muneer H. Al-Hammadi, Muhammad Hussain 0001, George Bebis |
Mach. Vis. Appl. | 1 |
| 2013 | Voice pathology detection and classification using MPEG-7 audio low-level features
Muhammad Ghulam, Moutasem Melhem |
INTERSPEECH | 1 |
| 2012 | Race recognition using local descriptorsabstractThis paper proposes a method for race recognition from face images using local descriptors. The proposed method uses two types of local descriptors: local binary pattern (LBP) and Weber local descriptors (WLD). First, LBP and WLD histograms are obtained separately from blocks of normalized face image. Kruskal-Wallis feature selection technique is applied to the histograms to select the significant bins for race recognition. Then the selected bins from the two histograms are concatenated block by block to produce the final feature set of the face image. Minimum city block distance is used as a classifier. The experiments are conducted using gray scale FERET images with five race groups. Experimental results show that the proposed method has superior race recognition accuracies for all the five race groups compared to LBP and WLD alone. Muhammad Ghulam, Muhammad Hussain 0001, Fatmah Alenezy, Anwar M. Mirza, George Bebis, Hatim A. Aboalsamh |
ICASSP | 1 |
| 2012 | Polynomial Correlation Filters for Human Face RecognitionabstractThis paper describes a nonlinear face recognition method based on polynomial spatial frequency image processing. This nonlinear method is known as the polynomial distance classifier correlation filter (PDCCF). PDCCF is a member of a well-known family of filters called correlation filters. Correlation filters are attractive because of their shift invariance and potential for distortion tolerant pattern recognition. PDCCF addresses more than one filter in the system, each one with a different form of non-linearity. Our experimental results on the Olivetti Research Laboratory (ORL) and Extended Yale B (EYB) face datasets show that PDCCF outperforms the principal component analysis (PCA), and the local binary pattern (LBP). Mohamed I. Alkanhal, Muhammad Ghulam |
ICMLA (1) | 2 |
| 2011 | Automatic voice disorder classification using vowel formantsabstractIn this paper, we propose an automatic voice disorder classification system using first two formants of vowels. Five types of voice disorder, namely, cyst, GERD, paralysis, polyp and sulcus, are used in the experiments. Spoken Arabic digits from the voice disordered people are recorded for input. First formant and second formant are extracted from the vowels [Fatha] and [Kasra], which are present in Arabic digits. These four features are then used to classify the voice disorder using two types of classification methods: vector quantization (VQ) and neural networks. In the experiments, neural network performs better than VQ. For female and male speakers, the classification rates are 67.86% and 52.5%, respectively, using neural networks. The best classification rate, which is 78.72%, is obtained for female sulcus disorder. Muhammad Ghulam, Mansour Alsulaiman, Awais Mahmood, Zulfiqar Ali 0001 |
ICME | 1 |
| 2010 | DPF-based japanese phoneme recognition using tandem MLNsabstractThis paper presents a method for automatic phoneme recognition for Japanese language using tandem MLNs. The method comprises three stages: (i) multilayer neural network (MLN) that converts acoustic features into distinctive phonetic features DPFs, (ii) MLN that combines DPFs and acoustic features as input and generates a 45 dimensional DPF vector with less context effect and (iii) the 45 dimensional feature vector generated by the second MLN are inserted into a hidden Markov model (HMM) based classifier to obtain more accurate phoneme strings from the input speech. From the experiments on Japanese Newspaper Article Sentences (JNAS), it is observed that the proposed method provides a higher phoneme correct rate and improves phoneme accuracy tremendously over the method based on a single MLN. Moreover, it requires fewer mixture components in HMMs. Mohammed Rokibul Alam Kotwal, Manoj Banik, Gazi Md. Moshfiqul Islam, M. Shahadat Hossain, Foyzul Hassan, Mohammad Mahedi Hasan, Muhammad Ghulam, Mohammad Nurul Huda |
HIS | 7 |
| 2010 | Study on pharyngeal and uvular consonants in foreign accented Arabic for ASR
Yousef Ajami Alotaibi, Muhammad Ghulam |
Comput. Speech Lang. | 2 |
| 2008 | Study on unique pharyngeal and uvular consonants in foreign accented ArabicabstractThis paper investigates the unique pharyngeal and uvular consonants of Arabic from the automatic speech recognition (ASR) point of view. Comparisons of the recognition error rates for these phonemes are analyzed in five experiments that involve different combinations of native and non-native Arabic speakers. The most three confusing consonants for every investigated consonant are uncovered and discussed. Results confirm that these Arabic distinct consonants are a major source of difficulty for ASR. While the recognition rate for certain of these unique consonants such as /H / can drop below 35 % when uttered by non-native speakers, there are advantages to including non-native speakers in ASR. Regional differences in the pronunciation of Modern Standard Arabic by native Arabic speakers require attention of Arabic ASR research. Yousef Ajami Alotaibi, Khondaker Abdullah Al Mamun, Muhammad Ghulam |
INTERSPEECH | 3 |
| 2007 | Distinctive phonetic feature (DPF) based phone segmentation using hybrid neural networksabstractSegmentation of speech into its corresponding phones has become very important issue in many speech processing areas such as speech recognition, speech analysis, speech synthesis, and speech database. In this paper, for accurate segmentation in speech recognition applications, we introduce Distinctive Phonetic Feature (DPF) based feature extraction using a twostage NN (Neural Networks) system consists of a RNN (Recurrent Neural Network) in the first stage and an MLN (Multi-Layer Neural Network) in the second stage. The RNN maps continuous acoustic features, Local Feature (LF), onto discrete DPF patterns, while the MLN constraints DPF context or dynamics in an utterance. The experiments are carried out using JNAS (Japanese Newspaper Article Sentences) continuous utterances that contains vowels and consonants. The proposed DPF based feature extractor provides good segmentation and high recognition rate with a reduced mixture-set of HMMs (Hidden Markov Models) by resolving co-articulation effect. Mohammad Nurul Huda, Muhammad Ghulam, Junsei Horikawa, Tsuneo Nitta |
INTERSPEECH | 2 |
| 2006 | A Pitch-Synchronous Peak-Amplitude Based Feature Extraction Method for Noise Robust ASRabstractIn this paper, we propose a novel pitch-synchronous auditory-based feature extraction method for robust automatic speech recognition (ASR). A pitch-synchronous zero-crossing peak-amplitude (PS-ZCPA)-based feature extraction method was proposed previously, and showed improved performance except while modulation enhancement was integrated together with Wiener filter (WF)-based noise reduction and auditory masking into it. However, since zero-crossing is not an auditory event, we propose a new pitch-synchronous peak-amplitude (PS-PA)-based method to make a feature extractor of ASR more auditory-like. We also examine the effect of WF-based noise reduction, modulation enhancement, and auditory masking into the proposed PS-PA method using Aurora-2J database. The experimental results showed the superiority of the proposed method over the PS-ZCPA method, and eliminated the problem due to the reconstruction of zero-crossings from modulated envelope. The highest relative performance over MFCC was achieved as 67.33% using the PS-PA method together with WF-based noise reduction, modulation enhancement, and auditory masking Muhammad Ghulam, Junsei Horikawa, Tsuneo Nitta |
ICASSP (1) | 1 |
| 2006 | Comparative study on contributions of pitch-synchronization and peak-amplitude towards robustness issue of ASRabstractWe proposed previously a novel pitch-synchronous peakamplitude (PS-PA) based feature extraction method, which achieved significant recognition accuracy for robust ASR [1]. It is well-known that an auditory neuron has pitch detection mechanism that can be useful for speech detection, and also peak-amplitudes in temporal pattern are robust to noise. In this paper, we conduct several experiments to find out relative contributions of pitch-synchronization (PS) and peakamplitudes (PA) on recognition accuracy of robust ASR. Experiments include methods with fixed and pitchsynchronous frame lengths, and that with traditional peakamplitudes and pitch-synchronous peak-amplitudes. The experimental results show that both PS and PA have strong contributions towards robust ASR and the effect of PS is higher than that of PA. Index Terms: speech recognition, feature extraction, pitch synchronization, peak amplitude Muhammad Ghulam, Junsei Horikawa, Tsuneo Nitta |
INTERSPEECH | 1 |
| 2005 | Pitch-Synchronous ZCPA (PS-ZCPA)-Based Feature Extraction with Auditory MaskingabstractA pitch-synchronous (PS) auditory feature extraction method, based on ZCPA (zero-crossings peak-amplitudes), has been proposed (Ghulam, M. et al., Proc. ICSLP04, 2004) and was shown to be more robust than the conventional ZCPA (Kim, D.S. et al., IEEE Trans. Speech Audio Process., vol.7, no.1, p.55-69, 1999). We examine the effect of auditory masking, both simultaneous and temporal, in the PS-ZCPA method. We also observe the effect of varying the number of histogram bins on the way to find out the optimum parameters of the proposed method. Experimental results demonstrate the improved performance of the PS-ZCPA method achieved by embedding auditory masking into it; for example, with both the masking methods embedded, the performance increases to 73.71% from the 69.92% obtained without masking for PS-ZCPA, while it showed little improvement with an increased number of histogram bins. Muhammad Ghulam, Takashi Fukuda, Junsei Horikawa, Tsuneo Nitta |
ICASSP (1) | 1 |
| 2005 | Designing multiple distinctive phonetic feature extractors for canonicalization by using clustering techniqueabstractAcoustic models of an HMM-based classifier include various types of hidden factors such as speaker-specific characteristics and acoustic environments. If there exist a canonicalization process that represses the decrease of differences in acoustic-likelihood among categories resulted from hidden factors, a robust ASR system can be realized. We have previously proposed the canonicalization process of featureparameters composed of three distinctive phonetic feature (DPF) extractors focused on a gender factor. This paper describes an attempt to design multiple DPF extractors corresponding to unspecific hidden factors, as well as to introduce a noise suppressor that is targeted for the canonicalization of a noise factor. In an experiment on Japanese version AURORA2 database (AURORA2-J), the proposed system achieved significant improvements when combining the canonicalization process with the noise reduction technique based on a two-stage Wiener filter. Takashi Fukuda, Muhammad Ghulam, Tsuneo Nitta |
INTERSPEECH | 2 |
| 2004 | A noise-robust feature extraction method based on pitch-synchronous ZCPA for ASRabstractIn this paper, we propose a novel feature extraction method based on an auditory nervous system for robust automatic speech recognition (ASR). In the proposed method, a pitchsynchronous mechanism is embedded in ZCPA (ZeroCrossings Peak-Amplitudes), which has previously been shown to outperform the conventional features in the presence of noise. A noise-robust non-delayed pitch determination algorithm (PDA) is also developed. In the experiment, the proposed pitch-synchronous ZCPA (PS-ZCPA) was proved more robust than the original ZCPA method. Moreover, a simple noise subtraction (NS) method is also integrated in the proposed method and the performance was evaluated using the Aurora-2J database. The experimental results showed the superiority of the proposed PS-ZCPA method with NS over the PS-ZCPA method without NS. Muhammad Ghulam, Takashi Fukuda, Junsei Horikawa, Tsuneo Nitta |
INTERSPEECH | 1 |
| 2003 | Voice quality normalization in an utterance for robust ASRabstractIn this paper, we propose a novel method of normalizing the voice quality in an utterance for both clean speech and speech contaminated by noise. The normalization method is applied to the N-best hypotheses from an HMM-based classifier, then an SM (Sub-space Method)-based verifier tests the hypotheses after normalizing the monophone scores together with the HMMbased likelihood score. The HMM-SM-based speech recognition system was proposed previously [1, 2] and successfully implemented on a speaker-independent word recognition task and an OOV word rejection task. We extend the proposed system to a connected digit string recognition task by exploring the effect of the voice quality normalization in an utterance for robust ASR and compare it with the HMM-based recognition systems with utterance-level normalization, word-level normalization, monophone-level normalization, and state-level normalization. Experimental results performed on connected 4digit strings showed that the word accuracy was significantly improved from 95.7% obtained by the typical HMM-based system with utterance-level normalization to 98.2% obtained by the HMM-SM-based system for clean speech, from 88.1% to 91.5% for noise-added speech with SNR=10dB, and from 72.4% to 76.4% for noise-added speech with SNR=5dB, while the other HMM-based systems also showed lower performances. Muhammad Ghulam, Takashi Fukuda, Tsuneo Nitta |
INTERSPEECH | 1 |
| 2002 | Confidence scoring for accurate HMM-based word recognition by using SM-based monophone score normalizationabstractIn this paper, we propose a novel confidence scoring method that is applied to N-best hypotheses output from an HMM-based classifier. In the first pass of the proposed method, the HMM-based classifier with monophone models outputs N-best hypotheses and boundaries of all the monophones in the hypotheses. In the second pass, an SM(sub-space method)-based verifier tests the hypotheses by comparing confidence scores. We discuss how to convert a monophone similarity score of SM into a likelihood score, how to normalize the variations of acoustic quality in an utterance, and how to combine an HMM-based likelihood of word level and an SM-based likelihood of monophone level. In the experiments performed on speaker-independent word recognition, the proposed confidence scoring method significantly improves correct word recognition rate from 95.3% obtained by the standard HMM classifier to 98.0%. Takaharu Sato, Muhammad Ghulam, Takashi Fukuda, Tsuneo Nitta |
ICASSP | 2 |
| 2002 | Improving performance of an HMM-based ASR system by using monophone-level normalized confidence measureabstractIn this paper, we propose a novel confidence scoring method that is applied to N-best hypotheses output from an HMM-based classifier. In the first pass of the proposed method, the HMM-based classifier with monophone models outputs N-best hypotheses (word candidates) and boundaries of all the monophones in the hypotheses. In the second pass, an SM (Sub-space Method)-based verifier tests the hypotheses by comparing confidence scores. We discuss how to convert a monophone similarity score of SM into a likelihood score, how to normalize the variations of acoustic quality in an utterance, how to combine an HMM-based likelihood of word level and an SM-based likelihood of monophone level, and also how to accept the correct words and reject OOV words. In the experiments performed on speaker-independent word recognition, the proposed confidence scoring method significantly reduced word error rate from 4.7% obtained by the standard HMM classifier to 2.0%, and it also reduced the equal error rate from 9.0% to 6.5% in an unknown word rejection task. Muhammad Ghulam, Takashi Fukuda, Takaharu Sato, Tsuneo Nitta |
INTERSPEECH | 1 |