Amir Hussain 0001

dblp:03/6772 · DBLP profile ↗
← Back
216ranked-venue papers
16as first author
117since 2021 · last 2026
0000-0002-8080-082XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 169 · 12 first-author · 81 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 13 since 2021Databases, data management, data science and information retrieval · 9 · 4 since 2021Computer networks · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 TADynFed: Dynamic modality-adaptive federated learning with tissue-aware disentanglement for cross-disease analysis
Saeed Iqbal, Xiaopin Zhong, Muhammad Attique Khan, Zongze Wu 0001, Nouf Almujally, Weixiang Liu, Amir Hussain 0001
Artif. Intell. Medicine7
2026 A novel Quantum Beta distributed multi-objective Particle Swarm Optimization algorithm for fake accounts detection
abstract
Detecting fake accounts on Online Social Networks is a pressing issue due to the rise in unethical online activities. This study presents a new Quantum Beta-behaved Multi-Objective Particle Swarm Optimization Algorithm (QB-MOPSO) for machine learning-based fake account detection. QB-MOPSO aims to enhance the learning process of a random forest algorithm by simultaneously minimizing feature dimensionality and classification error rates. It proposes a novel architecture that employs two optimization profiles: one improves exploratory behavior using a quantum-behaved equation, while the other enhances exploitation through a beta function. The main contributions of this study are as follows: the design of a novel Quantum Beta Distributed Multi-Objective Particle Swarm Optimization algorithm that integrates quantum-behaved exploration and beta-distributed exploitation, the application of this algorithm to enhance artificial intelligence–based fake account detection on Twitter datasets, and a comprehensive experimental evaluation demonstrating superior accuracy, F-measure, and MCC compared to existing methods. Experimental results on two Twitter datasets with 1982 and 928 accounts respectively show QB-MOPSO's effectiveness, achieving accuracy rates of about 99.19 % and 97.52 %. Comparisons with the original architecture demonstrate QB-MOPSO's ability to enhance the performance of the random forest algorithm.
Ahlem Aboud, Nizar Rokbani, Seyedali Mirjalili, Amir Hussain 0001, Adel M. Alimi
Eng. Appl. Artif. Intell.4
2026 A Novel Multistage Attention-Enhanced Mixture of Experts Model for Alzheimer's Disease Diagnosis
abstract
ABSTRACT Alzheimer's Disease (AD) is a progressive neurodegenerative disease diagnosed through cognitive impairment, and an early diagnosis is essential to improve treatment and care options. Current diagnostic approaches of AD, such as neuroimaging, cognitive assessments and biomarker research, are lengthy, vague and not sufficient to assess the early stages of AD. To address these problems, we introduced a novel deep learning model, ‘NeuroMixFormer’, which is based on a mixture‐of‐experts architecture for AD classification from MRI. The proposed multistage architecture employs a dynamic routing mechanism and four expert blocks per stage, each integrating dense connectivity with a spatial and channel attention module for feature extraction. To improve early feature learning, auxiliary classifiers are incorporated at intermediate stages of training. Evaluations on three datasets (ADNI, Mendeley and Kaggle Augmented Alzheimer's MRI) demonstrated the proposed model's superior performance over existing deep learning architectures and state‐of‐the‐art methods, achieving up to 99.48% accuracy on the Kaggle dataset, 90.28% on the Mendeley dataset and 99.86% on the ADNI dataset, respectively. Ablation studies confirmed the importance of dual‐attention mechanisms, and expert routing analysis showed clear specialisation patterns across AD stages, improving both classification accuracy and interpretability. These results underscore the effectiveness and generalisability of NeuroMixFormer in automated dementia detection, highlighting its potential to support early and precise AD diagnosis. However, the high computational cost and inference time associated with this high accuracy limit the practicality of the proposed approach in clinical settings.
Muhammad John Abbas, Muhammad Attique Khan, Veena Dillshad, Ahmed Ibrahim Alzahrani 0001, Nasser Alalwan, Ali Alamer, Yunyoung Nam, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.8
2026 Mamba- SBRNet : Real-Time Lightweight Student Behaviour Object Detection Model
abstract
ABSTRACT Detecting student behaviour objects in classroom environments is crucial for assessing educational progress, optimizing teaching strategies and improving student learning outcomes. With the ongoing advancement of educational informatization, analysing classroom behaviour has become an important tool for enhancing teaching quality and personalized learning. However, current student behaviour object detection models based on CNN and Transformer architectures face challenges such as large parameter sizes and high inference delays when deployed on edge devices in classrooms, limiting their practical application. To address these issues, this study proposes a lightweight student behaviour detection framework based on the Mamba architecture, aimed at balancing computational efficiency and detection accuracy. First, the framework based on the state‐space model (SSM) efficiently captures global dependencies, using local convolutions to enhance detection accuracy and scene understanding while maintaining real‐time performance. Second, the C2CGA module increases attention diversity through feature splitting, self‐attention, cascading and projected concatenation, deepening the network while reducing computational overhead. Finally, the A2CMoCA module aggregates multi‐scale features, improving the learning of small objects and occluded behaviours. Experiments on a self‐built classroom behaviour dataset (containing eight typical teaching behaviours) show that the proposed method achieves 91.5% detection accuracy while maintaining a lightweight design. Compared to the baseline model, its computational efficiency (5.9G FLOPs) is reduced by 56.6%, the parameter size is compressed to 3.65 M (a 39% reduction) and the inference speed is 3.2 ms, meeting the real‐time monitoring requirements in classroom teaching scenarios.
Le Zou, Yuanhang Xia, Fengling Jiang, Yimin Wu, Kia Dashtipour, Mandar Gogate, Amir Hussain 0001, Xiaofeng Wang 0009
Expert Syst. J. Knowl. Eng.8
2026 Hierarchical federated learning with paillier encryption: synergistic approach for secure analytics of sensitive healthcare data
Saeed Iqbal, Xiaopin Zhong, Muhammad Attique Khan, Zongze Wu 0001, Nouf Almujally, Weixiang Liu, Faisal Albalwy, Amir Hussain 0001
Expert Syst. Appl.8
2026 xMagNet: Dynamic magnification-aware fusion with uncertainty quantification for robust breast cancer histopathology
abstract
Histopathology image analysis faces challenges due to magnification variability, limiting robust tumor categorization. Existing deep learning models prioritize accuracy but neglect explainability, ethical biases, and real-world deployment. This study proposes xMagNet, a hybrid Transformer-Convolutional Neural Network (CNN) framework that synergizes technical rigor, clinical transparency, and ethical fairness for multi-magnification breast cancer diagnostics. xMagNet integrates a hybrid encoder combining Vision Transformers (ViT) for global tissue modeling at low magnifications (4x - 10x) and Separable Dilation Convolutions (SDC) for localized nuclear texture extraction at high magnifications (20x - 40x). Magnification-Aware Gating (MAG) dynamically balances ViT and SDC features via temperature-scaled sigmoid activation. A multi-task decoder employs Thresholded Grad-CAM (top 10% gradients) for explainable decision-making and Point-wise Reformation Blocks (PRB) for boundary preservation. Federated learning (FL) with momentum-enhanced aggregation and Sinkhorn divergence regularization ensures scanner/stain-invariant training across six institutions (Hamamatsu/Leica, H&E/IHC). Uncertainty-quantified predictions (Monte Carlo dropout) and adversarial debiasing mitigate demographic leakage. xMagNet achieves 97.8% F1-score for tumor segmentation on Camelyon16 and 93% Gleason AUC on PANDA, with 96.5% pathologist concordance via Grad-CAM. At 40x magnification, it detects micro-metastases with 94% sensitivity (vs. UNet++’s 89% and ResUNet’s 91%). Computational efficiency includes sub-second inference (0.42 sec/slide) and 2.3 x faster convergence than HoVer-Net. Ethical auditing reveals 3% fairness gaps ( ) and 73% domain shift reduction (MMD: 0.12 vs. FedAvg’s 0.45), validated on 15,000 whole-slide images (WSIs) from TCGA-BRCA, Camelyon16, and PANDA datasets. xMagNet bridges critical gaps in multi-magnification histopathology by harmonizing technical robustness (MAG fusion, bounded gradients) with clinical utility (HER2+/ER+ subtyping, Gleason grading) and ethical scalability. By achieving high accuracy, rapid inference, and equitable deployment, it advances AI-driven diagnostics toward trustworthy, deployable systems for breast, prostate, and metastatic cancer imaging. Code available at: xMagNet .
Saeed Iqbal, Muhammad Attique Khan, Leila Jamel Menzli, Adnan N. Qureshi, Imran Arshad Choudhry, Amir Hussain 0001
Neurocomputing6
2026 FedCapD: Federated class-incremental learning via capsule distillation and diffusion replay
abstract
Federated Class-Incremental Learning faces critical challenges including catastrophic forgetting, semantic drift, and privacy risks under non-IID data distributions. To address these, we propose FedCapD, a novel framework that unifies unsupervised task boundary detection via Bayesian nonparametric modeling, hierarchical semantic distillation through capsule alignment and GNN-based class propagation, secure gradient communication using homomorphic encryption and differential privacy, and diffusion-based generative replay for memory-efficient adaptation. By integrating structural reasoning with privacy-preserving learning, FedCapD achieves state-of-the-art performance across four diverse datasets-CheXpert, MIMIC-CXR-JPG, BraTS2021, and PHM2012-in terms of semantic consistency, encrypted distillation fidelity, cold-start accuracy, and utility efficiency under privacy constraints. The framework eliminates reliance on real-data storage and supports scalable, lifelong learning in regulated, resource-constrained environments such as healthcare AI and edge computing. Code is available at FCIL .
Saeed Iqbal, Muhammad Attique Khan, Sarra Ayouni, Abeer Aljohani, Amir Hussain 0001
Neurocomputing6
2026 Cross-modal invariant learning with latent diffusion for reliable medical diagnosis under dynamic shifts
abstract
Robust and reliable medical diagnosis using artificial intelligence is crucial, yet real-world clinical environments present significant challenges due to dynamic covariate shifts affecting multi-modal data (images, text, tabular). Existing methods, including the single-modal robust classifier LaDiNE, often fail under these complex, multi-modal shifts, lacking mechanisms for cross-modal invariance, dynamic modality fusion, and fine-grained uncertainty attribution. To address this gap, we propose DyMoLaDiNE (Dynamic Multi-Modal Latent Diffusion Nested-Ensembles), a framework designed for reliable medical diagnosis under dynamic multi-modal covariate shifts. DyMoLaDiNE introduces four key innovations: (1) a Cross-Modal Invariant Feature Extractor leveraging multi-modal Vision Transformers and contrastive learning to derive robust latent representations, (2) a Dynamic Modality Weighting Mechanism that adaptively adjusts modality contributions based on instance-specific reliability scores, (3) a Robust Multi-Modal Diffusion Ensemble utilizing conditional diffusion models conditioned on multi-modal inputs and reliability scores for flexible, calibrated density estimation, and (4) Modality-Attributed Uncertainty Quantification to decompose predictive uncertainty by input source. Extensive evaluations on diverse datasets (MedMD&RadMD, MultiCaRe, PadChest, TCIA RE-MIND, BRaTS, Camelyon16, PANDA) demonstrate that DyMoLaDiNE significantly outperforms (p 0.005) state-of-the-art methods (LDM, CMCL, CGMCL, CIIM, DTTL, FFL, ALDM, LaDiNE) in terms of classification accuracy, robustness under dynamic perturbations, confidence calibration (ECE), and precise uncertainty quantification (CPIW, CNPV), while providing superior modality attribution fidelity. Ablation studies confirm the necessity of each component. DyMoLaDiNE represents a significant advancement in trustworthy, robust multi-modal medical AI. Code supporting this study DyMoLaDiNE .
Saeed Iqbal, Xiaopin Zhong, Muhammad Attique Khan, Zongze Wu 0001, Nouf Almujally, Weixiang Liu, Amir Hussain 0001, Björn W. Schuller
Neurocomputing7
2026 Dynamic multi-kernel convolutional network with noise injected features for audio-only speech enhancement
Nasir Saleem, Sami Bourouis, Fazal-E. Wahab, Kia Dashtipour, Tughrul Arslan, Amir Hussain 0001
Neurocomputing6
2026 Core unlearning: A multi-modal gradient-efficient architecture for exact and approximate model rewriting
Saeed Iqbal, Xiaopin Zhong, Muhammad Attique Khan, Zongze Wu 0001, Nouf Almujally, Weixiang Liu, Amir Hussain 0001
Inf. Process. Manag.7
2026 Causal continual unlearning with disentangled anomaly representations for private industrial vision
Saeed Iqbal, Xiaopin Zhong, Muhammad Attique Khan, Zongze Wu 0001, Nouf Almujally, Weixiang Liu, Amir Hussain 0001
Inf. Process. Manag.7
2026 Underwater image enhancement with frequency-spatial cross-domain transformer and hybrid collaborative representation
Jing Yang 0041, Wenhan Zhu, Fengling Jiang, Amir Hussain 0001
Knowl. Based Syst.5
2026 Fourier fusion and dual-path attention enhancement network for medical image segmentation
Le Zou, Xiangxu Bu, Zhize Wu, Fengling Jiang, Lingma Sun, Kia Dashtipour, Mandar Gogate, Xiaofeng Wang 0009, Amir Hussain 0001
Multim. Syst.9
2026 An explainable machine learning model for detecting behavioral medication effects in motor subtypes of early Parkinson's disease based on acoustics speech signals
Zeineb Benmessaoud, Sonia BenHassen, Mohamed Neji, Amir Hussain 0001, Nouha Farhat, Emna Smaoui, Mariem Dammek, Mondher Frikha, Adel M. Alimi, Chokri Mhiri
Multim. Tools Appl.4
2026 Multimodal Cognitive Load Estimation With Radio Frequency Sensing and Pupillometry in Complex Auditory Environments
abstract
The detection of listening effort or cognitive load (CL) has been a major research challenge in recent years. Most conventional techniques utilise physiological or audio-visual sensors and are privacy-invasive and computationally complex. The challenges of synchronization, data alignment and accessibility limitations potentially increase the noise and error probability, compromising the accuracy of CL estimates. This innovative work presents a multi-modal, non-invasive and privacy-preserving approach that combines Radio Frequency (RF) and pupillometry sensing to address these challenges. Custom RF sensors are first designed and developed to capture blood flow changes in specific brain regions with high spatial resolution. Next, multi-modal fusion with pupillometry sensing is proposed and shown to offer a robust assessment of cognitive and listening effort through pupil size and pupil dilation. Our novel approach evaluates RF sensing to estimate CL from cerebral blood flow variations utilizing pupillometry as a baseline. A first-of-its-kind, multi-modal dataset is collected as a new benchmark resource in a controlled environment with participants to comprehend target speech with varying background noise levels. The framework is statistically evaluated using intraclass correlation for pupillometry data (average ICC> 0.95). The correlation between pupillometry and RF data is established through Pearson's correlation (average PCC> 0.79). Further, CL is classified into high and low categories based on RF data using K-means clustering. Future work involves integrating RF sensors with glasses to estimate listening effort for hearing-aid users and utilising RF measurements to optimize speech enhancement based on individual's listening effort and complexity of acoustic environment.
Usman Anwar, Adeel Hussain, Mandar Gogate, Kia Dashtipour, Tughrul Arslan, Amir Hussain 0001, Peter Lomax
IEEE J. Biomed. Health Informatics6
2026 DHSNet: Denoised-Modulated Hybrid-Semantic Scale-Aware Network for Low-Light Image Enhancement
abstract
Low-Light Image Enhancement (LLIE) methods based on either Retinex theory or deep learning still exhibit significant shortcomings in handling image corruptions, such as noise, artifacts, and color distortion. The primary issue is that both Retinex algorithms and existing networks may introduce or amplify these corruptions during enhancement. To address these limitations, we propose the Denoised-Modulated Hybrid-Semantic Scale-Aware Network (DHSNet), a novel one-stage LLIE method. DHSNet integrates a Signal-to-Noise Ratio (SNR)-based denoising mechanism and a Hybrid-Semantic Scale-Aware Module (HSM) to preprocess noise and fuse multi-scale features for robust image enhancement. Moreover, we introduce the Illumination Partial Attention Block (IPAB) to further improve illumination correction and nonlinear transformation capabilities. DHSNet effectively mitigates noise, preserves intricate details, and restores degraded structures. Extensive experiments on multiple LLIE datasets demonstrate that it outperforms state-of-the-art (SOTA) methods in both qualitative and quantitative metrics. Furthermore, DHSNet exhibits strong generalization in no-reference LLIE and low-light object detection tasks, underscoring its practical value for real-world applications.
Rentao Yang, Zhize Wu, Xiaofeng Wang 0009, Tong Xu 0001, Fengling Jiang, Amir Hussain 0001, Le Zou
IEEE Trans. Multim.6
2025 From 2D Images to 3D Model: Weakly Supervised Multi-View Face Reconstruction with Deep Fusion
abstract
While weakly supervised multi-view face reconstruction (MVR) is garnering increased attention, one critical issue still remains open: how to effectively interact and fuse multiple image information to reconstruct high-precision 3D models. In this regard, we propose a novel pipeline called Deep Fusion MVR (DF-MVR) to explore the feature correspondences between multi-view images and reconstruct high-precision 3D faces. Specifically, we present a novel multi-view feature fusion backbone that utilizes face masks to align features from multiple encoders and integrates one multi-layer attention mechanism to enhance feature interaction and fusion, resulting in one unified facial representation. Additionally, we develop one concise face mask mechanism that facilitates multi-view feature fusion and facial reconstruction by identifying common areas and guiding the network’s focus on critical facial features (e.g., eyes, brows, nose, and mouth). Experiments on Pixel-Face and Bosphorus datasets indicate the superiority of the proposed method. Without the 3D annotation, DF-MVR achieves relative 5.2% and 3.0% RMSE improvement over the existing weakly supervised MVRs, respectively, on Pixel-Face and Bosphorus datasets. Our code is available at https://github.com/weiguangzhao/DF_MVR.
Weiguang Zhao, Chaolong Yang, Jianan Ye, Rui Zhang 0012, Yuyao Yan, Xi Yang 0008, Bin Dong 0003, Amir Hussain 0001, Kaizhu Huang
ICME8
2025 Are Horses Always Strong and Donkeys Dumb? Animal Bias in Vision Language Models
abstract
Vision Language Models (VLMs), such as CLIP, are widely used for various multimodal tasks and offer significant advancements in image-text understanding. However, existing studies have revealed that VLMs inherit biases from their training data which lead to the reinforcement of harmful stereotypes and cultural misrepresentations. In the proposed work, we analyze the presence of biases associated with animals in the CLIP model. We introduce a novel taxonomy, called Animal Bias Taxonomy (ABT), which categorizes stereotyped associations of animals in three categories. We also curated an animal dataset from existing datasets and applied data-cleaning process on it to remove unwanted images. Using ABT, we evaluated the outputs of VLMs on animal dataset when prompted with animal-related stereotyped terms to assess whether CLIP propagates biased associations that align with cultural stereotypes. Our findings reveal that CLIP frequently exhibits skewed cultural interpretations, such as associating owls with wisdom. Our study underscores the necessity of bias evaluation in VLMs and calls for greater transparency and culturally diverse data curation to ensure fair and inclusive AI systems. The code is available at https://github.com/MohammadAnas5/Clip-sAnimalStereotyping
Mohammad Anas, Mohammad Nadeem, Shahab Saquib Sohail, Erik Cambria, Amir Hussain 0001
IJCNN5
2025 A Study on Speech Assessment with Visual Cues
Shafique Ahmed, Ryandhimas E. Zezario, Nasir Saleem, Amir Hussain 0001, Hsin-Min Wang, Yu Tsao 0001
INTERSPEECH4
2025 Towards Personalised Audio Visual Speech Enhancement
Mandar Gogate, Kia Dashtipour, Amir Hussain 0001
INTERSPEECH3
2025 Investigating Gender Bias in Text-to-Audio Generation Models
Aarish Shah Mohsin, Mohammad Nadeem, Shahab Saquib Sohail, Tughrul Arslan, Mandar Gogate, Nasir Saleem, Amir Hussain 0001
INTERSPEECH7
2025 Wearable RF Sensing System with Edge AI Inference for In-vivo Cognitive Load Classification
abstract
Cognitive load (CL) refers to the mental effort required to process information. Monitoring increased CL is important as it can indicate the onset of cognitive decline and neurodegeneration, allowing for early intervention and management. Traditional techniques using bio-physiological and audio-visual sensors are privacy-invasive and computationally complex. The synchronization, data alignment, and accessibility problems with these techniques can lead to increased noise and errors, reducing the accuracy of CL estimates. This paper presents a first-of-its-kind Radio Frequency (RF) based sensing system that effectively monitors and classifies cognitive load states by detecting cerebral blood flow variations through backscattered RF signal strength. The RF sensors are designed and miniaturized using microwave computational software and fabricated sensors are integrated with glasses to estimate in-vivo CL variations with on-edge processing and classification. The system is validated through user-oriented audio-only (AO) and audio-visual (AV) trials. Participants are tested on their ability to comprehend target speech with different levels of background noise, assessing the impact on CL. The statistical features from the RF reflection data are processed using machine learning (ML) and deep learning (DL) algorithms implemented on a Raspberry Pi. The participants’ CL is classified into high, medium and low categories separately for AO and AV trials. The Multilayer Perceptron (MLP) achieves an overall accuracy of 85% for AO trials and 66.2% for AV trials, with average training and testing times of 6 seconds and 0.001 seconds, respectively. The promising results suggest that the system is an effective, portable, and low-cost alternative for on-edge CL estimation.
Usman Anwar, Yinhuan Dong, Tughrul Arslan, Amir Hussain 0001, Peter Lomax
ISCAS4
2025 Arabic Short-Text Dataset for Sentiment Analysis of Tourism and Leisure Events
abstract
ABSTRACT The focus of this study is to present the detailed process of collecting a dataset of Arabic short‐text in the tourism context and annotating this dataset for the task of sentiment analysis using an automatic zero‐shot labelling technique utilising transformer‐based models. This is benchmarked against a baseline manual annotation approach utilising native Arab human annotators. This study also introduces an approach exploiting both manual/handcrafted and automatically generated annotations of the dataset tweets for the task of sentiment analysis as part of a cross‐domain approach using a model trained on sarcasm labels and vice versa. The total collected corpus size is 2293 tweets; after annotation, these tweets were labelled in a three‐way classification approach as either positive, negative or neutral. We run different experiments to provide benchmark results of Arabic sentiment classification. Comparative results on our dataset show that the highest performing baseline model when utilising manual labels was MARBERT, with an accuracy of up to 87%, which was pre‐trained for Arabic on a massive amount of data. It should be noted that this model enhanced its performance additionally after pre‐training on a dialectical Arabic and modern standard Arabic corpus. On the other hand, zero‐shot automatically generated labels achieved an 84% accuracy rate in predicting sarcasm classes from sentiment labels.
Seham Basabain, Ahmed Yassin Al-Dubai, Erik Cambria, Khalid Al-Omar, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.5
2025 A Novel Continual Learning and Adaptive Sensing State Response-Based Target Recognition and Long-Term Tracking Framework for Smart Industrial Applications
abstract
ABSTRACT Purpose With the rapid development of artificial intelligence technology, highly intelligent and unmanned factories have become an important trend. In the complex environments of smart factories, the long‐term tracking and inspection of specified targets, such as operators and special products, as well as comprehensive visual recognition and decision‐making capabilities throughout the whole production process, are critical components of automated unmanned factories. However, challenges such as target occlusion and disappearance frequently occur, complicating long‐term tracking. Currently, there is limited research specifically focused on developing robust and comprehensive long‐term visual tracking frameworks for unmanned factories, particularly those designed to integrate with embedded platforms and overcome various challenges. Methods We first construct three new benchmark datasets in the complex workshop environment of a smart factory (referred to as SF‐Complex3 data), which include challenging conditions such as complete occlusion and partial occlusion of targets. A brain memory‐inspired approach is used to determine uncertainty estimation parameters, including confidence, peak‐to‐sidelobe ratio and average peak‐to‐correlation energy, to develop a continual learning‐based adaptive model update method. Additionally, we design a lightweight target detection model to automatically detect and locate targets in the initial frame and during re‐detection. Finally, we integrate the algorithm with ground mobile robots and unmanned aerial vehicles‐based imaging and processing equipment to build a new visual detection and tracking framework, smart factory complex recognition and tracking. Results We conducted extensive tests on the benchmark UAV20L and SF‐Complex3 datasets. The proposed algorithm demonstrates an average performance improvement of 6% when addressing key challenging attributes, compared to state‐of‐the‐art tracking methods. Additionally, the algorithm was capable of running efficiently on embedded platforms, including mobile robots and UAVs, at a real‐time speed of 36.4 frames per second. Conclusions The proposed SFC‐RT framework effectively addresses the challenges of target loss and occlusion in long‐term tracking within complex smart factory environments. The framework meets the requirements for real‐time performance, robustness and lightweight design, making it well suited for practical deployment.
Gun Li, Jie Tang 0001, Yang Li 0074, Shenbing Fu, Weizhong Qian, Qinsheng Zhu, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.11
2025 Transfer Learning-Based Automatic Sentiment Annotation of a Twitter-Based Arabic Mental Illness (AMI) Dataset
abstract
ABSTRACT Sentiment analysis, crucial for discerning emotional tones in text, relies on manual annotation to train machine learning models and is considered the gold standard for creating annotated corpora. However, this process is time‐consuming, labour‐intensive, and prone to biases. This paper proposes an automatic annotation approach for the Twitter‐based Arabic Mental Illness (AMI) dataset, which encompasses both Modern Standard Arabic and Dialectal Arabic. The approach leverages transfer learning with existing manually annotated datasets and three advanced Arabic language models to automate annotation, thereby enriching Arabic as a low‐resource language with labelled sentiment data. Validation was conducted by comparing the automatically generated annotations to manual annotation on the same dataset, achieving strong inter‐annotator agreement with a Cohen's Kappa statistic of k = 0.8457. Additionally, various baseline models were evaluated on the AMI dataset, identifying AraBERT as the top performer with the highest F1 score and accuracy.
Arwa Diwali, Kawther Saeedi, Kia Dashtipour, Mandar Gogate, Zain U. Hussain, Adam Howard, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.7
2025 Federated learning-driven dual blockchain for data sharing and reputation management in Internet of medical things
abstract
Abstract In the Internet of Medical Things (IoMT), the vulnerability of federated learning (FL) to single points of failure, low‐quality nodes, and poisoning attacks necessitates innovative solutions. This article introduces a FL‐driven dual‐blockchain approach to address these challenges and improve data sharing and reputation management. Our approach comprises two blockchains: the Model Quality Blockchain (MQchain) and the Reputation Incentive Blockchain (RIchain). MQchain utilizes an enhanced Proof of Quality (PoQ) consensus algorithm to exclude low‐quality nodes from participating in aggregation, effectively mitigating single points of failure and poisoning attacks by leveraging node reputation and quality thresholds. In parallel, RIchain incorporates a reputation evaluation, incentive mechanism, and index query mechanism, allowing for rapid and comprehensive node evaluation, thus identifying high‐reputation nodes for MQchain. Security analysis confirms the theoretical soundness of the proposed method. Experimental evaluation using real medical datasets, specifically MedMNIST, demonstrates the remarkable resilience of our approach against attacks compared to three alternative methods.
Chenquan Gan, Xinghai Xiao, Qingyi Zhu, Deepak Kumar Jain 0001, Akanksha Saini, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.6
2025 A Novel Reciprocal Domain Adaptation Neural Network for Enhanced Diagnosis of Chronic Kidney Disease
abstract
ABSTRACT Chronic kidney disease (CKD) is a major global health concern caused mostly by high blood pressure and glucose levels. Detecting CKD early is critical for reducing its negative consequences since it can lead to increased mortality rates. With CKD's rising incidence expected to make it the fifth biggest cause of death by 2040, rapid advances in diagnostic approaches are required. This study presents the Reciprocal Domain Adaptation Network (RDAN) as a potential approach to the various issues of CKD diagnosis. RDAN is a neural network model that will help to traverse the complexity of CKD diagnosis by smoothly combining diverse data sets. RDAN consists of two critical units at its foundation: Mutual Model Adaptation (MMA) and Domain Model Learning. The MMA unit uses a powerful Global and Local Pyramid Pooling technique to extract rich features from a variety of data domains. Meanwhile, the DML unit uses semi‐supervised domain‐independent features combined with MMA features to improve representation learning. RDAN includes a reciprocal regularizer to promote cross‐domain knowledge transfer, maximising feature representation for accurate CKD identification. An analysis of RDAN's performance on a variety of real‐world datasets showed remarkable results in terms of accuracy (96.94%), precision (98.81%), recall (98.73%), F1‐Score (98.88%), and area under the curve (AUC—99.35%). These results highlight the unmatched expertise of RDAN in managing data bias, domain changes, and privacy issues related to CKD diagnosis. Beyond statistical measures, RDAN's implications promise revolutionary breakthroughs in early CKD identification and subsequent therapeutic therapies. RDAN stands out as a groundbreaking method for diagnosing CKD. It delivers exceptional accuracy and can be seamlessly applied in various clinical environments.
Saeed Iqbal, Adnan N. Qureshi, Musaed Alhussein, Khursheed Aurangzeb, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.6
2025 Federated Learning: Concepts, Challenges and Implementation
abstract
ABSTRACT Federated Learning (FL) has emerged as an innovative approach for distributed neural networks, allowing multiple clients to collaboratively train a model without centralising their data, thus preserving decentralisation and data privacy. This review provides a comprehensive discussion of FL's core concepts, including its components, key challenges, and distinctions from traditional machine learning. The paper outlines the various types of FL, highlighting applications in privacy‐sensitive fields like healthcare and finance. It also addresses recent advancements in self‐supervised learning, personalisation, and multi‐modal applications within FL, as well as the integration of blockchain technology for enhanced privacy. Key advantages of FL are discussed, such as reduced communication overhead through the transmission of model parameters instead of raw data, which minimises network load and enhances privacy protection. Furthermore, the paper explores emerging questions for FL development, including scalability, fairness, and system standardisation. Real‐world examples, such as Google Gboard and brain tumour segmentation, are presented to illustrate FL's practical impact. Finally, the paper discusses future directions, including potential integration with other AI techniques like reinforcement learning and transfer learning. This review provides valuable insights for researchers and professionals who are new to FL or seek a broader understanding of its ecosystem. While there are few studies that explore limited aspect of FL, this review adopts a holistic approach and covers all aspects of FL including foundational concepts, implementation, challenges faced by FL, and real‐world implementation. The broader scope, which spans FL from concepts to practical implementation, makes it particularly distinctive and a valuable contribution.
Naeem Khan, Shibli Nisar, Muhammad Asghar Khan, Muhammad Attique Khan, David Camacho, Yasar Abbas Ur Rehman, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.7
2025 MTFDN: An image copy-move forgery detection method based on multi-task learning
abstract
Abstract Image copy‐move forgery, where an image region is copied and pasted within the same image, is a simple yet widely employed manipulation. In this paper, we rethink copy‐move forgery detection from the perspective of multi‐task learning and summarize two characteristics of this problem: (1) Homology and (2) Manipulated traces. Consequently, we propose a multi‐task forgery detection network (MTFDN) for image copy‐move forgery localization and source/target distinguishment. The network consists of a hard‐parameter sharing feature extractor, global forged homology detection (GFHD) and local manipulated trace detection (LMTD) modules. The difference of feature distribution between the GFHD module and the LMTD module is significantly reduced by sharing parameters. Experimental results on several benchmark copy‐move forgery datasets demonstrate the effectiveness of our proposed MTFDN.
Peng Liang 0003, Hang Tu, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.3
2025 An Attention-Driven Hybrid Deep Neural Network for Enhanced Heart Disease Classification
abstract
ABSTRACT Heart disease continues to be a primary cause of mortality globally, highlighting the critical necessity for efficient early prediction and classification techniques. This study presents a new hybrid model attention‐based CNN‐Bi‐LSTM that integrates the SMOTE with an attention‐driven improved convolutional neural network‐recurrent neural network architecture to improve the classification of heart sounds, especially from imbalanced datasets. Heart sounds are difficult to classify because of their complex acoustic properties and the variability of their characteristics across frequency and temporal domains. The proposed model utilises an advanced CNN to effectively extract global and local features, in conjunction with a bidirectional long short‐term memory network to improve the architecture by capturing contextual information from both preceding and subsequent time sequences. The incorporation of spatial attention within the CNN and temporal attention in the RNN enables the model to concentrate on the most pertinent audio segments. To address the challenges presented by imbalanced and noisy datasets that may impede the efficacy of deep learning algorithms, our model employs SMOTE to improve data representation. The hybrid model outperformed popular models such as CNN, LSTM and CNN‐LSTM, achieving a classification accuracy of more than 97% on the PCG and PASCAL heart sound datasets. The findings demonstrate the model's reliability as an initial evaluation tool in clinical settings, thereby improving support for cardiovascular disease diagnosis.
Umesh Kumar Lilhore, Sarita Simaiya, Musaed Alhussein, Surjeet Dalal, Khursheed Aurangzeb, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.6
2025 Medical Domain Knowledge Collaborative Graph Learning for Healthcare Event Prediction
abstract
ABSTRACT Electronic health records have become more prevalent worldwide, and with this, the opportunity for more accurate and automated prediction of health events has grown. Such predictions are crucial for providing preventive and proactive healthcare to patients. Although various advanced methods have been explored, they often fail to fully leverage medical domain knowledge, understand interrelations between diseases and patients comprehensively, and efficiently integrate unstructured clinical notes into predictive models. To address these challenges, we propose the Medical Domain Knowledge Collaborative Graph Learning (MED‐CGL) model. MED‐CGL incorporates external medical knowledge bases to enhance the predictive power of unstructured clinical notes and extracts learnable features from the MIMIC‐III health record dataset using medical domain knowledge and collaborative graph learning. We introduce the Enhanced Medical Knowledge Integration (EMKI) module, which employs a novel attention mechanism to connect clinical notes with disease descriptions precisely. It also enhances the system's performance by integrating medical knowledge from the semantically labelled knowledge‐enhanced (SLAKE) dataset during the training phase. Furthermore, our model considers the complexities of unstructured clinical notes, providing a nuanced perspective on the interplay between diseases and patient profiles. Our experiments show that the MED‐CGL model exhibited outstanding performance in diagnosis prediction, achieving an F1 score of 27.32%, and in heart failure prediction, where it attained an accuracy of 91.39%. This significant improvement demonstrates the robustness and effectiveness of our model, which is further supported by our in‐depth ablation study.
Usman Naseem, Junaid Rashid, Haohui Lu, Dominic Ng, Zain U. Hussain, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.6
2025 Generative AI for Finance: Applications, Case Studies and Challenges
abstract
ABSTRACT Generative AI (GAI), which has become increasingly popular nowadays, can be considered a brilliant computational machine that can not only assist with simple searching and organising tasks but also possesses the capability to propose new ideas, make decisions on its own and derive better conclusions from complex inputs. Finance comprises various difficult and time‐consuming tasks that require significant human effort and are highly prone to errors, such as creating and managing financial documents and reports. Hence, incorporating GAI to simplify processes and make them hassle‐free will be consequential. Integrating GAI with finance can open new doors of possibility. With its capacity to enhance decision‐making and provide more effective personalised insights, it has the power to optimise financial procedures. In this paper, we address the research gap of the lack of a detailed study exploring the possibilities and advancements of the integration of GAI with finance. We discuss applications that include providing financial consultations to customers, making predictions about the stock market, identifying and addressing fraudulent activities, evaluating risks, and organising unstructured data. We explore real‐world examples of GAI, including Finance generative pre‐trained transformer (GPT), Bloomberg GPT, and so forth. We look closer at how finance professionals work with AI‐integrated systems and tools and how this affects the overall process. We address the challenges presented by comprehensibility, bias, resource demands, and security issues while at the same time emphasising solutions such as GPTs specialised in financial contexts. To the best of our knowledge, this is the first comprehensive paper dealing with GAI for finance.
Siva Sai, Keya Arunakar, Vinay Chamola, Amir Hussain 0001, Pranav Bisht
Expert Syst. J. Knowl. Eng.4
2025 A Novel Approach to Fire Detection With Enhanced Target Localisation and Recognition
abstract
ABSTRACT Real‐time monitoring of fires is crucial for safeguarding lives and property. However, current fire detection methods still suffer from issues such as redundant feature information, poor network generalisation capabilities and low perception of target location information. To address these challenges, a novel fire detection method called YOLO‐FDI has been proposed. This method utilises partial convolution and coordinate convolution with attention mechanisms and Alpha loss at different stages. Specifically, to enhance target localisation accuracy, an attention mechanism is integrated into the model to autonomously focus on fire‐affected areas. In terms of feature extraction, partial convolution is employed to reduce computational redundancy and memory access, improving performance and effectively extracting spatial features. During the feature fusion stage, coordinate convolution embeds feature information into coordinate data, further enhancing the coordinate perception capabilities of pixels on the feature map, thereby improving adaptability and accuracy in detecting fire targets. Additionally, the model utilises Alpha loss to enhance flexibility and robustness in fire object detection and recognition. Experimental results demonstrate the effectiveness of the proposed model based on three self‐constructed datasets. Compared to the baseline YOLOv7 model, its mAP has improved by 4.5 percentage points, 1.7 percentage points and 2.6 percentage points, respectively. This method demonstrates the capability to accurately represent fire targets and exhibits better stability and reliability in fire target detection, effectively reducing false positives and missed detections.
Le Zou, Fengling Jiang, Zhize Wu, Lingma Sun, Mandar Gogate, Kia Dashtipour, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.9
2025 Adaptive fuzzy convolution networks for uncertainty-aware image analysis in ambiguous environments
Saeed Iqbal, Xiaopin Zhong, Muhammad Attique Khan, Zongze Wu 0001, Amir Hussain 0001, Shrooq Alsenan, Weixiang Liu
Expert Syst. Appl.5
2025 PAFPT: Progressive aggregator with feature prompted transformer for underwater image enhancement
Jing Yang 0041, Shanbing Zhu, Shumin Bai, Fengling Jiang, Amir Hussain 0001
Expert Syst. Appl.6
2025 RI-L1Approx: A novel Resnet-Inception-based Fast L1-approximation method for face recognition
Supriya Bajpai, Gargi Mishra, Rachna Jain, Deepak Kumar Jain 0001, Dharmender Saini, Amir Hussain 0001
Neurocomputing6
2025 A novel pain sentiment detection system utilizing a PainCapsule model and textual facial patterns
Anay Ghosh, Saiyed Umer, Bibhas Chandra Dhara, Deepak Kumar Jain 0001, Ranjeet Kumar Rout, Amir Hussain 0001
Neurocomputing6
2025 TIxAI: A Trustworthiness Index for eXplainable AI in skin lesions classification
abstract
Skin cancer is one of the leading causes of mortality worldwide. Early diagnosis can ensure more effective patient treatment and outcomes, but, this is challenging due to the high similarity between different skin lesion types. There is a growing interest in developing Artificial Intelligence (AI)-based systems for automated skin lesion classification. However, current AI models are not transparent, leading to a lack of trust from clinicians who struggle to interpret and validate AI decisions. To this end, in this paper, a fine tuned EfficientNet-B0-based classifier is first developed to classify dermoscopic images of Melanoma (MEL), Nevus (NV) and Seborrheic Keratosis (SK) skin lesions gathered from the International Skin Imaging Collaboration (ISIC) dataset. Next, the explainability of the model is investigated. In particular, a new Trustworthiness Index for eXplainable AI, herein referred to as TIxAI , is proposed. The TIxAI is based on the difference between the relevance degree of the lesion and non-lesion areas, leading to the conclusion that the higher the TIxAI , the more trustworthy the classifier is expected to be. Experimental results support the use of the proposed TIxAI to assess and benchmark the reliability of classifiers also in other real-world applications.
Cosimo Ieracitano, Francesco Carlo Morabito, Amir Hussain 0001, Muhammad Suffian Nizami, Nadia Mammone
Neurocomputing3
2025 FairBias: Mitigating bias in medical image diagnosis with mixed noise and class imbalance
Saeed Iqbal, Xiaopin Zhong, Muhammad Attique Khan, Zongze Wu 0001, Nouf Almujally, Weixiang Liu, Amir Hussain 0001
Neurocomputing7
2025 Open-Pose 3D zero-shot learning: Benchmark and challenges
Weiguang Zhao, Guanyu Yang 0002, Rui Zhang 0012, Chenru Jiang, Chaolong Yang, Yuyao Yan, Amir Hussain 0001, Kaizhu Huang
Neural Networks7
2024 Unveiling Diplomatic Narratives: Analyzing United Nations Security Council Debates Through Metaphorical Cognition
Rui Mao 0010, Qian Liu 0012, Amir Hussain 0001, Erik Cambria
CogSci4
2024 Edged based audio-visual speech enhancement demonstrator
Song Chen 0005, Mandar Gogate, Kia Dashtipour, Jasper Kirton-Wingate, Adeel Hussain, Faiyaz Doctor, Tughrul Arslan, Amir Hussain 0001
INTERSPEECH8
2024 Real-Time Gaze-directed speech enhancement for audio-visual hearing-aids
Arif Reza Anway, Bryony Buck, Mandar Gogate, Kia Dashtipour, Michael A. Akeroyd, Amir Hussain 0001
INTERSPEECH6
2024 Arabic text classification based on analogical proportions
abstract
Abstract Text classification is the process of labelling a given set of text documents with predefined classes or categories. Existing Arabic text classifiers are either applying classic Machine Learning algorithms such as k‐NN and SVM or using modern deep learning techniques. The former are assessed using small text collections and their accuracy is still subject to improvement while the latter are efficient in classifying big data collections and show limited effectiveness in classifying small corpora with a large number of categories. This paper proposes a new approach to Arabic text classification to treat small and large data collections while improving the classification rates of existing classifiers. We first demonstrate the ability of analogical proportions (AP) (statements of the form ‘x is to as is to ’), which have recently been shown to be effective in classifying ‘structured’ data, to classify ‘unstructured’ text documents requiring preprocessing. We design an analogical model to express the relationship between text documents and their real categories. Next, based on this principle, we develop two new analogical Arabic text classifiers. These rely on the idea that the category of a new document can be predicted from the categories of three others, in the training set, in case the four documents build together a ‘valid’ analogical proportion on all or on a large number of components extracted from each of them. The two proposed classifiers (denoted AATC1 and AATC2) differ mainly in terms of the keywords extracted for classification. To evaluate the proposed classifiers, we perform an extensive experimental study using five benchmark Arabic text collections with small or large sizes, namely ANT (Arabic News Texts) v2.1 and v1.1, BBC‐Arabic, CNN‐Arabic and AlKhaleej‐2004. We also compare analogical classifiers with both classical ML‐based and Deep Learning‐based classifiers. Results show that AATC2 has the best average accuracy (78.78%) over all other classifiers and the best average precision (0.77) ranked first followed by AATC1 (0.73), NB (0.73) and SVM (0.72) for the ANT corpus v2.1. Besides, AATC1 shows the best average precisions (0.88) and (0.92), respectively for the BBC‐Arabic corpus and AlKhaleej‐2004, and the best average accuracy (85.64%) for CNN‐Arabic over all other classifiers. Results demonstrate the utility of analogical proportions for text classification. In particular, the proposed analogical classifiers are shown to significantly outperform a number of existing Arabic classifiers, and in many cases, compare favourably to the robust SVM classifier.
Myriam Bounhas, Bilel Elayeb, Amina Chouigui, Amir Hussain 0001, Erik Cambria
Expert Syst. J. Knowl. Eng.4
2024 A novel generative adversarial network-based super-resolution approach for face recognition
abstract
Abstract Face recognition is an essential feature required for a range of computer vision applications such as security, attendance systems, emotion detection, airport check‐in, and many others. The super‐resolution of subject images is an important and challenging element in numerous scenarios. At times the images are low resolution and need to be processed through super‐resolution techniques to gain more accurate results. For the problem of image super‐resolution, deep learning‐based face recognition systems have been explored in recent years; however, low‐resolution face recognition remains an arduous task. Generative adversarial network (GAN) based models are a promising approach to address this challenge. However, conventional GAN‐based models may generate images that differ significantly from an original high‐resolution image in the test set to the point that the identity of the target face may be changed. To address this shortcoming, we propose a novel U‐Net style generator architecture, where skip‐connections between the encoder and decoder layer can help in preserving the facial characteristics of the input image in the generated image, thus curbing the generator's ability to generate an entirely new image and training it to generate an image more similar in characteristics to the original image. In addition to statistical metrics like structural similarity index measure and Fréchet inception distance, we compute the pixel‐wise distance between the original and model‐generated images to ascertain that our model generates as close to the original images as possible. While we train the model for 4× super‐resolution (64 × 64 images to 256 × 256), our architecture can also be trained for an arbitrary resizing scale. Finally, the number of faces detected over high‐resolution images generated by our model is shown to be higher than state‐of‐the‐art high‐resolution image creation models for face recognition tasks.
Amit Chougule, Shreyas Kolte, Vinay Chamola, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.4
2024 Hate speech detection: A comprehensive review of recent works
abstract
Abstract There has been surge in the usage of Internet as well as social media platforms which has led to rise in online hate speech targeted on individual or group. In the recent years, hate speech has resulted in one of the challenging problems that can unfurl at a fast pace on digital platforms leading to various issues such as prejudice, violence and even genocide. Considering the acceptance of Artificial Intelligence (AI) and Natural Language Processing (NLP) techniques in varied application domains, it would be intriguing to consider these techniques for automated hate speech detection. In literature, there have been efforts to recognize and categorize hate speech using varied Machine Learning (ML) and Deep Learning (DL) techniques. Hence, considering the need and provocations for hate speech detection we aim to present a comprehensive review that discusses fundamental taxonomy as well as recent advances in the field of online hate speech identification. There is a significant amount of literature related to the initial phases of hate speech detection. The background section provides a detailed explanation of the previous research. The subsequent section that follows is dedicated to examining the recent literature published from the year 2020 onwards. The paper presents some of the hate speech datasets considered for hate speech detection. Furthermore, the paper discusses different data modalities, namely, textual hate speech detection, multi‐modal hate speech detection and multilingual hate speech detection. Apart from systematic review on hate speech detection, the paper also implement several multi‐label models to compare the performance of hate speech detection by employing classic ML technique namely, Logistic Regression and DL technique namely, Long Short‐Term Memory (LSTM) and a multiclass multi‐label architecture. In the implemented architecture, we have derived two new elements to quantify the hatefulness and intensity of hatred to improve the results for hate speech detection using Indonesian tweet dataset. Empirical Analysis of the model reveals that the implemented approach outperforms and is able to achieve improved results for the underlying dataset.
Ankita Gandhi, Param Ahir, Kinjal Adhvaryu, Pooja Shah, Ritika Lohiya, Erik Cambria, Soujanya Poria, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.8
2024 Multi-model deep learning system for screening human monkeypox using skin images
abstract
Abstract Purpose Human monkeypox (MPX) is a viral infection that transmits between individuals via direct contact with animals, bodily fluids, respiratory droplets, and contaminated objects like bedding. Traditional manual screening for the MPX infection is a time‐consuming process prone to human error. Therefore, a computer‐aided MPX screening approach utilizing skin lesion images to enhance clinical performance and alleviate the workload of healthcare providers is needed. The primary objective of this work is to devise an expert system that accurately classifies MPX images for the automatic detection of MPX subjects. Methods This work presents a multi‐modal deep learning system through the fusion of convolutional neural network (CNN) and machine learning algorithms, which effectively and autonomously detect MPX‐infected subjects using skin lesion images. The proposed framework, termed MPXCN‐Net is developed by fusing deep features of three pre‐trained CNNs: MobileNetV2, DarkNet19, and ResNet18. Three classifiers—K‐nearest neighbour, support vector machine (SVM), and ensemble classifier—with various kernel functions, are used to identify infected patients. To validate the efficacy of our proposed system, we employ a publicly accessible MPX skin lesion dataset. Results By amalgamating features extracted from all three CNNs and utilizing the medium Gaussian kernel of the SVM classifier, our proposed system achieves an outstanding average classification accuracy of 90.4%. Conclusions Developed MPXCN‐Net is suitable for testing with a large diversified dataset before being used in clinical settings.
Kapil Gupta 0002, Varun Bajaj, Deepak Kumar Jain 0001, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.4
2024 Bare-Bones particle Swarm optimization-based quantization for fast and energy efficient convolutional neural networks
abstract
Abstract Neural network quantization is a critical method for reducing memory usage and computational complexity in deep learning models, making them more suitable for deployment on resource‐constrained devices. In this article, we propose a method called BBPSO‐Quantizer, which utilizes an enhanced Bare‐Bones Particle Swarm Optimization algorithm, to address the challenging problem of mixed precision quantization of convolutional neural networks (CNNs). Our proposed algorithm leverages a new population initialization, a robust screening process, and a local search strategy to improve the search performance and guide the population towards a feasible region. Additionally, Deb's constraint handling method is incorporated to ensure that the optimized solutions satisfy the functional constraints. The effectiveness of our BBPSO‐Quantizer is evaluated on various state‐of‐the‐art CNN architectures, including VGG, DenseNet, ResNet, and MobileNetV2, using CIFAR‐10, CIFAR‐100, and Tiny ImageNet datasets. Comparative results demonstrate that our method delivers an excellent tradeoff between accuracy and computational efficiency.
Jihene Tmamna, Emna Ben Ayed, Rahma Fourati, Amir Hussain 0001, Mounir Ben Ayed
Expert Syst. J. Knowl. Eng.4
2024 A novel end-to-end deep convolutional neural network based skin lesion classification framework
abstract
Skin diseases are reported to contribute 1.79% of the global burden of disease. The accurate diagnosis of specific skin diseases is known to be a challenging task due, in part, to variations in skin tone, texture, body hair, etc. Classification of skin lesions using machine learning is a demanding task, due to the varying shapes, sizes, colors, and vague boundaries of some lesions. The use of deep learning for the classification of skin lesion images has been shown to help diagnose the disease at its early stages. Recent studies have demonstrated that these models perform well in skin detection tasks, with high accuracy and efficiency. Our paper proposes an end-to-end framework for skin lesion classification, and our contributions are two-fold. Firstly, two fundamentally different algorithms are proposed for segmenting and extracting features from images during image preprocessing. Secondly, we present a deep convolutional neural network model, S-MobileNet that aims to classify 7 different types of skin lesions. We used the HAM10000 dataset, which consists of 10000 dermatoscopic images from different populations and is publicly available through the International Skin Imaging Collaboration (ISIC) Archive. The image data was preprocessed to make it suitable for modeling. Exploratory data analysis (EDA) was performed to understand various attributes and their relationships within the dataset. A modified version of a Gaussian filtering algorithm and SFTA was applied for image segmentation and feature extraction. The processed dataset was then fed into the S-MobileNet model. This model was designed to be lightweight and was analysed in three dimensions: using the Relu Activation function, the Mish activation function, and applying compression at intermediary layers. In addition, an alternative approach for compressing layers in the S-MobileNet architecture was applied to ensure a lightweight model that does not compromise on performance. The model was trained using several experiments and assessed using various performance measures, including, loss, accuracy, precision, and the F1-score. Our results demonstrate an improvement in model performance when applying a preprocessing technique. The Mish activation function was shown to outperform Relu. Further, the classification accuracy of the compressed S-MobileNet was shown to outperform S-MobileNet. To conclude, our findings have shown that our proposed deep learning-based S-MobileNet model is the optimal approach for classifying skin lesion images in the HAM10000 dataset. In the future, our approach could be adapted and applied to other datasets, and validated to develop a skin lesion framework that can be utilised in real-time.
A. Razia Sulthana, Vinay Chamola, Zain U. Hussain, Faisal Albalwy, Amir Hussain 0001
Expert Syst. Appl.5
2024 Deep learning methods for early detection of Alzheimer's disease using structural MR images: a survey
Sonia Ben Hassen Neji, Mohamed Neji, Zain U. Hussain, Amir Hussain 0001, Adel M. Alimi, Mondher Frikha
Neurocomputing4
2024 Graph learning with label attention and hyperbolic embedding for temporal event prediction in healthcare
abstract
The digitization of healthcare systems has led to the proliferation of electronic health records (EHRs), serving as comprehensive repositories of patient information. However, the vast volume and complexity of EHR data present challenges in extracting meaningful insights. This paper addresses the need for automated analysis of EHRs by proposing a novel graph learning model with label attention (GLLA) for temporal event prediction. GLLA utilizes graph neural networks to capture intricate relationships between medical codes and patients, incorporating hierarchical structures and shared risk factors. Furthermore, it introduces the Label Attention and Attention-based Transformer (LAAT) algorithm to analyze unstructured clinical notes as a multi-label classification problem. Evaluation on the widely-used MIMIC III dataset demonstrates the efficacy of GLLA in enhancing diagnostic prediction performance. The contributions of this research include a comprehensive analysis of existing models, the identification of limitations, and the development of innovative approaches to improve the accuracy and effectiveness of EHR analysis. Ultimately, GLLA aims to advance healthcare decision-making, disease management strategies, and patient outcomes.
Usman Naseem, Surendrabikram Thapa, Qi Zhang 0020, Shoujin Wang, Junaid Rashid, Liang Hu 0004, Amir Hussain 0001
Neurocomputing7
2024 A binary particle swarm optimization-based pruning approach for environmentally sustainable and robust CNNs
Jihene Tmamna, Rahma Fourati, Emna Ben Ayed, Leandro A. Passos Junior, João Paulo Papa, Mounir Ben Ayed, Amir Hussain 0001
Neurocomputing7
2024 Overtaking Mechanisms Based on Augmented Intelligence for Autonomous Driving: Data Sets, Methods, and Challenges
abstract
The field of autonomous driving research has made significant strides towards achieving full automation, endowing vehicles with self-awareness and independent decision-making. However, integrating automation into vehicular operations presents formidable challenges, especially as these vehicles must seamlessly navigate public roads alongside other cars and pedestrians. An intriguing yet relatively underexplored domain within autonomous driving is overtaking. Overtaking involves a dynamic interplay of complex tasks, including precise steering and speed control, rendering it one of the most intricate operations for implementing augmented intelligence driving technologies. Surprisingly, the overtaking of autonomous vehicles remains largely uncharted territory in the context of augmented intelligence for autonomous systems. This void in knowledge beckons researchers to embark on explorations and investigations in this nascent field. Our review paper systematically synthesises overtaking methodologies hinging on computer vision techniques tailored for augmented intelligence autonomous driving scenarios in response to this pressing need. Our analysis encompasses an array of domains central to overtaking in augmented intelligence autonomous vehicles, encompassing Object Detection, Lane/Line Detection, Depth Estimation, Obstacle Detection, Segmentation, and Pedestrian Detection. We meticulously analyze each domain using well-established Multimodal datasets. We assess different models’ performance across various parameters by employing graphical structures, enabling visual comparative analyses. In object detection, YOLOv4 achieves a top performance with 0.90 mAP on the BDD100K dataset. For lane detection, CLRNET excels with the highest F1 score of around 0.96 on the LLAMAS dataset. ViT-Adapter-L leads in segmentation tasks, boasting an impressive mIoU score of 83 on Cityscapes. The Hierarchical Model achieves a superior mAP of 0.90 in road sign detection on the Tsinghua-Tencent Dataset. Steering angle computation sees InterFuser as the standout, achieving the highest driving score of approximately 74.0. This paper’s primary contributions include a comprehensive assessment of diverse models for each Multimodal dataset, aiding future research in this evolving domain.
Vinay Chamola, Amit Chougule, Aishwarya Sam, Amir Hussain 0001, F. Richard Yu
IEEE Internet Things J.4
2024 VLC-Assisted Safety Message Dissemination in Roadside Infrastructure-Less IoV Systems: Modeling and Analysis
abstract
Internet of Vehicles (IoV) is an emerging paradigm with significant potential to improve traffic efficiency and driving safety. Here, we focus on the design of a novel visible light communication (VLC)-assisted scheme to enable driving safety-related Internet of Vehicles (IoV) services that require ultrareliable and low-latency communications (URLLC). Specifically, the Vehicle-to-Vehicle (V2V) communication mode is adopted to satisfy the ultralow latency requirement of URLLC in roadside infrastructure-less IoV systems. In the outdoor V2V- VLC scenarios, the quality of the received optical signal is degraded by path loss, atmospheric turbulence and additive noise. In addition, the short-packet feature of URLLC introduces inevitable data decoding errors and imperfect channel state information (CSI). With this background, we aim to investigate the reliability performance of URLLC in outdoor V2V- VLC systems, which is described by the average packet loss probability under given user-plane transmission latency. First, we consider the ideal case of a perfect CSI at the receiver, and derive an analytical expression of average packet loss probability. Further, a closed-form approximation is provided to simplify the numerical calculation. Next, we extend the theoretical analysis to a practical V2V- VLC system with imperfect CSI at the receiver. Through numerical results, we validate the accuracy of our designed theoretical framework and propose ideas to enable driving safety-related IoV services in outdoor V2V- VLC systems.
Yuncong Xie, Dongyang Xu 0003, Keping Yu, Amir Hussain 0001, Mohsen Guizani
IEEE Internet Things J.5
2024 A change severity degree-based dynamic multi-objective optimization algorithm with adaptive response strategy
Najwa Kouka, Rahma Fourati, Raja Fdhila, Amir Hussain 0001, Adel M. Alimi
Inf. Sci.4
2024 DDformer: Dimension decomposition transformer with semi-supervised learning for underwater image enhancement
Zhi Gao 0005, Jing Yang 0041, Fengling Jiang, Xixiang Jiao, Kia Dashtipour, Mandar Gogate, Amir Hussain 0001
Knowl. Based Syst.7
2024 A novel IoT-based deep neural network for COVID-19 detection using a soft-attention mechanism
Zeineb Fki, Boudour Ammar, Rahma Fourati, Hela Fendri, Amir Hussain 0001, Mounir Ben Ayed
Multim. Tools Appl.5
2024 Sentiment Analysis Meets Explainable Artificial Intelligence: A Survey on Explainable Sentiment Analysis
abstract
Sentiment analysis can be used to derive knowledge that is connected to emotions and opinions from textual data generated by people. As computer power has grown, and the availability of benchmark datasets has increased, deep learning models based on deep neural networks have emerged as the dominant approach for sentiment analysis. While these models offer significant advantages, their lack of interpretability poses a major challenge in comprehending the rationale behind their reasoning and prediction processes, leading to complications in the models' explainability. Further, only limited research has been carried out into developing deep learning models that describe their internal functionality and behaviors. In this timely study, we carry out a first of its kind overview of key sentiment analysis techniques and eXplainable artificial intelligence (XAI) methodologies that are currently in use. Furthermore, we provide a comprehensive review of sentiment analysis explainability.
Arwa Diwali, Kawther Saeedi, Kia Dashtipour, Mandar Gogate, Erik Cambria, Amir Hussain 0001
IEEE Trans. Affect. Comput.6
2024 STIDNet: Identity-Aware Face Forgery Detection With Spatiotemporal Knowledge Distillation
abstract
The impressive development of facial manipulation techniques has raised severe public concerns. Identity-aware methods, especially suitable for protecting celebrities, are seen as one of promising face forgery detection approaches with additional reference video. However, without in-depth observation of fake video’s characteristics, most existing identity-aware algorithms are just naive imitation of face verification model and fail to exploit discriminative information. In this article, we argue that it is necessary to take both spatial and temporal perspectives into consideration for adequate inconsistency clues and propose a novel forgery detector named SpatioTemporal IDentity network (STIDNet). To effectively capture heterogeneous spatiotemporal information in a unified formulation, our STIDNet is following a knowledge distillation architecture that the student identity extractor receives supervision from a spatial information encoder (SIE) and a temporal information encoder (TIE) through multiteacher training. Specifically, a regional sensitive identity modelling paradigm is proposed in SIE by introducing facial blending augmentation but with uniform identity label, thus encourage model to focus on spatial discriminative region like outer face. Meanwhile, considering the strong temporal correlation between audio and talking face video, our TIE is devised in a cross-modal pattern that the audio information is introduced to supervise model exploiting temporal personalized movements. Benefit from knowledge transfer from SIE and TIE, STIDNet is able to capture individual’s essential spatiotemporal identity attributes and sensitive to even subtle identity deviation caused by manipulation. Extensive experiments indicate the superiority of our STIDNet compared with previous works. Moreover, we also demonstrate STIDNet is more suitable for real-world implementation in terms of model complexity and reference set size.
Mingqi Fang, Lingyun Yu 0002, Hongtao Xie 0001, Qingfeng Tan, Zhiyuan Tan 0001, Amir Hussain 0001, Zezheng Wang 0002, Zhihong Tian 0001
IEEE Trans. Comput. Soc. Syst.6
2024 Toxic Fake News Detection and Classification for Combating COVID-19 Misinformation
abstract
The emergence of COVID-19 has led to a surge in fake news on social media, with toxic fake news having adverse effects on individuals, society, and governments. Detecting toxic fake news is crucial, but little prior research has been done in this area. This study aims to address this gap and identify toxic fake news to save time spent on examining nontoxic fake news. To achieve this, multiple datasets were collected from different online social networking platforms such as Facebook and Twitter. The latest samples were obtained by collecting data based on the topmost keywords extracted from the existing datasets. The instances were then labeled as toxic/nontoxic using toxicity analysis, and traditional machine-learning (ML) techniques such as linear support vector machine (SVM), conventional random forest (RF), and transformer-based ML techniques such as bidirectional encoder representations from transformers (BERT) were employed to design a toxic-fake news detection (FND) and classification system. As per the experiments, the linear SVM method outperformed BERT SVM, RF, and BERT RF with an accuracy of 92% and -score, -score, and -score of 95%, 85%, and 87%, respectively. Upon comparison, the proposed approach has either suppressed or achieved results very close to the state-of-the-art techniques in the literature by recording the best values on performance metrics such as accuracy, F1-score, precision, and recall for linear SVM. Overall, the proposed methods have shown promising results and urge further research to restrain toxic fake news. In contrast to prior research, the presented methodology leverages toxicity-oriented attributes and BERT-based sequence representations to discern toxic counterfeit news articles from nontoxic ones across social media platforms.
Mudasir Ahmad Wani, Mohammed Ahmed El-Affendi, Kashish Ara Shakil, Ibrahem Mohammed Abuhaimed, Anand Nayyar, Amir Hussain 0001, Ahmed A. Abd El-Latif 0001
IEEE Trans. Comput. Soc. Syst.6
2024 Context-Aware Audio-Visual Speech Enhancement Based on Neuro-Fuzzy Modeling and User Preference Learning
abstract
It is estimated that by 2050 approximately one in ten individuals globally will experience disabling hearing impairment. In the presence of everyday reverberant noise, a substantial proportion of individual users encounter challenges in speech comprehension. This study introduces a novel application of neuro-fuzzy modeling that synergizes and fuses audio-visual speech enhancement (AV SE) with an initial user preference learning based framework. Specifically, our approach uniquely integrates multimodal AV speech data with innovative SE methods and fuzzy inferencing techniques. This integration is further enriched by incorporating a user-preference learning model that adapts to environmental and user-specific contexts, including signal-to-noise ratios, sound power, and the quality of visual information. The proposed framework facilitates the incorporation of clinical measures such as user cognitive load (or listening effort) with real-world uncertainty to steer the system outputs. We employ an adaptive fuzzy neural network to derive the most effective Sugeno fuzzy inference model, employing particle swarm optimization to ensure optimal SE by considering sound power, ambient noise levels, and visual quality. Experimental results utilize our new benchmark AV multitalker challenge dataset to demonstrate the superiority of our user preference-informed, context-aware AV SE approach in enhancing speech intelligibility and quality in challenging noisy conditions, marking a significant advancement over conventional methods while reducing energy consumption. The conclusion supports the ecological scalability of our approach and its potential for real-world applications, setting a new benchmark in AV SE research, paving the way for future assistive hearing and communication technologies.
Song Chen 0005, Jasper Kirton-Wingate, Faiyaz Doctor, Usama Arshad, Kia Dashtipour, Mandar Gogate, Zahid Halim, Ahmed Yassin Al-Dubai, Tughrul Arslan, Amir Hussain 0001
IEEE Trans. Fuzzy Syst.10
2024 SAR Target Incremental Recognition Based on Features With Strong Separability
abstract
With the rapid development of deep learning technology, many synthetic aperture radar (SAR) target recognition algorithms based on convolutional neural networks have achieved exceptional performance on various datasets. However, conventional neural networks are repeatedly iterated on a fixed dataset until convergence, and once they learn new tasks, a large amount of previously learned knowledge is forgotten, leading to a significant decline in performance on old tasks. This article presents an incremental learning method based on strong separability features (SSF-IL) to address the model’s forgetting of previously learned knowledge. The SSF-IL employs both intraclass and interclass scatter to compute the feature separability loss, in order to enhance the linear separability of features during incremental learning. In the process of learning new classes, an intraclass clustering loss is proposed to replace the conventional knowledge distillation. This loss function constrains the old class features to cluster around the saved class centers, maintaining the separability among the old class features. Finally, a classifier bias correction method based on boundary features is designed to reinforce the classifier’s decision boundary and reduce classification errors. SAR target incremental recognition experiments are conducted on the MSTAR dataset, and the results are compared with several existing incremental learning algorithms to demonstrate the effectiveness of the proposed algorithm.
Fei Gao 0005, Lingzhe Kong, Rongling Lang, Jinping Sun, Jun Wang 0041, Amir Hussain 0001, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 BBox-Free SAR Ship Instance Segmentation Method Based on Gaussian Heatmap
abstract
Recently, deep learning methods have been widely adopted for ship detection in synthetic aperture radar (SAR) images. However, many of the existing methods miss adjacent ship instances when detecting densely arranged ship targets in inshore scenes. Besides, they suffer from the lack of precision in the instance indication information and the confusion of multiple instances by a single mask head. In this paper, we propose a novel center point prediction algorithm, which detects the center points by finding a long distance variation relationship between two points. The whole prediction process is anchor-free and does not require additional bounding box (BBox) predictions for non-maximum suppression (NMS). Therefore, our algorithm is BBox-free and NMS-free, solving the problem of low recall rates when conducting NMS for densely arranged targets. Furthermore, to tackle the deficiency of position indication information in localization tasks, we introduce a feature fusion module with feature decoupling (FD). This module uses classification branch to provide guidance information for localization branch, while suppressing the influence of the gradient flow mixing, effectively improving the algorithm’s segmentation performance of ship contours. Finally, through principal component analysis (PCA) of the Gaussian distribution covariance matrix, we propose a loss function based on the distance between centroids and the difference of angle, called centroid and angle constraint (CAC). CAC guides the network in learning the criterion that a single dynamic mask head is only valid for a single instance. Experiments conducted on PSeg-SSDD and HRSID demonstrate the effectiveness and robustness of our method.
Fei Gao 0005, Fengjun Zhong, Jinping Sun, Amir Hussain 0001, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 ESPP: Efficient Sector-Based Charging Scheduling and Path Planning for WRSNs With Hexagonal Topology
abstract
Wireless Power Transfer (WPT) is a promising technology that can potentially mitigate the energy provisioning problem for sensor networks. In order to efficiently replenish energy for these battery-powered devices, designing appropriate scheduling and charging path planning algorithms is essential and challenging. Whilst previous studies have tackled this challenge, the conjoint influences of network topology, charging path planning, and energy threshold distribution in Wireless Rechargeable Sensor Networks (WRSNs) are still in their infancy. We mitigate the aforementioned problem by proposing novel algorithmic solutions to efficient sector-based on-demand charging scheduling and path planning. Specifically, we first propose a hexagonal cluster-based deployment of nodes such that finding an NP-Complete Hamiltonian path is feasible. Second, each cluster is divided into multiple sectors and a charging path planning algorithm is implemented to yield a Hamiltonian path, aimed at improving the Mobile Charging Vehicle (MCV) efficiency and charging throughput. Third, we propose an efficient algorithm to calculate theimportanceof nodes to be used for charging duration decision-making and prioritization. Fourth, a non-preemptive dynamic priority scheduling algorithm is proposed for charging tasks’ assignments and scheduling. Finally, extensive simulations have been conducted, revealing the significant advantages of our proposed algorithms in terms of energy efficiency, response time, dead nodes’ density, and queuing processing.
Abdulbary Naji, Ammar Hawbani, Xingfu Wang, Haithm M. Al-Gunid, Yunes Al-Dhabi, Ahmed Yassin Al-Dubai, Amir Hussain 0001, Liang Zhao 0004, Saeed H. Alsamhi
IEEE Trans. Sustain. Comput.7
2023 Deep Learning-Based Receiver Design for IoT Multi-User Uplink 5G-NR System
abstract
Designing an efficient receiver for multiple users transmitting orthogonal frequency-division multiplexing signals to the base station remain a challenging interference-limited problem in 5G-new radio (5G-NR) system. This can lead to stagnation of decoding performance at higher signal-to-noise-and-interference regimes. Further, the problem is exacerbated in future critical internet-of-thing (IoT) devices operating on smaller block size due to latency constraints and IoT users moving at varying speeds introducing Doppler shift and delay spread. In this work, we propose a novel deep learning (DL)-based U-net- and Resnet-inspired receiver for multi-user uplink transmission for a 5G-NR system that replaces only the signal demapping block of the receiver chain. Compared to traditional U-net frameworks, we propose a DL receiver with upsampling in the encoder that takes complex equalized symbols as input and downsampling in the decoder to output bit-wise log-likelihood ratios for multiple users. Further, residual skip connections are introduced in the decoder to facilitate stronger connections to the upsampling blocks. Finally, the DL receiver is optimized by maximizing the optimal bit-metric decoding rate. Comparative simulations show that our proposed DL receiver outperforms traditional 5G-NR receivers by considerable margins.
Ankit Gupta 0008, Abhijeet Bishnu, Tharmalingam Ratnarajah, Ahsan Adeel, Amir Hussain 0001, Mathini Sellathurai
GLOBECOM5
2023 Towards Deeper and Better Multi-view Feature Fusion for 3D Semantic Segmentation
Chaolong Yang, Yuyao Yan, Weiguang Zhao, Jianan Ye, Xi Yang 0008, Amir Hussain 0001, Bin Dong 0003, Kaizhu Huang
ICONIP (15)6
2023 5G-IoT Cloud based Demonstration of Real-Time Audio-Visual Speech Enhancement for Multimodal Hearing-aids
Ankit Gupta 0008, Abhijeet Bishnu, Mandar Gogate, Kia Dashtipour, Tughrul Arslan, Ahsan Adeel, Amir Hussain 0001, Tharmalingam Ratnarajah, Mathini Sellathurai
INTERSPEECH7
2023 Application for Real-time Audio-Visual Speech Enhancement
Mandar Gogate, Kia Dashtipour, Amir Hussain 0001
INTERSPEECH3
2023 Towards Two-point Neuron-inspired Energy-efficient Multimodal Open Master Hearing Aid
Adewale Adetomi, Khubaib Ahmed, Amir Hussain 0001, Tughrul Arslan, Ahsan Adeel
INTERSPEECH4
2023 Live Demonstration: Unlocking the Potential of Two-Point Neuronal Cells for Energy-Efficient Training of Deep Networks
abstract
This paper is the first live demonstration of the transformative computational potential of context-sensitive two-point layer 5 pyramidal cells (L5PCs). We will showcase a Multi-Processor System on Chip (MPSoC)-based implementation of a biologically plausible L5PC-driven deep neural network (DNN), termed multisensory cooperative computing (MCC). This will be shown to effectively process heterogeneous real-world audio-visual data consuming far less energy compared to state-of-the-art ‘point’ neuron-driven DNNs. Our approach opens new cross-disciplinary avenues for future on-chip DNN training implementations and posits a radical shift in current neuromorphic computing paradigms.
Ahsan Adeel, Adewale Adetomi, William A. Phillips, Khubaib Ahmed, Amir Hussain 0001, Tughrul Arslan
ISCAS6
2023 Wearable RF Sensing and Imaging System for Non-invasive Vascular Dementia Detection
abstract
Vascular dementia is the second most common form of dementia prevalent in old age groups, and is also one of the leading causes of mortality. Timely diagnosis and detection of vascular dementia is critical to avoid brain damage. Brain imaging is an essential tool for diagnosis and determines future treatment options available to the patient. Currently, Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), Computerized Tomography (CT) scan and carotid ultrasound are mainly used for brain imaging and vascular dementia diagnosis. However, these technologies are expensive, require extensive medical supervision and are not easily accessible. This can cause substantial delay in diagnosis and potentially lead to irreversible damage. This paper presents a first-of-its-kind, miniaturized octagonal monopole-patch antenna (OMPA) sensor for vascular dementia detection. It is designed to operate as part of a portable device and can effectively diagnose underlying causes like brain infarction, stroke and blood clots at the initial stage. The developed sensor designs are validated using microwave computational software and fabricated models are experimentally verified using artificial stroke and blood clot targets inside an artificial brain model. Simulated and measured reflection coefficient results are consistent, and target objects are detected successfully. The findings show that the prototype device is viable as an efficient, portable and low-cost alternate for vascular dementia detection.
Usman Anwar, Tughrul Arslan, Amir Hussain 0001, Peter Lomax
ISCAS3
2023 Live Demonstration: Cloud-based Audio-Visual Speech Enhancement in Multimodal Hearing-aids
abstract
Hearing loss is among the most serious public health problems, affecting as much as 20% of the worldwide population. Even cutting-edge multi-channel audio-only speech enhancement (SE) algorithms used in modern hearing aids face significant hurdles since they typically magnify noises while failing to boost speech understanding in crowded social environments. Recently, for the first time we proposed a novel integration of 5G cloud-radio access network, internet of things (IoT), and strong privacy algorithms to develop 5G IoT enabled hearing aid (HA) [1]. In this demonstration, we show the first-ever transceiver (PHY layer) model for cloud-based audio-visual (AV) SE, which meets the requirements for high data rate and low latency of forthcoming multi-modal HAs (such as Google glasses with integrated HAs). Even in highly noisy conditions like cafés, clubs, conferences, meetings, etc., the transceiver [2] transmits raw AV information from a hearing aid system to a cloud-based platform and obtains a clear signal. In Fig. 1, we illustrate an example of our cloud-based AV SE hearing aid demonstration. Herein, the left-side computer and Universal Software Radio Peripheral (USRP) x310 function as IoT systems (hearing aids), the right-side USRP serves as an access point or base station, and the right-side computer serves as a cloud-server for operating NN-based SE models. Please take note that the channel between the HA device and the cloud is defined as an uplink channel, whereas the channel between the access point (cloud) and the HA device is defined as a downlink channel. Given the time-varying sensitivity of the data received at HA devices, the uplink channel can therefore handle a variety of data rates. As a result, a customized long-term evolution (LTE)-based frame structure is developed for uplink transmission of data. It provides error-correction codes in the 1.4 MHz and 3 MHz bandwidths with a variety of modulations and code rates. Furthermore, the cloud access point simply supports a limited transmission rate because it only transmits audio data to the HA equipment. In order to support real-time AV SE, a modified frame structure for LTE with 1.4 MHz of bandwidth is developed. The AV SE algorithm receives cropped lip images of the target speaker and a noisy speech spectrogram, and it produces an ideal binary mask that lessens the noise-dominant regions while improving the speech-dominant areas. We use the depth-wise separable convolutions, reduced STFT window size of 32 ms, smaller STFT window shift of 8 ms, and 64 convolutions in the audio feature extraction layers of our Cochlea-Net [3] multi-modal AV SE neural network architecture to reduce processing latency. Furthermore, the visual feature extraction framework is employed. Our proposed architecture can handle streaming data frame-by-frame. Thus, the users will experience for the first time the real-world development of a physical layer transceiver that can perform AV SE in real-time under strict latency and data rate requirements. For this demonstration, we will bring two computers and two USRP x310 devices.
Abhijeet Bishnu, Ankit Gupta 0008, Mandar Gogate, Kia Dashtipour, Tughrul Arslan, Ahsan Adeel, Amir Hussain 0001, Mathini Sellathurai, Tharmalingam Ratnarajah
ISCAS7
2023 Live Demonstration: Real-time Multi-modal Hearing Assistive Technology Prototype
abstract
Hearing loss affects at least 1.5 billion people globally. The WHO estimates 83% of people who could benefit from hearing aids do not use them. Barriers to HA uptake are multifaceted but include ineffectiveness of current HA technology in noisy environments with multiple competing noise sources where human performance is known to be dependent upon input from both the aural and visual senses.
Mandar Gogate, Adeel Hussain, Kia Dashtipour, Amir Hussain 0001
ISCAS4
2023 Multimodal salient object detection via adversarial learning with collaborative generator
Zhengzheng Tu, Wenfang Yang, Kunpeng Wang 0005, Amir Hussain 0001, Bin Luo 0001, Chenglong Li 0002
Eng. Appl. Artif. Intell.4
2023 Diverse features discovery transformer for pedestrian attribute recognition
Aihua Zheng, Jiaxiang Wang 0001, Huaibo Huang, Ran He 0001, Amir Hussain 0001
Eng. Appl. Artif. Intell.6
2023 Randomized block-coordinate adaptive algorithms for nonconvex optimization problems
Yangfan Zhou 0004, Kaizhu Huang, Amir Hussain 0001, Xin Liu 0102
Eng. Appl. Artif. Intell.6
2023 A novel multimodal online news popularity prediction model based on ensemble learning
abstract
Abstract The prediction of news popularity is having substantial importance for the digital advertisement community in terms of selecting and engaging users. Traditional approaches are based on empirical data collected through surveys and applied statistical measures to prove a hypothesis. However, predicting news popularity based on statistical measures applied to past data is highly questionable. Therefore, in this article, we predict news popularity using machine learning classification models and deep residual neural network models. Articles are usually made up of textual content and in many cases, images are also used. Although it is evident that the appropriate amount of textual data is required to extract features and create models, image data is also helpful in gaining useful information. In this article, we present a novel multimodal online news popularity prediction model based on ensemble learning. This research work acts as a guide for extensive feature engineering, feature extraction, feature selection and effective modelling to create a robust news popularity Prediction Model. Three kinds of features—meta‐features, text features and image features are used to design an influential and robust model. The relative error performance measure Root Mean Squared logarithmic error (RMSLE) is used to quantify the popularity prediction error. Further, the RMSLE outcome shows 0.351 which is the lowest error value given by the proposed model. Further, the most important features are also sought out to show the dependence of the best‐fit model on text and image features.
Anuja Arora, Vikas Hassija, Shivam Bansal, Siddharth Yadav, Vinay Chamola, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.6
2023 Enhancing Arabic-text feature extraction utilizing label-semantic augmentation in few/zero-shot learning
abstract
Abstract A growing amount of research use pre‐trained language models to address few/zero‐shot text classification problems. Most of these studies neglect the semantic information hidden implicitly beneath the natural language names of class labels and develop a meta learner from the input texts solely. In this work, we demonstrate how label information can be utilized to extract enhanced feature representation of the input text from a Transformer‐based pre‐trained language model such as AraBERT. In addition, how this approach can improve performance when the data resources are scarce like in the Arabic language and the input text is short with little semantic information as is the case using tweets. The work also applies zero‐shot text classification to predict new classes with no training examples across different domains including sarcasm detection and sentiment analysis using the information in the last layer of a trained classifier in a transfer learning setting. Experiments show that our approach has a better performance for the few‐shot sentiment classification compared to baseline models and models trained without augmenting label information. Moreover, the zero‐shot implementation achieved an accuracy up to 0.874 in Arabic sarcasm detection from a model trained on a sentiment analysis task.
Seham Basabain, Erik Cambria, Khalid Al-Omar, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.4
2023 A Hurst-based diffusion model using time series characteristics for influence maximization in social networks
abstract
Abstract Online social networks have grown exponentially in the recent years while finding applications in real life like marketing, recommendation systems, and social awareness campaigns. An important research area in this field is Influence Maximization, which pertains to finding methods for maximizing the spread of information (influence) across an OSN. Existing works in IM widely use a pre‐defined edge propagation probability for node activation. Hurst exponent (H), which depicts the self‐similarity in the time series depicting a user's past interaction behaviour, has also been used as activation criteria. In this work, we propose a Time Series Characteristic based Hurst‐based Diffusion Model (TSC‐HDM), which calculates H based on the stationary or non‐stationary characteristic of the time series. TSC‐HDM selects a handful of seed nodes and activates a seed node's inactive successor only if H > 0.5. The proposed model has been tested on four real‐world OSN datasets. The results have been compared against four other IM models – Independent Cascade, Weighted Cascade, Trivalency, and Hurst‐based Influence Maximization. TSC‐HDM is found to have achieved as much as 590% higher expected influence spread as compared to the other models. Moreover, TSC‐HDM has attained 344% better average influence spread than other state‐of‐the‐art models namely LIR, A‐Greedy, LPIMA, Genetic Algorithm with Dynamic Probabilities, NeighborsRemove, DegreeDecrease, IGIM, IRR, and PHG.
Bhawna Saxena, Vikas Saxena, Nishit Anand, Vikas Hassija, Vinay Chamola, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.6
2023 Canonical cortical graph neural networks and its application for speech enhancement in audio-visual hearing aids
abstract
Despite the recent success of machine learning algorithms, most models face drawbacks when considering more complex tasks requiring interaction between different sources, such as multimodal input data and logical time sequences. On the other hand, the biological brain is highly sharpened in this sense, empowered to automatically manage and integrate such streams of information. In this context, this work draws inspiration from recent discoveries in brain cortical circuits to propose a more biologically plausible self-supervised machine learning approach. This combines multimodal information using intra-layer modulations together with Canonical Correlation Analysis, and a memory mechanism to keep track of temporal data, the overall approach termed Canonical Cortical Graph Neural networks. This is shown to outperform recent state-of-the-art models in terms of clean audio reconstruction and energy efficiency for a benchmark audio-visual speech dataset. The enhanced performance is demonstrated through a reduced and smother neuron firing rate distribution. suggesting that the proposed model is amenable for speech enhancement in future audio-visual hearing aid devices.
Leandro A. Passos Junior, João Paulo Papa, Amir Hussain 0001, Ahsan Adeel
Neurocomputing3
2023 Novel welch-transform based enhanced spectro-temporal analysis for cognitive microsleep detection using a single electrode EEG
Jash Shah, Amit Chougule, Vinay Chamola, Amir Hussain 0001
Neurocomputing4
2023 Adaptive Modulation Based on Nondata-Aided Error Vector Magnitude for Smart Systems in Smart Cities
abstract
A smart city involves big data transmission (BDT) between smart systems, which increases queue delays and leads to difficulty in enhancing the spectral efficiency. Adaptive modulation is an effective technique for enhancing data transmission rates in smart systems. However, traditional adaptive modulation approaches are not suitable for BDT in smart systems because the delays caused by the large amount of transmitted data lead to difficulty in evaluating the channel quality. In this paper, we propose a nondata-aided error vector (NDA-EVM) that can be employed in adaptive modulation over wireless channels. The proposed NDA-EVM can be used to evaluate the channel quality and symbol error rate (SER), which reflect the quality of service (QoS) of the system. We formulated the relationship between the NDA-EVM and SER, which provides a basis for designing adaptive modulation techniques for smart systems. To address the low average spectrum efficiency (ASE) caused by BDT queue delays, an adaptive modulation strategy based on the finite-state Markov chain (FSMC) of the NDA-EVM (i.e., NDA-EVM-AM) was designed. This method simplifies the adaptive modulation algorithm for smart systems to search for the optimal transfer probability in the FSMC matrix based on two typical states: the resident state and transient state. Moreover, we proposed an analytical procedure to describe queuing behavior to analyze the performance of the NDA-EVM-AM algorithm for smart systems in smart cities. The performance is compared with that of a conventional adaptive modulation algorithm through simulations. The results show that compared with traditional adaptive modulation, NDA-EVM-AM obtains a lower packet loss rate and higher spectral efficiency for smart systems.
Fan Yang 0031, Jie Huang 0018, Arpit Bhardwaj, Amir Hussain 0001, Ahmed A. Abd El-Latif 0001, Keping Yu
IEEE Internet Things J.4
2023 A novel approach of many-objective particle swarm optimization with cooperative agents based on an inverted generational distance indicator
Najwa Kouka, Fatma BenSaid, Raja Fdhila, Rahma Fourati, Amir Hussain 0001, Adel M. Alimi
Inf. Sci.5
2023 A Digital Twin-Assisted Intelligent Partial Offloading Approach for Vehicular Edge Computing
abstract
Vehicle Edge Computing (VEC) is a promising paradigm that exposes Mobile Edge Computing (MEC) to road scenarios. In VEC, task offloading can enable vehicles to offload the computing tasks to nearby Roadside Units (RSUs) that deploy computing capabilities. However, the highly dynamic network topology, strict low-delay constraints, and massive data of tasks of VEC pose significant challenges for implementing efficient offloading. Digital Twin-based VEC is emerging as a promising solution that enables real-time monitoring of the state of the VEC network through mapping and interaction between the physical and virtual worlds, thus assisting in making sound offload decisions in the physical world. Thus, this paper proposes an intelligent partial offloading scheme, namely, Digital Twin-Assisted Intelligent Partial Offloading (IGNITE). First, to find the optimal offloading space in advance, we combine the improved clustering algorithm with the Digital Twin (DT) technique, in which unreasonable decisions can be avoided by reducing the size of the decision space. Second, to reduce the overall cost of the system, Deep Reinforcement Learning (DRL) algorithm is employed to train the offloading strategy, allowing for automatic optimization of computational delay and vehicle service price. To improve the efficiency of cooperation between digital and physical spaces, a feedback mechanism is established. It can adjust the parameters of the clustering algorithm based on the final offloading results in this clustering. To the best of our knowledge, this is the first study on DT-assisted vehicle offloading that proposes a feedback mechanism, forming a complete closed loop as prediction-offloading-feedback. Extensive experiments demonstrate that IGNITE has significant advantages in terms of total system computational cost, total computational delay, and offloading success rate compared with its counterparts.
Liang Zhao 0004, Zijia Zhao, Enchao Zhang, Ammar Hawbani, Ahmed Yassin Al-Dubai, Zhiyuan Tan 0001, Amir Hussain 0001
IEEE J. Sel. Areas Commun.7
2023 An Incremental SAR Target Recognition Framework via Memory-Augmented Weight Alignment and Enhancement Discrimination
abstract
Synthetic Aperture Radar Automatic Target Recognition (SAR ATR) is one of the most important research directions in SAR image interpretation. While much existing research into SAR ATR has focused on deep learning technology, an equally important yet underexplored problem is its deployment in incremental learning scenarios. This letter proposes a new benchmark approach, termed Memory augmented weights alignment and Enhancement Discrimination Incremental Learning (MEDIL) algorithm to address this issue. Firstly, the attention mechanism is employed as part of the benchmark. Next, we discuss the problem of height deviation of weights at the fully connected layer and design a more suitable alignment of weights by guiding the memory module for contextual data processing. In addition, we leverage the incremental progressive sampling strategy to alleviate the imbalance between old and new classes during the training period. Finally, we propose to enhance the distinction among various classes with an angular penalty loss function to ensure the diversity of incremental instances. The proposed method is evaluated on MSTAR and OpenSARShip under different experimental settings. Experimental results demonstrate that our proposed approach can effectively solve catastrophic forgetting in SAR multiclass recognition problems.
Fei Gao 0005, Jun Wang 0041, Amir Hussain 0001, Huiyu Zhou 0001
IEEE Geosci. Remote. Sens. Lett.4
2023 A novel oversampling and feature selection hybrid algorithm for imbalanced data classification
Kuanching Li, Erfu Yang, Qingguo Zhou, Lihong Han, Amir Hussain 0001, Mingjiang Cai
Multim. Tools Appl.6
2023 An Enhanced Binary Particle Swarm Optimization (E-BPSO) algorithm for service placement in hybrid cloud platforms
Wissem Abbes, Zied Kechaou, Amir Hussain 0001, Abdulrahman M. Qahtani, Omar Almutiry, Habib Dhahri, Adel M. Alimi
Neural Comput. Appl.3
2023 WETM: A word embedding-based topic model with modified collapsed Gibbs sampling for short text
abstract
Short texts are a common source of knowledge, and the extraction of such valuable information is beneficial for several purposes. Traditional topic models are incapable of analyzing the internal structural information of topics. They are mostly based on the co-occurrence of words at the document level and are often unable to extract semantically relevant topics from short text datasets due to their limited length. Although some traditional topic models are sensitive to word order due to the strong sparsity of data, they do not perform well on short texts. In this paper, we propose a novel word embedding-based topic model (WETM) for short text documents to discover the structural information of topics and words and eliminate the sparsity problem. Moreover, a modified collapsed Gibbs sampling algorithm is proposed to strengthen the semantic coherence of topics in short texts. WETM extracts semantically coherent topics from short texts and finds relationships between words. Extensive experimental results on two real-world datasets show that WETM achieves better topic quality, topic coherence, classification, and clustering results. WETM also requires less execution time compared to traditional topic models.
Junaid Rashid, Jungeun Kim, Amir Hussain 0001, Usman Naseem
Pattern Recognit. Lett.3
2023 Guest Editorial Neurosymbolic AI for Sentiment Analysis
abstract
Neural network-based methods, especially deep learning, have been a burgeoning area in AI research and have been successful in tackling the expanding data volume as we move into a digital age. Today, the neural network-based methods are not only used for low-level cognitive tasks, such as recognizing objects and spotting keywords, but they have also been deployed in various industrial information systems to assist high-level decision-making. In natural language processing, there have been two milestones for the past decade: one is word2vec [1], a group of neural models that learn word embeddings (vector representations of words) from large datasets; and one is the most recent GPT-based models [2], which combine reinforcement learning with a generative transformer in order to enable multi-round end-to-end conversations. While producing highly accurate predictions on datasets and generating human-like utterances, those neural network-based artifacts provide little understanding of the internal features and representations of the data. Many problems and concerns subsequently emerge from this black-box issue. Because some of the problems and concerns are also relevant in the context of sentiment analysis.
Frank Z. Xing, Björn W. Schuller, Iti Chaturvedi, Erik Cambria, Amir Hussain 0001
IEEE Trans. Affect. Comput.5
2023 Recognizing British Sign Language Using Deep Learning: A Contactless and Privacy-Preserving Approach
abstract
Sign language is utilized by deaf-mute to communicate through hand movements, body postures, and facial emotions. The motions in sign language comprise a range of distinct hand and finger articulations that are occasionally synchronized with the head, face, and body. Automatic sign language recognition (SLR) is a highly challenging area and still remains in its infancy compared with speech recognition after almost three decades of research. Current wearable and vision-based systems for SLR are intrusive and suffer from the limitations of ambient lighting and privacy concerns. To the best of our knowledge, our work proposes the first contactless British sign language (BSL) recognition system using radar and deep learning (DL) algorithms. Our proposed system extracts the 2-D spatiotemporal features from the radar data and applies the state-of-the-art DL models to classify spatiotemporal features from BSL signs to different verbs and emotions, such as Help, Drink, Eat, Happy, Hate, and Sad. We collected and annotated a large-scale benchmark BSL dataset covering 15 different types of BSL signs. Our proposed system demonstrates highest classification performance with a multiclass accuracy of up to 90.07% at a distance of 141 cm from the subject using the VGGNet model.
Hira Hameed, Muhammad Usman 0003, Ahsen Tahir, Kashif Ahmad, Amir Hussain 0001, Muhammad Ali Imran 0001, Qammer H. Abbasi
IEEE Trans. Comput. Soc. Syst.5
2023 Guest Editorial Special Issue on Behavioral Modeling, Learning, and Adaptation in Cyber-Physical-Social Intelligence
abstract
The integration of artificial intelligence (AI) with cyber–physical–social systems (CPSS) creates new research opportunities and challenges with major societal implications. The behavioral and cognitive enhancement of intelligent systems promotes a productive and creative partnership and collaboration between humans and machines. Advancements in these areas enable adaptability, scalability, resiliency, safety, security, and usability that expand the horizons of CPSS.
Ying Tang 0001, Jiacun Wang 0001, Hui Yu 0001, Giancarlo Fortino, Fei-Yue Wang 0001, Amir Hussain 0001
IEEE Trans. Comput. Soc. Syst.6
2023 FastAdaBelief: Improving Convergence Rate for Belief-Based Adaptive Optimizers by Exploiting Strong Convexity
abstract
AdaBelief, one of the current best optimizers, demonstrates superior generalization ability over the popular Adam algorithm by viewing the exponential moving average of observed gradients. AdaBelief is theoretically appealing in which it has a data-dependent O(√T) regret bound when objective functions are convex, where T is a time horizon. It remains, however, an open problem whether the convergence rate can be further improved without sacrificing its generalization ability. To this end, we make the first attempt in this work and design a novel optimization algorithm called FastAdaBelief that aims to exploit its strong convexity in order to achieve an even faster convergence rate. In particular, by adjusting the step size that better considers strong convexity and prevents fluctuation, our proposed FastAdaBelief demonstrates excellent generalization ability and superior convergence. As an important theoretical contribution, we prove that FastAdaBelief attains a data-dependent O(logT) regret bound, which is substantially lower than AdaBelief in strongly convex cases. On the empirical side, we validate our theoretical analysis with extensive experiments in scenarios of strong convexity and nonconvexity using three popular baseline models. Experimental results are very encouraging: FastAdaBelief converges the quickest in comparison to all mainstream algorithms while maintaining an excellent generalization ability, in cases of both strong convexity or nonconvexity. FastAdaBelief is, thus, posited as a new benchmark model for the research community.
Yangfan Zhou 0004, Kaizhu Huang, Amir Hussain 0001, Xin Liu 0102
IEEE Trans. Neural Networks Learn. Syst.5
2023 A New Class of Efficient Adaptive Filters for Online Nonlinear Modeling
abstract
Nonlinear models are known to provide excellent performance in real-world applications that often operate in nonideal conditions. However, such applications often require online processing to be performed with limited computational resources. To address this problem, we propose a new class of efficient nonlinear models for online applications. The proposed algorithms are based on linear-in-the-parameters (LIPs) nonlinear filters using functional link expansions. In order to make this class of functional link adaptive filters (FLAFs) efficient, we propose low-complexity expansions and frequency-domain adaptation of the parameters. Among this family of algorithms, we also define the partitioned-block frequency-domain FLAF (FD-FLAF), whose implementation is particularly suitable for online nonlinear modeling problems. We assess and compare FD-FLAFs with different expansions providing the best possible tradeoff between performance and computational complexity. Experimental results prove that the proposed algorithms can be considered as an efficient and effective solution for online applications, such as the acoustic echo cancellation, even in the presence of adverse nonlinear conditions and with limited availability of computational resources.
Danilo Comminiello, Alireza Nezamdoust, Simone Scardapane, Michele Scarpiniti, Amir Hussain 0001, Aurelio Uncini
IEEE Trans. Syst. Man Cybern. Syst.5
2022 A Novel Frame Structure for Cloud-Based Audio-Visual Speech Enhancement in Multimodal Hearing-aids
abstract
In this paper, we design a first of its kind transceiver (PHY layer) prototype for cloud-based audio-visual (AV) speech enhancement (SE) complying with high data rate and low latency requirements of future multimodal hearing assistive technology. The innovative design needs to meet multiple challenging constraints including up/down link communications, delay of transmission and signal processing, and real-time AV SE models processing. The transceiver includes device detection, frame detection, frequency offset estimation, and channel estimation capabilities. We develop both uplink (hearing aid to the cloud) and downlink (cloud to hearing aid) frame structures based on the data rate and latency requirements. Due to the varying nature of uplink information (audio and lip-reading), the uplink channel supports multiple data rate frame structure, while the downlink channel has a fixed data rate frame structure. In addition, we evaluate the latency of different PHY layer blocks of the transceiver for developed frame structures using LabVIEW NXG. This can be used with software defined radio (such as Universal Software Radio Peripheral) for real-time demonstration scenarios.
Abhijeet Bishnu, Ankit Gupta 0008, Mandar Gogate, Kia Dashtipour, Ahsan Adeel, Amir Hussain 0001, Mathini Sellathurai, Tharmalingam Ratnarajah
HealthCom6
2022 AVSE Challenge: Audio-Visual Speech Enhancement Challenge
abstract
Audio-visual speech enhancement is the task of improving the quality of a speech signal when video of the speaker is available. It opens-up the opportunity of improving speech intelligibility in adverse listening scenarios that are currently too challenging for audio-only speech enhancement models. The Audio-Visual Speech Enhancement (AVSE) challenge aims to set the first benchmark in this area. We provide participants with datasets and scripts to test their audio-visual speech enhancement models under a common framework for both training and evaluation. The data is derived from real-world videos, and comprises noisy mixes, in which audio from target speaker is mixed with either a competing speaker or a noise signal. The submitted systems are evaluated by conducting AV intelligibility tests involving human participants. We expect this challenge to be a platform for advancing the field of audio-visual speech-enhancement and to provide further insight about the scope and limitations of current AV speech enhancement approaches.
Andrea Lorena Aldana Blanco, Cassia Valentini-Botinhao, Ondrej Klejch, Mandar Gogate, Kia Dashtipour, Amir Hussain 0001, Peter Bell 0001
SLT6
2022 A novel multiple kernel fuzzy topic modeling technique for biomedical data
abstract
BACKGROUND: Text mining in the biomedical field has received much attention and regarded as the important research area since a lot of biomedical data is in text format. Topic modeling is one of the popular methods among text mining techniques used to discover hidden semantic structures, so called topics. However, discovering topics from biomedical data is a challenging task due to the sparsity, redundancy, and unstructured format. METHODS: In this paper, we proposed a novel multiple kernel fuzzy topic modeling (MKFTM) technique using fusion probabilistic inverse document frequency and multiple kernel fuzzy c-means clustering algorithm for biomedical text mining. In detail, the proposed fusion probabilistic inverse document frequency method is used to estimate the weights of global terms while MKFTM generates frequencies of local and global terms with bag-of-words. In addition, the principal component analysis is applied to eliminate higher-order negative effects for term weights. RESULTS: Extensive experiments are conducted on six biomedical datasets. MKFTM achieved the highest classification accuracy 99.04%, 99.62%, 99.69%, 99.61% in the Muchmore Springer dataset and 94.10%, 89.45%, 92.91%, 90.35% in the Ohsumed dataset. The CH index value of MKFTM is higher, which shows that its clustering performance is better than state-of-the-art topic models. CONCLUSION: We have confirmed from results that proposed MKFTM approach is very efficient to handles to sparsity and redundancy problem in biomedical text documents. MKFTM discovers semantically relevant topics with high accuracy for biomedical documents. Its gives better results for classification and clustering in biomedical documents. MKFTM is a new approach to topic modeling, which has the flexibility to work with a variety of clustering methods.
Junaid Rashid, Jungeun Kim, Amir Hussain 0001, Usman Naseem, Sapna Juneja
BMC Bioinform.3
2022 Novel single and multi-layer echo-state recurrent autoencoders for representation learning
Naima Chouikhi, Boudour Ammar, Amir Hussain 0001, Adel M. Alimi
Eng. Appl. Artif. Intell.3
2022 A fuzzy-enhanced deep learning approach for early detection of Covid-19 pneumonia from portable chest X-ray images
Cosimo Ieracitano, Nadia Mammone, Mario Versaci, Giuseppe Varone, Abder-Rahman Ali, Antonio Armentano, Grazia Calabrese, Anna Ferrarelli, Lorena Turano, Carmela Tebala, Zain U. Hussain, Zakariya Sheikh, Aziz Sheikh, Giuseppe Sceni, Amir Hussain 0001, Francesco Carlo Morabito
Neurocomputing15
2022 A novel explainable machine learning approach for EEG-based brain-computer interface systems
Cosimo Ieracitano, Nadia Mammone, Amir Hussain 0001, Francesco Carlo Morabito
Neural Comput. Appl.3
2022 Improving aspect-level sentiment analysis with aspect extraction
Navonil Majumder, Rishabh Bhardwaj, Soujanya Poria, Alexander F. Gelbukh, Amir Hussain 0001
Neural Comput. Appl.5
2022 Does semantics aid syntax? An empirical study on named entity recognition and classification
Xiaoshi Zhong, Erik Cambria, Amir Hussain 0001
Neural Comput. Appl.3
2022 Cloud based scalable object recognition from video streams using orientation fusion and convolutional neural networks
Muhammad Usman Yaseen, Ashiq Anjum, Giancarlo Fortino, Antonio Liotta, Amir Hussain 0001
Pattern Recognit.5
2022 Ellipse Encoding for Arbitrary-Oriented SAR Ship Detection Based on Dynamic Key Points
abstract
In recent years, there has been growing interest in developing oriented bounding-box (OBB) based deep learning approaches to detect arbitrary-oriented ship targets in synthetic aperture radar (SAR) images. However, most existing OBB-based detection methods suffer from boundary discontinuity problems for bounding box angle prediction and key point regression challenges. In this paper, we present a novel OBB-based detection algorithm that utilizes ellipse encoding to effectively exploit the geometric and scattering properties of ship targets. Specifically, the ship contour is fitted by an OBB inscribed ellipse that is encoded as a set of distances between dynamic key points on the bow and target center. By combining the bow angle interval and the decoding process, the negative impact of the boundary discontinuity problem is avoided. In addition, we propose an elliptical Gaussian distribution heatmap and a pooling strategy termed double peaks max-pooling (DPM), to deal with the challenge of separating densely distributed ships in inshore scenes. The former can enhance the heatmap’s ship-side score gap between neighboring ship targets, while the latter can solve the problem of target center responses being suppressed after max-pooling. Simulation experiments conducted on the benchmark Rotating SAR Ship Detection Dataset (RSSDD) and Rotated Ship Detection Dataset in SAR Images (RSDD-SAR) demonstrate the superior performance of our method for ship target detection compared to several state-of-the-art OBB-based algorithms. Ablation experiments show that elliptical Gaussian distribution heatmap and DPM can further improve the inshore detection performance.
Fei Gao 0005, Yiyang Huo, Jinping Sun, Amir Hussain 0001, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 A Novel 3D Unsupervised Domain Adaptation Framework for Cross-Modality Medical Image Segmentation
abstract
We consider the problem of volumetric (3D) unsupervised domain adaptation (UDA) in cross-modality medical image segmentation, aiming to perform segmentation on the unannotated target domain (e.g. MRI) with the help of labeled source domain (e.g. CT). Previous UDA methods in medical image analysis usually suffer from two challenges: 1) they focus on processing and analyzing data at 2D level only, thus missing semantic information from the depth level; 2) one-to-one mapping is adopted during the style-transfer process, leading to insufficient alignment in the target domain. Different from the existing methods, in our work, we conduct a first of its kind investigation on multi-style image translation for complete image alignment to alleviate the domain shift problem, and also introduce 3D segmentation in domain adaptation tasks to maintain semantic consistency at the depth level. In particular, we develop an unsupervised domain adaptation framework incorporating a novel quartet self-attention module to efficiently enhance relationships between widely separated features in spatial regions on a higher dimension, leading to a substantial improvement in segmentation accuracy in the unlabeled target domain. In two challenging cross-modality tasks, specifically brain structures and multi-organ abdominal segmentation, our model is shown to outperform current state-of-the-art methods by a significant margin, demonstrating its potential as a benchmark resource for the biomedical and health informatics research community.
Zixian Su, Kaizhu Huang, Xi Yang 0008, Jie Sun 0024, Amir Hussain 0001, Frans Coenen
IEEE J. Biomed. Health Informatics6
2022 Exploiting Attention-Consistency Loss For Spatial-Temporal Stream Action Recognition
abstract
Currently, many action recognition methods mostly consider the information from spatial streams. We propose a new perspective inspired by the human visual system to combine both spatial and temporal streams to measure their attention consistency. Specifically, a branch-independent convolutional neural network (CNN) based algorithm is developed with a novel attention-consistency loss metric, enabling the temporal stream to concentrate on consistent discriminative regions with the spatial stream in the same period. The consistency loss is further combined with the cross-entropy loss to enhance the visual attention consistency. We evaluate the proposed method for action recognition on two benchmark datasets: Kinetics400 and UCF101. Despite its apparent simplicity, our proposed framework with the attention consistency achieves better performance than most of the two-stream networks, i.e., 75.7% top-1 accuracy on Kinetics400 and 95.7% on UCF101, while reducing 7.1% computational cost compared with our baseline. Particularly, our proposed method can attain remarkable improvements on complex action classes, showing that our proposed network can act as a potential benchmark to handle complicated scenarios in industry 4.0 applications.
Xiao-Bo Jin, Qiufeng Wang 0001, Amir Hussain 0001, Kaizhu Huang
ACM Trans. Multim. Comput. Commun. Appl.4
2021 Novel Artificial Immune Networks-based optimization of shallow machine learning (ML) classifiers
Summrina Kanwal, Amir Hussain 0001, Kaizhu Huang
Expert Syst. Appl.2
2021 A Hybrid-Domain Deep Learning-Based BCI For Discriminating Hand Motion Planning From EEG Sources
abstract
In this paper, a hybrid-domain deep learning (DL)-based neural system is proposed to decode hand movement preparation phases from electroencephalographic (EEG) recordings. The system exploits information extracted from the temporal-domain and time-frequency-domain, as part of a hybrid strategy, to discriminate the temporal windows (i.e. EEG epochs) preceding hand sub-movements (open/close) and the resting state. To this end, for each EEG epoch, the associated cortical source signals in the motor cortex and the corresponding time-frequency (TF) maps are estimated via beamforming and Continuous Wavelet Transform (CWT), respectively. Two Convolutional Neural Networks (CNNs) are designed: specifically, the first CNN is trained over a dataset of temporal (T) data (i.e. EEG sources), and is referred to as T-CNN; the second CNN is trained over a dataset of TF data (i.e. TF-maps of EEG sources), and is referred to as TF-CNN. Two sets of features denoted as T-features and TF-features, extracted from T-CNN and TF-CNN, respectively, are concatenated in a single features vector (denoted as TTF-features vector) which is used as input to a standard multi-layer perceptron for classification purposes. Experimental results show a significant performance improvement of our proposed hybrid-domain DL approach as compared to temporal-only and time-frequency-only-based benchmark approaches, achieving an average accuracy of [Formula: see text]%.
Cosimo Ieracitano, Francesco Carlo Morabito, Amir Hussain 0001, Nadia Mammone
Int. J. Neural Syst.3
2021 Leveraging label hierarchy using transfer and multi-task learning: A case study on patent classification
Segun Taofeek Aroyehun, Jason Angel, Navonil Majumder, Alexander F. Gelbukh, Amir Hussain 0001
Neurocomputing5
2021 Persuasive dialogue understanding: The baselines and negative results
Hui Chen 0023, Deepanway Ghosal, Navonil Majumder, Amir Hussain 0001, Soujanya Poria
Neurocomputing4
2021 Conceptual text region network: Cognition-inspired accurate scene text detection
Chenwei Cui 0001, Zhiyuan Tan 0001, Amir Hussain 0001
Neurocomputing4
2021 A novel context-aware multimodal framework for persian sentiment analysis
Kia Dashtipour, Mandar Gogate, Erik Cambria, Amir Hussain 0001
Neurocomputing4
2021 Lane-DeepLab: Lane semantic segmentation in automatic driving scenarios for high-definition maps
Fengling Jiang, Jing Yang 0041, Mandar Gogate, Kia Dashtipour, Amir Hussain 0001
Neurocomputing7
2021 A novel domain activation mapping-guided network (DA-GNT) for visual tracking
Zhengzheng Tu, Ajian Zhou, Chuang Gan 0003, Bo Jiang 0002, Amir Hussain 0001, Bin Luo 0001
Neurocomputing5
2021 Coarse-grained generalized zero-shot learning with efficient self-focus mechanism
Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas, Amir Hussain 0001
Neurocomputing5
2021 A novel few-shot learning method for synthetic aperture radar image recognition
Fei Gao 0005, Qingxu Xiong, Jinping Sun, Amir Hussain 0001, Huiyu Zhou 0001
Neurocomputing5
2021 Advances in machine translation for sign language: approaches, limitations, and challenges
Uzma Farooq, Mohd Shafry Mohd Rahim, Nabeel Sabir, Amir Hussain 0001, Adnan Abid
Neural Comput. Appl.4
2021 Improving generative adversarial networks with simple latent distributions
Shufei Zhang, Kaizhu Huang, Zhuang Qian, Rui Zhang 0012, Amir Hussain 0001
Neural Comput. Appl.5
2020 Dynamic Multi Objective Particle Swarm optimization with Cooperative Agents
abstract
Dynamic Multi-objective optimization problems (DMOPs) involve multiple objectives, constraints, and parameters that may change over time. For solving such types of problems, the conventional particle swarm optimization algorithm should not only be able to evolve near-optimal and diverse optimal solutions but also continually track the time-changing environment. To address these challenges, we propose a novel dynamic multiobjective particle swarm optimization approach with cooperative agents. In this strategy, the ability of tracking based on particle memory and multiple populations with sharing knowledge are combined to deal with environmental changing. If a change is detected, solutions with no improvement are re-evaluated and the worst solutions are replaced with the newly generated one. In addition, a movement strategy based on the sharing of the best knowledge is introduced to promote the population diversity. Experiments on several optimization problems are carried out to prove the performance of the proposed algorithm and the statistical results show that the proposed algorithm performs well with DMOPs.
Najwa Kouka, Raja Fdhila, Amir Hussain 0001, Adel M. Alimi
CEC3
2020 Deep Neural Network Driven Binaural Audio Visual Speech Separation
abstract
The central auditory pathway exploits the auditory signals and visual information sent by both ears and eyes to segregate speech from multiple competing noise sources and help disambiguate phonological ambiguity. In this study, inspired from this unique human ability, we present a deep neural network (DNN) that ingest the binaural sounds received at the two ears as well as the visual frames to selectively suppress the competing noise sources individually at both ears. The model exploits the noisy binaural cues and noise robust visual cues to improve speech intelligibility. The comparative simulation results in terms of objective metrics such as PESQ, STOI, SI-SDR and DBSTOI demonstrate significant performance improvement of the proposed audio-visual (AV) DNN as compared to the audio-only (A-only) variant of the proposed model. Finally, subjective listening tests with the real noisy AV ASPIRE corpus shows the superiority of the proposed AV DNN as compared to state-of-the-art approaches.
Mandar Gogate, Kia Dashtipour, Peter Bell 0001, Amir Hussain 0001
IJCNN4
2020 A Convolutional Neural Network based self-learning approach for classifying neurodegenerative states from EEG signals in dementia
abstract
In this paper, a novel deep learning based approach is proposed for the automatic classification of Electroencephalographic (EEG) signals of subjects diagnosed with the dementia of Alzheimer's disease (AD), Mild Cognitive Impairment (MCI) and Healthy Control (HC). Specifically, a custom Convolutional Neural Network (CNN) is designed to receive as input AD/MCI/HC EEG segments (epochs) of the same temporal width, and perform 2-way classification tasks: AD vs. HC, AD vs. MCI, MCI vs. HC. Our proposed architecture, termed EEG-CNN, is shown to exhibit remarkable abilities to self-learn relevant features directly from the EEG traces, avoiding the need for hand-crafted feature extraction engineering. Comparative experimental results demonstrate the promising performance of EEG-CNN, which is based on an analysis of the EEG time series only, reporting accuracies of 85.78 ± 2.18%, 69.03 ± 1.33%, 85.34 ± 1.86% in AD vs. HC, AD vs. MCI and MCI vs. HC classifications, respectively.
Cosimo Ieracitano, Nadia Mammone, Amir Hussain 0001, Francesco Carlo Morabito
IJCNN3
2020 Visual Speech In Real Noisy Environments (VISION): A Novel Benchmark Dataset and Deep Learning-Based Baseline System
abstract
In this paper, we present VIsual Speech In real nOisy eNvironments (VISION), a first of its kind audio-visual (AV) corpus comprising 2500 utterances from 209 speakers, recorded in real noisy environments including social gatherings, streets, cafeterias and restaurants. While a number of speech enhancement frameworks have been proposed in the literature that exploit AV cues, there are no visual speech corpora recorded in real environments with a sufficient variety of speakers, to enable evaluation of AV frameworks' generalisation capability in a wide range of background visual and acoustic noises. The main purpose of our AV corpus is to foster research in the area of AV signal processing and to provide a benchmark corpus that can be used for reliable evaluation of AV speech enhancement systems in everyday noisy settings. In addition, we present a baseline deep neural network (DNN) based spectral mask estimation model for speech enhancement. Comparative simulation results with subjective listening tests demonstrate significant performance improvement of the baseline DNN compared to state-of-the-art speech enhancement approaches.
Mandar Gogate, Kia Dashtipour, Amir Hussain 0001
INTERSPEECH3
2020 Inductive Generalized Zero-Shot Learning with Adversarial Relation Network
Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas, Amir Hussain 0001
ECML/PKDD (2)5
2020 A hybrid Persian sentiment analysis framework: Integrating dependency grammar based rules and deep neural networks
Kia Dashtipour, Mandar Gogate, Jingpeng Li 0001, Fengling Jiang, Amir Hussain 0001
Neurocomputing6
2020 A novel statistical analysis and autoencoder driven intelligent intrusion detection approach
Cosimo Ieracitano, Ahsan Adeel, Francesco Carlo Morabito, Amir Hussain 0001
Neurocomputing4
2020 A novel biologically-inspired target detection method based on saliency analysis for synthetic aperture radar (SAR) imagery
Fei Ma 0001, Fei Gao 0005, Jun Wang 0041, Amir Hussain 0001, Huiyu Zhou 0001
Neurocomputing4
2020 A non-parametric softmax for improving neural attention in time-series forecasting
Simone Totaro, Amir Hussain 0001, Simone Scardapane
Neurocomputing2
2020 Improving deep neural network performance by integrating kernelized Min-Max objective
Qiufeng Wang 0001, Rui Zhang 0012, Amir Hussain 0001, Kaizhu Huang
Neurocomputing4
2020 Novel deep neural network based pattern field classification architectures
Kaizhu Huang, Shufei Zhang, Rui Zhang 0012, Amir Hussain 0001
Neural Networks4
2020 A novel multi-modal machine learning based approach for automatic classification of EEG recordings in dementia
Cosimo Ieracitano, Nadia Mammone, Amir Hussain 0001, Francesco Carlo Morabito
Neural Networks3
2020 Encoding primitives generation policy learning for robotic arm to overcome catastrophic forgetting in sequential multi-tasks learning
Fangzhou Xiong, Zhiyong Liu 0001, Kaizhu Huang, Xu Yang 0004, Hong Qiao, Amir Hussain 0001
Neural Networks6
2020 A Novel Intelligent Computational Approach to Model Epidemiological Trends and Assess the Impact of Non-Pharmacological Interventions for COVID-19
abstract
The novel coronavirus disease 2019 (COVID-19) pandemic has led to a worldwide crisis in public health. It is crucial we understand the epidemiological trends and impact of non-pharmacological interventions (NPIs), such as lockdowns for effective management of the disease and control of its spread. We develop and validate a novel intelligent computational model to predict epidemiological trends of COVID-19, with the model parameters enabling an evaluation of the impact of NPIs. By representing the number of daily confirmed cases (NDCC) as a time-series, we assume that, with or without NPIs, the pattern of the pandemic satisfies a series of Gaussian distributions according to the central limit theorem. The underlying pandemic trend is first extracted using a singular spectral analysis (SSA) technique, which decomposes the NDCC time series into the sum of a small number of independent and interpretable components such as a slow varying trend, oscillatory components and structureless noise. We then use a mixture of Gaussian fitting (GF) to derive a novel predictive model for the SSA extracted NDCC incidence trend, with the overall model termed SSA-GF. Our proposed model is shown to accurately predict the NDCC trend, peak daily cases, the length of the pandemic period, the total confirmed cases and the associated dates of the turning points on the cumulated NDCC curve. Further, the three key model parameters, specifically, the amplitude (alpha), mean (mu), and standard deviation (sigma) are linked to the underlying pandemic patterns, and enable a directly interpretable evaluation of the impact of NPIs, such as strict lockdowns and travel restrictions. The predictive model is validated using available data from China and South Korea, and new predictions are made, partially requiring future validation, for the cases of Italy, Spain, the UK and the USA. Comparative results demonstrate that the introduction of consistent control measures across countries can lead to development of similar parametric models, reflected in particular by relative variations in their underlying sigma, alpha and mu values. The paper concludes with a number of open questions and outlines future research directions.
Jinchang Ren, Yijun Yan, Huimin Zhao 0001, Ping Ma 0002, Jaime Zabalza, Zain U. Hussain, Shaoming Luo, Sophia Zhao, Aziz Sheikh, Amir Hussain 0001, Huakang Li
IEEE J. Biomed. Health Informatics11
2019 Generalized Adversarial Training in Riemannian Space
abstract
Adversarial examples, referred to as augmented data points generated by imperceptible perturbations of input samples, have recently drawn much attention. Well-crafted adversarial examples may even mislead state-of-the-art deep neural network (DNN) models to make wrong predictions easily. To alleviate this problem, many studies have focused on investigating how adversarial examples can be generated and/or effectively handled. All existing works tackle this problem in the Euclidean space. In this paper, we extend the learning of adversarial examples to the more general Riemannian space over DNNs. The proposed work is important in that (1) it is a generalized learning methodology since Riemmanian space will be degraded to the Euclidean space in a special case; (2) it is the first work to tackle the adversarial example problem tractably through the perspective of Riemannian geometry; (3) from the perspective of geometry, our method leads to the steepest direction of the loss function, by considering the second order information of the loss function. We also provide a theoretical study showing that our proposed method can truly find the descent direction for the loss function, with a comparable computational time against traditional adversarial methods. Finally, the proposed framework demonstrates superior performance over traditional counterpart methods, using benchmark data including MNIST, CIFAR-10 and SVHN.
Shufei Zhang, Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICDM4
2019 A Time-Frequency based Machine Learning System for Brain States Classification via EEG Signal Processing
abstract
In the last decades, the use of Machine Learning (ML) algorithms have been widely employed to aid clinicians in the difficult diagnosis of neurological disorders, such as Alzheimer's disease (AD). In this context, here, a data-driven ML system for classifying Electroencephalographic (EEG) segments (i.e. epochs) of patients affected by AD, Mild Cognitive Impairment (MCI) and Healthy Control (HC) individuals, is introduced. Specifically, the proposed ML system consists of evaluating the average Time-Frequency Map (aTFM) related to a 19-channels EEG epoch and extracting some statistical coefficients (i.e. mean, standard deviation, skewness, kurtosis and entropy) from the main five conventional EEG sub-bands (or EEG-rhythms: delta, theta, alpha1, alpha2, beta). Afterwards, the time-frequency features vector is fed into an Autoeconder (AE), a Multi-Layer Perceptron (MLP), a Logistic Regression (LR) and a Support Vector Machine (SVM) based classifier to perform the 2-ways EEG epoch-classification tasks: AD vs HC and AD vs MCI. The performances of the proposed approach have been evaluated on a dataset of 189 EEG signals (63 AD, 63 MCI and 63 HC), recorded during an eye-closed resting condition at IRCCS Centro Neurolesi Bonino Pulejo of Messina (Italy). Experimental results reported that the 1-hidden layer MLP (MLP1) outperformed all the other developed learning systems as well as recently proposed state-of-the-art methods, achieving accuracy rate up to 95.76 % ± 0.0045 and 86.84% ± 0.0098 in AD vs HC and AD vs MCI classification, respectively.
Cosimo Ieracitano, Nadia Mammone, Alessia Bramanti, Silvia Marino, Amir Hussain 0001, Francesco Carlo Morabito
IJCNN5
2019 MPSSD: Multi-Path Fusion Single Shot Detector
abstract
Recent prevalent one stage detectors, such as single shot detector (SSD) and RetinaNet, are able to detect objects faster than two stage ones while maintaining comparable accuracy. To further boost the accuracy, many studies focus on enhancing the multi-scale feature pyramid. Most of these current proposals focus on strengthening features on one pyramid, ignoring the rich connection among different scale features. In contrast, we propose a novel multi-path design to fully utilize the localization and semantics information. First, we exploit the original SSD multi-scale features as our base pyramid. Then we fuse these features in different groups to generate multi-path feature pyramids. Finally, we combine these pyramids through a novel and effective aggregation module, to obtain the final informative pyramid for detection. Comparative experiments on benchmark PASCAL VOC and MS COCO datasets have shown that our proposed method outperforms many state-of-the-art detectors. As an illustrative example, for input image with size 512×512, we can achieve a mean Average Precision (mAP) of 81.8% on VOC2007 test and 33.1% mAP on COCO test-dev2015.
Shuyi Qu, Kaizhu Huang, Amir Hussain 0001, John Yannis Goulermas
IJCNN3
2019 A comparison of two methods of using a serious game for teaching marine ecology in a university setting
abstract
There is increasing interest in the use of serious games in STEM education. Interactive simulations and serious games can be used by students to explore systems where it would be impractical or unethical to perform real world studies or experiments. Simulations also have the capacity to reveal the internal workings of systems where these details are hidden in the real world. However, there is still much to be investigated about the best methods for using these games in the classroom so as to derive the maximum educational benefit. We report on an experiment to compare two different methods of using a serious game for teaching a complex concept in marine ecology, in a university setting: expert demonstration versus exploration-based learning. We created an online game based upon a mathematical simulation of fishery management, modelling how fish populations grow and shrink in the presence of stock removal through fishing. The player takes on the role of a fishery manager, who must set annual catch quotas, making these as high as possible to maximise profit, without exceeding sustainable limits and causing the stock to collapse. There are two versions of the game. The “white-box” or “teaching” game gives the player full information about all model parameters and actual levels of stock in the ocean, something which is impossible to measure in reality. The “black-box” or “testing” game displays only the limited information that is available to fishery managers in the real world, and is used to test the player's understanding of how to use that information to solve the problem of estimating the optimal catch quota. Our study addresses the question of whether students are likely to learn better by freely exploring the teaching game themselves, or by viewing a demonstration of the game being played expertly by the lecturer. We conducted an experiment with two groups of students, one using free, self-directed exploration and the other viewing an expert demonstration. Both groups were then assessed using the black box testing game, and completed a questionnaire. Our results show a statistically significant benefit for expert demonstration over free exploration. Qualitative analysis of the responses to the questionnaire demonstrates that students saw benefits to both teaching approaches, and many would have preferred a combination of expert demonstration with exploration of the game. The research was carried out among a mix of undergraduate and taught postgraduate science students. Future research challenges include extending the current study to larger cohorts and exploring the potential effectiveness of serious games and interactive simulation-based teaching methods in a range of STEM subjects in both university and school settings.
Omair Z. Ameerbakhsh, Savi Maharaj, Amir Hussain 0001, Bruce J. McAdam
Int. J. Hum. Comput. Stud.3
2019 Bi-level multi-objective evolution of a Multi-Layered Echo-State Network Autoencoder for data representations
Naima Chouikhi, Boudour Ammar, Amir Hussain 0001, Adel M. Alimi
Neurocomputing3
2019 A Convolutional Neural Network approach for classification of dementia stages based on 2D-spectral representation of EEG recordings
Cosimo Ieracitano, Nadia Mammone, Alessia Bramanti, Amir Hussain 0001, Francesco Carlo Morabito
Neurocomputing4
2019 A novel deep learning driven, low-cost mobility prediction approach for 5G cellular networks: The case of the Control/Data Separation Architecture (CDSA)
abstract
One of the fundamental goals of mobile networks is to enable uninterrupted access to wireless services without compromising the expected quality of service (QoS). This paper reports a number of significant contributions. First, a novel analytical model is proposed for holistic handover (HO) cost evaluation, that integrates signaling overhead, latency, call dropping, and radio resource wastage. The developed mathematical model is applicable to several cellular architectures, but the focus here is on the Control/Data Separation Architecture (CDSA). Second, data-driven HO prediction is proposed and evaluated as part of the holistic cost, for the first time, through novel application of a recurrent deep learning architecture, specifically, a stacked long-short-term memory (LSTM) model. Finally, simulation results and preliminary analysis reveal different cases where non-predictive and predictive deep neural networks can be effectively utilized, based on HO management requirements. Both analytical and machine learning models are evaluated with a benchmark, real-world dataset measuring human behaviors and interactions. Numerical and comparative simulation results demonstrate the potential of our proposed deep learning-driven HO management framework, as a future benchmark for the mobile networking and machine learning communities.
Metin Öztürk, Mandar Gogate, Oluwakayode Onireti, Ahsan Adeel, Amir Hussain 0001, Muhammad Ali Imran 0001
Neurocomputing5
2019 Robust pixelwise saliency detection via progressive graph rankings
Bo Jiang 0002, Zhengzheng Tu, Amir Hussain 0001, Jin Tang 0001
Neurocomputing4
2019 Saliency detection via multi-view graph based saliency optimization
abstract
Saliency detection is an important problem in computer vision and pattern recognition area. Many works have been proposed for addressing the saliency detection task. As a popular method, graph based saliency optimization has been widely studied. However, previous works have universally focussed on single graph optimization which fails to consider multi-view feature representation of image content . In this paper, we first provide a general framework for traditional graph based saliency optimization models. Then, we extend the general framework to the multi-view case and propose our general multi-view graph based saliency optimization model. Finally, we present a particular implementation of our general model and derive an effective updating algorithm to solve it. Experimental results using several benchmark datasets demonstrate the effectiveness of our proposed saliency model.
Yun Xiao 0003, Bo Jiang 0002, Aihua Zheng, Aiwu Zhou, Amir Hussain 0001, Jin Tang 0001
Neurocomputing5
2019 Guided Policy Search for Sequential Multitask Learning
abstract
Policy search in reinforcement learning (RL) is a practical approach to interact directly with environments in parameter spaces, that often deal with dilemmas of local optima and real-time sample collection. A promising algorithm, known as guided policy search (GPS), is capable of handling the challenge of training samples using trajectory-centric methods. It can also provide asymptotic local convergence guarantees. However, in its current form, the GPS algorithm cannot operate in sequential multitask learning scenarios. This is due to its batch-style training requirement, where all training samples are collectively provided at the start of the learning process. The algorithm’s adaptation is thus hindered for real-time applications, where training samples or tasks can arrive randomly. In this paper, the GPS approach is reformulated, by adapting a recently proposed, lifelong-learning method, and elastic weight consolidation. Specifically, Fisher information is incorporated to impart knowledge from previously learned tasks. The proposed algorithm, termed sequential multitask learning-GPS, is able to operate in sequential multitask learning settings and ensuring continuous policy learning, without catastrophic forgetting. Pendulum and robotic manipulation experiments demonstrate the new algorithms efficacy to learn control policies for handling sequentially arriving training samples, delivering comparable performance to the traditional, and batch-based GPS algorithm. In conclusion, the proposed algorithm is posited as a new benchmark for the real-time RL and robotics research community.
Fangzhou Xiong, Biao Sun 0005, Xu Yang 0004, Hong Qiao, Kaizhu Huang, Amir Hussain 0001, Zhiyong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2018 Improving Deep Neural Network Performance with Kernelized Min-Max Objective
Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICONIP (1)4
2018 DNN Driven Speaker Independent Audio-Visual Mask Estimation for Speech Separation
abstract
Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on target speaker while filtering out other noises. In this study, we propose a novel deep neural network (DNN) based audiovisual (AV) mask estimation model. The proposed AV mask estimation model contextually integrates the temporal dynamics of both audio and noise-immune visual features for improved mask estimation and speech separation. For optimal AV features extraction and ideal binary mask (IBM) estimation, a hybrid DNN architecture is exploited to leverages the complementary strengths of a stacked long short term memory (LSTM) and convolution LSTM network. The comparative simulation results in terms of speech quality and intelligibility demonstrate significant performance improvement of our proposed AV mask estimation model as compared to audio-only and visual-only mask estimation approaches for both speaker dependent and independent scenarios.
Mandar Gogate, Ahsan Adeel, Ricard Marxer, Jon Barker, Amir Hussain 0001
INTERSPEECH5
2018 Three-Dimensional Local Energy-Based Shape Histogram (3D-LESH): A Novel Feature Extraction Technique
abstract
In this paper, we present a novel feature extraction technique, termed Three-Dimensional Local Energy-Based Shape Histogram (3D-LESH), and exploit it to detect breast cancer in volumetric medical images. The technique is incorporated as part of an intelligent expert system that can aid medical practitioners making diagnostic decisions. Analysis of volumetric images, slice by slice, is cumbersome and inefficient. Hence, 3D-LESH is designed to compute a histogram-based feature set from a local energy map, calculated using a phase congruency (PC) measure of volumetric Magnetic Resonance Imaging (MRI) scans in 3D space. 3D-LESH features are invariant to contrast intensity variations within different slices of the MRI scan and are thus suitable for medical image analysis. The contribution of this article is manifold. First, we formulate a novel 3D-LESH feature extraction technique for 3D medical images to analyse volumetric images. Further, the proposed 3D-LESH algorithmis, for the first time, applied to medical MRI images. The final contribution is the design of an intelligent clinical decision support system (CDSS) as a multi-stage approach, combining novel 3D-LESH feature extraction with machine learning classifiers, to detect cancer from breast MRI scans. The proposed system applies contrast-limited adaptive histogram equalisation (CLAHE) to the MRI images before extracting 3D-LESH features. Furthermore, a selected subset of these features is fed into a machine-learning classifier, namely, a support vector machine (SVM), an extreme learning machine (ELM) or an echo state network (ESN) classifier, to detect abnormalities and distinguish between different stages of abnormality. We demonstrate the performance of the proposed technique by its application to benchmark breast cancer MRI images. The results indicate high-performance accuracy of the proposed system (98%±0.0050, with an area under a receiver operating charactertistic curve value of 0.9900 ± 0.0050) with multiple classifiers. When compared with the state-of-the-art wavelet-based feature extraction technique, statistical analysis provides conclusive evidence of the significance of our proposed 3D-LESH algorithm.
Summrina Kanwal Wajid, Amir Hussain 0001, Kaizhu Huang
Expert Syst. Appl.2
2018 Semi-supervised learning for big social data analysis
Amir Hussain 0001, Erik Cambria
Neurocomputing1
2018 A new two-layer mixture of factor analyzers with joint factor loading model for the classification of small dataset problems
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas, Amir Hussain 0001
Neurocomputing5
2018 Spatial-temporal representatives selection and weighted patch descriptor for person re-identification
Aihua Zheng, Foqin Wang, Amir Hussain 0001, Jin Tang 0001, Bo Jiang 0002
Neurocomputing3
2018 Multi-modal Fusion
Huaping Liu 0001, Amir Hussain 0001, Shuliang Wang 0001
Inf. Sci.2
2018 Extracting online information from dual and multiple data streams
Zeeshan Khawar Malik, Amir Hussain 0001, Q. M. Jonathan Wu
Neural Comput. Appl.2
2018 Applications of Deep Learning and Reinforcement Learning to Biological Data
abstract
Rapid advances in hardware-based technologies during the past decades have opened up new possibilities for life scientists to gather multimodal data in various application domains, such as omics, bioimaging, medical imaging, and (brain/body)-machine interfaces. These have generated novel opportunities for development of dedicated data-intensive machine learning techniques. In particular, recent research in deep learning (DL), reinforcement learning (RL), and their combination (deep RL) promise to revolutionize the future of artificial intelligence. The growth in computational power accompanied by faster and increased data storage, and declining computing costs have already allowed scientists in various fields to apply these techniques on data sets that were previously intractable owing to their size and complexity. This paper provides a comprehensive survey on the application of DL, RL, and deep RL techniques in mining biological data. In addition, we compare the performances of DL techniques when applied to different data sets across various application domains. Finally, we outline open issues in this challenging research area and discuss future development perspectives.
Mufti Mahmud, M. Shamim Kaiser, Amir Hussain 0001, Stefano Vassanelli
IEEE Trans. Neural Networks Learn. Syst.3
2017 Benchmarking Multimodal Sentiment Analysis
Erik Cambria, Devamanyu Hazarika, Soujanya Poria, Amir Hussain 0001, R. B. V. Subramanyam
CICLing (2)4
2017 Adaptation of Sentiment Analysis Techniques to Persian Language
Kia Dashtipour, Amir Hussain 0001, Alexander F. Gelbukh
CICLing (2)2
2017 Improve Deep Learning with Unsupervised Objective
Shufei Zhang, Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICONIP (1)4
2017 Customer churn prediction in the telecommunication sector using a rough set approach
Adnan Amin, Sajid Anwar 0001, Awais Adnan, Khalid Alawfi, Amir Hussain 0001, Kaizhu Huang
Neurocomputing6
2017 Ensemble application of convolutional neural networks and multiple kernel learning for multimodal sentiment analysis
Soujanya Poria, Haiyun Peng, Amir Hussain 0001, Newton Howard, Erik Cambria
Neurocomputing3
2017 Group sparse regularization for deep neural networks
Simone Scardapane, Danilo Comminiello, Amir Hussain 0001, Aurelio Uncini
Neurocomputing3
2017 Affective Reasoning for Big Social Data Analysis
abstract
This special section focuses on the introduction, presentation, and discussion of novel techniques that further develop and apply affective reasoning tools and techniques for big social data analysis. A key motivation for this special section, in particular, is to explore the adoption of novel affective reasoning frameworks and cognitive learning systems to go beyond a mere word-level analysis of natural language text and provide novel concept-level tools and techniques that allow a more efficient passage from (unstructured) natural language to (structured) machine-processable affective data, in potentially any domain. The selected papers aim to address the wide spectrum of issues related to affective computing research and, hence, better grasp the current limitations and opportunities related to this fast-evolving branch of artificial intelligence. Out of the 29 submissions received, 5 were accepted to appear in the special section. One of the accepted papers underwent 3 rounds of revisions, the rest were revised twice. The papers appearing in this issue are briefly summarized.
Erik Cambria, Amir Hussain 0001, Alessandro Vinciarelli
IEEE Trans. Affect. Comput.2
2017 Multilayered Echo State Machine: A Novel Architecture and Algorithm
abstract
In this paper, we present a novel architecture and learning algorithm for a multilayered echo state machine (ML-ESM). Traditional echo state networks (ESNs) refer to a particular type of reservoir computing (RC) architecture. They constitute an effective approach to recurrent neural network (RNN) training, with the (RNN-based) reservoir generated randomly, and only the readout trained using a simple computationally efficient algorithm. ESNs have greatly facilitated the real-time application of RNN, and have been shown to outperform classical approaches in a number of benchmark tasks. In this paper, we introduce a novel criteria for integrating multiple layers of reservoirs within the ML-ESM. The addition of multiple layers of reservoirs are shown to provide a more robust alternative to conventional RC networks. We demonstrate the comparative merits of this approach in a number of applications, considering both benchmark datasets and real world applications.
Zeeshan Khawar Malik, Amir Hussain 0001, Q. M. Jonathan Wu
IEEE Trans. Cybern.2
2016 Convolutional MKL Based Multimodal Emotion Recognition and Sentiment Analysis
abstract
Technology has enabled anyone with an Internet connection to easily create and share their ideas, opinions and content with millions of other people around the world. Much of the content being posted and consumed online is multimodal. With billions of phones, tablets and PCs shipping today with built-in cameras and a host of new video-equipped wearables like Google Glass on the horizon, the amount of video on the Internet will only continue to increase. It has become increasingly difficult for researchers to keep up with this deluge of multimodal content, let alone organize or make sense of it. Mining useful knowledge from video is a critical need that will grow exponentially, in pace with the global growth of content. This is particularly important in sentiment analysis, as both service and product reviews are gradually shifting from unimodal to multimodal. We present a novel method to extract features from visual and textual modalities using deep convolutional neural networks. By feeding such features to a multiple kernel learning classifier, we significantly outperform the state of the art of multimodal emotion recognition and sentiment analysis on different datasets.
Soujanya Poria, Iti Chaturvedi, Erik Cambria, Amir Hussain 0001
ICDM4
2016 Learning Latent Features with Infinite Non-negative Binary Matrix Tri-factorization
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICONIP (1)4
2016 An online generalized eigenvalue version of Laplacian Eigenmaps for visual big data
Zeeshan Khawar Malik, Amir Hussain 0001, Q. M. Jonathan Wu
Neurocomputing2
2016 Fusing audio, visual and textual clues for sentiment analysis from multimodal content
Soujanya Poria, Erik Cambria, Newton Howard, Guang-Bin Huang, Amir Hussain 0001
Neurocomputing5
2015 Efficient text localization in born-digital images by local contrast-based segmentation
abstract
Text localization in born-digital images is usually performed using methods designed for scene text images. Based on the observation that text strokes in born-digital images mostly have complete contours and the pixels on the contours have high contrast compared with the adjacent non-text pixels, we propose a method to extract candidate text components using local contrast. First, the image is segmented into smooth and non-smooth regions. After removing non-text smooth regions, the remaining smooth regions are merged with non-smooth regions to form a candidate text image, which is binarized into high-value and low-value connected components (CCs). The CCs undergo CC filtering, line grouping and line classification to give the text localization result. Experimental results on the born-digital dataset of ICDAR2013 robust reading competition demonstrate the efficiency and superiority of the proposed method.
Amir Hussain 0001, Cheng-Lin Liu 0001
ICDAR3
2015 Local energy-based shape histogram feature extraction technique for breast cancer diagnosis
Summrina Kanwal Wajid, Amir Hussain 0001
Expert Syst. Appl.2
2015 Towards an intelligent framework for multimodal affective data analysis
Soujanya Poria, Erik Cambria, Amir Hussain 0001, Guang-Bin Huang
Neural Networks3
2014 Dependency-Based Semantic Parsing for Concept-Level Text Analysis
Soujanya Poria, Basant Agarwal, Alexander F. Gelbukh, Amir Hussain 0001, Newton Howard
CICLing (1)4
2014 Classification of Fish Ectoparasite Genus Gyrodactylus SEM Images Using ASM and Complex Network Model
Rozniza Ali, Bo Jiang 0002, Mustafa Man, Amir Hussain 0001, Bin Luo 0001
ICONIP (3)4
2014 Special issue on International Symposium on Neural Networks
Zhigang Zeng, Amir Hussain 0001, Qinglai Wei
Neurocomputing2
2014 EmoSenticSpace: A novel framework for affective common-sense reasoning
Soujanya Poria, Alexander F. Gelbukh, Erik Cambria, Amir Hussain 0001, Guang-Bin Huang
Knowl. Based Syst.4
2014 Affective neural networks and cognitive learning systems for big data analysis
Amir Hussain 0001, Erik Cambria, Björn W. Schuller, Newton Howard
Neural Networks1
2013 Cognitive computation: A case study in cognitive control of autonomous systems and some future directions
abstract
Cognitive computation is an emerging discipline linking together neurobiology, cognitive psychology and artificial intelligence. Springer Neuroscience has launched a journal in this exciting multidisciplinary topic, which seeks to publish biologically inspired theoretical, computational, experimental and integrative accounts of all aspects of natural and artificial cognitive systems. In this keynote, we outline and build on some of the pioneering work of the late Professor John Taylor, who was also founding Advisory Board Chair of Cognitive Computation, specifically his proposal on how to create a cognitive machine equipped with multi-modal cognitive capabilities. In this context, we first present a novel modular cognitive control framework for autonomous systems that could potentially realize the required cognitive action-selection and learning capabilities in Professor Taylor's envisaged cognitive machine. An ongoing case study in autonomous vehicle control is described, as a benchmark problem, with encouraging preliminary results in a range of realistic driving scenarios - and with significant potential fuel and emission economy implications, compared to conventional control systems. Finally, possible future avenues are explored, including our ongoing work aimed at developing a general modular cognitive framework incorporating multiple modalities, including vision, motor action, language and emotion, required for enabling multi-modal social cognitive and affective behavioral capabilities in future autonomous agents.
Amir Hussain 0001
IJCNN1
2012 Off-Line Handwritten Arabic Word Recognition Using SVMs with Normalized Poly Kernel
Abdulrahman Alalshekmubarak, Amir Hussain 0001, Qiufeng Wang 0001
ICONIP (2)2
2012 The Use of ASM Feature Extraction and Machine Learning for the Discrimination of Members of the Fish Ectoparasite Genus Gyrodactylus
Rozniza Ali, Amir Hussain 0001, James E. Bron, Andrew P. Shinn
ICONIP (4)2
2012 Towards IMACA: Intelligent Multimodal Affective Conversational Agent
Amir Hussain 0001, Erik Cambria, Thomas Mazzocco, Marco Grassi, Qiufeng Wang 0001, Tariq S. Durrani
ICONIP (1)1
2012 Retrieval of Semantic Concepts Based on Analysis of Texts for Automatic Construction of Ontology
Reshmy Krishnan, Amir Hussain 0001, Sherimon Puliprathu Cherian
ICONIP (1)2
2012 Decoding Network Activity from LFPs: A Computational Approach
Mufti Mahmud, Davide Travalin, Amir Hussain 0001
ICONIP (1)3
2012 A Novel Road Traffic Sign Detection and Recognition Approach by Introducing CCM and LESH
Usman Zakir, Asima Usman, Amir Hussain 0001
ICONIP (3)3
2012 Towards a Chinese Common and Common Sense Knowledge Base for Sentiment Analysis
Erik Cambria, Amir Hussain 0001, Tariq S. Durrani, Jiajun Zhang 0001
IEA/AIE2
2012 Clustering Social Networks Using Interaction Semantics and Sentics
Praphul Chandra, Erik Cambria, Amir Hussain 0001
ISNN (1)3
2012 Sentic Maxine: Multimodal Affective Fusion and Emotional Paths
Isabelle Hupont, Erik Cambria, Eva Cerezo Bagdasari, Amir Hussain 0001, Sandra Baldassarri
ISNN (2)4
2012 Affective Common Sense Knowledge Acquisition for Sentiment Analysis
Erik Cambria, Yunqing Xia, Amir Hussain 0001
LREC3
2012 Sentic PROMs: Application of sentic computing to the development of a novel unified framework for measuring health-care quality
Erik Cambria, Tim Benson, Chris Eckl, Amir Hussain 0001
Expert Syst. Appl.4
2012 Novel logistic regression models to aid the diagnosis of dementia
Thomas Mazzocco, Amir Hussain 0001
Expert Syst. Appl.2
2012 Sentic Computing for social media marketing
Erik Cambria, Marco Grassi, Amir Hussain 0001, Catherine Havasi
Multim. Tools Appl.3
2011 SimConnector: An Approach to Testing Disaster-Alerting Systems Using Agent Based Simulation Models
Muaz A. Niazi, Qasim Siddique, Amir Hussain 0001, Giancarlo Fortino
FedCSIS3
2011 Multi-stage classification of Gyrodactylus species using machine learning and feature selection techniques
abstract
This study explores the use of multi-stage machine learning based classifiers and feature selection techniques in the classification and identification of fish parasites. Accurate identification of pathogens is a key to their control and as a proof of concept, the monogenean worm genus Gyrodactylus, economically important pathogens of cultured fish species, an ideal test-bed for the selected techniques. Gyrodactylus salaris is a notifiable pathogen of salmonids and a semi-automated / automated method permitting its confident species discrimination from other non-pathogenic species is sought to assist disease diagnostics during periods of a suspected outbreak. This study will assist pathogen management in wild and cultured fish stocks, providing improvements in fish health and welfare and accompanying economic benefits. Multi-stage classification is proposed as a solution to this problem because use of a single classifier is not sufficient to ensure that all the species are accurately classified. The results show that Linear Discriminant Analysis (LDA) with 21 features is the best classifier for performing the initial classification of Gyrodactylus species. This first stage classification which allocates specimens to species-groups is then followed by a second or subsequent round of classification using additional classifiers to allocate species to their true class within the species-groups.
Rozniza Ali, Amir Hussain 0001, James E. Bron, Andrew P. Shinn
ISDA2
2011 Sentic Medoids: Organizing Affective Common Sense Knowledge in a Multi-Dimensional Vector Space
Erik Cambria, Thomas Mazzocco, Amir Hussain 0001, Chris Eckl
ISNN (3)3
2011 Guest Editorial Data-Based Control, Modeling, and Optimization
abstract
The 21 papers in this special section focus on data-based control, modeling, and optimization.
Tianyou Chai, Zhongsheng Hou, Frank L. Lewis, Amir Hussain 0001, Dongbin Zhao
IEEE Trans. Neural Networks4
2010 SenticSpace: Visualizing Opinions and Sentiments in a Multi-dimensional Vector Space
Erik Cambria, Amir Hussain 0001, Catherine Havasi, Chris Eckl
KES (4)2
2009 Controlled and Automatic Processing in Animals and Machines with Application to Autonomous Vehicle Control
Kevin N. Gurney, Amir Hussain 0001, Jonathan M. Chambers, Rudwan Abdullah
ICANN (1)2
2009 Brain inspired cognitive systems (BICS)
Amir Hussain 0001, Igor Aleksander, Leslie S. Smith, Ron Chrisley
Neurocomputing1
2009 Special issue on non-linear and non-conventional speech processing
Mohamed Chetouani, Marcos Faúndez-Zanuy, Amir Hussain 0001, Bruno Gas, Jean-Luc Zarader, Kuldip K. Paliwal
Speech Commun.3
2008 Emergent Common Functional Principles in Control Theory and the Vertebrate Brain: A Case Study with Autonomous Vehicle Control
Amir Hussain 0001, Kevin N. Gurney, Rudwan Abdullah, Jonathan M. Chambers
ICANN (2)1
2008 Autonomous intelligent cruise control using a novel multiple-controller framework incorporating fuzzy-logic-based switching and tuning
Rudwan Abdullah, Amir Hussain 0001, Kevin Warwick, Ali S. Zayed
Neurocomputing2
2008 Engineering of intelligent systems (ICEIS 2006)
Amir Hussain 0001, Simone G. O. Fiori, Ijaz Mansoor Qureshi, Tariq S. Durrani, Muhammad Mansoor Ahmed, K. Fukushima
Neurocomputing1
2007 Fuzzy Logic Based Switching and Tuning Supervisor for a Multi-variable Multiple Controller
abstract
This paper presents a novel fuzzy-logic based switching and tuning supervisor for an intelligent multiple-controller framework. The fuzzy logic based supervisor operates at the highest level of the system and makes a switching decision, on the basis of the required performance measure, between two non-linear fixed structure controllers, namely a conventional Proportional-Integral-Derivative (PID) controller, or a PID structure based zero and pole placement controller. The fuzzy supervisor also adaptively tunes the parameters of the controllers. The proposed methodology is used to simultaneously control the throttle and brake systems of a validated nonlinear vehicle model. Sample simulation results are used to demonstrate the effectiveness of the multiple-controller with respect to tracking desired vehicle speed changes and achieving the desired speed of response, whilst penalising excessive control action.
Rudwan Abdullah, Amir Hussain 0001, Marios M. Polycarpou
FUZZ-IEEE2
2006 A new biclustering technique based on crossing minimization
Ahsan Abdullah 0002, Amir Hussain 0001
Neurocomputing2
2006 Brain inspired cognitive systems (BICS 2004)
Amir Hussain 0001, Leslie S. Smith, Igor Alexander
Neurocomputing1
2006 A novel multiple-controller incorporating a radial basis function neural network based generalized learning model
Ali S. Zayed, Amir Hussain 0001, Rudwan Abdullah
Neurocomputing2
2005 Biclustering Gene Expression Data in the Presence of Noise
Ahsan Abdullah 0002, Amir Hussain 0001
ICANN (1)2
2005 A New RBF Neural Network Based Non-linear Self-tuning Pole-Zero Placement Controller
Rudwan Abdullah, Amir Hussain 0001, Ali S. Zayed
ICANN (2)2
2005 Non-linear Predictive Models for Speech Processing
Mohamed Chetouani, Amir Hussain 0001, Marcos Faúndez-Zanuy, Bruno Gas
ICANN (2)2
2005 New Neural Network Based Mobile Location Estimation in a Metropolitan Area
Muhammad Javed 0003, Amir Hussain 0001, Alexander Neskovic, Evan H. Magill
ICANN (2)2
2003 A recurrent multiscale architecture for long-term memory prediction task
abstract
In the past few years, researchers have been extensively studying the application of recurrent neural networks (RNNs) to solving tasks where detection of long term dependencies is required. This paper proposes an original architecture termed the Recurrent Multiscale Network, RMN, to deal with these kinds of problems. Its most relevant properties are concerned with maintaining conventional RNNs' capability of information storing whilst simultaneously attempting to reduce their typical drawback occurring when they are trained by gradient descent algorithms, namely the vanishing gradient effect. This is achieved through RMN which preprocesses the original signal separating information at different temporal scales through an adequate DSP tool, and handling each information level with an autonomous recurrent architecture; the final goal is achieved by a nonlinear reconstruction section. This network has shown a markedly improved generalization performance over conventional RNNs, in its application to time series prediction tasks where long range dependencies are involved.
Stefano Squartini, Amir Hussain 0001, Francesco Piazza
ICASSP (2)2
2003 Attempting to reduce the vanishing gradient effect through a novel recurrent multiscale architecture
abstract
This paper proposes a possible solution to the vanishing gradient problem in recurrent neural networks, occurring when such networks are applied to solving tasks where detection of long term dependencies is required. The main idea consists of pre-processing the signal (a time series typically) through a discrete wavelet decomposition, in order to separate the short term information from the long term ones, and treating each scale by different recurrent neural networks. The partial results concerning all the sequences at diverse time/frequency resolutions are combined through an adaptive nonlinear structure in order to achieve the final goal. This new preprocessing based approach is distinct from the other one reported in literature to-date, as it tends to mitigate the effects of the problem under study avoiding relevant changing in network's architecture and learning techniques. The overall system (called recurrent multiscale network, RMN) is described and its performances tested through typical tasks namely the latching problem and time series prediction.
Stefano Squartini, Amir Hussain 0001, Francesco Piazza
IJCNN2
2002 Stochastic resonance and finite resolutions in a leaky integrate-and-fire neuron
Nhamo Mtetwa, Leslie S. Smith, Amir Hussain 0001
ESANN3
2002 Stochastic Resonance and Finite Resolution in a Network of Leaky Integrate-and-Fire Neurons
Nhamo Mtetwa, Leslie S. Smith, Amir Hussain 0001
ICANN3
2000 Intelligibility assessment of a multi-band speech enhancement scheme
abstract
A series of psychoacoustic experiments are described which attempt to assess the capability of a multi-microphone sub-band adaptive (MMSBA) signal processing scheme for improving the intelligibility of speech corrupted with recorded automobile noise. For the experimental cases considered, the proposed MMSBA scheme employing diverse sub-band processing is shown to deliver a statistically significant improvement in terms of both speech intelligibility and perceived quality when compared with both the wideband processed and the noisy unprocessed case.
Amir Hussain 0001
ICASSP1
1999 Intelligibility improvements using diverse sub-band processing applied to noisy speech
Amir Hussain 0001, Douglas R. Campbell
EUROSPEECH1
1999 Real-time speech modeling using computationally efficient locally recurrent neural networks (CERNs)
abstract
A general class of Computationally Efficient locally Recurrent Networks (CERN) is described for real-time adaptive signal processing. The structure of the CERN is based on linear-in-the-parameters single-hiddenlayered feedforward neural networks such as the Radial Basis Function (RBF) network, the Volterra Neural Network (VNN) and the recently developed Functionally Expanded Neural Network (FENN), adapted to employ local output feedback. The corresponding learning algorithms are described and key structural and computational complexity comparisons are made between the CERN and conventional Recurrent Neural Networks. A speech signal is used, which shows that a Recurrent FENN based adaptive CERN predictor can significantly outperform the corresponding feedforward FENN and conventionally employed linear adaptive filtering models.
John J. Soraghan, Amir Hussain 0001, Ivy Shim
EUROSPEECH2
1999 Multi-Sensor Neural-Network Processing of Noisy Speech
abstract
In this paper, a novel Artificial Neural-Network (ANN) based multi-sensor multi-band adaptive signal-processing scheme is described for enhancing acoustic-speech corrupted by real noise and reverberation. Numerically robust adaptation-algorithms are employed for the ANN based sub-band filters; and, new simulation experiments are reported using real-reverberant automobile data which demonstrate that the proposed speech-enhancement system is capable of outperforming conventional linear filtering-based wide-band and multi-band noise-cancellation schemes.
Amir Hussain 0001
Int. J. Neural Syst.1
1999 Speech-Intelligibility Improvements Using a Binaural Adaptive-Scheme Based Conceptually on Human Auditory Processing
abstract
A speech enhancement scheme is presented using diverse processing in sub-bands spaced according to a human-cochlear describing function. The binaural adaptive scheme decomposes the wide-band input signals into a number of band-limited signals, superficially similar to the treatment the human ears perform on incoming signals. The results of a series of intelligibility and formal listening tests are presented in which acoustic speech signals corrupted with recorded automobile noise were presented to 15 normal hearing volunteer subjects. For the experimental cases considered, the proposed binaural adaptive sub-band processing scheme delivers a statistically significant improvement in terms of both speech-intelligibility and perceived quality when compared with both the conventional wide-band processed and the noisy unprocessed case. The scheme is capable of extension to a potentially more flexible sub-band processing method based on a constrained artificial neural network (ANN).
Amir Hussain 0001, Douglas R. Campbell
Int. J. Neural Syst.1
1998 Binaural sub-band adaptive speech enhancement using artificial neural networks
Amir Hussain 0001, Douglas R. Campbell
Speech Commun.1
1997 A new neural network structure for temporal signal processing
abstract
A new two-layer linear-in-the-parameters feedforward network termed the functionally expanded neural network (FENN) is presented, together with its design strategy and learning algorithm. It is essentially a hybrid neural network incorporating a variety of non-linear basis functions within its single hidden layer which emulate other universal approximators employed in the conventional multi-layered perceptron (MLP), radial basis function (RBF) and Volterra neural networks (VNN). The FENN's output error surface is shown to be uni-modal allowing high speed single run learning. A simple strategy based on an iterative pruning retraining scheme coupled with statistical model validation tests is proposed for pruning the FENN. Both simulated chaotic (Mackey-Glass time series) and real-world noisy, highly nonstationary (sunspot) time series are used to illustrate the superior modeling and prediction performance of the FENN compared with other previously reported, more complex neural network based predictor models.
Amir Hussain 0001
ICASSP1
1997 A new metric for selecting sub-band processing in adaptive speech enhancement systems
abstract
A multi-microphone adaptive speech enhancement system employing diverse sub-band processing is presented. A new robust metric is developed, which is capable of real-time implementation, in order to automatically select the best form of processing within each sub-band. It is based on an adaptively estimated inter-channel Magnitude Squared Coherence (MSC) relationship, which is used to detect the level of correlation between in-band signals from multiple sensors during noise-alone periods in intermittent speech. This paper reports recent results of comparative experiments with simulated anechoic data extended to include simulated reverberant data. The results demonstrate that the method is capable of significantly outperforming conventional noise cancellation schemes.
Amir Hussain 0001, Douglas R. Campbell, Tom J. Moir
EUROSPEECH1
1997 A new adaptive functional-link neural-network-based DFE for overcoming co-channel interference
abstract
A new approach for the decision feedback equalizer (DFE) based on the functional-link neural network is described. The structure is applied to the problem of adaptive equalization in the presence of intersymbol interference (ISI), additive white Gaussian noise, and co-channel interference (CCI). It is shown through simulation results for a severe amplitude distorted co-channel system that the decision feedback functional-link equalizer (DFFLE) provides significantly superior bit-error rate (BER) performance characteristics compared to the conventional DFE, the linear transversal equalizer (LTE), the nonlinear radial basis function (RBF) neural-network-based structures and the feed-forward functional-link equalizer (FFLE)-based structures. The DFFLE is also shown to have a significantly simpler computational requirement relative to the RBF and the FFLE.
Amir Hussain 0001, John J. Soraghan, Tariq S. Durrani
IEEE Trans. Commun.1