Shaikh Anowarul Fattah

dblp:39/1646 · DBLP profile ↗
← Back
39ranked-venue papers
9as first author
11since 2021 · last 2025
0000-0001-8090-2327ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 8 since 2021Systems, architecture and hardware · 9 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
YearPublicationVenuePosition
2025 SONICS: Synthetic Or Not - Identifying Counterfeit Songs
abstract
The recent surge in AI-generated songs presents exciting possibilities and challenges. These innovations necessitate the ability to distinguish between human-composed and synthetic songs to safeguard artistic integrity and protect human musical artistry. Existing research and datasets in fake song detection only focus on singing voice deepfake detection (SVDD), where the vocals are AI-generated but the instrumental music is sourced from real songs. However, these approaches are inadequate for detecting contemporary end-to-end artificial songs where all components (vocals, music, lyrics, and style) could be AI-generated. Additionally, existing datasets lack music-lyrics diversity, long-duration songs, and open-access fake songs. To address these gaps, we introduce SONICS, a novel dataset for end-to-end Synthetic Song Detection (SSD), comprising over 97k songs (4,751 hours) with over 49k synthetic songs from popular platforms like Suno and Udio. Furthermore, we highlight the importance of modeling long-range temporal dependencies in songs for effective authenticity detection, an aspect entirely overlooked in existing methods. To utilize long-range patterns, we introduce SpecTTTra, a novel architecture that significantly improves time and memory efficiency over conventional CNN and Transformer-based models. For long songs, our top-performing variant outperforms ViT by 8% in F1 score, is 38% faster, and uses 26% less memory, while also surpassing ConvNeXt with a 1% F1 score gain, 20% speed boost, and 67% memory reduction.
Md Awsafur Rahman, Zaber Ibn Abdul Hakim, Najibul Haque Sarker, Bishmoy Paul, Shaikh Anowarul Fattah
ICLR5
2025 S-LIME: An Explainable SVM-Based Framework for Human Activity Recognition Using LIME
abstract
Human Activity Recognition (HAR) has emerged as a pivotal application in healthcare, assistive technology, and smart environments, leveraging wearable sensor data to classify human actions. While Support Vector Machines (SVMs) are widely used for HAR due to their robustness in handling high-dimensional data, their lack of interpretability restricts their deployment in real-world applications requiring transparency and trust. To bridge this gap, we propose S-LIME (SVM-LIME), an explainable AI (XAI) framework that integrates Local Interpretable Model-Agnostic Explanations (LIME) with SVMs to enhance model interpretability in HAR systems. S-LIME enables domain experts, medical practitioners, and endusers to understand the feature contributions that drive SVM-based activity classifications, making AI-driven HAR more transparent and accountable. We evaluate S-LIME on benchmark HAR datasets, demonstrating its ability to generate human-interpretable insights while maintaining high classification accuracy. The framework provides instance-level feature explanations, highlighting the key sensor signals influencing activity classification. Additionally, we introduce LIME-based feature ranking, which identifies and optimizes critical sensor signals, further improving model performance. Our findings show that S-LIME achieves state-of-the-art accuracy while providing an interpretable decision-making process, paving the way for reliable, transparent, and ethically sound HAR models.
Muhammad Masud Karim, Shaharier Kabir, Shaikh Anowarul Fattah
TENCON3
2025 SSGRAM: 3-D Spectral-Spatial Feature Network Enhanced by Graph Attention Map for Hyperspectral Image Classification
abstract
Convolutional Neural Networks (CNN) and Graph Neural Networks (GNN) are two widely used architectures in Hyperspectral Image (HSI) Classification. Most CNN models tend to heavily rely on neighboring pixels of the target pixel, leading to performance degradation when target pixels differ from their surrounding pixels. To address this limitation, this paper introduces a novel approach to Hyperspectral Image (HSI) classification using a combination of CNN and GNN networks to complement the drawbacks of CNN models. We propose a CNN-based 3D-Spectral Spatial Feature Network (3D-S2FN) to extract features in spectral, spatial, and spectral-spatial spaces, utilizing pyramid squeeze attention to prioritize their importance. Additionally, we introduce a Graph Attention Feature Processor (GAFP) module that evaluates the relevance of neighboring HSI pixels to the target pixel, generating a graph attention map (GRAM). This GRAM is applied to the CNN’s intermediate features to mitigate neighborhood bias. Unlike conventional GNNs, the proposed GAFP module directly processes the data within a local window, without requiring dimensionality reduction methods that would diminish pixel-level detail. By integrating features from the GAFP and the attention-enhanced 3D-S2FN, the proposed SSGRAM method achieves superior classification performance, as demonstrated through extensive experiments on four publicly available datasets.
Bishmoy Paul, Shaikh Anowarul Fattah, Mohammad Adnan Rajib, Mohammad Saquib
IEEE Trans. Geosci. Remote. Sens.2
2024 Classical to Quantum Neural Network Transfer Learning Approach for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) has gained utmost attention among researchers for its vital role in Human-Computer Interaction, enabling systems to understand and respond to the emotional states of the users. In this work, the concept of transfer learning has been extended to the novel domain of hybrid classical-quantum neural networks for enhancing SER performance. The classical part involves preprocessing raw speech signals to generate three-channel log mel-spectrograms and the pretrained VGG 19 network for extracting useful emotional features. The quantum part is built upon a dressed quantum neural network (QNN), comprising multiple parallel QNN blocks, for accurate emotion classification. Extensive experiments on the publicly accessible TESS dataset reveal that the proposed hybrid transfer learning approach achieves state of the art accuracy of 98.93 %. An ablation study demonstrates the significant performance improvement attributed to the incorporation of QNN blocks. The integration of classical and quantum neural networks provides a robust and efficient solution, leveraging the strengths of both computational paradigms to enhance SER based applications.
Sheikh Iftekhar Ahmed, Shaswata Mahernob Sarkar, Shaikh Anowarul Fattah, Mohammad Saquib
TENCON3
2024 Advancements in Efficient and Sustainable Wireless Charging for Electric Vehicles
abstract
The transition to Electric Vehicles (EVs) is crucial for sustainable transportation. However, the adoption of traditional plug-in charging systems remains limited due to their time-consuming nature and infrastructure demands. This study presents a Wireless Power Transfer (WPT) system leveraging inductive coupling to enhance EV charging efficiency, compatibility, and safety. The system converts AC to DC with power factor correction, then transforms it into high-frequency AC to energize the transmitter coil. The receiver coil generates alternating current (AC), which then converts to direct current (DC) for charging the battery. The proposed WPT model, developed using MATLAB Simulink Optimizer, demonstrates a charging efficiency ranging from 90 % to 93 %, resulting in energy losses as low as 10 % to 15 % and improved user convenience by reducing the need for manual plug-in operations. The developed model assumes a standard 60 kWh EV battery can be charged from 20 % to 80 % in approximately 3 hours using the proposed WPT system, compared to 5–6 hours with conventional plug-in chargers.
Abu Shufian, Md. Tanvir Rahman, Sowrov Komar Shib, Mehrab Islam Omi, Saniat Rahman Zishan, Shaikh Anowarul Fattah, Mohammad Saquib
TENCON6
2024 Semi-Supervised Semantic Depth Estimation using Symbiotic Transformer and NearFarMix Augmentation
abstract
In computer vision, depth estimation is crucial for domains like robotics, autonomous vehicles, augmented reality, and virtual reality. Integrating semantics with depth enhances scene understanding through reciprocal information sharing. However, the scarcity of semantic information in datasets poses challenges. Existing convolutional approaches with limited local receptive fields hinder the full utilization of the symbiotic potential between depth and semantics. This paper introduces a dataset-invariant semi-supervised strategy to address the scarcity of semantic information. It proposes the Depth Semantics Symbiosis module, leveraging the Symbiotic Transformer for achieving comprehensive mutual awareness by information exchange within both local and global contexts. Additionally, a novel augmentation, NearFarMix is introduced to combat overfitting and compensate both depth-semantic tasks by strategically merging regions from two images, generating diverse and structurally consistent samples with enhanced control. Extensive experiments on NYU-Depth-V2 and KITTI datasets demonstrate the superiority of our proposed techniques in indoor and outdoor environments.
Md Awsafur Rahman, Shaikh Anowarul Fattah
WACV2
2024 MD-CardioNet: A Multi-Dimensional Deep Neural Network for Cardiovascular Disease Diagnosis From Electrocardiogram
abstract
Automated classification of cardiovascular diseases from electrocardiogram (ECG) signals using deep learning has gained significant interest due to its wide range of applications. However, existing deep learning approaches often overlook inter-channel shared information or lose time-sequence dependent information when considering 1D and 2D ECG representations, respectively. Moreover, besides considering spatial dimension, it is necessary to understand the context of the signals from a global feature space. We propose MD-CardioNet, an efficient deep learning architecture that captures temporal, spatial, and volumetric features from multi-lead ECG signals using multidimensional (1D, 2D, and 3D) convolutions to address these challenges. Sequential feature extractors capture time-dependent information, while a 2D convolution is applied to form an image representation from the multi-channel ECG signal, extracting inter-channel features. Additionally, a volumetric feature extraction network is designed to incorporate intra-channel, inter-channel, and inter-filter global space information. To reduce computational complexity, we introduce a practical knowledge distillation framework that reduces the number of trainable parameters by up to eight times ( from 4,304,910 parameters to 94,842 parameters) while maintaining satisfactory performance compatible with the other existing approaches. The proposed architecture is evaluated on a large publicly available dataset containing ECG signals from over 10,000 patients, achieving an accuracy of 97.3% in classifying six heartbeat rhythms. Our results surpass the performance of some state-of-the-art approaches. This paper presents a novel deep-learning approach for ECG classification that addresses the limitations of existing methods. The experimental results highlight the robustness and accuracy of MD-CardioNet in cardiovascular disease classification, offering valuable insights for future research in this field.
Md Toki Tahmid, Muhammad Ehsanul Kader, Tanvir Mahmud, Shaikh Anowarul Fattah
IEEE J. Biomed. Health Informatics4
2023 Artifact: A Large-Scale Dataset With Artificial And Factual Images For Generalizable And Robust Synthetic Image Detection
abstract
Synthetic image generation has opened up new opportunities but has also created threats in regard to privacy, authenticity, and security. Detecting fake images is of paramount importance to prevent illegal activities, and previous research has shown that generative models leave unique patterns in their synthetic images that can be exploited to detect them. However, the fundamental problem of generalization remains, as even state-of-the-art detectors encounter difficulty when facing generators never seen during training. To assess the generalizability and robustness of synthetic image detectors in the face of real-world impairments, this paper presents a large-scale dataset1named ArtiFact, comprising diverse generators, object categories, and real-world challenges. Moreover, the proposed multi-class classification scheme, combined with a filter stride reduction strategy addresses social platform impairments and effectively detects synthetic images from both seen and unseen generators. The proposed solution significantly outperforms other top teams by 8.34% on Test 1, 1.26% on Test 2, and 15.08% on Test 3 in the IEEE VIP Cup challenge at ICIP 2022, as measured by the accuracy metric.
Md Awsafur Rahman, Bishmoy Paul, Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Shaikh Anowarul Fattah
ICIP5
2023 A Parallel Quantum Feature Encoding Scheme for Effective Classical Data Classification in Quantum Convolutional Neural Networks
abstract
Quantum machine learning is one of the most exciting new avenues in the world of artificial intelligence, especially because of the enormous computational power of quantum computers and the promise of the development of near error-free quantum computers in the not-so-distant future. For quantum algorithms to be used in real-life applications, quantum computers must be able to work with classical data. One of the key steps in quantum algorithms dealing with classical data is the encoding of classical data points to quantum states, which can then be processed by quantum gates. It is known that the type of encoding technique that works best for a particular network is dependent on the dataset being used. In this paper, a new parallel structure is proposed utilizing two encoding techniques, namely amplitude encoding and angle encoding, for effective classical data classification via quantum neural network. The paper further proposes a maximally expressible and entangled ansatz used to design a simple Quantum Convolutional Neural Network (QCNN) with only 32 parameters, that is used in the latter stages of the network and is kept the same across all encoding instances so that a comparison between the different encoding methods is possible. Extensive experimentation is carried out on two publicly available image datasets, namely MNIST and Fashion MNIST. The results show that the proposed method achieves better results than any of the encoding techniques deployed alone for binary classification.
Raisa Mashtura, Jishnu Mahmud, Shaikh Anowarul Fattah, Mohammad Saquib
TENCON3
2023 STOW-Net: Spatio-Temporal Operation Based Deep Learning Network for Classifying Wavelet Transformed Motor Imagery EEG Signals
abstract
Brain-computer interface (BCI) systems rely on capturing characteristics of human brain activity from the electroencephalography (EEG) signals, especially for the reliable classification of motor imagery tasks. For multi-channel EEG signals, it is crucial to precisely capture the spatio-temporal variation along with the frequency characteristics. Hence, instead of directly operating on raw EEG data, in this paper, discrete wavelet transform (DWT) is first applied to the motor-imagery multi-channel EEG data and then a deep learning architecture is designed incorporating spatial-temporal operations, which operates on the DWT-transformed EEG signal. In the proposed architecture, temporal convolution followed by spatial convolution is performed on the DWT-operated MI-EEG signal, and this part is termed as SAT-net. Next, by considering all channels together convolutional operation is performed to reduce the number of channels and this part is termed as SOC-net. Finally, a fully connected layer is used to classify the MI-EEG data from the derived feature vector. Extensive experimentation is performed on multiple subjects taken from the MI-based EEG dataset BCI Competition IV 2a. It is found that the proposed model offers a classification accuracy of 84.65%, consistently providing better classification performance than that obtained by some state-of-the-art methods.
Md. Shohel Rana, Shaikh Anowarul Fattah, Mohammad Saquib
TENCON2
2021 CovTANet: A Hybrid Tri-Level Attention-Based Network for Lesion Segmentation, Diagnosis, and Severity Prediction of COVID-19 Chest CT Scans
abstract
Rapid and precise diagnosis of COVID-19 is one of the major challenges faced by the global community to control the spread of this overgrowing pandemic. In this article, a hybrid neural network is proposed, named CovTANet, to provide an end-to-end clinical diagnostic tool for early diagnosis, lesion segmentation, and severity prediction of COVID-19 utilizing chest computer tomography (CT) scans. A multiphase optimization strategy is introduced for solving the challenges of complicated diagnosis at a very early stage of infection, where an efficient lesion segmentation network is optimized initially, which is later integrated into a joint optimization framework for the diagnosis and severity prediction tasks providing feature enhancement of the infected regions. Moreover, for overcoming the challenges with diffused, blurred, and varying shaped edges of COVID lesions with novel and diverse characteristics, a novel segmentation network is introduced, namely tri-level attention-based segmentation network. This network has significantly reduced semantic gaps in subsequent encoding-decoding stages, with immense parallelization of multiscale features for faster convergence providing considerable performance improvement over traditional networks. Furthermore, a novel tri-level attention mechanism has been introduced, which is repeatedly utilized over the network, combining channel, spatial, and pixel attention schemes for faster and efficient generalization of contextual information embedded in the feature map through feature recalibration and enhancement operations. Outstanding performances have been achieved in all three tasks through extensive experimentation on a large publicly available dataset containing 1110 chest CT-volumes, which signifies the effectiveness of the proposed scheme at the current stage of the pandemic.
Tanvir Mahmud, Md. Jahin Alam, Sakib Chowdhury, Shams Nafisa Ali, Md Maisoon Rahman, Shaikh Anowarul Fattah, Mohammad Saquib
IEEE Trans. Ind. Informatics6
2020 ResCovNet: A Deep Learning-Based Architecture For COVID-19 Detection From Chest CT Scan Images
abstract
Automatic disease detection using machine learning-based techniques from X-ray and computed tomography (CT) can play a major role in the frontline to assist medical professionals during the current outbreak of COVID-19. Fast diagnosis of the disease is the key to reduce the uncontrollable spread of this life-threatening disease, where machine learning-based applications can contribute greatly by predicting the situation of patients so that professionals can decide accordingly. The major drawbacks of detecting COVID-19 are its similarities with different types of pneumonia, and the absence of properly labeled data. Considering the ResNet152V2 as a backbone network, an efficient architecture, namely ResCovNet is proposed to detect COVID-19 accurately from chest CT scan images by separating it from three types of pneumonia and normal cases. Otsu's thresholding is applied in the pre-processing step to strengthen the features for the classification network. With the use of proposed architecture, a very satisfactory classification accuracy of 88.1% is achieved to separate COVID-19 from all other four classes. Evaluating the performance of this study by 3-fold cross-validation, and comparison with related works prove that this adroit algorithm provides an effective way to be implemented as a diagnostic tool in the COVID-19 screening.
Ankan Ghosh Dastider, Mohseu Rashid Subah, Farhan Sadik, Tanvir Mahmud, Shaikh Anowarul Fattah
TENCON5
2020 Transfer Learning Based Method for COVID-19 Detection From Chest X-ray Images
abstract
Radiology examination of chest radiography or chest X-ray (CXR), is currently performed manually by radiologists. With the onset of the COVID-19 pandemic, there is now a need to automate this process which is currently one of the key methods of primary detection of the SARS-Cov-2 virus. This will lead to shorter diagnosis time and less human error. In this study, we try to perform three-class image classification on a dataset of chest X-rays of confirmed COVID-19 patients(408 images), confirmed pneumonia patients(4273 images), and chest X-rays of healthy people(1590 images). In total the dataset consists of 6271 people. We aim to use a Convolutional Neural Network(CNN) and transfer learning to perform this image classification task. Our model is based on a pre-trained InceptionV3 network with weights trained on the ImageNet dataset. We fine-tune the layers of the Inception network to train it to our specific task. We try fine-tuning the network to different extents by freezing a different number of layers and then comparing accuracy for each variation of the network. To evaluate the performance of our network we use several metrics which include Classification accuracy, Precision, Sensitivity, and Specificity. Our proposed method achieves an accuracy of 96.33% on a 3-class classification task (Normal, COVID-19, Pneumonia) and an accuracy of 99.39% on a 2-class (COVID and Non-COVID) classification task.
Nayeeb Rashid, Md Adnan Faisal Hossain, Mumtahina Islam Sukanya, Tanvir Mahmud, Shaikh Anowarul Fattah
TENCON6
2020 A Multi-Model Based Ensembling Approach to Detect COVID-19 from Chest X-Ray Images
abstract
Since the onset of COVID-19, radiographic image analysis coupled with artificial intelligence (AI) has become popular due to insufficient RT-PCR test kits. In this paper, an automated AI-assisted COVID-19 diagnosis scheme is proposed utilizing the ensembling approach of multiple convolutional neural networks (CNNs). Two different strategies have been carried out for ensembling: A feature level fusionbased ensembling method and a decision level ensembling method. Several traditional CNN architectures are tested and finally in the ensembling operation, MobileNet, InceptionV3, DenseNet201, DenseNet121 and Xception are used. To handle the computational complexity of multiple networks, transfer learning strategy is incorporated through ImageNet pre-trained weight initialization. For feature-level ensembling scheme, global averages of the convolutional feature maps generated from multiple networks are aggregated and undergo through fully connected layers for combined optimization. Additionally, for decision level ensembling scheme, final prediction generated from multiple networks are converged into a single prediction by utilizing the maximum voting criterion. Both strategies perform better than any individual network. Outstanding performances have been achieved through extensive experimentation on a public database with 96% accuracy on 3-class (COVID-19/normal/pneumonia) diagnosis and 89.21% on 4-class (COVID-19/normal/viral pneumonia/bacterial pneumonia) diagnosis.
Oishy Saha, Jarin Tasnim, Md. Tanvir Raihan, Tanvir Mahmud, Istak Ahmmed, Shaikh Anowarul Fattah
TENCON6
2019 Unmanned Floating Waste Collecting Robot
abstract
This paper presents the design of a cost-effective remote controlled floating waste removing robot which can be used in canals, ponds, rivers or in oceans. As the usage of plastic is growing in an unregulated way in many countries, toxins from these elements cause an imbalance in the ecosystem and are threatening to human health leading to cancers, birth defects and immune system problems. A prototype of the robot is built in which there are two propellers which are connected with two DC motors to move forward, backward, left and right- mobile application based Bluetooth control system to control the robot from a distance and a robotic hand to make it easy to collect the trashes. A conveyor belt arrangement guiding the trash to the collection box and a sensor arrangement for protection against overloading have been used. Both manufacturing and maintenance cost are kept very affordable. The prototype itself can collect trash weighing upto 10 kg, clean an area of about 3000 square centimeters drawing only 45 watts from the battery. The robot has the capability of working for four hours continuously without the necessity of charging. Floating waste collecting robot can be a savior for the endangered aquatic animals.
Abir Ahsan Akib, Faiza Tasnim, Disha Biswas, Maeesha Binte Hashem, Kristi Rahman, Arnab Bhattacharjee, Shaikh Anowarul Fattah
TENCON7
2019 Phonocardiogram Heartbeat Segmentation and Autoregressive Modeling for Person Identification
abstract
With the rapid advancement of biosensors and increasing demand of more secured biometric authentication system, cardiac signals are getting special attention. Because of its very simple acquisition technique, phonocardiogram (PCG) signal is getting popularity in this field. This paper presents an automatic person identification scheme based on autoregressive modeling of PCG beats. On a given PCG recording first preprocesing and then wavelet denoising are applied. In order to perform beat by beat operation, a segmentation scheme is proposed using the Hilbert envelope which extracts the PCG beat containing the first and second heart sounds. Next reflection coefficients are extracted by employing the AR Burg modeling of the PCG beat. Finally the AR Burg reflection coefficients are used in ensemble bag trees classifier to identify a person. Performance of the proposed method is being tested on PCG signals of 50 different person taken from a publicly available PCG dataset and very satisfactory identification performance is achieved.
Imran Fahad, Md Ashiqur Rahman Apu, Abesh Ghosh, Shaikh Anowarul Fattah
TENCON4
2019 Sleep Apnea Detection Based on Rician Modeling of Feature Variation in Multiband EEG Signal
abstract
Sleep apnea, a serious sleep disorder affecting a large population, causes disruptions in breathing during sleep. In this paper, an automatic apnea detection scheme is proposed using single lead electroencephalography (EEG) signal to discriminate apnea patients and healthy subjects as well as to deal with the difficult task of classifying apnea and nonapnea events of an apnea patient. A unique multiband subframe based feature extraction scheme is developed to capture the feature variation pattern within a frame of EEG data, which is shown to exhibit significantly different characteristics in apnea and nonapnea frames. Such within-frame feature variation can be better represented by some statistical measures and characteristic probability density functions. It is found that use of Rician model parameters along with some statistical measures can offer very robust feature qualities in terms of standard performance criteria, such as Bhattacharyya distance and geometric separability index. For the purpose of classification, proposed features are used in K Nearest Neighbor classifier. From extensive experimentations and analysis on three different publicly available databases it is found that the proposed method offers superior classification performance in terms of sensitivity, specificity, and accuracy.
Arnab Bhattacharjee, Suvasish Saha, Shaikh Anowarul Fattah, Wei-Ping Zhu 0001, M. Omair Ahmad
IEEE J. Biomed. Health Informatics3
2018 Automatic Handwritten words on Touchscreen to Text file converter
abstract
This paper proposes a system for converting handwritten words and numbers into a text file. Our system uses A CNN based method to identify letter and digits. This Automatic system requires image preprocessing, classifying into letters and digits and saving the letter into a text file. The system consists of a touchscreen as a user interface, an Arduino board (microcontroller ATmega 2560) and MATLAB. Any type of handwriting is tested with the classifying process; we got 89% accuracy using our own dataset. Both letter and digits can be recognized and converted into text file using this process. The proposed system will lessen the labor of creating electronic documents and provide easy preservation of data.
Bidya Debnath, Adrita Anika, Mohammed Abid Abrar, Tanney Chowdhury, Rajat Chakraborty, Asir Intisar Khan, Shaikh Anowarul Fattah, Celia Shahnaz
TENCON7
2018 Configurable Digital Hearing Aid System with Reduction of Noise for Speech Enhancement Using Spectral Subtraction Method and Frequency Dependent Amplification
abstract
The greatest challenge for hearing impaired people is the realization of speech in a noisy environment. For the compensation of hearing losses, hearing aid devices are used by hearing impaired people because of being programmable to regulate gain values with different frequencies, reduction of noise for speech enhancement and improved signal to noise ratio (SNR) which have made it better in performance than analog hearing aid. The speech signal is always deteriorated by background noise in a noisy environment which degrades the original speech. To develop the speech quality, perceptibility and degree of listener's exhaustion, speech enhancement is done using various methods. In this paper, we propose a configurable digital hearing aid system with reduction of noise for speech enhancement using spectral subtraction method using MATLAB. In this work, configurable digital hearing aid includes a combination of frequency dependent amplification of speech and noise reduction filter for background noise reduction which will provide more flexibility to hearing impaired people along with improvement of quality of the speech signal.
Biswajit Saha, Shakil Khan, Celia Shahnaz, Shaikh Anowarul Fattah, Mohammad Tariqul Islam 0003, Asir Intisar Khan
TENCON4
2018 Real-Time American Sign Language Recognition Using Skin Segmentation and Image Category Classification with Convolutional Neural Network and Deep Learning
abstract
A real-time sign language translator is an important milestone in facilitating communication between the deaf community and the general public. We hereby present the development and implementation of an American Sign Language (ASL) fingerspelling translator based on skin segmentation and machine learning algorithms. We present an automatic human skin segmentation algorithm based on color information. The YCbCr color space is employed because it is typically used in video coding and provides an effective use of chrominance information for modeling the human skin color. We model the skin-color distribution as a bivariate normal distribution in the CbCr plane. The performance of the algorithm is illustrated by simulations carried out on images depicting people of different ethnicity. Then Convolutional Neural Network (CNN) is used to extract features from the images and Deep Learning Method is used to train a classifier to recognize Sign Language.
Shadman Shahriar Nitol, Ashraf Siddiquee, Tanveerul Islam, Abesh Ghosh, Rajat Chakraborty, Asir Intisar Khan, Celia Shahnaz, Shaikh Anowarul Fattah
TENCON8
2018 A Low-cost Venturi Tube Spirometer for the Diagnosis of COPD
abstract
The paper presents the design methodology and implementation of a low-cost spirometer prototype that can detect pulmonary diseases. By sensing the pressure difference of the expiratory air passing through the 3D printed venturi tube, the system calculates forced vital capacity (FVC). Through free campaigns, the performance and efficiency of this venturi tube spirometer were tested in different locations and the accuracy rate was highly promising. With this great efficiency and accuracy, this model diagnoses chronic obstructive pulmonary disease (COPD). Apart from measuring FVC, the model also measures forced expiratory volume in one second (FEV1) and from their ratio, it identifies COPD. In underprivileged areas, each and every hospital cannot afford spirometers available in the market because of their high price. This low-cost spirometer model ensures proper detection of pulmonary diseases in those underprivileged areas and it is definitely going to add a new dimension to the diagnosis of COPD.
Parama Sridevi, Pritam Kundu, Tahmida Islam, Celia Shahnaz, Shaikh Anowarul Fattah
TENCON5
2018 Biometric Authentication Using CNN Features of Dorsal Vein Pattern Extracted from NIR Image
abstract
Biometric authentication is a process of identifying and differentiating individuals for security purpose that counts on the distinctive physiological and behavioral characteristics of a person for verification. Dorsal vein pattern is one of the prospective biometrics and we have made successful use of this biometric feature for authentication purpose. This paper describes a novel design and its implementation for identifying individuals on the basis of their vein pattern of dorsal hand. The prime goal has been to establish a method with better Correct Recognition Rate (CRR), low False Acceptance Rate (FAR) and low False Rejection Rate (FRR). For our paper we have used Near Infrared (NIR) image of dorsal hand as it provides better resolution of vein pattern in the image than visible light. This recognition system consists of several steps; they are denoising of the input image using Laplacian Scale Mixture Modeling (LSM), Region of Interest (ROI) extraction from denoised image using valley point detection method, contrast enhancement using pyramid based edge aware filtering method, contrast limited adaptive histogram equalization, binarization, several serial morphological operations and finally authentication using neural networks and Support Vector Machine (SVM) classifier. The experiment was initially performed on a database of 16 distinct subjects where there are 10 raw images of each subject taken in different conditions. The accuracy obtained from experimental results is 96.63%. The scale of the experiment can be extended for more users, which we believe will yield similar success and thus this method of biometric authentication using dorsal vein pattern can be used in security purpose.
Rafat Jamal Tazim, Md. Messal Monem Miah, Sanzida Sayedul Surma, Mohammad Tariqul Islam 0003, Celia Shahnaz, Shaikh Anowarul Fattah
TENCON6
2017 Low-cost smart electric wheelchair with destination mapping and intelligent control features
abstract
Wheelchairs provide immense impact for the physically challenged as they offer mobility with relative ease. But in cases like severe disability or seeking relief from physical exertion, there is a necessity for automated electric wheelchairs which are very expensive. In this paper, an innovative design and implementation of a robust and user-friendly smart electric wheelchair are presented, where the product cost is minimized considering the affordability of a large group of people. A notable feature of the proposed scheme is `Destination Mapping' by which the smart wheelchair learns the destinations accessed by the user and autonomously reaches those destinations afterward using speech recognition. There are also various smart features such as rough surface detection, torque adjustments, slope and obstacle detection which promise a comfortable and safe ride. The proposed system takes input from a joystick, microphone, rotary encoders and sonar sensors, processes them through a microcontroller and eventually drives the motors through a motor driver. After much study on selecting suitable motors for implementation, windshield wiper motors were used as they produce quite enough torque at low rpm and are locally available as recycled parts. Experimentation has been carried out with 50 disabled persons of various age groups and subsequent improvements were undertaken based on their feedback over the timeframe of 1 year. Satisfactory results were obtained which implies that the proposed wheelchair can really serve as a low-cost smart alternative to aid the physically challenged, which has a tremendous potential impact on society.
Celia Shahnaz, Ahmed Maksud, Shaikh Anowarul Fattah, Sayeed Shafayet Chowdhury
ISTAS3
2017 Smart-hat: Safe and smooth walking assistant for elderly people
abstract
Elderly people in our society may be visually challenged, physically weak and thus face troubles during day-to-day life movement. Due to poor vision, often they fail to detect obstacles around them and hurt themselves. Due to aging, abnormal involuntary movements and lack of a timely rescue system, the magnitude of injury and other adverse effects worsens. One possible way to reduce these old age sufferings could be to design an assistive device which can take corrective measures based on position, neighborhood information and the event of fall (if any). The aim of this research is to develop a cost effective, comfortable and easily producible smart-hat for safe and smooth movement of the elderly people. Since wearing a hat is a common phenomenon for elderly people, smart-hat will get ease of acceptance by them. Some of the key features to be offered by the smart-hat are — obstacle detection, providing aid in the movement in a straight path and fall detection. No such technology exists at present that incorporates all these features in a single device. The proposed device uses a very simple methodology. This system takes reading from an accelerometer and differentiates a fall from other normal activities using an adaptive threshold set in triaxial directions. This device aids an elderly person to walk straight using the angular displacement of head movement obtained from a gyroscope. Lastly, few sonar sensors are attached to the hat to sense obstacles around the user covering a 360-degree range. During an event of fall, the system keeps sending request for help to pre-assigned caregivers by using a GSM module, until any of them responds. Also, the device makes wireless arrangement for opening the lock of the front door for the ease of rescue operation when the victim is at home. The device showed an overall success rate of nearly 86.25%. The implications of this safe and smooth movement assistant would be to reduce the number of accidents, decrease the time of arrival of medical attention, increase supervision capability, thus reducing injuries as well as mortality rate; which in turn will ensure independent living of the target group.
Celia Shahnaz, Md. Julker Neyen Sampad, Dimitri Adhikary, Sajid Mahfuz Uchayash, Marjana Mahdia, Sayeed Shafayet Chowdhury, Shaikh Anowarul Fattah
ISTAS7
2017 Noise Robust Formant Frequency Estimation Method Based on Spectral Model of Repeated Autocorrelation of Speech
abstract
In this paper, a noise robust formant frequency estimation scheme is developed based on a spectral model matching algorithm. Considering the vocal tract as an autoregressive system, a spectral model of repeated autocorrelation function (RACF) of band-limited speech signal is proposed. It is shown that because of the repeated autocorrelation operation on band-limited signal, the proposed model can exhibit prominent formant characteristics. First from given noisy speech observations, an adaptive band selection criterion is developed. Next, on each resulting band-limited noisy speech signal, a repeated autocorrelation operation is carried out, which not only reduces the effect of noise but also strengthens the dominant poles corresponding to the formant frequencies. Finally, spectrum of the RACF is computed and instead of direct spectral peak picking, a model fitting scheme is introduced to find out model parameters which lead to formant estimation. The proposed algorithm has been tested on natural vowels as well as some naturally spoken sentences in the presence of different environmental noises. It is found that the proposed scheme provides better formant estimation accuracy in comparison to some of the existing methods at low levels of signal-to-noise ratio.
Abu Shafin Mohammad Mahdee Jameel, Shaikh Anowarul Fattah, Rajib Goswami, Wei-Ping Zhu 0001, M. Omair Ahmad
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Seizure detection exploiting EMD-wavelet analysis of EEG signals
abstract
In his paper a method of seizure detection has been proposed based on the Discrete Wavelet Transform (DWT) analysis of the dominant Intrinsic mode function(IMF) resulting from the Empirical Mode Decomposition(EMD) of the EEG signals. Considering the normalized energy, Fourier spectrum and cross-correlation coefficient analysis, only the 4th Level DWT coefficients of the dominant IMF is found reasonable for feature computation. In order to reduce the dimension of the feature vector, Higher order statistics of these coefficients are employed to form he feature vector. The reduced feature vector thus formed is found effective for distinguishing seizure and non-seizure EEG signals when fed to a k-nearest neighborhood (k-NN) classifier. Extensive simulations are carried out using a benchmark EEG dataset. It is shown that the proposed method is capable of producing greater sensitivity, specificity and accuracy in comparison to that obtained by a sate-of-the-art method using the same EEG dataset and classifier.
Celia Shahnaz, R. H. Md. Rafi, Shaikh Anowarul Fattah, Wei-Ping Zhu 0001, M. Omair Ahmad
ISCAS3
2014 Speech emotion recognition based on entropy of enhanced wavelet coefficients
abstract
This paper presents a speaker-independent speech emotion recognition method, where emotional features are derived from the Teager energy (TE) operated wavelet coefficients of speech signal. Due to TE operation, the enhanced detail as well as approximate Wavelet coefficients thus obtained is then used to compute entropy. Entropy values of TE operated detail and approximate wavelet coefficients not only reduces feature dimension but also form an effective feature vector for distinguishing different emotions when fed to a Euclidean distance based classifier. Extensive simulations are carried out using EMO-DB German speech emotion database containing four class emotions, such as angry, happy, sad and neutral. Simulation results show that the proposed method is capable of outperforming an existing speaker-independent emotion recognition method thus solving a four-class emotion recognition problem in terms of higher recognition accuracy with lower computation.
Sharifa Sultana, Celia Shahnaz, Shaikh Anowarul Fattah, Istak Ahmmed, Wei-Ping Zhu 0001, M. Omair Ahmad
ISCAS3
2013 Identification of motor neuron disease using wavelet domain features extracted from EMG signal
abstract
Amyotrophic lateral sclerosis (ALS) is a common fatal motor neuron disease that assails the nerve cells in the brain. As the nervous system controls the muscle activity, the electromyography (EMG) signals can be viewed and examined in order to detect the vital features of the ALS disease in individuals. In this paper, the discrete wavelet transform (DWT) based features, which are extracted from a frame of EMG data, are introduced to classify the normal person and the ALS patients. From each frame of EMG data, instead of using a large number of DWT coefficients, the DWT coefficients with higher values as well as their mean and maxima are proposed to be used, which drastically reduces the feature dimension. It is shown that the proposed feature vector offers a high within class compactness and between class separations. For the purpose of classification, the K-nearest neighborhood classifier is employed. In order to demonstrate the classification performance, an EMG database consisted of 5 normal subjects and 5 ALS patients is considered and it is found that the proposed method is capable of distinctly separating the ALS patients from the normal persons.
Shaikh Anowarul Fattah, Abul Barkat Mollah Sayeed Ud Doulah, Md. Asif Iqbal, Celia Shahnaz, Wei-Ping Zhu 0001, M. Omair Ahmad
ISCAS1
2013 A detection method of nasalised vowels based on an acoustic parameter derived from phase spectrum
abstract
In this paper, a phase spectrum based acoustic parameter is presented for the detection of nasalized vowels from the mixture of oral and nasalized vowels of normal speakers. Acoustic analysis shows that during the event of nasalization, although additional formants (resonances) at various frequency locations are introduced, the introduction of a new formant in low frequency region around 250 Hz is found to remain consistent irrespective of female or male speakers in the modified group delay derived from the phase spectrum. By exploiting and verifying this fact on the band-limited modified group delay spectrum capable of resolving two closely spaced formants, an acoustic parameter RMGD is derived. Utilizing RMGD, the problem of detecting nasalized vowels is solved based on a threshold based scheme or a Euclidean distance based classifier. Simulation Results on TIMIT database show that the proposed method even with a simple classifier is superior in performance in comparison to that of the methods using Mel-frequency cepstral coefficients as a feature and Hidden Markov Modeling or Support Vector Machine as a classifier.
Celia Shahnaz, Shamima Najnin, Shaikh Anowarul Fattah, Wei-Ping Zhu 0001, M. Omair Ahmad
ISCAS3
2012 Detection of voice disorders based on wavelet and prosody-related properties
abstract
This paper presents an approach to detect voice disorders based on wavelet and prosody-related voice properties. First, several statistical measures of the normalized energy contents of the Discrete Wavelet Transform (DWT) coefficients over all voice frames are determined. Then, similar statistical measures of some prosody-related voice properties, such as mean pitch, jitter and shimmer are also computed over all the frames. In order to form a feature vector to be used in both training and testing phases, a set of statistical measure of the normalized energy contents of the DWT coefficients is combined with a set of statistical measure of the extracted prosody-related voice properties. Here, the voice samples under consideration are assumed to be of two categories, namely healthy and disordered thus formulating the problem in the proposed method as a two-class problem to be solved. Finally, the feature vector as obtained above is fed to an Euclidean Distance based classifier to detect the disordered voice. By performing extensive simulations, it is shown that the statistical analysis based on wavelet and prosody-related properties are able to provide effective detection of a variety of voice disorders from the mixture of healthy and disordered voices.
Celia Shahnaz, Shaikh Anowarul Fattah, Upal Mahbub, Wei-Ping Zhu 0001, M. Omair Ahmad
ISCAS2
2011 A mutual information based approach for evaluating the quality of clustering
abstract
In this paper, a new method for evaluating the quality of clustering of genes is proposed based on mutual information criterion. Instead of using the conventional histogram-based modeling method to assess clustering performance, we derive a normalized mutual information criterion utilizing the Gaussian kernel density estimator. In the computation of the mutual information, we propose to use only cluster-centroids instead of involving all the members, which offers a huge computational savings. The proposed algorithm not only considers the cluster size but also takes into consideration the homogeneity within a cluster. One major advantage of the proposed algorithm is that, it is capable of estimating an appropriate number of clusters. Extensive experimentation has been carried out on some synthetic data as well as the most widely used Yeast cell cycle gene expression data. Under various clustering conditions it is found that the proposed method provides an excellent performance in terms of measuring the quality of cluster and identifying the true number of cluster.
Shaikh Anowarul Fattah, Chia-Chun Lin, Sun-Yuan Kung
ICASSP1
2009 A Time-frequency Domain Formant Frequency Estimation Scheme for Noisy Speech Signals
abstract
Formant frequency is a one of the most important speech feature, which has widespread applications in speech recognition, synthesis, and compression. In this paper, a new time-frequency domain scheme for the estimation of formant frequencies from noise-corrupted speech signals is presented. In order to overcome the adverse effect of noise, instead of conventional autocorrelation function (ACF), a repeated ACF (RACF) of the noisy speech is employed. Exploiting the characteristics of the zero lag, a set of equations containing the lower lags of the RACF of the noisy speech is used to estimate the formant frequencies. In order to avoid estimation errors that may occur in the case of weak formants, a frequency-domain algorithm is introduced utilizing the RACF of the observed speech. Formant frequency estimation accuracy is measured for different natural and synthetic vowels in noisy environments and even at low levels of signal-to-noise ratio, a better performance is obtained by the proposed scheme in comparison to some of the existing methods.
Shaikh Anowarul Fattah, Wei-Ping Zhu 0001, M. Omair Ahmad
ISCAS1
2008 An algorithm for ARMA model parameter estimation from noisy observations
abstract
This paper presents a new algorithm for the parameter estimation of minimum-phase autoregressive moving average (ARMA) systems from noise-corrupted observations. In order to estimate the AR parameters of the ARMA system, based on a repeated autocorrelation function (ACF) of the observed data, a set of zero lag compensated equations has been developed. For the estimation of the MA parameters, first, a noise-subtraction algorithm is proposed to reduce the effect of noise from the ACF of the residual signal which is obtained by filtering the noisy ARMA signal via the estimated AR parameters. The MA parameters are then estimated by using a spectral factorization corresponding to the noise-compensated ACF of the residual signal. Computer simulations are carried out for ARMA systems of different orders under noisy environments and simulation results demonstrate a superior identification performance in terms of estimation accuracy and consistency.
Shaikh Anowarul Fattah, Wei-Ping Zhu 0001, M. Omair Ahmad
ISCAS1
2007 An Approach to Formant Frequency Estimation at Low Signal-to-Noise Ratio
abstract
A new approach for the formant frequency estimation of the voiced speech segments in the presence of noise is presented in this paper. A correlation model for the voiced speech is proposed considering the vocal-tract system as an autoregressive moving average (ARMA) model with a periodic impulse-train excitation. It is shown that the formant frequencies can be directly obtained from the model parameters. An adaptive residue-based least-squares optimization algorithm is proposed to estimate the model parameters, which overcomes the failure of conventional correlation based techniques in estimating formant frequencies at a low signal-to-noise ratio (SNR). The proposed algorithm has been tested on synthetic and natural vowels as well as voiced segments of some naturally spoken sentences from TIMIT database in presence of white Gaussian or babble noises. The experimental results show that the proposed method is more robust to noise than some existing methods even at a low SNR of 0 dB.
Shaikh Anowarul Fattah, Wei-Ping Zhu 0001, M. Omair Ahmad
ICASSP (4)1
2007 An Identification Technique for Noisy ARMA Systems in Correlation Domain
abstract
In this paper, an identification technique for the minimum-phase autoregressive moving average (ARMA) systems using only the noise-corrupted observations is presented. In order to obtain a more accurate estimate of the AR parameters in the noisy environment, a repeated autocorrelation function (RACF) of the observed data is employed in the modified least-squares Yule-Walker equations. It has been found that at a very low signal-to-noise ratio (SNR), the effect of the additive noise can be significantly reduced if a twice-RACF is employed instead of the conventional ACF. Prior to the MA part identification, a noise-compensation scheme is proposed which operates on the noise-contaminated residual signal. The MA parameters are extracted from the noise-compensated power spectrum of the residual signal using the spectral factorization. ARMA systems of different orders and some natural speech signals are tested and computer simulations demonstrate a superior identification results even at a very low SNR.
Shaikh Anowarul Fattah, Wei-Ping Zhu 0001, M. Omair Ahmad
ISCAS1
2006 A blind identification technique for noisy ARMA systems
abstract
In this paper, a new model for the ramp-cepstrum of the one-sided autocorrelation function of a noise-free autoregressive moving average (ARMA) signal is presented. The proposed blind identification technique can estimate the parameters of ARMA systems in both noise-free and noisy environments without using the input observations. It is shown that, utilizing the proposed ARMA ramp-cepstrum model in accordance with a residue-based least-squares optimization technique, both AR and MA parameters of ARMA systems can be directly obtained. The proposed method is tested on synthetic ARMA systems of different orders and also on some natural speech signals. Simulation results demonstrate the efficacy of the proposed identification scheme at low to high SNR levels
Shaikh Anowarul Fattah, Wei-Ping Zhu 0001, M. Omair Ahmad
ISCAS1
2005 An approach to ARMA system identification at a very low signal-to-noise ratio
abstract
A new approach for the identification of minimum-phase autoregressive moving average (ARMA) systems in the presence of heavy noise is presented in this paper. A damped sinusoidal (DS) model for the autocorrelation function of a noise-free ARMA signal is proposed to estimate the AR parameters, which overcomes the failure of conventional correlation based techniques in estimating the AR parameters of an ARMA system at a very low signal-to-noise ratio (SNR). The MA parameters of the ARMA system are then estimated by using Durbin's method along with an optimum order selection criterion. Both white noise and periodic impulse train excitations are considered for the application of the proposed method to system identification as well as to speech processing. Computer simulations are carried out based on both synthetic ARMA systems and natural speech signals, showing superior identification results even at an SNR of -5 dB for which most of the existing methods would fail.
Shaikh Anowarul Fattah, Wei-Ping Zhu 0001, M. Omair Ahmad
ICASSP (4)1
2003 Identification of noisy AR systems using damped sinusoidal model of autocorrelation function
abstract
This letter presents a novel method for minimum-phase autoregressive (AR) system identification at a very low SNR using damped sinusoidal model representation of the autocorrelation function of the noise-free AR signal with guaranteed stability. The new model parameters are estimated solely from the given noisy observations. Then AR parameters are obtained directly from the estimates of the damped sinusoidal model parameters. The simulation results show that the proposed method can estimate the AR system parameters with high accuracy even at an SNR as low as -5dB.
Md. Kamrul Hasan 0001, Shaikh Anowarul Fattah, M. Rezwan Khan
IEEE Signal Process. Lett.2
2002 Identification of AR systems at a very low SNR using damped cosine model of autocorrelation function
abstract
This paper presents a new method for autoregressive (AR) system identification at a very low signal to noise ratio (SNR) using damped cosine model for the autocorrelation function of the noise-free, AR signal. The AR parameters are obtained directly from the estimated damped cosine model parameters. The simulation results show that the method can estimate the system parameters with high accuracy even at an SNR as low as −5 dB.
Shaikh Anowarul Fattah, Md. Kamrul Hasan 0001, M. Rezwan Khan
ICASSP1