Shovan Barma

dblp:157/9846 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0001-8822-7362ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Real-Time Implementation of Accelerated HCP-MMA for Deep Learning-Based ECG Arrhythmia Classification Using Contour-Based Visualization
abstract
This study presents a real-time implementation of an accelerated Hurst Contour Projection from MultiscaleMultifractal Analysis (HCP-MMA) for deep learning-based ECG arrhythmia classification. Traditional heart rate variability analyses rely on fixed time scales and predefined parameters, limiting their ability to capture intricate scaling patterns and leading to diagnostic inconsistencies. HCP-MMA converts complex multifractal properties into a contour-based representation, enhancing interpretability for automated classification. However, the high computational cost of MMA hinders real-time processing. To address this, a runtime-optimized parallel computing pipeline is introduced, incorporating singular value decomposition (SVD) and vectorized processing, achieving a $730\times$ speedup over the baseline implementation on an Intel-based system. The proposed HCP-MMA framework, integrated with AlexNet, achieved over 98% classification accuracy across three benchmark datasets (PhysioNet, MIT-BIH, CU), with an F1-score of up to 99.3%. Runtime optimizations enabled real-time deployment on Raspberry Pi 5, demonstrating a $\sim 199\times$ speedup over baseline MMA computation on embedded hardware, with an average inference time of 0.0668 seconds per image, a memory footprint of approximately 220 MB, and a model size of $\sim \text{122}$ MB. Statistical validation using ANOVA and Tukey's HSD tests (p $< 0.05$) confirmed the approach's robustness and generalizability. By bridging computational efficiency with real-time adaptability, this method not only advances automated ECG diagnostics but also paves the way for scalable deployment in wearable monitoring, telemedicine, and multifractal analysis of complex physiological time-series.
Basab Bijoy Purkayastha, Shovan Barma
IEEE J. Biomed. Health Informatics2
2024 Attention Dynamics: Estimating Attention Levels of ADHD using Swin Transformer
Debashis Das Chakladar, Anand Shankar, Foteini Liwicki, Shovan Barma, Rajkumar Saini
ICPR (11)4
2024 Identifying Suitable Anatomical Planes in 3D MRI for AD Classification by Employing ViT Pipeline
abstract
This work aims to identify the suitable anatomical planes in 3D MRI for discriminating different types of Alzheimer's disease (AD) - cognitively normal (CN), mild cognitive impairment (MCI), and final stage Alzheimer's (AD). Such classification tasks in deep learning (DL) framework displayed noticeable results. Existing works consider global feature selection-based techniques in convolutional neural networks (CNN) and most of them consider all 3D images as input, which creates insignificant use of data that may not be suitable. Therefore, this work aims to identify suitable images from different anatomical planes of 3D MRI and emphasize on local-level features by involving the self-attention mechanism in Vision Transformer (ViT). For validation, MRI data from benchmark ADNI database has been taken into account. Aiming to find suitable MRI scans, 10–12 slices from each of three different anatomical planes - axial$(\mathrm{A}_{\mathrm{x}})$, coronal$(\mathrm{C}_{\mathrm{r}})$, and sagittal$(\mathrm{S}_{\mathrm{g}})$are segregated. Further, the images are fed into ViT by framing three cases (a) individual planes$\mathrm{A}_{\mathrm{x}}, \mathrm{C}_{\mathrm{r}},\mathrm{S}_{\mathrm{g}}, (\mathrm{b})$planes in pair$\left(\mathbf{A}_{x_{-}} \mathbf{C}_{\mathbf{r}}\right),\left(\mathbf{A}_{\mathbf{x}_{-}} \mathbf{S}_{\mathrm{g}}\right)$, and$(\mathrm{C}_{\mathrm{r}_{-}}\mathrm{S}_{\mathrm{g}})$and (c) all planes together$\left(\mathbf{A}_{\mathbf{x}_{-}} \mathbf{C}_{\mathrm{r}_{-}} \mathrm{S}_{\mathrm{g}}\right)$. The analysis has been performed by measuring accuracy$(A_{c})$, sensitivity$(S_{e})$, specificity$(S_{p})$, and Fl-score for a different set of inputs; while CNN model as a baseline. The results show that the ViT outperforms the CNN model, while the input pair$\mathrm{A}_{\mathrm{x}-}\mathrm{C}_{\mathrm{I}}$achieved accuracy as high as 93.3%. Thus, the analysis shows that choosing pair of planes could be suitable for MRI-based AD while involving DL-based techniques.
Chandita Barman, Shovan Barma
TENCON2
2024 MicrosMobiNet: A Deep Lightweight Network With Hierarchical Feature Fusion Scheme for Microscopy Image Analysis in Mobile-Edge Computing
abstract
In recent advancements of lightweight deep architectures for edge devices, most of the works follow a typical MobileNet pipeline designed for computer vision tasks which is not very appropriate for microscopy image analysis. Certainly, the design of the dedicated lightweight network for highly complex microscopy image analysis has not been attempted so far. Therefore, this work proposes a new deep lightweight network, “MicrosMobiNet” having multiscale feature extraction mechanism for bright-field microscopy image analysis on a mobile-edge computing framework. It consists of three key attributes—depth-wise separable convolution for making the network lightweight, multiple kernels with hierarchical feature fusion to extract complex features, and residual connection to keep network deep. Experimental validations have been conducted by two different microscopy image data sets—plant (potato tuber) and histopathology (cancer cell) generated by two different image generation modalities. In the experiment, multiclass and multilabel classification tasks have been evaluated by measuring accuracy, F1-score, and error. In the ablation study, the key attributes of the network have been verified. The results and analysis show that the MicrosMobiNet can achieve classification accuracy up to 98.43% and 96.25% for plant and cancer cells with minimum error 8.38% and 10.03%, respectively. In a comparative study, the MicrosMobiNet outperforms the existing lightweight state-of-the-art methods with fewer parameters (1.9M) and FLOPs count (42M). Finally, the new network has been implemented on an edge device, Smartphone (Android platform) which is working satisfactorily with high speed (140 ms) and very low memory (7.4 MB). Hence, the network exhibits its superiority in bright-field microscopy image analysis on mobile-edge computing platforms in a lightweight deep learning framework.
Sumona Biswas, Shovan Barma
IEEE Internet Things J.2
2024 Feature Fusion GAN Based Virtual Staining on Plant Microscopy Images
abstract
Virtual staining of microscopy specimens using GAN-based methods could resolve critical concerns of manual staining process as displayed in recent studies on histopathology images. However, most of these works use basic-GAN framework ignoring microscopy image characteristics and their performance were evaluated based on structural and error statistics (SSIM and PSNR) between synthetic and ground-truth without considering any color space although virtual staining deals with color transformation. Besides, major aspects of staining, like color, contrast, focus, image-realness etc. were totally ignored. However, modifications of GAN architecture for virtual staining might be suitable by incorporating microscopy image features. Further, its successful implementation need to be examined by considering various aspects of staining process. Therefore, we designed, a new feature-fusion-GAN for virtual staining followed by performance assessment by framing a state-of-the-art multi-evaluation framework that includes numerous metrics in -qualitative (based on histogram-correlation of color and brightness); quantitative (SSIM and PSNR); focus aptitude (Brenner metrics and Spectral-Moments); and influence on perception (semantic perceptual influence score). For, experimental validation cell boundaries were highlighted by two different staining reagents, Safranin-O and Toluidine-Blue-O on plant microscopy images of potato tuber. We evaluated virtually stained image quality w.r.t ground-truth in RGB and YCbCr color spaces based on defined metrics and results are found very consistent. Further, impact of feature fusion has been demonstrated. Collectively, this study could be a baseline towards guiding architectural upgrading of deep pipelines for virtual staining of diverse microscopy modalities followed by future benchmark methodology or protocols.
Sumona Biswas, Shovan Barma
IEEE ACM Trans. Comput. Biol. Bioinform.2
2023 Seizure Type Detection Using EEG Signals Based on Phase Synchronization and Deep Learning
abstract
Epileptic seizure occurs due to the intricate reorganization of neural networks in the brain that can be identified by using Electroencephalogram (EEG) signals. Several attempts have been made at its automatic detection by involving several machine learning algorithms, but fewer efforts have been made at the discrimination of its types. Eventually, accurate identification of different types of seizures can play an important role in clinical care, diagnosis, and preference for propitious drugs. However, its discrimination is very challenging due to indiscernible variation and distinct preeminent synchronization among them. Meanwhile, deep learning (DL), that automatically identifies feature vectors from input, has shown notable performance in image classification and could be suitable. However, its effective performance relies on how the 2D images are generated from 1D EEG followed by its feeding in the DL pipeline. Certainly, during a seizure, significant changes in phase synchronization among EEG channels can be observed, which can be exploited in the discrimination of seizure types. Therefore, in this work, 2D images were generated based on a phase synchronization matrix by measuring mean phase coherence among each pair of common EEG channels and fed into a convolution neural network (CNN) to classify three seizure types (absence, complex partial, and myoclonic seizures). For validation, an EEG dataset from the Temple University Hospital was used. The classification performance was evaluated in terms of accuracy, sensitivity, specificity, and weighted F1-score which reached up to 83.30%, 91.43%, 82.90%, and 83.03% respectively, which is significantly high. Further, through a 5-fold cross-validation, the proposed method shows its robustness.
Anand Shankar, Debaleena Chakraborty, Manob Jyoti Saikia, Samarendra Dandapat, Shovan Barma
BSN5
2022 Single-Channel Selection for EEG-Based Emotion Recognition Using Brain Rhythm Sequencing
abstract
Recently, electroencephalography (EEG) signals have shown great potential for emotion recognition. Nevertheless, multichannel EEG recordings lead to redundant data, computational burden, and hardware complexity. Hence, efficient channel selection, especially single-channel selection, is vital. For this purpose, a technique termed brain rhythm sequencing (BRS) that interprets EEG based on a dominant brain rhythm having the maximum instantaneous power at each 0.2 s timestamp has been proposed. Then, dynamic time warping (DTW) is used for rhythm sequence classification through the similarity measure. After evaluating the rhythm sequences for the emotion recognition task, the representative channel that produces impressive accuracy can be found, which realizes single-channel selection accordingly. In addition, the appropriate time segment for emotion recognition is estimated during the assessments. The results from the music emotion recognition (MER) experiment and three emotional datasets (SEED, DEAP, and MAHNOB) indicate that the classification accuracies achieve 70-82% by single-channel data with a 10 s time length. Such performances are remarkable when considering minimum data sources as the primary concerns. Furthermore, the individual characteristics in emotion recognition are investigated based on the channels and times found. Therefore, this study provides a novel method to solve single-channel selection for emotion recognition.
Jia Wen Li 0001, Shovan Barma, Peng Un Mak, Fei Chen 0011, Ming Tao Li, Mang I Vai, Sio-Hang Pun
IEEE J. Biomed. Health Informatics2
2022 Seizure Types Classification by Generating Input Images With in-Depth Features From Decomposed EEG Signals for Deep Learning Pipeline
abstract
Electroencephalogram (EEG) based seizure types classification has not been addressed well, compared to seizure detection, which is very important for the diagnosis and prognosis of epileptic patients. The minuscule changes reflected in EEG signals among different seizure types make such tasks more challenging. Therefore, in this work, underlying features in EEG have been explored by decomposing signals into multiple subcomponents which have been further used to generate 2D input images for deep learning (DL) pipeline. The Hilbert vibration decomposition (HVD) has been employed for decomposing the EEG signals by preserving phase information. Next, 2D images have been generated considering the first three subcomponents having high energy by involving continuous wavelet transform and converting them into 2D images for DL inputs. For classification, a hybrid DL pipeline has been constructed by combining the convolution neural network (CNN) followed by long short-term memory (LSTM) for efficient extraction of spatial and time sequence information. Experimental validation has been conducted by classifying five types of seizures and seizure-free, collected from the Temple University EEG dataset (TUH v1.5.2). The proposed method has achieved the highest classification accuracy up to 99% along with an F1-score of 99%. Further analysis shows that the HVD-based decomposition and hybrid DL model can efficiently extract in-depth features while classifying different types of seizures. In a comparative study, the proposed idea demonstrates its superiority by displaying the uppermost performance.
Anand Shankar, Samarendra Dandapat, Shovan Barma
IEEE J. Biomed. Health Informatics3
2016 An efficient image retrieval scheme for colour enhancement of embedded and distributed surveillance images
Kashif Iqbal, Michael O. Odetayo, Anne E. James, Rahat Iqbal, Neeraj Kumar 0001, Shovan Barma
Neurocomputing6
2016 A New Binary-Halved Clustering Method and ERT Processor for ASSR System
abstract
This paper presents an automatic speech–speaker recognition (ASSR) system implemented in a chip which includes a built-in extraction, recognition, and training (ERT) core. For VLSI design (here, ASSR system), the hardware cost and time complexity are always the important issues which are improved in this proposed design in two levels: 1) algorithmic and 2) architecture. At the algorithm level, a newly binary-halved clustering (BHC) is proposed to achieve low time complexity and low memory requirement. In addition, at the architecture level, a new ERT core is proposed and implemented based on data dependence and reuse mechanism to reduce the time and hardware cost as well. Finally, the chip implementation is synthesized, placed, and routed using TSMC 90-nm technology library. To verify the performance of the proposed BHC method, a case study is performed based on nine speakers. Moreover, the validation of the ASSR system is examined in two parts: 1) speech recognition and 2) speaker recognition. The results show that the proposed system can achieve 93.38% and 87.56% of recognition rates during speech and speaker recognition, respectively. Furthermore, the proposed ASSR chip includes 396k gate counts, and consumes power in 8.74 mW. Such results demonstrate that the performance of the proposed ASSR system is superior to the conventional systems.
Chih-Hung Chou, Ta-Wen Kuan, Shovan Barma, Bo-Wei Chen, Wen Ji 0003, Chih-Hsiang Peng, Jhing-Fa Wang
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Quantitative Measurement of Split of the Second Heart Sound (S2)
abstract
This study proposes a quantitative measurement of split of the second heart sound (S2) based on nonstationary signal decomposition to deal with overlaps and energy modeling of the subcomponents of S2. The second heart sound includes aortic (A2) and pulmonic (P2) closure sounds. However, the split detection is obscured due to A2-P2 overlap and low energy of P2. To identify such split, HVD method is used to decompose the S2 into a number of components while preserving the phase information. Further, A2s and P2s are localized using smoothed pseudo Wigner-Ville distribution followed by reassignment method. Finally, the split is calculated by taking the differences between the means of time indices of A2s and P2s. Experiments on total 33 clips of S2 signals are performed for evaluation of the method. The mean ± standard deviation of the split is 34.7 ± 4.6 ms. The method measures the split efficiently, even when A2-P2 overlap is ≤ 20 ms and the normalized peak temporal ratio of P2 to A2 is low (≥ 0.22). This proposed method thus, demonstrates its robustness by defining split detectability (SDT), the split detection aptness through detecting P2s, by measuring up to 96 percent. Such findings reveal the effectiveness of the method as competent against the other baselines, especially for A2-P2 overlaps and low energy P2.
Shovan Barma, Bo-Wei Chen, Ka Lok Man, Jhing-Fa Wang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 Game theory based no-reference perceptual quality assessment for stereoscopic images
Feng Jiang 0001, K. Bharanitharan, Shovan Barma, Debin Zhao
J. Supercomput.3