VLDB 2026 Research / reviewers in the wild / expert
Muhammad E. H. Chowdhury
dblp:184/6641 · also Muhammad Enamul Hoque Chowdhury
· DBLP profile ↗
44ranked-venue papers
0as first author
43since 2021 · last 2026
0000-0003-0744-8206ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spectral-spatial integration of hyperspectral imaging for glioblastoma classification with cross-modal attention and swin transformer
Mansura Naznine, Muhammad E. H. Chowdhury, Sawal Ali, Mohd Faisal Ibrahim, Mamun Bin Ibne Reaz, Asma Begam Mohamed Meeran, Rajendran Sankaran |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Point cloud-based three-dimensional segmentation of teeth from Cone Beam Computed Tomography images
Mehrin Newaz, Muhammad E. H. Chowdhury, Rusab Sarmun |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Enhancing multi-class satellite image classification with MRCL-ELM: a hybrid explainable deep learning approachabstractAbstract Satellite image classification has many important applications that play a crucial role in the areas of urban planning, agriculture, as well as environmental monitoring. Nevertheless, the high accuracy and interpretability of deep learning models with such complex datasets is still a big challenge. To solve this, a new hybrid deep-learning architecture, MRCL-ELM is proposed to improve the performance of satellite image classification. The model uses EfficientNetB0 as its building block in terms of ability to learn rich features in an efficient manner with optimization of the network depth and size to minimize memory and processing requirements. It combines the Multi-Residual Convolutional (MRC) networks to learn spatial features robustly with the aid of multiple residual paths, in each block, in MRC, to enhance the learning of features and gradient flow. It uses a Long Short-Term Memory (LSTM) time series modeling layer, and an Extreme Learning Machine (ELM) to quickly and non-iteratively classify data and is therefore lightweight, accurate, and more scalable than other common deep learning architectures. To enhance the interpretability of the proposed model, Local Interpretable Model-agnostic Explanations (LIME) explains individual predictions by testing small variations in the input, whereas SHapley Additive exPlanations (SHAP) provides feature importance scores throughout the model, along with improving model interpretability and trust. The proposed model provides the highest possible results, with 98.33% accuracy on the EuroSAT dataset and 98.10% accuracy on the UC Merced Land Use dataset, being higher than the use of existing Convolutional Neural Networks (CNN) and transformer-based techniques. The training using fixed random seeds and 5-fold cross-validation is used to ensure robustness. Lastly, MRCL-ELM was implemented as a real-time web-based application and tested with real-life Google Maps imagery, and thus needs real-time, precise, and interpretable satellite image classification for the end-users. Md Ashik Ahmmed, Rashel Mahmud Rabbi, Md Shafiuzzaman, Md. Faysal Ahamed, Md. Nahiduzzaman, Muhammad E. H. Chowdhury |
Neural Comput. Appl. | 6 |
| 2026 | A comprehensive review of U-Net architectures for medical image segmentation: emerging trends and federated learning perspectivesabstractAbstract U-Net has become a fundamental method in medical image segmentation with its architecture evolving to tackle complex segmentation tasks across modalities such as Magnetic Resonance Imaging (MRI), Computed Tomography (CT) scans and microscopic images. This review offers a comprehensive analysis of various U-Net architectures including U-Net++, Residual U-Net, Dense U-Net and more recent transformer-based models like TransUNet and Swin-UNet. These architectures introduce optimizations like improved feature propagation, gradient flow and self-attention mechanisms, significantly enhancing segmentation accuracy. Despite its widespread success, U-Net faces limitations in handling data imbalance, computational complexity and challenges with multi-modal and large-scale data. The integration of federated learning with U-Net addresses privacy concerns by enabling secure, collaborative model training across healthcare institutions while maintaining data confidentiality. This review also highlights the applications of U-Net variants in key medical imaging tasks, including brain segmentation, retinal vessel segmentation, cell and nuclei segmentation, prostate segmentation and skin lesion segmentation, which play a critical role in early disease detection and treatment planning. In addition to exploring these applications, the paper examines benchmark datasets, loss functions and evaluation metrics essential for assessing U-Net architectures. The review also discusses emerging trends like self-supervised learning, lightweight models for resource-constrained environments and interpretability techniques, offering potential directions for future research. By presenting both the advancements and limitations of U-Net, this review provides valuable insights into how these models can be optimized and applied to real-world medical applications. Md. Fahim Hossen, Turja Majumder, Md. Fahmidun Nabi, Md. Faysal Ahamed, Fariya Bintay Shafi, Rusab Sarmun, Muhammad E. H. Chowdhury |
Neural Comput. Appl. | 8 |
| 2026 | A comprehensive review of convolutional neural networks: foundations, enhancements and applicationsabstractAbstract Convolutional Neural Networks (CNNs) have emerged as a cornerstone in the field of deep learning, demonstrating remarkable performance across various domains, including computer vision and natural language processing. Their widespread acceptance on both academic and industrial levels has spurred much research and development. This study provides a comprehensive overview of recent developments in convolutional neural network (CNN) architectures by analyzing their foundational concepts, structural enhancements, and various applications such as image classification, medical imaging, and autonomous systems. Starting with conventional CNN architectures and their fundamental elements, the paper uses a systematic approach to look into recent advances. Key architectural innovations, including advanced activation functions, novel pooling strategies, and optimized convolutional techniques, are discussed. The study also explores hybrid architectures that integrate CNNs with transformers and recurrent neural networks to enhance contextual and sequential learning. Advances in training processes, such as enhanced loss functions and regularization approaches, have been studied for enhancing model performance. The study highlights CNN advancements that improve accuracy, reduce computational costs, and enhance model generalization. It underscores the effectiveness of CNNs in critical domains. Findings reveal that improved feature extraction techniques and interpretability methods, including Gradient-weighted Class Activation Mapping (Grad-CAM) and Graph-CNNs, contribute significantly to CNN performance. Finally, the paper identifies open challenges and outlines potential research directions, providing insights into the future development of CNNs. Md. Himel Reza, Md. Noman Biswas Sibly, Shaikh Golam Rabbani, Shafayetul Huda Sadi, Md. Faysal Ahamed, Fariya Bintay Shafi, Rusab Sarmun, Muhammad E. H. Chowdhury |
Neural Comput. Appl. | 8 |
| 2026 | Can deep learning-based segmentation and classification improve the detection of renal cortical abnormalities?abstractAbstract Renal cortical abnormalities frequently signify severe kidney conditions, rendering their precise diagnosis crucial for clinical management and treatment strategy formulation. Nonetheless, manual evaluation of nuclear renal imaging is arduous and prone to considerable inter-observer variability, resulting in conflicting results. This paper presents a fully automated method for the differential detection of renal cortical anomalies via deep learning-based segmentation and classification. Initially, we generated an innovative compilation of rigorously annotated renal nuclear images from 613 patients. Among them 193 patients are primarily diagnosed with kidney scar. Utilizing this dataset, we devised a novel segmentation method to precisely identify and outline renal areas. The proposed DenseNet121_Self-ONN_FPN model combines the DenseNet121 backbone, Self-Organizing Neural Network (Self-ONN) layers in the Feature Pyramid Network (FPN) for enhanced performance in segmentation tasks achieving impressive results: an Accuracy of 98.74%, Intersection over Union (IoU) of 86.47%, Dice Similarity Coefficient (DSC) of 92.74%, precision of 92.61%, recall of 92.88%, F1-score of 99.29%, False Negative Rate (FNR) of 7.12%, and False Positive Rate (FPR) of 0.71%. We optimized a modified DenseNet205 model for the classification of renal cortical anomalies. We employed Contrast Limited Adaptive Histogram Equalization (CLAHE) and Gamma correction as a pre-processing measure to enhance image contrast and model efficacy. The model attained exceptional results, exceeding state-of-the-art techniques with an accuracy of 96.91%, precision of 96.98%, sensitivity of 96.91%, F1-score of 96.86%, and specificity of 95.87%. Furthermore, we used ScoreCAM explainable AI to produce heatmaps for the classification network, offering critical insights into the model’s decision-making process and guaranteeing transparency for clinical application. This automated pipeline overcomes the constraints of manual image analysis by improving precision, efficiency, and reliability, while establishing a new standard for renal segmentation and classification. Tariq O. Abbas, Mansura Naznine, Muhammad E. H. Chowdhury |
Neural Comput. Appl. | 4 |
| 2026 | FootSegONN: an ensemble of Self-ONN-based models for diabetic foot ulcer segmentationabstractAbstract Diabetes mellitus is a chronic metabolic disorder characterized by persistent high blood sugar levels due to insulin deficiencies, leading to complications such as diabetic foot ulcers (DFUs). DFUs are associated with high morbidity and a significant risk of amputation, necessitating precise monitoring and management. Manual measurement of ulcer areas is labor-intensive and error-prone, prompting the need for automated, computer-aided methods. Deep learning (DL) techniques have shown promise in this domain, enhancing the accuracy and efficiency of ulcer detection and segmentation. This study investigates various methodologies for DFU segmentation, focusing on advanced DL models. We propose FootSegONN, an EfficientNet-based encoder and Self-organized Operational Neural Network (Self-ONN) and Feature Pyramid Network-based decoder for foot ulcer segmentation, evaluated on a publicly available chronic wound dataset consisting of 1010 diabetic foot images. Self-ONNs address the limitations of conventional Convolutional neural networks by achieving ultimate heterogeneity and boosting network diversity while maintaining computational efficiency. To ensure robust validation, fivefold cross-validation was applied, and more than 10 different segmentation models were utilized. The STAPLE algorithm is employed to combine mask predictions from top-performing models, and a post-processing approach is investigated to enhance performance further. The proposed method achieved a state-of-the-art Dice score of 91.55%. Furthermore, the combined FootSegONN model achieved a Dice score of 88.97% on the external validation dataset. Gradient-weighted class activation map (Grad-CAM) visualization was utilized to assess the model’s interpretability. Md. Shaheenur Islam Sumon, Saadia Binte Alam, Rashedur Rahman, Rusab Sarmun, Md. Mezbah Ahmed Mahedi, Zaid Bin Mahbub, Rumana Habib, Muhammad E. H. Chowdhury |
Neural Comput. Appl. | 8 |
| 2026 | Floating waste detection using deep learning: a comparative study of YOLO, RT-DETR, and faster R-CNNabstractAbstract Floating waste in inland water bodies poses severe threats to aquatic ecosystems, water quality, and public health. The accurate and timely detection of such waste is essential for enabling autonomous cleanup sys-tems like unmanned surface vehicles (USVs). However, detecting floating waste remains challenging due to the small size of debris, water surface reflections, glare, and complex backgrounds. This study presents a comparative evaluation of state-of-the-art deep learning-based object detection models—YOLO (v8–v10), Faster R-CNN, and Real-Time Detection Transformer (RT-DETR)—using the FloW-Img dataset, which is specifically designed for floating waste detection from USV perspectives. To enhance detection performance, we also explored four ensemble strategies: Weighted Box Fusion (WBF), Non-Maximum Suppression (NMS), Soft-NMS, and Non-Maximum Weighted (NMW). Our experiments show that the ensemble of RT-DETR-X and Faster R-CNN using WBF achieves the best results, with a mean Average Precision (mAP50) of 89.081%. This performance surpasses all previously reported methods on the same dataset, including YOLO-Float and Cascade R-CNN. The findings demonstrate the effectiveness of deep learning ensembles in improving small object detection in challenging water environments. This comparative study contributes valuable insights for developing robust, real-time, and scalable solutions for environmental monitoring and automated waste management systems. Md. Shaheenur Islam Sumon, Muhammad E. H. Chowdhury, Jawad-Ul Kabir Chowdhury, Azad Ashraf, Saad Bin Abul Kashem, Molla E. Majid, Mohammad Nashbat, Amith Khandakar, Mazhar Hasan-Zia, Ali K. Ansaruddin Kunju |
Neural Comput. Appl. | 2 |
| 2026 | 3D Foot Kinetics Estimation From Distributed VGRF From Smart Insoles via 1D Domain TransformationabstractUnderstanding foot kinetics is fundamental to analyzing human locomotion, offering critical insights into mechanical loads exerted on the feet. While vertical ground reaction force (vGRF) is widely used in biomechanics research, comprehensive 3D kinetic measurements, including ground reaction force (GRF), ground reaction moment (GRM), and center of pressure (CoP) along the anterior-posterior and medial-lateral axes, provide deeper insights for various applications. Smart insoles, though portable, cost-effective, and user-friendly, primarily capture vGRF and often generate lower-quality data than force plates and instrumented treadmills. This study leverages deep learning-based domain transformation to generate instrumented treadmill-level 3D-GRF&M-CoP from distributed vGRF signals recorded by smart insoles for healthy subjects. Additionally, a multi-segment analysis is performed to identify the most relevant plantar regions for each kinetic parameter. The proposed approach is rigorously evaluated against treadmill data and benchmarked against state-of-the-art methods, accounting for subject variations and walking speeds. Key contributions include: (1) transforming distributed vGRF into 3D-GRF&M-CoP using 1D-segmentation models, (2) enhancing insole vGRF to treadmill quality, (3) optimizing insole pressure sensor layout for efficient 3D kinetics estimation, and (4) introducing Ke2KeNet, a novel deep learning model that outperforms current 1D-segmentation benchmarks. Sakib Mahmud, Muhammad E. H. Chowdhury, Faycal Bensaali |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Machine-agnostic automated lumbar MRI segmentation using a cascaded model based on generative neurons
Promit Basak, Rusab Sarmun, Saidul Kabir, Israa Al-Hashimi, Enamul Hoque Bhuiyan, Anwarul Hasan, Muhammad Salman Khan 0001, Muhammad E. H. Chowdhury |
Expert Syst. Appl. | 8 |
| 2025 | Advanced deep learning and large language models: Comprehensive insights for cancer detection
Yassine Habchi, Hamza Kheddar, Yassine Himeur, Adel Belouchrani, Erchin Serpedin, Fouad Khelifi, Muhammad E. H. Chowdhury |
Image Vis. Comput. | 7 |
| 2025 | Deep learning-based beat-to-beat arterial blood pressure estimation using distant radar signalsabstractAbstract Maintaining constant vigilance over arterial blood pressure (ABP) is crucial for diagnosing hypertension and other critical cardiovascular diseases. While traditional cuff-based approaches are non-invasive, they have limitations in providing continuous blood pressure monitoring. In contrast, complex ABP monitoring systems, while accurate, are primarily suitable for clinical settings due to their intrusive nature. This study introduces a groundbreaking method for generating arterial blood pressure (ABP) waveforms using remote radar signals and deep learning (DL) techniques. This approach eliminates the need for invasive procedures, wearable biosensors, and costly equipment typically associated with ABP recording. We introduce MultiResLinkNet, a segmentation model based on a one-dimensional convolutional neural network (1D CNN), specifically designed to synthesize arterial blood pressure (ABP) directly from raw radar waveforms. We trained and evaluated the end-to-end DL framework using a publicly available benchmark radar dataset containing raw radar data and corresponding physiological signals from 30 subjects across various scenarios, including Resting, Valsalva, Apnea, Tilt-up, and Tilt-down. The proposed MultiResLinkNet excelled in ABP segmentation, outperforming state-of-the-art networks in combined and individual scenarios, and produced the best average temporal and spectral correlations as well as the lowest temporal and spectral errors in nearly all scenarios’ data. Furthermore, qualitative evaluation demonstrated a strong resemblance between the synthesized and ground truth ABP waveforms. Our novel approach enables remote monitoring of critical patients continuously, especially those undergoing surgery, by predicting ABP waveforms from non-contact radar signals. This breakthrough offers significant advantages, facilitating continuous ABP monitoring without the need for invasive procedures or cumbersome wearable sensors. Chowdhury Farhan Ahmed, Md Kamal Hosain, Md. Shafayet Hossain, Muhammad E. H. Chowdhury, Sakib Mahmud, Muhammad Ashad Kabir, Abdulrahman Alqahtani, Anwarul Hasan |
Neural Comput. Appl. | 4 |
| 2025 | Improving pediatric trauma care: an automated system for wrist trauma detection using GELANabstractAbstract Trauma is a major cause of disability among children, requiring swift and accurate diagnosis for effective treatment. This paper introduces an automated method that uses deep learning to detect and categorize fractures in children using X-ray images. The system makes use of the GRAZPEDWRI-DX dataset, which consists of 20,327 annotated X-ray images of pediatric wrist fractures. Our architecture, which is built upon the generalized efficient layer aggregation network (GELAN), effectively tackles the issues of class imbalance and image resolution. As a result, it achieves state-of-the-art performance in both trauma and severity detection. Our proposed framework surpassed the most advanced techniques, showcasing exceptional precision and effectiveness, achieving a mean average precision (mAP50) score of 74.1%, 95%, and 85.5% for Task A (trauma detection), Task B (fracture detection), and Task C (fracture severity detection), respectively. The results of our study highlight the capacity of deep learning to improve the diagnosis of pediatric trauma, decrease the burden on radiologists, and boost patient outcomes. Promit Basak, Adam Mushtak, Mohamed Ouda, Sadia Farhana Nobi, Anwarul Hasan, Muhammad E. H. Chowdhury |
Neural Comput. Appl. | 6 |
| 2025 | Decoding silent speech: a machine learning perspective on data, methods, and frameworksabstractAbstract At the nexus of signal processing and machine learning (ML), silent speech recognition (SSR) has evolved as a game-changing technology that allows for communication without audible voice. This study offers a thorough overview of SSR, tracing its evolution from early waveform analysis to the most recent ML methods. We start by examining current SSR techniques using ML and determining the essential conditions for efficient SSR systems. After that, we look at the datasets and data collection techniques currently employed in SSR research, highlighting the difficulties posed by the variety of articulatory movements and the scarcity of data. Examining state-of-the-art SSR frameworks, the paper covers important topics such signal processing, feature extraction, ML techniques for decoding and optimizing and assessing the performance of SSR models. We emphasize how deep learning (DL) and ML models have evolved to increase SSR resilience and accuracy. The field's proposed procedures are examined, with an emphasis on sophisticated feature extraction and classification methods. Modern SSR techniques are compared in terms of performance, highlighting the advantages and disadvantages of different models. There is also discussion of ethical issues, especially those pertaining to privacy and consent. The integration of multimodal information—visual cues, electromyography signals, and neuroimaging data—to improve SSR systems is covered in this work. We investigate the functions of transfer learning and domain adaptation in handling cross-subject variability. Lastly, the study offers suggestions and future prospects for SSR research, providing practitioners, engineers, and academics with a road map. As SSR continues to push the frontiers of human–machine interaction, our study aims to increase our collective understanding of the technological advances and societal effects of SSR in the ML age. Adiba Tabassum Chowdhury, Mehrin Newaz, Purnata Saha, Mohannad Natheef AbuHaweeleh, Sara Mohsen, Diala Bushnaq, Malek Chabbouh, Raghad Aljindi, Shona Pedersen, Muhammad E. H. Chowdhury |
Neural Comput. Appl. | 10 |
| 2025 | VisioDECT: a novel approach to drone detection using CBAM-integrated YOLO and GELAN-E modelsabstractAbstract Unmanned aerial vehicles have revolutionized logistics, environmental monitoring, and aerial surveillance. Their widespread use has created security concerns, specifically regarding illegal spying, smuggling, and hazardous substance movement. Maintaining public safety and protecting sensitive locations requires effective drone detection and payload assessment. Our article proposes a vision-based system for real-time drone identification and classification utilizing YOLOv5, YOLOv8, and GELAN-E deep learning models, enhanced with novel attention mechanisms and interpretability techniques. By integrating the Convolutional Block Attention Module into the YOLO architecture, which is named AttnYOLO, the system enhances feature extraction and focuses on the most relevant regions in an image. This improvement in spatial and channel attention significantly boosts detection performance, particularly for small and occluded drones. Additionally, we employ Gradient-weighted Class Activation Mapping (EigenCAM) for visualizing model focus during detection, increasing the system’s transparency and interpretability. VisioDECT comprises 20,924 annotated photographs of six drone models in overcast, sunny, and evening circumstances. Under cloudy conditions, DenseNet201 achieved 100% classification accuracy, while DarkNet53 and InceptionV3 reached 99.99% and 99.9%, respectively. In evening scenarios, InceptionV3 had 100% accuracy, followed by DarkNet53 with 99.98%. Our proposed model, GELAN-E, excelled in detection and classification. In overcast settings, GELAN-E outperformed YOLOv8 with an accuracy of 0.988, a recall of 0.994, and a mAP50-95 score of 0.688. For evening conditions, GELAN-E achieved a higher mAP50-95 score of 0.642 compared to YOLOv8. These results demonstrate that the inclusion of attention mechanisms, along with visual interpretability, enhances drone detection performance, particularly in low-light and challenging environments, making this system ideal for real-time drone detection in civilian and military applications . Md. Sakib Bin Islam, Muhammad E. H. Chowdhury, Mazhar Hasan-Zia, Saad Bin Abul Kashem, Molla E. Majid, Ali K. Ansaruddin Kunju, Amith Khandakar, Azad Ashraf, Mohammad Nashbat |
Neural Comput. Appl. | 2 |
| 2025 | Deep learning and vision transformers-based framework for breast cancer and subtype identificationabstractAbstract Breast cancer, marked by uncontrolled cell growth in breast tissue, is the most common cancer among women and a second-leading cause of cancer-related deaths. Among its types, ductal and lobular carcinomas are the most prevalent, with invasive ductal carcinoma accounting for about 70–80% of cases and invasive lobular carcinoma for about 10–15%. Accurate identification is crucial for effective treatment but can be time-consuming and prone to interobserver variability. AI can rapidly analyze pathological images, providing precise, cost-effective identification, thus reducing the pathologists’ workload. This study utilizes a deep learning framework for advanced, automatic breast cancer detection and subtype identification. The framework comprises three key components: detecting cancerous patches, identifying cancer subtypes (ductal and lobular carcinoma), and predicting patient-level outcomes from whole slide images (WSI). The validation process includes visualization using Score-CAM to highlight cancer-affected areas prominently. Datasets include 111 WSIs (85 malignant from the Warwick HER2 dataset and 26 benign from pathologists). For subtype detection, there are 57 ductal and 8 lobular carcinoma cases. A total of 28,428 annotated patches were reviewed by two expert pathologists. Four pre-trained models—DenseNet-201, MobileNetV2, an ensemble of these two, and a Vision Transformer-based model—were fine-tuned and tested on the patches. Patient-level results were predicted using a majority voting technique based on the percentage of each patch type in the WSI. The Vision Transformer-based model outperformed other models in patch classification, achieving an accuracy of 96.74% for cancerous patch detection and 89.78% for cancer subtype classification. For WSI-based cancer classification, the majority voting method attained an F1-score of 99.06 and 96.13% for WSI-based cancer subtype classification. The proposed deep learning-based framework for advanced breast cancer detection and subtype identification yielded promising results. This advanced framework shows great promise in medical practice, offering an economical, efficient solution for generating accurate, clinically relevant results and enhancing diagnostic accuracy in hospitals, research centers, and pathology laboratories. Nonetheless, further studies are needed to validate its effectiveness across various environments and larger datasets. Ishrat Jahan 0003, Muhammad E. H. Chowdhury, Semir Vranic, Rafif Mahmood Al Saady, Saidul Kabir, Zahid Hasan Pranto, Sabiha Jahan Mim, Sadia Farhana Nobi |
Neural Comput. Appl. | 2 |
| 2025 | Optimizing energy efficiency through precise occupancy detection: A tailored CNN architecture for smart buildings and beyondabstractAbstract Occupancy detection is crucial for various applications, including smart buildings, security systems, and energy management. This paper introduces a novel convolutional neural network (CNN) architecture based on an image encoding approach for accurate occupancy detection. Our network effectively extracts relevant features from occupancy images by leveraging deep learning and image processing techniques, enabling reliable and real-time detection. We employed an image encoding method that converts environmental time-series data into 2D image representations—either grayscale or RGB-like—depending on the input requirements of the CNN model. This transformation captures spatial and temporal characteristics of the data, allowing the network to learn more expressive occupancy-related patterns from raw 1D input. Additionally, we developed a custom CNN architecture optimized for the encoded images, enabling the network to identify key features and understand complex spatial relationships. We evaluated the performance of our CNN through extensive testing on well-known occupancy datasets. The results highlight the superiority of our approach, outperforming existing techniques in accuracy, precision, recall, and F1-score. Our model achieved impressive accuracies of 98.45%, 99.05%, and 97.32% across the three datasets used in this study. Aya Nabil Sayed, Sakib Mahmud, Faycal Bensaali, Muhammad E. H. Chowdhury, Yassine Himeur |
Neural Comput. Appl. | 4 |
| 2025 | Enhancing waste sorting and recycling efficiency: robust deep learning-based approach for classification and detectionabstractAbstract Given the severity of waste pollution as a major environmental concern, intelligent and sustainable waste management is becoming increasingly crucial in both developed and developing countries. The material composition and volume of urban solid waste are key considerations in processing, managing, and utilizing city waste. Deep learning technologies have emerged as viable solutions to address waste management issues by reducing labor costs and automating complex tasks. However, the limited number of trash image categories and the inadequacy of existing datasets have constrained the proper evaluation of machine learning model performance across a large number of waste classes. In this paper, we present robust waste image classification and object detection studies using deep learning models, utilizing 28 distinct recyclable categories of waste images comprising a total of 10,406 images. For the waste classification task, we proposed a novel dual-stream network that outperformed several state-of-the-art models, achieving an overall classification accuracy of 83.11%. Additionally, we introduced the GELAN-E (generalized efficient layer aggregation network) model for waste object detection tasks, obtaining a mean average precision (mAP50) of 63%, surpassing other state-of-the-art detection models. These advancements demonstrate significant progress in the field of intelligent waste management, paving the way for more efficient and effective solutions. Faizul Rakib Sayem, Md. Sakib Bin Islam, Mansura Naznine, Mohammad Nashbat, Mazhar Hasan-Zia, Ali K. Ansaruddin Kunju, Amith Khandakar, Azad Ashraf, Molla E. Majid, Saad Bin Abul Kashem, Muhammad E. H. Chowdhury |
Neural Comput. Appl. | 11 |
| 2025 | Self-CephaloNet: a two-stage novel framework using operational neural network for cephalometric analysisabstractAbstract Cephalometric analysis is essential for the diagnosis and treatment planning of orthodontics. In lateral cephalograms, however, the manual detection of anatomical landmarks is a time-consuming procedure. Deep learning solutions hold the potential to address the time constraints associated with certain tasks; however, concerns regarding their performances have been observed. To address this critical issue, we propose an end-to-end cascaded deep learning framework (Self-CephaloNet) for the task, which demonstrates benchmark performance over the ISBI 2015 dataset in predicting 19 cephalometric landmarks. Due to their adaptive nodal capabilities, Self-ONN (self-operational neural networks) demonstrates superior learning performance for complex feature spaces over conventional convolutional neural networks. To leverage this attribute, we introduce a novel self-bottleneck in the HRNetV2 (high-resolution network) backbone, which has exhibited benchmark performance on our landmark detection task. Our first-stage result surpasses previous studies, showcasing the efficacy of our singular end-to-end deep learning model, which achieves a remarkable 70.95% success rate in detecting cephalometric landmarks within a 2-mm range for the Test1 and Test2 datasets which are part of ISBI 2015 dataset. Moreover, the second stage significantly improves overall performance, yielding an impressive 82.25% average success rate for the datasets above within the same 2-mm distance. Furthermore, external validation has been conducted using the PKU cephalogram dataset. Our model demonstrates a commendable success rate of 75.95% within the 2-mm range. Md. Shaheenur Islam Sumon, Khandaker Reajul Islam, Md Sakib Abrar Hossain, Tanzila Rafique, Ranjit Ghosh, Gazi Shamim Hassan, Kanchon Kanti Podder, Noha Barhom, Faleh Tamimi, Muhammad E. H. Chowdhury |
Neural Comput. Appl. | 10 |
| 2025 | Diffuse glioma classification with deep learning and explainability: addressing challenges in histopathology image analysis
Ishrat Jahan 0003, Rafif Mahmood Al Saady, Semir Vranic, Muhammad E. H. Chowdhury |
Soft Comput. | 4 |
| 2024 | Restoration of magnetohydrodynamic-corrupted 12-lead electrocardiogram to enhance cardiac monitoring during magnetic resonance imaging
Sakib Mahmud, Muhammad E. H. Chowdhury, Moajjem Hossain Chowdhury, Abdulrahman Alqahtani, Zaid Bin Mahbub, Faycal Bensaali, Serkan Kiranyaz |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Restoration of motion-corrupted EEG signals using attention-guided operational CycleGAN
Sakib Mahmud, Muhammad E. H. Chowdhury, Serkan Kiranyaz, Nasser Al-Emadi, Anas M. Tahir, Md. Shafayet Hossain, Amith Khandakar, Somaya Al-Máadeed |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Enhancing intima-media complex segmentation with a multi-stage feature fusion-based novel deep learning framework
Rusab Sarmun, Saidul Kabir, Johayra Prithula, Abdulrahman Alqahtani, Sohaib Bassam Zoghoul, Israa Al-Hashimi, Adam Mushtak, Muhammad E. H. Chowdhury |
Eng. Appl. Artif. Intell. | 8 |
| 2024 | A novel deep learning technique for morphology preserved fetal ECG extraction from mother ECG using 1D-CycleGAN
Promit Basak, A. H. M. Nazmus Sakib, Muhammad E. H. Chowdhury, Nasser Al-Emadi, Huseyin Cagatay Yalcin, Shona Pedersen, Sakib Mahmud, Serkan Kiranyaz, Somaya Al-Máadeed |
Expert Syst. Appl. | 3 |
| 2024 | Deep learning in computed tomography pulmonary angiography imaging: A dual-pronged approach for pulmonary embolism detectionabstractThe increasing reliance on Computed Tomography Pulmonary Angiography (CTPA) for Pulmonary Embolism (PE) diagnosis presents challenges and a pressing need for improved diagnostic solutions. The primary objective of this study is to leverage deep learning techniques to enhance the Computer Assisted Diagnosis (CAD) of PE. With this aim, we propose a classifier-guided detection approach that effectively leverages the classifier’s probabilistic inference to direct the detection predictions, marking a novel contribution in the domain of automated PE diagnosis. Our classification system includes an Attention-Guided Convolutional Neural Network (AG-CNN) that uses local context by employing an attention mechanism. This approach emulates a human expert's attention by looking at both global appearances and local lesion regions before making a decision. The classifier demonstrates robust performance on the FUMPE dataset, achieving an AUROC of 0.927, sensitivity of 0.862, specificity of 0.879, and an F1-score of 0.805 with the Inception-v3 backbone architecture. Moreover, AG-CNN outperforms the baseline DenseNet-121 model, achieving an 8.1% AUROC gain. While previous research has mostly focused on finding PE in the main arteries, our use of cutting-edge object detection models and ensembling techniques greatly improves the accuracy of detecting small embolisms in the peripheral arteries. Finally, our proposed classifier-guided detection approach further refines the detection metrics, contributing new state-of-the-art to the community: mAP50, sensitivity, and F1-score of 0.846, 0.901, and 0.779, respectively, outperforming the former benchmark with a significant 3.7% improvement in mAP50. Our research aims to elevate PE patient care by integrating AI solutions into clinical workflows, highlighting the potential of human-AI collaboration in medical diagnostics. Fabiha Bushra, Muhammad E. H. Chowdhury, Rusab Sarmun, Saidul Kabir, Menatalla Said, Sohaib Bassam Zoghoul, Adam Mushtak, Israa Al-Hashimi, Abdulrahman Alqahtani, Anwarul Hasan |
Expert Syst. Appl. | 2 |
| 2024 | The utility of a deep learning-based approach in Her-2/neu assessment in breast cancerabstractHER-2/neu is a protein present on the surface of specific cancer cells and has been linked to the development and progression of certain cancer types. It is present in 15 to 20% of breast cancers and is clinically significant due to the availability of multiple anti-Her2 treatment options. Immunohistochemistry (IHC) is the most commonly used method to evaluate and quantify the expression of Her-2/neu. Although IHC is well-standardized in clinical practice, it is still subjected to inter-observer variability. Automating Her-2/neu scoring can improve accuracy, efficiency, consistency, and cost-effectiveness while reducing pathologists' workload. A deep learning-based automatic framework was utilized for the automatic detection of Her-2/neu score from whole slide images (WSI). The framework consists of three phases: identification of tumor patches, scoring of tumor patches, and Her-2/neu score prediction for whole slide images (WSI) based on the distribution of each score. This work used the dataset from the University of Warwick HER2 challenge contest. Two expert pathologists evaluated all 86 WSIs and assigned Her-2/neu scores to them. In addition, patches were generated from 50 WSIs and annotated individually by the pathologists. A total of 6641 extracted patches were generated out of which, 947 were labeled as 0, 327 as 1+, 1401 as 2+, 2950 as 3+, and 1016 were marked for discarding. Four pre-trained image classification models, namely DenseNet201, GoogleNet, MobileNet_v2, and a Vision Transformer based model, were fine-tuned, and tested on the generated patches. In order to predict the Her-2/neu score of the entire WSI, a random forest classifier was trained to predict the Her-2/neu score from the percentages of patches of each score present in the whole slide image. In patch classification performances, the vision transformer-based model outperformed the other models by achieving an accuracy of 92.6% on tumor patch classification and 91.15% on patch score classification. The random forest classifier achieved an accuracy of 88% on four scores (0, 1+, 2+ and 3+) classification and 96% on three score classification (0/1+, 2+ and 3+). The proposed deep learning-based framework for the automatic detection and evaluation of Her-2/neu expression in breast cancer obtained encouraging results. This framework has the potential to be used as a prognostic tool, providing a cost-effective and time-efficient alternative for generating clinically relevant results. However, additional research is required to assess the applicability of this pipeline in different contexts. Saidul Kabir, Semir Vranic, Rafif Mahmood Al Saady, Muhammad Salman Khan 0001, Rusab Sarmun, Abdulrahman Alqahtani, Tariq O. Abbas, Muhammad E. H. Chowdhury |
Expert Syst. Appl. | 8 |
| 2024 | Automated grading of prenatal hydronephrosis severity from segmented kidney ultrasounds using deep learning
Sakib Mahmud, Tariq O. Abbas, Muhammad E. H. Chowdhury, Adam Mushtak, Saidul Kabir, Sreekumar Muthiyal, Alaa Koko, Ahmed Balla Abdalla Altyeb, Abdulrahman Alqahtani, Amith Khandakar, Sheikh Mohammed Shariful Islam |
Expert Syst. Appl. | 3 |
| 2024 | Wearable wrist to finger photoplethysmogram translation through restoration using super operational neural networks based 1D-CycleGAN for enhancing cardiovascular monitoringabstractPhysiological signals, such as the Photoplethysmogram (PPG) collected through wearable devices, consistently encounter significant motion artifacts. Current signal processing techniques, and even state-of-the-art machine learning algorithms, frequently struggle to effectively restore the inherent bodily signals amidst the array of randomly generated distortions. This often leads to the modification or even the degradation of the underlying physiological information. To enhance heart rate estimation from wrist PPG (wPPG) signals, this study introduces the Translation Through Restoration GAN (TTR-GAN). TTR-GAN comprises cascaded dual-stage 1D Cycle Generative Adversarial Networks (1D-CycleGANs) constructed using Super-ONNs. In the first phase, corrupted wPPG waveforms are blindly restored using a 1D-CycleGAN-based restoration framework. Subsequently, in the second phase, the restored wPPG waveforms are translated into clean finger PPG (fPPG) signals through a 1D-CycleGAN-based signal-to-signal translation or synthesis framework. Both the restorer and translator GANs undergo independent evaluation using robust temporal, spectral, and clinical metrics. The application of the multipass restoration scheme to the wPPG signals resulted in significantly lower entropy compared to the raw wPPGs, indicating reduced irregularity. Using the proposed PRTX metric to evaluate the translational ability of the multichannel translator CycleGAN, we achieved a substantial improvement of 35.88% in wrist-to-finger PPG translation. The correlation between the pulse rate and pulse rate variations estimated from the generated fPPG signals and the heart rate and heart rate variability readings from the ground truth ECG improved by approximately 10.4% and 14.7%, respectively, when compared to the raw wPPG signals. The proposed TTR-GAN can be implemented in wearable devices to obtain reliable real-time cardiovascular data during daily activities. Sakib Mahmud, Muhammad E. H. Chowdhury, Serkan Kiranyaz, Malisha Islam Tapotee, Purnata Saha, Anas M. Tahir, Amith Khandakar, Abdulrahman Alqahtani |
Expert Syst. Appl. | 2 |
| 2024 | A Versatile and Wireless Multichannel Capacitive EMG Measurement System for Digital HealthcareabstractTransforming existing electromyography (EMG) measurement system into portable and wearable devices is key to fuel the revolution of digital healthcare and rehabilitation. Conventional EMG measurement systems that rely on invasive needle electrodes and non-invasive wet and dry contact electrodes are impractical to telehealth applications. Existing capacitive electromyography (cEMG) measurement systems presented by various research groups are typically designed with multi-stage analog front-end circuitry and complex data acquisition (DAQ) systems. This paper proposed a simple and versatile multichannel wireless cEMG measurement system. It consists of flexible cEMG biomedical sensors, a low-noise wireless DAQ module, and a moving average of the squared data (MASq) signal processing techniques. The proposed flexible capacitive biomedical sensor can be insulated by porous and non-porous materials. Only two electrodes are needed to acquire raw EMG signals while achieving low common-mode noise. Overall, the system achieves a high mean pulse signal-to-noise ratio (PSNR) of 13.6 (polyimide film) and 8.3 (micropore). It recorded a linear correlation between the mean root-mean-square (RMS) of the MASq data and muscle strength with a step size of 1 kg. The total power consumption of this system is 41 mW with five EMG input channels, averaging 8 mW each. This low-noise and low power consumption characteristic is ideal for battery-based wearable devices. Charn Loong Ng, Mamun Bin Ibne Reaz, Maria Liz Crespo, Andres Cicuttin, Mohd Ibrahim Shapiai, Sawal Ali, Muhammad E. H. Chowdhury |
IEEE Internet Things J. | 7 |
| 2024 | Enhance data availability and network consistency using artificial neural network for IoT
Mujahid Tabassum, Sundresan Perumal, Saad Bin Abul Kashem, Ponnan Suresh, Chinmay Chakraborty, Muhammad E. H. Chowdhury, Amith Khandakar |
Multim. Tools Appl. | 6 |
| 2024 | Robust and novel attention guided MultiResUnet model for 3D ground reaction force and moment prediction from foot kinematicsabstractAbstract Ground reaction force and moment (GRF&M) measurements are vital for biomechanical analysis and significantly impact the clinical domain for early abnormality detection for different neurodegenerative diseases. Force platforms have become the de facto standard for measuring GRF&M signals in recent years. Although the signal quality achieved from these devices is unparalleled, they are expensive and require laboratory setup, making them unsuitable for many clinical applications. For these reasons, predicting GRF&M from cheaper and more feasible alternatives has become a topic of interest. Several works have been done on predicting GRF&M from kinematic data captured from the subject’s body with the help of motion capture cameras. The problem with these solutions is that they rely on markers placed on the whole body to capture the movements, which can be very infeasible in many practical scenarios. This paper proposes a novel deep learning-based approach to predict 3D GRF&M from only 5 markers placed on the shoe. The proposed network “Attention Guided MultiResUNet” can predict the force and moment signals accurately and reliably compared to the techniques relying on full-body markers. The proposed deep learning model is tested on two publicly available datasets containing data from 66 healthy subjects to validate the approach. The framework has achieved an average correlation coefficient of 0.96 for 3D ground reaction force prediction and 0.86 for 3D ground reaction momentum prediction in cross-dataset validation. The framework can provide a cheaper and more feasible alternative for predicting GRF&M in many practical applications. Md. Ahasan Atick Faisal, Sakib Mahmud, Muhammad E. H. Chowdhury, Amith Khandakar, Mosabber Uddin Ahmed, Abdulrahman Alqahtani, Mohammed Alhatou |
Neural Comput. Appl. | 3 |
| 2024 | A novel approach for Parkinson's disease detection using Vold-Kalman order filtering and machine learning algorithmsabstractAbstract Parkinson’s disease (PD) is the second most common neurological disorder caused by damage to dopaminergic neurons. Therefore, it is important to develop systems for early and automatic diagnosis of PD. For this purpose, a study that will contribute to the development of systems for the automatic diagnosis of PD is presented. The Electroencephalography (EEG) signals were decomposed into sub-bands using adaptive decomposition methods, such as empirical mode decomposition, variational mode decomposition, and Vold-Kalman order filtering (VKF). Various features were extracted from the sub-band decomposed signals, and the significant ones were determined by Chi-squared test. These important features were applied as input to support vector machine (SVM), fitch neural network (FNN), k-nearest neighbours (KNN), and decision trees (DT), machine learning (ML) models and classification was performed. We analysed the performance of ML models by obtaining accuracy, sensitivity, specificity, positive predictive value, negative predictive values, F1-score, false-positive rate, kappa statistics, and area under the curve. The classification process was performed for two cases: PD ON-HC and PD OFF-HC groups. The most successful method in this study was the VKF method, which was applied for the first time in this field with the approach specified for both cases. In both instances, the SVM algorithm was employed as the ML model, with classifier performance criterion values close to 100%. The results obtained in this study seem to be successful compared to the results of recent research on the diagnosis of PD. Fatma Latifoglu, Sultan Penekli, Firat Orhanbulucu, Muhammad E. H. Chowdhury |
Neural Comput. Appl. | 4 |
| 2023 | NDDNet: a deep learning model for predicting neurodegenerative diseases from gait pattern
Md. Ahasan Atick Faisal, Muhammad E. H. Chowdhury, Zaid Bin Mahbub, Shona Pedersen, Mosabber Uddin Ahmed, Amith Khandakar, Mohammed Alhatou, Mohammad Nabil, Iffat Ara, Enamul Hoque Bhuiyan, Sakib Mahmud, Mohammed AbdulMoniem |
Appl. Intell. | 2 |
| 2023 | PCovNet+: A CNN-VAE anomaly detection framework with LSTM embeddings for smartwatch-based COVID-19 detection
Farhan Fuad Abir, Muhammad E. H. Chowdhury, Malisha Islam Tapotee, Adam Mushtak, Amith Khandakar, Sakib Mahmud, Anwarul Hasan |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Fetal ECG extraction from maternal ECG using deeply supervised LinkNet++ model
Arafat Rahman, Sakib Mahmud, Muhammad E. H. Chowdhury, Huseyin Cagatay Yalcin, Amith Khandakar, Onur Mutlu, Zaid Bin Mahbub, Reema Youssef Kamal, Shona Pedersen |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | RamanNet: a generalized neural network architecture for Raman spectrum analysisabstractAbstract Raman spectroscopy provides a vibrational profile of the molecules and thus can be used to uniquely identify different kinds of materials. This sort of molecule fingerprinting has thus led to the widespread application of Raman spectrum in various fields like medical diagnosis, forensics, mineralogy, bacteriology, virology, etc. Despite the recent rise in Raman spectra data volume, there has not been any significant effort in developing generalized machine learning methods targeted toward Raman spectra analysis. We examine, experiment, and evaluate existing methods and conjecture that neither current sequential models nor traditional machine learning models are satisfactorily sufficient to analyze Raman spectra. Both have their perks and pitfalls; therefore, we attempt to mix the best of both worlds and propose a novel network architecture RamanNet. RamanNet is immune to the invariance property in convolutional neural networks (CNNs) and at the same time better than traditional machine learning models for the inclusion of sparse connectivity. This has been achieved by incorporating shifted multi-layer perceptrons (MLP) at the earlier levels of the network to extract significant features across the entire spectrum, which are further refined by the inclusion of triplet loss in the hidden layers. Our experiments on 4 public datasets demonstrate superior performance over the much more complex state-of-the-art methods, and thus, RamanNet has the potential to become the de facto standard in Raman spectra data analysis. Nabil Ibtehaz, Muhammad E. H. Chowdhury, Amith Khandakar, Serkan Kiranyaz, Mohammad Sohel Rahman, Susu M. Zughaier |
Neural Comput. Appl. | 2 |
| 2023 | MLMRS-Net: Electroencephalography (EEG) motion artifacts removal using a multi-layer multi-resolution spatially pooled 1D signal reconstruction networkabstractAbstract Electroencephalogram (EEG) signals suffer substantially from motion artifacts when recorded in ambulatory settings utilizing wearable sensors. Because the diagnosis of many neurological diseases is heavily reliant on clean EEG data, it is critical to eliminate motion artifacts from motion-corrupted EEG signals using reliable and robust algorithms. Although a few deep learning-based models have been proposed for the removal of ocular, muscle, and cardiac artifacts from EEG data to the best of our knowledge, there is no attempt has been made in removing motion artifacts from motion-corrupted EEG signals:In this paper, a novel 1D convolutional neural network (CNN) called multi-layer multi-resolution spatially pooled (MLMRS) network for signal reconstruction is proposed for EEG motion artifact removal. The performance of the proposed model was compared with ten other 1D CNN models: FPN, LinkNet, UNet, UNet+, UNetPP, UNet3+, AttentionUNet, MultiResUNet, DenseInceptionUNet, and AttentionUNet++ in removing motion artifacts from motion-contaminated single-channel EEG signal. All the eleven deep CNN models are trained and tested using a single-channel benchmark EEG dataset containing 23 sets of motion-corrupted and reference ground truth EEG signals from PhysioNet. Leave-one-out cross-validation method was used in this work. The performance of the deep learning models is measured using three well-known performance matrices viz. mean absolute error (MAE)-based construction error, the difference in the signal-to-noise ratio (ΔSNR), and percentage reduction in motion artifacts (η). The proposedMLMRS-Netmodel has shown the best denoising performance, producing an average ΔSNR,η, and MAE values of 26.64 dB, 90.52%, and 0.056, respectively, for all 23 sets of EEG recordings. The results reported using the proposed model outperformed all the existing state-of-the-art techniques in terms of averageηimprovement. Sakib Mahmud, Md. Shafayet Hossain, Muhammad E. H. Chowdhury, Mamun Bin Ibne Reaz |
Neural Comput. Appl. | 3 |
| 2023 | BIO-CXRNET: a robust multimodal stacking machine learning technique for mortality risk prediction of COVID-19 patients using chest X-ray images and clinical dataabstractAbstract Nowadays, quick, and accurate diagnosis of COVID-19 is a pressing need. This study presents a multimodal system to meet this need. The presented system employs a machine learning module that learns the required knowledge from the datasets collected from 930 COVID-19 patients hospitalized in Italy during the first wave of COVID-19 (March–June 2020). The dataset consists of twenty-five biomarkers from electronic health record and Chest X-ray (CXR) images. It is found that the system can diagnose low- or high-risk patients with an accuracy, sensitivity, and F1-score of 89.03%, 90.44%, and 89.03%, respectively. The system exhibits 6% higher accuracy than the systems that employ either CXR images or biomarker data. In addition, the system can calculate the mortality risk of high-risk patients using multivariate logistic regression-based nomogram scoring technique. Interested physicians can use the presented system to predict the early mortality risks of COVID-19 patients using the web-link: Covid-severity-grading-AI. In this case, a physician needs to input the following information: CXR image file, Lactate Dehydrogenase (LDH), Oxygen Saturation (O2%), White Blood Cells Count, C-reactive protein, and Age. This way, this study contributes to the management of COVID-19 patients by predicting early mortality risk. Tawsifur Rahman, Muhammad E. H. Chowdhury, Amith Khandakar, Zaid Bin Mahbub, Md Sakib Abrar Hossain, Abraham Alhatou, Eynas Abdalla, Sreekumar Muthiyal, Khandaker F. Islam, Saad Bin Abul Kashem, Muhammad Salman Khan 0001, Susu M. Zughaier, Muhammad Maqsud Hossain |
Neural Comput. Appl. | 2 |
| 2023 | Robust Peak Detection for Holter ECGs by Self-Organized Operational Neural NetworksabstractAlthough numerous R-peak detectors have been proposed in the literature, their robustness and performance levels may significantly deteriorate in low-quality and noisy signals acquired from mobile electrocardiogram (ECG) sensors, such as Holter monitors. Recently, this issue has been addressed by deep 1-D convolutional neural networks (CNNs) that have achieved state-of-the-art performance levels in Holter monitors; however, they pose a high complexity level that requires special parallelized hardware setup for real-time processing. On the other hand, their performance deteriorates when a compact network configuration is used instead. This is an expected outcome as recent studies have demonstrated that the learning performance of CNNs is limited due to their strictly homogenous configuration with the sole linear neuron model. This has been addressed by operational neural networks (ONNs) with their heterogenous network configuration encapsulating neurons with various nonlinear operators. In this study, to further boost the peak detection performance along with an elegant computational efficiency, we propose 1-D Self-Organized ONNs (Self-ONNs) with generative neurons. The most crucial advantage of 1-D Self-ONNs over the ONNs is their self-organization capability that voids the need to search for the best operator set per neuron since each generative neuron has the ability to create the optimal operator during training. The experimental results over the China Physiological Signal Challenge-2020 (CPSC) dataset with more than one million ECG beats show that the proposed 1-D Self-ONNs can significantly surpass the state-of-the-art deep CNN with less computational complexity. Results demonstrate that the proposed solution achieves a 99.10% F1-score, 99.79% sensitivity, and 98.42% positive predictivity in the CPSC dataset, which is the best R-peak detection performance ever achieved. Moncef Gabbouj, Serkan Kiranyaz, Junaid Malik, Muhammad Uzair Zahid, Turker Ince, Muhammad E. H. Chowdhury, Amith Khandakar, Anas M. Tahir |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Osegnet: Operational Segmentation Network for Covid-19 Detection Using Chest X-Ray ImagesabstractCoronavirus disease 2019 (COVID-19) has been diagnosed automatically using Machine Learning algorithms over chest X-ray (CXR) images. However, most of the earlier studies used Deep Learning models over scarce datasets bearing the risk of overfitting. Additionally, previous studies have revealed the fact that deep networks are not reliable for classification since their decisions may originate from irrelevant areas on the CXRs. Therefore, in this study, we propose Operational Segmentation Network (OSegNet) that performs detection by segmenting COVID-19 pneumonia for a reliable diagnosis. To address the data scarcity encountered in training and especially in evaluation, this study extends the largest COVID-19 CXR dataset: QaTa-COV19 with 121,378 CXRs including 9258 COVID-19 samples with their corresponding ground-truth segmentation masks that are publicly shared with the research community. Consequently, OSegNet has achieved a detection performance with the highest accuracy of 99.65% among the state-of-the-art deep models with 98.09% precision. Aysen Degerli, Serkan Kiranyaz, Muhammad E. H. Chowdhury, Moncef Gabbouj |
ICIP | 3 |
| 2022 | Custom Hardware Architectures for Deep Learning on Portable Devices: A ReviewabstractThe staggering innovations and emergence of numerous deep learning (DL) applications have forced researchers to reconsider hardware architecture to accommodate fast and efficient application-specific computations. Applications, such as object detection, image recognition, speech translation, as well as music synthesis and image generation, can be performed with high accuracy at the expense of substantial computational resources using DL. Furthermore, the desire to adopt Industry 4.0 and smart technologies within the Internet of Things infrastructure has initiated several studies to enable on-chip DL capabilities for resource-constrained devices. Specialized DL processors reduce dependence on cloud servers, improve privacy, lessen latency, and mitigate bandwidth congestion. As we reach the limits of shrinking transistors, researchers are exploring various application-specific hardware architectures to meet the performance and efficiency requirements for DL tasks. Over the past few years, several software optimizations and hardware innovations have been proposed to efficiently perform these computations. In this article, we review several DL accelerators, as well as technologies with emerging devices, to highlight their architectural features in application-specific integrated circuit (IC) and field-programmable gate array (FPGA) platforms. Finally, the design considerations for DL hardware in portable applications have been discussed, along with some deductions about the future trends and potential research directions to innovate DL accelerator architectures further. By compiling this review, we expect to help aspiring researchers widen their knowledge in custom hardware architectures for DL. Kh Shahriya Zaman, Mamun Bin Ibne Reaz, Sawal Ali, Ahmad Ashrif A. Bakar, Muhammad E. H. Chowdhury |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Reliable Covid-19 Detection using Chest X-Ray ImagesabstractCoronavirus disease 2019 (COVID-19) has emerged the need for computer-aided diagnosis with automatic, accurate, and fast algorithms. Recent studies have applied Machine Learning algorithms for COVID-19 diagnosis over chest X-ray (CXR) images. However, the data scarcity in these studies prevents a reliable evaluation with the potential of overfitting and limits the performance of deep networks. Moreover, these networks can discriminate COVID-19 pneumonia usually from healthy subjects only or occasionally, from limited pneumonia types. Thus, there is a need for a robust and accurate COVID-19 detector evaluated over a large CXR dataset. To address this need, in this study, we propose a reliable COVID-19 detection network: ReCovNet, which can discriminate COVID-19 pneumonia from 14 different thoracic diseases and healthy subjects. To accomplish this, we have compiled the largest COVID-19 CXR dataset: QaTa-COV19 with 124,616 images including 4603 COVID-19 samples. The proposed ReCovNet achieved a detection performance with 98.57% sensitivity and 99.77% specificity. Aysen Degerli, Mete Ahishali, Serkan Kiranyaz, Muhammad E. H. Chowdhury, Moncef Gabbouj |
ICIP | 4 |
| 2021 | Convolutional Sparse Support Estimator-Based COVID-19 Recognition From X-Ray ImagesabstractCoronavirus disease (COVID-19) has been the main agenda of the whole world ever since it came into sight. X-ray imaging is a common and easily accessible tool that has great potential for COVID-19 diagnosis and prognosis. Deep learning techniques can generally provide state-of-the-art performance in many classification tasks when trained properly over large data sets. However, data scarcity can be a crucial obstacle when using them for COVID-19 detection. Alternative approaches such as representation-based classification [collaborative or sparse representation (SR)] might provide satisfactory performance with limited size data sets, but they generally fall short in performance or speed compared to the neural network (NN)-based methods. To address this deficiency, convolution support estimation network (CSEN) has recently been proposed as a bridge between representation-based and NN approaches by providing a noniterative real-time mapping from query sample to ideally SR coefficient support, which is critical information for class decision in representation-based techniques. The main premises of this study can be summarized as follows: 1) A benchmark X-ray data set, namely QaTa-Cov19, containing over 6200 X-ray images is created. The data set covering 462 X-ray images from COVID-19 patients along with three other classes; bacterial pneumonia, viral pneumonia, and normal. 2) The proposed CSEN-based classification scheme equipped with feature extraction from state-of-the-art deep NN solution for X-ray images, CheXNet, achieves over 98% sensitivity and over 95% specificity for COVID-19 recognition directly from raw X-ray images when the average performance of 5-fold cross validation over QaTa-Cov19 data set is calculated. 3) Having such an elegant COVID-19 assistive diagnosis performance, this study further provides evidence that COVID-19 induces a unique pattern in X-rays that can be discriminated with high accuracy. Mehmet Yamac, Mete Ahishali, Aysen Degerli, Serkan Kiranyaz, Muhammad E. H. Chowdhury, Moncef Gabbouj |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | A New Wearable ECG Monitor Evaluation and Experimental Analysis: Proof of ConceptabstractElectrocardiogram (ECG) is an electrical activity of the heart, which can be recorded by placing electrodes near heart or on the limbs. ECG is a vital body signal, which reflects the heart health condition. This paper presents a new wearable ECG system, which can be used for long-term rhythm monitoring with the potential of increased sensitivity to detect intermittent or subclinical arrhythmia. This study presents the design and development of a wearable pervasive healthcare monitoring system by ECG measurement systems and internet of things (IoT) platform. In this design, non-intrusive healthcare system was designed based on wireless body area network (WBAN) for wide area coverage with minimum battery power to support wireless transmission. Data were transmitted via Wi-Fi to the personalized mobile system. These were integrated into a comfortable, easy to wear, and ergonomically designed armband ECG sensor system, which can acquire an ECG signal from the upper arm of the user over a period of 72 hours. Khalid Abualsaud, Muhammad E. H. Chowdhury, Abdurrazzak Gehani, Elias Yaacoub, Tamer Khattab, Jamal Hammad |
IWCMC | 2 |