Md Tanvir Islam

dblp:261/1342 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
12since 2021 · last 2026
0009-0007-9405-5684ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 SpikeRain: Towards Energy-Efficient Single Image Deraining with Spiking Neural Networks
abstract
With the rapid deployment of vision systems on edge devices, energy-efficient and temporally aware image deraining models are increasingly needed. We propose SpikeRain, a spiking neural network (SNN) that achieves competitive deraining performance with substantially lower computational cost than conventional artificial neural networks (ANNs). Unlike ANN-based approaches with dense activations and high memory demands, SpikeRain leverages the event-driven sparse-firing nature of spiking neurons for efficient temporal integration and contextual learning. Built on an encoder-decoder framework, SpikeRain incorporates three spiking native modules: a Dense Spiking Residual Block (DSRB) for temporal integration and feature reuse, a Multi-Dimensional Spiking Attention (MDSA) module to model temporal channel spatial dependencies, and an Adaptive Residual Feature Enhancement (ARFE) block with gated attention to refine salient features. Experiments on synthetic and real-world benchmarks show that SpikeRain achieves state-of-the-art PSNR and SSIM while reducing parameters by approximately 40% and FLOPs by approximately 89%, with energy efficiency on par with existing SNN methods. These results highlight the potential of SNNs for real-time low-power image restoration on neuromorphic platforms. SpikeRain code is available on GitHub.
Md Tanvir Islam, Inzamamul Alam, Sambit Bakshi, Khan Muhammad 0001, Javier Del Ser, Sangtae Ahn
WACV1
2026 Resource-Efficient Neural Network for Crop Damage Classification in Precision Agriculture
abstract
Timely and accurate crop damage classification (CDC) is vital for informed decision-making in the industry of precision agriculture. Traditional manual methods are slow and unreliable, whereas recent deep learning models, although accurate, are often too computationally intensive for resource-constrained environments. In this study, we present LNetCDC, a lightweight attention-based convolutional neural network tailored for CDC. The architecture integrates an EchoBlock for efficient feature extraction, combined with residual pathways enhanced by“Channelwise Refine”and“Dual Gate Attention”modules to emphasize critical spatial and channelwise features. Also, dilated convolutions are incorporated into deeper layers to capture multiscale contextual patterns. We evaluated our LNetCDC on a benchmark crop damage dataset, where it outperformed existing state-of-the-art (SOTA) models in terms of both accuracy and efficiency. Notably, it achieves around 2.3% gain in accuracy with only 0.86 million parameters compared with 1.13 million in the prior SOTA model for CDC. These results demonstrate the effectiveness and suitability of LNetCDC for real-time deployment on industrial edge devices.
Md Tanvir Islam, Shehzad Ali, Abdul Khader Jilani Saudagar, Mohammad Hijji, Yazeed Alkhrijah, Khan Muhammad 0001, Victor Hugo C. de Albuquerque
IEEE Trans. Ind. Informatics1
2025 IARD: Intruder Activity Recognition Dataset for Threat Detection
abstract
Home security and surveillance systems are rapidly evolving, with Artificial Intelligence (AI) playing a transformative role in enhancing safety and threat detection. While several AI methods and datasets for intruder-related risk assessment exist, they predominantly focus on face detection and recognition, leaving a significant gap in addressing high-risk scenarios involving malicious intent, such as theft or harm. The lack of dedicated datasets for recognizing complex intruder activities, such as carrying weapons or engaging in destructive actions like kicking doors or breaking locks, limits the development of robust solutions. This work bridges this gap by introducing the Intruder Activity Recognition Dataset (IARD), a video dataset specifically designed to recognize four critical intruder activities: Armed Intruder, Door Kick, Intruder Inside and Lock Breaking. Leveraging IARD, we thoroughly benchmark various state-of-the-art methods, among which a Vision Transformer is found to achieve an impressive 93.3% accuracy in recognizing intruder actions. Our contribution highlights the potential of IARD in advancing AI-driven surveillance systems, providing a foundational dataset and benchmark for recognizing complex intruder activities.
Shehzad Ali, Md Tanvir Islam, Ikhyun Lee, Saeed Anwar, Javier Del Ser, Khan Muhammad 0001
CIKM2
2025 SpecGuard: Spectral Projection-Based Advanced Invisible Watermarking
abstract
Watermarking embeds imperceptible patterns into images for authenticity verification. However, existing methods often lack robustness against various transformations primarily including distortions, image regeneration, and adversarial perturbation, creating real-world challenges. In this work, we introduce SpecGuard, a novel watermarking approach for robust and invisible image watermarking. Unlike prior approaches, we embed the message inside hidden convolution layers by converting from the spatial domain to the frequency domain using spectral projection of a higher frequency band that is decomposed by wavelet projection. Spectral projection employs Fast Fourier Transform approximation to transform spatial data into the frequency domain efficiently. In the encoding phase, a strength factor enhances resilience against diverse attacks, including adversarial, geometric, and regeneration-based distortions, ensuring the preservation of copyrighted information. Meanwhile, the decoder leverages Parseval's theorem to effectively learn and extract the watermark pattern, enabling accurate retrieval under challenging transformations. We evaluate the proposed SpecGuard based on the embedded watermark's invisibility, capacity, and robustness. Comprehensive experiments demonstrate the proposed SpecGuard outperforms the state-of-the-art models. To ensure reproducibility, the full code is released on \href{https://github.com/inzamamulDU/SpecGuard_ICCV_2025}{\textcolor{blue}{\textbf{GitHub}}}.
Inzamamul Alam, Md Tanvir Islam, Simon S. Woo, Khan Muhammad 0001
ICCV2
2025 Leveraging DermoGrabcut Segmentation for Improved CNN-Based Skin Lesion Classification
Md Tanvir Islam, Yunfei Yin, Dayong Deng, Md Minhazul Islam, Syed Murtoza Mushrul Pasha
ICIC (26)1
2025 AI-Driven Detection and Mitigation of Backdoor Data Infiltration Among IoT Networks
abstract
The increasing popularity of Internet of Things (IoT) networks has resulted in exposure to cyberattacks, and in particular backdoor data infiltration. These attacks constitute serious threats to the confidentiality, integrity, and availability of IoT systems. This research proposes an AI-based system for detection and mitigation of backdoor data infiltration in IoT networks. The proposed system employs advanced machine learning models, including deep learning and anomaly detection methods to improve the security of IoT ecosystems. A comprehensive IoT network traffic dataset was developed that incorporates both normal and malicious (backdoor) activities. The machine learning models were then trained to distinguish between normal and malicious traffic, achieving overall high results for accuracy (96.7%), precision (97.2%), recall (95.8%) and with a low false positive rate (3.2%). A mitigation framework in the system was introduced for the purpose of responding to incidents of backdoor data infiltration that deploys automated responses such as device isolation and remote firmware restoration upon attack detection. The performance of the proposed system was evaluated in simulated and real-world IoT environments to demonstrate effectiveness in terms of the ability to detect and mitigate threats while minimizing negative impacts on network performance. A comparative analysis between proposed AI system and traditional intrusion detection systems, which demonstrated improved detection accuracy, precision, recall and lower false negatives. Therefore, a novel AI-based framework was created that proved to be a robust and scalable design in response to cyber threats in IoT networks that could potentially be integrated into existing IoT security frameworks.
Syed Murtoza Mushrul Pasha, Chunxiao Ye, Shahidur Rahoman Sohag, Muhammad Mahin Ali, Md Tanvir Islam
IJCNN5
2025 ROAD-6: A Diverse Dataset for Unexpected Hazard Recognition in Autonomous Vehicles
Shehzad Ali, Md Tanvir Islam, Minh-Son Dao, Ikhyun Lee, Shuai Liu 0009, Khan Muhammad 0001
ICMR2
2025 SpecXNet: A Dual-Domain Convolutional Network for Robust Deepfake Detection
abstract
The increasing realism of content generated by GANs and diffusion models has made deepfake detection significantly more challenging. Existing approaches often focus solely on spatial or frequency-domain features, limiting their generalization to unseen manipulations. We propose the Spectral Cross-Attentional Network (SpecXNet), a dual-domain architecture for robust deepfake detection. The core \textbf{Dual-Domain Feature Coupler (DDFC)} decomposes features into a local spatial branch for capturing texture-level anomalies and a global spectral branch that employs Fast Fourier Transform to model periodic inconsistencies. This dual-domain formulation allows SpecXNet to jointly exploit localized detail and global structural coherence, which are critical for distinguishing authentic from manipulated images. We also introduce the \textbf{Dual Fourier Attention (DFA)} module, which dynamically fuses spatial and spectral features in a content-aware manner. Built atop a modified XceptionNet backbone, we embed the DDFC and DFA modules within a separable convolution block. Extensive experiments on multiple deepfake benchmarks show that SpecXNet achieves state-of-the-art accuracy, particularly under cross-dataset and unseen manipulation scenarios, while maintaining real-time feasibility. Our results highlight the effectiveness of unified spatial-spectral learning for robust and generalizable deepfake detection. To ensure reproducibility, we released the full code on \href{https://github.com/inzamamulDU/SpecXNet}{\textcolor{blue}{\textbf{GitHub}}}.
Inzamamul Alam, Md Tanvir Islam, Simon S. Woo
ACM Multimedia2
2025 Towards Hazardous Activity Recognition for A Novel Real-World Dataset
abstract
Detecting hazardous activities is essential for ensuring safety. However, existing datasets often lack coverage of the nuanced and diverse hazards present in indoor environments, which hinders the development of a specialized model. To address this, we introduce the Real-World Hazardous Activities Dataset (RHAD), a novel and diverse video dataset specifically curated for recognizing hazardous activities in real-world indoor settings. Leveraging RHAD, we introduce HazardNet, a hybrid deep-learning architecture designed for hazardous activity recognition. HazardNet integrates local and global spatial-temporal representation modules to effectively capture complex patterns, enabling a robust understanding of the activity. We perform comprehensive evaluations by benchmarking against a range of state-of-the-art activity recognition models. Experimental results show that our proposed model performs significantly better, surpassing the latest model, VideoMamba, with a 9.2% accuracy gain. Moreover, by providing the dataset and an effective recognition model, our work lays the foundation for further research, paving the way for enhanced safety measures and preventive interventions. The dataset and code are available at https://github.com/ShehzadCS18/RHAD.
Shehzad Ali, Md Tanvir Islam, Ikhyun Lee, Mingfu Xiong, Minh-Son Dao, Saeed Anwar, Sambit Bakshi, Khan Muhammad 0001
ACM Multimedia2
2025 Resource constraint crop damage classification using depth channel shuffling
abstract
Accurate crop damage classification is crucial for timely interventions, loss reduction, and resource optimization in agriculture. However, datasets and models for binary classification of damaged versus non-damaged crops remain scarce. To address this, we conducted an extensive study on crop damage classification using deep learning, focusing on the challenges posed by imbalanced datasets common in agriculture. We began by preprocessing the “Consultative Group for International Agricultural Research (CGIAR)” dataset to enhance data quality and balance class distributions. We created the new “Crop Damage Classification (CDC)” dataset tailored for binary classification of “Damaged” versus “Non-damaged” crops, serving as an effective training medium for deep learning models. Using the CDC dataset, we benchmarked the state-of-the-art models to evaluate their effectiveness in classifying crop damage. Leveraging the depth channel shuffling technique of ShuffleNetV2, we proposed a lightweight model “Light Crop Damage Classifier (LightCDC)” , reducing the parameters from 1.40 million to 1.13 million while achieving an accuracy of 89.44%. LightCDC outperformed existing classification and ensemble models in terms of model size, parameter count, inference time, and accuracy. Furthermore, we tested LightCDC under adverse conditions like blur, low light, and fog, validating its robustness for real-world scenarios. Thus, our contributions include a refined dataset and an efficient model tailored for crop damage classification, which is essential for timely interventions and improved crop management in resource-constrained precision agriculture. To ensure reproducibility, we released the code and dataset on GitHub . • New Crop Damage Classification (CDC) dataset created for binary classification. • Custom lightweight model, LightCDC, achieved 89.44% accuracy with 1.13M parameters. • LightCDC model excels in accuracy, reduced size, and faster inference times. • Study addresses imbalanced datasets, improving real-time precision agriculture.
Md Tanvir Islam, Safkat Shahrier Swapnil, Md. Masum Billal, Asif Karim, Niusha Shafiabady, Md. Mehedi Hassan
Eng. Appl. Artif. Intell.1
2024 LoLI-Street: Benchmarking Low-Light Image Enhancement and Beyond
Md Tanvir Islam, Inzamamul Alam, Simon S. Woo, Saeed Anwar, Ikhyun Lee, Khan Muhammad 0001
ACCV (5)1
2024 HazeSpace2M: A Dataset for Haze Aware Single Image Dehazing
abstract
Reducing the atmospheric haze and enhancing image clarity is crucial for computer vision applications. The lack of real-life hazy ground truth images necessitates synthetic datasets, which often lack diverse haze types, impeding effective haze type classification and dehazing algorithm selection. This research introduces the HazeSpace2M dataset, a collection of over 2 million images designed to enhance dehazing through haze type classification. HazeSpace2M includes diverse scenes with 10 haze intensity levels, featuring Fog, Cloud, and Environmental Haze (EH). Using the dataset, we introduce a technique of haze type classification followed by specialized dehazers to clear hazy images. Unlike conventional methods, our approach classifies haze types before applying type-specific dehazing, improving clarity in real-life hazy images. Benchmarking with state-of-the-art (SOTA) models, ResNet50 and AlexNet achieve 92.75\% and 92.50\% accuracy, respectively, against existing synthetic datasets. However, these models achieve only 80% and 70% accuracy, respectively, against our Real Hazy Testset (RHT), highlighting the challenging nature of our HazeSpace2M dataset. Additional experiments show that haze type classification followed by specialized dehazing improves results by 2.41% in PSNR, 17.14% in SSIM, and 10.2\% in MSE over general dehazers. Moreover, when testing with SOTA dehazing models, we found that applying our proposed framework significantly improves their performance. These results underscore the significance of HazeSpace2M and our proposed framework in addressing atmospheric haze in multimedia processing. Complete code and dataset is available on \href{https://github.com/tanvirnwu/HazeSpace2M} {\textcolor{blue}{\textbf{GitHub}}}.
Md Tanvir Islam, Nasir Rahim, Saeed Anwar, Sambit Bakshi, Khan Muhammad 0001
ACM Multimedia1