EDBT 2026 Demo / reviewers in the wild / expert
Swalpa Kumar Roy
dblp:166/4544
· DBLP profile ↗
45ranked-venue papers
15as first author
37since 2021 · last 2026
0000-0002-6580-3977ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 12 first-author · 26 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Text-Guided Convolutional Adapter for the Diffusion ModelabstractWe introduce the Nexus Adapters, novel text-guided efficient adapters to the diffusion-based framework for the Structure Preserving Conditional Generation (SPCG). Recently, structure-preserving methods have achieved promising results in conditional image generation by using a base model for prompt conditioning and an adapter for structure input, such as sketches or depth maps. These approaches are highly inefficient and sometimes require equal parameters in the adapter compared to the base architecture. It is not always possible to train the model since the diffusion model is itself costly, and doubling the parameter is highly inefficient. In these approaches, the adapter is not aware of the input prompt; therefore, it is optimal only for the structural input but not for the input prompt. To overcome the above challenges, we proposed two efficient adapters, Nexus Prime and Slim, which are guided by prompts and structural inputs. Each Nexus Block incorporates cross-attention mechanisms to enable rich multimodal conditioning. Therefore, the proposed adapter has a better understanding of the input prompt while preserving the structure. We conducted extensive experiments on the proposed models and demonstrated that the Nexus Prime adapter significantly enhances performance, requiring only 8M additional parameters compared to the baseline, T2I-Adapter. Furthermore, we also introduced a lightweight Nexus Slim adapter with 18M fewer parameters than the T2I-Adapter, which still achieved state-of-the-art results. Code: https://github.com/arya-domain/Nexus-Adapters Aryan Das, Koushik Biswas, Swalpa Kumar Roy, Badri Narayana Patro, Vinay Kumar Verma |
WACV | 3 |
| 2026 | Uncertainty-Aware Vision-Language Segmentation for Medical ImagingabstractWe introduce a novel uncertainty-aware multimodal segmentation framework that leverages both radiological images and associated clinical text for precise medical diagnosis. We propose a Modality Decoding Attention Block (MoDAB) with a lightweight State Space Mixer (SSMix) to enable efficient cross-modal fusion and long-range dependency modelling. To guide learning under ambiguity, we propose the Spectral-Entropic Uncertainty (SEU) Loss, which jointly captures spatial overlap, spectral consistency, and predictive uncertainty in a unified objective. In complex clinical circumstances with poor image quality, this formulation improves model reliability. Extensive experiments on various publicly available medical datasets, QATA-COVID19, MosMed++, and Kvasir-SEG, demonstrate that our method achieves superior segmentation performance while being significantly more computationally efficient than existing State-of-the-Art (SoTA) approaches. Our results highlight the importance of incorporating uncertainty modelling and structured modality alignment in vision-language medical segmentation tasks. Code: https://github.com/arya-domain/UA-VLS Aryan Das, Tanishq Rachamalla, Koushik Biswas, Swalpa Kumar Roy, Vinay Kumar Verma |
WACV | 4 |
| 2026 | RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage UnderstandingabstractAccurate road damage detection is crucial for timely infrastructure maintenance and public safety, but existing vision-only datasets and models lack the rich contextual understanding that textual information can provide. To address this limitation, we introduce RoadBench, the first multimodal benchmark for comprehensive road damage understanding. This dataset pairs high-resolution images of road damages with detailed textual descriptions, providing a richer context for model training. We also present RoadCLIP, a novel vision-language model that builds upon CLIP by integrating domain-specific enhancements. It includes a disease-aware positional encoding that captures spatial patterns of road defects and a mechanism for injecting road-condition priors to refine the model’s understanding of road damages. We further employ a GPT-driven data generation pipeline to expand the image–text pairs in Road-Bench, greatly increasing data diversity without exhaustive manual annotation. Experiments demonstrate that Road-CLIP achieves state-of-the-art performance on road damage recognition tasks, significantly outperforming existing vision-only models by 19.2%. These results highlight the advantages of integrating visual and textual information for enhanced road condition analysis, setting new benchmarks for the field and paving the way for more effective infrastructure monitoring through multimodal learning. Xi Xiao 0003, Yunbei Zhang, Janet Wang, Yuxiang Wei 0004, Hengjia Li, Yanshu Li, Xiao Wang 0004, Swalpa Kumar Roy, Tianyang Wang 0004 |
WACV | 9 |
| 2026 | Revealing the human-like similarities in automated facial expression recognition: an empirical investigation using eXplainable artificial intelligence
Sayan Kumar Bhowmick, Asit Barman, Swalpa Kumar Roy, Paramartha Dutta |
Multim. Tools Appl. | 3 |
| 2025 | TD-RD: A Top-Down Benchmark with Real-Time Framework for Road Damage DetectionabstractObject detection has witnessed remarkable advancements over the past decade, largely driven by breakthroughs in deep learning and the proliferation of large-scale datasets. However, the domain of road damage detection remains relatively underexplored, despite its critical significance for applications such as infrastructure maintenance and road safety. This paper addresses this gap by introducing a novel top-down benchmark that offers a complementary perspective to existing datasets, specifically tailored for road damage detection. Our proposed Top-Down Road Damage Detection Dataset (TD-RD) includes three primary categories of road damage—cracks, potholes, and patches—captured from an top-down viewpoint. The dataset consists of 7,088 high-resolution images, encompassing 12,882 annotated instances of road damage. Additionally, we present a novel real-time object detection framework, TD-YOLOV10, designed to handle the unique challenges posed by the TD-RD dataset. Comparative studies with state-of-the-art models demonstrate competitive baseline results. By releasing TD-RD, we aim to accelerate research in this crucial area. A sample of the dataset will be made publicly available upon the paper’s acceptance. Xi Xiao 0003, Zhengji Li, Houjie Lin, Swalpa Kumar Roy, Tianyang Wang 0004, Min Xu 0009 |
ICASSP | 6 |
| 2025 | Spatial-spectral morphological mamba for hyperspectral image classification
Muhammad Ahmad 0002, Muhammad Hassaan Farooq Butt, Adil Khan 0001, Manuel Mazzara, Salvatore Distefano, Swalpa Kumar Roy, Jocelyn Chanussot, Danfeng Hong |
Neurocomputing | 7 |
| 2025 | MixerSENet: A Lightweight Framework for Efficient Hyperspectral Image ClassificationabstractIn this paper, a novel framework, MixerSENet, is introduced for hyperspectral image (HSI) classification, designed to address the challenges of computational efficiency and limited labeled data. The proposed model processes hyperspectral image patches while maintaining consistent size and resolution throughout the network, effectively decoupling the mixing of spatial and channel dimensions. Notably, MixerSENet is lightweight and computationally efficient, requiring fewer parameters compared to traditional models, making it suitable for resource-constrained environments. A squeeze and excitation block is incorporated into the model to refine feature extraction, enhancing the network’s ability to capture more informative features. Experimental results on two benchmark datasets demonstrate that MixerSENet achieves superior performance, reaching an overall accuracy (OA) of 82.47% on Houston13 dataset and 96.70% on the Qingyun dataset, outperforming state-of-the-art methods including 3D-CNN, HybridKAN, HSIFormer, SimPoolFormer, and MorphMamba. Furthermore, a detailed analysis of computational efficiency shows that MixerSENet achieves a favorable balance between accuracy and efficiency, with only 53,146 parameters and an low inference time, confirming its practicality for real-world applications. At publication, source code will be publicly available at https://github.com/mqalkhatib/MixerSENet. Mohammed Q. Alkhatib, Swalpa Kumar Roy, Ali Jamali |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | MSHCCT: A Multiscale Compact Convolutional Network for High-Resolution Aerial Scene ClassificationabstractThe growing popularity of vision transformers (ViTs) in remote sensing image classification is due to their ability to effectively capture long-range dependencies. However, their high computational cost and memory footprint limit their applicability, particularly for small-scale datasets and resource-constrained environments. To address these challenges, we propose the multiscale multihead compact convolutional transformer (MSHCCT), a lightweight yet powerful model that integrates convolutional tokenization with small-scale ViTs to enhance multiscale feature representation while maintaining computational efficiency. Despite a modest increase in parameters and training time, MSHCCT achieves superior classification accuracy and robustness on high-resolution aerial scenes. Importantly, our approach eliminates the need for model pretraining, additional datasets, or multisensor data fusion, ensuring a computationally efficient and practical solution for remote sensing applications. The code will be made publicly available athttps://github.com/aj1365/MSHCCT Ali Jamali, Swalpa Kumar Roy, Bing Lu 0003, Leila Hashemi Beni, Nafiseh Kakhani, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Are Vision-xLSTM-embedded U-Nets better at segmenting medical images?
Pallabi Dutta, Soham Bose, Swalpa Kumar Roy, Sushmita Mitra |
Neural Networks | 3 |
| 2025 | C2FNet: Cross-Probabilistic Weak Supervision Learning for High-Resolution Land Cover EnhancementabstractAutonomous large-scale high-resolution land cover (HRLC) mapping remains a major challenge in remote sensing due to the scarcity of reliable training data and resolution mismatches between available labels and input of massive of emerging imagery. Existing global land cover products often suffer from coarse spatial resolution and label noise, limiting their utility for fine-scale urban analysis and environmental monitoring. This article presents C2FNet, a novel Coarse-to-Fine Network designed to generate HRLC maps from noisy, coarse-resolution labels using a weak supervision strategy of the cross-probability. The C2FNet consists of three key modules: 1) edge resolution refinement backbones (ERRBs), which preserve spatial detail via multiscale feature extraction through parallel convolutional branches; 2) unsupervised dynamic shuffle and diagonal annotation (UDSDA), which enhances training reliability by identifying confident regions through spatial-consistency analysis and confidence estimation; and 3) a contrasting self-supervised loss (C2F-Loss) that integrates cross-entropy and cosine similarity terms to mitigate supervision noise and resolution gaps. Evaluations of three benchmark datasets that encompass diverse urban and rural landscapes show that C2FNet achieves state-of-the-art (SoA) performance, with 80.01% overall accuracy (OA) and a Cohen’s kappa score of 0.7567, outperforming SoA models with weak supervision. The dataset and code are available athttp://drive.google.com/file/d/1X_Fz7LQIeix3rV3K29FBfKiU1WMdROe-/view Boaz Mwubahimana, Jianguo Yan, Dingruibo Miao, Zhuohong Li, Maurice Mugabowindekwe, Swalpa Kumar Roy, Xiao Huang 0003, Elias Nyandwi, Tuyishimire Joseph, Eric Habineza, Fidele Mwizerwa, Hafashimana Athanase, Gaspard Rwanyiziri |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Wavelet-Infused Convolution-Transformer for Efficient Segmentation in Medical ImagesabstractRecent medical image segmentation methods extract the characteristics of anatomical structures only from the spatial domain, ignoring the distinctive patterns present in the spectral representation. This study aims to develop a novel segmentation architecture that leverages both spatial and spectral characteristics for better segmentation outcomes. This research introduces the wavelet-infused convolutional Transformer (WaveCoformer), a computationally effective framework to fuse information from both spatial and spectral domains of medical images. Fine-grained textural features are captured from the wavelet components by the convolution module. A transformer block identifies the relevant activation maps within the volumes, followed by self-attention to effectively learn long-range dependencies to capture the global context of the target regions. A cross-attention mechanism effectively combines the distinctive features acquired by both modules to produce a comprehensive and robust representation of the input data. WaveCoformer outperforms related state-of-the-art networks in publicly available Synapse and Adrenal tumor segmentation datasets, with a mean Dice score of 83.86% and 79%, respectively. The model is feasible for deployment in resource-constrained environments with rapid medical image analysis due to its computationally efficient nature and improved segmentation performance. The code is available at:https://github.com/duttapallabi2907/WaveCoformer. Pallabi Dutta, Sushmita Mitra, Swalpa Kumar Roy |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Multi-dimension Transformer with Attention-based Filtering for Medical Image SegmentationabstractThe accurate segmentation of medical images is crucial for diagnosing and treating diseases. Recent studies demonstrate that vision transformer-based methods have significantly improved performance in medical image segmentation, primarily due to their superior ability to establish global relationships among features and adaptability to various inputs. However, these methods struggle with the low signal-to-noise ratio inherent to medical images. Additionally, the effective utilization of channel and spatial information, which are essential for medical image segmentation, is limited by the representation capacity of self-attention. To address these challenges, we propose a Multi-dimension Transformer with Attention-based Filtering (MDT-AF), which redesigns the patch embedding and self-attention mechanism for medical image segmentation. MDT-AF incorporates an attention-based feature filtering mechanism into the patch embedding blocks and employs a coarse-to-fine process to mitigate the impact of a low signal-to-noise ratio. To better capture complex structures in medical images, MDT-AF extends self-attention and introduces an interaction mechanism to build and enhance feature relationships between dimensions, which can achieve richer feature representations across the spatial and channel dimensions. Experimental results on three public medical image segmentation benchmarks show that MDT-AF achieves state-of-the-art (SOTA) performance. Xi Xiao 0003, Qizhen Lan, Xuanyao Huang, Qing Tian 0003, Swalpa Kumar Roy, Tianyang Wang 0004 |
ICTAI | 7 |
| 2024 | PolSARConvMixer: A Channel and Spatial Mixing Convolutional Algorithm for PolSAR Data ClassificationabstractGiven the exceptional effectiveness of deep Convolutional Neural Networks (CNNs) in computer vision, there has been a recent surge of interest in employing CNNs for various applications in image classification. Additionally, scientists are exploring the potential of vision transformers for Earth observation applications, owing to their recent tremendous success. However, a major challenge with vision transformers is their increased demand for training data compared to CNN classifiers. Furthermore, vision transformers exhibit quadratic complexity and necessitate substantial hardware resources. In the context of PolSAR image classification, we propose the PolSARConvMixer—a fundamental framework that segregates the mixing of spatial and channel dimensions, maintains uniform size and resolution across the network and directly processes PolSAR image patches as input. Our experiments on two PolSAR data benchmarks, namely Flevoland and San Francisco, demonstrate the significant superiority of the developed PolSARConvMixer over several other algorithms, including AlexNet, ResNet, FNet, a 2D CNN, and PolSARFormer. Ali Jamali, Swalpa Kumar Roy, Bing Lu 0003, Avik Bhattacharya, Pedram Ghamisi |
IGARSS | 2 |
| 2024 | Spatial-Gated Multilayer Perceptron for Land Use and Land Cover MappingabstractDue to its capacity to recognize detailed spectral differences, hyperspectral data have been extensively used for precise Land Use Land Cover (LULC) mapping. However, recent multi-modal methods have shown their superior classification performance over the algorithms that use single data sets. On the other hand, Convolutional Neural Networks (CNNs) are models extensively utilized for the hierarchical extraction of features. Vision transformers (ViTs), through a self-attention mechanism, have recently achieved superior modeling of global contextual information compared to CNNs. However, to harness their image classification strength, ViTs require substantial training datasets. In cases where the available training data is limited, current advanced multi-layer perceptrons (MLPs) can provide viable alternatives to both deep CNNs and ViTs. In this paper, we developed the SGU-MLP, a deep learning algorithm that effectively combines MLPs and spatial gating units (SGUs) for precise Land Use Land Cover (LULC) mapping using multi-modal data from multi-spectral, LiDAR, and hyperspectral data. Results illustrated the superiority of the developed SGU-MLP classification algorithm over several CNN and CNN-ViT-based models, including HybridSN, ResNet, iFormer, EfficientFormer, and CoAtNet. The SGU-MLP classification model consistently outperformed the benchmark CNN and CNN-ViT-based algorithms. The code will be made publicly available at https: //github.com/aj1365/SGUMLP. Ali Jamali, Swalpa Kumar Roy, Danfeng Hong, Peter M. Atkinson, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Attention Graph Convolutional Network for Disjoint Hyperspectral Image ClassificationabstractConvolutional Neural Networks (CNNs) are employed extensively in remote sensing due to their capacity to capture intricate features from a broad range of object patterns, irrespective of object size, shape or color. These networks excel at extracting high-frequency spectral information such as angles, edges and outlines. The classification boundary zone, however, becomes hazy for CNNs because they learn characteristics by means of a fixed shape kernel concentrated on the central pixel, and can perform poorly in image classification at class boundaries. Additionally, CNNs are not designed to capture global relations. Thus, in this letter, we propose an Attention Graph Convolutional Network (Attention-GCN) as a solution to the aforementioned shortcomings. The developed model illustrated a high level of superiority over several CNN and ViT-based models. For example, in the Augsburg data benchmark, the developed algorithm exhibited an average accuracy of 61.11%, substantially outperforming other models such as HybridSN, iFormer, Efficient Former, GCN, CoAtNet, 2D-CNN, 3D-CNN, and ResNet by approximately 9, 13, 14, 15, 18, 24, 25 and 29 percentage points, respectively. The code will be made publicly available at https://github.com/aj1365/AGCN. Ali Jamali, Swalpa Kumar Roy, Danfeng Hong, Peter M. Atkinson, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Neighborhood Attention Makes the Encoder of ResUNet Stronger for Accurate Road ExtractionabstractIn the domain of remote sensing image interpretation, road extraction from high-resolution aerial imagery has already been a hot research topic. Although deep CNNs have presented excellent results for semantic segmentation, the efficiency and capabilities of vision transformers are yet to be fully researched. As such, for accurate road extraction, a deep semantic segmentation neural network that utilizes the abilities of residual learning, HetConvs, UNet, and vision transformers, which is called ResUNetFormer, is proposed in this letter. The developed ResUNetFormer is evaluated on various cutting-edge deep learning-based road extraction techniques on the public Massachusetts road dataset. Statistical and visual results demonstrate the superiority of the ResUNetFormer over the state-of-the-art CNNs and vision transformers for segmentation. The code will be made available publicly at https://github.com/aj1365/ResUNetFormer. Ali Jamali, Swalpa Kumar Roy, Jonathan Li 0001, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | TL2GH²T: Triple-Path Local-to-Global Network With Hybrid Head Transformer for Hyperspectral Change DetectionabstractWith the aid of transformers, significant progress has been achieved in hyperspectral image change detection (HSI-CD) in recent times. Nonetheless, most contemporary detection methods fail to incorporate diverse diagnostic features extracted from hyperspectral (HS) images. In addition, relying solely on algebraic-based techniques to extract information of difference is insufficient for achieving satisfactory detection performance. In this regard, we propose an innovative triple-path local-to-global network (TL2GN), complemented by a hybrid head transformer (HybridHT), called TL2GH2T, tailored for HSI-CD tasks. To be specific, TL2GH2T first investigates spatial, spectral, and spatial–spectral features from a local-to-global perspective. Then, a novel spatial and spectral token fusion (SSTF) module is developed to integrate the above three tokenized features, producing discriminative features from two HS images separately. Moreover, drawing inspiration from chromosomal crossover mechanisms, we propose a HybridHT. Its goal is to simultaneously learn cross correlation and self-correlation information of bitemporal features from a global perspective, producing highly discriminative distinctions. Our approach, validated through extensive experimentation on four varied HS benchmarks, exhibits exceptional performance in HSI-CD, outperforming contemporary methods in both visual and quantitative evaluations. Zhonghao Chen, Swalpa Kumar Roy, Hongmin Gao 0001, Yao Ding 0010, Xiongwu Xiao, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Cross Hyperspectral and LiDAR Attention Transformer: An Extended Self-Attention for Land Use and Land Cover ClassificationabstractThe successes of attention-driven deep models like the Vision Transformer (ViT) have sparked interest in cross-domain exploration. However, current transformer-based techniques in remote sensing primarily focus on single-modal data, limiting their potential to exploit the growing array of multimodal Earth observation data fully. Enhancing these models for multimodal integration is crucial for comprehensive remote sensing applications. To achieve this, we extend the traditional self-attention mechanism by introducing Cross Hyperspectral and LiDAR (Cross-HL) attention. We present a novel multimodal deep learning framework that effectively fuses remote sensing (RS) data, intending to improve land use and land cover (LULC) recognition. To enhance the accurate exchange of information across different modalities, we fuse their patch projections using the Cross-HL self-attention module. In this process, LiDAR patch tokens serve as queries (Q), while keys (K) and values (V) are derived from HS patch tokens. To demonstrate the superiority of Cross-HL in the proposed multimodal deep learning framework, we conducted extensive experiments on three multimodal RS benchmark datasets: Houston, Trento, and MUUFL. These datasets contain hyperspectral and light detection and ranging (LiDAR) data. The source code for Cross-HL will be made available publicly at https://github.com/AtriSukul1508/Cross-HL. Swalpa Kumar Roy, Atri Sukul, Ali Jamali, Juan Mario Haut, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Mutual Channels Loss and Channel-Wise Attention Aided Convolutional Autoencoder for Hyperspectral Image UnmixingabstractHyperspectral imaging is a valuable tool for analyzing and understanding remote sensing data. Our research paper presents a pioneering method for hyperspectral unmixing through the utilization of a convolutional autoencoder (CAE). To train the CAE, our approach incorporates mutual channels loss (MCL) as a loss function, and we implement channel-wise attention within the CAE architecture. Furthermore, we conduct experiments using two datasets, namely the Samson and Apex hyperspectral datasets, to compare the outcomes of our approach against those achieved by state-of-the-art methods. Our findings demonstrate that our suggested approach achieves significant improvements in terms of accuracy, efficiency and being robust as compared to existing methods. These results highlight the potential of our approach for improving hyperspectral unmixing in a range of remote sensing applications. Nithish Reddy Banda, Mrinmoy Ghorai, Swalpa Kumar Roy |
IGARSS | 3 |
| 2023 | Hyperspectral Image Classification Based On 3d Sharpened Cosine Similarity OperationabstractHyperspectral remote sensing images contain abundant spectral information and have a broad application in various areas, such as precision agriculture and geological exploration. However, due to the existence of many bands and the correlation between these bands, it is easy to cause the curse of dimensionality. Therefore, how to efficiently extract spectral and spatial features should be addressed. In this paper, a novel and efficient sharpened cosine similarity operation is applied in hyperspectral images classification to enhance the ability of feature extraction. To validate the efficiency of the proposed method, experiments are conducted on a real hyperspectral dataset, e.g., the University of Pavia (UP). Quantitative and qualitative results demonstrate that sharpened cosine distance operation can efficiently extract discriminative features and achieve a better classification performance. Hongjing Wu, Swalpa Kumar Roy, Weimin Huang 0001 |
IGARSS | 3 |
| 2023 | TAttMSRecNet: Triplet-attention and multiscale reconstruction network for band selection in hyperspectral images
Utpal Nandi, Swalpa Kumar Roy, Danfeng Hong, Xin Wu 0001, Jocelyn Chanussot |
Expert Syst. Appl. | 2 |
| 2023 | Local Window Attention Transformer for Polarimetric SAR Image ClassificationabstractConvolutional neural networks (CNNs) have recently found great attention in image classification since deep CNNs have exhibited excellent performance in computer vision. Owing to their immense success, of late, scientists are exploring the functionality of transformers in Earth observation applications. Nevertheless, the primary issue with transformers is that they demand significantly more training data than CNN classifiers. Thus, the use of these transformers in remote sensing is considered challenging, notably in utilizing polarimetric synthetic aperture radar (PolSAR) data, due to the insufficient number of existing labeled data. In this letter, we develop and propose a vision transformer (ViT)-based framework that utilizes 3-D and 2-D CNNs as feature extractors and, in addition, local window attention (LWA) for the effective classification of PolSAR data. Extensive experimental results demonstrated that the developed modelPolSARFormerobtained better classification accuracy than the state-of-the-art vision Swin Transformer and FNet algorithms. ThePolSARFormeroutperformed the Swin Transformer and FNet by the margin of 5.86% and 17.63%, in terms of average accuracy (AA) in the San Francisco data benchmark. Moreover, the results over the Flevoland dataset illustrated that thePolSARFormerexceeds several other algorithms, including the ResNet (97.49%), Swin Transformer (96.54%), FNet (95.28%), 2-D CNN (94.57%), and AlexNet (91.83%), with a kappa index (KI) of 99.30%. The code will be made available publicly athttps://github.com/aj1365/PolSARFormer. Ali Jamali, Swalpa Kumar Roy, Avik Bhattacharya, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Learning Tensor Low-Rank Representation for Hyperspectral Anomaly DetectionabstractRecently, low-rank representation (LRR) methods have been widely applied for hyperspectral anomaly detection, due to their potentials in separating the backgrounds and anomalies. However, existing LRR models generally convert 3-D hyperspectral images (HSIs) into 2-D matrices, inevitably leading to the destruction of intrinsic 3-D structure properties in HSIs. To this end, we propose a novel tensor low-rank and sparse representation (TLRSR) method for hyperspectral anomaly detection. A 3-D TLR model is expanded to separate the LR background part represented by a tensorial background dictionary and corresponding coefficients. This representation characterizes the multiple subspace property of the complex LR background. Based on the weighted tensor nuclear norm and the$L_{F,1}$sparse norm, a dictionary is designed to make its atoms more relevant to the background. Moreover, a principal component analysis (PCA) method can be assigned as one preprocessing step to exact a subset of HSI bands, retaining enough the HSI object information and reducing computational time of the postprocessing tensorial operations. The proposed model is efficiently solved by the well-designed alternating direction method of multipliers (ADMMs). A comparison with the existing algorithms via experiments establishes the competitiveness of the proposed method with the state-of-the-art competitors in the hyperspectral anomaly detection task. Qiang Wang 0001, Danfeng Hong, Swalpa Kumar Roy, Jocelyn Chanussot |
IEEE Trans. Cybern. | 4 |
| 2023 | Parameter-Free Attention Network for Spectral-Spatial Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) comprise plenty of information in the spatial and spectral domain, which is highly beneficial for performing classification tasks in a very accurate way. Recently, attention mechanisms have been widely used in HSI classification due to their ability to extract relevant spatial and spectral features. Notwithstanding their positive results, most of the attentional strategies usually introduce a significant number of parameters to be trained, making the models more complex and increasing the computational load. In this paper, we develop a new parameter-free attention network for HSI classification. The main advantage of our model is that it does not add parameters to the original network (as opposed to other state-of-the-art approaches), whilst providing higher classification accuracies. Extensive experimental validations and quantitative comparisons are conducted –using different benchmark HSIs– to illustrate these advantages. Code is available on https://github.com/mhaut/Free2Resnet. Mercedes Eugenia Paoletti, Xuanwen Tao, Lirong Han, Zhaoyue Wu, Sergio Moreno-Álvarez, Swalpa Kumar Roy, Antonio Plaza, Juan Mario Haut |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Multiscale Neighborhood Attention Transformer With Optimized Spatial Pattern for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) provide hundreds of continuous spectral bands and have been widely used for the fine identification of targets with similar appearances. In earlier studies, convolutional neural networks (CNNs) have been an effective method for HSIs classification due to their powerful feature extraction capabilities. Recently, self-attention-based vision transformer (ViT) architecture has been widely explored to fully represent global information. However, most existing transformer-based models primarily focus on global relationships and lack the ability to capture the multi-scale features which are crucial for HSIs classification. This limitation results in inferior performance for transformer-based methods compared to state-of-the-art CNN-based models. To solve this problem, a novel network called multi-scale neighborhood attention transformer (MSNAT) is proposed in this paper. Unlike previous transformer-based models, MSNAT emphasizes the neighborhood pixels within a local window size and extracts multi-scale spatial information by using different local window sizes. In addition, a spatial transformation module is integrated to generate optimized spatial input. The effectiveness of the proposed MSNAT is verified on three real hyperspectral datasets including University of Pavia (UP), University of Houston (UH), and University of Trento (UT). Experimental results demonstrate that the proposed MSNAT method outperforms both CNNs and existing transformer-based models, achieving state-of-the-art classification performance with an overall accuracy of 93.34%, 86.26%, and 96.63% on UP, UH, and UT, respectively. The source code will be available at https://github.com/xinqiao123/MSNAT. Swalpa Kumar Roy, Weimin Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Multimodal Fusion Transformer for Remote Sensing Image ClassificationabstractVision transformers (ViTs) have been trending in image classification tasks due to their promising performance when compared to convolutional neural networks (CNNs). As a result, many researchers have tried to incorporate ViTs in hyperspectral image (HSI) classification tasks. To achieve satisfactory performance, close to that of CNNs, transformers need fewer parameters. ViTs and other similar transformers use an external classification (CLS) token which is randomly initialized and often fails to generalize well, whereas other sources of multimodal datasets, such as light detection and ranging (LiDAR) offer the potential to improve these models by means of a CLS. In this paper, we introduce a new multimodal fusion transformer (MFT) network which comprises a multihead cross patch attention (mCrossPA) for HSI land-cover classification. Our mCrossPA utilizes other sources of complementary information in addition to the HSI in the transformer encoder to achieve better generalization. The concept of tokenization is used to generate CLS and HSI patch tokens, helping to learn a distinctive representation in a reduced and hierarchical feature space. Extensive experiments are carried out on widely used benchmark datasets i.e., the University of Houston, Trento, University of Southern Mississippi Gulfpark (MUUFL), and Augsburg. We compare the results of the proposed MFT model with other state-of-the-art transformers, classical CNNs, and conventional classifiers models. The superior performance achieved by the proposed model is due to the use of multihead cross patch attention. The source code will be made available publicly at https://github.com/AnkurDeria/MFT. Swalpa Kumar Roy, Ankur Deria, Danfeng Hong, Behnood Rasti, Antonio Plaza, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Spectral-Spatial Morphological Attention Transformer for Hyperspectral Image ClassificationabstractIn recent years, convolutional neural networks (CNNs) have drawn significant attention for the classification of hyperspectral images (HSIs). Due to their self-attention mechanism, the vision transformer (ViT) provides promising classification performance compared to CNNs. Many researchers have incorporated ViT for HSI classification purposes. However, its performance can be further improved because the current version does not use spatial–spectral features. In this article, we present a new morphological transformer (morphFormer) that implements a learnable spectral and spatial morphological network, where spectral and spatial morphological convolution operations are used (in conjunction with the attention mechanism) to improve the interaction between the structural and shape information of the HSI token and theCLStoken. Experiments conducted on widely used HSIs demonstrate the superiority of the proposed morphFormer over the classical CNN models and state-of-the-art transformer models. The source will be made available publicly athttps://github.com/mhaut/morphFormer. Swalpa Kumar Roy, Ankur Deria, Chiranjibi Shah, Juan Mario Haut, Qian Du 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Reinforcement Learning for Neural Architecture Search in Hyperspectral UnmixingabstractIn this letter, a novel neural architecture search (NAS) method based on reinforcement learning, called RLNAS, is devised to realize the automatic architecture design in the field of hyperspectral unmixing (HU). This method first train the search network in the constructed self-supervised datasets based on hyperspectral images. The block-based searching and weight-sharing strategies are then introduced to reduce the computational cost in the training phase. The final optimal architecture is obtained by optimizing the multi-objective reward function to balance the trade-off between accuracy and computational efficiency. Compared with the state-of-the-art unmixing algorithms, the proposed RLNAS method can yield better unmixing results on synthetic and real hyperspectral datasets, which verifies its effectiveness and superiority. In addition, the proposed method offers promising potential of the NAS for HU. Zhu Han 0002, Danfeng Hong, Lianru Gao, Swalpa Kumar Roy, Bing Zhang 0001, Jocelyn Chanussot |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Separable Attention Network in Single- and Mixed-Precision Floating Point for Land-Cover Classification of Remote Sensing ImagesabstractLand-cover information is of paramount importance in a wide range of environmental and socioeconomic applications. Deep learning (DL) provides a large variety of potential models for extracting useful information from raw images. However, remote sensing image (RSI) classification remains a challenging goal due to the intrinsic features of the data, such as the high sample variability and lack of labeled data. This provides a challenge to the reliability of deep classifiers. In particular, convolution-based models are greatly affected by overfitting and vanishing gradient problems. To overcome these drawbacks, this letter presents a new attention-based architecture, including attention modular blocks. These blocks divide their input feature maps into several groups and split them along the channel dimension and then combine them to create an attention mask encoding global contextual information. The mask is applied to obtain a refined feature representation, strengthening those features that affect most significantly the classification and attenuating the rest. Our new method reduces significantly the number of trainable parameters. Our results, obtained using several widely used RSIs, demonstrate that the new method exhibits higher classification performance when compared to several state-of-the-art methods. Mercedes Eugenia Paoletti, Juan Mario Haut, Tayeb Alipourfard, Swalpa Kumar Roy, Eligius M. T. Hendrix, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Lightweight Heterogeneous Kernel Convolution for Hyperspectral Image Classification With Noisy LabelsabstractConvolutional neural networks (CNNs) have exhibited commendable performance in the hyperspectral images (HSIs) classification task with manually annotated limited available training data for supervision. The accurate classification of pixel-wise land covers using traditional CNNs is often hampered by the presence of wrong (noisy) labels in the training data and can easily be overfitted to the label noises. However, training on noisy labeled data inevitably suffers from performance degradation since CNNs tend to overfit the label noises. To overcome this problem, we propose a lightweight heterogeneous kernel convolution (HetConv3D) for HSI classification with noisy labels, whereHetConv3Duses two different types of convolutional kernels, i.e., spectral and spatial domains, and fuses them to produce the final feature maps that are less prompted to the noises and also reduces the computation time. The experiments are conducted using three well-known HSI datasets, i.e., Kennedy Space Center (KSC), Salinas Scene (SA), and University of Pavia (UP), and results are compared with traditional supervised classification methods, including support vector machine (SVM), random forest (RF), CNN3D, ContextNet, MS3DNet, lightweight dual-channel residual network (DCRN), and HetConv3DNet. The superior performance exhibited by the proposed modelHetConv3D-HSIconfirms the importance of learning a fusion of spatial and spectral kernel features. The source code will be made available publicly athttps://github.com/purbayankar/HetConv3DNet. Swalpa Kumar Roy, Danfeng Hong, Purbayan Kar, Xin Wu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Hyperspectral Unmixing Using Transformer NetworkabstractTransformers have intrigued the vision research community with their state-of-the-art performance in natural language processing. With their superior performance, transformers have found their way in the field of hyperspectral image classification and achieved promising results. In this article, we harness the power of transformers to conquer the task of hyperspectral unmixing and propose a novel deep neural network-based unmixing model with transformers. A transformer network captures nonlocal feature dependencies by interactions between image patches, which are not employed in CNN models, and hereby has the ability to enhance the quality of the endmember spectra and the abundance maps. The proposed model is a combination of a convolutional autoencoder and a transformer. The hyperspectral data is encoded by the convolutional encoder. The transformer captures long-range dependencies between the representations derived from the encoder. The data are reconstructed using a convolutional decoder. We applied the proposed unmixing model to three widely used unmixing datasets, i.e., Samson, Apex, and Washington DC mall and compared it with the state-of-the-art in terms of root mean squared error and spectral angle distance. The source code for the proposed model will be made publicly available at https://github.com/preetam22n/DeepTrans-HSU. Preetam Ghosh, Swalpa Kumar Roy, Bikram Koirala, Behnood Rasti, Paul Scheunders |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Hyperspectral and LiDAR Data Classification Using Joint CNNs and Morphological Feature LearningabstractConvolutional Neural Networks (CNNs) have been extensively utilized for Hyperspectral (HSI) as well as Light Detection and Ranging (LiDAR) data Classification. However, CNNs have not been much explored for joint HSI and LiDAR image classification. Therefore, this article proposes a joint feature learning (HSI and LiDAR) and fusion mechanism using CNN and Spatial Morphological blocks which generates highly accurate land-cover maps. The CNN model comprises three Conv3D layers and is directly applied to the HSIs for extracting discriminative spectral-spatial feature representation. On the contrary, the spatial morphological block is able to capture the information relevant to the height or shape of the different land-cover regions from LiDAR data. The LiDAR features are extracted using morphological dilation and erosion layers which increase the robustness of the proposed model by considering elevation information as an additional feature. Finally, both the obtained features from CNNs and spatial morphological blocks are combined using an additive operation prior to the classification. Extensive experiments are shown with widely used HSIs and LiDAR datasets, i.e., University of Houston (UH), Trento, and MUUFL Gulfport scene. The reported results show that the proposed model significantly outperforms traditional methods and other state-of-the-art deep learning models. The source code for the proposed model will be made available publicly at https://github.com/AnkurDeria/HSI+LiDAR. Swalpa Kumar Roy, Ankur Deria, Danfeng Hong, Muhammad Ahmad 0002, Antonio Plaza, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Generative Adversarial Minority Oversampling for Spectral-Spatial Hyperspectral Image ClassificationabstractRecently, convolutional neural networks (CNNs) have exhibited commendable performance for hyperspectral image (HSI) classification. Generally, an important number of samples are needed for each class to properly train CNNs. However, existing HSI data sets suffer from a significant class imbalance problem, where many classes do not have enough samples to characterize the spectral information. The performance of existing CNN models is biased toward the majority classes, which possess more samples for the training. This article addresses this issue of imbalanced data in HSI classification. In particular, a new3D-HyperGAMOmodel is proposed, which uses generative adversarial minority oversampling. The proposed3D-HyperGAMOautomatically generates more samples for minority classes at training time, using the existing samples of that class. The samples are generated in the form of a 3-D hyperspectral patch. A different classifier from the generator and the discriminator is used in the3D-HyperGAMOmodel, which is trained using both original and generated samples to determine the classes of newly generated samples to which they actually belong. The generated data are combined classwise with the original training data set to learn the network parameters of the class. Finally, the trained 3-D classifier network validates the performance of the model using the test set. Four benchmark HSI data sets, namely, Indian Pines (IP), Kennedy Space Center (KSC), University of Pavia (UP), and Botswana (BW), have been considered in our experiments. The proposed model shows outstanding data generation ability during the training, which significantly improves the classification performance over the considered data sets. The source code is available publicly athttps://github.com/mhaut/3D-HyperGAMO. Swalpa Kumar Roy, Juan Mario Haut, Mercedes Eugenia Paoletti, Shiv Ram Dubey, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Revisiting Deep Hyperspectral Feature Extraction Networks via Gradient Centralized ConvolutionabstractThe hyperspectral images are composed of a variety of textures across the different bands which increase the spectral similarity and make it difficult to predict the pixel-wise labels without inducing additional complexity at the feature level. To extract robust and discriminative features from the different regions of land cover, the hyperspectral research community is still seeking such type of convolutions which can efficiently deal with fine-grained texture information during the feature extraction phase, which often overlook this aspect by vanilla convolution. To overcome the above shortcoming, this article proposes a generalized gradient centralized 3D convolution (G2C-Conv3D) operation, which is a weighted combination between the vanilla and gradient centralized 3D convolutions (GC-Conv3D) to extract both theintensity-levelsemantic information andgradient-levelinformation. This can be easily plugged into the existing HSI feature extraction networks to boost the performance of accurate prediction for land-cover types. To validate the feasibility of the proposedG2C-Conv3D, we have considered the existing CNN3D, MS3DNet, ContextNet, and SSRN feature extraction models and as well as CAE3D, VAE3D, and SAE3D autoencoder (AE) networks, respectively. All these networks are embedded withG2C-Conv3Dconvolution to implement both generalized gradient centralized feature extraction networks (G2C-FE) and generalized gradient centralized AE networks (G2C-AE) for fine-grained spectral–spatial feature learning. In addition,G2C-Conv2Dis also considered with few networks. The extensive experiments are conducted on four most widely used hyperspectral datasets i.e., IP, KSC, UH, and UP, respectively, and compared with the nine methods. The results demonstrate that the proposedG2C-Conv3Dcan effectively enhance the feature learning ability of the existing networks and both the qualitative and quantitative results show the superiority and effectiveness of the proposedG2C-Conv3D. The source codes will be publicly available athttps://github.com/danfenghong/G2C-Conv3D-HSI. Swalpa Kumar Roy, Purbayan Kar, Danfeng Hong, Xin Wu 0001, Antonio Plaza, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Multiple Incremental Kernel Convolution for Land Cover Classification of Remotely Sensed ImagesabstractLand cover classification of remotely sensed images is an extremely important and challenging task. During the last two decades, several methods have been proposed to deal with this problem. In particular, convolutional neural network (CNN)-based methods for land cover classification have enjoyed high popularity due to their strong feature extraction and characterization abilities. However, most CNNs-based methods use relatively small kernels (usually, 3 x 3 pixels in size). Increasing the size of the kernel introduces a lot of parameters and renders considerable computational overloads. To address this issue and allow for the processing of large image datasets, the pyramidal convolution (PyConv) network has been adopted. PyConv network contains several levels of kernels with varying scales and depths, and shows significant improvements in the task of visual recognition. In this paper, we evaluate the performance of the PyConv network on the UCMERCED dataset. Our experimental results reveal that the considered approach exhibits good performance and high efficiency in the task of land cover classification. Xuanwen Tao, Lirong Han, Mercedes Eugenia Paoletti, Swalpa Kumar Roy, Javier Plaza, Juan Mario Haut, Antonio Plaza |
IGARSS | 4 |
| 2021 | DARecNet-BS: Unsupervised Dual-Attention Reconstruction Network for Hyperspectral Band SelectionabstractDue to the existence of noise and spectral redundancies in hyperspectral images (HSIs), the band selection (BS) is highly required and can be achieved through the attention mechanism. However, existing BS methods fail to consider global interaction between the spectral information and spatial information in a nonlinear fashion. In this letter, we propose an end-to-end unsupervised dual-attention reconstruction network for BS (DARecNet-BS). The proposed network employs a dual-attention mechanism, i.e., position attention module (PAM) and channel attention module (CAM), to recalibrate the feature maps and subsequently uses a 3-D reconstruction network to restore the original HSI. This way, the long-range nonlinear contextual information in spectral and spatial directions is captured, and the informative band subset can be selected. Experiments are conducted on three well-known hyperspectral data sets, i.e., Indian Pines (IP), University of Pavia (UP), and Salinas (SA), to compare existing BS approaches, and the proposedDARecNet-BScan effectively select less redundant bands with comparable or better classification accuracy. The source code will be made publicly available athttps://github.com/ucalyptus/DARecNet-BS. Swalpa Kumar Roy, Sayantan Das 0002, Tiecheng Song, Bhabatosh Chanda |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Attention-Based Adaptive Spectral-Spatial Kernel ResNet for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) provide rich spectral-spatial information with stacked hundreds of contiguous narrowbands. Due to the existence of noise and band correlation, the selection of informative spectral-spatial kernel features poses a challenge. This is often addressed by using convolutional neural networks (CNNs) with receptive field (RF) having fixed sizes. However, these solutions cannot enable neurons to effectively adjust RF sizes and cross-channel dependencies when forward and backward propagations are used to optimize the network. In this article, we present an attention-based adaptive spectral-spatial kernel improved residual network (A2S2K-ResNet) with spectral attention to capture discriminative spectral-spatial features for HSI classification in an end-to-end training fashion. In particular, the proposed network learns selective 3-D convolutional kernels to jointly extract spectral-spatial features using improved 3-D ResBlocks and adopts an efficient feature recalibration (EFR) mechanism to boost the classification performance. Extensive experiments are performed on three well-known hyperspectral data sets, i.e., IP, KSC, and UP, and the proposed A2S2K-ResNet can provide better classification results in terms of overall accuracy (OA), average accuracy (AA), and Kappa compared with the existing methods investigated. The source code will be made available at https://github.com/suvojit- 0×55aa/A2S2K-ResNet. Swalpa Kumar Roy, Suvojit Manna, Tiecheng Song, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | FuSENet: fused squeeze-and-excitation network for spectral-spatial hyperspectral image classificationabstractDeep learning‐based approaches have become very prominent in recent years due to its outstanding performance as compared to the hand‐extracted feature‐based methods. Convolutional neural network (CNN) is a type of deep learning architecture to deal with the image/video data. Residual network and squeeze and excitation network (SENet) are among recent developments in CNN for image classification. However, the performance of SENet depends on the squeeze operation done by global pooling, which sometimes may lead to poor performance. In this study, the authors propose a bilinear fusion mechanism over different types of squeeze operation such as global pooling and max pooling. The excitation operation is performed using the fused output of squeeze operation. They used to model the proposed fused SENet with the residual unit and name it as FuSENet . Here the classification experiments are performed over benchmark hyperspectral image datasets. The experimental results confirm the superiority of the proposed FuSENet method with respect to the state‐of‐the‐art methods. The source code of the complete system is made publicly available at https://github.com/swalpa/FuSENet . Swalpa Kumar Roy, Shiv Ram Dubey, Subhrasankar Chatterjee, Bidyut B. Chaudhuri |
IET Image Process. | 1 |
| 2020 | HybridSN: Exploring 3-D-2-D CNN Feature Hierarchy for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is widely used for the analysis of remotely sensed images. Hyperspectral imagery includes varying bands of images. Convolutional neural network (CNN) is one of the most frequently used deep learning-based methods for visual data processing. The use of CNN for HSI classification is also visible in recent works. These approaches are mostly based on 2-D CNN. On the other hand, the HSI classification performance is highly dependent on both spatial and spectral information. Very few methods have used the 3-D-CNN because of increased computational complexity. This letter proposes a hybrid spectral CNN (HybridSN) for HSI classification. In general, the HybridSN is a spectral-spatial 3-D-CNN followed by spatial 2-D-CNN. The 3-D-CNN facilitates the joint spatial-spectral feature representation from a stack of spectral bands. The 2-D-CNN on top of the 3-D-CNN further learns more abstract-level spatial representation. Moreover, the use of hybrid CNNs reduces the complexity of the model compared to the use of 3-D-CNN alone. To test the performance of this hybrid approach, very rigorous HSI classification experiments are performed over Indian Pines, University of Pavia, and Salinas Scene remote sensing data sets. The results are compared with the state-of-the-art hand-crafted as well as end-to-end deep learning-based methods. A very satisfactory performance is obtained using the proposed HybridSN for HSI classification. The source code can be found at https://github.com/gokriznastic/HybridSN. Swalpa Kumar Roy, Shiv Ram Dubey, Bidyut B. Chaudhuri |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Local jet pattern: a robust descriptor for texture classification
Swalpa Kumar Roy, Bhabatosh Chanda, Bidyut B. Chaudhuri, Dipak Kumar Ghosh, Shiv Ram Dubey |
Multim. Tools Appl. | 1 |
| 2020 | Local bit-plane decoded convolutional neural network features for biomedical image retrieval
Shiv Ram Dubey, Swalpa Kumar Roy, Soumendu Chakraborty, Snehasis Mukherjee, Bidyut B. Chaudhuri |
Neural Comput. Appl. | 2 |
| 2020 | Lightweight Spectral-Spatial Squeeze-and- Excitation Residual Bag-of-Features Learning for Hyperspectral ClassificationabstractOf late, convolutional neural networks (CNNs) find great attention in hyperspectral image (HSI) classification since deep CNNs exhibit commendable performance for computer vision-related areas. CNNs have already proved to be very effective feature extractors, especially for the classification of large data sets composed of 2-D images. However, due to the existence of noisy or correlated spectral bands in the spectral domain and nonuniform pixels in the spatial neighborhood, HSI classification results are often degraded and unacceptable. However, the elementary CNN models often find intrinsic representation of pattern directly when employed to explore the HSI in the spectral-spatial domain. In this article, we design an end-to-end spectral-spatial squeeze-and-excitation (SE) residual bag-of-feature (S3EResBoF) learning framework for HSI classification that takes as input raw 3-D image cubes without engineering and builds a codebook representation of transform feature by motivating the feature maps facilitating classification by suppressing useless feature maps based on patterns present in the feature maps. To boost the classification performance and learn the joint spatial-spectral features, every residual block is connected to every other 3-D convolutional layer through an identity mapping followed by an SE block, thereby facilitating the rich gradients through backpropagation. Additionally, we introduce batch normalization on every convolutional layer (ConvBN) to regularize the convergence of the network and scale invariant BoF quantization for the measure of classification. The experiments conducted using three well-known HSI data sets and compared with the state-of-the-art classification methods reveal that S3EResBoF provides competitive performance in terms of both classification and computation time. Swalpa Kumar Roy, Subhrasankar Chatterjee, Siddhartha Bhattacharyya 0001, Bidyut B. Chaudhuri, Jan Platos |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | diffGrad: An Optimization Method for Convolutional Neural NetworksabstractStochastic gradient descent (SGD) is one of the core techniques behind the success of deep neural networks. The gradient provides information on the direction in which a function has the steepest rate of change. The main problem with basic SGD is to change by equal-sized steps for all parameters, irrespective of the gradient behavior. Hence, an efficient way of deep network optimization is to have adaptive step sizes for each parameter. Recently, several attempts have been made to improve gradient descent methods such as AdaGrad, AdaDelta, RMSProp, and adaptive moment estimation (Adam). These methods rely on the square roots of exponential moving averages of squared past gradients. Thus, these methods do not take advantage of local change in gradients. In this article, a novel optimizer is proposed based on the difference between the present and the immediate past gradient (i.e., diffGrad). In the proposed diffGrad optimization technique, the step size is adjusted for each parameter in such a way that it should have a larger step size for faster gradient changing parameters and a lower step size for lower gradient changing parameters. The convergence analysis is done using the regret bound approach of the online learning framework. In this article, thorough analysis is made over three synthetic complex nonconvex functions. The image categorization experiments are also conducted over the CIFAR10 and CIFAR100 data sets to observe the performance of diffGrad with respect to the state-of-the-art optimizers such as SGDM, AdaGrad, AdaDelta, RMSProp, AMSGrad, and Adam. The residual unit (ResNet)-based convolutional neural network (CNN) architecture is used in the experiments. The experiments show that diffGrad outperforms other optimizers. Also, we show that diffGrad performs uniformly well for training CNN using different activation functions. The source code is made publicly available at https://github.com/shivram1987/diffGrad. Shiv Ram Dubey, Soumendu Chakraborty, Swalpa Kumar Roy, Snehasis Mukherjee, Satish Kumar Singh, Bidyut B. Chaudhuri |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Affine Differential Local Mean ZigZag Pattern for Texture ClassificationabstractThe texture classification is a significant problem in the area of pattern recognition. This work proposes a novel Affine Differential Local Mean ZigZag Pattern (ADLMZP) descriptor for texture classification. The proposed method has two manifolds: first Local Mean ZigZag Pattern (LMZP) map is calculated by thresholding the 3 × 3 patch neighbor intensity values with respect to path mean but in a ZigZag weighting fashion, which provides a well discriminated descriptor compared to other local binary descriptors. The local micropattern is obtained by comparing neighbor intensity values with respect to path mean which makes the descriptor robust against noise and illumination variations. Secondly, in order to make it invariant to affine changes, we incorporated an affine differential transformation along with affine gradient magnitude information of a texture image which is differed from Euclidean Gradient. The final ADLMZP descriptor is generated by concatenating the histograms of all Affine Differential Local Mean ZigZag maps. The results are computed over well known KTH-TIPS, Brodatz, and CUReT texture datasets and compared with the state-of-the-art texture classification methods. Swalpa Kumar Roy, Dipak Kumar Ghosh, Rajat Kumar Pal, Bidyut B. Chaudhuri |
TENCON | 1 |
| 2018 | Local directional ZigZag pattern: A rotation invariant descriptor for texture classification
Swalpa Kumar Roy, Bhabatosh Chanda, Bidyut B. Chaudhuri, Soumitro Banerjee, Dipak Kumar Ghosh, Shiv Ram Dubey |
Pattern Recognit. Lett. | 1 |