EDBT 2026 Demo / reviewers in the wild / expert
Biplab Banerjee
dblp:87/9571
· DBLP profile ↗
112ranked-venue papers
8as first author
76since 2021 · last 2026
0000-0001-8371-8138ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 2 first-author · 32 since 2021Artificial intelligence and machine learning · 46 · 34 since 2021Applied, interdisciplinary, general and emerging computing · 38 · 6 first-author · 23 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QMC-Net: Data-Aware Quantum Representations for Remote Sensing Image Classification
Md Aminur Hossain, Ayush V. Patel, Biplab Banerjee |
ICPR (9) | 3 |
| 2026 | X-JEPA: A Novel Joint Learning Cross-Modal Predictive Alignment Framework for Remote Sensing Image RetrievalabstractThe growing scale and heterogeneity of remote sensing (RS) imagery demand robust, scalable frameworks for content-based image retrieval across sensor modalities. We introduce X-JEPA, a novel predictive self-supervised architecture explicitly designed for cross-modal remote sensing image retrieval (RS-CMIR), and the first to extend joint embedding predictive paradigms beyond unimodal domains. Unlike prior contrastive or reconstruction-based methods, X-JEPA formulates representation learning as a latent forecasting task: predicting the semantic embedding of a target modality given context from another. To enforce modality-invariant alignment, we propose a geometry-aware Prediction Space Alignment (PSA) loss, which captures the structure of the latent space without requiring pixel-level reconstruction or modality pairing. We evaluate X-JEPA on two large-scale benchmarks—BEN-14K (Sentinel-1/Sentinel-2) and fMoW (RGB/Sentinel) across both unimodal and cross-modal retrieval tasks. X-JEPA consistently outperforms state-of-the-art self-supervised baselines, including MAE, SatMAE, CrossMAE, CSMAE-SESD, CROMA, SkySense, DeCUR, and REJEPA, achieving up to 11.0% F1-score improvement in cross-modal retrieval and 9.8% in unimodal settings. Despite its high retrieval accuracy, the model remains lightweight, requiring fewer parameters and yielding 8–10% F1-score gains on average, establishing a new state-of-the-art for scalable, sensor-agnostic RS-CMIR.1 Shabnam Choudhury, Yash Salunkhe, Vaibhav Rajan, Subhasis Chaudhuri, Biplab Banerjee |
WACV | 5 |
| 2026 | Vision-informed Semantic Text Alignment for Open-set Recognition in Remote SensingabstractExisting Open-Set Recognition (OSR) methods struggle in remote sensing (RS) as their reliance on unimodal visual features fails to resolve the severe inter-class similarity inherent in overhead imagery. To address this, we propose ViSTA-RS, a novel multimodal framework that leverages semantic context from language to disambiguate visually similar scenes. Our approach first constructs semantically-rich class prototypes by jointly encoding images with generated text captions using a Vision-Language Model. We then introduce a reconstruction-based mechanism where an image’s visual embedding is expressed as a weighted combination of these semantic prototypes. The magnitude of the reconstruction error serves as a robust novelty score, with a statistically principled threshold determined by Extreme Value Theory (EVT). This alignment of multimodal semantics with prototype reconstruction is uniquely suited for the fine-grained nature of RS data. On four challenging benchmarks, ViSTA-RS sets a new state-of-the-art, improving the AUROC for unknown detection by a significant 6.7% over leading baselines while maintaining high accuracy on known classes. Siddhant Gole, Akash Pal, Ankit Jha, Subhasis Chaudhuri, Biplab Banerjee |
WACV | 5 |
| 2026 | Multi fault detection and root cause analysis of wind turbine using Multivariate time series data based on autoencoder
Manisha Galphade, Valmik B. Nikam, Biplab Banerjee, Nilkamal More, Arvind W. Kiwelekar |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | X-CLPA: A contrastive learning and prototypical alignment-based crossmodal remote sensing image retrieval
Aparna H, Biplab Banerjee, Avik Hati |
Expert Syst. Appl. | 2 |
| 2026 | SCOPE: Segmenting common objects with prompt-conditioned encoding and SAM distillationabstractCo-segmentation aims to identify and segment common objects across a set of related images, requiring consistent semantic understanding despite contextual variations. While foundation models such as SAM have demonstrated remarkable success across diverse vision tasks, their potential for co-segmentation remains largely unexplored. In this paper, we propose a novel framework that distills feature-level knowledge from the Segment Anything Model (SAM) into a Swin backbone, enhancing semantic consistency and generalization. To better align with the co-segmentation objective, we integrate learnable prompts into the Swin backbone. The resulting hierarchical features are processed using intra-image and inter-image attention mechanisms to capture correlations within and across the images. These features are further refined by a noise suppression module, and the SAM decoder is used to produce high-quality segmentation masks. We train the network using a contrastive loss tailored for co-segmentation, alongside conventional objectives. Extensive experiments on four challenging benchmarks-PASCAL-VOC, Internet, iCoseg, and MSRC-demonstrate that our method achieves state-of-the-art performance. Shruthi Akkala, Tanisha Chawada, Saikat Dutta 0002, Subhasis Chaudhuri, Biplab Banerjee |
Pattern Recognit. Lett. | 5 |
| 2026 | UA-JEPA: Uncertainty-aware joint-embedding predictive learning for remote sensing image retrieval
Aniket Junghare, Biplab Banerjee |
Pattern Recognit. Lett. | 2 |
| 2026 | AerOSeg++: Scale-Aware and Texture-Guided Open-Vocabulary Segmentation with SAM Features for Remote Sensing ImagesabstractRemote sensing image segmentation poses significant challenges in generalizing to unseen categories during the evaluation phase. Existing open-vocabulary segmentation methods, primarily designed for natural images, struggle to cope with the spatial complexity, scale variation, and high-resolution characteristics of remote sensing imagery. Specifically, scale variations during inference can degrade performance, as the model tends to overfit to fixed-scale patterns encountered during training. This also affects the model’s ability to recognize unseen or novel class objects appearing in varying sizes or resolutions during testing. These limitations increase the need for developing open-vocabulary segmentation methods addressing the challenges of geospatial images. In this work, we introduce AerOSeg ++, an open-vocabulary segmentation method in remote sensing, focusing on scale-invariant feature learning. We first compute robust image-text correlation features using rotated input images and domain-specific prompts. These are refined via spatial and class refinement blocks, guided by SAM features to enhance spatial consistency. To upscale the refined correlation features, we propose a multi-scale decoder framework that fuses fine-grained texture features with SAM-derived features. By leveraging texture information across multiple receptive fields, AerOSeg++ effectively captures scale-consistent patterns, facilitating accurate segmentation of objects across varying spatial resolutions. Additionally, our training pipeline incorporates ScaleDrop, a computationally efficient parameter-free feature rescaling module ensuring scale-invariant feature representation learning. Our proposed model has shown significant performance gains compared to the state-of-the-art open-vocabulary methods when evaluated on three benchmark datasets for remote sensing—iSAID, DLRSD, and OpenEarthMap. These results highlight the effectiveness of our scale-invariant design and texture-guided multi-scale feature upsampling in handling the challenges of open-vocabulary segmentation in remote sensing imagery. Saikat Dutta 0002, Akhil Vasim, Seyed Hamid Rezatofighi, Biplab Banerjee |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIPabstractWe introduce Low-Shot Open-Set Domain Generalization (LSOSDG), a novel paradigm unifying low-shot learning with open-set domain generalization (ODG). While prompt-based methods using models like CLIP have advanced DG, they falter in low-data regimes (e.g., 1-shot) and lack precision in detecting open-set samples with finegrained semantics related to training classes. To address these challenges, we propose OSLoPrompt, an advanced prompt-learning framework for CLIP with two core innovations. First, to manage limited supervision across source domains and improve DG, we introduce a domainagnostic prompt-learning mechanism that integrates adaptable domain-specific cues and visually guided semantic attributes through a novel cross-attention module, besides being supported by learnable domain- and class-generic visual prompts to enhance cross-modal adaptability. Second, to improve outlier rejection during inference, we classify unfamiliar samples as “unknown” and train specialized prompts with systematically synthesized pseudo-open samples that maintain fine-grained relationships to known classes, generated through a targeted query strategy with off-the-shelf foundation models. This strategy enhances feature learning, enabling our model to detect open samples with varied granularity more effectively. Extensive evaluations across five benchmarks demonstrate that OSLO- Prompt establishes a new state-of-the-art in LSOSDG, significantly outperforming existing methods.1 Mohamad Hassan N C, Divyam Gupta, Mainak Singha, Sai Bhargav Rongali, Ankit Jha, Muhammad Haris Khan, Biplab Banerjee |
CVPR | 7 |
| 2025 | When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven ApproachabstractGeneralized Class Discovery (GCD) clusters base and novel classes in a target domain, using supervision from a source domain with only base classes. Current methods often falter with distribution shifts and typically require access to target data during training, which can sometimes be impractical. To address this issue, we introduce the novel paradigm of Domain Generalization in GCD (DG-GCD), where only source data is available for training, while the target domain—with a distinct data distribution—remains unseen until inference. To this end, our solution, DG2CD-Net, aims to construct a domain-independent, discriminative embedding space for GCD. The core innovation is an episodic training strategy that enhances cross-domain generalization by adapting a base model on tasks derived from source and synthetic domains generated by a foundation model. Each episode focuses on a cross-domain GCD task, diversifying task setups over episodes and combining openset domain adaptation with a novel margin loss and representation learning for optimizing the feature space progressively. To capture the effects of fine-tunings on the base model, we extend task arithmetic by adaptively weighting the local task vectors concerning the fine-tuned models based on their GCD performance on a validation distribution. This episodic update mechanism boosts the adaptability of the base model to unseen targets. Experiments across three datasets confirm that DG2CD-Net outperforms existing GCD methods customized for DG-GCD. Vaibhav Rathore, Shubhranil B, Saikat Dutta 0002, Sarthak Mehrotra, Zsolt Kira, Biplab Banerjee |
CVPR | 6 |
| 2025 | Hyperbolic Uncertainty-Aware Few-Shot Incremental Point Cloud Segmentationabstract3D point cloud segmentation is essential across a range of applications; however, conventional methods often struggle in evolving environments, particularly when tasked with identifying novel categories under limited supervision. Few-Shot Learning (FSL) and Class Incremental Learning (CIL) have been adapted previously to address these challenges in isolation, yet the combined paradigm of Few-Shot Class Incremental Learning (FSCIL) remains largely unexplored for point cloud segmentation. To address this gap, we introduce Hyperbolic Ideal Prototypes Optimization (HIPO), a novel framework that harnesses hyperbolic embeddings for FSCIL in 3D point clouds. HIPO employs the Poincaré Hyperbolic Sphere as its embedding space, integrating Ideal Prototypes enriched by CLIP-derived class semantics, to capture the hierarchical structure of 3D data. By enforcing orthogonality among prototypes and maximizing representational margins, HIPO constructs a resilient embedding space that mitigates forgetting and enables the seamless integration of new classes, thereby effectively countering overfitting. Extensive evaluations on S3DIS, ScanNetv2, and cross-dataset scenarios demonstrate HIPO’s strong performance, significantly surpassing existing approaches in both in-domain and cross-dataset FSCIL tasks for 3D point cloud segmentation. Tanuj Sur, Samrat Mukherjee, Kaizer Rahaman, Subhasis Chaudhuri, Muhammad Haris Khan, Biplab Banerjee |
CVPR | 6 |
| 2025 | ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social MediaabstractAakash Kumar Agarwal, Saprativa Bhattacharjee, Mauli Rastogi, Jemima S. Jacob, Biplab Banerjee, Rashmi Gupta, Pushpak Bhattacharyya. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Aakash Kumar Agarwal, Saprativa Bhattacharjee, Mauli Rastogi, Jemima Jacob, Biplab Banerjee, Pushpak Bhattacharyya |
EMNLP | 5 |
| 2025 | UIDAPLE: Unsupervised Incremental Domain Adaptation through Adaptive Prompt LearningabstractContinual learning poses significant challenges for deep neural networks, notably catastrophic forgetting, particularly when faced with shifting data distributions that compromise previously acquired knowledge. This paper tackles these issues within the Unsupervised Incremental Domain Adaptation (UIDA) framework, where the initial source domain is labeled, but subsequent domains are not. Existing methods often struggle with limited cross-domain generalization and adaptation capabilities. As a remedy, we introduce UIDAPLE, a novel approach that utilizes a unified prompt across all domains, leveraging the foundation model CLIP to obviate the need for isolated domain treatments. Specifically, UIDAPLE implements supervised prompt learning in the labeled source domain and extends this learning to unlabeled domains through confidence-based adaptation. We also present an efficient parameter alignment strategy that maintains semantic coherence across domains, effectively balancing stability and plasticity to combat catastrophic forgetting. Extensive evaluations on two benchmark datasets reveal that UIDAPLE markedly surpasses other UIDA techniques in performance. Samrat Mukherjee, Tanuj Sur, Saurish Seksaria, Subhasis Chaudhuri, Gemma Roig, Biplab Banerjee |
ICASSP | 6 |
| 2025 | Spatially-Aware Cross-Modal Contrastive Learning for Low-Shot HSI ClassificationabstractClassifying hyperspectral images (HSI) with limited supervision is challenging due to their high dimensionality and complex spectral features, which frequently result in overfitting, especially under extremely low supervision. Existing self-supervised methods for HSI data focus predominantly on spectral attributes, neglecting the spatial details crucial for effective HSI classification. To address this, we introduce the Cross-Modal Spatial Contrastive (CM-SCON) framework, a novel self-supervised approach that employs co-registered, unlabeled HSI and LiDAR data. CM-SCON leverages LiDAR’s spatial context to enhance the spectral discriminability of the HSI encoder. Central to our method is a pair of cross-modal pretext tasks that merges cross-modal patch reconstruction with a contrastive learning objective, significantly boosting the HSI encoder’s effectiveness, which adeptly handles downstream land-cover classification tasks, even with minimal labeled data. Extensive evaluations on the Houston-13, 18, and Trento benchmark datasets show that CM-SCON outperforms existing baselines for within-dataset and cross-dataset evaluation scenarios. Akhil Vasim, Pankhi Kashyap, Shabnam Choudhury, Biplab Banerjee |
ICASSP | 4 |
| 2025 | DuET: Dual Incremental Object Detection via Exemplar-Free Task ArithmeticabstractReal-world object detection systems, such as those in autonomous driving and surveillance, must continuously learn new object categories and simultaneously adapt to changing environmental conditions. Existing approaches, Class Incremental Object Detection (CIOD) and Domain Incremental Object Detection (DIOD) only address one aspect of this challenge. CIOD struggles in unseen domains, while DIOD suffers from catastrophic forgetting when learning new classes, limiting their real-world applicability. To overcome these limitations, we introduce Dual Incremental Object Detection (DuIOD), a more practical setting that simultaneously handles class and domain shifts in an exemplar-free manner. We propose DuET, a Task Arithmetic-based model merging framework that enables stable incremental learning while mitigating sign conflicts through a novel Directional Consistency Loss. Unlike prior methods, DuET is detector-agnostic, allowing models like YOLO11 and RT-DETR to function as real-time incremental object detectors. To comprehensively evaluate both retention and adaptation, we introduce the Retention-Adaptability Index (RAI), which combines the Average Retention Index (Avg RI) for catastrophic forgetting and the Average Generalization Index for domain adaptability into a common ground. Extensive experiments on the Pascal Series and Diverse Weather Series demonstrate DuET's effectiveness, achieving a +13.12% RAI improvement while preserving 89.3% Avg RI on the Pascal Series (4 tasks), as well as a +11.39% RAI improvement with 88.57% Avg RI on the Diverse Weather Series (3 tasks), outperforming existing methods. Munish Monga, Vishal M. Chudasama, Pankaj Wasnik, Biplab Banerjee |
ICCV | 4 |
| 2025 | FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language ModelsabstractIn federated learning, textual prompt tuning adapts Vision-Language Models (e.g., CLIP) by tuning lightweight input tokens (or prompts) on local client data, while keeping network weights frozen. After training, only the prompts are shared by the clients with the central server for aggregation. However, textual prompt tuning suffers from overfitting to known concepts, limiting its generalizability to unseen concepts. To address this limitation, we propose Multimodal Visual Prompt Tuning (FedMVP) that conditions the prompts on multimodal contextual information - derived from the input image and textual attribute features of a class. At the core of FedMVP is a PromptFormer module that synergistically aligns textual and visual features through a cross-attention mechanism. The dynamically generated multimodal visual prompts are then input to the frozen vision encoder of CLIP, and trained with a combination of CLIP similarity loss and a consistency loss. Extensive evaluation on 20 datasets, spanning three generalization settings, demonstrates that FedMVP not only preserves performance on in-distribution classes and domains, but also displays higher generalizability to unseen classes and domains, surpassing state-of-the-art methods by a notable margin of +1.57% - 2.26%. Code is available at https://github.com/mainaksingha01/FedMVP. Mainak Singha, Subhankar Roy, Sarthak Mehrotra, Ankit Jha, Moloud Abdar, Biplab Banerjee, Elisa Ricci 0001 |
ICCV | 6 |
| 2025 | HIDISC: A Hyperbolic Framework for Domain Generalization with Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) aims to classify test-time samples into either seen categories—available during training—or novel ones, without relying on label supervision. Most existing GCD methods assume simultaneous access to labeled and unlabeled data during training and arising from the same domain, limiting applicability in open-world scenarios involving distribution shifts. Domain Generalization with GCD (DG-GCD) lifts this constraint by requiring models to generalize to unseen domains containing novel categories, without accessing target-domain data during training.
The only prior DG-GCD method, DG$^2$CD-Net~\cite{dg2net}, relies on episodic training with multiple synthetic domains and task vector aggregation, incurring high computational cost and error accumulation. We propose \textsc{HiDISC}, a hyperbolic representation learning framework that achieves domain and category-level generalization without episodic simulation. To expose the model to minimal but diverse domain variations, we augment the source domain using GPT-guided diffusion, avoiding overfitting while maintaining efficiency.
To structure the representation space, we introduce \emph{Tangent CutMix}, a curvature-aware interpolation that synthesizes pseudo-novel samples in tangent space, preserving manifold consistency. A unified loss—combining penalized Busemann alignment, hybrid hyperbolic contrastive regularization, and adaptive outlier repulsion—facilitates compact, semantically structured embeddings. A learnable curvature parameter further adapts the geometry to dataset complexity.
\textsc{HiDISC} achieves state-of-the-art results on PACS~\cite{pacs}, Office-Home~\cite{officehome}, and DomainNet~\cite{domainnet}, consistently outperforming the existing Euclidean and hyperbolic (DG)-GCD baselines. Vaibhav Rathore, Divyam Gupta, Biplab Banerjee |
NeurIPS | 3 |
| 2025 | Foundation Models and Adaptive Feature Selection: A Synergistic Approach to Video Question AnsweringabstractThis paper tackles the intricate challenge of video question-answering (VideoQA). Despite notable progress, current methods fall short of effectively integrating questions with video frames and semantic object-level abstractions to create question-aware video representations. We introduce Local - Global Question Aware Video Embedding (LGQAVE), which incorporates three major innovations to integrate multi-modal knowledge better and emphasize semantic visual concepts relevant to specific questions. LGQAVE moves beyond traditional ad-hoc frame sampling by utilizing a cross-attention mechanism that precisely identifies the most relevant frames concerning the questions. It captures the dynamics of objects within these frames using distinct graphs, grounding them in question semantics with the miniGPT model. These graphs are processed by a question-aware dynamic graph transformer (Q-DGT), which refines the outputs to develop nuanced global and local video representations. An additional cross-attention module integrates these local and global embeddings to generate the final video embeddings, which a language model uses to generate answers. Extensive evaluations across multiple benchmarks demonstrate that LGQAVE significantly outperforms existing models in delivering accurate multi-choice and open-ended answers. Sai Bhargav Rongali, Mohamad Hassan N C, Ankit Jha, Neha Bhargava, Saurabh Prasad, Biplab Banerjee |
WACV | 6 |
| 2025 | Towards molecular structure discovery from cryo-ET density volumes via modelling auxiliary semantic prototypesabstractCryo-electron tomography (cryo-ET) is confronted with the intricate task of unveiling novel structures. General class discovery (GCD) seeks to identify new classes by learning a model that can pseudo-label unannotated (novel) instances solely using supervision from labeled (base) classes. While 2D GCD for image data has made strides, its 3D counterpart remains unexplored. Traditional methods encounter challenges due to model bias and limited feature transferability when clustering unlabeled 2D images into known and potentially novel categories based on labeled data. To address this limitation and extend GCD to 3D structures, we propose an innovative approach that harnesses a pretrained 2D transformer, enriched by an effective weight inflation strategy tailored for 3D adaptation, followed by a decoupled prototypical network. Incorporating the power of pretrained weight-inflated Transformers, we further integrate CLIP, a vision-language model to incorporate textual information. Our method synergizes a graph convolutional network with CLIP's frozen text encoder, preserving class neighborhood structure. In order to effectively represent unlabeled samples, we devise semantic distance distributions, by formulating a bipartite matching problem for category prototypes using a decoupled prototypical network. Empirical results unequivocally highlight our method's potential in unveiling hitherto unknown structures in cryo-ET. By bridging the gap between 2D GCD and the distinctive challenges of 3D cryo-ET data, our approach paves novel avenues for exploration and discovery in this domain. Ashwin R. Nair, Xingjian Li 0002, Bhupendra Solanki, Souradeep Mukhopadhyay, Ankit Jha, Mostofa Rafid Uddin, Mainak Singha, Biplab Banerjee, Min Xu 0009 |
Briefings Bioinform. | 8 |
| 2025 | RS3Lip: Consistency for remote sensing image classification on part embeddings using self-supervised learning and CLIP
Ankit Jha, Mainak Singha, Avigyan Bhattacharya, Biplab Banerjee |
Comput. Vis. Image Underst. | 4 |
| 2024 | COSMo: CLIP Talks on Open-Set Multi-Target Domain Adaptation
Munish Monga, Sachin Kumar Giroh, Ankit Jha, Mainak Singha, Biplab Banerjee, Jocelyn Chanussot |
BMVC | 5 |
| 2024 | Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain GeneralizationabstractWe delve into Open Domain Generalization (ODG), marked by domain and category shifts between training's labeled source and testing's unlabeled target domains. Existing solutions to ODG face limitations due to constrained generalizations of traditional CNN backbones and errors in detecting target open samples in the absence of prior knowledge. Addressing these pitfalls, we introduce ODG-CLIP, harnessing the semantic prowess of the vision-language model, CLIP. Our framework brings forth three primary innovations: Firstly, distinct from prevailing paradigms, we conceptualize ODG as a multi-class classification challenge encompassing both known and novel categories. Central to our approach is modeling a unique prompt tailored for detecting unknown class samples, and to train this, we employ a readily accessible stable diffusion model, elegantly generating proxy images for the open class. Secondly, aiming for domain-tailored classification (prompt) weights while ensuring a balance of precision and simplicity, we devise a novel visual stylecentric prompt learning mechanism. Finally, we infuse images with class-discriminative knowledge derived from the prompt space to augment the fidelity of CLIP's visual embeddings. We introduce a novel objective to safeguard the continuity of this infused semantic intel across domains, especially for the shared classes. Through rigorous testing on diverse datasets, covering closed and open-set DG contexts, ODG-CLIP demonstrates clear supremacy, consistently outpacing peers with performance boosts between 8%-16%. Code will be available at https://github.com/mainaksingha01/ODG-CLIP. Mainak Singha, Ankit Jha, Shirsha Bose, Ashwin R. Nair, Moloud Abdar, Biplab Banerjee |
CVPR | 6 |
| 2024 | Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning
Mainak Singha, Ankit Jha, Divyam Gupta, Pranav Singla, Biplab Banerjee |
ECCV (24) | 5 |
| 2024 | Enhancing the Domain Robustness of Self-Supervised pre-Training with Synthetic ImagesabstractWe present a novel method for improving the adaptability of self-supervised (SSL) pre-trained models across different domains. Our approach uses synthetic images that are generated using an auxiliary diffusion model, namely InstructPix2Pix. More specifically, starting from a real image, we prompt the diffusion model to generate synthetic versions of that image in the style of the target domains. This allows us to generate a diverse set of multi-domain images that share the same semantics as real images. Integrating these synthetic images into the training dataset enhances the model’s capacity to generalize to other domains. We pre-trained different SSL methods on Imagenet-100 with and without the synthetic images and evaluated their performance on three multi-domain datasets, DomainNet, PACS, and Office-Home. Our results show significant improvements in all datasets and methods, encouraging new research in the direction of leveraging synthetic data to improve the robustness of pre-trained models. Code is available at https://github.com/has97/Diffusion_pre-training. Mohamad Hassan N C, Avigyan Bhattacharya, Victor G. T. da Costa, Biplab Banerjee, Elisa Ricci 0001 |
ICASSP | 4 |
| 2024 | SPDG-Net: Semantics Preserving Domain Augmentation through Style Interpolation for Multi-Source Domain GeneralizationabstractThis paper focuses on domain generalization (DG), addressing the challenge of robust classifier learning from multiple source domains for generalizing to unseen ones. DG suffers from limited source domain diversity, which may hinder model generalization. Recent studies explore domain-augmentation strategies but struggle to maintain semantics while altering image styles and generate only a few pseudo domains. To tackle this, we introduce Semantics Preserving DG Network (SPDG-Net). SPDG-Net is a triplet-conditioned U-Net-based GAN that synthesizes various pseudo domains, preserving image semantics through a cycle-consistency constraint. Moreover, our style interpolation-based domain generation produces fine-grained synthetic domains, unlike existing models. We also propose recognizing image styles alongside object classes to reduce model bias. Our experiments across benchmark datasets consistently outperform recent literature in DG. Advait Kumar, Shirsha Bose, Mohamad Hassan N C, Biplab Banerjee |
ICASSP | 4 |
| 2024 | CrossVG: Visual Grounding in Remote Sensing with Modality-Guided InteractionsabstractVisual grounding aims to use a natural language expression to find specific objects in an image, whether in a bounding box or a segmentation mask. The vision research community has extensively investigated the objective. Nevertheless, the existing benchmark datasets and methodologies predominantly emphasize natural images rather than remote sensing images. Remote sensing images differ from realistic images in that they encompass expansive scenes and provide geographical spatial information about ground objects. Current approaches address this issue by extending the fundamental object detection framework. Yet, the effectiveness of visual feature models based on these predetermined locations may be limited since they only partially leverage the visual context and attribute information offered by the text query. This restricts the platitudinous interaction in the visual-linguistic setting. To circumvent this constraint, we suggest a novel architecture called CrossVG based on visual and language-guided cross-modality interactions to build multi-modal correspondence. Through experiments, we demonstrate that a simple stack of transformer encoder layers can substitute complex fusion modules with better-performing alternatives. We validate the efficacy of our suggested model and exhibit SOTA performance using the benchmark dataset RSVGD. Shabnam Choudhury, Pratham Kurkure, Priyanka Talwar, Biplab Banerjee |
IGARSS | 4 |
| 2024 | Enhancing Crop Type Classification from Multi-Frequency Dual-Pol SAR Data by Probabilistic Fusion of Gaussian ProcessesabstractThis paper proposes a novel multivariate Gaussian Process Regression (GPR) approach for multi-class crop classification. We have trained and validated the proposed model utilising backscatter information from E-SAR C- and L-band dual-polarimetric data acquired during the AGRISAR 2006 campaign. Further, we use the Product of Experts (PoE) fusion strategy to combine decisions from the proposed Gaussian Process (GP) models trained and validated independently over C- and L-band data to analyze the changes in the classification performance. The synergistic C- and L- band information show an improved classification accuracy during various phenological stages of major crop types by (a) 4 to 37 % for VV-VH backscatter intensity channels and (b) 1 to 39 % for HH-HV backscatter intensity channels. Swarnendu Sekhar Ghosh, Avik Bhattacharya, Dipankar Mandal, Biplab Banerjee, Narayanarao Bhogapurapu, Paul Siqueira |
IGARSS | 5 |
| 2024 | Shape-prior Free Space-time Neural Radiance Field for 4D Semantic Reconstruction of Dynamic Scene from Sparse-View RGB VideosabstractMany applications in Augmented/Virtual Reality or robotics require precise geometry modeling of individual elements in a dynamic scene under a sparse-view camera setup, without any prior information about their semantic labels or shapes. In our research, we introduce a 3D shape prior-free Neural Radiance Field-based technique for detailed geometry reconstruction under human-object interactions, offering an explicit surface reconstruction with semantic labels for reconstructed geometry. Our approach harnesses the capabilities of an Invertible Neural Network to learn a deformation function that effectively connects local (current-frame input) and canonical spaces for each of the components under motion. The deformation process is guided by temporal constraints from multi-frame, facilitating the precise reconstruction of the complex interactions between humans and objects. Our experimental evaluations highlight the effectiveness of our framework, demonstrating its ability to accurately represent both the comprehensive object-compositional scene and individual components over state-of-the-art methods, under complex interactions between the scene entities. This research is deemed to mark a significant stride in semantic 3D geometry modeling within dynamic interactive environments, relying solely on sparse multi-view RGB data. Sandika Biswas, Biplab Banerjee, Seyed Hamid Rezatofighi |
IROS | 2 |
| 2024 | TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic SceneabstractDespite advancements in Neural Implicit models for 3D surface reconstruction, handling dynamic environments with interactions between arbitrary rigid, non-rigid, or deformable entities remains challenging. The generic reconstruction methods adaptable to such dynamic scenes often require additional inputs like depth or optical flow or rely on pre-trained image features for reasonable outcomes. These methods typically use latent codes to capture frame-by-frame deformations. Another set of dynamic scene reconstruction methods, are entity-specific, mostly focusing on humans, and relies on template models. In contrast, some template-free methods bypass these requirements and adopt traditional LBS (Linear Blend Skinning) weights for a detailed representation of deformable object motions,
although they involve complex optimizations leading to lengthy training times. To this end, as a remedy, this paper introduces TFS-NeRF, a template-free 3D semantic NeRF for dynamic scenes captured from sparse or single-view RGB videos, featuring interactions among two entities and more time-efficient than other LBS-based approaches. Our framework uses an Invertible Neural Network (INN) for LBS prediction, simplifying the training process. By disentangling the motions of interacting entities and optimizing per-entity skinning weights, our method efficiently generates accurate, semantically separable geometries. Extensive experiments demonstrate that our approach produces high-quality reconstructions of both deformable and non-deformable objects in complex interactions, with improved
training efficiency compared to existing methods. The code and models will be available on our github page. Sandika Biswas, Qianyi Wu, Biplab Banerjee, Seyed Hamid Rezatofighi |
NeurIPS | 3 |
| 2024 | Learning Class and Domain Augmentations for Single-Source Open-Domain GeneralizationabstractSingle-source open-domain generalization (SS-ODG) addresses the challenge of labeled source domains with supervision during training and unlabeled novel target domains during testing. The target domain includes both known classes from the source domain and samples from previously unseen classes. Existing techniques for SS-ODG primarily focus on calibrating source-domain classifiers to identify open samples in the target domain. However, these methods struggle with visually fine-grained open-closed data, often misclassifying open samples as closed-set classes. Moreover, relying solely on a single source domain restricts the model’s ability to generalize. To overcome these limitations, we propose a novel framework called SODG-Net that simultaneously synthesizes novel domains and generates pseudo-open samples using a learning-based objective, in contrast to the ad-hoc mixing strategies commonly found in the literature. Our approach enhances generalization by diversifying the styles of known class samples using a novel metric criterion and generates diverse pseudo-open samples to train a unified and confident multiclass classifier capable of handling both open and closed-set data. Extensive experimental evaluations conducted on multiple benchmarks consistently demonstrate the superior performance of SODG-Net compared to the literature. Prathmesh Bele, Valay Bundele, Avigyan Bhattacharya, Ankit Jha, Gemma Roig, Biplab Banerjee |
WACV | 6 |
| 2024 | StyLIP: Multi-Scale Style-Conditioned Prompt Learning for CLIP-based Domain GeneralizationabstractLarge-scale foundation models, such as CLIP, have demonstrated impressive zero-shot generalization performance on downstream tasks, leveraging well-designed language prompts. However, these prompt learning techniques often struggle with domain shift, limiting their generalization capabilities. In our study, we tackle this issue by proposing StyLIP, a novel approach for Domain Generalization (DG) that enhances CLIP’s classification performance across domains. Our method focuses on a domain-agnostic prompt learning strategy, aiming to disentangle the visual style and content information embedded in CLIP’s pre-trained vision encoder, enabling effortless adaptation to novel domains during inference. To achieve this, we introduce a set of style projectors that directly learn the domain-specific prompt tokens from the extracted multi-scale style features. These generated prompt embeddings are subsequently combined with the multi-scale visual content features learned by a content projector. The projectors are trained in a contrastive manner, utilizing CLIP’s fixed vision and text backbones. Through extensive experiments conducted in five different DG settings on multiple benchmark datasets, we consistently demonstrate that StyLIP outperforms the current state-of-the-art (SOTA) methods. Shirsha Bose, Ankit Jha, Enrico Fini, Mainak Singha, Elisa Ricci 0001, Biplab Banerjee |
WACV | 6 |
| 2024 | Domain Adaptive 3D Shape Retrieval from Monocular ImagesabstractIn this work, we address the novel and challenging problem of domain adaptive 3D shape retrieval from single 2D images (DA-IBSR). While the existing image-based 3D shape retrieval (IBSR) problem focuses on modality alignment for retrieving a matchable 3D shape from a shape repository given a 2D image query, it does not consider any distribution shift between the training and testing image-shape pairs, making the performance of off-the-shelves IBSR methods subpar. In contrast, the proposed DA-IBSR addresses the non-trivial problem of modality shift as well distribution shift across training and test sets. To address these issues, we propose an end-to-end trainable model called DAIS-NET. Our objective is to align the images and shapes separately from both domains while simultaneously learn a shared embedding space for the 2D and 3D modalities. The former problem is addressed by separately employing maximum mean discrepancy loss across the 2D images and 3D shapes of the two domains. To address the modality alignment, we incorporate the notion of negative sample mining and employ triplet loss to bridge the gap between positive 2D-3D pairs (of same class) and increase the separation between negative 2D-3D pairs (of different class). Additionally, we employ an entropy minimization strategy to align the unlabeled target domain data in the semantic space. To evaluate our proposed approach, we define the experimental setting of DA-IBSR on the following benchmarks: SHREC’14 ↔ Pix3D and ShapeNet ↔ SHREC’14. Considering the novelty of the problem statement, we have demonstrated that the issue of domain gap is prevalent by comparing our method with the existing literature. Additionally, through extensive evaluations, we demonstrate the capability of DAIS-NET to successfully mitigate this domain gap in image based 3D shape retrieval. Harsh Pal, Ritwik Khandelwal, Shivam Pande, Biplab Banerjee, Srikrishna Karanam |
WACV | 4 |
| 2024 | Digital image noise removal towards soybean and cotton plant disease using image processing filters
Vaishali G. Bhujade, Vijay Sambhe, Biplab Banerjee |
Expert Syst. Appl. | 3 |
| 2024 | DARK: Few-Shot Remote-Sensing Colorization Using Label-Conditioned Color InjectionabstractSatellite image colorization is a broad challenging problem in the domain of remote sensing (RS) having huge potential applications. The problem becomes even more complicated under the few-shot setting yet it has barely been studied to date. In this paper, we propose a colorization framework for the RS scene for synthesizing optical images from their panchromatic (PAN) counterparts using color injection and attention fusion mechanism. Our proposed model ensures that the synthesized optical images are coherent with the structural variability of panchromatic images while constraining the realistic appearance in the optical domain from a few training image pairs. To accomplish the same, we introduce a novel Dual Attention fusion of Receptive Kernels (DARK) which considers the spatial nuances along with color injection conditioned on prior label allocation. DARK is a multi spectral-spatial feature generator that selectively accentuates important cross-spatial features based on attention fusion. We also employ a prior distribution constraint on color embedding generation for introducing vibrant yet diverse variance in a color generation. Our approach achieves state-of-the-art results on the publicly available EuroSAT and PatternNet datasets while demonstrating significant speedups. We showcase our results quantitatively by comparing the PSNR, mean squared error(MSE), and cosine similarity of generated images and qualitatively via visual perception. Rupak Bose, Anshul Shrivastava, Biplab Banerjee, Subhasis Chaudhuri |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Boosting cross-modal retrieval in remote sensing via a novel unified attention networkabstractWith the rapid advent and abundance of remote sensing data in different modalities, cross-modal retrieval tasks have gained importance in the research community. Cross-modal retrieval belongs to the research paradigm in which the query is of one modality and the retrieved output is of the other modality. In this paper, the remote sensing (RS) data modalities considered are the earth observation optical data (aerial photos) and the corresponding hand-drawn sketches. The main challenge of the cross-modal retrieval research objective for optical remote sensing images and the corresponding sketches is the distribution gap between the shared embedding space of the modalities. Prior attempts to resolve this issue have not yielded satisfactory outcomes regarding accurately retrieving cross-modal sketch-image RS data. The state-of-the-art architectures used conventional convolutional architectures, which focused on local pixel-wise information about the modalities to be retrieved. This limits the interaction between the sketch texture and the corresponding image, making these models susceptible to overfitting datasets with particular scenarios. To circumvent this limitation, we suggest establishing multi-modal correspondence using a novel architecture of the combined self and cross-attention algorithms, SPCA-Net to minimize the modality gap by employing attention mechanisms for the query and other modalities. Efficient cross-modal retrieval is achieved through the suggested attention architecture, which empirically emphasizes the global information of the relevant query modality and bridges the domain gap through a unique pairwise cross-attention network. In addition to the novel architecture, this paper introduces a unique loss function, label-specific supervised contrastive loss, tailored to the intricacies of the task and to enhance the discriminative power of the learned embeddings. Extensive evaluations are conducted on two sketch-image remote sensing datasets, Earth-on-Canvas and RSketch. Under the same experimental conditions, the performance metrics of our proposed model beat the state-of-the-art architectures by significant margins of 16.7%, 18.9%, 33.7%, and 40.9% correspondingly. Shabnam Choudhury, Devansh Saini, Biplab Banerjee |
Neural Networks | 3 |
| 2023 | GOPro: Generate and Optimize Prompts in CLIP using Self-Supervised Learning
Mainak Singha, Ankit Jha, Biplab Banerjee |
BMVC | 3 |
| 2023 | Domain Adaptive Few-Shot Open-Set LearningabstractFew-shot learning has made impressive strides in addressing the crucial challenges of recognizing unknown samples from novel classes in target query sets and managing visual shifts between domains. However, existing techniques fall short when it comes to identifying target outliers under domain shifts by learning to reject pseudo-outliers from the source domain, resulting in an incomplete solution to both problems. To address these challenges comprehensively, we propose a novel approach called Domain Adaptive Few-Shot Open Set Recognition (DA-FSOS) and introduce a meta-learning-based architecture named DAFOS-Net. During training, our model learns a shared and discriminative embedding space while creating a pseudo-open-space decision boundary, given a fully-supervised source domain and a label-disjoint few-shot target domain. To enhance data density, we use a pair of conditional adversarial networks with tunable noise variances to augment both domains’ closed and pseudo-open spaces. Furthermore, we propose a domain-specific batch-normalized class prototypes alignment strategy to align both domains globally while ensuring class-discriminativeness through novel metric objectives. Our training approach ensures that DAFOS-Net can generalize well to new scenarios in the target domain. We present three benchmarks for DA-FSOS based on the Office-Home, mini-ImageNet/CUB, and DomainNet datasets and demonstrate the efficacy of DAFOS-Net through extensive experimentation. Debabrata Pal, Deeptej More, Sai Bhargav, Dipesh Tamboli, Vaneet Aggarwal, Biplab Banerjee |
ICCV | 6 |
| 2023 | Spatio-Temporal Change Detection Model For Deforestation Analysis Using Advanced Deep Learning ApproachesabstractSatellite imagery is widely available in this era of internet. Satellite images are available in multispectral, hyperspectral images, Lidar data form, and other types of imaging dataset are also available. Huge information exists in these imageries. In this study, multispectral optical satellite images are used. These images have spectral, spatial, and temporal properties. This data has velocity, volume and variety. Though enough research is available to process this type of data, as this data is a Bigdata, the processing and analysis of these images is compute intensive, there are research challenges in solving the problems associated with this domain. This research study explores the classification accuracy of change detection and compute time efficiency for processing huge satellite imaging datasets.As a use case for this research, the spatio-temporal change detection for deforestation detection employing top tree canopy has been chosen. We have considered one of India's densely populated metropolitan areas, Mumbai, which has several industrial areas, and rapidly expanding population. As a result of the growing urbanization, the tree canopy of the area is gradually decreasing. We have chosen this deforestation analysis, and addressed accuracy and compute performance for the satellite imagery dataset as a research problem.This research has following major two contributions:a)Dataset creation for Mumbai region from Sentinel 2 dataset of 2500 images.b)The Spatio-Temporal Change Detection (STCD) model, explores the hybrid approach using transfer learning and self-attention models for better accuracy and improved compute performance for deforestation change detection with tree canopy. Nilkamal More, Valmik B. Nikam, Biplab Banerjee, T. P. Singh, Abhay Bambole |
IGARSS | 3 |
| 2023 | Semi-Supervised Learning for Hyperspectral Images by Non Parametrically Predicting View AssignmentCRediTabstractHyperspectral image (HSI) classification is gaining a lot of momentum in present time because of high inherent spectral information within the images. However, these images suffer from the problem of curse of dimensionality and usually require a large number samples for tasks such as classification, especially in supervised setting. Recently, to effectively train the deep learning models with minimal labelled samples, the unlabeled samples are also being leveraged in self-supervised and semi-supervised setting. In this work, we leverage the idea of semi-supervised learning to assist the discriminative self-supervised pretraining of the models. The proposed method takes different augmented views of the unlabeled samples as input and assigns them the same pseudo-label corresponding to the labelled sample from the downstream task. We train our model on two HSI datasets, anemly Houston dataset (from data fusion contest, 2013) and Pavia university dataset, and show that the proposed approach performs better than self-supervised approach and supervised training. Shivam Pande, Nassim Ait Ali Braham, Yi Wang 0072, Conrad M. Albrecht, Biplab Banerjee, Xiao Xiang Zhu 0001 |
IGARSS | 5 |
| 2023 | Visual Question Answering in Remote Sensing with Cross-Attention and Multimodal Information BottleneckabstractIn this research, we deal with the problem of visual question answering (VQA) in remote sensing. While remotely sensed images contain information significant for the task of identification and object detection, they pose a great challenge in their processing because of high dimensionality, volume and redundancy. Furthermore, processing image information jointly with language features adds additional constraints, such as mapping the corresponding image and language features. To handle this problem, we propose a cross attention based approach combined with information maximization. The CNN-LSTM based cross-attention highlights the information in the image and language modalities and establishes a connection between the two, while information maximization learns a low dimensional bottleneck layer, that has all the relevant information required to carry out the VQA task. We evaluate our method on two VQA remote sensing datasets of different resolutions. For the high resolution dataset, we achieve an overall accuracy of 79.11% and 73.87% for the two test sets while for the low resolution dataset, we achieve an overall accuracy of 85.98%. Jayesh Songara, Shivam Pande, Shabnam Choudhury, Biplab Banerjee, Rajbabu Velmurugan |
IGARSS | 4 |
| 2023 | Transfomer Based Hyperspectral Dimensionality Reduction with Gabor Kernel CNN for Feature ExtractionabstractDimensionality Reduction (DR) algorithms are used to identify sparse representation of given dataset. The proposed MT-CGF (Multiheaded Transformer with CNN and Gabor Filter) model uses the Multi-head Transformer-based Attention technique [1] for channel attention, which helps identify the most important bands in the dataset. The attention mask generated by the Transformer model is then used to select the top 'n' optimal bands. CNN with Gabor Filter [2], extract complex spatial and spectral features from the dimensionality reduced dataset. Gabor Filters are linear filters that closely resemble the human visual system and possess optimal localization properties in both spatial and frequency domains. Overall, the proposed MT-CGF model combines the strengths of multiple techniques to perform dimensionality reduction and feature extraction from hyperspectral data. By identifying the most important bands it reduces the dataset's dimensionality without significantly compromising its accuracy. It becomes more manageable and easier to analyze, and it can also help to reduce the risk of overfitting. Harshula Tulapurkar, B. Krishna Mohan, Biplab Banerjee |
IGARSS | 3 |
| 2023 | USIM-DAL: Uncertainty-aware Statistical Image Modeling-based Dense Active Learning for Super-resolutionabstractDense regression is a widely used approach in computer vision for tasks such as image super-resolution, enhancement, depth estimation, etc. However, the high cost of annotation and labeling makes it challenging to achieve accurate results. We propose incorporating active learning into dense regression models to address this problem. Active learning allows models to select the most informative samples for labeling, reducing the overall annotation cost while improving performance. Despite its potential, active learning has not been widely explored in high-dimensional computer vision regression tasks like super-resolution. We address this research gap and propose a new framework called USIM-DAL that leverages the statistical properties of colour images to learn informative priors using probabilistic deep neural networks that model the heteroscedastic predictive distribution allowing uncertainty quantification. Moreover, the aleatoric uncertainty from the network serves as a proxy for error that is used for active learning. Our experiments on a wide variety of datasets spanning applications in natural images (visual genome, BSD100), medical imaging (histopathology slides), and remote sensing (satellite images) demonstrate the efficacy of the newly proposed USIM-DAL and superiority over several dense regression active learning methods. Vikrant Rangnekar, Uddeshya Upadhyay, Zeynep Akata, Biplab Banerjee |
UAI | 4 |
| 2023 | Contrastive Learning of Semantic Concepts for Open-set Cross-domain RetrievalabstractWe consider the problem of image retrieval where query images during testing belong to classes and domains both unseen during training. This requires learning a feature space that has the ability to generalize across both classes and domains together. To this end, we propose semantic contrastive concept network (SCNNet), a new learning framework that helps take a step towards class and domain generalization in a principled fashion. Unlike existing methods that rely on global object representations, SCNNet proposes to learn local feature vectors to facilitate unseen-class generalization. To this end, SCNNet’s key innovations include (a) a novel trainable local concept extraction module that learns an orthonormal set of basis vectors, and (b) computes local features for any unseen-class data as a linear combination of the learned basis set. Next, to enable unseen-domain generalization, SCNNet proposes to generate supervisory signals from an adjacent data modality, i.e., natural language, by mining freely available textual label information associated with images. SCNNet derives these signals from our novel trainable semantic ordinal distance constraints that ensure semantic consistency between pairs of images sampled from different domains. Both the proposed modules above enable end-to-end training of the SC-NNet, resulting in a model that helps establish state-of-the-art performance on the standard DomainNet, PACS, and Sketchy benchmark datasets with average Prec@200 improvements of 42.6%, 6.5%, and 13.6% respectively over the most recently reported results. Aishwarya Agarwal, Srikrishna Karanam, Balaji Vasan Srinivasan, Biplab Banerjee |
WACV | 4 |
| 2023 | GAF-Net: Improving the Performance of Remote Sensing Image Fusion using Novel Global Self and Cross Attention LearningabstractThe notion of self and cross-attention learning has been found to substantially boost the performance of remote sensing (RS) image fusion. However, while the self-attention models fail to incorporate the global context due to the limited size of the receptive fields, cross-attention learning may generate ambiguous features as the feature extractors for all the modalities are jointly trained. This results in the generation of redundant multi-modal features, thus limiting the fusion performance. To address these issues, we propose a novel fusion architecture called Global Attention based Fusion Network (GAF-Net), equipped with novel self and cross-attention learning techniques. We introduce the within-modality feature refinement module through global spectral-spatial attention learning using the query-key-value processing where both the global spatial and channel contexts are used to generate two channel attention masks. Since it is non-trivial to generate the cross-attention from within the fusion network, we propose to leverage two auxiliary tasks of modality-specific classification to produce highly discriminative cross-attention masks. Finally, to ensure non-redundancy, we propose to penalize the high correlation between attended modality-specific features. Our extensive experiments on five benchmark datasets, including optical, multispectral (MS), hyperspectral (HSI), light detection and ranging (LiDAR), synthetic aperture radar (SAR), and audio modalities establish the superiority of GAF-Net concerning the literature. Ankit Jha, Shirsha Bose, Biplab Banerjee |
WACV | 3 |
| 2023 | MORGAN: Meta-Learning-based Few-Shot Open-Set Recognition via Generative Adversarial NetworkabstractIn few-shot open-set recognition (FSOSR) for hyperspectral images (HSI), one major challenge arises due to the simultaneous presence of spectrally fine-grained known classes and outliers. Prior research on generative FSOSR cannot handle such a situation due to their inability to approximate the open space prudently. To address this issue, we propose a method, Meta-learning-based Open-set Recognition via Generative Adversarial Network (MORGAN), that can learn a finer separation between the closed and the open spaces. MORGAN seeks to generate class-conditioned adversarial samples for both the closed and open spaces in the few-shot regime using two GANs by judiciously tuning noise variance while ensuring discriminability using a novel Anti-Overlap Latent (AOL) regularizer. Adversarial samples from low noise variance amplify known class data density, and we use samples from high noise variance to augment "known-unknowns". A first-order episodic strategy is adapted to ensure stability in the GAN training. Finally, we introduce a combination of metric losses which push these augmented "known-unknowns" or outliers to disperse in the open space while condensing known class distributions. Extensive experiments on four benchmark HSI datasets indicate that MORGAN achieves state-of-the-art FSOSR performance consistently.1 Debabrata Pal, Shirsha Bose, Biplab Banerjee, Yogananda V. Jeppu |
WACV | 3 |
| 2023 | Multi-head attention with CNN and wavelet for classification of hyperspectral image
Harshula Tulapurkar, Biplab Banerjee, Krishna Mohan Buddhiraju |
Neural Comput. Appl. | 2 |
| 2023 | Self-supervision assisted multimodal remote sensing image classification with coupled self-looping convolution networks
Shivam Pande, Biplab Banerjee |
Neural Networks | 2 |
| 2023 | MAML-SR: Self-adaptive super-resolution networks via multi-scale optimized attention-aware meta-learning
Debabrata Pal, Shirsha Bose, Deeptej More, Ankit Jha, Biplab Banerjee, Yogananda V. Jeppu |
Pattern Recognit. Lett. | 5 |
| 2023 | MDFS-Net: Multidomain Few Shot Classification for Hyperspectral Images With Support Set ReconstructionabstractDeep neural networks are highly specialized for a given task and visual domain, which can limit their practical use. To address this issue, recent studies have proposed to learn universal feature extractors that can be used across multiple domains simultaneously, inspired by the success of transfer learning. However, these universal features are still inferior to specialized networks. In the context of hyperspectral image (HSI) classification, the lack of labeled training samples due to high cost and the restriction to single-domain learning further complicates the problem. To overcome these challenges, we propose a solution that combines the problems of multi-domain learning (MDL) and few-shot learning (FSL) for HSI classification. Our goal is to train a highly shareable network with all domains in the low-shot training regime. We call our network the Multi-Domain Few-Shot (MDFS) network, which shares the majority of model parameters (specifically, convolution and dense layer parameters) across domains while keeping domain-specific batch-normalization layers separate to capture domain characteristics. To address the overfitting issue in few-shot models, we supplement the main classification task with an auxiliary self-supervised task. We test our proposed method on five benchmark HSI datasets and find that MDFS-Net consistently outperforms relevant baselines convincingly. Our approach offers a promising solution for HSI classification in remote sensing by enabling the design of a unified classification system that can work with multiple HSI sites (domains) and fewer labeled samples. Ankit Jha, Biplab Banerjee |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Extreme Value Meta-Learning for Few-Shot Open-Set Recognition of Hyperspectral ImagesabstractRecent advancements in prototype-based Few-Shot Open-Set Recognition (FSOSR) approaches reject outliers based on the high metric distances from theknownclass prototypes and fail to distinguish spectrally fine-grained land cover outliers. Learning only the Euclidean distance fit spherical distributions ignores the essential distribution parameters like shift, shape, and scale. The conventional meta-training of FSOSR also ignores the topological consistency of theknownclasses impacting reduced closed and open accuracy in the meta-testing phase. Moreover, the existing hyperspectral outlier detection methods do not provide intuition about the rejected outlier’s land cover category. To tackle the aforesaid problems, we introduceExtreme Value Meta-Learning(EVML), where we fit Weibull distributions per known class based on the limited support-set distances from respective prototypes. A newly proposed Prototypical OpenMax (P-OpenMax) layer leverages these meta-trained Weibull models and calibrates the query distances to reject fine-grained outliers. Then, to learn the topological consistency, we split all the samples in an episode into four parts, including the prototype and its sameknownclass queries, otherknownclass queries, and the remainingknown-unknownqueries. A novel open quadruplet loss ensures that a prototype’s same-class queries reside closer than the otherknown-class andknown-unknownqueries. Finally, we coarse classify the detected outliers into major land cover categories and perform cross-dataset incremental FSOSR to enhance robustness over unknown geographical regions. We validate the efficacy of EVML over four benchmark hyperspectral datasets. Debabrata Pal, Shirsha Bose, Biplab Banerjee, Yogananda V. Jeppu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Teaching CNNs to Mimic Human Visual Cognitive Process & Regularise Texture-Shape BiasabstractRecent experiments in computer vision demonstrate texture bias as the primary reason for supreme results in models employing Convolutional Neural Networks (CNNs), conflicting with early works claiming that these networks identify objects using shape. It is believed that the cost function forces the CNN to take a greedy approach and develop a proclivity for local information like texture to increase accuracy, thus failing to explore any global statistics. We propose CognitiveCNN, a new intuitive architecture, inspired from feature integration theory in psychology to utilise human-interpretable feature like shape, texture, edges etc. to reconstruct, and classify the image. We define novel metrics to quantify the "relevance" of "abstract information" present in these modalities using attention maps. We further introduce a regularisation method which ensures that each modality like shape, texture etc. gets proportionate influence in a given task, as it does for reconstruction; and perform experiments to show the resulting boost in accuracy and robustness, be-sides imparting explainability to these CNNs for achieving superior performance in object recognition. Satyam Mohla, Anshul Nasery, Biplab Banerjee |
ICASSP | 3 |
| 2022 | Zero-Shot Sketch Based Image Retrieval Using Graph TransformerabstractThe performance of a zero-shot sketch-based image retrieval (ZS-SBIR) task is primarily affected by two challenges. The substantial domain gap between image and sketch features needs to be bridged, while at the same time the side information has to be chosen tactfully. Existing literature has shown that varying the semantic side information greatly affects the performance of ZS-SBIR. To this end, we propose a novel graph transformer based zero-shot sketch-based image retrieval (GTZSR) framework for solving ZS-SBIR tasks which uses a novel graph transformer to preserve the topology of the classes in the semantic space and propagates the context-graph of the classes within the embedding features of the visual space. To bridge the domain gap between the visual features, we propose minimizing the Wasserstein distance between images and sketches in a learned domain-shared space. We also propose a novel compatibility loss that further aligns the two visual domains by bridging the domain gap of one class with respect to the domain gap of all other classes in the training set. Experimental results obtained on the extended Sketchy, TU-Berlin, and QuickDraw datasets exhibit sharp improvements over the existing state-of-the-art methods in both ZS-SBIR and generalized ZS-SBIR. Sumrit Gupta, Ushasi Chaudhuri, Biplab Banerjee, Saurabh Kumar 0005 |
ICPR | 3 |
| 2022 | RSINet: Inpainting Remotely Sensed Images Using Triple GAN FrameworkabstractWe tackle the problem of image inpainting in the remote sensing domain. Remote sensing images possess high resolution and geographical variations, that render the conventional inpainting methods less effective. This further entails the requirement of models with high complexity to sufficiently capture the spectral, spatial and textural nuances within an image, emerging from its high spatial variability. To this end, we propose a novel inpainting method that individually focuses on each aspect of an image such as edges, colour and texture using a task specific GAN. Moreover, each individual GAN also incorporates the attention mechanism that explicitly extracts the spectral and spatial features. To ensure consistent gradient flow, the model uses residual learning paradigm, thus simultaneously working with high and low level features. We evaluate our model, alongwith previous state of the art models, on the two well known remote sensing datasets, Open Cities AI and Earth on Canvas, and achieve competitive performance. The code can be referred here: https://github.com/advaitkumar3107/RSINet. Advait Kumar, Dipesh Tamboli, Shivam Pande, Biplab Banerjee |
IGARSS | 4 |
| 2022 | Feedback Convolution Based Autoencoder for Dimensionality Reduction in Hyperspectral ImagesabstractHyperspectral images (HSI) possess a very high spectral res-olution (due to innumerous bands), which makes them invalu-able in the remote sensing community for landuse/land cover classification. However, the multitude of bands forces the algorithms to consume more data for better performance. To tackle this, techniques from deep learning are often explored, most prominently convolutional neural networks (CNN) based autoencoders. However, one of the main limitations of conventional CNNs is that they only have forward connections. This prevents them to generate robust representations since the information from later layers is not used to refine the earlier layers. Therefore, we introduce a 1D-convolutional autoencoder based on feedback connections for hyperspec-tral dimensionality reduction. Feedback connections create self-updating loops within the network, which enable it to use future information to refine past layers. Hence, the low dimensional code has more refined information for efficient classification. The performance of our method is evaluated on Indian pines 2010 and Indian pines 1992 HSI datasets, where it surpasses the existing approaches. Shivam Pande, Biplab Banerjee |
IGARSS | 2 |
| 2022 | Semantics-Driven Generative Replay for Few-Shot Class Incremental LearningabstractWe deal with the problem of few-shot class incremental learning (FSCIL), which requires a model to continuously recognize new categories for which limited training data are available. Existing FSCIL methods depend on prior knowledge to regularize the model parameters for combating catastrophic forgetting. Devising an effective prior in a low-data regime, however, is not trivial. The memory-replay based approaches from the fully-supervised class incremental learning (CIL) literature cannot be used directly for FSCIL as the generative memory-replay modules of CIL are hard to train from few training samples. However, generative replay can tackle both the stability and plasticity of the models simultaneously by generating a large number of class-conditional samples. Convinced by this fact, we propose a generative modeling-based FSCIL framework using the paradigm of memory-replay in which a novel conditional few-shot generative adversarial network (GAN) is incrementally trained to produce visual features while ensuring the stability-plasticity trade-off through novel loss functions and combating the mode-collapse problem effectively. Furthermore, the class-specific synthesized visual features from the few-shot GAN are constrained to match the respective latent semantic prototypes obtained from a well-defined semantic space. We find that the advantages of this semantic restriction is two-fold, in dealing with forgetting, while making the features class-discernible. The model requires a single per-class prototype vector to be maintained in a dynamic memory buffer. Experimental results on the benchmark and large-scale CiFAR-100, CUB-200, and Mini-ImageNet confirm the superiority of our model over the current FSCIL state of the art. Aishwarya Agarwal, Biplab Banerjee, Fabio Cuzzolin, Subhasis Chaudhuri |
ACM Multimedia | 2 |
| 2022 | Few-Shot Open-Set Recognition of Hyperspectral Images with Outlier Calibration NetworkabstractWe tackle the few-shot open-set recognition (FSOSR) problem in the context of remote sensing hyperspectral image (HSI) classification. Prior research on OSR mainly considers an empirical threshold on the class prediction scores to reject the outlier samples. Further, recent endeavors in few-shot HSI classification fail to recognize outliers due to the ‘closed-set’ nature of the problem and the fact that the entire class distributions are unknown during training. To this end, we propose to optimize a novel outlier calibration network (OCN) together with a feature extraction module during the meta-training phase. The feature extractor is equipped with a novel residual 3D convolutional block attention network (R3CBAM) for enhanced spectral-spatial feature learning from HSI. Our method rejects the outliers based on OCN prediction scores barring the need for manual thresholding. Finally, we propose to augment the query set with synthesized support set features during the similarity learning stage in order to combat the data scarcity issue of few-shot learning. The superiority of the proposed model is showcased on four benchmark HSI datasets.1 Debabrata Pal, Valay Bundele, Renuka Sharma, Biplab Banerjee, Yogananda V. Jeppu |
WACV | 4 |
| 2022 | FRIDA - Generative feature replay for incremental domain adaptation
Sayan Rakshit, Anwesh Mohanty, Ruchika Chavhan, Biplab Banerjee, Gemma Roig, Subhasis Chaudhuri |
Comput. Vis. Image Underst. | 4 |
| 2022 | BDA-SketRet: Bi-level domain adaptation for zero-shot SBIR
Ushasi Chaudhuri, Ruchika Chavan, Biplab Banerjee, Anjan Dutta 0001, Zeynep Akata |
Neurocomputing | 3 |
| 2022 | A Zero-Shot Sketch-Based Intermodal Object Retrieval Scheme for Remote Sensing ImagesabstractDomain-agnostic data retrieval has lately become essential amidst the availability of large-scale data from different types of sensors. However, the unavailability of a sufficient amount of samples of certain classes during training curtails the utility of existing retrieval models in remote sensing (RS) applications. Here, we propose a novel framework for zero-shot intermodal data retrieval of RS data. Thereupon, we design an encoder–decoder structure that ensures enhanced overlapping among the two data domains utilizing cross-triplet and cross-projection loss functions. Furthermore, we propose a sketch-based representation of the RS databaseEarth on Canvaswith diverse classes. We perform a thorough benchmarking of this data set and demonstrate that the proposed framework outperforms state-of-the-art methods for zero-shot sketch-based retrieval framework for RS data. Ushasi Chaudhuri, Biplab Banerjee, Avik Bhattacharya, Mihai Datcu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Attention-Driven Graph Convolution Network for Remote Sensing Image RetrievalabstractGraph convolution networks (GCNs) are useful in remote sensing (RS) image retrieval. It is found to be effective because, in a graph representation, the relative geometrical interactions between different regions (or segments) are appropriately captured, along with their region-wise features in their region adjacency graphs. Also, the attention mechanism has often been applied to the nodes to highlight the essential features in each node. In this regard, a significant amount of high-frequency information is missed since each image segment is effectively summarized within a single node. To account for this and increase the learning capacity, we propose to attend over the edge/adjacency matrix to highlight the interactions among meaningful regions that contribute to supervised learning from images. We exploit this novel edge attention mechanism together with node attention to highlight essential image context by allowing more importance to the meaningful neighboring regions that highlight a relevant node. We implement the proposed context-attended GCN framework for image retrieval on the benchmarked UC-Merced and the PatternNet datasets. We observe a notable improvement in the results compared to the state of the art. Ushasi Chaudhuri, Biplab Banerjee, Avik Bhattacharya, Mihai Datcu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Dual-Path Morph-UNet for Road and Building Segmentation From Satellite ImagesabstractBuilding footprints and road network detection have gained significant attention for map preparation, humanitarian aid dissemination, disaster management, to name a few. Traditionally, morphological filters excel at extracting shape features from remotely sensed images and have been widely used in the literature. However, the structural element (SE) dimension selection impedes these classical and learning-based methods utilizing any morphological operators. To overcome this aspect, we propose a novel framework to extract road and building from remote sensing (RS) images by exploiting morphological networks. The method predominantly aims at learning an optimized SE to capture variably-sized building and road footprints. We substitute convolutions with 2-D morphological operations in the basic building blocks of the network architecture (Dual-path Morph-UNet) to manage the intricate task of optimizing the SE in addition to the actual segmentation task. The dual-path framework incorporates parallel residual and dense paths in an encoder-decoder architecture, which permits learning of higher-level feature representations with fewer parameters. Finally, we implement the proposed framework on the benchmarked Massachusetts roads and buildings dataset and demonstrate superior results than the state-of-the-art (SOTA). In addition, the proposed network consists of$10\times $less learnable parameters than the SOTA methods. Moni Shankar Dey, Ushasi Chaudhuri, Biplab Banerjee, Avik Bhattacharya |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Exploring Transformer and Multilabel Classification for Remote Sensing Image CaptioningabstractHigh-resolution remote sensing images are now available with the progress of remote sensing technology. With respect to popular remote sensing tasks like scene classification, image captioning provides comprehensible information about such images by summarizing the image content in human-readable text. Most existing remote sensing image captioning methods are based on deep learning-based encoder-decoder frameworks, using Convolutional Neural Network or Recurrent Neural Network as the backbone of such frameworks. Such frameworks show a limited capability to analyze sequential data and cope with the lack of captioned remote sensing training images. Recently introduced Transformer architecture exploits self-attention to obtain superior performance for sequence-analysis tasks. Inspired by this, in this work, we employ a Transformer as an encoder-decoder for remote sensing image captioning. Moreover, to deal with the limited training data, an auxiliary decoder is used that further helps the encoder in the training process. The auxiliary decoder is trained for multi-label scene classification due to its conceptual similarity to image captioning and capability of highlighting semantic classes. To the best of our knowledge, this is the first work exploiting multi-label classification to improve remote sensing image captioning. Experimental results on the UC Merced caption data set show the efficacy of the proposed method. The implementation details can be found in https://gitlab.lrz.de/ai4eo/captioningMultilabel. Hitesh Kandala, Sudipan Saha, Biplab Banerjee, Xiao Xiang Zhu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | SPN: Stable Prototypical Network for Few-Shot Learning-Based Hyperspectral Image ClassificationabstractWe tackle the problem of few-shot image classification in the context of remote sensing hyperspectral images (HSIs). Due to the difficulties in collecting a large number of labeled training samples, the few-shot classification techniques hold much prominence in remote sensing in general. One of the bottlenecks in designing few-shot learning (FSL) systems arises from the fact that the model is likely to overfit in the presence of few training samples and the complex spectral feature distributions of the land-cover classes. To this end, we introduce a stable prototypical network (SPN) for FSL by judiciously incorporating dropout and DropBlock-based regularizers jointly within the framework and averaging model parameters using the Monte Carlo approximation. Besides, a novel variance loss term to reduce the uncertainty of the network is considered together with the cross-entropy-based classification loss to train the model in an end-to-end manner. The experimental analysis on three benchmark HSI datasets confirms the SPN’s superior performance. Debabrata Pal, Valay Bundele, Biplab Banerjee, Yogananda V. Jeppu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | SSMTReID-Net: Multi-Target Unsupervised Domain Adaptation for Person Re-Identification
Anwesh Mohanty, Biplab Banerjee, Rajbabu Velmurugan |
Pattern Recognit. Lett. | 2 |
| 2022 | Zero-Shot Cross-Modal Retrieval for Remote Sensing Images With Minimal SupervisionabstractThe performance of a deep-learning-based model primarily relies on the diversity and size of the training dataset. However, obtaining such a large amount of labeled data for practical remote sensing applications is expensive and labor-intensive. Training protocols have been previously proposed for few-shot learning (FSL) and zero-shot learning (ZSL). However, FSL is not compatible with handling unobserved class data at the inference phase, while ZSL requires many training samples of the seen classes. In this work, we propose a novel training protocol for image retrieval and name it aslabel-deficit zero-shot learning(LDZSL). We use this novel LDZSL training protocol for the challenging task of cross-sensor data retrieval in remote sensing. This protocol uses very few labeled data samples of the seen classes during training and interprets unobserved class data samples at the inference phase. This strategy is critical as some data modalities are hard to annotate without domain experts. This work proposes a novel bi-level Siamese network to perform the LDZSL cross-sensor retrieval of multispectral and SAR images. We utilize the available geo-referenced SAR and multispectral data to domain align the embedding features of the two modalities. We experimentally demonstrate the proposed model’s efficacy using the So2Sat dataset compared to the existing state-of-the-art models of the ZSL framework trained under a reduced training set. We also show the generalizability of the proposed model using a sketch-based image retrieval task. Experimental results on the Earth on Canvas dataset exhibit comparative performance over the literature. Ushasi Chaudhuri, Rupak Bose, Biplab Banerjee, Avik Bhattacharya, Mihai Datcu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | SemGIF: A Semantics Guided Incremental Few-shot Learning Framework with Generative Replay
S. Divakar Bhat, Biplab Banerjee, Subhasis Chaudhuri |
BMVC | 2 |
| 2021 | Two Headed Dragons: Multimodal Fusion And Cross Modal TransactionsabstractAs the field of remote sensing is evolving, we witness the accumulation of information from several modalities, such as multispectral (MS), hyperspectral (HSI), LiDAR etc. Each of these modalities possess its own distinct characteristics and when combined synergistically, perform very well in the recognition and classification tasks. However, fusing multiple modalities in remote sensing is cumbersome due to highly disparate domains. Furthermore, the existing methods do not facilitate cross-modal interactions. To this end, we propose a novel transformer based fusion method for HSI and LiDAR modalities. The model is composed of stacked auto encoders that harness the cross key-value pairs for HSI and LiDAR, thus establishing a communication between the two modalities, while simultaneously using the CNNs to extract the spectral and spatial information from HSI and LiDAR. We test our model on Houston (Data Fusion Contest – 2013) and MUUFL Gulfport datasets and achieve competitive results. Rupak Bose, Shivam Pande, Biplab Banerjee |
ICIP | 3 |
| 2021 | Attention-Driven Cross-Modal Remote Sensing Image RetrievalabstractIn this work, we address a cross-modal retrieval problem in remote sensing (RS) data. A cross-modal retrieval problem is more challenging than the conventional uni-modal data retrieval frameworks as it requires learning of two completely different data representations to map onto a shared feature space. For this purpose, we chose a photo-sketch RS database. We exploit the data modality comprising more spatial information (sketch) to extract the other modality features (photo) with cross-attention networks. This sketch-attended photo features are more robust and yield better retrieval results. We validate our proposal by performing experiments on the benchmarked Earth on Canvas dataset. We show a boost in the overall performance in comparison to the existing literature. Besides, we also display the Grad-CAM visualizations of the trained model's weights to highlight the framework's efficacy. Ushasi Chaudhuri, Biplab Banerjee, Avik Bhattacharya, Mihai Datcu |
IGARSS | 2 |
| 2021 | Attention Based Convolution Autoencoder for Dimensionality Reduction in Hyperspectral ImagesabstractHyperspectral images (HSIs) are being actively used for land use/land cover classification owing to their high spectral resolution. However, this leads to the problem of high dimensionality, making the algorithms data hungry. To resolve these issues, deep learning techniques, such as convolution neural networks (CNNs) based autoencoders, are used. However, traditional CNNs tend to focus on all the features irrespective of their importance, leading to weaker representations. To overcome this, we incorporate attention modules in our autoencoder architecture. These attention modules explicitly focus on more important wavelengths, leading to better transformation of the features in the low dimension. In the proposed method, the attention driven encoder transforms high dimension features to low dimensions, considering their relative importance, while the CNN based decoder reconstructs the original features. We evaluate our method on Indian pines 2010 and Indian pines 1992 hyperspectral datasets, where it surpasses the previous approaches. Shivam Pande, Biplab Banerjee |
IGARSS | 2 |
| 2021 | Bidirectional GRU Based Autoencoder for Dimensionality Reduction in Hyperspectral ImagesabstractHyperspectral images (HSI) are being extensively used in land use/land cover classification because they possess high spectral resolution. Although, this leads to better reflectance distinguishability, the problem of high dimensionality also occurs, making the algorithms data greedy. To counter it, deep learning models, such as autoencoders, are being used. To exploit the contiguous nature of HSIs, sequential models like recurrent neural network (RNNs) are adopted. However, for longer sequences, RNNs exhibit vanishing gradients. Also, they fail to incorporate the future information, limiting their scope. Hence, we propose a Bidirectional Gated Recurrent Unit based autoencoder (BiGRUAE), to project the high dimensional features to a low dimensional space. The bidirectional nature captures the information, both from past and future states, while the gating mechanism of GRU prevents the vanishing gradient. We evaluate our method on two hyperspectral datasets, namely, Indian pines 2010 and Salinas, where our method surpasses the benchmark methods. Shivam Pande, Biplab Banerjee |
IGARSS | 2 |
| 2021 | Trusting Small Training Dataset for Supervised Change DetectionabstractDeep learning (DL) based supervised change detection (CD) models require large labeled training data. Due to the difficulty of collecting labeled multi-temporal data, unsupervised methods are preferred in the CD literature. However, unsupervised methods cannot fully exploit the potentials of data-driven deep learning and thus they are not absolute alternative to the supervised methods. This motivates us to look deeper into the supervised DL methods and investigate how they can be adopted intelligently for CD by minimizing the requirement of labeled training data. Towards this, in this work we show that geographically diverse training dataset can yield significant improvement over less diverse training datasets of the same size. We propose a simple confidence indicator for verifying the trustworthiness/confidence of supervised models trained with small labeled dataset. Moreover, we show that for the test cases where supervised CD model is found to be less confident/trustworthy, unsupervised methods often produce better result than the supervised ones. Sudipan Saha, Biplab Banerjee, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2021 | ADA-AT/DT: An Adversarial Approach for Cross-Domain and Cross-Task Knowledge TransferabstractWe deal with the problem of cross-task and cross-domain knowledge transfer in the realm of scene understanding for autonomous vehicles. We consider the scenario where supervision is available for a pair of tasks in a source domain while it is available for only one of the tasks in the target domain. Given that, the goal is to perform inference for the task in the target which is devoid of any training information. We argue that the only reported work in learning across tasks and domains (AT/DT) [26] faces the problem of domain shift between the source and target domains, hindering predictions on the target domain when the transfer of knowledge is learned on a statistically different yet related source domain. As a remedy, we develop a novel framework called ADA-AT/DT based on the adversarial training strategy to ensure that the domain-gaps are minimized for the common cross-domain supervised task. This, in effect, helps in realizing a domain-independent task-transfer function that eventually helps in performing improved inference in the target domain. We demonstrate that our proposed method significantly outperforms [26] by using models with 81% fewer trainable parameters. In addition, we perform experiments on a transformation mapping similar to U-Net to ensure maximum exploitation of features for task transfer. Extensive experiments have been performed on four different domains (Synthia, CityScapes, Carla, and KITTI) for two visual tasks (depth estimation and semantic segmentation) to confirm the superiority of our method. Ruchika Chavhan, Ankit Jha, Biplab Banerjee, Subhasis Chaudhuri |
WACV | 3 |
| 2021 | Improved Landcover Classification using Online Spectral Data Hallucination
Saurabh Kumar 0005, Biplab Banerjee, Subhasis Chaudhuri |
Neurocomputing | 2 |
| 2021 | BiophyNet: A Regression Network for Joint Estimation of Plant Area Index and Wet Biomass From SAR DataabstractIn this study, we propose a sequence-to-sequence neural network architecture to jointly estimate the plant area index (PAI) and wet biomass of canola and soybean. The PAI and wet biomass have considerable importance for crop growth stage mapping and monitoring. RADARSAT-2 quad-pol data along within situmeasurements of canola and soybean obtained from the SMAPVEX16 campaign over Manitoba, Canada, are utilized for evaluating the efficiency and accuracy of the proposed estimation methodology. The analysis indicates promising results for the two crops with a correlation coefficient$(r)$in the range of 0.69–0.87. The results also confirm intercorrelation between the PAI and wet biomass for canola and soybean. Subhadip Dey, Ushasi Chaudhuri, Dipankar Mandal, Avik Bhattacharya, Biplab Banerjee, Heather McNairn |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2021 | Empowering Knowledge Distillation via Open Set Recognition for Robust 3D Point Cloud Classification
Ayush Bhardwaj, Sakshee Pimpale, Saurabh Kumar 0005, Biplab Banerjee |
Pattern Recognit. Lett. | 4 |
| 2021 | Adaptive hybrid attention network for hyperspectral image classification
Shivam Pande, Biplab Banerjee |
Pattern Recognit. Lett. | 2 |
| 2020 | SD-MTCNN: Self-Distilled Multi-Task CNN
Ankit Jha, Awanish Kumar, Biplab Banerjee, Vinay P. Namboodiri |
BMVC | 3 |
| 2020 | Multi-source Open-Set Deep Adversarial Domain Adaptation
Sayan Rakshit, Dipesh Tamboli, Pragati Shuddhodhan Meshram, Biplab Banerjee, Gemma Roig, Subhasis Chaudhuri |
ECCV (26) | 4 |
| 2020 | LEt-SNE: A Hybrid Approach to Data Embedding and Visualization Of Hyperspectral ImageryabstractHyperspectral Imagery (and Remote Sensing in general) captured from UAVs or satellites are highly voluminous in nature due to the large spatial extent and wavelengths captured by them. Since analyzing these images requires a huge amount of computational time and power, various dimensionality reduction techniques have been used for feature reduction. Some popular techniques among these falter when applied to Hyperspectral Imagery due to the famed curse of dimensionality. In this paper, we propose a novel approach, LEt-SNE, which combines graph based algorithms like t-SNE and Laplacian Eigenmaps into a model parameterized by a shallow feed forward network. We introduce a new term, Compression Factor, that enables our method to combat the curse of dimensionality. The proposed algorithm is suitable for manifold visualization and sample clustering with labelled or unlabelled data. We demonstrate that our method is competitive with current state-of-the-art methods on hyperspectral remote sensing datasets in public domain. Megh Shukla, Biplab Banerjee, Krishna Mohan Buddhiraju |
ICASSP | 2 |
| 2020 | Saliency-Driven Class Impressions For Feature Visualization Of Deep Neural NetworksabstractIn this paper, we propose a data-free method of extracting Impressions of each class from the classifier's memory. The Deep Learning regime empowers classifiers to extract distinct patterns (or features) of a given class from training data, which is the basis on which they generalize to unseen data. Before deploying these models on critical applications, it is very useful to visualize the features considered to be important for classification. Existing visualization methods develop high confidence images consisting of both background and foreground features. This makes it hard to judge what the important features of a given class are. In this work, we propose a saliency-driven approach to visualize discriminative features that are considered most important for a given task. Another drawback of existing methods is that, confidence of the generated visualizations is increased by creating multiple instances of the given class. We restrict the algorithm to develop a single object per image, which helps further in extracting features of high confidence, and also results in better visualizations. We further demonstrate the generation of negative images as naturally fused images of two or more classes. Our code is available at: https://giChub.com/val-iisc/Saliency-driven-Class-Impressions. Sravanti Addepalli, Dipesh Tamboli, Venkatesh Babu Radhakrishnan, Biplab Banerjee |
ICIP | 4 |
| 2020 | MT-UNET: A Novel U-Net Based Multi-Task Architecture For Visual Scene UnderstandingabstractWe tackle the problem of deep end-to-end multi-task learning (MTL) for jointly performing image segmentation and depth estimation from monocular images. It is proven already that learning several related tasks together helps in attaining improved performance per task than training them autonomously. To this end, we follow the typical U-Net based encoder-decoder architecture (MT-UNet) where the densely connected deep convolutional neural network (CNN) based feature encoder is shared among the tasks while the soft attention based task-specific decoder modules produce the desired outputs. Additionally, we encourage cross-talk (CT) between the tasks by introducing cross-task skip connections at the decoder end with adaptive weight learning for the task-specific loss functions in the final cost measure. We validate the proposed framework on the challenging CityScapes and NYUv2 datasets, where our method sharply outperforms the current state-of-the-art. Ankit Jha, Awanish Kumar, Shivam Pande, Biplab Banerjee, Subhasis Chaudhuri |
ICIP | 4 |
| 2020 | A Novel Actor Dual-Critic Model for Remote Sensing Image CaptioningabstractWe deal with the problem of generating textual captions from optical remote sensing (RS) images using the notion of deep reinforcement learning. Due to the high inter-class similarity in reference sentences describing remote sensing data, jointly encoding the sentences and images encourages prediction of captions that are semantically more precise than the ground truth in many cases. To this end, we introduce an Actor Dual-Critic training strategy where a second critic model is deployed in the form of an encoder-decoder RNN to encode the latent information corresponding to the original and generated captions. While all actor-critic methods use an actor to predict sentences for an image and a critic to provide rewards, our proposed encoder-decoder RNN guarantees high-level comprehension of images by sentence-to-image translation. We observe that the proposed model generates sentences on the test data highly similar to the ground truth and is successful in generating even better captions in many critical cases. Extensive experiments on the benchmark Remote Sensing Image Captioning Dataset (RSICD) and the UCM-captions dataset confirm the superiority of the proposed approach in comparison to the previous state-of-the-art where we obtain a gain of sharp increments in both the ROUGE-L and CIDEr measures. Ruchika Chavhan, Biplab Banerjee, Xiao Xiang Zhu 0001, Subhasis Chaudhuri |
ICPR | 2 |
| 2020 | PoseCVAE: Anomalous Human Activity DetectionabstractAnomalous human activity detection is the task of identifying human activities that differ from the usual. Existing techniques, in general, try to deploy some samples from an open-set (anomalous activities can not be represented as a closed set) to define the discriminator. However, it is non-trivial to obtain novel activity instances. To this end, we propose PoseCVAE, a novel anomalous human activity detection strategy using the notion of generative modeling. We adopt a hybrid training strategy comprising of self-supervised and unsupervised learning. The self-supervised learning helps the encoder and decoder to learn better latent space representation of human pose trajectories. We train our framework to predict future pose trajectory given a normal track of past poses, i.e., the goal is to learn a conditional posterior distribution that represents normal training data. To achieve this we use a novel adaptation of a conditional variational autoencoder (CVAE) and refer it as PoseCVAE. Future pose prediction will be erroneous if the given poses are sampled from a distribution different from the learnt posterior, which is indeed the case with abnormal activities. To further separate the abnormal class, we imitate abnormal poses in the encoded space by sampling from a distinct mixture of Gaussians (MoG). We use a binary cross-entropy (BCE) loss as a novel addition to the standard CVAE loss function to achieve this. We test our framework on three publicly available datasets and achieve comparable performance to existing unsupervised methods that exploit pose information. Yashswi Jain, Ashvini Kumar Sharma, Rajbabu Velmurugan, Biplab Banerjee |
ICPR | 4 |
| 2020 | Distilling Spikes: Knowledge Distillation in Spiking Neural NetworksabstractSpiking Neural Networks (SNN) are energy-efficient computing architectures that exchange spikes for processing information, unlike classical Artificial Neural Networks (ANN). Due to this, SNNs are better suited for real-life deployments. However, similar to ANNs, SNNs also benefit from deeper architectures to obtain improved performance. Furthermore, like the deep ANNs, the memory, compute and power requirements of SNNs also increase with model size, and model compression becomes a necessity. Knowledge distillation is a model compression technique that enables transferring the learning of a large machine learning model to a smaller model with minimal loss in performance. In this paper, we propose techniques for knowledge distillation in spiking neural networks for the task of image classification. We present ways to distill spikes from a larger SNN, also called the teacher network, to a smaller one, also called the student network, while minimally impacting the classification accuracy. We demonstrate the effectiveness of the proposed method with detailed experiments on three standard datasets while proposing novel distillation methodologies and loss functions. We also present a multi-stage knowledge distillation technique for SNNs using an intermediate network to obtain higher performance from the student network. Our approach is expected to open up new avenues for deploying high performing large SNN models on resource-constrained hardware platforms. Ravi Kumar Kushawaha, Saurabh Kumar 0005, Biplab Banerjee, Rajbabu Velmurugan |
ICPR | 3 |
| 2020 | Dimensionality Reduction Using 3D Residual Autoencoder for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) are actively used for land-use/land-cover classification. However, HSIs suffer from the problem of high dimensionality and high spectral-spatial variability leading to requirement of large number of training samples. Deep learning offers several approaches to handle the aforementioned problems but is limited with its own problem of vanishing gradient that creeps in deeper networks (especially CNNs). In this paper, we propose an autoencoder (3D ResAE) that uses 3D convolutions and residual blocks to project the high-dimensional HSI features to a low-dimensional space. 3D convolutions effectively handle the spectral-spatial characteristics whereas, residual block adds an identity mapping, thereby tackling the issue of vanishing gradient. Furthermore, 3D deconvolutions are used to reconstruct the original features, while the network is trained in a semi-supervised manner. Our proposed method is tested on Indian pines and Salinas hyperspectral datasets and the results clearly demonstrate its effectiveness in classification. Shivam Pande, Biplab Banerjee |
IGARSS | 2 |
| 2020 | Leaf Counting in Rice (Oryza Sativa L.) Using Object Detection: A Deep Learning ApproachabstractLeaf count is one of the crucial tasks in plant phenotyping, and leaves are the basic unit of plant architecture involved in photosynthesis, growth, and yield of a plant. Therefore, the total number of leaves per plant is considered as one of the essential physio-morphological plant traits for phenotyping. The current work proposes to estimate the total number of leaves of a rice plant by detecting their leaves tips. A rice plant has a single tip for a single leaf. Hence, this proposed framework counts the total number of leaves by counting the number of leaves tips equal to the number of leaves. You Only Look Once (YOLO) algorithm is used for the detection of the leaves tips as an object. This hypothesis builds a basis for counting the total number of leaves in a plant like rice, and similar field crops such as wheat (Triticum aestivum L), maize (Zea mays L.), sorghum (Sorghum bicolor), barley (Hordeum vulgare L.). The model detected leaves of a rice plant (RGB images) by detecting corresponding leaves tips with YOLO having average accuracy up to 82% and IOU around 0.53-0.60 and estimates the number of leaves in a plant by counting predicted bounding boxes around tips. The model also performed well with the wheat crop. Mukesh Kumar Vishal, Biplab Banerjee, Rohit Saluja, Raju Dhandapani, Viswanathan Chinnusamy, Sudhir Kumar 0003, Rabi N. Sahoo, J. Adinarayana |
IGARSS | 2 |
| 2020 | Generalized Zero-Shot Learning using Generated Proxy Unseen Samples and Entropy SeparationabstractThe recent generative model-driven Generalized Zero-shot Learning (GZSL) techniques overcome the prevailing issue of the model bias towards the seen classes by synthesizing the visual samples of the unseen classes through leveraging the corresponding semantic prototypes. Although such approaches significantly improve the GZSL performance due to data augmentation, they violate the principal assumption of GZSL regarding the unavailability of semantic information of unseen classes during training. In this work, we propose to use a generative model (GAN) for synthesizing the visual proxy samples while strictly adhering to the standard assumptions of the GZSL. The aforementioned proxy samples are generated by exploring the early training regime of the GAN. We hypothesize that such proxy samples can effectively be used to characterize the average entropy of the label distribution of the samples from the unseen classes. Further, we train a classifier on the visual samples from the seen classes and proxy samples using entropy separation criterion such that an average entropy of the label distribution is low and high, respectively, for the visual samples from the seen classes and the proxy samples. Such entropy separation criterion generalizes well during testing where the samples from the unseen classes exhibit higher entropy than the entropy of the samples from the seen classes. Subsequently, low and high entropy samples are classified using supervised learning and ZSL rather than GZSL. We show the superiority of the proposed method by experimenting on AWA1, CUB, HMDB51, and UCF101 datasets. Omkar Gune, Biplab Banerjee, Subhasis Chaudhuri, Fabio Cuzzolin |
ACM Multimedia | 2 |
| 2020 | Multimodal Noisy Segmentation based fragmented burn scars identification in Amazon RainforestabstractDetection of burn marks due to wildfires in inaccessible rain forests is important for various disaster management and ecological studies. Diverse cropping patterns and the fragmented nature of arable landscapes amidst similar looking land patterns often thwart the precise mapping of burn scars. Recent advances in remote-sensing and availability of multimodal data offer a viable time-sensitive solution to classical methods, which often requires human expert intervention. However, computer vision based segmentation methods have not been used, largely due to lack of labelled datasets. In this work we present AmazonNET - a convolutional based network that allows extracting of burn patters from multimodal remote sensing images. The network consists of UNet- a well-known encoder decoder type of architecture with skip connections commonly used in biomedical segmentation. The proposed framework utilises stacked RGB-NIR channels to segment burn scars from the pastures by training on a new weakly labelled noisy dataset from Amazonia. Our model illustrates superior performance by correctly identifying partially labelled burn scars and rejecting incorrectly labelled samples, demonstrating our approach as one of the first to effectively utilise deep learning based segmentation models in multimodal burn scar identification. Satyam Mohla, Sidharth Mohla, Anupam Guha, Biplab Banerjee |
SMC | 4 |
| 2020 | CrossATNet - a novel cross-attention based framework for sketch-based image retrieval
Ushasi Chaudhuri, Biplab Banerjee, Avik Bhattacharya, Mihai Datcu |
Image Vis. Comput. | 2 |
| 2020 | CMIR-NET : A deep learning based model for cross-modal retrieval in remote sensing
Ushasi Chaudhuri, Biplab Banerjee, Avik Bhattacharya, Mihai Datcu |
Pattern Recognit. Lett. | 2 |
| 2019 | Generalized Zero-shot Learning using Open Set Recognition
Omkar Gune, Amit More, Biplab Banerjee, Subhasis Chaudhuri |
BMVC | 3 |
| 2019 | Crop Phenology Classification Using A Representation Learning Network From Sentinel-1 SAR DataabstractThis work deals with the classification of wheat phenology by regressing the synthetic aperture radar (SAR) backscatter coefficients (VV, VH) to vegetation water content (VWC) and plant area index (PAI) through a representation learning network. The representation network architecture consists of a pair (VV, VH) of two regression layers (VWC, PAI) which finally converge to a classification (crop phenology) layer. The study was conducted with the Sentinel-1 C-band SAR data acquired during the SMAPVEX16 campaign in Manitoba, Canada. Using this framework, the wheat phenology was classified to an accuracy of 86.67%. However, in comparison, the classification accuracy reduced by ~ 20% while using only the backscatter coefficients of (VV, VH) polarization channels. The results obtained from this study justifies the potential of using a representation learning scheme for crop phenology classification with SAR data. Subhadip Dey, Dipankar Mandal, Vineet Kumar 0004, Biplab Banerjee, Juan M. Lopez-Sanchez, Heather McNairn, Avik Bhattacharya |
IGARSS | 4 |
| 2019 | Assessment of Sentinel-1 and Sentinel-2 Satellite Imagery for Crop Classification in Indian Region During Kharif and Rabi Crop CyclesabstractReal-time monitoring of agricultural crops is an important exercise because of it’s huge impact on agri-business and agricultural policy management. Identification of crops during multiple crop growth stages can help formulate better agricultural policies and management strategies. In this context, the objective of this article is to evaluate the potential of Sentinel-1 Synthetic Aperture Radar (SAR) and Sentinel-2 optical imagery in crop classification for an Indian region. A multi-class classification algorithm based on the support vector machine (SVM) is applied to the temporal features extracted from the above mentioned satellite data sets. The experiments are conducted for Kharif and Rabi crop cycles with major crops in the region. The experiments suggest that the joint use of optical and radar imagery results in better classification accuracy compared to using them individually. An overall accuracy of 89% and 96% is obtained for Kharif and Rabi crops, respectively. Aniruddha Mahapatra, Saurav Basu, Biplab Banerjee |
IGARSS | 4 |
| 2019 | Siamese graph convolutional network for content based remote sensing image retrieval
Ushasi Chaudhuri, Biplab Banerjee, Avik Bhattacharya |
Comput. Vis. Image Underst. | 2 |
| 2019 | Graph convolutional network for multi-label VHR remote sensing scene recognition
Nagma Khan, Ushasi Chaudhuri, Biplab Banerjee, Subhasis Chaudhuri |
Neurocomputing | 3 |
| 2019 | Hierarchical Metric Learning for Optical Remote Sensing Scene CategorizationabstractWe address the problem of scene classification from optical remote sensing (RS) images based on the paradigm of hierarchical metric learning. Ideally, supervised metric learning strategies learn a projection from a set of training data points so as to minimize intraclass variance while maximizing the interclass separability to the class label space. However, standard metric learning techniques do not incorporate the class interaction information in learning the transformation matrix, which is often considered to be a bottleneck while dealing with fine-grained visual categories. As a remedy, we propose to organize the classes in a hierarchical fashion by exploring their visual similarities and subsequently learn separate distance metric transformations for the classes present at the nonleaf nodes of the tree. We employ an iterative maximum-margin clustering strategy to obtain the hierarchical organization of the classes. Experiment results obtained on the large-scale NWPU-RESISC45 and the popular UC-Merced data sets demonstrate the efficacy of the proposed hierarchical metric learning-based RS scene recognition strategy in comparison to the standard approaches. Akashdeep Goel, Biplab Banerjee, Aleksandra Pizurica |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Structure Aligning Discriminative Latent Embedding for Zero-Shot Learning
Omkar Gune, Biplab Banerjee, Subhasis Chaudhuri |
BMVC | 2 |
| 2018 | Class Specific Coders for Hyper-Spectral Image ClassificationabstractIn this paper, we introduce the paradigm of class specific coders (CSC) for classification of hyper-spectral images (HSI). Apparently, CSC are defined as a set of distinct encoder-decoder (henceforth called a coder) networks where a given coder is trained on the samples of a particular class. In contrast to auto-encoders (AE) which learn an identity mapping of data in an unsupervised fashion, the CSC model, on the other hand, learns re-constructive mappings for all possible pairs of training samples for each class in separate coders. Further, for reducing redundancy, it is ensured that the latent space dimensions of the coders are orthogonal to each other. Once the CSC are trained, we introduce a feature encoding by concatenating the latent representations of all the class specific coders. Experimental results obtained on the benchmark Botswana and Indian Pines datasets ensure that the CSC model outperforms a number of AE variants significantly. Sanatan Sharma, Akashdeep Goel, Omkar Gune, Biplab Banerjee, Subhasis Chaudhuri |
ICIP | 4 |
| 2018 | Discriminative Latent Visual Space For Zero-Shot Object ClassificationabstractIn this paper We deal with the problem of zero-shot visual recognition. The standard zero-shot learning (ZSL) pipeline is based on the idea of learning a functional mapping from a visual embedding space to an auxiliary semantic space for a set of seen categories. In the testing phase, the task is to recognize a set of novel categories which are semantically linked to the already known ones. Although such a pipeline is inherently supervised, there exists very few endeavours in the context of ZSL that enforce discrimination in learning this mapping. In this work, we propose a novel encoder-decoder network to explore the possibility of learning an intermediate latent space for the visual features, which is deemed to be simultaneously reconstructive and discriminative. By reaching a trade-off between the joint (re)construction of the visual and the semantic embedding spaces, while ensuring separability among the known classes, the proposed model better generalizes to the unknown categories. Experimental results obtained on challenging datasets, such as AwA, CUB, and ImageNet-2, establish the efficacy of such a discriminative latent space for the standard ZSL setup. Abhinaba Roy, Biplab Banerjee, Vittorio Murino |
ICPR | 2 |
| 2018 | Discriminative body part interaction mining for mid-level action representation and classification
Abhinaba Roy, Biplab Banerjee, Vittorio Murino |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Scene Recognition From Optical Remote Sensing Images Using Mid-Level Deep Feature MiningabstractWe solve the problem of scene recognition from very high-resolution optical satellite remote sensing (RS) images by exploring the notion of mid-level feature mining. The existing mid-level feature extraction techniques are based on applying feature encodings over a set of discriminatively selected localized feature descriptors from the images. Such techniques inherently suffer from two shortcomings: 1) the local descriptors are not enough discriminative, since they are mostly based on scale invariant feature transform (SIFT) like ad hoc features and 2) the definition of a robust ranking function to select discriminative local features is nontrivial. As a remedy, we propose a pattern mining-based approach for an efficient discovery of mid-level visual elements, which considers convolutional neural network features of the category-independent region proposals extracted from the images as the local descriptors. While the region proposals depict better semantic information than the SIFT like features, the proposed pattern mining strategy can efficiently highlight the correlations between such local descriptors and the class labels. Experimental results suggest that the proposed technique outperforms a number of existing mid-level feature descriptors for the standard optical RS data sets. Biplab Banerjee, Subhasis Chaudhuri |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Efficient pooling of image based CNN features for action recognition in videosabstractIn this paper, we propose a new video representation incorporating image based deep features and an efficient pooling strategy for the purpose of action recognition. The Convolutional Neural Network (CNN) based features have very recently emerged as the new state of the art for image classification. Several attempts have been made to extend such CNN models for videos by explicitly focusing on the temporal evolution of the frames. Feature pooling is one such approach which represents video sequences in terms of some statistical properties of the feature dimensions over the frames. However, traditional pooling strategies including max or average pooling explicitly fail to capture the temporal progression of the frame-level contents. In contrast to previous pooling techniques, we propose a two-level video representation which separately focuses on the entire video as well as a number of video sub-volumes. In both levels, we introduce a generic time series pooling on the frame-level deep CNN features efficiently. Further, a self-tuning spectral clustering is considered to highlight video snippets which are highly probable to contain significant sub-action sequences. We validate the proposed feature encoding on the challenging KTH-actions and UCF-50 datasets and find that the proposed encoding outperforms traditional pooling based feature representations by substantial margin. Biplab Banerjee, Vittorio Murino |
ICASSP | 1 |
| 2017 | A Novel Dictionary Learning based Multiple Instance Learning Approach to Action Recognition from VideosabstractIn this paper we deal with the problem of action recognition from unconstrained videos under the notion of multiple instance learning (MIL).The traditional MIL paradigm considers the data items as bags of instances with the constraint that the positive bags contain some class-specific instances whereas the negative bags consist of instances only from negative classes.A classifier is then further constructed using the bag level annotations and a distance metric between the bags.However, such an approach is not robust to outliers and is time consuming for a moderately large dataset.In contrast, we propose a dictionary learning based strategy to MIL which first identifies class-specific discriminative codewords, and then projects the bag-level instances into a probabilistic embedding space with respect to the selected codewords.This essentially generates a fixedlength vector representation of the bags which is specifically dominated by the properties of the class-specific instances.We introduce a novel exhaustive search strategy using a support vector machine classifier in order to highlight the class-specific codewords.The standard multiclass classification pipeline is followed henceforth in the new embedded feature space for the sake of action recognition.We validate the proposed framework on the challenging KTH and Weizmann datasets, and the results obtained are promising and comparable to representative techniques from the literature. Abhinaba Roy, Biplab Banerjee, Vittorio Murino |
ICPRAM | 2 |
| 2016 | An unsupervised hidden Markov random field based segmentation of polarimetric SAR imagesabstractThis paper proposes an iterative unsupervised Markov Random Field (MRF) based segmentation technique for polarimetric Synthetic Aperture Radar (SAR) image using the optimized scattering mechanism similarity parameters. Parameter estimation for the MRF model is generally performed from the available training data in order to perform tasks including semantic image segmentation. Since the current scenario is entirely unsupervised, the parameter estimation is performed iteratively using the Expectation Maximization (EM) technique considering the classes are distributed according to Gaussian functions. Further, we model the pairwise potential of the MRF cost function using a weighted combination of the similarity parameters. Results obtained on a fully polarimetric SAR data establishes the potential of such unsupervised random field models for analyzing SAR data effectively. Biplab Banerjee, Shaunak De, Surendar Manickam, Avik Bhattacharya |
IGARSS | 1 |
| 2016 | Domain Adaptation in the Absence of Source Domain Labeled Samples - A Coclustering-Based ApproachabstractWe propose a novel coclustering-based domain-adaptation algorithm for simultaneously generating classification maps for a set of remote sensing (RS) multitemporal images in this letter. Unsupervised domain-adaptation techniques consider two different but related domains: a source domain with ample number of labeled samples and a target domain with no labeled data. The task at hand is to build an inference model exploring the available data that is expected to work consistently well in both the domains. This is a challenging problem, since the probability distributions governing both the domains are substantially different leading to the violation of the probably approximate correct assumptions of statistical learning theory. We consider an even complex scenario in this letter by assuming the absence of source domain training samples in the learning process. Our algorithm broadly consists of two stages: first, data from both the domains are projected into a common subspace using geodesic flow kernel in a Grassmannian manifold, and we further propose an iterative coclustering technique to obtain the consistent clustering outcomes for both the domains in the newly defined space. In line with the traditional domain-adaptation approaches, we also consider that both the domains contain the same set of semantic land-cover classes. The proposed method is simple, scalable, and results in highly precise clustering outputs for standard RS data sets. Biplab Banerjee, Krishna Mohan Buddhiraju |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | A New Self-Training-Based Unsupervised Satellite Image Classification Technique Using Cluster Ensemble StrategyabstractThis letter addresses the problem of unsupervised land-cover classification of remotely sensed multispectral satellite images from the perspective of cluster ensembles and self-learning. The cluster ensembles combine multiple data partitions generated by different clustering algorithms into a single robust solution. A cluster-ensemble-based method is proposed here for the initialization of the unsupervised iterative expectation-maximization (EM) algorithm which eventually produces a better approximation of the cluster parameters considering a certain statistical model is followed to fit the data. The method assumes that the number of land-cover classes is known. A novel method for generating a consistent labeling scheme for each clustering of the consensus is introduced for cluster ensembles. A maximum likelihood classifier is henceforth trained on the updated parameter set obtained from the EM step and is further used to classify the rest of the image pixels. The self-learning classifier, although trained without any external supervision, reduces the effect of data overlapping from different clusters which otherwise a single clustering algorithm fails to identify. The clustering performance of the proposed method on a medium resolution and a very high spatial resolution image have effectively outperformed the results of the individual clustering of the ensemble. Biplab Banerjee, Francesca Bovolo, Avik Bhattacharya, Lorenzo Bruzzone, Subhasis Chaudhuri, B. Krishna Mohan |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | A Novel Graph-Matching-Based Approach for Domain Adaptation in Classification of Remote Sensing Image PairabstractThis paper addresses the problem of land-cover classification of remotely sensed image pairs in the context of domain adaptation. The primary assumption of the proposed method is that the training data are available only for one of the images (source domain), whereas for the other image (target domain), no labeled data are available. No assumption is made here on the number and the statistical properties of the land-cover classes that, in turn, may vary from one domain to the other. The only constraint is that at least one land-cover class is shared by the two domains. Under these assumptions, a novel graph theoretic cross-domain cluster mapping algorithm is proposed to detect efficiently the set of land-cover classes which are common to both domains as well as the additional or missing classes in the target domain image. An interdomain graph is introduced, which contains all of the class information of both images, and subsequently, an efficient subgraph-matching algorithm is proposed to highlight the changes between them. The proposed cluster mapping algorithm initially clusters the target domain data into an optimal number of groups given the available source domain training samples. To this end, a method based on information theory and a kernel-based clustering algorithm is proposed. Considering the fact that the spectral signature of land-cover classes may overlap significantly, a postprocessing step is applied to refine the classification map produced by the clustering algorithm. Two multispectral data sets with medium and very high geometrical resolution and one hyperspectral data set are considered to evaluate the robustness of the proposed technique. Two of the data sets consist of multitemporal image pairs, while the remaining one contains images of spatially disjoint geographical areas. The experiments confirm the effectiveness of the proposed framework in different complex scenarios. Biplab Banerjee, Francesca Bovolo, Avik Bhattacharya, Lorenzo Bruzzone, Subhasis Chaudhuri, Krishna Mohan Buddhiraju |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | An ant colony optimization based inter domain cluster mapping for domain adaptation in remote sensingabstractA novel ant colony optimization based domain adaptation for satellite images has been proposed in the paper. Given a source domain and a target domain image, it has been considered here that we have labeled training data for the source domain image. The goal is to classify the target domain image for which no prior information is available. The proposed method exploits the advantages of ant colony optimization for performing the adaptation process. The target domain data is first over-clustered and a cross-domain cluster matching strategy based on ant movement is followed next to match the target domain clusters to the source domain land-cover classes. Experiments suggest good matching capabilities of the proposed technique. Shakti Sharma, Krishna Mohan Buddhiraju, Biplab Banerjee |
IGARSS | 3 |
| 2013 | A Novel Graph Based Clustering Technique for Hybrid Segmentation of Multi-spectral Remotely Sensed Images
Biplab Banerjee, Surender Varma G., Krishna Mohan Buddhiraju |
ACIVS | 1 |
| 2013 | Radon transform based edge detection for SAR imageryabstractEdge detection in SAR images has always been a challenge due to the effects of random interference of coherent signal. Unlike optical images, boundary delineation of regions is relatively ineffective for SAR images. The usual edge detectors, successful with incoherent images, yield poor results when applied to radar images, especially those with a small number of looks. Radon spectrum can be used to find the local orientations in an image which embed the edge information [1]. In this paper, we demonstrate the capability of Radon transform to efficiently extract edges from a SAR image covering agricultural landform. One approach to edge detection in polarimetric SAR images is to perform the edge detection separately for each of the polarization channels and subsequently combine the results using a fusion operation [2] [3]. We have made use of HH, HV and VV channels and fused the edge maps of these individual bands with Boolean ‘AND’ operator. This type of fusion has preserved the prominent edge information reasonably while suppressing spurious ones. Surender Varma G., Biplab Banerjee, Arnab Muhuri, Avik Bhattacharya, Krishna Mohan Buddhiraju |
IGARSS | 2 |
| 2012 | Satellite image segmentation: A novel adaptive mean-shift clustering based approachabstractSegmentation of satellite images using a novel adaptive non parametric mean-shift clustering algorithm is proposed in this paper. Image segmentation refers to the process of splitting up an image into its constituent objects. It is also an important step in bridging the semantic gap between low level image interpretation and high level visual analysis. Mean-shift technique is based on the concept of kernel density estimation. It has been applied successfully in diverse vision related tasks including segmentation. The performance of the mean shift algorithm is greatly affected by the size of the parzen window and the terminating criteria. These two issues have been taken care of here in a purely statistical framework. The efficiency of this newly developed adaptive clustering has been judged for segmentation of any initially oversegmented satellite image. The notion of object based image analysis is preserved by initially over segmenting the image by watershed technique. Extensive experiments on several multispectral satellite images have confirmed the effectivity of this proposed approach in comparison to some widely used state of the art segmentation methods. Biplab Banerjee, Surender Varma G., Krishna Mohan Buddhiraju |
IGARSS | 1 |
| 2012 | Representation of desert sand dunes by surface orientations using Radon transformabstractDesert regions are typically characterized by the texture of sand that particular place and the orientation of dunes. We have attempted to visualize and quantify the sand dunes by local surface orientations calculated by Radon transform of sliding window region that covers entire image. This kind of representation can be used as a pre-processing step for segmentation and change detection of sand dunes. Hough transform, though traditionally used [1] to extract linear features is not very suitable for natural images. Radon transform on the other hand is robust to noise [2][3] comparatively and can effectively be used to find dune orientations. First of all an edge image is generated using canny operator or morphological operators and then local orientations are computed. Populations residing near desert land scapes have Livestock and marginal cultivation. To protect their interest's administrative setup need to have accurate understanding of how the directions of sand dunes are changing with respect to time. It is seen that for desert environments, remotely sensed satellite imagery can reliably be used, to identify regions that are nontextured and with various dune orientations using radon transform. Surender Varma G., Biplab Banerjee, Krishna Mohan Buddhiraju |
IGARSS | 2 |