Shirsha Bose

dblp:312/8014 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Finding Dino: A Plug-and-Play Framework for Zero-Shot Detection of Out-of-Distribution Objects Using Prototypes
abstract
8474
Poulami Sinhamahapatra, Franziska Schwaiger, Shirsha Bose, Karsten Roscher, Stephan Günnemann
WACV3
2024 Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization
abstract
We delve into Open Domain Generalization (ODG), marked by domain and category shifts between training's labeled source and testing's unlabeled target domains. Existing solutions to ODG face limitations due to constrained generalizations of traditional CNN backbones and errors in detecting target open samples in the absence of prior knowledge. Addressing these pitfalls, we introduce ODG-CLIP, harnessing the semantic prowess of the vision-language model, CLIP. Our framework brings forth three primary innovations: Firstly, distinct from prevailing paradigms, we conceptualize ODG as a multi-class classification challenge encompassing both known and novel categories. Central to our approach is modeling a unique prompt tailored for detecting unknown class samples, and to train this, we employ a readily accessible stable diffusion model, elegantly generating proxy images for the open class. Secondly, aiming for domain-tailored classification (prompt) weights while ensuring a balance of precision and simplicity, we devise a novel visual stylecentric prompt learning mechanism. Finally, we infuse images with class-discriminative knowledge derived from the prompt space to augment the fidelity of CLIP's visual embeddings. We introduce a novel objective to safeguard the continuity of this infused semantic intel across domains, especially for the shared classes. Through rigorous testing on diverse datasets, covering closed and open-set DG contexts, ODG-CLIP demonstrates clear supremacy, consistently outpacing peers with performance boosts between 8%-16%. Code will be available at https://github.com/mainaksingha01/ODG-CLIP.
Mainak Singha, Ankit Jha, Shirsha Bose, Ashwin R. Nair, Moloud Abdar, Biplab Banerjee
CVPR3
2024 SPDG-Net: Semantics Preserving Domain Augmentation through Style Interpolation for Multi-Source Domain Generalization
abstract
This paper focuses on domain generalization (DG), addressing the challenge of robust classifier learning from multiple source domains for generalizing to unseen ones. DG suffers from limited source domain diversity, which may hinder model generalization. Recent studies explore domain-augmentation strategies but struggle to maintain semantics while altering image styles and generate only a few pseudo domains. To tackle this, we introduce Semantics Preserving DG Network (SPDG-Net). SPDG-Net is a triplet-conditioned U-Net-based GAN that synthesizes various pseudo domains, preserving image semantics through a cycle-consistency constraint. Moreover, our style interpolation-based domain generation produces fine-grained synthetic domains, unlike existing models. We also propose recognizing image styles alongside object classes to reduce model bias. Our experiments across benchmark datasets consistently outperform recent literature in DG.
Advait Kumar, Shirsha Bose, Mohamad Hassan N C, Biplab Banerjee
ICASSP2
2024 StyLIP: Multi-Scale Style-Conditioned Prompt Learning for CLIP-based Domain Generalization
abstract
Large-scale foundation models, such as CLIP, have demonstrated impressive zero-shot generalization performance on downstream tasks, leveraging well-designed language prompts. However, these prompt learning techniques often struggle with domain shift, limiting their generalization capabilities. In our study, we tackle this issue by proposing StyLIP, a novel approach for Domain Generalization (DG) that enhances CLIP’s classification performance across domains. Our method focuses on a domain-agnostic prompt learning strategy, aiming to disentangle the visual style and content information embedded in CLIP’s pre-trained vision encoder, enabling effortless adaptation to novel domains during inference. To achieve this, we introduce a set of style projectors that directly learn the domain-specific prompt tokens from the extracted multi-scale style features. These generated prompt embeddings are subsequently combined with the multi-scale visual content features learned by a content projector. The projectors are trained in a contrastive manner, utilizing CLIP’s fixed vision and text backbones. Through extensive experiments conducted in five different DG settings on multiple benchmark datasets, we consistently demonstrate that StyLIP outperforms the current state-of-the-art (SOTA) methods.
Shirsha Bose, Ankit Jha, Enrico Fini, Mainak Singha, Elisa Ricci 0001, Biplab Banerjee
WACV1
2023 GAF-Net: Improving the Performance of Remote Sensing Image Fusion using Novel Global Self and Cross Attention Learning
abstract
The notion of self and cross-attention learning has been found to substantially boost the performance of remote sensing (RS) image fusion. However, while the self-attention models fail to incorporate the global context due to the limited size of the receptive fields, cross-attention learning may generate ambiguous features as the feature extractors for all the modalities are jointly trained. This results in the generation of redundant multi-modal features, thus limiting the fusion performance. To address these issues, we propose a novel fusion architecture called Global Attention based Fusion Network (GAF-Net), equipped with novel self and cross-attention learning techniques. We introduce the within-modality feature refinement module through global spectral-spatial attention learning using the query-key-value processing where both the global spatial and channel contexts are used to generate two channel attention masks. Since it is non-trivial to generate the cross-attention from within the fusion network, we propose to leverage two auxiliary tasks of modality-specific classification to produce highly discriminative cross-attention masks. Finally, to ensure non-redundancy, we propose to penalize the high correlation between attended modality-specific features. Our extensive experiments on five benchmark datasets, including optical, multispectral (MS), hyperspectral (HSI), light detection and ranging (LiDAR), synthetic aperture radar (SAR), and audio modalities establish the superiority of GAF-Net concerning the literature.
Ankit Jha, Shirsha Bose, Biplab Banerjee
WACV2
2023 MORGAN: Meta-Learning-based Few-Shot Open-Set Recognition via Generative Adversarial Network
abstract
In few-shot open-set recognition (FSOSR) for hyperspectral images (HSI), one major challenge arises due to the simultaneous presence of spectrally fine-grained known classes and outliers. Prior research on generative FSOSR cannot handle such a situation due to their inability to approximate the open space prudently. To address this issue, we propose a method, Meta-learning-based Open-set Recognition via Generative Adversarial Network (MORGAN), that can learn a finer separation between the closed and the open spaces. MORGAN seeks to generate class-conditioned adversarial samples for both the closed and open spaces in the few-shot regime using two GANs by judiciously tuning noise variance while ensuring discriminability using a novel Anti-Overlap Latent (AOL) regularizer. Adversarial samples from low noise variance amplify known class data density, and we use samples from high noise variance to augment "known-unknowns". A first-order episodic strategy is adapted to ensure stability in the GAN training. Finally, we introduce a combination of metric losses which push these augmented "known-unknowns" or outliers to disperse in the open space while condensing known class distributions. Extensive experiments on four benchmark HSI datasets indicate that MORGAN achieves state-of-the-art FSOSR performance consistently.1
Debabrata Pal, Shirsha Bose, Biplab Banerjee, Yogananda V. Jeppu
WACV2
2023 MAML-SR: Self-adaptive super-resolution networks via multi-scale optimized attention-aware meta-learning
Debabrata Pal, Shirsha Bose, Deeptej More, Ankit Jha, Biplab Banerjee, Yogananda V. Jeppu
Pattern Recognit. Lett.2
2023 Extreme Value Meta-Learning for Few-Shot Open-Set Recognition of Hyperspectral Images
abstract
Recent advancements in prototype-based Few-Shot Open-Set Recognition (FSOSR) approaches reject outliers based on the high metric distances from theknownclass prototypes and fail to distinguish spectrally fine-grained land cover outliers. Learning only the Euclidean distance fit spherical distributions ignores the essential distribution parameters like shift, shape, and scale. The conventional meta-training of FSOSR also ignores the topological consistency of theknownclasses impacting reduced closed and open accuracy in the meta-testing phase. Moreover, the existing hyperspectral outlier detection methods do not provide intuition about the rejected outlier’s land cover category. To tackle the aforesaid problems, we introduceExtreme Value Meta-Learning(EVML), where we fit Weibull distributions per known class based on the limited support-set distances from respective prototypes. A newly proposed Prototypical OpenMax (P-OpenMax) layer leverages these meta-trained Weibull models and calibrates the query distances to reject fine-grained outliers. Then, to learn the topological consistency, we split all the samples in an episode into four parts, including the prototype and its sameknownclass queries, otherknownclass queries, and the remainingknown-unknownqueries. A novel open quadruplet loss ensures that a prototype’s same-class queries reside closer than the otherknown-class andknown-unknownqueries. Finally, we coarse classify the detected outliers into major land cover categories and perform cross-dataset incremental FSOSR to enhance robustness over unknown geographical regions. We validate the efficacy of EVML over four benchmark hyperspectral datasets.
Debabrata Pal, Shirsha Bose, Biplab Banerjee, Yogananda V. Jeppu
IEEE Trans. Geosci. Remote. Sens.2