Soumen Basu

dblp:319/2555 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-3915-7545ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Focus on Texture: Rethinking Pre-training in Masked Autoencoders for Medical Image Classification
Chetan Madan, Aarjav Satia, Soumen Basu, Pankaj Gupta 0005, Usha Dutta, Chetan Arora 0001
MICCAI (4)3
2025 LQ-Adapter: ViT-Adapter with Learnable Queries for Gallbladder Cancer Detection from Ultrasound Images
abstract
We focus on the problem of Gallbladder Cancer (GBC) detection from Ultrasound (US) images. The problem presents unique challenges to modern Deep Neural Network (DNN) techniques due to low image quality arising from noise, textures, and viewpoint variations. Tackling such challenges would necessitate precise localization performance by the DNN to identify the discerning features for the downstream malignancy prediction. While several techniques have been proposed in the recent years for the problem, all of these methods employ complex custom architectures. Inspired by the success of foundational models for natural image tasks, along with the use of adapters to fine-tune such models for the custom tasks, we investigate the merit of one such design, ViT-Adapter, for the GBC detection problem. We observe that ViT-Adapter relies pre-dominantly on a primitive CNN-based spatial prior module to inject the localization information via cross-attention, which is inefficient for our problem due to the small pathology sizes, and variability in their appearances due to non-regular structure of the malignancy. In response, we propose, LQ-Adapter, a modified Adapter design for ViT, which improves localization information by leveraging learnable content queries over the basic spatial prior module. Our method surpasses existing approaches, enhancing the mean IoU (mIoU) scores by 5.4%, 5.8%, and 2.7% over ViT-Adapters, DINO, and FocalNet-DINO, respectively on the US image-based GBC detection dataset, and establishing a new state-of-the-art (SOTA). Additionally, we validate the applicability and effectiveness of LQ-Adapter on the Kvasir-Seg dataset for polyp detection from colonoscopy images. Superior performance of our design on this problem as well showcases its capability to handle diverse medical imaging tasks across different datasets. Source code and trained models are publicly released.
Chetan Madan, Mayuna Gupta, Soumen Basu, Pankaj Gupta 0005, Chetan Arora 0001
WACV3
2024 FocusMAE: Gallbladder Cancer Detection from Ultrasound Videos with Focused Masked Autoencoders
abstract
In recent years, automated Gallbladder Cancer (GBC) detection has gained the attention of researchers. Current state-of-the-art (SOTA) methodologies relying on ultra-sound sonography (US) images exhibit limited generalization, emphasizing the need for transformative approaches. We observe that individual US frames may lack sufficient information to capture disease manifestation. This study advocates for a paradigm shift towards video-based GBC detection, leveraging the inherent advantages of spatiotemporal representations. Employing the Masked Autoencoder (MAE) for representation learning, we address shortcomings in conventional image-based methods. We propose a novel design called FocusMAE to systematically bias the selection of masking tokens from high-information regions, fostering a more refined representation of malignancy. Additionally, we contribute the most extensive US video dataset for GBC detection. We also note that, this is the first study on US video-based GBC detection. We validate the proposed methods on the curated dataset, and report a new SOTA accuracy of 96.4% for the GBC detection problem, against an accuracy of 84% by current Image-based SOTA – GBCNet and RadFormer, and 94.7% by Video-based SOTA – AdaMAE. We further demonstrate the generality of the proposed FocusMAE on a public CT-based Covid detection dataset, reporting an improvement in accuracy by 3.3% over current baselines. Project page with source code, trained models, and data is available at: https://gbc-iitd.github.io/focusmae.
Soumen Basu, Mayuna Gupta, Chetan Madan, Pankaj Gupta 0005, Chetan Arora 0001
CVPR1
2023 Gall Bladder Cancer Detection from US Images with only Image Level Labels
Soumen Basu, Ashish Papanai, Pankaj Gupta 0005, Chetan Arora 0001
MICCAI (1)1
2023 How Reliable are the Metrics Used for Assessing Reliability in Medical Imaging?
Soumen Basu, Chetan Arora 0001
MICCAI (3)2
2023 RadFormer: Transformers with global-local attention for interpretable and accurate Gallbladder Cancer detection
Soumen Basu, Pratyaksha Rana, Pankaj Gupta 0005, Chetan Arora 0001
Medical Image Anal.1
2022 Surpassing the Human Accuracy: Detecting Gallbladder Cancer from USG Images with Curriculum Learning
abstract
We explore the potential of CNN-based models for gall-bladder cancer (GBC) detection from ultrasound (USG) images as no prior study is known. USG is the most common diagnostic modality for GB diseases due to its low cost and accessibility. However, USG images are challenging to analyze due to low image quality, noise, and varying viewpoints due to the handheld nature of the sensor. Our exhaustive study of state-of-the-art (SOTA) image classification techniques for the problem reveals that they often fail to learn the salient GB region due to the presence of shadows in the USG images. SOTA object detection techniques also achieve low accuracy because of spurious textures due to noise or adjacent organs. We propose GBCNet to tackle the challenges in our problem. GBCNet first extracts the regions of interest (ROIs) by detecting the GB (and not the cancer), and then uses a new multi-scale, second-order pooling architecture specializing in classifying GBC. To effectively handle spurious textures, we propose a curriculum inspired by human visual acuity, which reduces the texture biases in GBCNet. Experimental results demonstrate that GBC-Net significantly outperforms SOTA CNN models, as well as the expert radiologists. Our technical innovations are generic to other USG image analysis tasks as well. Hence, as a validation, we also show the efficacy of GBCNet in detecting breast cancer from USG images. Project page with source code, trained models, and data is available at https://GBC-iitd.github.io/GBCnet.
Soumen Basu, Pratyaksha Rana, Pankaj Gupta 0005, Chetan Arora 0001
CVPR1
2022 Unsupervised Contrastive Learning of Image Representations from Ultrasound Videos with Hard Negative Mining
Soumen Basu, Somanshu Singla, Pratyaksha Rana, Pankaj Gupta 0005, Chetan Arora 0001
MICCAI (4)1