VLDB 2026 Research / reviewers in the wild / expert
Sandipan Sarma
dblp:336/3663
· DBLP profile ↗
7ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0003-4619-3058ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Funnel-HOI: top-down perception for zero-shot HOI detection
Sandipan Sarma, Agney Talwarr, Arijit Sur |
Mach. Vis. Appl. | 1 |
| 2024 | X-CAUNET: Cross-Color Channel Attention with Underwater Image-Enhancing TransformerabstractUnderwater image enhancement is essential to mitigate the environment-centric noise in images, such as haziness, color degradation, etc. With most existing works focused on processing an RGB image as a whole, the explicit context that can be mined from each color channel separately goes unaccounted for, ignoring the effects produced by the wavelength of light in underwater conditions. In this work, we propose a framework called X-CAUNET that addresses this research gap by using cross-attention transformers. The input image is split into three channels (R-G-B), local context is captured using convolutional layers with different receptive field sizes, and a message-passing mechanism allows for context correlation between them. To maintain consistency, another transformer is used on the original image to aggregate global context, and a weighted combination of all the outputs enhances the input degraded image. Extensive experiments demonstrate we achieve state-of-the-art PSNR and SSIM with 2.66% and 2.11% relative gains. Code is available at: https://github.com/Alik033/X-CAUNET. Alik Pramanick, Sandipan Sarma, Arijit Sur |
ICASSP | 2 |
| 2024 | Boosting Zero-Shot Human-Object Interaction Detection with Vision-Language TransferabstractHuman-Object Interaction (HOI) detection is a crucial task that involves localizing interactive human-object pairs and identifying the actions being performed. Most existing HOI detectors are supervised in nature and lack the ability of zero-shot discovery of unseen interactions. Recently, transformer-based methods have superseded the traditional CNN detectors by aggregating image-wide context but still suffer from the long-tail distribution problem in HOI. In this work, our primary focus is improving HOI detection in images, particularly in zero-shot scenarios. We use an end-to-end transformer-based object detector to localize human-object pairs and yield visual features of actions and objects. Moreover, we adopt the text encoder from a popular visual-language model called CLIP with a novel prompting mechanism to extract semantic information for unseen actions and objects. Finally, we learn a strong visual-semantic alignment and achieve state-of-the-art performance on the challenging HICO-DET dataset across five zero-shot settings, with up to 70.88% relative gains. Code is available at https://github.com/sandipan211/ZSHOI-VLT. Sandipan Sarma, Pradnesh Kalkar, Arijit Sur |
ICASSP | 1 |
| 2024 | Zero-Shot Underwater Gesture Recognition
Sandipan Sarma, Gundameedi Sai Ram Mohan, Hariansh Sehgal, Arijit Sur |
ICPR (7) | 1 |
| 2024 | DiRaC-I: Identifying Diverse and Rare Training Classes for Zero-Shot LearningabstractZero-Shot Learning (ZSL) is an extreme form of transfer learning that aims at learning from a few “seen classes” to have an understanding about the “unseen classes” in the wild. Given a dataset in ZSL research, most existing works use a predetermined, disjoint set of seen-unseen classes to evaluate their methods. These seen (training) classes might be sub-optimal for ZSL methods to appreciate the diversity and rarity of an object domain. Inspired by strategies like active learning, it is intuitive that intelligently selecting the training classes can improve ZSL performance. In this work, we propose a framework called Diverse and Rare Class Identifier (DiRaC-I) which, given an attribute-based dataset, can intelligently yield the most suitable “seen classes” for training ZSL models. DiRaC-I has two main goals – constructing a diversified set of seed classes, and using them to initialize a visual-semantic mining algorithm for acquiring the classes capturing both diversity and rarity in the object domain adequately. These classes can then be used as “seen classes” to train ZSL models for image classification. We simulate a real-world scenario where visual samples of novel object classes in the wild are available to neither DiRaC-I nor the ZSL models during training and conducted extensive experiments on two benchmark data sets for zero-shot image classification — CUB and SUN. Our results demonstrate DiRaC-I helps ZSL models to achieve significant classification accuracy improvements – specifically, up to 8% for CUB and up to 5% for SUN dataset. Additionally, while recognizing classes exhibiting rare attributes we also observe a performance boost for ZSL models, which is up to 10% and 7% for CUB and SUN datasets, respectively. Sandipan Sarma, Arijit Sur |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Zero-Shot Learning for Computer Vision ApplicationsabstractHuman beings possess the remarkable ability to recognize unseen concepts by integrating their visual perception of known concepts with some high-level descriptions. However, the best-performing deep learning frameworks today are supervised learners that struggle to recognize concepts without training on their labeled visual samples. Zero-shot learning (ZSL) has recently emerged as a solution that mimics humans and leverages multimodal information to transfer knowledge from seen to unseen concepts. This study aims to emphasize the practicality of ZSL, unlocking its potential across four different applications in computer vision, namely -- object recognition, object detection, action recognition, and human-object interaction detection. Several task-specific challenges are identified and addressed in the presented research hypotheses. Zero-shot frameworks are proposed to attain state-of-the-art performance, elucidating some future research directions as well. Sandipan Sarma |
ACM Multimedia | 1 |
| 2022 | Resolving Semantic Confusions for Improved Zero-Shot Detection
Sandipan Sarma, Arijit Sur |
BMVC | 1 |