EDBT 2026 Demo / reviewers in the wild / expert
Sudipta Roy 0002
dblp:309/4008
· DBLP profile ↗
24ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0001-5161-9311ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NoMoColor: Unified Noise Modulation for Enhanced Diffusion-based Image Colorization (Student Abstract)abstractWe present a language-based noise modulation module for diffusion models that improves image color generation under textual guidance. Unlike standard approaches that inject noise uniformly, our method leverages semantic cues from text to selectively control the noise injection process, preserving local details and enhancing color accuracy even when descriptions are ambiguous or incomplete. Applied to language guided image colorization, this targeted modulation leads to more faithful and visually consistent results. The proposed module is lightweight, generalizable, and can be integrated into existing diffusion pipelines, offering a simple yet effective step toward more controllable text-to-image generation. Ankan Deria, Dwarikanath Mahapatra, Murari Mondal, Sudipta Roy 0002 |
AAAI | 4 |
| 2026 | How Reasoning Influences Intersectional Biases in Vision-Language Models (Student Abstract)abstractVision-Language Models (VLMs) are increasingly deployed across downstream tasks, yet their training data often encode social biases that surface in outputs. Unlike humans, who interpret images through contextual and social cues, VLMs process them through statistical associations, often leading to reasoning that diverges from human reasoning. By analyzing how a VLM reasons, we can understand how inherent biases are perpetuated and can adversely affect downstream performance. To examine this gap, we systematically analyze social biases in five open-source VLMs for an occupation prediction task, on the FairFace dataset. Across 32 occupations and three different prompting styles, we elicit both predictions and reasoning. Our findings show that the biased reasoning patterns systematically underlie intersectional disparities, highlighting the need to align VLM reasoning with human values before downstream deployment. Adit Desai, Sudipta Roy 0002, Mohna Chakraborty |
AAAI | 2 |
| 2026 | VALIANT: Prompt Instability for Active Learning in Black-Box Medical ImagingabstractThe deployment of large, black-box foundation models for medical image classification is often hindered by the high cost of acquiring large, task-specific labeled datasets for fine-tuning. While active learning (AL) presents a promising solution, many state-of-the-art AL methods are computationally expensive or require full access to internal model parameters. We present VALIANT (Visual Adaptation and Learning Integration for Active learNing Tasks), a new active learning framework designed to efficiently adapt black-box foundation models by overcoming these limitations. VALIANT introduces a lightweight Visual Prompt Decoder (VIPD), trained via unsupervised Zero-Order Optimization (ZOO), to generate task-specific visual prompts without internal model access. Our core contribution is a perturbation-based ranking strategy that leverages this VIPD to formulate a computationally efficient, gradient-aware informativeness metric. This metric, which we term prompt instability, identifies the most impactful samples for the labeling budget. VALIANT further enhances this process by incorporating anatomical information from unsupervised segmentation maps to generate more discriminative visual prompts. Extensive evaluations on multiple medical datasets demonstrate VALIANT’s superior performance and significant reduction in labeling costs compared to a range of existing active learning techniques, positioning it as a scalable and practical solution for medical image analysis. Dwarikanath Mahapatra, Behzad Bozorgtabar, Sudipta Roy 0002, Muhammad Imran Razzak, Mauricio Reyes 0001 |
AAAI | 3 |
| 2026 | Disentangled generative uncertainty-aware multi-modal diffusion segmentation of medical images
Dwarikanath Mahapatra, Sudipta Roy 0002, Mauricio Reyes 0001 |
Medical Image Anal. | 2 |
| 2025 | Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM CaptioningabstractDespite significant advances in inference-time search for vision–language models (VLMs), existing approaches remain both computationally expensive and prone to unpenalized, low-confidence generations which often lead to persistent hallucinations. We introduce \textbf{Value-guided Inference with Margin-based Reward (ViMaR)}, a two-stage inference framework that improves both efficiency and output fidelity by combining a temporal-difference value model with a margin-aware reward adjustment. In the first stage, we perform a single pass to identify the highest-value caption among diverse candidates. In the second stage, we selectively refine only those segments that were overlooked or exhibit weak visual grounding, thereby eliminating frequently rewarded evaluations. A calibrated margin-based penalty discourages low-confidence continuations while preserving descriptive richness. Extensive experiments across multiple VLM architectures demonstrate that ViMaR generates captions that are significantly more reliable, factually accurate, detailed, and explanatory, while achieving over 4$\times$ speedup compared to existing value-guided methods. Specifically, we show that ViMaR trained solely on LLaVA Mistral-7B \textit{generalizes effectively to guide decoding in stronger unseen models}. To further validate this, we adapt ViMaR to steer generation in both LLaVA-OneVision-Qwen2-7B and Qwen2.5-VL-3B, leading to consistent improvements in caption quality and demonstrating robust cross-model guidance. This cross-model generalization highlights ViMaR's flexibility and modularity, positioning it as a scalable and transferable inference-time decoding strategy. Furthermore, when ViMaR-generated captions are used for self-training, the underlying models achieve substantial gains across a broad suite of visual comprehension benchmarks, underscoring the potential of fast, accurate, and self-improving VLM pipelines.
Code: https://github.com/ankan8145/ViMaR Ankan Deria, Adinath Madhavrao Dukre, Sara Atito Ali Ahmed, Sudipta Roy 0002, Muhammad Awais 0001, Muhammad Haris Khan, Muhammad Imran Razzak |
NeurIPS | 5 |
| 2025 | Multimodal Fusion Learning with Dual Attention for Medical ImagingabstractMultimodal fusion learning has shown significant promise in classifying various diseases such as skin cancer and brain tumors. However, existing methods face three key limitations. First, they often lack generalizability to other diagnosis tasks due to their focus on a particular disease. Second, they do not fully leverage multiple health records from diverse modalities to learn robust complementary information. And finally, they typically rely on a single attention mechanism, missing the benefits of multiple attention strategies within and across various modalities. To address these issues, this paper proposes a dual robust information fusion attention mechanism (DRIFA) that leverages two attention modules - i.e., multi-branch fusion attention module and the multimodal information fusion attention module. DRIFA can be integrated with any deep neural network, forming a multimodal fusion learning framework denoted as DRIFA-Net. We show that the multi-branch fusion attention of DRIFA learns enhanced representations for each modality, such as dermoscopy, pap smear, MRI, and CT-scan, whereas multimodal information fusion attention module learns more refined multimodal shared representations - improving the network's generalization across multiple tasks and enhancing overall performance. Additionally, to estimate the uncertainty of DRIFA-Net predictions, we have employed an ensemble Monte Carlo dropout strategy. Extensive experiments on five publicly available datasets with diverse modalities demonstrate that our approach consistently outperforms state-of-the-art methods. The code is available at https://github.com/misti1203/DRIFA-Net. Joy Dhar, Nayyar Abbas Zaidi, Maryam Haghighat, Sudipta Roy 0002, Puneet Goyal, Azadeh Alavi |
WACV | 4 |
| 2025 | Self-Supervised Anomaly Segmentation via Diffusion Models with Dynamic Transformer UNetabstractA robust anomaly detection mechanism should possess the capability to effectively remediate anomalies, restoring them to a healthy state, while preserving essential healthy information. Despite the efficacy of existing generative models in learning the underlying distribution of healthy reference data, they face primary challenges when it comes to efficiently repair larger anomalies or anomalies situated near high pixel-density regions. In this paper, we introduce a self-supervised anomaly detection method based on a diffusion model that samples from multi-frequency, four-dimensional simplex noise and makes predictions using our proposed Dynamic Transformer UNet (DTUNet). This simplex-based noise function helps address primary problems to some extent and is scalable for three-dimensional and colored images. In the evolution of ViT, our developed architecture serving as the backbone for the diffusion model, is tailored to treat time and noise image patches as tokens. We incorporate long skip connections bridging the shallow and deep layers, along with smaller skip connections within these layers. Furthermore, we integrate a partial diffusion Markov process, which reduces sampling time, thus enhancing scalability. Our method surpasses existing generative-based anomaly detection methods across three diverse datasets, which include BrainMRI, Brats2021, and the MVtec dataset. It achieves an average improvement of +10.1% in Dice coefficient, +10.4% in IOU, and +9.6% in AUC. Our source code is made publicly available on Github. Komal Kumar, Snehashis Chakraborty, Dwarikanath Mahapatra, Behzad Bozorgtabar, Sudipta Roy 0002 |
WACV | 5 |
| 2025 | Weakly supervised learning based bone abnormality detection from musculoskeletal x-rays
Komal Kumar, Snehashis Chakraborty, Kalyan Tadepalli, Sudipta Roy 0002 |
Multim. Tools Appl. | 4 |
| 2025 | Multi-Label Generalized Zero Shot Chest X-Ray Classification by Combining Image-Text Information With Feature DisentanglementabstractIn fully supervised learning-based medical image classification, the robustness of a trained model is influenced by its exposure to the range of candidate disease classes. Generalized Zero Shot Learning (GZSL) aims to correctly predict seen and novel unseen classes. Current GZSL approaches have focused mostly on the single-label case. However, it is common for chest X-rays to be labelled with multiple disease classes. We propose a novel multi-modal multi-label GZSL approach that leverages feature disentanglement andmulti-modal information to synthesize features of unseen classes. Disease labels are processed through a pre-trained BioBert model to obtain text embeddings that are used to create a dictionary encoding similarity among different labels. We then use disentangled features and graph aggregation to learn a second dictionary of inter-label similarities. A subsequent clustering step helps to identify representative vectors for each class. The multi-modal multi-label dictionaries and the class representative vectors are used to guide the feature synthesis step, which is the most important component of our pipeline, for generating realistic multi-label disease samples of seen and unseen classes. Our method is benchmarked against multiple competing methods and we outperform all of them based on experiments conducted on the publicly available NIH and CheXpert chest X-ray datasets. Dwarikanath Mahapatra, Antonio Jimeno-Yepes, Behzad Bozorgtabar, Sudipta Roy 0002, ZongYuan Ge, Mauricio Reyes 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Corrections to "Multi-Label Generalized Zero Shot Chest X-Ray Classification By Combining Image-Text Information With Feature Disentanglement"abstractPresents corrections to the paper, (Corrections to "Multi-Label Generalized Zero Shot Chest X-Ray Classification By Combining Image-Text Information With Feature Disentanglement"). Dwarikanath Mahapatra, Antonio Jimeno-Yepes, Behzad Bozorgtabar, Sudipta Roy 0002, ZongYuan Ge, Mauricio Reyes 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | SEANet: Rethinking Skip-Connections Design in Encoder-Decoder Networks via Synergistic Spatial-Spectral Fusion for LDCT Denoising
Vandan Gorade, Dwarikanath Mahapatra, Sudipta Roy 0002 |
ICPR (12) | 4 |
| 2024 | Confidence-Guided Semi-supervised Learning for Generalized Lesion Localization in X-Ray Images
Vandan Gorade, Komal Kumar, Snehashis Chakraborty, Dwarikanath Mahapatra, Sudipta Roy 0002 |
MICCAI (1) | 6 |
| 2024 | ALFREDO: Active Learning with FeatuRe disEntangelement and DOmain adaptation for medical image classification
Dwarikanath Mahapatra, Ruwan B. Tennakoon, Yasmeen M. George, Sudipta Roy 0002, Behzad Bozorgtabar, ZongYuan Ge, Mauricio Reyes 0001 |
Medical Image Anal. | 4 |
| 2024 | Unleashing the power of explainable AI: sepsis sentinel's clinical assistant for early sepsis identification
Snehashis Chakraborty, Komal Kumar, Kalyan Tadepalli, Balakrishna Pailla, Sudipta Roy 0002 |
Multim. Tools Appl. | 5 |
| 2024 | Exploring deep learning for carotid artery plaque segmentation: atherosclerosis to cardiovascular risk biomarkers
Pankaj Kumar Jain, Kalyan Tadepalli, Sudipta Roy 0002 |
Multim. Tools Appl. | 3 |
| 2024 | A weighted ensemble transfer learning approach for melanoma classification from skin lesion images
Himanshi Meswal, Deepika Kumar, Sudipta Roy 0002 |
Multim. Tools Appl. | 4 |
| 2024 | Forward attention-based deep network for classification of breast histopathology image
Sudipta Roy 0002, Pankaj Kumar Jain, Kalyan Tadepalli, Balakrishna Pailla Reddy |
Multim. Tools Appl. | 1 |
| 2023 | Class Specific Feature Disentanglement and Text Embeddings for Multi-label Generalized Zero Shot CXR Classification
Dwarikanath Mahapatra, Antonio Jimeno-Yepes, Shiba Kuanar, Sudipta Roy 0002, Behzad Bozorgtabar, Mauricio Reyes 0001, ZongYuan Ge |
MICCAI (2) | 4 |
| 2023 | Number plate recognition from enhanced super-resolution using generative adversarial network
Anwesh Kabiraj, Debojyoti Pal, Debayan Ganguly, Kingshuk Chatterjee, Sudipta Roy 0002 |
Multim. Tools Appl. | 5 |
| 2021 | A deep survey on supervised learning based human detection and activity classification methods
Muhammad Attique Khan, Mamta Mittal, Lalit Mohan Goyal, Sudipta Roy 0002 |
Multim. Tools Appl. | 4 |
| 2020 | Data hiding in virtual bit-plane using efficient Lucas number sequences
Biswajita Datta, Koushik Dutta, Sudipta Roy 0002 |
Multim. Tools Appl. | 3 |
| 2019 | Multi-bit robust image steganography based on modular arithmetic
Biswajita Datta, Sudipta Roy 0002, Subhranil Roy, Samir Kumar Bandyopadhyay |
Multim. Tools Appl. | 2 |
| 2019 | Blood vessel segmentation of retinal image using Clifford matched filter and Clifford convolution
Somasis Roy, Anirban Mitra, Sudipta Roy 0002, Sanjit Kumar Setua |
Multim. Tools Appl. | 3 |
| 2017 | An improved brain MR image binarization method as a preprocessing for abnormality detection and features extraction
Sudipta Roy 0002, Debnath Bhattacharyya, Samir Kumar Bandyopadhyay, Tai-Hoon Kim |
Frontiers Comput. Sci. | 1 |