EDBT 2026 Demo / reviewers in the wild / expert
Dwarikanath Mahapatra
dblp:50/6718
· DBLP profile ↗
57ranked-venue papers
28as first author
31since 2021 · last 2026
0000-0001-9749-7858ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 15 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 31 · 18 first-author · 19 since 2021Artificial intelligence and machine learning · 15 · 6 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NoMoColor: Unified Noise Modulation for Enhanced Diffusion-based Image Colorization (Student Abstract)abstractWe present a language-based noise modulation module for diffusion models that improves image color generation under textual guidance. Unlike standard approaches that inject noise uniformly, our method leverages semantic cues from text to selectively control the noise injection process, preserving local details and enhancing color accuracy even when descriptions are ambiguous or incomplete. Applied to language guided image colorization, this targeted modulation leads to more faithful and visually consistent results. The proposed module is lightweight, generalizable, and can be integrated into existing diffusion pipelines, offering a simple yet effective step toward more controllable text-to-image generation. Ankan Deria, Dwarikanath Mahapatra, Murari Mondal, Sudipta Roy 0002 |
AAAI | 2 |
| 2026 | VALIANT: Prompt Instability for Active Learning in Black-Box Medical ImagingabstractThe deployment of large, black-box foundation models for medical image classification is often hindered by the high cost of acquiring large, task-specific labeled datasets for fine-tuning. While active learning (AL) presents a promising solution, many state-of-the-art AL methods are computationally expensive or require full access to internal model parameters. We present VALIANT (Visual Adaptation and Learning Integration for Active learNing Tasks), a new active learning framework designed to efficiently adapt black-box foundation models by overcoming these limitations. VALIANT introduces a lightweight Visual Prompt Decoder (VIPD), trained via unsupervised Zero-Order Optimization (ZOO), to generate task-specific visual prompts without internal model access. Our core contribution is a perturbation-based ranking strategy that leverages this VIPD to formulate a computationally efficient, gradient-aware informativeness metric. This metric, which we term prompt instability, identifies the most impactful samples for the labeling budget. VALIANT further enhances this process by incorporating anatomical information from unsupervised segmentation maps to generate more discriminative visual prompts. Extensive evaluations on multiple medical datasets demonstrate VALIANT’s superior performance and significant reduction in labeling costs compared to a range of existing active learning techniques, positioning it as a scalable and practical solution for medical image analysis. Dwarikanath Mahapatra, Behzad Bozorgtabar, Sudipta Roy 0002, Muhammad Imran Razzak, Mauricio Reyes 0001 |
AAAI | 1 |
| 2026 | Disentangled generative uncertainty-aware multi-modal diffusion segmentation of medical images
Dwarikanath Mahapatra, Sudipta Roy 0002, Mauricio Reyes 0001 |
Medical Image Anal. | 1 |
| 2025 | DiMPLE - Disentangled Multi-Modal Prompt Learning: Enhancing Out-of-Distribution Alignment with Invariant and Spurious Feature SeparationabstractWe introduce DiMPLe (Disentangled Multi-Modal Prompt Learning), a novel approach to disentangle invariant and spurious features across vision and language modalities in multi-modal learning. Spurious correlations in visual data often hinder out-of-distribution (OOD) performance. Unlike prior methods focusing solely on image features, DiMPLe disentangles features within and across modalities while maintaining consistent alignment, enabling better generalization to novel classes and robustness to distribution shifts. Our method combines three key objectives: (1) mutual information minimization between invariant and spurious features, (2) spurious feature regularization, and (3) contrastive learning on invariant features. Extensive experiments demonstrate DiMPLe demonstrates superior performance compared to CoOp-OOD, when averaged across 11 diverse datasets, and achieves absolute gains of 15.27 in base class accuracy and 44.31 in novel class accuracy. Umaima Rahman, Mohammad Yaqub, Dwarikanath Mahapatra |
ICCV | 3 |
| 2025 | Self-Supervised Anomaly Segmentation via Diffusion Models with Dynamic Transformer UNetabstractA robust anomaly detection mechanism should possess the capability to effectively remediate anomalies, restoring them to a healthy state, while preserving essential healthy information. Despite the efficacy of existing generative models in learning the underlying distribution of healthy reference data, they face primary challenges when it comes to efficiently repair larger anomalies or anomalies situated near high pixel-density regions. In this paper, we introduce a self-supervised anomaly detection method based on a diffusion model that samples from multi-frequency, four-dimensional simplex noise and makes predictions using our proposed Dynamic Transformer UNet (DTUNet). This simplex-based noise function helps address primary problems to some extent and is scalable for three-dimensional and colored images. In the evolution of ViT, our developed architecture serving as the backbone for the diffusion model, is tailored to treat time and noise image patches as tokens. We incorporate long skip connections bridging the shallow and deep layers, along with smaller skip connections within these layers. Furthermore, we integrate a partial diffusion Markov process, which reduces sampling time, thus enhancing scalability. Our method surpasses existing generative-based anomaly detection methods across three diverse datasets, which include BrainMRI, Brats2021, and the MVtec dataset. It achieves an average improvement of +10.1% in Dice coefficient, +10.4% in IOU, and +9.6% in AUC. Our source code is made publicly available on Github. Komal Kumar, Snehashis Chakraborty, Dwarikanath Mahapatra, Behzad Bozorgtabar, Sudipta Roy 0002 |
WACV | 3 |
| 2025 | Multi-Label Generalized Zero Shot Chest X-Ray Classification by Combining Image-Text Information With Feature DisentanglementabstractIn fully supervised learning-based medical image classification, the robustness of a trained model is influenced by its exposure to the range of candidate disease classes. Generalized Zero Shot Learning (GZSL) aims to correctly predict seen and novel unseen classes. Current GZSL approaches have focused mostly on the single-label case. However, it is common for chest X-rays to be labelled with multiple disease classes. We propose a novel multi-modal multi-label GZSL approach that leverages feature disentanglement andmulti-modal information to synthesize features of unseen classes. Disease labels are processed through a pre-trained BioBert model to obtain text embeddings that are used to create a dictionary encoding similarity among different labels. We then use disentangled features and graph aggregation to learn a second dictionary of inter-label similarities. A subsequent clustering step helps to identify representative vectors for each class. The multi-modal multi-label dictionaries and the class representative vectors are used to guide the feature synthesis step, which is the most important component of our pipeline, for generating realistic multi-label disease samples of seen and unseen classes. Our method is benchmarked against multiple competing methods and we outperform all of them based on experiments conducted on the publicly available NIH and CheXpert chest X-ray datasets. Dwarikanath Mahapatra, Antonio Jimeno-Yepes, Behzad Bozorgtabar, Sudipta Roy 0002, ZongYuan Ge, Mauricio Reyes 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Corrections to "Multi-Label Generalized Zero Shot Chest X-Ray Classification By Combining Image-Text Information With Feature Disentanglement"abstractPresents corrections to the paper, (Corrections to "Multi-Label Generalized Zero Shot Chest X-Ray Classification By Combining Image-Text Information With Feature Disentanglement"). Dwarikanath Mahapatra, Antonio Jimeno-Yepes, Behzad Bozorgtabar, Sudipta Roy 0002, ZongYuan Ge, Mauricio Reyes 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Prompt-Driven Latent Domain Generalization for Medical Image ClassificationabstractDeep learning models for medical image analysis easily suffer from distribution shifts caused by dataset artifact bias, camera variations, differences in the imaging station, etc., leading to unreliable diagnoses in real-world clinical settings. Domain generalization (DG) methods, which aim to train models on multiple domains to perform well on unseen domains, offer a promising direction to solve the problem. However, existing DG methods assume domain labels of each image are available and accurate, which is typically feasible for only a limited number of medical datasets. To address these challenges, we propose a unified DG framework for medical image classification without relying on domain labels, called Prompt-driven Latent Domain Generalization (PLDG). PLDG consists of unsupervised domain discovery and prompt learning. This framework first discovers pseudo domain labels by clustering the bias-associated style features, then leverages collaborative domain prompts to guide a Vision Transformer to learn knowledge from discovered diverse domains. To facilitate cross-domain knowledge learning between different prompts, we introduce a domain prompt generator that enables knowledge sharing between domain prompts and a shared prompt. A domain mixup strategy is additionally employed for more flexible decision margins and mitigates the risk of incorrect domain assignments. Extensive experiments on three medical image classification tasks and one debiasing task demonstrate that our method can achieve comparable or even superior performance than conventional DG algorithms without relying on domain labels. Our code is publicly available at https://github.com/SiyuanYan1/PLDG/tree/main. Siyuan Yan, Chi Liu 0002, Lie Ju, Dwarikanath Mahapatra, Brigid Betz-Stablein, Victoria Mar, Monika Janda, H. Peter Soyer, ZongYuan Ge |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Combining Graph Transformers Based Multi-Label Active Learning and Informative Data Augmentation for Chest Xray ClassificationabstractInformative sample selection in active learning (AL) helps a machine learning system attain optimum performance with minimum labeled samples, thus improving human-in-the-loop computer-aided diagnosis systems with limited labeled data. Data augmentation is highly effective for enlarging datasets with less labeled data. Combining informative sample selection and data augmentation should leverage their respective advantages and improve performance of AL systems. We propose a novel approach to combine informative sample selection and data augmentation for multi-label active learning. Conventional informative sample selection approaches have mostly focused on the single-label case which do not perform optimally in the multi-label setting. We improve upon state-of-the-art multi-label active learning techniques by representing disease labels as graph nodes, use graph attention transformers (GAT) to learn more effective inter-label relationships and identify most informative samples. We generate transformations of these informative samples which are also informative. Experiments on public chest xray datasets show improved results over state-of-the-art multi-label AL techniques in terms of classification performance, learning rates, and robustness. We also perform qualitative analysis to determine the realism of generated images. Dwarikanath Mahapatra, Behzad Bozorgtabar, ZongYuan Ge, Mauricio Reyes 0001, Jean-Philippe Thiran |
AAAI | 1 |
| 2024 | SEANet: Rethinking Skip-Connections Design in Encoder-Decoder Networks via Synergistic Spatial-Spectral Fusion for LDCT Denoising
Vandan Gorade, Dwarikanath Mahapatra, Sudipta Roy 0002 |
ICPR (12) | 3 |
| 2024 | Confidence-Guided Semi-supervised Learning for Generalized Lesion Localization in X-Ray Images
Vandan Gorade, Komal Kumar, Snehashis Chakraborty, Dwarikanath Mahapatra, Sudipta Roy 0002 |
MICCAI (1) | 5 |
| 2024 | GANDALF: Graph-based transformer and Data Augmentation Active Learning Framework with interpretable features for multi-label chest Xray classification
Dwarikanath Mahapatra, Behzad Bozorgtabar, ZongYuan Ge, Mauricio Reyes 0001 |
Medical Image Anal. | 1 |
| 2024 | ALFREDO: Active Learning with FeatuRe disEntangelement and DOmain adaptation for medical image classification
Dwarikanath Mahapatra, Ruwan B. Tennakoon, Yasmeen M. George, Sudipta Roy 0002, Behzad Bozorgtabar, ZongYuan Ge, Mauricio Reyes 0001 |
Medical Image Anal. | 1 |
| 2023 | Attention-Conditioned Augmentations for Self-Supervised Anomaly Detection and LocalizationabstractSelf-supervised anomaly detection and localization are critical to real-world scenarios in which collecting anomalous samples and pixel-wise labeling is tedious or infeasible, even worse when a wide variety of unseen anomalies could surface at test time. Our approach involves a pretext task in the context of masked image modeling, where the goal is to impose agreement between cluster assignments obtained from the representation of an image view containing saliency-aware masked patches and the uncorrupted image view. We harness the self-attention map extracted from the transformer to mask non-salient image patches without destroying the crucial structure associated with the foreground object. Subsequently, the pre-trained model is fine-tuned to detect and localize simulated anomalies generated under the guidance of the transformer's self-attention map. We conducted extensive validation and ablations on the benchmark of industrial images and achieved superior performance against competing methods. We also show the adaptability of our method to the medical images of the chest X-rays benchmark. Behzad Bozorgtabar, Dwarikanath Mahapatra |
AAAI | 2 |
| 2023 | Towards Trustable Skin Cancer Diagnosis via Rewriting Model's DecisionabstractDeep neural networks have demonstrated promising performance on image recognition tasks. However, they may heavily rely on confounding factors, using irrelevant artifacts or bias within the dataset as the cue to improve performance. When a model performs decision-making based on these spurious correlations, it can become untrustable and lead to catastrophic outcomes when deployed in the realworld scene. In this paper, we explore and try to solve this problem in the context of skin cancer diagnosis. We introduce a human-in-the-loop framework in the model training process such that users can observe and correct the model's decision logic when confounding behaviors happen. Specifically, our method can automatically discover confounding factors by analyzing the co-occurrence behavior of the samples. It is capable of learning confounding concepts using easily obtained concept exemplars. By mapping the black-box model's feature representation onto an explainable concept space, human users can interpret the concept and intervene via first order-logic instruction. We systematically evaluate our method on our newly crafted, well-controlled skin lesion dataset and several public skin lesion datasets. Experiments show that our method can effectively detect and remove confounding factors from datasets without any prior knowledge about the category distribution and does not require fully annotated concept labels. We also show that our method enables the model to focus on clinical-related concepts, improving the model's performance and trustworthiness during model inference. Siyuan Yan, Dwarikanath Mahapatra, Shekhar Chandra, Monika Janda, H. Peter Soyer, ZongYuan Ge |
CVPR | 4 |
| 2023 | AMAE: Adaptation of Pre-trained Masked Autoencoder for Dual-Distribution Anomaly Detection in Chest X-Rays
Behzad Bozorgtabar, Dwarikanath Mahapatra, Jean-Philippe Thiran |
MICCAI (1) | 2 |
| 2023 | Class Specific Feature Disentanglement and Text Embeddings for Multi-label Generalized Zero Shot CXR Classification
Dwarikanath Mahapatra, Antonio Jimeno-Yepes, Shiba Kuanar, Sudipta Roy 0002, Behzad Bozorgtabar, Mauricio Reyes 0001, ZongYuan Ge |
MICCAI (2) | 1 |
| 2023 | EPVT: Environment-Aware Prompt Vision Transformer for Domain Generalization in Skin Lesion Recognition
Siyuan Yan, Chi Liu 0002, Lie Ju, Dwarikanath Mahapatra, Victoria Mar, Monika Janda, H. Peter Soyer, ZongYuan Ge |
MICCAI (7) | 5 |
| 2023 | Probabilistic Integration of Object Level Annotations in Chest X-ray ClassificationabstractMedical image datasets and their annotations are not growing as fast as their equivalents in the general domain. This makes translation from the newest, more data-intensive methods that have made a large impact on the vision field increasingly more difficult and less efficient. In this paper, we propose a new probabilistic latent variable model for disease classification in chest X-ray images. Specifically we consider chest X-ray datasets that contain global disease labels, and for a smaller subset contain object level expert annotations in the form of eye gaze patterns and disease bounding boxes. We propose a two-stage optimization algorithm which is able to handle these different label granularities through a single training pipeline in a two-stage manner. In our pipeline global dataset features are learned in the lower level layers of the model. The specific details and nuances in the fine-grained expert object-level annotations are learned in the final layers of the model using a knowledge distillation method inspired by conditional variational inference. Subsequently, model weights are frozen to guide this learning process and prevent overfitting on the smaller richly annotated data subsets. The proposed method yields consistent classification improvement across different back-bones on the common benchmark datasets Chest X-ray14 and MIMIC-CXR. This shows how two-stage learning of labels from coarse to fine-grained, in particular with object level annotations, is an effective method for more optimal annotation usage. Tom van Sonsbeek, Xiantong Zhen, Dwarikanath Mahapatra, Marcel Worring |
WACV | 3 |
| 2023 | Graph Node Based Interpretability Guided Sample Selection for Active LearningabstractWhile supervised learning techniques have demonstrated state-of-the-art performance in many medical image analysis tasks, the role of sample selection is important. Selecting the most informative samples contributes to the system attaining optimum performance with minimum labeled samples, which translates to fewer expert interventions and cost. Active Learning (AL) methods for informative sample selection are effective in boosting performance of computer aided diagnosis systems when limited labels are available. Conventional approaches to AL have mostly focused on the single label setting where a sample has only one disease label from the set of possible labels. These approaches do not perform optimally in the multi-label setting where a sample can have multiple disease labels (e.g. in chest X-ray images). In this paper we propose a novel sample selection approach based on graph analysis to identify informative samples in a multi-label setting. For every analyzed sample, each class label is denoted as a separate node of a graph. Building on findings from interpretability of deep learning models, edge interactions in this graph characterize similarity between corresponding interpretability saliency map model encodings. We explore different types of graph aggregation to identify informative samples for active learning. We apply our method to public chest X-ray and medical image datasets, and report improved results over state-of-the-art AL techniques in terms of model performance, learning rates, and robustness. Dwarikanath Mahapatra, Alexander Pollinger, Mauricio Reyes 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2022 | Anomaly Detection and Localization Using Attention-Guided Synthetic Anomaly and Test-Time Adaptation
Behzad Bozorgtabar, Dwarikanath Mahapatra, Jean-Philippe Thiran |
BMVC | 2 |
| 2022 | LifeLonger: A Benchmark for Continual Disease Classification
Mohammad Mahdi Derakhshani, Ivona Najdenkoska, Tom van Sonsbeek, Xiantong Zhen, Dwarikanath Mahapatra, Marcel Worring, Cees Snoek |
MICCAI (2) | 5 |
| 2022 | Interpretability-Guided Inductive Bias For Deep Learning Based Medical Image
Dwarikanath Mahapatra, Alexander Pollinger, Mauricio Reyes 0001 |
Medical Image Anal. | 1 |
| 2022 | Gated fusion network for SAO filter and inter frame prediction in Versatile Video Coding
Shiba Kuanar, Vassilis Athitsos, Dwarikanath Mahapatra, Kamisetty Ramamohan Rao |
Signal Process. Image Commun. | 3 |
| 2022 | Improving Medical Images Classification With Label Noise Using Dual-Uncertainty EstimationabstractDeep neural networks are known to be data-driven and label noise can have a marked impact on model performance. Recent studies have shown great robustness to classic image recognition even under a high noisy rate. In medical applications, learning from datasets with label noise is more challenging since medical imaging datasets tend to have instance-dependent noise (IDN) and suffer from high observer variability. In this paper, we systematically discuss the two common types of label noise in medical images - disagreement label noise from inconsistency expert opinions and single-target label noise from biased aggregation of individual annotations. We then propose an uncertainty estimation-based framework to handle these two label noise amid the medical image classification task. We design a dual-uncertainty estimation approach to measure the disagreement label noise and single-target label noise via improved Direct Uncertainty Prediction and Monte-Carlo-Dropout. A boosting-based curriculum training procedure is later introduced for robust learning. We demonstrate the effectiveness of our method by conducting extensive experiments on three different diseases with synthesized and real-world label noise: skin lesions, prostate cancer, and retinal diseases. We also release a large re-engineered database that consists of annotations from more than ten ophthalmologists with an unbiased golden standard dataset for evaluation and benchmarking. The dataset is available at https://mmai.group/peoples/julie/. Lie Ju, Xin Wang 0094, Lin Wang 0027, Dwarikanath Mahapatra, Quan Zhou 0004, Tongliang Liu, ZongYuan Ge |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Self-Supervised Generalized Zero Shot Learning for Medical Image Classification Using Novel Interpretable Saliency MapsabstractIn many real world medical image classification settings, access to samples of all disease classes is not feasible, affecting the robustness of a system expected to have high performance in analyzing novel test data. This is a case of generalized zero shot learning (GZSL) aiming to recognize seen and unseen classes. We propose a GZSL method that uses self supervised learning (SSL) for: 1) selecting representative vectors of disease classes; and 2) synthesizing features of unseen classes. We also propose a novel approach to generate GradCAM saliency maps that highlight diseased regions with greater accuracy. We exploit information from the novel saliency maps to improve the clustering process by: 1) Enforcing the saliency maps of different classes to be different; and 2) Ensuring that clusters in the space of image and saliency features should yield class centroids having similar semantic information. This ensures the anchor vectors are representative of each class. Different from previous approaches, our proposed approach does not require class attribute vectors which are essential part of GZSL methods for natural images but are not available for medical images. Using a simple architecture the proposed method outperforms state of the art SSL based GZSL performance for natural images as well as multiple types of medical images. We also conduct many ablation studies to investigate the influence of different loss terms in our method. Dwarikanath Mahapatra, ZongYuan Ge, Mauricio Reyes 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2022 | Multi-path dilated convolution network for haze and glow removal in nighttime images
Shiba Kuanar, Dwarikanath Mahapatra, Monalisa Bilas, Kamisetty Ramamohan Rao |
Vis. Comput. | 2 |
| 2021 | Relational Subsets Knowledge Distillation for Long-Tailed Retinal Diseases Recognition
Lie Ju, Xin Wang 0094, Lin Wang 0027, Tongliang Liu, Tom Drummond, Dwarikanath Mahapatra, ZongYuan Ge |
MICCAI (8) | 7 |
| 2021 | Synergic Adversarial Label Learning for Grading Retinal Diseases via Knowledge Distillation and Multi-Task LearningabstractThe need for comprehensive and automated screening methods for retinal image classification has long been recognized. Well-qualified doctors annotated images are very expensive and only a limited amount of data is available for various retinal diseases such as diabetic retinopathy (DR) and age-related macular degeneration (AMD). Some studies show that some retinal diseases such as DR and AMD share some common features like haemorrhages and exudation but most classification algorithms only train those disease models independently when the only single label for one image is available. Inspired by multi-task learning where additional monitoring signals from various sources is beneficial to train a robust model. We propose a method called synergic adversarial label learning (SALL) which leverages relevant retinal disease labels in both semantic and feature space as additional signals and train the model in a collaborative manner using knowledge distillation. Our experiments on DR and AMD fundus image classification task demonstrate that the proposed method can significantly improve the accuracy of the model for grading diseases by 5.91% and 3.69% respectively. In addition, we conduct additional experiments to show the effectiveness of SALL from the aspects of reliability and interpretability in the context of medical imaging application. Lie Ju, Xin Wang 0094, Huimin Lu 0001, Dwarikanath Mahapatra, C. Paul Bonnington, ZongYuan Ge |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Interpretability-Driven Sample Selection Using Self Supervised Learning for Disease Classification and SegmentationabstractIn supervised learning for medical image analysis, sample selection methodologies are fundamental to attain optimum system performance promptly and with minimal expert interactions (e.g. label querying in an active learning setup). In this article we propose a novel sample selection methodology based on deep features leveraging information contained in interpretability saliency maps. In the absence of ground truth labels for informative samples, we use a novel self supervised learning based approach for training a classifier that learns to identify the most informative sample in a given batch of images. We demonstrate the benefits of the proposed approach, termed Interpretability-Driven Sample Selection (IDEAL), in an active learning setup aimed at lung disease classification and histopathology image segmentation. We analyze three different approaches to determine sample informativeness from interpretability saliency maps: (i) an observational model stemming from findings on previous uncertainty-based sample selection approaches, (ii) a radiomics-based model, and (iii) a novel data-driven self-supervised approach. We compare IDEAL to other baselines using the publicly available NIH chest X-ray dataset for lung disease classification, and a public histopathology segmentation dataset (GLaS), demonstrating the potential of using interpretability information for sample selection in active learning systems. Results show our proposed self supervised approach outperforms other approaches in selecting informative samples leading to state of the art performance with fewer samples. Dwarikanath Mahapatra, Alexander Pollinger, Ling Shao 0001, Mauricio Reyes 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2021 | MoNuSAC2020: A Multi-Organ Nuclei Segmentation and Classification ChallengeabstractDetecting various types of cells in and around the tumor matrix holds a special significance in characterizing the tumor micro-environment for cancer prognostication and research. Automating the tasks of detecting, segmenting, and classifying nuclei can free up the pathologists' time for higher value tasks and reduce errors due to fatigue and subjectivity. To encourage the computer vision research community to develop and test algorithms for these tasks, we prepared a large and diverse dataset of nucleus boundary annotations and class labels. The dataset has over 46,000 nuclei from 37 hospitals, 71 patients, four organs, and four nucleus types. We also organized a challenge around this dataset as a satellite event at the International Symposium on Biomedical Imaging (ISBI) in April 2020. The challenge saw a wide participation from across the world, and the top methods were able to match inter-human concordance for the challenge metric. In this paper, we summarize the dataset and the key findings of the challenge, including the commonalities and differences between the methods developed by various participants. We have released the MoNuSAC2020 dataset to the public. Ruchika Verma, Neeraj Kumar 0002, Abhijeet Patil, Nikhil Cherian Kurian, Swapnil Rane, Simon Graham, Quoc Dang Vu, Mieke Zwager, Shan E Ahmed Raza, Nasir M. Rajpoot, Xiyi Wu, Huai Chen, Lisheng Wang, Hyun Jung, G. Thomas Brown, Shuolin Liu, Seyed Alireza Fatemi Jahromi, Aliasghar Khani, Ehsan Montahaei, Mahdieh Soleymani Baghshah, Hamid Behroozi, Pavel Semkin, Alexandr Rassadin, Prasad Dutande, Romil Lodaya, Ujjwal Baid, Bhakti Baheti, Sanjay N. Talbar, Amirreza Mahbod, Rupert Ecker, Isabella Ellinger, Bin Dong 0006, Zhengyu Xu, Yuehan Yao, Ming Feng, Kele Xu, Hasib Zunair, A. Ben Hamza, Steven M. Smiley, Tang-Kai Yin, Qi-Rui Fang, Shikhar Srivastava 0001, Dwarikanath Mahapatra, Lubomira Trnavska, Hanyun Zhang, Priya Lakshmi Narayanan, Justin Law, Yinyin Yuan, Abhiroop Tejomay, Aditya Mitkari, Dinesh Koka, Vikas Ramachandra, Lata Kini, Amit Sethi |
IEEE Trans. Medical Imaging | 47 |
| 2020 | Pathological Retinal Region Segmentation From OCT Images Using Geometric Relation Based AugmentationabstractMedical image segmentation is important for computer aided diagnosis. Pixelwise manual annotations of large datasets require high expertise and is time consuming. Conventional data augmentations have limited benefit by not fully representing the underlying distribution of the training set, thus affecting model robustness when tested on images captured from different sources. Prior work leverages synthetic images for data augmentation ignoring the interleaved geometric relationship between different anatomical labels. We propose improvements over previous GAN-based medical image synthesis methods by jointly encoding the intrinsic relationship of geometry and shape. Latent space variable sampling results in diverse generated images from a base image and improves robustness. Augmented datasets using our method for automatic segmentation of retinal optical coherence tomography (OCT) images outperform existing methods on the public RETOUCH dataset having images captured from different acquisition procedures. Ablation studies and visual analysis also demonstrate benefits of integrating geometry and diversity. Dwarikanath Mahapatra, Behzad Bozorgtabar, Ling Shao 0001 |
CVPR | 1 |
| 2020 | SALAD: Self-supervised Aggregation Learning for Anomaly Detection on X-Rays
Behzad Bozorgtabar, Dwarikanath Mahapatra, Guillaume Vray, Jean-Philippe Thiran |
MICCAI (1) | 2 |
| 2020 | Structure Preserving Stain Normalization of Histopathology Images Using Self Supervised Semantic Guidance
Dwarikanath Mahapatra, Behzad Bozorgtabar, Jean-Philippe Thiran, Ling Shao 0001 |
MICCAI (5) | 1 |
| 2020 | Improving multi-label chest X-ray disease diagnosis by exploiting disease and health labels dependencies
ZongYuan Ge, Dwarikanath Mahapatra, Xiaojun Chang, Zetao Chen, Lianhua Chi, Huimin Lu 0001 |
Multim. Tools Appl. | 2 |
| 2020 | ExprADA: Adversarial domain adaptation for facial expression analysis
Behzad Bozorgtabar, Dwarikanath Mahapatra, Jean-Philippe Thiran |
Pattern Recognit. | 2 |
| 2020 | Training data independent image registration using generative adversarial networks and domain adaptation
Dwarikanath Mahapatra, ZongYuan Ge |
Pattern Recognit. | 1 |
| 2019 | SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-MotionabstractDespite well-established baselines, learning of scene depth and ego-motion from monocular video remains an ongoing challenge, specifically when handling scaling ambiguity issues and depth inconsistencies in image sequences. Much prior work uses either a supervised mode of learning or stereo images. The former is limited by the amount of labeled data, as it requires expensive sensors, while the latter is not always readily available as monocular sequences. In this work, we demonstrate the benefit of using geometric information from synthetic images, coupled with scene depth information, to recover the scale in depth and ego-motion estimation from monocular videos. We developed our framework using synthetic image-depth pairs and unlabeled real monocular images. We had three training objectives: first, to use deep feature alignment to reduce the domain gap between synthetic and monocular images to yield more accurate depth estimation when presented with only real monocular images at test time. Second, we learn scene specific representation by exploiting self-supervision coming from multi-view synthetic images without the need for depth labels. Third, our method uses single-view depth and pose networks, which are capable of jointly training and supervising one another mutually, yielding consistent depth and ego-motion estimates. Extensive experiments demonstrate that our depth and ego-motion models surpass the state-of-the-art, unsupervised methods and compare favorably to early supervised deep models for geometric understanding. We validate the effectiveness of our training objectives against standard benchmarks thorough an ablation study. Behzad Bozorgtabar, Mohammad Saeed Rad, Dwarikanath Mahapatra, Jean-Philippe Thiran |
ICCV | 3 |
| 2019 | Low Dose Abdominal CT Image Reconstruction: An Unsupervised Learning Based ApproachabstractIn medical practice, the X-ray Computed tomography-based scans expose a high radiation dose and lead to the risk of prostate or abdomen cancers. On the other hand, the low-dose CT scan can reduce radiation exposure to the patient. But the reduced radiation dose degrades image quality for human perception, and adversely affects the radiologist's diagnosis and prognosis. In this paper, we introduce a GAN based auto-encoder network to de-noise the CT images. Our network first maps CT images to low dimensional manifolds and then restore the images from its corresponding manifold representations. Our reconstruction algorithm separately calculates perceptual similarity, learns the latent feature maps, and achieves more accurate and visually pleasing reconstructions. We also showed the effectiveness of our model on a number of patient abdomen CT images, and compare our results with existing deep learning and iterative reconstruction methods. Experimental results demonstrate that our model outperforms other state-of-the-art methods in terms of PSNR, SSIM, and statistical properties of the image regions. https://github.com/ShibaPrasad/CT-Image-Reconstruction. Shiba Kuanar, Vassilis Athitsos, Dwarikanath Mahapatra, Kamisetty Ramamohan Rao, Zahid Akhtar, Dipankar Dasgupta |
ICIP | 3 |
| 2019 | Adversarial Pulmonary Pathology Translation for Pairwise Chest X-Ray Data Augmentation
Yunyan Xing, ZongYuan Ge, Dwarikanath Mahapatra, Jarrel Seah, Meng Law, Tom Drummond |
MICCAI (6) | 4 |
| 2019 | Informative sample generation using class aware generative adversarial networks for classification of chest Xrays
Behzad Bozorgtabar, Dwarikanath Mahapatra, Hendrik von Tengg-Kobligk, Alexander Pollinger, Lukas Ebner, Jean-Philippe Thiran, Mauricio Reyes 0001 |
Comput. Vis. Image Underst. | 2 |
| 2018 | Efficient Active Learning for Image Classification and Segmentation Using a Sample Selection and Conditional Generative Adversarial Network
Dwarikanath Mahapatra, Behzad Bozorgtabar, Jean-Philippe Thiran, Mauricio Reyes 0001 |
MICCAI (2) | 1 |
| 2017 | Image Super Resolution Using Generative Adversarial Networks and Local Saliency Maps for Retinal Image Analysis
Dwarikanath Mahapatra, Behzad Bozorgtabar, Sajini Hewavitharanage, Rahil Garnavi |
MICCAI (3) | 1 |
| 2017 | Semi-supervised Segmentation of Optic Cup in Retinal Fundus Images Using Variational Autoencoder
Suman Sedai, Dwarikanath Mahapatra, Sajini Hewavitharanage, Stefan Maetschke, Rahil Garnavi |
MICCAI (2) | 2 |
| 2017 | Semi-supervised learning and graph cuts for consensus based medical image segmentation
Dwarikanath Mahapatra |
Pattern Recognit. | 1 |
| 2016 | Combining multiple expert annotations using semi-supervised learning and graph cuts for medical image segmentation
Dwarikanath Mahapatra |
Comput. Vis. Image Underst. | 1 |
| 2016 | Image Registration Based on Autocorrelation of Local StructureabstractRegistration of images in the presence of intra-image signal fluctuations is a challenging task. The definition of an appropriate objective function measuring the similarity between the images is crucial for accurate registration. This paper introduces an objective function that embeds local phase features derived from the monogenic signal in the modality independent neighborhood descriptor (MIND). The image similarity relies on the autocorrelation of local structure (ALOST) which has two important properties: 1) low sensitivity to space-variant intensity distortions (e.g., differences in contrast enhancement in MRI); 2) high distinctiveness for 'salient' image features such as edges. The ALOST method is quantitatively compared to the MIND approach based on three different datasets: thoracic CT images, synthetic and real abdominal MR images. The proposed method outperformed the NMI and MIND similarity measures on these three datasets. The registration of dynamic contrast enhanced and post-contrast MR images of patients with Crohn's disease led to relative contrast enhancement measures with the highest correlation (r=0.56) to the Crohn's disease endoscopic index of severity. Dwarikanath Mahapatra, Jeroen A. W. Tielbeek, Jaap Stoker, Lucas J. van Vliet, Frans Vos |
IEEE Trans. Medical Imaging | 2 |
| 2014 | A Real-Time Smart Assistant for Video Surveillance Through Handheld DevicesabstractIn a remote surveillance system, a high resolution surveillance camera streams its video to a user's handheld device. Such devices are unable to make use of the high resolution video due to their limited display size and bandwidth. In this paper, we propose a method to assist the mobile operator of the surveillance camera in focusing on sensitive regions of the video. Our system automatically identifies relevant regions. We introduce a pan and zoom strategy to ensure that the operator is able to see fine details in these areas while maintaining contextual knowledge. Regions of interest are identified using foreground detection as well as face and body detection. The efficacy of the proposed method is demonstrated through a user study. Our proposed method was reported to be more useful than two comparable approaches for getting an understanding of the activities in a surveillance scene while maintaining context. Hao Kuang, Benjamin Guthier, Mukesh Saini, Dwarikanath Mahapatra, Abdulmotaleb El Saddik |
ACM Multimedia | 4 |
| 2014 | Analyzing Training Information From Random Forests for Improved Image SegmentationabstractLabeled training data are used for challenging medical image segmentation problems to learn different characteristics of the relevant domain. In this paper, we examine random forest (RF) classifiers, their learned knowledge during training and ways to exploit it for improved image segmentation. Apart from learning discriminative features, RFs also quantify their importance in classification. Feature importance is used to design a feature selection strategy critical for high segmentation and classification accuracy, and also to design a smoothness cost in a second-order MRF framework for graph cut segmentation. The cost function combines the contribution of different image features like intensity, texture, and curvature information. Experimental results on medical images show that this strategy leads to better segmentation accuracy than conventional graph cut algorithms that use only intensity information in the smoothness cost. Dwarikanath Mahapatra |
IEEE Trans. Image Process. | 1 |
| 2013 | Semi-Supervised and Active Learning for Automatic Segmentation of Crohn's Disease
Dwarikanath Mahapatra, Peter J. Schüffler, Jeroen A. W. Tielbeek, Frans Vos, Joachim M. Buhmann |
MICCAI (2) | 1 |
| 2013 | Automatic Detection and Segmentation of Crohn's Disease Tissues From Abdominal MRIabstractWe propose an information processing pipeline for segmenting parts of the bowel in abdominal magnetic resonance images that are affected with Crohn's disease. Given a magnetic resonance imaging test volume, it is first oversegmented into supervoxels and each supervoxel is analyzed to detect presence of Crohn's disease using random forest (RF) classifiers. The supervoxels identified as containing diseased tissues define the volume of interest (VOI). All voxels within the VOI are further investigated to segment the diseased region. Probability maps are generated for each voxel using a second set of RF classifiers which give the probabilities of each voxel being diseased, normal or background. The negative log-likelihood of these maps are used as penalty costs in a graph cut segmentation framework. Low level features like intensity statistics, texture anisotropy and curvature asymmetry, and high level context features are used at different stages. Smoothness constraints are imposed based on semantic information (importance of each feature to the classification task) derived from the second set of learned RF classifiers. Experimental results show that our method achieves high segmentation accuracy with Dice metric values of 0.90 ± 0.04 and Hausdorff distance of 7.3 ± 0.8 mm. Semantic information and context features are an integral part of our method and are robust to different levels of added noise. Dwarikanath Mahapatra, Peter J. Schüffler, Jeroen A. W. Tielbeek, Jesica Makanyanga, Jaap Stoker, Stuart A. Taylor, Frans Vos, Joachim M. Buhmann |
IEEE Trans. Medical Imaging | 1 |
| 2012 | Integrating Segmentation Information for Improved MRF-Based Elastic Image RegistrationabstractIn this paper, we propose a method to exploit segmentation information for elastic image registration using a Markov-random-field (MRF)-based objective function. MRFs are suitable for discrete labeling problems, and the labels are defined as the joint occurrence of displacement fields (for registration) and segmentation class probability. The data penalty is a combination of the image intensity (or gradient information) and the mutual dependence of registration and segmentation information. The smoothness is a function of the interaction between the defined labels. Since both terms are a function of registration and segmentation labels, the overall objective function captures their mutual dependence. A multiscale graph-cut approach is used to achieve subpixel registration and reduce the computation time. The user defines the object to be registered in the floating image, which is rigidly registered before applying our method. We test our method on synthetic image data sets with known levels of added noise and simulated deformations, and also on natural and medical images. Compared with other registration methods not using segmentation information, our proposed method exhibits greater robustness to noise and improved registration accuracy. Dwarikanath Mahapatra, Ying Sun 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Orientation Histograms as Shape Priors for Left Ventricle Segmentation Using Graph Cuts
Dwarikanath Mahapatra, Ying Sun 0001 |
MICCAI (3) | 1 |
| 2010 | An MRF framework for joint registration and segmentation of natural and perfusion imagesabstractRegistration and segmentation provide complementary information about each other. In this paper we propose a method for the joint registration and segmentation (JRS) of images using Markov random fields (MRFs). The use of MRFs allows us to formulate the problem as one of labeling and apply fast discrete optimization techniques like graph cuts. Graph cuts is able to overcome the limitations of previously used active contour frameworks namely, large number of iterations, risk of being trapped in local minima, and sensitivity to initialization. The labels in the MRF formulation indicate joint occurrence of displacement vectors and segmentation class and the energy formulation is able to capture their mutual dependency. Experiments on real patient perfusion data and natural images show that JRS gives better performance than conventional registration and segmentation methods. Dwarikanath Mahapatra, Ying Sun 0001 |
ICIP | 1 |
| 2010 | Joint Registration and Segmentation of Dynamic Cardiac Perfusion Images Using MRFs
Dwarikanath Mahapatra, Ying Sun 0001 |
MICCAI (1) | 1 |
| 2008 | Illumination invariant tracking in office environments using neurobiology-saliency based particle filterabstractBackground subtraction is a commonly employed approach for tracking in scenarios where the ambience is more or less constant in terms of illumination and number of objects. However in office environments, where the illumination can very easily change by switching off or on lights, the background subtraction method can lead to erroneous tracking. In this paper we propose a neurobiology-saliency based particle filter approach that uses low-level features like color, luminance and edge information along with motion cues to track a single person. We have tested our method on clips showing a single person carrying out a range of activities expected in an office environment. Our method performs better than a background subtraction method using a Kalman filter, in terms of the number of frames showing correct tracking and change detection for automatic initialization of tracks. Dwarikanath Mahapatra, Mukesh Saini, Ying Sun 0001 |
ICME | 1 |
| 2008 | Nonrigid Registration of Dynamic Renal MR Images Using a Saliency Based MRF Model
Dwarikanath Mahapatra, Ying Sun 0001 |
MICCAI (1) | 1 |