EDBT 2026 Demo / reviewers in the wild / expert
Arindam Dutta
dblp:42/8025
· DBLP profile ↗
17ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visibility guided Self-Supervised Occlusion-Resilient Human Pose EstimationabstractOcclusion remains a significant challenge for existing human pose estimation algorithms, often resulting in inaccurate and anatomically implausible predictions. Although recent occlusion-robust methods report strong performance, they typically rely heavily on supervised learning and privileged information, such as multiview data or temporal sequences. Furthermore, these models often fail under domain changes. Domain-adaptive human pose estimation seeks to mitigate this issue; however, when occlusions are present in the target domain, a common occurrence in real-world applications, performance of these algorithms deteriorates significantly. To address these challenges, we propose VisOR, a novel Visibility guided Self-Supervised algorithm for Occlusion-Resilient Human Pose Estimation. VisOR achieves robustness to both domain shifts and occlusions by integrating contextual reasoning with iterative pseudo-label refinement. It mitigates the overfitting to noisy labels from occluded regions via a visibility-driven curriculum learning strategy, which progressively introduces the model to increasingly occluded training samples. Additionally, VisOR is regularized by a learned human pose prior that maintains anatomical plausibility throughout the adaptation process. Recognizing the scarcity of human pose datasets with realistic occlusions, we introduce BOW Blended Occlusions in-the-Wild, a rigorously constructed context-aware synthetic benchmark designed to evaluate the occlusion resilience of human pose estimation algorithms. BOW offers a diverse range of context-aware occlusions across both indoor and outdoor environments, simulating real-world conditions. Through extensive experiments, we demonstrate that VisOR outperforms current state-of-the-art methods by ∼ 7% in challenging occluded human pose estimation benchmarks and provides a baseline performance on BOW, against existing algorithms. Arindam Dutta, Sarosij Bose, Rohit Kundu, Calvin-Khang Ta, Saketh Bachu, Konstantinos Karydis, Amit K. Roy-Chowdhury |
WACV | 1 |
| 2026 | Pose Guided Unsupervised Domain Adaptation for Human Body Part SegmentationabstractExisting algorithms for human body part segmentation have shown promising results on challenging datasets, primarily relying on end-to-end supervision. However, these algorithms exhibit severe performance drops in the face of domain shifts, leading to inaccurate segmentation masks. To tackle this issue, we introduce POSTURE: Pose Guided Unsupervised Domain Adaptation for Human Body Part Segmentation - an innovative pseudo-labelling approach 0designed to improve segmentation performance on the unlabeled target data. Distinct from conventional domain adaptive methods for general semantic segmentation, POSTURE stands out by considering the underlying structure of the human body and uses anatomical guidance from pose keypoints to drive the adaptation process. This strong inductive prior translates to impressive performance improvements, averaging 8% over existing state-of-the-art domain adaptive semantic segmentation methods across three benchmark datasets. Furthermore, the inherent flexibility of our proposed approach facilitates seamless extension to source-free settings (SF-POSTURE), effectively mitigating potential privacy and computational concerns, with negligible drop in performance. Arindam Dutta, Rohit Lal, Yash Garg, Calvin-Khang Ta, Dripta S. Raychaudhuri, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 1 |
| 2025 | Towards Source-Free Machine UnlearningabstractAs machine learning becomes more pervasive and data privacy regulations evolve, the ability to remove private or copyrighted information from trained models is becoming an increasingly critical requirement. Existing unlearning methods often rely on the assumption of having access to the entire training dataset during the forgetting process. However, this assumption may not hold true in practical scenarios where the original training data may not be accessible, i.e., the source-free setting. To address this challenge, we focus on the source-free unlearning scenario, where an unlearning algorithm must be capable of removing specific data from a trained model without requiring access to the original training dataset. Building on recent work, we present a method that can estimate the Hessian of the unknown remaining training data, a crucial component required for efficient unlearning. Leveraging this estimation technique, our method enables efficient zero-shot unlearning while providing robust theoretical guarantees on the unlearning performance, while maintaining performance on the remaining data. Extensive experiments over a wide range of datasets verify the efficacy of our method. Sk Miraj Ahmed, Umit Yigit Basaran, Dripta S. Raychaudhuri, Arindam Dutta, Rohit Kundu, Fahim Faisal Niloy, Basak Guler, Amit K. Roy-Chowdhury |
CVPR | 4 |
| 2025 | Uncertainty-Aware Diffusion-Guided Refinement of 3D ScenesabstractReconstructing 3D scenes from a single image is a fundamentally ill-posed task due to the severely under-constrained nature of the problem. Consequently, when the scene is rendered from novel camera views, existing single image to 3D reconstruction methods render incoherent and blurry views. This problem is exacerbated when the unseen regions are far away from the input camera. In this work, we address these inherent limitations in existing single image-to-3D scene feedforward networks. To alleviate the poor performance due to insufficient information beyond the input image's view, we leverage a strong generative prior in the form of a pre-trained latent video diffusion model, for iterative refinement of a coarse scene represented by optimizable Gaussian parameters. To ensure that the style and texture of the generated images align with that of the input image, we incorporate on-the-fly Fourier-style transfer between the generated images and the input image. Additionally, we design a semantic uncertainty quantification module that calculates the per-pixel entropy and yields uncertainty maps used to guide the refinement process from the most confident pixels while discarding the remaining highly uncertain ones. We conduct extensive experiments on real-world scene datasets, including in-domain RealEstate-10K and out-of-domain KITTI-v2, showing that our approach can provide more realistic and high-fidelity novel view synthesis results compared to existing state-of-the-art methods. Sarosij Bose, Arindam Dutta, Sayak Nag, Junge Zhang, Konstantinos Karydis, Amit K. Roy-Chowdhury |
ICCV | 2 |
| 2025 | CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-Consistency from a Single ImageabstractReconstructing clothed humans from a single image is a fundamental task in computer vision with wide-ranging applications. Although existing monocular clothed human reconstruction solutions have shown promising results, they often rely on the assumption that the human subject is in an occlusion-free environment. Thus, when encountering in-the-wild occluded images, these algorithms produce multiview inconsistent and fragmented reconstructions. Additionally, most algorithms for monocular 3D human reconstruction leverage geometric priors such as SMPL annotations for training and inference, which are extremely challenging to acquire in real-world applications. To address these limitations, we propose CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-ConsistEncy from a Single Image, a novel pipeline designed to reconstruct occlusion-resilient 3D humans with multiview consistency from a single occluded image, without requiring either ground-truth geometric prior annotations or 3D supervision. Specifically, CHROME leverages a multiview diffusion model to first synthesize occlusion-free human images from the occluded input, compatible with off-the-shelf pose control to explicitly enforce cross-view consistency during synthesis. A 3D reconstruction model is then trained to predict a set of 3D Gaussians conditioned on both the occluded input and synthesized views, aligning cross-view details to produce a cohesive and accurate 3D representation. CHROME achieves significant improvements in terms of both novel view synthesis (upto 3 db PSNR) and geometric reconstruction under challenging conditions. Arindam Dutta, Meng Zheng 0002, Zhongpai Gao, Benjamin Planche, Anwesa Choudhuri, Terrence Chen, Amit K. Roy-Chowdhury, Ziyan Wu 0001 |
ICCV | 1 |
| 2025 | VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation Under Real Occlusions
Yash Garg, Saketh Bachu, Arindam Dutta, Rohit Lal, Sarosij Bose, Calvin-Khang Ta, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
ICCV | 3 |
| 2025 | Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language ModelsabstractVision-language models (VLMs) have improved significantly in their capabilities, but their complex architecture makes their safety alignment challenging. In this paper, we reveal an uneven distribution of harmful information across the intermediate layers of the image encoder and show that skipping a certain set of layers and exiting early can increase the chance of the VLM generating harmful responses. We call it as “Image enCoder Early-exiT” based vulnerability (ICET). Our experiments across three VLMs: LLaVA-1.5, LLaVA-NeXT, and Llama 3.2 show that performing early exits from the image encoder significantly increases the likelihood of generating harmful outputs. To tackle this, we propose a simple yet effective modification of the Clipped-Proximal Policy Optimization (Clip-PPO) algorithm for performing layer-wise multi-modal RLHF for VLMs. We term this as Layer-Wise PPO (L-PPO). We evaluate our L-PPO algorithm across three multi-modal datasets and show that it consistently reduces the harmfulness caused by early exits. Saketh Bachu, Erfan Shayegani, Rohit Lal, Trishna Chakraborty, Arindam Dutta, Chengyu Song, Yue Dong 0002, Nael B. Abu-Ghazaleh, Amit K. Roy-Chowdhury |
ICML | 5 |
| 2025 | ODES: Online Domain Adaptation with Expert Guidance for Medical Image Segmentation
Md Shazid Islam, Sayak Nag, Arindam Dutta, Sk Miraj Ahmed, Fahim Faisal Niloy, Shreyangshu Bera, Amit K. Roy-Chowdhury |
MICCAI (4) | 3 |
| 2025 | STRIDE: Single-Video Based Temporally Continuous Occlusion-Robust 3D Pose EstimationabstractAccurately estimating 3D human poses is crucial for fields like action recognition, gait recognition, and virtual/augmented reality. However, predicting human poses under severe occlusion remains a persistent and significant challenge. Existing image-based estimators struggle with heavy occlusions due to a lack of temporal context, resulting in inconsistent predictions, while video-based models, despite benefiting from temporal data, face limitations with prolonged occlusions over multiple frames. Additionally, existing algorithms often struggle to generalize unseen videos. Addressing these challenges, we propose STRIDE (Single-video based TempoRally contInuous Occlusion-Robust 3D Pose Estimation), a novel Test-Time Training (TTT) approach to fit a human motion prior for estimating 3D human poses for each video. Our proposed approach handles occlusions not encountered during the model's training by refining a sequence of noisy initial pose estimates into accurate, temporally coherent poses at test time, effectively overcoming the limitations of existing methods. Our flexible, model-agnostic framework allows us to use any off-the-shelf 3D pose estimation method to improve robustness and temporal consistency. We validate STRIDE's efficacy through comprehensive experiments on multiple challenging datasets where it not only outperforms existing single-image and video-based pose estimation models but also showcases superior handling of substantial occlusions, achieving fast, robust, accurate, and temporally consistent 3D pose estimates. Code is made publicly available at https://github.com/take2rohit/stride Rohit Lal, Saketh Bachu, Yash Garg, Arindam Dutta, Calvin-Khang Ta, Hannah Dela Cruz, Dripta S. Raychaudhuri, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
WACV | 4 |
| 2025 | DEGAST3D: Learning Deformable 3D Graph Similarity to Track Plant Cells in Unregistered Time Lapse ImagesabstractTracking plant cells in three-dimensional (3D) tissue captured through light microscopy presents significant challenges due to the large number of densely packed cells, non-uniform growth patterns, and variations in cell division planes across different cell layers. In addition, images of deeper tissue layers are often noisy, and systemic imaging errors further exacerbate the complexity of the task. In this paper, we propose a novel learning-based method DEGAST3D: Learning Deformable 3D GrAph Similarity to Track Plant Cells in Unregistered Time Lapse Images exploits the tightly packed 3D cell structure of plant cells to create a three-dimensional graph for accurate cell tracking. We also propose a novel algorithm for cell division detection and an effective three-dimensional registration, improving state-of-the-art algorithms. On a public dataset, our novel cell pair matching method outperforms the baseline by $6.83 \%$, $5.96 \%$, $6.40 \%$ in precision, recall, and F-1 score, respectively. On the same dataset, our proposed novel cell division technique improves the results of the baseline method by $15.38 \%$ and $14.78 \%$ in terms of recall and F1-score, respectively. Md Shazid Islam, Arindam Dutta, Calvin-Khang Ta, Kevin Rodriguez, Christian Michael, Mark S. Alber, G. Venugopala Reddy, Amit K. Roy-Chowdhury |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | POISE: Pose Guided Human Silhouette Extraction under OcclusionsabstractHuman silhouette extraction is a fundamental task in computer vision with applications in various downstream tasks. However, occlusions pose a significant challenge, leading to incomplete and distorted silhouettes. To address this challenge, we introduce POISE: Pose Guided Human Silhouette Extraction under Occlusions, a novel self-supervised fusion framework that enhances accuracy and robustness in human silhouette prediction. By combining initial silhouette estimates from a segmentation model with human joint predictions from a 2D pose estimation model, POISE leverages the complementary strengths of both approaches, effectively integrating precise body shape information and spatial information to tackle occlusions. Furthermore, the self-supervised nature of POISE eliminates the need for costly annotations, making it scalable and practical. Extensive experimental results demonstrate its superiority in improving silhouette extraction under occlusions, with promising results in downstream tasks such as gait recognition. The code for our method is available https://github.com/take2rohit/poise. Arindam Dutta, Rohit Lal, Dripta S. Raychaudhuri, Calvin-Khang Ta, Amit K. Roy-Chowdhury |
WACV | 1 |
| 2023 | Prior-guided Source-free Domain Adaptation for Human Pose EstimationabstractDomain adaptation methods for 2D human pose estimation typically require continuous access to the source data during adaptation, which can be challenging due to privacy, memory, or computational constraints. To address this limitation, we focus on the task of source-free domain adaptation for pose estimation, where a source model must adapt to a new target domain using only unlabeled target data. Although recent advances have introduced source-free methods for classification tasks, extending them to the regression task of pose estimation is non-trivial. In this paper, we present Prior-guided Self-training (POST), a pseudo-labeling approach that builds on the popular Mean Teacher framework to compensate for the distribution shift. POST leverages prediction-level and feature-level consistency between a student and teacher model against certain image transformations. In the absence of source data, POST utilizes a human pose prior that regularizes the adaptation process by directing the model to generate more accurate and anatomically plausible pose pseudo-labels. Despite being simple and intuitive, our framework can deliver significant performance gains compared to applying the source model directly to the target data, as demonstrated in our extensive experiments and ablation studies. In fact, our approach achieves comparable performance to recent state-of-the-art methods that use source data for adaptation. Dripta S. Raychaudhuri, Calvin-Khang Ta, Arindam Dutta, Rohit Lal, Amit K. Roy-Chowdhury |
ICCV | 3 |
| 2022 | Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous ProcessorabstractRF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design. Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue |
ISCAS | 17 |
| 2019 | Fault Classification in Three Phase Self-Excited Induction Generators using Deep Neural NetworksabstractIn this paper, an algorithm is proposed for the classification of faults in three phase self-excited induction generators using deep neural networks, on the basis of their voltage and current waveforms. Three phase self-exited induction generators, which are mostly used in wind power stations, are often connected to the national grid. Therefore, the transient stability analysis of this machine, prior and post symmetrical and unsymmetrical short circuit faults is one of the main concerns in power system security and operation. In this study, voltage and current waveforms of faults have been simulated in the Simulink environment, for different conditions of fault. Following this, visual time-frequency representations of the fault signals called scalograms are created, using continuous wavelet transform. Finally, a deep convolutional neural network is used for classification of the fault signals. Experimental results show that a final accuracy of 82.14% as well as real-time inference is achieved on the validation set, using the proposed scheme. Sohom Mukherjee, Arindam Dutta, Saradindu Ghosh |
TENCON | 2 |
| 2019 | Convolutional Neural Networks for Noise Classification and Denoising of ImagesabstractThe goal of this paper is to find whether a convolutional neural network (CNN) performs better than the existing blind algorithms for image denoising, and, if yes, whether the noise statistics has an effect on the performance gap. For automatic identification of noise distribution, we used two different convolutional neural networks, VGG-16 and Inception-v3, and it was found that Inception-v3 identifies the noise distribution more accurately over a set of nine possible distributions, namely, Gaussian, log-normal, uniform, exponential, Poisson, salt and pepper, Rayleigh, speckle and Erlang. Next, for each of these noisy image sets, we compared the performance of FFDNet, a CNN based denoising method, with noise clinic, a blind denoising algorithm. It was found that CNN based denoising outperforms blind denoising in general, with an average improvement of 16% in peak signal to noise ratio (PSNR). The improvement is however very prominent for salt and pepper type noise with a PSNR difference of 72%, whereas for noise distributions such as Gaussian, FFDNet could achieve only a 2% improvement over noise clinic. The results indicate that for developing a CNN based optimum denoising platform, consideration of noise distribution is necessary. Dibakar Sil, Arindam Dutta, Aniruddha Chandra |
TENCON | 2 |
| 2016 | Learning approach for classification of GENEActiv accelerometer data for unique activity identificationabstractRecent popular emphasis on exercise for personal wellbeing has created a demand for techniques which monitor and classify human activities. Previous studies have shown promising results in applying various classification and feature extraction methods for identifying unique physical activities on various datasets. We apply learning techniques to GENEactiv accelerometer recordings to identify and monitor a wide range of daily activities. The dataset is composed of 92 participants, of ages 20-65, performing 25 unique activities, both ambulatory and non-ambulatory. The algorithm identified 130 different time and frequency domain features and selected the most efficient features with the sequential forward selection algorithm. With classification in two stages with both Gaussian mixture model (GMM) and hidden Markov model (HMM) we have combined the activities with similar features. We have also shown a comparative study between the two classifiers. We achieved an accuracy of 95.5% while classifying 10 unique activities with HMM and 89.7% while classifying 9. The most efficient result is obtained using HMM in 2-D feature space, where it is able to classify 15 unique activities at an accuracy of 90.12%. Arindam Dutta, Owen Ma, Matthew P. Buman, Daniel W. Bliss |
BSN | 1 |
| 2016 | Comparing Gaussian Mixture Model and Hidden Markov Model to Classify Unique Physical Activities from Accelerometer Sensor DataabstractWith the recent interest in physical therapy through sufficient physical activity, considerable efforts have been made to monitor and classify daily human activities, especially for people who need physical rehabilitation. In our previous study, we designed a classifier to identify 25 unique physical activities performed by 92 healthy participants between the ages of 20 and 65. In this study, with the use of a GENEActiv accelerometer to monitor a wide range of daily activities, we present a learning approach to identify unique activities performed by a varied group of participants with various health conditions. The dataset is comprised of 99 senior participants and 23 participants who are significantly taller in height than the general population, performing 8 unique activities. We have extracted 130 different features in time and frequency domain and selected the most efficient features with the sequential forward selection algorithm. With two stages of classification, the first is utilized for combining similar classes, and the second determines the final decision. We have tested two classifiers for our learning approach, the Gaussian mixture models (GMMs) and the hidden Markov models (HMMs) and compared their performances. We have improved the GMM classifier from our previous study and it has shown more promising results for this dataset. We achieved an accuracy of 88.92% when classifying the 8 unique activities with GMM and 93.5% with HMM when classifying 7 activities. Arindam Dutta, Owen Ma, Meynard John Toledo, Matthew P. Buman, Daniel W. Bliss |
ICMLA | 1 |