EDBT 2026 Demo / reviewers in the wild / expert
Calvin-Khang Ta
dblp:321/9960
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0001-7191-7981ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visibility guided Self-Supervised Occlusion-Resilient Human Pose EstimationabstractOcclusion remains a significant challenge for existing human pose estimation algorithms, often resulting in inaccurate and anatomically implausible predictions. Although recent occlusion-robust methods report strong performance, they typically rely heavily on supervised learning and privileged information, such as multiview data or temporal sequences. Furthermore, these models often fail under domain changes. Domain-adaptive human pose estimation seeks to mitigate this issue; however, when occlusions are present in the target domain, a common occurrence in real-world applications, performance of these algorithms deteriorates significantly. To address these challenges, we propose VisOR, a novel Visibility guided Self-Supervised algorithm for Occlusion-Resilient Human Pose Estimation. VisOR achieves robustness to both domain shifts and occlusions by integrating contextual reasoning with iterative pseudo-label refinement. It mitigates the overfitting to noisy labels from occluded regions via a visibility-driven curriculum learning strategy, which progressively introduces the model to increasingly occluded training samples. Additionally, VisOR is regularized by a learned human pose prior that maintains anatomical plausibility throughout the adaptation process. Recognizing the scarcity of human pose datasets with realistic occlusions, we introduce BOW Blended Occlusions in-the-Wild, a rigorously constructed context-aware synthetic benchmark designed to evaluate the occlusion resilience of human pose estimation algorithms. BOW offers a diverse range of context-aware occlusions across both indoor and outdoor environments, simulating real-world conditions. Through extensive experiments, we demonstrate that VisOR outperforms current state-of-the-art methods by ∼ 7% in challenging occluded human pose estimation benchmarks and provides a baseline performance on BOW, against existing algorithms. Arindam Dutta, Sarosij Bose, Rohit Kundu, Calvin-Khang Ta, Saketh Bachu, Konstantinos Karydis, Amit K. Roy-Chowdhury |
WACV | 4 |
| 2026 | Pose Guided Unsupervised Domain Adaptation for Human Body Part SegmentationabstractExisting algorithms for human body part segmentation have shown promising results on challenging datasets, primarily relying on end-to-end supervision. However, these algorithms exhibit severe performance drops in the face of domain shifts, leading to inaccurate segmentation masks. To tackle this issue, we introduce POSTURE: Pose Guided Unsupervised Domain Adaptation for Human Body Part Segmentation - an innovative pseudo-labelling approach 0designed to improve segmentation performance on the unlabeled target data. Distinct from conventional domain adaptive methods for general semantic segmentation, POSTURE stands out by considering the underlying structure of the human body and uses anatomical guidance from pose keypoints to drive the adaptation process. This strong inductive prior translates to impressive performance improvements, averaging 8% over existing state-of-the-art domain adaptive semantic segmentation methods across three benchmark datasets. Furthermore, the inherent flexibility of our proposed approach facilitates seamless extension to source-free settings (SF-POSTURE), effectively mitigating potential privacy and computational concerns, with negligible drop in performance. Arindam Dutta, Rohit Lal, Yash Garg, Calvin-Khang Ta, Dripta S. Raychaudhuri, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 4 |
| 2025 | Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph GenerationabstractScene Graph Generation (SGG) aims to represent visual scenes by identifying objects and their pairwise relationships, providing a structured understanding of image content. However, inherent challenges like long-tailed class distributions and prediction variability necessitate uncertainty quantification in SGG for its practical viability. In this paper, we introduce a novel Conformal Prediction based framework, adaptive to any existing SGG method, for quantifying their predictive uncertainty by constructing well-calibrated prediction sets over their generated scene graphs. These scene graph prediction sets are designed to achieve statistically rigorous coverage guarantees under exchangeability assumptions. Additionally, to ensure the prediction sets contain the most practically interpretable scene graphs, we propose an effective MLLM-based post-processing strategy for selecting the most visually and semantically plausible scene graphs within each set. We show that our proposed approach can produce diverse possible scene graphs from an image, assess the reliability of SGG methods, and improve overall SGG performance. Sayak Nag, Udita Ghosh, Calvin-Khang Ta, Sarosij Bose, Amit K. Roy-Chowdhury |
CVPR | 3 |
| 2025 | VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation Under Real Occlusions
Yash Garg, Saketh Bachu, Arindam Dutta, Rohit Lal, Sarosij Bose, Calvin-Khang Ta, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
ICCV | 6 |
| 2025 | STRIDE: Single-Video Based Temporally Continuous Occlusion-Robust 3D Pose EstimationabstractAccurately estimating 3D human poses is crucial for fields like action recognition, gait recognition, and virtual/augmented reality. However, predicting human poses under severe occlusion remains a persistent and significant challenge. Existing image-based estimators struggle with heavy occlusions due to a lack of temporal context, resulting in inconsistent predictions, while video-based models, despite benefiting from temporal data, face limitations with prolonged occlusions over multiple frames. Additionally, existing algorithms often struggle to generalize unseen videos. Addressing these challenges, we propose STRIDE (Single-video based TempoRally contInuous Occlusion-Robust 3D Pose Estimation), a novel Test-Time Training (TTT) approach to fit a human motion prior for estimating 3D human poses for each video. Our proposed approach handles occlusions not encountered during the model's training by refining a sequence of noisy initial pose estimates into accurate, temporally coherent poses at test time, effectively overcoming the limitations of existing methods. Our flexible, model-agnostic framework allows us to use any off-the-shelf 3D pose estimation method to improve robustness and temporal consistency. We validate STRIDE's efficacy through comprehensive experiments on multiple challenging datasets where it not only outperforms existing single-image and video-based pose estimation models but also showcases superior handling of substantial occlusions, achieving fast, robust, accurate, and temporally consistent 3D pose estimates. Code is made publicly available at https://github.com/take2rohit/stride Rohit Lal, Saketh Bachu, Yash Garg, Arindam Dutta, Calvin-Khang Ta, Hannah Dela Cruz, Dripta S. Raychaudhuri, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
WACV | 5 |
| 2025 | DEGAST3D: Learning Deformable 3D Graph Similarity to Track Plant Cells in Unregistered Time Lapse ImagesabstractTracking plant cells in three-dimensional (3D) tissue captured through light microscopy presents significant challenges due to the large number of densely packed cells, non-uniform growth patterns, and variations in cell division planes across different cell layers. In addition, images of deeper tissue layers are often noisy, and systemic imaging errors further exacerbate the complexity of the task. In this paper, we propose a novel learning-based method DEGAST3D: Learning Deformable 3D GrAph Similarity to Track Plant Cells in Unregistered Time Lapse Images exploits the tightly packed 3D cell structure of plant cells to create a three-dimensional graph for accurate cell tracking. We also propose a novel algorithm for cell division detection and an effective three-dimensional registration, improving state-of-the-art algorithms. On a public dataset, our novel cell pair matching method outperforms the baseline by $6.83 \%$, $5.96 \%$, $6.40 \%$ in precision, recall, and F-1 score, respectively. On the same dataset, our proposed novel cell division technique improves the results of the baseline method by $15.38 \%$ and $14.78 \%$ in terms of recall and F1-score, respectively. Md Shazid Islam, Arindam Dutta, Calvin-Khang Ta, Kevin Rodriguez, Christian Michael, Mark S. Alber, G. Venugopala Reddy, Amit K. Roy-Chowdhury |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | POISE: Pose Guided Human Silhouette Extraction under OcclusionsabstractHuman silhouette extraction is a fundamental task in computer vision with applications in various downstream tasks. However, occlusions pose a significant challenge, leading to incomplete and distorted silhouettes. To address this challenge, we introduce POISE: Pose Guided Human Silhouette Extraction under Occlusions, a novel self-supervised fusion framework that enhances accuracy and robustness in human silhouette prediction. By combining initial silhouette estimates from a segmentation model with human joint predictions from a 2D pose estimation model, POISE leverages the complementary strengths of both approaches, effectively integrating precise body shape information and spatial information to tackle occlusions. Furthermore, the self-supervised nature of POISE eliminates the need for costly annotations, making it scalable and practical. Extensive experimental results demonstrate its superiority in improving silhouette extraction under occlusions, with promising results in downstream tasks such as gait recognition. The code for our method is available https://github.com/take2rohit/poise. Arindam Dutta, Rohit Lal, Dripta S. Raychaudhuri, Calvin-Khang Ta, Amit K. Roy-Chowdhury |
WACV | 4 |
| 2023 | Prior-guided Source-free Domain Adaptation for Human Pose EstimationabstractDomain adaptation methods for 2D human pose estimation typically require continuous access to the source data during adaptation, which can be challenging due to privacy, memory, or computational constraints. To address this limitation, we focus on the task of source-free domain adaptation for pose estimation, where a source model must adapt to a new target domain using only unlabeled target data. Although recent advances have introduced source-free methods for classification tasks, extending them to the regression task of pose estimation is non-trivial. In this paper, we present Prior-guided Self-training (POST), a pseudo-labeling approach that builds on the popular Mean Teacher framework to compensate for the distribution shift. POST leverages prediction-level and feature-level consistency between a student and teacher model against certain image transformations. In the absence of source data, POST utilizes a human pose prior that regularizes the adaptation process by directing the model to generate more accurate and anatomically plausible pose pseudo-labels. Despite being simple and intuitive, our framework can deliver significant performance gains compared to applying the source model directly to the target data, as demonstrated in our extensive experiments and ablation studies. In fact, our approach achieves comparable performance to recent state-of-the-art methods that use source data for adaptation. Dripta S. Raychaudhuri, Calvin-Khang Ta, Arindam Dutta, Rohit Lal, Amit K. Roy-Chowdhury |
ICCV | 2 |
| 2022 | Poisson2Sparse: Self-supervised Poisson Denoising from a Single Image
Calvin-Khang Ta, Abhishek Aich, Akash Gupta 0001, Amit K. Roy-Chowdhury |
MICCAI (8) | 1 |
| 2022 | GAMA: Generative Adversarial Multi-Object Scene AttacksabstractThe majority of methods for crafting adversarial attacks have focused on scenes with a single dominant object (e.g., images from ImageNet). On the other hand, natural scenes include multiple dominant objects that are semantically related. Thus, it is crucial to explore designing attack strategies that look beyond learning on single-object scenes or attack single-object victim classifiers. Due to their inherent property of strong transferability of perturbations to unknown models, this paper presents the first approach of using generative models for adversarial attacks on multi-object scenes. In order to represent the relationships between different objects in the input scene, we leverage upon the open-sourced pre-trained vision-language model CLIP (Contrastive Language-Image Pre-training), with the motivation to exploit the encoded semantics in the language space along with the visual space. We call this attack approach Generative Adversarial Multi-object Attacks (GAMA). GAMA demonstrates the utility of the CLIP model as an attacker's tool to train formidable perturbation generators for multi-object scenes. Using the joint image-text features to train the generator, we show that GAMA can craft potent transferable perturbations in order to fool victim classifiers in various attack settings. For example, GAMA triggers ~16% more misclassification than state-of-the-art generative approaches in black-box settings where both the classifier architecture and data distribution of the attacker are different from the victim. Our code is available here: https://abhishekaich27.github.io/gama.html Abhishek Aich, Calvin-Khang Ta, Akash Gupta 0001, Chengyu Song, Srikanth V. Krishnamurthy, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
NeurIPS | 2 |
| 2022 | Combined computational modeling and experimental analysis integrating chemical and mechanical signals suggests possible mechanism of shoot meristem maintenanceabstractStem cell maintenance in multilayered shoot apical meristems (SAMs) of plants requires strict regulation of cell growth and division. Exactly how the complex milieu of chemical and mechanical signals interact in the central region of the SAM to regulate cell division plane orientation is not well understood. In this paper, simulations using a newly developed multiscale computational model are combined with experimental studies to suggest and test three hypothesized mechanisms for the regulation of cell division plane orientation and the direction of anisotropic cell expansion in the corpus. Simulations predict that in the Apical corpus, WUSCHEL and cytokinin regulate the direction of anisotropic cell expansion, and cells divide according to tensile stress on the cell wall. In the Basal corpus, model simulations suggest dual roles for WUSCHEL and cytokinin in regulating both the direction of anisotropic cell expansion and cell division plane orientation. Simulation results are followed by a detailed analysis of changes in cell characteristics upon manipulation of WUSCHEL and cytokinin in experiments that support model predictions. Moreover, simulations predict that this layer-specific mechanism maintains both the experimentally observed shape and structure of the SAM as well as the distribution of WUSCHEL in the tissue. This provides an additional link between the roles of WUSCHEL, cytokinin, and mechanical stress in regulating SAM growth and proper stem cell maintenance in the SAM. Mikahl Banwarth-Kuhn, Kevin Rodriguez, Christian Michael, Calvin-Khang Ta, Alexander Plong, Eric Bourgain-Chang, Ali Nematbakhsh, Amit K. Roy-Chowdhury, G. Venugopala Reddy, Mark S. Alber |
PLoS Comput. Biol. | 4 |