VLDB 2026 Research / reviewers in the wild / expert
Ketul Shah
dblp:220/4323
· DBLP profile ↗
10ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VRAgent: Self-Refining Agent for Zero-Shot Multimodal Video RetrievalabstractRecent advances in Vision-Language Models (VLMs) and Large Language Models (LLMs) have demonstrated remarkable zero-shot capabilities while becoming increasingly accessible. They have enabled zero-shot text-to-video search, yet they still struggle on real-world queries that demand temporally aligned reasoning across vision and speech. We present VRAgent, an agentic retrieval framework that leverages a central LLM as a planner which (i) decomposes a free-form user query into a tool-instruction set spanning visual, dialogue and other modality-specific retrievers, and (ii) iteratively self-refines this plan by scoring the outputs and rewriting the next tool-instruction set. The resulting closed-loop optimization acts at test time and requires no additional training data or gradient updates.VRAgent is modular by design—adding a new modality is as simple as adding a corresponding foundation model into the toolbox. On our newly proposed MM-MSRVTT and TVR-1200 multimodal benchmarks, VRAgent improves average recall by +8.3% and +1.7% over the best zero-shot baselines, while on single-modality MSR-VTT and DiDeMo it obtains consistent gains of +3.7% and +4.1%. An interactive variant that asks the user up to two multiple-choice questions pushes average recall to 79.7% on MSR-VTT, underscoring the value of on-the-fly human feedback. Ketul Shah, Pankaj Nathani, Rama Chellappa, Fabian Caba Heilbron |
WACV | 1 |
| 2025 | AeroGen: Ground-to-Air Generalization for Action RecognitionabstractWe address the problem of action recognition from aerial views using only ground-based videos for training. Due to the viewpoint-induced domain shift, models trained solely on ground videos exhibit significant performance degradation when naively applied to aerial videos. To mitigate this performance gap, we introduce a domain generalization technique that addresses the viewpoint-induced domain shift. Our method uses available real ground videos to generate additional synthetic training data from both ground and air viewpoints, for improving generalization to aerial video-based action recognition. Specifically, we perform 3D human mesh estimation from the ground videos, then render synthetic videos from alternate viewpoints with additional appearance randomizations. To further align the ground-air and syntheticreal domains, we propose a Dual Domain Alignment loss by enforcing consistency in predictions between the original ground videos and augmented videos from each domain. In order to facilitate research on the problem of ground-to-air generalization for human action recognition, we also create a benchmark by combining parts of the NTU-60, UAV-Human and NEC-DRONE datasets. We demonstrate the effectiveness of our approach on these new benchmarks along with the existing RoCoG-Ground $\rightarrow$ RoCoG-Air benchmark, and also perform extensive ablations. Ketul Shah, Anshul Shah 0001, Arun V. Reddy, Aniket Roy, Celso de Melo, Rama Chellappa |
FG | 1 |
| 2025 | DIFFUSE2ADAPT: Controlled Diffusion for Synthetic-to-Real Domain AdaptationabstractSynthetic data generated from graphics engines has been shown to be effective for learning, while also being a cost-effective alternative to annotating real-world data. However, models trained on synthetic data often suffer from performance degradation when applied to real-world data due to the domain gap. In this paper, we propose Diffuse2Adapt, a novel unsupervised domain adaptation (UDA) approach that leverages controlled diffusion models to bridge the synthetic-to-real domain gap. Our method utilizes text-to-image generative models to translate synthetic images to the target domain while preserving class semantics. We introduce two methods to reduce the domain gap: (1) incorporating target domain context extracted from multimodal language models, and (2) capturing target domain style via learned textual tokens. Extensive experiments are performed on three synthetic-to-real domain adaptation benchmarks, VisDA-2017, S2RDA-49, and S2RDA-MS-39. Diffuse2Adapt outperforms state-of-the-art methods by +3.00% on VisDA-2017, +2.55% on S2RDA-49 and +1.01% on S2RDA-MS-39. Code will be released. Ketul Shah, Arushi Sinha, Arun V. Reddy, Aniket Roy, Rama Chellappa |
ICIP | 1 |
| 2025 | Cap2Aug: Caption Guided Image data AugmentationabstractVisual recognition in a low-data regime is challenging and often prone to overfitting. To mitigate this issue, several data augmentation strategies have been proposed. However, standard transformations, e.g., rotation, cropping, and flip-ping provide limited semantic variations. To this end, we propose Cap2Aug, an image-to-image diffusion model-based data augmentation strategy using image captions to condition the image synthesis step. We generate a caption for an image and use this caption as an additional input for an image-to-image diffusion model. This increases the semantic diversity of the augmented images due to caption conditioning compared to the usual data augmentation techniques. We show that Cap2Aug is particularly effective where only a few samples are available for an object class. However, naively generating the synthetic images is not adequate due to the domain gap between real and synthetic images. Thus, we employ a maximum mean discrepancy loss to align the synthetic images to the real images to minimize the domain gap. We evaluate our method on few-shot classification and image classification with long-tail class distribution tasks. Cap2Aug achieves state-of-the-art performance on both tasks while evaluated on eleven benchmarks. Code: https://github.com/aniket004/Cap_2_Aug.git Aniket Roy, Anshul Shah 0001, Ketul Shah, Rama Chellappa |
WACV | 3 |
| 2024 | Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingabstractIn this work, we tackle the problem of unsupervised domain adaptation (UDA) for video action recognition. Our approach, which we call UNITE, uses an image teacher model to adapt a video student model to the target domain. UNITE first employs self-supervised pretraining to promote discriminative feature learning on target domain videos using a teacher-guided masked distillation objective. We then perform self-training on masked target data, using the video student model and image teacher model together to generate improved pseudolabels for unlabeled target videos. Our self-training process successfully leverages the strengths of both models to achieve strong transfer performance across domains. We evaluate our approach on multiple video domain adaptation benchmarks and observe significant improvements upon previously reported results. Arun V. Reddy, William Paul, Corban Rivera, Ketul Shah, Celso de Melo, Rama Chellappa |
CVPR | 4 |
| 2023 | HaLP: Hallucinating Latent Positives for Skeleton-based Self-Supervised Learning of ActionsabstractSupervised learning of skeleton sequence encoders for action recognition has received significant attention in recent times. However, learning such encoders without labels continues to be a challenging problem. While prior works have shown promising results by applying contrastive learning to pose sequences, the quality of the learned representations is often observed to be closely tied to data augmentations that are used to craft the positives. However, augmenting pose sequences is a difficult task as the geometric constraints among the skeleton joints need to be enforced to make the augmentations realistic for that action. In this work, we propose a new contrastive learning approach to train models for skeleton-based action recognition without labels. Our key contribution is a simple module, HaLP - to Hallucinate Latent Positives for contrastive learning. Specifically, HaLP explores the latent space of poses in suitable directions to generate new positives. To this end, we present a novel optimization formulation to solve for the synthetic positives with an explicit control on their hardness. We propose approximations to the objective, making them solvable in closed form with minimal overhead. We show via experiments that using these generated positives within a standard contrastive learning framework leads to consistent improvements across benchmarks such as NTU-60, NTU-120, and PKU-II on tasks like linear evaluation, transfer learning, and kNN evaluation. Our code can be found at https://github.com/anshulbshah/HaLP. Anshul Shah 0001, Aniket Roy, Ketul Shah, Shlok Kumar Mishra, David Jacobs 0001, Anoop Cherian, Rama Chellappa |
CVPR | 3 |
| 2023 | Synthetic-to-Real Domain Adaptation for Action Recognition: A Dataset and Baseline PerformancesabstractHuman action recognition is a challenging problem, particularly when there is high variability in factors such as subject appearance, backgrounds and viewpoint. While deep neural networks (DNNs) have been shown to perform well on action recognition tasks, they typically require large amounts of high-quality labeled data to achieve robust performance across a variety of conditions. Synthetic data has shown promise as a way to avoid the substantial costs and potential ethical concerns associated with collecting and labeling enormous amounts of data in the real-world. However, synthetic data may differ from real data in important ways. This phenomenon, known as domain shift, can limit the utility of synthetic data in robotics applications. To mitigate the effects of domain shift, substantial effort is being dedicated to the development of domain adaptation (DA) techniques. Yet, much remains to be understood about how best to develop these techniques. In this paper, we introduce a new dataset called Robot Control Gestures (RoCoG-v2). The dataset is composed of both real and synthetic videos from seven gesture classes, and is intended to support the study of synthetic-to-real domain shift for video-based action recognition. Our work expands upon existing datasets by focusing the action classes on gestures for human-robot teaming, as well as by enabling investigation of domain shift in both ground and aerial views. We present baseline results using state-of-the-art action recognition and domain adaptation algorithms and offer initial insight on tackling the synthetic-to-real and ground-to-air domain shifts. Instructions on accessing the dataset can be found at https://github.com/reddyav1/RoCoG-v2. Arun V. Reddy, Ketul Shah, William Paul, Rohita Mocharla, Judy Hoffman, Kapil D. Katyal, Dinesh Manocha, Celso de Melo, Rama Chellappa |
ICRA | 2 |
| 2023 | Multi-View Action Recognition using Contrastive LearningabstractIn this work, we present a method for RGB-based action recognition using multi-view videos. We present a supervised contrastive learning framework to learn a feature embedding robust to changes in viewpoint, by effectively leveraging multi-view data. We use an improved supervised contrastive loss and augment the positives with those coming from synchronized viewpoints. We also propose a new approach to use classifier probabilities to guide the selection of hard negatives in the contrastive loss, to learn a more discriminative representation. Negative samples from confusing classes based on posterior are weighted higher. We also show that our method leads to better domain generalization compared to the standard supervised training based on synthetic multi-view data. Extensive experiments on real (NTU-60, NTU-120, NUMA) and synthetic (RoCoG) data demonstrate the effectiveness of our approach. Ketul Shah, Anshul Shah 0001, Chun Pong Lau 0001, Celso de Melo, Rama Chellappa |
WACV | 1 |
| 2022 | FeLMi : Few shot Learning with hard MixupabstractLearning from a few examples is a challenging computer vision task. Traditionally,meta-learning-based methods have shown promise towards solving this problem.Recent approaches show benefits by learning a feature extractor on the abundantbase examples and transferring these to the fewer novel examples. However, thefinetuning stage is often prone to overfitting due to the small size of the noveldataset. To this end, we propose Few shot Learning with hard Mixup (FeLMi)using manifold mixup to synthetically generate samples that helps in mitigatingthe data scarcity issue. Different from a naïve mixup, our approach selects the hardmixup samples using an uncertainty-based criteria. To the best of our knowledge,we are the first to use hard-mixup for the few-shot learning problem. Our approachallows better use of the pseudo-labeled base examples through base-novel mixupand entropy-based filtering. We evaluate our approach on several common few-shotbenchmarks - FC-100, CIFAR-FS, miniImageNet and tieredImageNet and obtainimprovements in both 1-shot and 5-shot settings. Additionally, we experimented onthe cross-domain few-shot setting (miniImageNet → CUB) and obtain significantimprovements. Aniket Roy, Anshul Shah 0001, Ketul Shah, Prithviraj Dhar, Anoop Cherian, Rama Chellappa |
NeurIPS | 3 |
| 2020 | Improved Modeling of 3D Shapes with Multi-view Depth MapsabstractWe present a simple yet effective general-purpose framework for modeling 3D shapes by leveraging recent advances in 2D image generation using CNNs. Using just a single depth image of the object, we can output a dense multi-view depth map representation of 3D objects. Our simple encoder-decoder framework, comprised of a novel identity encoder and class-conditional viewpoint generator, generates 3D consistent depth maps. Our experimental results demonstrate the two-fold advantage of our approach. First, we can directly borrow architectures that work well in the 2D image domain to 3D. Second, we can effectively generate high-resolution 3D shapes with low computational memory. Our quantitative evaluations show that our method is superior to existing depth map methods for reconstructing and synthesizing 3D objects and is competitive with other representations, such as point clouds, voxel grids, and implicit functions. Code and other material will be made available at http://multiview-shapes. umiacs.io. Kamal Gupta 0002, Susmija Jabbireddy, Ketul Shah, Abhinav Shrivastava, Matthias Zwicker |
3DV | 3 |