Anshul Shah 0001

dblp:250/5430 · also Anshul B. Shah · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 AeroGen: Ground-to-Air Generalization for Action Recognition
abstract
We address the problem of action recognition from aerial views using only ground-based videos for training. Due to the viewpoint-induced domain shift, models trained solely on ground videos exhibit significant performance degradation when naively applied to aerial videos. To mitigate this performance gap, we introduce a domain generalization technique that addresses the viewpoint-induced domain shift. Our method uses available real ground videos to generate additional synthetic training data from both ground and air viewpoints, for improving generalization to aerial video-based action recognition. Specifically, we perform 3D human mesh estimation from the ground videos, then render synthetic videos from alternate viewpoints with additional appearance randomizations. To further align the ground-air and syntheticreal domains, we propose a Dual Domain Alignment loss by enforcing consistency in predictions between the original ground videos and augmented videos from each domain. In order to facilitate research on the problem of ground-to-air generalization for human action recognition, we also create a benchmark by combining parts of the NTU-60, UAV-Human and NEC-DRONE datasets. We demonstrate the effectiveness of our approach on these new benchmarks along with the existing RoCoG-Ground $\rightarrow$ RoCoG-Air benchmark, and also perform extensive ablations.
Ketul Shah, Anshul Shah 0001, Arun V. Reddy, Aniket Roy, Celso de Melo, Rama Chellappa
FG2
2025 GaitContour: Efficient Gait Recognition Based on a Contour-Pose Representation
abstract
Gait recognition holds the promise to robustly identify subjects based on walking patterns instead of appearance information. In recent years, this field has been dominated by learning methods based on two input formats: silhouette images and sparse keypoints. Compared to image-based approaches, keypoint-based methods can achieve significantly higher efficiency due to their sparsity. However, sparsity also results in information loss, thereby reducing performance. In this work, we propose a novel, keypoint-based Contour-Pose representation, which compactly encodes both body shape and parts information. We further propose a local-to-global architecture, called GaitContour, to leverage this novel representation and efficiently compute subject embedding in two stages. The first stage consists of a local transformer that extracts features from five different body regions. The second stage then aggregates the regional features to estimate a global human gait representation. Such a design significantly reduces the complexity of the attention operation and improves both efficiency and performance. Through large scale experiments, Gait-Contour is shown to perform significantly better than previous keypoint-based methods. Furthermore, the ContourPose representation also achieves new SoTA performances on fusion-based gait recognition methods.
Yuxiang Guo 0001, Anshul Shah 0001, Jiang Liu 0014, Ayush Gupta 0001, Rama Chellappa, Cheng Peng 0008
WACV2
2025 Cap2Aug: Caption Guided Image data Augmentation
abstract
Visual recognition in a low-data regime is challenging and often prone to overfitting. To mitigate this issue, several data augmentation strategies have been proposed. However, standard transformations, e.g., rotation, cropping, and flip-ping provide limited semantic variations. To this end, we propose Cap2Aug, an image-to-image diffusion model-based data augmentation strategy using image captions to condition the image synthesis step. We generate a caption for an image and use this caption as an additional input for an image-to-image diffusion model. This increases the semantic diversity of the augmented images due to caption conditioning compared to the usual data augmentation techniques. We show that Cap2Aug is particularly effective where only a few samples are available for an object class. However, naively generating the synthetic images is not adequate due to the domain gap between real and synthetic images. Thus, we employ a maximum mean discrepancy loss to align the synthetic images to the real images to minimize the domain gap. We evaluate our method on few-shot classification and image classification with long-tail class distribution tasks. Cap2Aug achieves state-of-the-art performance on both tasks while evaluated on eleven benchmarks. Code: https://github.com/aniket004/Cap_2_Aug.git
Aniket Roy, Anshul Shah 0001, Ketul Shah, Rama Chellappa
WACV2
2023 HaLP: Hallucinating Latent Positives for Skeleton-based Self-Supervised Learning of Actions
abstract
Supervised learning of skeleton sequence encoders for action recognition has received significant attention in recent times. However, learning such encoders without labels continues to be a challenging problem. While prior works have shown promising results by applying contrastive learning to pose sequences, the quality of the learned representations is often observed to be closely tied to data augmentations that are used to craft the positives. However, augmenting pose sequences is a difficult task as the geometric constraints among the skeleton joints need to be enforced to make the augmentations realistic for that action. In this work, we propose a new contrastive learning approach to train models for skeleton-based action recognition without labels. Our key contribution is a simple module, HaLP - to Hallucinate Latent Positives for contrastive learning. Specifically, HaLP explores the latent space of poses in suitable directions to generate new positives. To this end, we present a novel optimization formulation to solve for the synthetic positives with an explicit control on their hardness. We propose approximations to the objective, making them solvable in closed form with minimal overhead. We show via experiments that using these generated positives within a standard contrastive learning framework leads to consistent improvements across benchmarks such as NTU-60, NTU-120, and PKU-II on tasks like linear evaluation, transfer learning, and kNN evaluation. Our code can be found at https://github.com/anshulbshah/HaLP.
Anshul Shah 0001, Aniket Roy, Ketul Shah, Shlok Kumar Mishra, David Jacobs 0001, Anoop Cherian, Rama Chellappa
CVPR1
2023 STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural Videos
abstract
We address the problem of extracting key steps from un-labeled procedural videos, motivated by the potential of Augmented Reality (AR) headsets to revolutionize job training and performance. We decompose the problem into two steps: representation learning and key steps extraction. We propose a training objective, Bootstrapped Multi-Cue Contrastive (BMC2) loss to learn discriminative representations for various steps without any labels. Different from prior works, we develop techniques to train a light-weight temporal module which uses off-the-shelf features for self supervision. Our approach can seamlessly leverage information from multiple cues like optical flow, depth or gaze to learn discriminative features for key-steps, making it amenable for AR applications. We finally extract key steps via a tunable algorithm that clusters the representations and samples. We show significant improvements over prior works for the task of key step localization and phase classification. Qualitative results demonstrate that the extracted key steps are meaningful and succinctly represent various steps of the procedural tasks. Our code can be found at https://github.com/anshulbshah/STEPs.
Anshul Shah 0001, Ben Lundell, Harpreet Sawhney, Rama Chellappa
ICCV1
2023 Multi-View Action Recognition using Contrastive Learning
abstract
In this work, we present a method for RGB-based action recognition using multi-view videos. We present a supervised contrastive learning framework to learn a feature embedding robust to changes in viewpoint, by effectively leveraging multi-view data. We use an improved supervised contrastive loss and augment the positives with those coming from synchronized viewpoints. We also propose a new approach to use classifier probabilities to guide the selection of hard negatives in the contrastive loss, to learn a more discriminative representation. Negative samples from confusing classes based on posterior are weighted higher. We also show that our method leads to better domain generalization compared to the standard supervised training based on synthetic multi-view data. Extensive experiments on real (NTU-60, NTU-120, NUMA) and synthetic (RoCoG) data demonstrate the effectiveness of our approach.
Ketul Shah, Anshul Shah 0001, Chun Pong Lau 0001, Celso de Melo, Rama Chellappa
WACV2
2022 Max-Margin Contrastive Learning
abstract
Standard contrastive learning approaches usually require a large number of negatives for effective unsupervised learning and often exhibit slow convergence. We suspect this behavior is due to the suboptimal selection of negatives used for offering contrast to the positives. We counter this difficulty by taking inspiration from support vector machines (SVMs) to present max-margin contrastive learning (MMCL). Our approach selects negatives as the sparse support vectors obtained via a quadratic optimization problem, and contrastiveness is enforced by maximizing the decision margin. As SVM optimization can be computationally demanding, especially in an end-to-end setting, we present simplifications that alleviate the computational burden. We validate our approach on standard vision benchmark datasets, demonstrating better performance in unsupervised representation learning over state-of-the-art, while having better empirical convergence properties.
Anshul Shah 0001, Suvrit Sra, Rama Chellappa, Anoop Cherian
AAAI1
2022 Learning visual representations for transfer learning by suppressing texture
Shlok Kumar Mishra, Anshul Shah 0001, Ankan Bansal, Janit Anjaria, Abhinav Shrivastava, Abhishek Sharma 0001, David Jacobs 0001
BMVC2
2022 FeLMi : Few shot Learning with hard Mixup
abstract
Learning from a few examples is a challenging computer vision task. Traditionally,meta-learning-based methods have shown promise towards solving this problem.Recent approaches show benefits by learning a feature extractor on the abundantbase examples and transferring these to the fewer novel examples. However, thefinetuning stage is often prone to overfitting due to the small size of the noveldataset. To this end, we propose Few shot Learning with hard Mixup (FeLMi)using manifold mixup to synthetically generate samples that helps in mitigatingthe data scarcity issue. Different from a naïve mixup, our approach selects the hardmixup samples using an uncertainty-based criteria. To the best of our knowledge,we are the first to use hard-mixup for the few-shot learning problem. Our approachallows better use of the pseudo-labeled base examples through base-novel mixupand entropy-based filtering. We evaluate our approach on several common few-shotbenchmarks - FC-100, CIFAR-FS, miniImageNet and tieredImageNet and obtainimprovements in both 1-shot and 5-shot settings. Additionally, we experimented onthe cross-domain few-shot setting (miniImageNet → CUB) and obtain significantimprovements.
Aniket Roy, Anshul Shah 0001, Ketul Shah, Prithviraj Dhar, Anoop Cherian, Rama Chellappa
NeurIPS2
2022 Pose and Joint-Aware Action Recognition
abstract
Recent progress on action recognition has mainly focused on RGB and optical flow features. In this paper, we approach the problem of joint-based action recognition. Unlike other modalities, constellation of joints and their motion generate models with succinct human motion information for activity recognition. We present a new model for joint-based action recognition, which first extracts motion features from each joint separately through a shared motion encoder before performing collective reasoning. Our joint selector module re-weights the joint information to select the most discriminative joints for the task. We also propose a novel joint-contrastive loss that pulls together groups of joint features which convey the same action. We strengthen the joint-based representations by using a geometry-aware data augmentation technique which jitters pose heatmaps while retaining the dynamics of the action. We show large improvements over the current state-of-the-art joint-based approaches on JHMDB, HMDB, Charades, AVA action recognition datasets. A late fusion with RGB and Flow-based approaches yields additional improvements. Our model also outperforms the existing baseline on Mimetics, a dataset with out-of-context actions.
Anshul Shah 0001, Shlok Kumar Mishra, Ankan Bansal, Jun-Cheng Chen, Rama Chellappa, Abhinav Shrivastava
WACV1
2021 The CS1 Reviewer App: Choose Your Own Adventure or Choose for Me!
abstract
We present the CS1 Reviewer App - an online tool for an introductory Python course that allows students to solve customized problem sets on many concepts in the course. Currently, the app's questions focus on code tracing by presenting a block of Python code and asking students to predict the output of the code. The tool tracks a student's response history to maintain a "mastery level" that represents a student's knowledge of a concept. We also provide an option of answering auto-generated quizzes based on the student's mastery across concepts. As a result, the tool provides students a choice between creating their own learning experience or leveraging our question selection algorithm. The app is supported on traditional webpages and mobile devices, providing a convenient way for students to study a variety of concepts. Students in the CS1 course at Duke University used this tool during the Spring and Fall 2020 semesters. In this paper, we explore trends in usage, feedback and suggestions from students, and avenues of future work based on student experiences.
Anshul Shah 0001, Jonathan Liu, Kristin Stephens-Martinez, Susan H. Rodger
ITiCSE (1)1
2019 Bringing Alive Blurred Moments
abstract
We present a solution for the goal of extracting a video from a single motion blurred image to sequentially reconstruct the clear views of a scene as beheld by the camera during the time of exposure. We first learn motion representation from sharp videos in an unsupervised manner through training of a convolutional recurrent video autoencoder network that performs a surrogate task of video reconstruction. Once trained, it is employed for guided training of a motion encoder for blurred images. This network extracts embedded motion information from the blurred image to generate a sharp video in conjunction with the trained recurrent video decoder. As an intermediate step, we also design an efficient architecture that enables real-time single image deblurring and outperforms competing methods across all factors: accuracy, speed, and compactness. Experiments on real scenes and standard datasets demonstrate the superiority of our framework over the state-of-the-art and its ability to generate a plausible sequence of temporally consistent sharp frames.
Kuldeep Purohit, Anshul Shah 0001, A. N. Rajagopalan 0001
CVPR2
2018 Learning Based Single Image Blur Detection and Segmentation
abstract
This paper addresses the problem of obtaining a blur-based segmentation map from a single image affected by motion or defocus blur. Since traditional hand-designed priors have fundamental limitations, we utilise deep neural networks to learn features related to blur and enable a pixel-level blur classification. Our approach mitigates the ambiguities present in blur detection task by introducing joint learning of global context and local features into the framework. Specifically, we train two sub-networks to perform the task at global (image) and local (patch) levels. We aggregate the pixel-level probabilities estimated by two networks and feed them to a MRF based framework which returns a refined and dense segmentation-map of the image with respect to blur. We also demonstrate via both qualitative and quantitative evaluation, that our approach performs favorably against state-of-the-art blur detection or segmentation works, and show its utility to applications of automatic image matting and blur magnification.
Kuldeep Purohit, Anshul Shah 0001, A. N. Rajagopalan 0001
ICIP2