VLDB 2026 Research / reviewers in the wild / expert
Shubh Maheshwari
dblp:210/5261
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Computer animation and physical simulation · 67% Geometric modeling and processing · 33% | |
| Artificial intelligence
1 paper |
Video understanding and tracking · 77% 3D vision · 23% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Geometric modeling and processing › shape deformation
deformation transfer |
0.7 | 1 | 2023 | Transfer4D: A Framework for Frugal Motion Capture and Deformation Transfer · CVPR 2023 |
Computer animation and physical simulation
motion capture |
0.7 | 1 | 2023 | Transfer4D: A Framework for Frugal Motion Capture and Deformation Transfer · CVPR 2023 |
Computer animation and physical simulation
motion retargeting |
0.7 | 1 | 2023 | Transfer4D: A Framework for Frugal Motion Capture and Deformation Transfer · CVPR 2023 |
Computer vision › Video understanding and tracking
action recognition |
0.5 | 1 | 2021 | Quo Vadis, Skeleton Action Recognition? · Int. J. Comput. Vis. 2021 |
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition |
0.5 | 1 | 2021 | Quo Vadis, Skeleton Action Recognition? · Int. J. Comput. Vis. 2021 |
Computer vision › 3D vision
pose estimation |
0.1 | 1 | 2021 | Quo Vadis, Skeleton Action Recognition? · Int. J. Comput. Vis. 2021 |
Computer vision › 3D vision › 3d shape representation
skeleton representation |
0.1 | 1 | 2021 | Quo Vadis, Skeleton Action Recognition? · Int. J. Comput. Vis. 2021 |
Methods — techniques the papers use, named apart from their topics
skinning decomposition · 0.7skeleton extraction · 0.7non-rigid reconstruction · 0.7survey · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Morag - Multi-Fusion Retrieval Augmented Generation for Human MotionabstractWe introduce MoRAG, a novel multi-part fusion based retrieval-augmented generation strategy for text-based human motion generation. The method enhances motion diffusion models by leveraging additional knowledge obtained through an improved motion retrieval process. By effectively prompting large language models (LLMs), we address spelling errors and rephrasing issues in motion retrieval. Our approach utilizes a multi-part retrieval strategy to improve the generalizability of motion retrieval across the language space. We create diverse samples through the spatial composition of the retrieved motions. Furthermore, by utilizing low-level, part-specific motion information, we can construct motion samples for unseen text descriptions. Our experiments demonstrate that our framework can serve as a plug-and-play module, improving the performance of motion diffusion models. Code, pre-trained models, and sample videos are available at motion-rag. github.io. Sai Shashank Kalakonda, Shubh Maheshwari, Ravi Kiran Sarvadevabhatla |
WACV | 2 |
| 2023 | Transfer4D: A Framework for Frugal Motion Capture and Deformation TransferabstractAnimating a virtual character based on a real performance of an actor is a challenging task that currently requires expensive motion capture setups and additional effort by expert animators, rendering it accessible only to large production houses. The goal of our work is to democratize this task by developing a frugal alternative termed “Transfer4D” that uses only commodity depth sensors and further reduces animators' effort by automating the rigging and animation transfer process. Our approach can transfer motion from an incomplete, single-view depth video to a semantically similar target mesh, unlike prior works that make a stricter assumption on the source to be noise-free and watertight. To handle sparse, incomplete videos from depth video inputs and variations between source and target objects, we propose to use skeletons as an intermediary representation between motion capture and transfer. We propose a novel unsupervised skeleton extraction pipeline from a single-view depth sequence that incorporates additional geometric information, resulting in superior performance in motion reconstruction and transfer in comparison to the contemporary methods and making our approach generic. We use non-rigid reconstruction to track motion from the depth sequence, and then we rig the source object using skinning decomposition. Finally, the rig is embedded into the target object for motion retargeting. Shubh Maheshwari, Rahul Narain, Ramya Hebbalaguppe |
CVPR | 1 |
| 2023 | Action-GPT: Leveraging Large-scale Language Models for Improved and Generalized Action GenerationabstractWe introduce Action-GPT, a plug-and-play framework for incorporating Large Language Models (LLMs) into text-based action generation models. Action phrases in current motion capture datasets contain minimal and to-the-point information. By carefully crafting prompts for LLMs, we generate richer and fine-grained descriptions of the action. We show that utilizing these detailed descriptions instead of the original action phrases leads to better alignment of text and motion spaces. We introduce a generic approach compatible with stochastic (e.g. VAE-based) and deterministic (e.g. MotionCLIP) text-to-motion models. In addition, the approach enables multiple text descriptions to be utilized. Our experiments show (i) noticeable qualitative and quantitative improvement in the quality of synthesized motions, (ii) benefits of utilizing multiple LLM-generated descriptions, (iii) suitability of the prompt function, and (iv) zero-shot generation capabilities of the proposed approach. Code and pretrained models are available at https://actiongpt.github.io. Sai Shashank Kalakonda, Shubh Maheshwari, Ravi Kiran Sarvadevabhatla |
ICME | 2 |
| 2023 | DSAG: A Scalable Deep Framework for Action-Conditioned Multi-Actor Full Body Motion SynthesisabstractWe introduce DSAG, a controllable deep neural framework for action-conditioned generation of full body multiactor variable duration actions. To compensate for incompletely detailed finger joints in existing large-scale datasets, we introduce full body dataset variants with detailed finger joints. To overcome shortcomings in existing generative approaches, we introduce dedicated representations for encoding finger joints. We also introduce novel spatiotemporal transformation blocks with multi-head self attention and specialized temporal processing. The design choices enable generations for a large range in body joint counts (24 - 52), frame rates (13 - 50), global body movement (inplace, locomotion) and action categories (12 - 120), across multiple datasets (NTU-120, HumanAct12, UESTC, Human3.6M). Our experimental results demonstrate DSAG’s significant improvements over state-of-the-art, its suitability for action-conditioned generation at scale. Debtanu Gupta, Shubh Maheshwari, Sai Shashank Kalakonda, Manasvi Vaidyula, Ravi Kiran Sarvadevabhatla |
WACV | 2 |
| 2022 | MUGL: Large Scale Multi Person Conditional Action Generation with LocomotionabstractWe introduce MUGL, a novel deep neural model for large-scale, diverse generation of single and multi-person pose-based action sequences with locomotion. Our controllable approach enables variable-length generations customizable by action category, across more than 100 categories. To enable intra/inter-category diversity, we model the latent generative space using a Conditional Gaussian Mixture Variational Autoencoder. To enable realistic generation of actions involving locomotion, we decouple local pose and global trajectory components of the action sequence. We incorporate duration-aware feature representations to enable variable-length sequence generation. We use a hybrid pose sequence representation with 3D pose sequences sourced from videos and 3D Kinect-based sequences of NTU-RGBD-120. To enable principled comparison of generation quality, we employ suitably modified strong baselines during evaluation. Although smaller and simpler compared to baselines, MUGL provides better quality generations, paving the way for practical and controllable large-scale human action generation. Shubh Maheshwari, Debtanu Gupta, Ravi Kiran Sarvadevabhatla |
WACV | 1 |
| 2021 | Quo Vadis, Skeleton Action Recognition?
Pranay Gupta, Anirudh Thatipelli, Aditya Aggarwal, Shubh Maheshwari, Neel Trivedi, Ravi Kiran Sarvadevabhatla |
Int. J. Comput. Vis. | 4 |