VLDB 2026 Research / reviewers in the wild / expert
Shashikant Verma
dblp:261/7964
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | L3D-Pose: Lifting Pose for 3D Avatars from a Single Camera in the WildabstractWhile 2D pose estimation has advanced our ability to interpret body movements in animals and primates, it is limited by the lack of depth information, constraining its application range. 3D pose estimation provides a more comprehensive solution by incorporating spatial depth, yet creating extensive 3D pose datasets for animals is challenging due to their dynamic and unpredictable behaviours in natural settings. To address this, we propose a hybrid approach that utilizes rigged avatars and the pipeline to generate synthetic datasets to acquire the necessary 3D annotations for training. Our method introduces a simple attention-based MLP network for converting 2D poses to 3D, designed to be independent of the input image to ensure scalability for poses in natural environments. Additionally, we identify that existing anatomical keypoint detectors are insufficient for accurate pose retargeting onto arbitrary avatars. To overcome this, we present a lookup table based on a deep pose estimation method using a synthetic collection of diverse actions rigged avatars perform. Our experiments demonstrate the effectiveness and efficiency of this lookup table-based retargeting approach. Overall, we propose a comprehensive framework with systematically synthesized datasets for lifting poses from 2D to 3D and then utilize this to re-target motion from wild settings onto arbitrary avatars. The L3D-Pose dataset can be found at https://soumyaratnadebnath.github.io/L3D-Pose Soumyaratna Debnath, Harish Katti, Shashikant Verma, Shanmuganathan Raman |
ICASSP | 3 |
| 2025 | Darts: Deformable Animation Ready Templates for Clothing HumansabstractAccurate 3D modeling of humans and high-fidelity garments is crucial in computer vision and graphics, impacting gaming, virtual, and augmented reality applications. While recent data-driven approaches have progressed in estimating segregated geometries for clothed humans, they often struggle with the seamless integration required for physics-based simulations. We introduce Deformable Animation Ready Templates (DARTs) to address these challenges, which enhance template-based garment reconstruction. Our framework employs a robust feature-line regressor network to establish precise deformation constraints guided by input image characteristics. Additionally, we present a novel differentiable Constrained Rigid Deformation Layer (CRDL) that facilitates effective template deformation while preserving the essential geometry of the garment. Our experiments demonstrate that DARTs can generate templates for physics-based simulation, allowing for seamless garment animations influenced by dynamic environmental factors. With minor adjustments, our templates can accommodate various clothing categories, promoting diversity in animated garment modeling. Shashikant Verma, Shanmuganathan Raman |
ICIP | 1 |
| 2025 | GMOT-Mamba: Mamba-Based Model Prediction For Generic Multiple Object TrackingabstractWe introduce GMOT-Mamba, a novel Mamba-based model prediction framework for Generic Multiple Object Tracking (GMOT) in video sequences. Our approach features a Weighted Feature Pooling (WFP) layer, which processes encoded target states, and an innovative encoder-decoder architecture that leverages Vision-Mamba (ViM) to predict filter weights. We train our model on combinations of large-scale datasets to capture strong priors and discriminative features necessary for generic object tracking. Through extensive experiments and ablation studies, we demonstrate the effectiveness of our approach, showcasing its competitive performance against state-of-the-art GMOT methods while outperforming SOT methods in both accuracy and inference speed. Our findings underscore the potential of Mamba for enhancing model prediction in visual tracking applications. Shashikant Verma, Nicu Sebe, Shanmuganathan Raman |
ICIP | 1 |
| 2024 | SemFaceEdit: Semantic Face Editing on Generative Radiance Manifolds
Shashikant Verma, Shanmuganathan Raman |
ICPR (6) | 1 |
| 2024 | GraphFill: Deep Image Inpainting using GraphsabstractWe present a novel coarser-to-finer approach for deep graphical image inpainting that utilizes GraphFill, a graph neural network-based deep learning framework, and a lightweight generative baseline network. We construct a pyramidal graph for the input-masked image by reducing it into superpixels, each representing a node in the graph. The proposed pyramidal approach facilitates the transfer of global context from coarser to finer pyramid levels, enabling GraphFill to estimate plausible information for unknown node values in the graph. The estimated information is used to fill in the masked region, which a Refine Network then refines. Furthermore, we propose a resolution-robust pyramidal graph construction method, allowing for efficient inpainting of high-resolution images with relatively fewer computations. Our proposed GAN-based network is trained in adversarial settings on Places365 and CelebA-HQ datasets and demonstrates competitive performance compared to existing methods while using fewer learning parameters. We conduct thorough ablation studies to evaluate the effectiveness of each component in the GraphFill Network for improved performance. Our proposed lightweight model for image inpainting is efficient in real-world scenarios, as it can be easily deployed on mobile devices with limited resources. Shashikant Verma, Roopa Sheshadri, Shanmuganathan Raman |
WACV | 1 |
| 2022 | DMD-Net: Deep Mesh Denoising NetworkabstractWe present Deep Mesh Denoising Network (DMD-Net), an end-to-end deep learning framework, for solving the mesh denoising problem. DMD-Net consists of a Graph Convolutional Neural Network in which aggregation is performed in both the primal as well as the dual graph. This is realized in the form of an asymmetric two-stream network, which contains a primal-dual fusion block that enables communication between the primal-stream and the dual-stream. We develop a Feature Guided Transformer (FGT) paradigm, which consists of a feature extractor, a transformer, and a denoiser. The feature extractor estimates the local features, that guide the transformer to compute a transformation, which is applied to the noisy input mesh to obtain a useful intermediate representation. This is further processed by the denoiser to obtain the denoised mesh. Our network is trained on a large scale dataset of 3D objects. We perform exhaustive ablation studies to demonstrate that each component in our network is essential for obtaining the best performance. We show that our method obtains competitive or better results when compared with the state-of-the-art mesh denoising algorithms. We demonstrate that our method is robust to various kinds of noise. We observe that even in the presence of extremely high noise, our method achieves excellent performance. Aalok Gangopadhyay, Shashikant Verma, Shanmuganathan Raman |
ICPR | 2 |