EDBT 2026 Demo / reviewers in the wild / expert
Avinash Sharma 0001
dblp:08/4811-1
· DBLP profile ↗
25ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0001-5013-5024ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 8 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NGD: Neural Gradient Based Deformation for Monocular Garment Reconstruction
Soham Dasgupta, Shanthika Naik, Preet Savalia, Sujay Kumar Ingle, Avinash Sharma 0001 |
ICCV | 5 |
| 2024 | Learning Based Infinite Terrain Generation with Level of DetailingabstractInfinite terrain generation is an important use case for computer graphics, games and simulations. However, current techniques are often procedural which reduces their realism. We introduce a learning-based generative framework for infinite terrain generation along with a novel learning-based approach for level-of-detailing of terrains. Our framework seamlessly integrates with quad-tree-based terrain rendering algorithms. Our approach leverages image completion techniques for infinite generation and progressive super-resolution for terrain enhancement. Notably, we propose a novel quad-tree-based training method for terrain enhancement which enables seamless integration with quad-tree-based rendering algorithms while minimizing the errors along the edges of the enhanced terrain. Comparative evaluations against existing techniques demonstrate our framework’s ability to generate highly realistic terrain with effective level-of-detailing. Aryamaan Jain, Avinash Sharma 0001, Krishnan Sundara Rajan |
3DV | 2 |
| 2024 | MANUS: Markerless Grasp Capture Using Articulated 3D GaussiansabstractUnderstanding how we grasp objects with our hands has important applications in areas like robotics and mixed re-ality. However, this challenging problem requires accurate modeling of the contact between hands and objects. To capture grasps, existing methods use skeletons, meshes, or parametric models that does not represent hand shape accu-rately resulting in inaccurate contacts. We present MANUS, a method for Markerless Hand-Object Grasp Capture using Articulated 3D Gaussians. We build a novel articulated 3D Gaussians representation that extends 3D Gaussian splatting [29] for high-fidelity representation of articulating hands. Since our representation uses Gaussian primitives optimized from the multi-view pixel-aligned losses, it enables us to efficiently and accurately estimate contacts between the hand and the object. For the most accurate results, our method requires tens of camera views that current datasets do not provide. We therefore build MANUS-Grasps, a new dataset that contains hand-object grasps viewed from 50+ cameras across 30+ scenes, 3 subjects, and comprising over 7M frames. In addition to extensive qualitative results, we also show that our method outper-forms others on a quantitative contact evaluation method that uses paint transfer from the object to the hand. Chandradeep Pokhariya, Ishaan Nikhil Shah, Angela Xing, Zekun Li 0002, Avinash Sharma 0001, Srinath Sridhar 0002 |
CVPR | 6 |
| 2024 | WordRobe: Text-Guided Generation of Textured 3D Garments
Astitva Srivastava, Pranav Manu, Amit Raj, Varun Jampani, Avinash Sharma 0001 |
ECCV (1) | 5 |
| 2024 | ATPPNet: Attention based Temporal Point cloud Prediction NetworkabstractPoint cloud prediction is an important yet challenging task in the field of autonomous driving. The goal is to predict future point cloud sequences that maintain object structures while accurately representing their temporal motion. These predicted point clouds help in other subsequent tasks like object trajectory estimation for collision avoidance or estimating locations with the least odometry drift. In this work, we present ATPPNet, a novel architecture that predicts future point cloud sequences given a sequence of previous time step point clouds obtained with LiDAR sensor. ATPPNet leverages Conv-LSTM along with channel-wise and spatial attention dually complemented by a 3D-CNN branch for extracting an enhanced spatio-temporal context to recover high quality fidel predictions of future point clouds. We conduct extensive experiments on publicly available datasets and report impressive performance outperforming the existing methods. We also conduct a thorough ablative study of the proposed architecture and provide an application study that highlights the potential of our model for tasks like odometry estimation. Kaustab Pal, Aditya Sharma 0001, Avinash Sharma 0001, K. Madhava Krishna |
ICRA | 3 |
| 2023 | SHARP: Shape-Aware Reconstruction of People in Loose Clothing
Sai Sagar Jinka, Astitva Srivastava, Chandradeep Pokhariya, Avinash Sharma 0001, P. J. Narayanan |
Int. J. Comput. Vis. | 4 |
| 2022 | Deep Generative Framework for Interactive 3D Terrain Authoring and ManipulationabstractAutomated generation and (user) authoring of realistic virtual terrain is most sought for by the multimedia applications like VR models and gaming. The most common representation adopted for terrain is Digital Elevation Model (DEM). In this paper, we propose a novel realistic terrain authoring framework powered by a combination of VAE and generative conditional GAN model. Our framework is an example-based method that attempts to overcome the limitations of existing methods by learning a latent space from a real-world terrain dataset. This latent space allows us to generate multiple variants of terrain from a single input as well as interpolate between terrains while keeping the generated terrains close to real-world data distribution. We also developed an interactive tool that lets the user generate diverse terrains with minimal inputs. We perform a thorough qualitative and quantitative analysis and provide a comparison with other SOTA methods. Shanthika Naik, Aryamaan Jain, Avinash Sharma 0001, Krishnan Sundara Rajan |
IGARSS | 3 |
| 2022 | Multiple Kernel Learning for Modeling Resting State EEG Connectomes using Structural Connectivity of the BrainabstractAn active area of research in cognitive science is characterizing the relationship between brain structure and the observed functional activations. Recent graph diffusion models have had great success in mapping whole-brain, resting-state dynamics measured using functional Magnetic Resonance Imaging (fMRI) to the brain structure derived using diffusion and T1 brain imaging. Here we test the application of one such graph diffusion method called the Multiple Kernel Learning (MKL) model. MKL model, formulated as a reaction-diffusion system using Wilson-Cowan equations, combines multiple diffusion kernels at different scales to predict functional connectome (FC) arising from a fixed structural connectome (SC). Our simulation results demonstrate that the MKL model successfully mapped the relationship between SC and FC from five different Electroen-cephalogram (EEG) bands (delta, theta, alpha, beta, and gamma). We used simultaneously acquired EEG-fMRI and NODDI dataset of 17 participants. The correlation between predicted FC and ground truth FC was higher for EEG bands than for fMRI data. The prediction accuracy peaked for the alpha band, and the highest frequency band, gamma had the lowest prediction accuracy. To the best of our knowledge, this is the first such end-to-end application of multiple kernel graph diffusion framework for modeling EEG data. One of the important features of MKL model is its ability to incorporate structural connectivity features into the generative model that predicts the EEG functional connectivity. P. L. Ammar Ahmed, Archi Yadav, Avinash Sharma 0001, Raju S. Bapi |
IJCNN | 3 |
| 2022 | Multiple GraphHeat Networks for Structural to Functional Brain MappingabstractOver the last decade, there has been growing interest in learning the mapping from structural connectivity (SC) to functional connectivity (FC) of the brain. The spontaneous brain activity fluctuations during the resting-state as captured by functional MRI (rsfMRI) contain rich non-stationary dynamics over a relatively fixed structural connectome. Among the modeling approaches, graph diffusion-based methods with single and multiple diffusion kernels approximating static or dynamic functional connectivity have shown promise in predicting the FC given the SC. However, these methods are computationally expensive, not scalable, and fail to capture the complex dynamics underlying the whole process. Recently, deep learning methods such as GraphHeat networks along with graph diffusion have been shown to handle complex relational structures while preserving global information. In this paper, we propose multiple GraphHeat networks (M-GHN), a novel approach for mapping SC-FC. M-GHN enables us to model multiple heat kernel diffusion over the brain graph for approximating the complex Reaction Diffusion phenomenon. We argue that the proposed deep learning method overcomes the scalability and computational inefficiency issues but can still learn the SC-FC mapping successfully. Training and testing were done using the rsfMRI data of 100 participants from the human connectome project (HCP), and the results establish the viability of the proposed model. On the HCP dataset of 100 participants, the M-GHN achieves a high Pearson correlation of 0.747. Furthermore, experiments demonstrate that M-GHN outperforms the existing methods in learning the complex nature of human brain function. Subba Reddy Oota, Archi Yadav, Arpita Dash, Raju S. Bapi, Avinash Sharma 0001 |
IJCNN | 5 |
| 2022 | xCloth: Extracting Template-free Textured 3D Clothes from a Monocular ImageabstractExisting approaches for 3D garment reconstruction either assume a predefined template for the garment geometry (restricting them to fixed clothing styles) or yield vertex-colored meshes (lacking high-frequency textural details). Our novel framework co-learns geometric and semantic information of garment surface from the input monocular image for template-free textured 3D garment digitization. More specifically, we propose to extend PeeledHuman representation to predict the pixel-aligned, layered depth and semantic maps to extract 3D garments. The layered representation is further exploited to UV parametrize the arbitrary surface of the extracted garment without any human intervention to form a UV atlas. The texture is then imparted on the UV atlas in a hybrid fashion by first projecting pixels from the input image to UV space for the visible region, followed by inpainting the occluded regions. Thus, we are able to digitize arbitrarily loose clothing styles while retaining high-frequency textural details from a monocular image. We achieve high-fidelity 3D garment reconstruction results on three publicly available datasets and generalization on internet images. Astitva Srivastava, Chandradeep Pokhariya, Sai Sagar Jinka, Avinash Sharma 0001 |
ACM Multimedia | 4 |
| 2022 | Robust 3D Garment Digitization from Monocular 2D Images for 3D Virtual Try-On SystemsabstractIn this paper, we develop a robust 3D garment digitization solution that can generalize well on real-world fashion catalog images with cloth texture occlusions and large body pose variations. We assumed fixed topology parametric template mesh models for known types of garments (e.g., T-shirts, Trousers) and perform mapping of high-quality texture from an input catalog image to UV map panels corresponding to the parametric mesh model of the garment. We achieve this by first predicting a sparse set of 2D landmarks on the boundary of the garments. Subsequently, we use these landmarks to perform Thin-Plate-Spline-based texture transfer on UV map panels. Subsequently, we employ a deep texture inpainting network to fill the large holes (due to view variations & self-occlusions) in TPS output to generate consistent UV maps. Furthermore, to train the supervised deep networks for landmark prediction & texture inpainting tasks, we generated a large set of synthetic data with varying texture and lighting imaged from various views with the human present in a wide variety of poses. Additionally, we manually annotated a small set of fashion catalog images crawled from online fashion e-commerce platforms to finetune. We conduct thorough empirical evaluations and show impressive qualitative results of our proposed 3D garment texture solution on fashion catalog images. Such 3D garment digitization helps us solve the challenging task of enabling 3D Virtual Try-on. Sahib Majithia, Sandeep N. Parameswaran, Sadbhavana Babar, Vikram Garg, Astitva Srivastava, Avinash Sharma 0001 |
WACV | 6 |
| 2021 | GlocalNet: Class-aware Long-term Human Motion SynthesisabstractSynthesis of long-term human motion skeleton sequences is essential to aid human-centric video generation [8] with potential applications in Augmented Reality, 3D character animations, pedestrian trajectory prediction, etc. Long-term human motion synthesis is a challenging task due to multiple factors like, long-term temporal dependencies among poses, cyclic repetition across poses, bi-directional and multi-scale dependencies among poses, variable speed of actions, and a large as well as partially overlapping space of temporal pose variations across multiple class/types of human activities. This paper aims to address these challenges to synthesize a long-term (> 6000 ms) human motion trajectory across a large variety of human activity classes (> 50). We propose a two-stage activity generation method to achieve this goal, where the first stage deals with learning the long-term global pose dependencies in activity sequences by learning to synthesize a sparse motion trajectory while the second stage addresses the generation of dense motion trajectories taking the output of the first stage. We demonstrate the superiority of the proposed method over SOTA methods using various quantitative evaluation metrics on publicly available datasets. Neeraj Battan, Yudhik Agrawal, Sai Soorya Rao, Aman Goel, Avinash Sharma 0001 |
WACV | 5 |
| 2021 | Action Quality Assessment Using Siamese Network-Based Deep Metric LearningabstractAutomated vision-based score estimation models can be used to provide an alternate opinion to avoid judgment bias. Existing works have learned score estimation models by regressing the video representation to ground truth score provided by judges. However, such regression-based solutions lack interpretability in terms of giving reasons for the awarded score. One solution to make the scores more explicable is to compare the given action video with a reference video, which would capture the temporal variations vis-á-vis the reference video and map those variations to the final score. In this work, we propose a new action scoring system termed as Reference Guided Regression (RGR), which comprises (1) a Deep Metric Learning Module that learns similarity between any two action videos based on their ground truth scores given by the judges, and (2) a Score Estimation Module that uses the first module to find the resemblance of a video with a reference video to give the assessment score. The proposed scoring model is tested for Olympics Diving and Gymnastic vaults and the model outperforms the existing state-of-the-art scoring models. Hiteshi Jain, Gaurav Harit, Avinash Sharma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | PeeledHuman: Robust Shape Representation for Textured 3D Human Body ReconstructionabstractWe introduce PeeledHuman - a novel shape representation of the human body that is robust to self-occlusions. PeeledHuman encodes the human body as a set of Peeled Depth and RGB maps in 2D, obtained by performing raytracing on the 3D body model and extending each ray beyond its first intersection. This formulation allows us to handle self-occlusions efficiently compared to other representations. Given a monocular RGB image, we learn these Peeled maps in an end-to-end generative adversarial fashion using our novel framework - PeelGAN. We train PeelGAN using a 3D Chamfer loss and other 2D losses to generate multiple depth values per-pixel and a corresponding RGB field per-vertex in a dual-branch setup. In our simple non-parametric solution, the generated Peeled Depth maps are back-projected to 3D space to obtain a complete textured 3D shape. The corresponding RGB maps provide vertex-level texture details. We compare our method with current parametric and non-parametric methods in 3D reconstruction and find that we achieve state-of-the-art-results. We demonstrate the effectiveness of our representation on publicly available BUFF and MonoPerfCap datasets as well as loose clothing data collected by our calibrated multi-Kinect setup. Sai Sagar Jinka, Rohan Chacko, Avinash Sharma 0001, P. J. Narayanan |
3DV | 3 |
| 2020 | AFN: Attentional Feedback Network Based 3D Terrain Super-Resolution
Ashish Kubade, Diptiben Patel, Avinash Sharma 0001, Krishnan Sundara Rajan |
ACCV (1) | 3 |
| 2020 | Feedback Neural Network Based Super-Resolution of DEM for Generating High Fidelity FeaturesabstractHigh resolution Digital Elevation Models(DEMs) are an important requirement for many applications like modelling water flow, landslides, avalanches etc. Yet publicly available DEMs have low resolution for most parts of the world. Despite tremendous success in image super-resolution task using deep learning solutions, there are very few works that have used these powerful systems on DEMs to generate HRDEMs. Motivated from feedback neural networks, we propose a novel neural network architecture that learns to add high frequency details iteratively to low resolution DEM, turning it into a high resolution DEM without compromising its fidelity. Our experiments confirm that without any additional modality such as aerial images(RGB), our network DSRFB achieves RMSEs of 0.59 to 1.27 across 4 different terrains having diverse geographical structures. Ashish Kubade, Avinash Sharma 0001, Krishnan Sundara Rajan |
IGARSS | 2 |
| 2018 | Deep Textured 3D Reconstruction of Human Bodies
Abbhinav Venkat, Sai Sagar Jinka, Avinash Sharma 0001 |
BMVC | 3 |
| 2018 | Towards View-Invariant Intersection Recognition from Videos using Deep Network EnsemblesabstractThis paper strives to answer the following question: Is it possible to recognize an intersection when seen from different road segments that constitute the intersection? An intersection or a junction typically is a meeting point of three or four road segments. Its recognition from a road segment that is transverse to or 180 degrees apart from its previous sighting is an extremely challenging and yet a very relevant problem to be addressed from the point of view of both autonomous driving as well as loop detection. This paper formulates this as a problem of video recognition and proposes a novel LSTM based Siamese style deep network for video recognition. For what is indeed a challenging problem and the limited annotated dataset available we show competitive results of recognizing intersections when approached from diverse viewpoints or road segments. Specifically, we tabulate effective recognition accuracy even as the approaches to the intersection being compared are disparate both in terms of viewpoints and weather/illumination conditions. We show competitive results on both synthetic yet highly realistic data mined from the gaming platform GTA as well as on real world data made available through Mapillary. Gunshi Gupta, Avinash Sharma 0001, K. Madhava Krishna |
IROS | 3 |
| 2018 | Fast Multi Model Motion Segmentation on Road ScenesabstractWe propose a novel motion clustering formulation over spatio-temporal depth images obtained from stereo sequences that segments multiple motion models in the scene in an unsupervised manner. The motion models are obtained at frame rates that compete with the speed of the stereo depth computation. This is possible due to a decoupling framework that first delineates spatial clusters and subsequently assigns motion labels to each of these cluster with analysis of a novel motion graph model. A principled computation of the weights of the motion graph that signifies the relative shear and stretch between possible clusters lends itself to a high fidelity segmentation of the motion models in the scene. The fidelity is vindicated through accuracies reaching 89.61% on KITTI and complex native sequences. Mahtab Sandhu, Nazrul Haque, Avinash Sharma 0001, K. Madhava Krishna, Shanti Medasani |
Intelligent Vehicles Symposium | 3 |
| 2017 | Multi-trajectory pose correspondences using scale-dependent topological analysis of pose-graphsabstractThis paper considers the problem of finding pose matches between trajectories of multiple robots in their respective coordinate frames or equivalent matches between trajectories obtained during different sessions. Pose correspondences between trajectories are mediated by common landmarks represented in a topological map lacking distinct metric coordinates. Despite such lack of explicit metric level associations, we mine preliminary pose level correspondences between trajectories through a novel multi-scale heat-kernel descriptor and correspondence graph framework. These serve as an improved initialization for ICP (Iterative Closest Point) to yield dense pose correspondences. We perform extensive analysis of the proposed method under varying levels of pose and landmark noise and showcase its superiority in obtaining pose matches in comparison with standard ICP like methods. To the best of our knowledge, this is the first work of the kind that brings in elements from spectral graph theory to solve the problem of pose correspondences in a multi-robotic setting and differentiates itself from other works. Sayantan Datta, Avinash Sharma 0001, K. Madhava Krishna |
IROS | 2 |
| 2016 | Image Annotation using Multi-scale Hypergraph Heat Diffusion FrameworkabstractThe task of automatic image annotation involves assigning relevant multiple labels/tags to query images based on their visual content. One of the key challenge in multi-label image annotation task is the class imbalance problem where frequently occurring labels suppress the participation of rarely occurring labels. In this paper, we propose to exploit the multi-scale behavior in hypergraph heat diffusion framework for the automatic image annotation task. The proposed novel technique enables to model the higher order relationship among images in the feature space and provides a multi-scale label diffusion mechanism to address the class imbalance problem in the data. Venkatesh N. Murthy, Avinash Sharma 0001, Visesh Chari, R. Manmatha |
ICMR | 2 |
| 2014 | RoadEye: A System for Personalized Retrieval of Dynamic Road ConditionsabstractAwareness of dynamically changing road conditions is crucial for a safe and quality driving experience, as well as, in augmenting trip planning. This work addresses the problem of keeping users informed in a timely and personalized manner about road conditions arising from both scheduled and ad hoc events. We propose Road Eye, a system for personalized retrieval of dynamic road conditions. The key contribution of Road Eye is the psi R-tree, which is a novel R-tree-based index augmented with linked lists for facilitating quick and personalized retrieval of user-queried road conditions. Our performance study indicates that the psi R-tree is indeed effective in retrieving dynamic road conditions with reduced query response times and disk I/Os. Anirban Mondal, Avinash Sharma 0001, Abhishek Tripathi, Atul Singh, Nischal M. Piratla |
MDM (1) | 2 |
| 2011 | Topologically-robust 3D shape matching based on diffusion geometry and seed growingabstract3D Shape matching is an important problem in computer vision. One of the major difficulties in finding dense correspondences between 3D shapes is related to the topological discrepancies that often arise due to complex kinematic motions. In this paper we propose a shape matching method that is robust to such changes in topology. The algorithm starts from a sparse set of seed matches and outputs dense matching. We propose to use a shape descriptor based on properties of the heat-kernel and which provides an intrinsic scale-space representation. This descriptor incorporates (i) heat-flow from already matched points and (ii) self diffusion. At small scales the descriptor behaves locally and hence it is robust to global changes in topology. Therefore, it can be used to build a vertex-to-vertex matching score conditioned by an initial correspondence set. This score is then used to iteratively add new correspondences based on a novel seed-growing method that iteratively propagates the seed correspondences to nearby vertices. The matching is farther densified via an EM-like method that explores the congruency between the two shape embeddings. Our method is compared with two recently proposed algorithms and we show that we can deal with substantial topological differences between the two shapes. Avinash Sharma 0001, Radu Horaud, Jan Cech, Edmond Boyer |
CVPR | 1 |
| 2010 | Learning Shape Segmentation Using Constrained Spectral Clustering and Probabilistic Label Transfer
Avinash Sharma 0001, Etienne von Lavante, Radu Horaud |
ECCV (5) | 1 |
| 2008 | Projected Texture for Object Classification
Avinash Sharma 0001, Anoop M. Namboodiri |
ECCV (3) | 1 |