Parag Chaudhuri

dblp:57/5346 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-1706-5703ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Learning Assisted Interactive Modelling with Rough Freehand 3D Sketch Strokes
abstract
Freehand 3D sketches are a great medium to ideate and create visual content. However, generating 3D models from such rough sketches remains an unsolved, non-trivial task. We present an end-to-end interactive framework for rapid, incremental modelling from sparse, irregular 3D sketches. At the core of our solution, is sketchTransformer, a twostaged transformer network architecture, that fits parametric surface patches to a set of sketch strokes. We devise a novel pseudo height field representation that enables the sketchTransformer to handle noise and sparseness in the input strokes. Our method interactively evolves the surface model while maintaining smooth joins between nearby patches. We implement two frontends for our framework, one on the desktop and another as a mobile AR application, to illustrate how our method complements a standard 3D modelling pipelines. Our framework robustly handles a large variety of input 3D strokes that competing methods cannot parse adequately.
Sukanya Bhattacharjee, Parag Chaudhuri
3DV2
2025 HandRT: Simultaneous Hand Shape and Appearance Reconstruction With Pose Tracking From Monocular RGB-D Video
abstract
We propose a method to reconstruct a personalized hand avatar, representing the user's hand shape and appearance, from a monocular RGB-D video of a hand performing unknown hand poses under unknown illumination. Our method, HandRT, jointly optimizes hand pose, shape, appearance, and lighting parameters using a physically-based shading model in a differentiable rendering framework incorporating Monte Carlo path tracing. HandRT extends our previous work, Intrinsic Hand Avatar, by relaxing the assumption of a known coarse hand pose and utilizing depth data in the optimization. Specifically, we introduce an articulated registration energy based on iterative closest point over the depth point cloud that enables reconstruction from unknown hand poses via tracking across frames. Thus, we can reconstruct the avatar from arbitrary poses with high accuracy without relying on an off-the-shelf 2D joint detector at each frame. Further, HandRT is capable of precisely tracking the reconstructed avatar from either RGB or RGB-D input. Our evaluation demonstrates that our method outperforms existing hand avatar reconstruction methods on all commonly used metrics while producing significantly accurate mesh compared to state-of-the-art hand mesh recovery methods by a large margin on public and our captured datasets.
Pratik Kalshetti, Parag Chaudhuri
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 SPRINT: Script-agnostic Structure Recognition in Tables
Dhruv Kudale, Badri Vishal Kasuba, Venkatapathy Subramanian, Parag Chaudhuri, Ganesh Ramakrishnan
ICDAR (5)4
2024 Intrinsic Hand Avatar: Illumination-aware Hand Appearance and Shape Reconstruction from Monocular RGB Video
abstract
Reconstructing a user-specific hand avatar is essential for a personalized experience in augmented and virtual reality systems. Current state-of-the-art avatar reconstruction methods use implicit representations to capture detailed geometry and appearance combined with neural rendering. However, these methods rely on a complicated multi-view setup, do not explicitly handle environment lighting leading to baked-in illumination and self-shadows, and require long hours for training. We present a method to reconstruct a hand avatar from a monocular RGB video of a user’s hand in arbitrary hand poses captured under real-world environment lighting. Specifically, our method jointly optimizes shape, appearance, and lighting parameters using a realistic shading model in a differentiable rendering framework incorporating Monte Carlo path tracing. Despite relying on physically-based rendering, our method can complete the reconstruction within minutes. In contrast to existing work, our method disentangles intrinsic properties of the underlying appearance and environment lighting, leading to realistic self-shadows. We compare our method with state-of-the-art hand avatar reconstruction methods and observe that it outperforms them on all commonly used metrics. We also evaluate our method on our captured dataset to emphasize its generalization capability. Finally, we demonstrate applications of our intrinsic hand avatar on novel pose synthesis and relighting. We plan to release our code to aid further research.
Pratik Kalshetti, Parag Chaudhuri
WACV2
2024 Textron: Weakly Supervised Multilingual Text Detection through Data Programming
abstract
Several recent deep learning (DL) based techniques perform considerably well on image-based multilingual text detection. However, their performance relies heavily on the availability and quality of training data. There are numerous types of page-level document images consisting of information in several modalities, languages, fonts, and layouts. This makes text detection a challenging problem in the field of computer vision (CV), especially for low-resource or handwritten languages. Furthermore, there is a scarcity of word-level labeled data for text detection, especially for multilingual settings and Indian scripts that incorporate both printed and handwritten text. Conventionally, Indian script text detection requires training a DL model on plenty of labeled data, but to the best of our knowledge, no relevant datasets are available. Manual annotation of such data requires a lot of time, effort, and expertise. In order to solve this problem, we propose Textron, a Data Programming-based approach, where users can plug various text detection methods into a weak supervision-based learning frame-work. One can view this approach to multilingual text detection as an ensemble of different CV-based techniques and DL approaches. Textron can leverage the predictions of DL models pre-trained on a significant amount of language data in conjunction with CV-based methods to improve text detection in other languages. We demonstrate that Textron can improve the detection performance for documents written in Indian languages, despite the absence of corresponding labeled data. Further, through extensive experimentation, we show improvement brought about by our approach over the current State-of-the-art (SOTA) models, especially for handwritten Devanagari text. Code and dataset has been made available at https://github.Com/IITB-LEAP-OCR/TEXTRON
Dhruv Kudale, Badri Vishal Kasuba, Venkatapathy Subramanian, Parag Chaudhuri, Ganesh Ramakrishnan
WACV4
2023 TACTFUL: A Framework for Targeted Active Learning for Document Analysis
Venkatapathy Subramanian, Sagar Poudel, Parag Chaudhuri, Ganesh Ramakrishnan
ICDAR (5)3
2023 Remeshing-free Graph-based Finite Element Method for Fracture Simulation
abstract
Fracture produces new mesh fragments that introduce additional degrees of freedom in the system dynamics. Existing finite element method (FEM) based solutions suffer from an explosion in computational cost as the system matrix size increases. We solve this problem by presenting a graph-based FEM model for fracture simulation that is remeshing-free and easily scales to high-resolution meshes. Our algorithm models fracture on the graph induced in a volumetric mesh with tetrahedral elements. We relabel the edges of the graph using a computed damage variable to initialize and propagate fracture. We prove that non-linear, hyper-elastic strain energy is expressible entirely in terms of the edge lengths of the induced graph. This allows us to reformulate the system dynamics for the relabeled graph without changing the size of system dynamics matrix and thus prevents the computational cost from blowing up. The fractured surface has to be reconstructed explicitly only for visualization purposes. We simulate standard laboratory experiments from structural mechanics and compare the results with corresponding real-world experiments. We fracture objects made of a variety of brittle and ductile materials, and show that our technique offers stability and speed that is unmatched in current literature.
Avirup Mandal, Parag Chaudhuri, Subhasis Chaudhuri
Comput. Graph. Forum2
2022 Deep Interactive Surface Creation from 3D Sketch Strokes
abstract
We present a deep neural framework that allows users to create surfaces from a stream of sparse 3D sketch strokes. Our network consists of a global surface estimation module followed by a local surface refinement. This facilitates in the incremental prediction of surfaces. Thus, our proposed method works with 3D sketch strokes and estimate a surface interactively in real time. We compare the proposed method with various state-of-the-art methods and show its efficacy for surface fitting. Further, we integrate our method into an existing Blender based 3D content creation pipeline to show its usefulness in 3D modelling.
Sukanya Bhattacharjee, Parag Chaudhuri
IJCAI2
2022 Simulating Fracture in Anisotropic Materials Containing Impurities
abstract
Fracture simulation of real-world materials is an exceptionally challenging problem due to complex material properties like anisotropic elasticity and the presence of material impurities. We present a graph-based finite element method to simulate dynamic fracture in anisotropic materials. We further enhance this model by developing a novel probabilistic damage mechanics for modelling materials with impurities using a random graph-based formulation. We demonstrate how this formulation can be used by artists for directing and controlling fracture. We simulate and render fractures for a diverse set of materials to demonstrate the potency and robustness of our methods.
Avirup Mandal, Parag Chaudhuri, Subhasis Chaudhuri
MIG2
2022 Local Scale Adaptation to Hand Shape Model for Accurate and Robust Hand Tracking
abstract
Abstract The accuracy of hand tracking algorithms depends on how closely the geometry of the mesh model resembles the user's hand shape. Most existing methods rely on a learned shape space model; however, this fails to generalize to unseen hand shapes with significant deviations from the training set. We introduce local scale adaptation to augment this data‐driven shape model and thus enable modeling hands of substantially different sizes. We also present a framework to calibrate our proposed hand shape model by registering it to depth data and achieve accurate and robust tracking. We demonstrate the capability of our proposed adaptive shape model over the most widely used existing hand model by registering it to subjects from different demographics. We also validate the accuracy and robustness of our tracking framework on challenging public hand datasets where we improve over state‐of‐the‐art methods. Our adaptive hand shape model and tracking framework offer a significant boost towards generalizing the accuracy of hand tracking.
Pratik Kalshetti, Parag Chaudhuri
Comput. Graph. Forum2
2020 A Survey on Sketch Based Content Creation: from the Desktop to Virtual and Augmented Reality
abstract
Abstract Sketching is one of the most natural ways for representing any object pictorially. It is however, challenging to convert sketches to 3D content that is suitable for various applications like movies, games and computer aided design. With the advent of more accessible Virtual Reality (VR) and Augmented Reality (AR) technologies, sketching can potentially become a more powerful yet easy‐to‐use modality for content creation. In this state‐of‐the‐art report, we aim to present a comprehensive overview of techniques related to sketch based content creation, both on the desktop and in VR/AR. We discuss various basic concepts related to static and dynamic content creation using sketches. We provide a structured review of various aspects of content creation including model generation, coloring and texturing, and finally animation. We try to motivate the advantages that VR/AR based sketching techniques and systems can offer into making sketch based content creation a more accessible and powerful mode of expression. We also discuss and highlight various unsolved challenges that current sketch based techniques face with the goal of encouraging future research in the domain.
Sukanya Bhattacharjee, Parag Chaudhuri
Comput. Graph. Forum2
2019 OCR On-the-Go: Robust End-to-End Systems for Reading License Plates & Street Signs
abstract
We work on the problem of recognizing license plates and street signs automatically in challenging conditions such as chaotic traffic. We leverage state-of-the-art text spotters to generate a large amount of noisy labeled training data. The data is filtered using a pattern derived from domain knowledge. We augment training and testing data with interpolated boxes and annotations that makes our training and testing robust. We further use synthetic data during training to increase the coverage of the training data. We train two different models for recognition. Our baseline is a conventional Convolution Neural Network (CNN) encoder followed by a Recurrent Neural Network (RNN) decoder. As our first contribution, we bypass the detection phase by augmenting the baseline with an Attention mechanism in the RNN decoder. Next, we build in the capability of training the model end-to-end on scenes containing license plates by incorporating inception based CNN encoder that makes the model robust to multiple scales. We achieve improvements of 3.75% at the sequence level, over the baseline model. We present the first results of using multi-headed attention models on text recognition in images and illustrate the advantages of using multiple-heads over a single head. We observe gains as large as 7.18% by incorporating multi-headed attention. We also experiment with multi-headed attention models on French Street Name Signs dataset (FSNS) and a new Indian Street dataset that we release for experiments. We observe that such models with multiple attention masks perform better than the model with single-headed attention on three different datasets with varying complexities. Our models outperform state-of-the-art methods on FSNS and IIIT-ILST Devanagari datasets by 1.1% and 8.19% respectively.
Rohit Saluja, Ayush Maheshwari, Ganesh Ramakrishnan, Parag Chaudhuri, Mark J. Carman
ICDAR4
2019 Sub-Word Embeddings for OCR Corrections in Highly Fusional Indic Languages
abstract
Texts in Indic Languages contain a large proportion of out-of-vocabulary (OOV) words due to frequent fusion using conjoining rules (of which there are around 4000 in Sanskrit). OCR errors further accentuate this complexity for the error correction systems. Variations of sub-word units such as n-grams, possibly encapsulating the context, can be extracted from the OCR text as well as the language text individually. Some of the sub-word units that are derived from the texts in such languages highly correlate to the word conjoining rules. Signals such as frequency values (on a corpus) associated with such sub-word units have been used previously with log-linear classifiers for detecting errors in Indic OCR texts. We explore two different encodings to capture such signals and augment the input to Long Short Term Memory (LSTM) based OCR correction models, that have proven useful in the past for jointly learning the language as well as OCR-specific confusions. The first type of encoding makes direct use of sub-word unit frequency values, derived from the training data. The formulation results in faster convergence and better accuracy values of the error correction model on four different languages with varying complexities. The second type of encoding makes use of trainable sub-word embeddings. We introduce a new procedure for training fastText embeddings on the sub-word units and further observe a large gain in F-Scores, as well as word-level accuracy values.
Rohit Saluja, Mayur Punjabi, Mark J. Carman, Ganesh Ramakrishnan, Parag Chaudhuri
ICDAR5
2017 Error Detection and Corrections in Indic OCR Using LSTMs
abstract
Conventional approaches to spell checking suggest spelling corrections using proximity-based matches to a known vocabulary. For highly inflectional Indian languages, any off-the-shelf vocabulary is significantly incomplete, since a large fraction of words in Indic documents are generated using word conjoining rules. Therefore, a tremendous manual effort is needed in spell-correcting words in Indic OCR documents. Moreover, in a spell checking system, a vocabulary may suggest multiple alternatives to the incorrect word. The ranking of these corrective suggestions is improved using language models. Owing to corpus resource scarcity, however, Indian languages lack reliable language models. Thus, learning the character (or n-gram) confusions or error patterns of the OCR system can be helpful in correcting the Out of Vocabulary (OOV) words in OCR documents. We adopt a Long Short-Term Memory (LSTM) based character level language model with a fixed delay for discriminative language modeling in the context of OCR errors for jointly addressing the problems of error detection and correction in Indic OCR. For words that need not be corrected in the OCR output, our model simply abstains from suggesting any changes. We present extensive results to validate the performance of our model on four Indian languages with different inflectional complexities. We achieve F-Scores above 92.4% and decreases in Word Error Rates (WER) of at least 26.7% across the four languages.
Rohit Saluja, Devaraj Adiga, Parag Chaudhuri, Ganesh Ramakrishnan, Mark J. Carman
ICDAR3
2013 Wetting of Porous Solids
abstract
This paper presents a simple, three stage method to simulate the mechanics of wetting of porous solid objects, like sponges and cloth, when they interact with a fluid. In the first stage, we model the absorption of fluid by the object when it comes in contact with the fluid. In the second stage, we model the transport of absorbed fluid inside the object, due to diffusion, as a flow in a deforming, unstructured mesh. The fluid diffuses within the object depending on saturation of its various parts and other body forces. Finally, in the third stage, oversaturated parts of the object shed extra fluid by dripping. The simulation model is motivated by the physics of imbibition of fluids into porous solids in the presence of gravity. It is phenomenologically capable of simulating wicking and imbibition, dripping, surface flows over wet media, material weakening, and volume expansion due to wetting. The model is inherently mass conserving and works for both thin 2D objects like cloth and for 3D volumetric objects like sponges. It is also designed to be computationally efficient and can be easily added to existing cloth, soft body, and fluid simulation pipelines.
Saket Patkar, Parag Chaudhuri
IEEE Trans. Vis. Comput. Graph.2
2011 Tracing specular light paths in point-based scenes
Rhushabh Goradia, Sriram Kashyap, Parag Chaudhuri, Sharat Chandran
Vis. Comput.3
2010 Real time ray tracing of point-based models
abstract
Mirroring the development of rendering algorithms for polygonal models, z-buffer style rendering for point-based models has given way recently to more advanced methods. A fast raycasting based approach [Wald and Seidel 2005] shows shadows, but does not demonstrate reflective effects. The more general raytracing approach [Linsen et al. 2007] is substantially slower.
Sriram Kashyap, Rhushabh Goradia, Parag Chaudhuri, Sharat Chandran
SI3D3
2009 Fast EMG-data driven skin deformation
abstract
Abstract Virtual characters are being modeled and animated with increasing accuracy and photorealism in current games and virtual reality simulations. However, the detailed modeling of dynamic skin deformations due to movement of the muscles during articulation is still prohibitively expensive for real‐time simulations. We present in this paper a fast approach for modeling such deformations driven by Electromyography (EMG) data. We demonstrate the method on the muscles of the upper leg during a gait cycle. A muscle map, created from a 3D muscle model, is applied as a deformation map to deform the skin surface as per the underlying muscle geometry. EMG signals recorded from a gait cycle are then used to dynamically vary the weights in the deformation map during animation thus producing the dynamic skin deformations. The whole operation is done on the GPU to make it very fast and suitable forreal‐time simulations. Copyright © 2009 John Wiley & Sons, Ltd.
Mustafa Kasap, Parag Chaudhuri, Nadia Magnenat-Thalmann
Comput. Animat. Virtual Worlds2
2008 Self adaptive animation based on user perspective
Parag Chaudhuri, George Papagiannakis, Nadia Magnenat-Thalmann
Vis. Comput.1
2007 Reusing view-dependent animation
Parag Chaudhuri, Prem Kumar Kalra, Subhashis Banerjee
Vis. Comput.1
2004 An Efficient Central Path Algorithm for Virtual Navigation
abstract
We give an efficient, scalable, and simple algorithm for computation of a central path for navigation in closed virtual environments. The algorithm requires less preprocessing and produces paths of high visual fidelity. The algorithm enables computing paths at multiple resolutions. The algorithm is based on a distance from boundary field computed on a hierarchical subdivision of the free space inside the closed 3D object. We also present a progressive version of our algorithm based on a local search strategy thus giving navigable paths in a localized region of interest.
Parag Chaudhuri, Rohit Khandekar, Deepak Sethi, Prem Kumar Kalra
Computer Graphics International1
2004 A System for View-Dependent Animation
abstract
Abstract In this paper, we present a novel system for facilitating the creation of stylized view‐dependent 3D animation. Our system harnesses the skill and intuition of a traditionally trained animator by providing a convivial sketch based 2D to 3D interface. A base mesh model of the character can be modified to match closely to an input sketch, with minimal user interaction. To do this, we recover the best camera from the intended view direction in the sketch using robust computer vision techniques. This aligns the mesh model with the sketch. We then deform the 3D character in two stages ‐ first we reconstruct the best matching skeletal pose from the sketch and then we deform the mesh geometry. We introduce techniques to incorporate deformations in the view‐dependent setting. This allows us to set up view‐dependent models for animation. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Three‐Dimensional Graphics and Realism ‐ Animation Our system takes as input a sketch (a), and a base mesh model (b), then recovers a camera to orient the base mesh (c), then reconstructs the skeleton pose (d), and finally deforms the mesh to find the best possible match with the sketch (e). image
Parag Chaudhuri, Prem Kumar Kalra, Subhashis Banerjee
Comput. Graph. Forum1
2004 A measure for mesh compression of time-variant geometry
abstract
Abstract We present a novel measure for compression of time‐variant geometry. Compression of time‐variant geometry has become increasingly relevant as transmission of high quality geometry streams is severely limited by network bandwidth. Some work has been done on such compression schemes, but none of them give a measure for prioritizing the loss of information from the geometry stream while doing a lossy compression. In this paper we introduce a cost function which assigns a cost to the removal of particular geometric primitives during compression, based upon their importance in preserving the complete animation. We demonstrate that the use of this measure visibly enhances the performance of existing compression schemes. Copyright © 2004 John Wiley & Sons, Ltd.
Prasun Mathur, Chhavi Upadhyay, Parag Chaudhuri, Prem Kumar Kalra
Comput. Animat. Virtual Worlds3