VLDB 2026 Research / reviewers in the wild / expert
Ramya Hebbalaguppe
dblp:145/2287 · also Ramya S. Hebbalaguppe
· DBLP profile ↗
22ranked-venue papers
6as first author
13since 2021 · last 2026
0009-0006-1186-6311ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Refine and Align: Confidence Calibration Through Multi-Agent Interaction in VQAabstractIn the context of Visual Question Answering (VQA) and Agentic AI, calibration refers to how closely an AI system's confidence in its answers reflects their actual correctness. This aspect becomes especially important when such systems operate autonomously and must make decisions under visual uncertainty. While modern VQA systems, powered by advanced vision-language models (VLMs), are increasingly used in high-stakes domains like medical diagnostics and autonomous navigation due to their improved accuracy, the reliability of their confidence estimates remains under-examined. Particularly, these systems often produce overconfident responses. To address this, we introduce AlignVQA, a debate-based multi-agent framework, in which diverse specialized VLM -- each following distinct prompting strategies -- generate candidate answers and then engage in two-stage interaction: generalist agents critique, refine and aggregate these proposals. This debate process yields confidence estimates that more accurately reflect the model’s true predictive performance. We find that more calibrated specialized agents produce better aligned confidences. Furthermore, we introduce a novel differentiable calibration-aware loss function called AlignCal designed to fine-tune the specialized agents by minimizing an upper bound on the calibration error. This objective explicitly improves the fidelity of each agent’s confidence estimates. Empirical results across multiple benchmark VQA datasets substantiate the efficacy of our approach, demonstrating substantial reductions in calibration discrepancies. Jai Bardhan, Ishita Jain, Ramya Hebbalaguppe, Rohan Raju Dhanakshirur, Lovekesh Vig |
AAAI | 4 |
| 2026 | Curve Skeletonization in Continuous domain for Meshes and Point CloudsabstractAdvancements in 3D curve skeletonization are accelerating progress across a wide range of applications. However, developing robust skeletonization algorithms that capture intricate object details remains challenging. Skeletonization via Local Separators (LS) offers an efficient graph-based approach but suffers from representation inaccuracies due to its discrete nature. To address this, we introduce CSCD, a novel framework for Curve Skeletonization in the Continuous Domain, generalizing LS to manifolds. Specifically, we present two realizations: CSCD-M for meshes and CSCD-PC for point clouds. CSCD-M leverages the intrinsic triangulation of a mesh for resilience to noise and improved topological preservation, while CSCD-PC employs tufted Laplacians for enhanced robustness. To our knowledge, CSCD-M is the first intrinsic method for curve skeletonization. Our results show CSCD-M matches LS performance across diverse meshes and outperforms LS (TOG’21) on benchmarks like Thingi10k dataset. CSCD-PC qualitatively outperforms CoverageAxis++ (Eurographics’24) and EPCS (CAG’23). Finally, we demonstrate the efficacy of CSCD in a few downstream tasks: object classification, shape segmentation, identifying handles, tunnels, and constrictions in objects.Project Website: https://cscd-skel.pages.dev Jai Bardhan, Ramya Hebbalaguppe, Aravind Udupa |
WACV | 2 |
| 2025 | Better Features, Better Calibration: A Simple Fix for Overconfident Networks
Soumya Suvra Ghosal, Ramya Hebbalaguppe, Dinesh Manocha |
ECML/PKDD (1) | 2 |
| 2025 | Prompting without Panic: Attribute-Aware, Zero-Shot, Test-Time Calibration
Ramya Hebbalaguppe, Tamoghno Kandar, Abhinav Nagpal, Chetan Arora 0001 |
ECML/PKDD (6) | 1 |
| 2024 | Calibration Transfer via Knowledge Distillation
Ramya Hebbalaguppe, Mayank Baranwal, Kartik Anand, Chetan Arora 0001 |
ACCV (8) | 1 |
| 2024 | LoMOE: Localized Multi-Object Editing via Multi-Diffusion
Goirik Chakrabarty, Aditya Chandrasekar, Ramya Hebbalaguppe, Prathosh A. P. |
ACM Multimedia | 3 |
| 2023 | Transfer4D: A Framework for Frugal Motion Capture and Deformation TransferabstractAnimating a virtual character based on a real performance of an actor is a challenging task that currently requires expensive motion capture setups and additional effort by expert animators, rendering it accessible only to large production houses. The goal of our work is to democratize this task by developing a frugal alternative termed “Transfer4D” that uses only commodity depth sensors and further reduces animators' effort by automating the rigging and animation transfer process. Our approach can transfer motion from an incomplete, single-view depth video to a semantically similar target mesh, unlike prior works that make a stricter assumption on the source to be noise-free and watertight. To handle sparse, incomplete videos from depth video inputs and variations between source and target objects, we propose to use skeletons as an intermediary representation between motion capture and transfer. We propose a novel unsupervised skeleton extraction pipeline from a single-view depth sequence that incorporates additional geometric information, resulting in superior performance in motion reconstruction and transfer in comparison to the contemporary methods and making our approach generic. We use non-rigid reconstruction to track motion from the depth sequence, and then we rig the source object using skinning decomposition. Finally, the rig is embedded into the target object for motion retargeting. Shubh Maheshwari, Rahul Narain, Ramya Hebbalaguppe |
CVPR | 3 |
| 2023 | Calibrating Deep Neural Networks using Explicit Regularisation and Dynamic Data PruningabstractDeep neural networks (DNNS) are prone to miscalibrated predictions, often exhibiting a mismatch between the predicted output and the associated confidence scores. Contemporary model calibration techniques mitigate the problem of overconfident predictions by pushing down the confidence of the winning class while increasing the confidence of the remaining classes across all test samples. However, from a deployment perspective an ideal model is desired to (i) generate well calibrated predictions for high-confidence samples with predicted probability say > 0.95 and (ii) generate a higher proportion of legitimate high-confidence samples. To this end, we propose a novel regularization technique that can be used with classification losses, leading to state-of-the-art calibrated predictions at test time; From a deployment standpoint in safety critical applications, only high-confidence samples from a well-calibrated model are of interest, as the remaining samples have to undergo manual inspection. Predictive confidence reduction of these potentially "high-confidence samples" is a downside of existing calibration approaches. We mitigate this via proposing a dynamic traintime data pruning strategy which prunes low confidence samples every few epochs, providing an increase in confident yet calibrated samples. We demonstrate state-of-the-art calibration performance across image classification benchmarks, reducing training time without much compromise in accuracy. We provide insights into why our dynamic pruning strategy that prunes low confidence training samples leads to an increase in high-confidence samples at test time. Rishabh Patra, Ramya Hebbalaguppe, Tirtharaj Dash, Gautam Shroff, Lovekesh Vig |
WACV | 2 |
| 2023 | A review on monocular tracking and mapping: from model-based to data-driven methods
Nivesh Gadipudi, I. Elamvazuthi, Lila Iznita Izhar, Lokender Tiwari, Ramya Hebbalaguppe, Cheng-Kai Lu, Arockia Selvakumar Arockia Doss |
Vis. Comput. | 5 |
| 2022 | A Stitch in Time Saves Nine: A Train-Time Regularizing Loss for Improved Neural Network CalibrationabstractDeep Neural Networks (dnns) are known to make over-confident mistakes, which makes their use problematic in safety-critical applications. State-of-the-art (sota) calibration techniques improve on the confidence of predicted labels alone, and leave the confidence of non-max classes (e.g. top-2, top-5) uncalibrated. Such calibration is not suitable for label refinement using post-processing. Further, most sota techniques learn a few hyper-parameters post-hoc, leaving out the scope for image, or pixel specific calibration. This makes them unsuitable for calibration under domain shift, or for dense prediction tasks like semantic segmentation. In this paper, we argue for intervening at the train time itself, so as to directly produce calibrated dnn models. We propose a novel auxiliary loss function: Multi-class Difference in Confidence and Accuracy (mdca), to achieve the same. mdca can be used in conjunction with other application/task specific loss functions. We show that training with mdca leads to better calibrated models in terms of Expected Calibration Error (ece), and Static Calibration Error (sce) on image classification, and segmentation tasks. We report ece (sce) score of 0.72 (1.60) on the cifar 100 dataset, in comparison to 1.90 (1.71) by the sota. Under domain shift, a ResNet-18 model trained on pacs dataset using mdca gives a average ece (sce) score of 19.7 (9.7) across all domains, compared to 24.2 (11.8) by the sota. For segmentation task, we report a 2× reduction in calibration error on pascal-voc dataset in comparison to Focal Loss [32]. Finally, mdca training improves calibration even on imbalanced data, and for natural language classification tasks. Ramya Hebbalaguppe, Jatin Prakash, Neelabh Madan, Chetan Arora 0001 |
CVPR | 1 |
| 2022 | A Novel Data Augmentation Technique for Out-of-Distribution Sample Detection Using Compounded Corruptions
Ramya Hebbalaguppe, Soumya Suvra Ghosal, Jatin Prakash, Harshad Khadilkar, Chetan Arora 0001 |
ECML/PKDD (3) | 1 |
| 2021 | Empirical Study of Data-Free Iterative Knowledge Distillation
Het Shah, Ashwin Vaswani, Tirtharaj Dash, Ramya Hebbalaguppe, Ashwin Srinivasan 0001 |
ICANN (3) | 4 |
| 2021 | 3DPoseLite: A Compact 3D Pose Estimation Using Node EmbeddingsabstractEfficient pose estimation finds utility in Augmented Reality (AR) and other computer vision applications such as autonomous navigation and robotics, to name a few. A compact and accurate pose estimation methodology is of paramount importance for on-device inference in such applications. Our proposed solution 3DPoseLite, estimates pose of generic objects by utilizing a compact node embedding representation, unlike computationally expensive multi-view and point-cloud representations. The neural network outputs a 3D pose, taking RGB image and its corresponding graph (obtained by skeletonizing the 3D meshes [31]) as inputs. Our approach utilizes node2vec framework to learn low-dimensional representations for nodes in a graph by optimizing a neighborhood preserving objective. We achieve a space and time reduction by a factor of 11 × and 3 × respectively, with respect to the state-of-the-art approach, Pose-FromShape [50], on benchmark Pascal3D dataset [48]. We also test the performance of our model on unseen data using Pix3D dataset. Meghal Dani, Karan Narain, Ramya Hebbalaguppe |
WACV | 3 |
| 2020 | An Empirical Study of Iterative Knowledge Distillation for Neural Network Compression
Sharan Yalburgi, Tirtharaj Dash, Ramya Hebbalaguppe, Srinidhi Hegde, Ashwin Srinivasan 0001 |
ESANN | 3 |
| 2020 | Variational Student: Learning Compact and Sparser Networks In Knowledge Distillation FrameworkabstractThe holy grail in deep neural network research is porting the memory- and computation-intensive network models on embedded platforms with a minimal compromise in model accuracy. To this end, we propose Variational Student where we reap the benefits of compressibility of the knowledge distillation framework, and sparsity inducing abilities of variational inference (VI) techniques. Essentially, we build an accurate and sparse student network, whose sparsity is induced by the variational parameters found via optimizing a loss function based on VI, leveraging the knowledge learnt by an accurate but complex pre-trained teacher network. Further, for sparsity enhancement, we also employ a Block Sparse Regularizer on a concatenated tensor of teacher and student network weights. We benchmark our results on MLP and CNN variants and illustrate an improved performance in lowering the memory footprint up to ~ 213× without a need to retrain the teacher network. Srinidhi Hegde, Ranjitha Prasad, Ramya Hebbalaguppe, Vishwajeet Kumar |
ICASSP | 3 |
| 2020 | SmartOverlays: A Visual Saliency Driven Label Placement for Intelligent Human-Computer InterfacesabstractIn augmented reality (AR), the computer generated labels assist in understanding a scene by addition of contextual information. However, naive label placement often results in clutter and occlusion impairing the effectiveness of AR visualization. For label placement, the main objectives to be satisfied are, non-occlusion to the scene of interest, the proximity of labels to the object, and, temporally coherent labels in a video/live feed. We present a novel method for the placement of labels corresponding to objects of interest in a video/live feed that satisfies the aforementioned objectives. Our proposed framework, SmartOverlays, first identifies the objects and generates corresponding labels using a YOLOv2 [28] in a video frame; at the same time, Saliency Attention Model (SAM) [7] learns eye fixation points that aid in predicting saliency maps; finally, computes Voronoi partitions of the video frame, choosing the centroids of objects as seed points, to place labels for satisfying the proximity constraints with the object of interest. In addition, our approach incorporates tracking the detected objects in a frame to facilitate temporal coherence between frames that enhances the readability of labels. We measure the effectiveness of SmartOverlays framework using three objective metrics: (a) Label Occlusion over Saliency (LOS), (b) temporal jitter metric to quantify jitter in the label placement, (c) computation time for label placement. Srinidhi Hegde, Jitender Maurya, Ramya Hebbalaguppe, Aniruddha Kalkar |
WACV | 3 |
| 2018 | Real Time Hand Segmentation on Frugal Headmounted Device for Gestural InterfaceabstractWith the resurgence of Head Mounted Displays (HMDs), in-air gestures form a natural and intuitive interaction mode of communication. HMDs such as Microsoft Hololens, Daqri smart-glasses etc., have on-board processors with additional sensors, making the device expensive. Our goal, therefore, is to enable mass-market reach by extending the interaction space around mobile devices for Augmented/Virtual reality: with just frugal head-mounts such as Google Cardboard and Wearality with a smartphone. One necessary step for a feasible human-computer interaction is gesture recognition, preceded by a reliable hand segmentation in egocentric view. We propose a technique for real-time hand segmentation that utilize only RGB camera in an off-the-shelf mobile device. The novelty of our work lies in coming up with a filtering technique that term Multi Orientation Matched filter, used for hand segmentation that works on-device, even in situations of skin-like background. We have extensively tested our hand segmentation method on public datasets. We provide comparisons of hand segmentation with the existing methods under the same limitations. We demonstrate that our method outperforms in term of computational time and comparable results in term of accuracy. Further, we demonstrate that using our solution a zoom gesture classification can work in real-time on android smartphones. Jitender Maurya, Ramya Hebbalaguppe, Puneet Gupta 0002 |
ICIP | 2 |
| 2018 | Where to Place: A Real-Time Visual Saliency Based Label Placement for Augmented Reality ApplicationsabstractTextual overlays/labels add contextual information in Augmented Reality (AR) applications. The spatial placement of labels is a challenging task due to constraints that labels (i) are not occluding the object/scene of interest, and, (ii) are optimally placed for better interpretation of scene. To this end, we present a novel method for optimal placement of labels for AR. We formulate this method by an objective function that minimizes both occlusion with visually salient regions in scenes of interest, and the temporal jitter for facilitating coherence in real-time AR applications. The main focus of proposed algorithm is real-time label placement on low-end android phones/tablets. The sophisticated state-of-the-art algorithms for optimal positioning of textual label work only on the images and are often inefficient for real-time performance on those devices. We demonstrate the efficiency of our method by porting the algorithm on a smart-phone/tablet. Further, we capture objective and subjective metrics to determine the efficacy of the method; objective metrics include computation time taken for determining the label location and Label Occlusion over Saliency (LOS) score over salient regions in the scene. Subjective metrics include position, temporal coherence in the overlay, color and responsiveness. Neel Rakholia, Srinidhi Hegde, Ramya Hebbalaguppe |
ICIP | 3 |
| 2017 | Telecom Inventory Management via Object Recognition and Localisation on Google Street View ImagesabstractWe present a novel method to update assets for telecommunication infrastructure using google street view (GSV) images. The problem is formulated as a object recognition task, followed by use of triangulation to estimate the object coordinates from sensor plane coordinates, To this end, we have explored different state-of-the-art object recognition techniques both from feature engineering and using deep learning namely HOG descriptors with SVM, Deformable parts model (DPM), and Deep learning (DL) using faster RCNNs. While HOG+SVM has proved to be robust human detector, DPM which is based on probabilistic graphical models and DL which is a non-linear classifier have proved their versatility in different types of object recognition problems. Asset recognition from the street view images however pose unique challenge as they could be installed on the ground in various poses, orientations and with occlusions, objects camouflaged in the background and in some cases inter class variation is small. We present comparative performance of these techniques for specific use-case involving telecom equipment for highest precision and recall. The blocks of proposed pipeline are detailed and compared to traditional inventory management methods. Ramya Hebbalaguppe, Ehtesham Hassan, Hiranmay Ghosh |
WACV | 1 |
| 2017 | Robust Hand Gestural Interaction for Smartphone Based AR/VR ApplicationsabstractThe future of user interfaces will be dominated by hand gestures. In this paper, we explore an intuitive hand gesture based interaction for smartphones having a limited computational capability. To this end, we present an efficient algorithm for gesture recognition with First Person View (FPV), which focuses on recognizing a four swipe model (Left, Right, Up and Down) for smartphones through single monocular camera vision. This can be used with frugal AR/VR devices such as Google Cardboard1andWearality2in building AR/VR based automation systems for large scale deployments, by providing a touch-less interface and real-time performance. We take into account multiple cues including palm color, hand contour segmentation, and motion tracking, which effectively deals with FPV constraints put forward by a wearable. We also provide comparisons of swipe detection with the existing methods under the same limitations. We demonstrate that our method outperforms both in terms of gesture recognition accuracy and computational time. Shreyash Mohatta, Ramakrishna Perla, Ehtesham Hassan, Ramya Hebbalaguppe |
WACV | 5 |
| 2016 | Reduction of false alarms triggered by spiders/cobwebs in surveillance camera networksabstractThe percentage of false alarms caused by spiders in automated surveillance can range from 20-50%. False alarms increase the workload of surveillance personnel validating the alarms and the maintenance labor cost associated with regular cleaning of webs. We propose a novel, cost effective method to detect false alarms triggered by spiders/webs in surveillance camera networks. This is accomplished by building a spider classifier intended to be a part of the surveillance video processing pipeline. The proposed method uses a feature descriptor obtained by early fusion of blur and texture. The approach is sufficiently efficient for real-time processing and yet comparable in performance with more computationally costly approaches like SIFT with bag of visual words aggregation. The proposed method can eliminate 98.5% of false alarms caused by spiders in a data set supplied by an industry partner, with a false positive rate of less than 1%. Ramya Hebbalaguppe, Kevin McGuinness, Jogile Kuklyte, Rami Albatal, Cem Direkoglu, Noel E. O'Connor |
ICIP | 1 |
| 2016 | Automatic Container Code Recognition via Spatial Transformer Networks and Connected Component Region ProposalsabstractContainer identification and recognition is still performed manually or in a semi-automatic fashion in multiple ports globally. This results in errors and inefficiencies in port operations. The problem of automatic container identification and recognition is challenging as the ISO standard only prescribes the pattern of the code and does not specify other parameters such as the foreground and background colors, font type and size, orientation of characters (horizontal or vertical) so on. Additionally, the corrugated surface of container body makes the two dimensional projection of the text on three dimensional containers slanted and jagged. We propose a solution in the form of an end-to-end pipeline that uses Region Proposals generated based on Connected Components for text detection in conjunction with Spatial Transformer NeConnected Componentstworks for text recognition. We demonstrate via our experimental results that the pipeline is reliable and robust even in situations when the code characters are highly distorted and outperforms the state-of-the-art results for text detection and recognition over the containers. We achieve text coverage rate of 100% and text recognition rate of 99.64%. Ramya Hebbalaguppe, Ehtesham Hassan, Lovekesh Vig |
ICMLA | 3 |