EDBT 2026 Demo / reviewers in the wild / expert
Vigneshwaran Subbaraju
dblp:68/11081
· DBLP profile ↗
25ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-4276-5939ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Computer networks · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Personalized Image Privacy Advisors via Federated Daisy-ChainingabstractImage sharing on social media has become routine but poses serious privacy risks, as users may unknowingly expose sensitive information. This necessitates an image privacy advisor that assigns personalized privacy risk scores, helping users decide whether to share images publicly or not. However, centralized training of such models risks user data exposure and loss of ownership, as data must be uploaded to a central server. To safeguard user privacy, we adopt Federated Learning (FL), which enables collaborative model training without sharing raw data. Despite its advantages, FL faces challenges such as data heterogeneity from diverse user privacy preferences, limited annotations per user, and communication overhead. To address these issues, we propose CFedDC, a personalized FL algorithm combined with PIONet, a mobile friendly model with 14.5 × fewer trainable parameters and 93.09% lower memory footprint than centralized baselines. CFedDC mitigates data heterogeneity through clustering and cluster aware regularization with stability, and tackles data scarcity using a daisy-chaining knowledge transfer mechanism. Comprehensive experimental evaluations demonstrate that our proposed method achieves well-aligned personalized user privacy scores, outperforming existing centralized and FL-based image privacy models. Sourasekhar Banerjee, Vengateswaran Subramaniam, Debaditya Roy, Vigneshwaran Subbaraju, Monowar Bhuyan |
WACV | 4 |
| 2026 | Detecting Social Engagement of Elderly From Lifelog Image-streams to Identify Effective Cues for Autobiographic RecallabstractLifelog images captured automatically by wearable cam-eras serve as effective cues that induce Autobiographic Memory Recall (AMR) of social interactions. This is very useful for personalized memory interventions. However, manual selection of images for such therapy imposes significant load on the caregivers who need to browse through a voluminous collection of images. To reduce this load, auto-mated tools that identify moments involving significant engagement of the camera wearer in social interactions are needed. To achieve this, we reannotate images extracted from public lifelog datasets for the presence of non-verbal social signals and the perceived engagement of the lifelogger during interactions. We use this data to develop models and explore how social signals and the detected intensity of social engagement are helpful for predicting AMR. We show that understanding visual social engagement can enhance AMR prediction, demonstrating the potential of the models in reducing caregivers’ effort. Vengateswaran Subramaniam, Vigneshwaran Subbaraju, Debaditya Roy, Pramath Krishna, Thivya Kandappu, Qianli Xu |
WACV | 2 |
| 2025 | Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understandingabstract3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pure language-based 3D grounding, there has been limited exploration of 3D-ERU, which also incorporates human pointing gestures. To address this gap, we introduce a data augmentation framework– Imputer, and use it to curate a new benchmark dataset– ImputeRefer for 3D-ERU, by incorporating human pointing gestures into existing 3D scene datasets that only contain language instructions. We also propose Ges3ViG, a novel model for 3D-ERU that achieves ~30% improvement in accuracy as compared to other 3D-ERU models and ~9% compared to other purely language-based 3D grounding models. Our code and dataset are available at https://github.com/AtharvMane/Ges3ViG. Atharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju, Sougata Sen, Sanjay E. Sarma, Archan Misra |
CVPR | 3 |
| 2025 | Predicting Event Memorability Using Personalized Federated LearningabstractLifelog images are very useful as memory cues for recalling past events. Estimating the level of event memory recall induced by a given lifelog image (event memorability), is useful for selecting images for cognitive interventions. Previous works for predicting event memorability follow a centralized model training paradigm that requires several users to share their lifelog images. This risks violating the privacy of individual lifeloggers. Alternatively, a personal model trained with a lifelogger's own data guarantees privacy. However, it imposes significant effort on the lifelogger to provide a large enough sample of self-rated images to develop a well-performing model for event memorability. Therefore, we propose a clustered personalized federated learning setup FedMEM that avoids sharing raw images but still enables collaborative learning via model sharing. For an enhanced learning performance in the presence of data heterogeneity, FedMEM evaluates similarity among users to group them into clusters. We demonstrate that our approach furnishes high-performing personalized models compared to the state-of-the-art. Sourasekhar Banerjee, Debaditya Roy, Vigneshwaran Subbaraju, Monowar Bhuyan |
WACV | 3 |
| 2025 | NeuroViG - Integrating Event Cameras for Resource-Efficient Video GroundingabstractSpatio-Temporal Video Grounding (STVG) - the task of identifying the target object in the field-of-view, that the language instruction refers to - is a fundamental vision-language task. Current STVG approaches typically utilize feeds from an RGB camera that is assumed to be always-on and process the video frames using complex neural network pipelines. As a result, they often impose prohibitive system overheads (energy, latency) on pervasive devices. To address this, we propose NeuroViG with two key innovations: (a) leveraging on event streams from a low-power neuromorphic event camera sensor to perform selective triggering of the more energy-hungry RGB camera for STVG, and (b) augmenting the STVG model with a lightweight Adaptive Frame Selector (AFS) that bypasses complex transformer-based operations for a majority of video frames, thereby enabling its execution on a pervasive Jetson AGX device. We have also introduced modifications to the neural network processing pipeline such that the system can offer tunable tradeoffs between accuracy and energy/latency. Our proposed NeuroViG system allows us to reduce the STVG energy overhead and latency by ~ 4x and ~ 3.8x, respectively, for less than 1% loss in accuracy. Dulanga Weerakoon, Vigneshwaran Subbaraju, Joo-Hwee Lim, Archan Misra |
WACV | 2 |
| 2024 | Poster: Towards Efficient Spatio-Temporal Video Grounding in Pervasive Mobile DevicesabstractAs the use of pervasive devices expands into complex collaborative tasks such as cognitive assistants and interactive AR/VR companions, they are equipped with a myriad of sensors facilitating natural interactions, such as voice commands. Spatio-Temporal Video Grounding (STVG), the task of identifying the target object in the field-of-view referred to in a language instruction, is a key capability needed for such systems. However, current STVG models tend to be resource-intensive, relying on multiple cross-attentional transformers applied to each video frame. This results in runtime complexity that increases linearly with video length. Furthermore, deploying these models on mobile devices while maintaining a low-latency poses additional challenges. Hence, this paper explores the latency and energy requirements for implementing STVG models on a pervasive device. Dulanga Weerakoon, Vigneshwaran Subbaraju, Joo-Hwee Lim, Archan Misra |
MobiSys | 2 |
| 2024 | PrivObfNet: A Weakly Supervised Semantic Segmentation Model for Data ProtectionabstractThe use of social media has made it easy to communicate and share information over the internet. However, it also brings issues such as data privacy leakage, which can be exploited by recipients with malicious intentions to harm the sender. In this paper, we propose a deep neural network that analyzes user’s image for privacy sensitive content and automatically locates sensitive regions for obfuscation. Our approach relies solely on image level annotations and learns to (a) predict an overall privacy score, (b) detect sensitive attributes and (c) demarcate the sensitive regions for obfuscation, in a given input image. We validated the performance of our proposed method on three large datasets, VISPR, PASCAL VOC 2012 and MS COCO 2014, in terms of privacy score, attribute prediction and obfuscation performance. On the VISPR dataset, we achieved a Pearson correlation of 0.88 and a Spearman correlation of 0.86, outperforming previous methods. On PASCAL VOC 2012 and MS COCO 2014, our model achieved a mean IOU of 71.5% and 43.9% respectively, and is among the state-of-the-art techniques using weakly supervised semantic segmentation learning. Chiat-Pin Tay, Vigneshwaran Subbaraju, Thivya Kandappu |
WACV | 2 |
| 2024 | Manufacturing domain instruction comprehension using synthetic data
Kritika Johari, Christopher Tay Zi Tong, Rishabh Bhardwaj, Vigneshwaran Subbaraju, Jung-Jae Kim 0001, U-Xuan Tan |
Vis. Comput. | 4 |
| 2023 | Demo Abstract: VGGlass - Demonstrating Visual Grounding and Localization Synergy with a LiDAR-enabled Smart-GlassabstractThis work demonstrates the VGGlass system, which simultaneously interprets human instructions for a target acquisition task and determines the precise 3D positions of both user and the target object. This is achieved by utilizing LiDARs mounted in the infrastructure and a smart glass device worn by the user. Key to our system is the union of LiDAR-based localization termed LiLOC and a multi-modal visual grounding approach termed RealG(2)In-Lite. To demonstrate the system, we use Intel RealSense L515 cameras and a Microsoft HoloLens 2, as the user devices. VGGlass is able to: a) track the user in real-time in a global coordinate system, and b) locate target objects referred by natural language and pointing gestures. Darshana Rathnayake, Dulanga Weerakoon, Meera Radhakrishnan, Vigneshwaran Subbaraju, Inseok Hwang 0001, Archan Misra |
SenSys | 4 |
| 2022 | SoftSkip: Empowering Multi-Modal Dynamic Pruning for Single-Stage Referring ComprehensionabstractSupporting real-time referring expression comprehension (REC) on pervasive devices is an important capability for human-AI collaborative tasks. Model pruning techniques, applied to DNN models, can enable real-time execution even on resource-constrained devices. However, existing pruning strategies are designed principally for uni-modal applications, and suffer a significant loss of accuracy when applied to REC tasks that require fusion of textual and visual inputs. We thus present a multi-modal pruning model, LGMDP, which uses language as a pivot to dynamically and judiciously select the relevant computational blocks that need to be executed. LGMDP also introduces a new SoftSkip mechanism, whereby 'skipped' visual scales are not completely eliminated but approximated with minimal additional computation. Experimental evaluation, using 3 benchmark REC datasets and an embedded device implementation, shows that LGMDP can achieve 33% latency savings, with an accuracy loss 0.5% - 2%. Dulanga Weerakoon, Vigneshwaran Subbaraju, Archan Misra |
ACM Multimedia | 2 |
| 2021 | Predicting Event Memorability from Contextual Visual SemanticsabstractEpisodic event memory is a key component of human cognition. Predicting event memorability,i.e., to what extent an event is recalled, is a tough challenge in memory research and has profound implications for artificial intelligence. In this study, we investigate factors that affect event memorability according to a cued recall process. Specifically, we explore whether event memorability is contingent on the event context, as well as the intrinsic visual attributes of image cues. We design a novel experiment protocol and conduct a large-scale experiment with 47 elder subjects over 3 months. Subjects’ memory of life events is tested in a cued recall process. Using advanced visual analytics methods, we build a first-of-its-kind event memorability dataset (called R3) with rich information about event context and visual semantic features. Furthermore, we propose a contextual event memory network (CEMNet) that tackles multi-modal input to predict item-wise event memorability, which outperforms competitive benchmarks. The findings inform deeper understanding of episodic event memory, and open up a new avenue for prediction of human episodic memory. Source code is available at https://github.com/ffzzy840304/Predicting-Event-Memorability. Qianli Xu, Fen Fang, Ana Garcia del Molino, Vigneshwaran Subbaraju, Joo-Hwee Lim |
NeurIPS | 4 |
| 2021 | Discriminant Spatial Filtering Method (DSFM) for the identification and analysis of abnormal resting state brain activities
Abhay M. S. Aradhya, Vigneshwaran Subbaraju, Suresh Sundaram 0002, Narasimhan Sundararajan |
Expert Syst. Appl. | 2 |
| 2021 | PrivacyPrimer: Towards Privacy-Preserving Episodic Memory Support For Older AdultsabstractBuilt-in pervasive cameras have become an integral part of mobile/wearable devices and enabled a wide range of ubiquitous applications with their ability to be "always-on". In particular, life-logging has been identified as a means to enhance the quality of life of older adults by allowing them to reminisce about their own life experiences. However, the sensitive images captured by the cameras threaten individuals' right to have private social lives and raise concerns about privacy and security in the physical world. This threat gets worse when image recognition technologies can link images to people, scenes, and objects, hence, implicitly and unexpectedly reveal more sensitive information such as social connections. In this paper, we first examine life-log images obtained from 54 older adults to extract (a) the artifacts or visual cues, and (b) the context of the image that influences an older life-logger's ability to recall the life events associated with a life-log image. We call these artifacts and contextual cues "stimuli". Using the set of stimuli extracted, we then propose a set of obfuscation strategies that naturally balances the trade-off between reminiscability and privacy (revealing social ties) while selectively obfuscating parts of the images. More specifically, our platform yields privacy-utility tradeoff by compromising, on average, modest 13.4% reminiscability scores while significantly improving privacy guarantees -- around 40% error in cloud estimation. Thivya Kandappu, Vigneshwaran Subbaraju, Qianli Xu |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2021 | Lifelog Image Retrieval Based on Semantic Relevance MappingabstractLifelog analytics is an emerging research area with technologies embracing the latest advances in machine learning, wearable computing, and data analytics. However, state-of-the-art technologies are still inadequate to distill voluminous multimodal lifelog data into high quality insights. In this article, we propose a novel semantic relevance mapping ( SRM ) method to tackle the problem of lifelog information access. We formulate lifelog image retrieval as a series of mapping processes where a semantic gap exists for relating basic semantic attributes with high-level query topics. The SRM serves both as a formalism to construct a trainable model to bridge the semantic gap and an algorithm to implement the training process on real-world lifelog data. Based on the SRM, we propose a computational framework of lifelog analytics to support various applications of lifelog information access, such as image retrieval, summarization, and insight visualization. Systematic evaluations are performed on three challenging benchmarking tasks to show the effectiveness of our method. Qianli Xu, Ana Garcia del Molino, Jie Lin 0001, Fen Fang, Vigneshwaran Subbaraju, Liyuan Li, Joo-Hwee Lim |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | Gesture Enhanced Comprehension of Ambiguous Human-to-Robot InstructionsabstractThis work demonstrates the feasibility and benefits of using pointing gestures, a naturally-generated additional input modality, to improve the multi-modal comprehension accuracy of human instructions to robotic agents for collaborative tasks.We present M2Gestic, a system that combines neural-based text parsing with a novel knowledge-graph traversal mechanism, over a multi-modal input of vision, natural language text and pointing. Via multiple studies related to a benchmark table top manipulation task, we show that (a) M2Gestic can achieve close-to-human performance in reasoning over unambiguous verbal instructions, and (b) incorporating pointing input (even with its inherent location uncertainty) in M2Gestic results in a significant (30%) accuracy improvement when verbal instructions are ambiguous. Dulanga Weerakoon, Vigneshwaran Subbaraju, Nipuni Karumpulli, Qianli Xu, U-Xuan Tan, Joo-Hwee Lim, Archan Misra |
ICMI | 2 |
| 2020 | PrivAttNet: Predicting Privacy Risks in Images Using Visual AttentionabstractVisual privacy concerns associated with image sharing is a critical issue that need to be addressed to enable safe and lawful use of online social platforms. Users of social media platforms often suffer from no guidance in sharing sensitive images in public, and often face with social and legal consequences. Given the recent success of visual attention based deep learning methods in measuring abstract phenomena like image memorability, we are motivated to investigate whether visual attention based methods could be useful in measuring psychophysical phenomena like “privacy sensitivity”. In this paper we propose PrivAttNet - a visual attention based approach, that can be trained end-to-end to estimate the privacy sensitivity of images without explicitly detecting sensitive objects and attributes present in the image. We show that our PrivAttNet model outperforms various SOTA and baseline strategies - a 1.6 fold reduction in L1 - error over SOTA and 7%-10% improvement in Spearman-rank correlation between the predicted and ground truth sensitivity scores. Additionally, the attention maps from PrivAttNet are found to be useful in directing the users to the regions that are responsible for generating the privacy risk score. Thivya Kandappu, Vigneshwaran Subbaraju |
ICPR | 3 |
| 2020 | Annapurna: An automated smartwatch-based eating detection and food journaling system
Sougata Sen, Vigneshwaran Subbaraju, Archan Misra, Rajesh Krishna Balan, Youngki Lee 0001 |
Pervasive Mob. Comput. | 2 |
| 2018 | I4S: capturing shopper's in-store interactionsabstractIn this paper, we present I4S, a system that identifies item interactions of customers in a retail store through sensor data fusion from smartwatches, smartphones and distributed BLE beacons. To identify these interactions, I4S builds a gesture-triggered pipeline that (a) detects the occurrence of "item picks", and (b) performs fine-grained localization of such pickup gestures. By analyzing data collected from 31 shoppers visiting a midsized stationary store, we show that we can identify person-independent picking gestures with a precision of over 88%, and identify the rack from where the pick occurred with 91%+ precision (for popular racks). Sougata Sen, Archan Misra, Vigneshwaran Subbaraju, Karan Grover, Meera Radhakrishnan, Rajesh Krishna Balan, Youngki Lee 0001 |
UbiComp | 3 |
| 2018 | Personalized Serious Games for Cognitive Intervention with Lifelog Visual AnalyticsabstractThis paper presents a novel serious game app and a method to cre- ate and integrate personalized game content based on lifelog visual analytics. The main objective is to extract personalized content from visual lifelogs, integrate it into mobile games, and evaluate the effect of personalization on user experience. First, a suite of visual analysis methods is proposed to extract semantic informa- tion from visual lifelogs and discover the association among the lifelog entities. The outcome is dataset that contains augmented and personal lifelog images. Next, a mobile game app is developed that makes use of the dataset as game content. Finally, an experiment is conducted to evaluate user gameplay behaviors in the wild over three months, where a mixture of generic and personalized game content is deployed. It is observed that user adherence is heightened by personalized game content as compared to generic content. Also observed is a higher enjoyment level in personalized than generic game content. The result provides the first empirical evidence of the effect of personalized games on user adherence and preference for cognitive intervention. This work paves the way for effective cognitive training with user-generated content. Qianli Xu, Vigneshwaran Subbaraju, Chee How Cheong, Aijing Wang, Kathleen Kang, Munirah Bashir, Yanhong Dong, Liyuan Li, Joo-Hwee Lim |
ACM Multimedia | 2 |
| 2018 | Annapurna: Building a Real-World Smartwatch-Based Automated Food JournalabstractWe describe the design and implementation of a smartwatch-based, completely unobtrusive, food journaling system, where the smartwatch helps to intelligently capture useful images of food that an individual consumes throughout the day. The overall system, called Annapurna, is based on three key components: (a) a smartwatch-based gesture recognizer to identify eating gestures, (b) a smartwatch-based image capturer that obtains a small set of relevant and useful images with a low energy overhead, and (c) a server-based image filtering engine that removes irrelevant uploaded images, and then catalogs them through a portal. Our primary challenge is to make the system robust to the huge diversity in natural eating habits and food choices. We show how we address this by an appropriate coupling between a smartwatch's camera sensor and inertial sensor-based tracking of eating gestures, thereby helping to capture multiple likely-to-be-useful images with low energy overhead. Through a series of real-world, in-the-wild studies, we demonstrate the end-to-end working of Annapurna, which captures useful images in over 95% of all natural eating episodes. Sougata Sen, Vigneshwaran Subbaraju, Archan Misra, Rajesh Krishna Balan, Youngki Lee 0001 |
WOWMOM | 2 |
| 2017 | MiPAL: Multiple-instance passive aggressive learning for identification of attention deficit hyperactive disorder from fMRIabstractThis paper proposes a new algorithm for the multiple instance learning problem (MIL) and investigates its application for detecting Attention Deficit Hyperactive Disorder (ADHD) from resting-state functional Magnetic Resonance Imaging data. The core component of many kernel-based MIL algorithms is usually an SVM-like batch optimization framework, hence scaling to large datasets like fMRI is often difficult. On the other hand, a family of on-line kernel classification algorithms widely known as “perceptron-like” kernel classifiers demonstrate efficient and accurate solutions. This paper presents MiPAL - Multiple-instance Passive Aggressive Learning algorithm, based on such a perceptron-like kernel classifier. First, MiPAL builds a labeller by inputting negative bags into PA algorithm. Second, this labeller helps to train a separate PA classifier to predict binary class labels such that least-negative instances are regarded as positive. Due to the on-line PA algorithm's fast adaptation, the impact of invalid positive support-vectors could be attenuated by the new, accurate support-set over time. Our experimental results reveal performance gains in several MIL datasets including state-of-the-art performance in Muskl, Fox, and comparable accuracy in the preprocessed ADHD-200 dataset. Prabhash Kumarasinghe, Suresh Sundaram 0002, Vigneshwaran Subbaraju |
IJCNN | 3 |
| 2017 | Identifying differences in brain activities and an accurate detection of autism spectrum disorder using resting state functional-magnetic resonance imaging : A spatial filtering approach
Vigneshwaran Subbaraju, Mahanand Belathur Suresh, Suresh Sundaram 0002, Narasimhan Sundararajan |
Medical Image Anal. | 1 |
| 2015 | Using infrastructure-provided context filters for efficient fine-grained activity sensingabstractWhile mobile and wearable sensing can capture unique insights into fine-grained activities (such as gestures and limb-based actions) at an individual level, their energy overheads are still prohibitive enough to prevent them from being executed continuously. In this paper, we explore practical alternatives to addressing this challenge-by exploring how cheap infrastructure sensors or information sources (e.g., BLE beacons) can be harnessed with such mobile/wearable sensors to provide an effective solution that reduces energy consumption without sacrificing accuracy. The key idea is that many fine-grained activities that we desire to capture are specific to certain location, movement or background context: infrastructure sensors and information sources (e.g., BLE beacons) offer practical and cheap ways to identify such context. In this paper, we first explore how various infrastructure, mobile & wearable sensors can be used to identify fine-grained location/movement context (e.g., transiting through a door). We then show, using a couple of illustrative examples (specifically, the detection of `switch pressing' before exiting a room and the identification of `water drinking' after approaching a water cooler) to show that such background context can be predicted, with sufficient accuracy, with sufficient lead time to enable a `triggered' model for mobile/wearable sensing of such microscopic, transient gestures and activities. Moreover, such `triggered' sensing also helps to improve the accuracy of such microscopic gesture recognition, by reducing the set of candidate activity labels. Empirical experiments show that we are able to identify 82.2% of switch-pressing and 91.73% of water-drinking activities in a campus lab setting, with a significant reduction in active sensing time (up to 92.9% compared to continuous sensing). Vigneshwaran Subbaraju, Sougata Sen, Archan Misra, Satyadip Chakraborti, Rajesh Krishna Balan |
PerCom | 1 |
| 2015 | Accurate detection of autism spectrum disorder from structural MRI using extended metacognitive radial basis function network
Vigneshwaran Subbaraju, Suresh Sundaram 0002, Narasimhan Sundararajan, Mahanand Belathur Suresh |
Expert Syst. Appl. | 1 |
| 2012 | Demo: context driven advertisement optimizerabstractNo abstract available. Azeem J. Khan, Vigneshwaran Subbaraju, Archan Misra, Srinivasan Seshan |
MobiSys | 2 |