VLDB 2026 Research / reviewers in the wild / expert
Sougata Sen
dblp:150/9980
· DBLP profile ↗
20ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-2466-0025ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 10 · 3 first-author · 6 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understandingabstract3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pure language-based 3D grounding, there has been limited exploration of 3D-ERU, which also incorporates human pointing gestures. To address this gap, we introduce a data augmentation framework– Imputer, and use it to curate a new benchmark dataset– ImputeRefer for 3D-ERU, by incorporating human pointing gestures into existing 3D scene datasets that only contain language instructions. We also propose Ges3ViG, a novel model for 3D-ERU that achieves ~30% improvement in accuracy as compared to other 3D-ERU models and ~9% compared to other purely language-based 3D grounding models. Our code and dataset are available at https://github.com/AtharvMane/Ges3ViG. Atharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju, Sougata Sen, Sanjay E. Sarma, Archan Misra |
CVPR | 4 |
| 2025 | WiP Paper: Navigator: A Microcontroller-based Assistive Smartglass for Guided Navigation
Saket Nerurkar, Vyomesh Bhatt, Surjya Ghosh, Sougata Sen |
EWSN | 4 |
| 2024 | Improving Continuous Emotion Annotation in Video Platforms via Physiological Response ProfilingabstractMany video applications (e.g., gaming, meeting, tutoring) aim to improve the user's interaction experience based on continuously inferred user emotion. To infer user emotion, these apps typically deploy machine learning models, trained with continuously collected emotion ground truth labels. However, as continuous annotations are generally collected as emotion self-reports during video consumption (using an auxiliary device), they incur significant annotation effort. To address this problem, we propose PResUP, a framework that creates users' profile using physiological responses (e.g., GSR or galvanic skin re-sponse) and deploys an LSTM network to identify the opportune probing moments for emotion ground truth (emotion self-report) collection instead of continuous annotation. We evaluate the proposed approach on a large-scale publicly available dataset (CASE) containing the physiological signals of subjects during video consumption. The evaluation of PRes UP reveals that it reduces the probing rate by 30.3 % (on average), detects the opportune probing moments with a TPR (True Positive Rate) of 86.1 %, and yet maintains the quality of the emotion annotations as observed in the continuous self-reports. Furthermore, we evaluated the generalizability of PResUP on another public dataset (K-emocon); which reveals an average probing rate reduction of 25.71 %. These results underscore the efficiency of PResUP in reducing the continuous emotion annotation overhead. Swarnali Banik, Sougata Sen, Snehanshu Saha, Surjya Ghosh |
ACII | 2 |
| 2024 | Towards Estimating Missing Emotion Self-reports Leveraging User Similarity: A Multi-task Learning ApproachabstractThe Experience Sampling Method (ESM) is widely used to collect emotion self-reports to train machine learning models for emotion inference. However, as ESM studies are time-consuming and burdensome, participants often withdraw in between. This unplanned withdrawal compels the researchers to discard the dropout participants’ data, significantly impacting the quality and quantity of the self-reports. To address this problem, we leverage only the self-reporting similarity across participants (unlike prior works that apply different machine learning approaches on additional modalities) for missing self-report estimation. In specific, we propose a Multi-task Learning (MTL) framework, MUSE, that constructs the missing self-reports of the dropout participants. We evaluate MUSE in two in-the-wild studies (N1=24, N2=30) of 6-week and 8-week duration, during which the participants reported four emotions (happy, sad, stressed, relaxed) using a smartphone application. The evaluation reveals that MUSE estimates the missing emotion self-reports with an average AUCROC of 84% (Study I) and 82% (Study II). A follow-up evaluation of MUSE for an emotion inference (downstream) task reveals no significant difference in emotion inference performance when estimated self-reports are used. These findings underscore the utility of MUSE in estimating missing self-reports in ESM studies and the applicability of MUSE for downstream tasks (e.g., emotion inference). Surjya Ghosh, Salma Mandi, Sougata Sen, Bivas Mitra, Pradipta De |
CHI | 3 |
| 2024 | Poster: HarvNet: Battery-Free Device Network SimulatorabstractUbiquitous sensing technologies, leveraging a network of interconnected sensors and devices, offer multifaceted benefits to society. However, the use of batteries has persistently posed a challenge to their advancement. In recent years, researchers have explored establishing communication between battery-free nodes. The primary objective of this work is to expedite the testing of algorithms for multiple battery-free interconnected nodes by isolating the algorithm from the underlying hardware. Isolating the hardware allows faster tuning of algorithms, and enables testing on a large scale. We developed a novel python-based simulation framework, HarvNet. Simulations using real-world power traces demonstrated a success rate of 81% in establishing connections between nodes which closely emulates the success rate achieved using hardware. Hrishikesh Govindrao Kusneniwar, Sougata Sen, Josiah D. Hester |
MobiSys | 2 |
| 2024 | Poster: An Automated Method to Detect Tooth Brushing Activity with Smartwatch SensorsabstractOral diseases affect an estimated 3.5 billion people globally, posing significant health challenges. According to the World Health Organization (WHO), adopting self-care practices and maintaining personal oral hygiene can substantially mitigate the prevalence of dental caries. While smartwatches have previously been utilized to track activities of daily living (ADL), their widespread availability has yet to be harnessed for accurately identifying tooth brushing activity among other common ADL. In this Work in Progress (WIP), we demonstrate how motion sensors integrated into smartwatches can effectively distinguish tooth brushing from seven other very similar ADL. We present our initial results that show a promising 94% accuracy with 84% sensitivity. Blake Schleter, Marina Avdonina, Rishiraj Adhikary, Dheryta Jaisinghani, Sougata Sen |
MobiSys | 5 |
| 2024 | Self-SLAM: A Self-supervised Learning Based Annotation Method to Reduce Labeling Overhead
Alfiya M. Shaikh, Hrithik Nambiar, Kshitish Ghate, Swarnali Banik, Sougata Sen, Surjya Ghosh, Vaskar Raychoudhury, Niloy Ganguly, Snehanshu Saha |
ECML/PKDD (9) | 5 |
| 2024 | NIR-sighted: A Programmable Streaming Architecture for Low-Energy Human-Centric Vision ApplicationsabstractHuman studies often rely on wearable lifelogging cameras that capture videos of individuals and their surroundings to aid in visual confirmation or recollection of daily activities like eating, drinking, and smoking. However, this may include private or sensitive information that may cause some users to refrain from using such monitoring devices. Also, short battery lifetime and large form factors reduce applicability for long-term capture of human activity. Solving this triad of interconnected problems is challenging due to wearable embedded systems’ energy, memory, and computing constraints. Inspired by this critical use case and the unique design problem, we developed NIR-sighted, an architecture for wearable video cameras that navigates this design space via three key ideas: (i) reduce storage and enhance privacy by discarding masked pixels and frames, (ii) enable programmers to generate effective masks with low computational overhead, and (iii) enable the use of small MCUs by moving masking and compression off-chip. Combined together in an end-to-end system, NIR-sighted’s masking capabilities and off-chip compression hardware shrinks systems, stores less data, and enables programmer-defined obfuscation to yield privacy enhancement. The user’s privacy is enhanced significantly as nowhere in the pipeline is any part of the image stored before it is obfuscated. We design a wearable camera called NIR-sightedCam based on this architecture; it is compact and can record IR and grayscale video at 16 and 20+ fps, respectively, for 26 hours nonstop (59 hours with IR disabled) at a fraction of comparable platforms power draw. NIR-sightedCam includes a low-power Field Programmable Gate Array that implements our mJPEG compress/obfuscate hardware, Blindspot. We additionally show the potential for privacy-enhancing function and clinical utility via an in-lab eating study, validated by a nutritionist. John Mamish, Rawan Alharbi, Sougata Sen, Shashank Holla, Panchami Kamath, Yaman Sangar, Nabil Alshurafa, Josiah D. Hester |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | AffectPro: Towards Constructing Affective Profile Combining Smartphone Typing Interaction and Emotion Self-reporting PatternabstractThe ubiquity of smartphones and the widespread usage of text entry by soft keyboard in different instant messaging applications (e.g., WhatsApp, FB messenger) have opened the possibilities of inferring emotions from longitudinal typing data. To build this emotion inference engine, we apply machine learning models on features extracted from user’s typing patterns (not content). However, one major challenge encountered while developing the emotion inference model is the requirement of individual training data as typing patterns are often person-specific. In this paper, we investigate the possibility of combining typing pattern with emotion self-reporting to identify a group of similar users so that the training data among these users can be shared to fulfill the requirement of personalized dataset. We develop a framework AffectPro, which quantifies the typing interaction behavior (e.g., typing speed, error rate) and self-reporting pattern (e.g., emotion state transition probability) to construct the affective profiles of users. We evaluated AffectPro in a 6-week in-the-wild study involving 28 users, who used an Android application encompassing a custom keyboard to perform all their typing activities, and to report their instantaneous emotions. We extracted different typing signatures and self-report behavior details from the collected dataset (≈ 5000 typing sessions, ≈ 108 hours of typing data) to construct the affective profile of users. Our results demonstrate similarity across users in terms of typing signature, emotion self-reporting pattern, and a combination of both; which can be leveraged to share training data among similar users to overcome the challenges of personalized data collection. Satchit Hari, Ajay Bhardwaj, Sayan Sarcar, Sougata Sen, Surjya Ghosh |
ICMI | 4 |
| 2022 | ActiSight: Wearer Foreground Extraction Using a Practical RGB-Thermal WearableabstractWearable cameras provide an informative view of wearer activities, context, and interactions. Video obtained from wearable cameras is useful for life-logging, human activity recognition, visual confirmation, and other tasks widely utilized in mobile computing today. Extracting foreground information related to the wearer and separating irrelevant background pixels is the fundamental operation underlying these tasks. However, current wearer foreground extraction methods that depend on image data alone are slow, energy-inefficient, and even inaccurate in some cases, making many tasks–like activity recognition–challenging to implement in the absence of significant computational resources. To fill this gap, we built ActiSight, a wearable RGB-Thermal video camera that uses thermal information to make wearer segmentation practical for body-worn video. Using ActiSight, we collected a total of 59 hours of video from 6 participants, capturing a wide variety of activities in a natural setting. We show that wearer foreground extracted with ActiSight achieves a high dice similarity score while significantly lowering execution time and energy cost when compared with an RGB-only approach. Rawan Alharbi, Sougata Sen, Ada Ng, Nabil Alshurafa, Josiah D. Hester |
PerCom | 2 |
| 2021 | Exploring the Challenges of Using Food Journaling Apps: A Case-study with Young Adults
Tejal Lalitkumar Karnavat, Jaskaran Singh Bhatia, Surjya Ghosh, Sougata Sen |
MobiQuitous | 4 |
| 2021 | Recurring verification of interaction authenticity within bluetooth networksabstractAlthough user authentication has been well explored, device-to-device authentication - specifically in Bluetooth networks - has not seen the same attention. We propose Verification of Interaction Authenticity (VIA) - a recurring authentication scheme based on evaluating characteristics of the communications (interactions) between devices. We adapt techniques from wireless traffic analysis and intrusion-detection systems to develop behavioral models that capture typical, authentic device interactions (behavior); these models enable recurring verification of device behavior. To evaluate our approach we produced a new dataset consisting of more than 300 Bluetooth network traces collected from 20 Bluetooth-enabled smart-health and smart-home devices. In our evaluation, we found that devices can be correctly verified at a variety of granularities, achieving an F1-score of 0.86 or better in most cases. Travis Peters, Timothy J. Pierson, Sougata Sen, José Camacho 0001, David Kotz |
WISEC | 3 |
| 2021 | VibeRing: Using vibrations from a smart ring as an out-of-band channel for sharing secret keys
Sougata Sen, David Kotz |
Pervasive Mob. Comput. | 1 |
| 2020 | Continuous Detection of Physiological Stress with Commodity HardwareabstractTimely detection of an individual’s stress level has the potential to improve stress management, thereby reducing the risk of adverse health consequences that may arise due to mismanagement of stress. Recent advances in wearable sensing have resulted in multiple approaches to detect and monitor stress with varying levels of accuracy. The most accurate methods, however, rely on clinical-grade sensors to measure physiological signals; they are often bulky, custom made, and expensive, hence limiting their adoption by researchers and the general public. In this article, we explore the viability of commercially available off-the-shelf sensors for stress monitoring. The idea is to be able to use cheap, nonclinical sensors to capture physiological signals and make inferences about the wearer’s stress level based on that data. We describe a system involving a popular off-the-shelf heart rate monitor, the Polar H7; we evaluated our system with 26 participants in both a controlled lab setting with three well-validated stress-inducing stimuli and in free-living field conditions. Our analysis shows that using the off-the-shelf sensor alone, we were able to detect stressful events with an F 1-score of up to 0.87 in the lab and 0.66 in the field, on par with clinical-grade sensors. Varun Mishra 0001, Gunnar Pope, Sarah E. Lord, Stephanie Lewia, Byron Lowens, Kelly Caine, Sougata Sen, Ryan J. Halter, David Kotz |
ACM Trans. Comput. Heal. | 7 |
| 2020 | Annapurna: An automated smartwatch-based eating detection and food journaling system
Sougata Sen, Vigneshwaran Subbaraju, Archan Misra, Rajesh Krishna Balan, Youngki Lee 0001 |
Pervasive Mob. Comput. | 1 |
| 2018 | I4S: capturing shopper's in-store interactionsabstractIn this paper, we present I4S, a system that identifies item interactions of customers in a retail store through sensor data fusion from smartwatches, smartphones and distributed BLE beacons. To identify these interactions, I4S builds a gesture-triggered pipeline that (a) detects the occurrence of "item picks", and (b) performs fine-grained localization of such pickup gestures. By analyzing data collected from 31 shoppers visiting a midsized stationary store, we show that we can identify person-independent picking gestures with a precision of over 88%, and identify the rack from where the pick occurred with 91%+ precision (for popular racks). Sougata Sen, Archan Misra, Vigneshwaran Subbaraju, Karan Grover, Meera Radhakrishnan, Rajesh Krishna Balan, Youngki Lee 0001 |
UbiComp | 1 |
| 2018 | Annapurna: Building a Real-World Smartwatch-Based Automated Food JournalabstractWe describe the design and implementation of a smartwatch-based, completely unobtrusive, food journaling system, where the smartwatch helps to intelligently capture useful images of food that an individual consumes throughout the day. The overall system, called Annapurna, is based on three key components: (a) a smartwatch-based gesture recognizer to identify eating gestures, (b) a smartwatch-based image capturer that obtains a small set of relevant and useful images with a low energy overhead, and (c) a server-based image filtering engine that removes irrelevant uploaded images, and then catalogs them through a portal. Our primary challenge is to make the system robust to the huge diversity in natural eating habits and food choices. We show how we address this by an appropriate coupling between a smartwatch's camera sensor and inertial sensor-based tracking of eating gestures, thereby helping to capture multiple likely-to-be-useful images with low energy overhead. Through a series of real-world, in-the-wild studies, we demonstrate the end-to-end working of Annapurna, which captures useful images in over 95% of all natural eating episodes. Sougata Sen, Vigneshwaran Subbaraju, Archan Misra, Rajesh Krishna Balan, Youngki Lee 0001 |
WOWMOM | 1 |
| 2017 | Cloud-based query evaluation for energy-efficient mobile sensing
Tianli Mo, Lipyeow Lim, Sougata Sen, Archan Misra, Rajesh Krishna Balan, Youngki Lee 0001 |
Pervasive Mob. Comput. | 3 |
| 2015 | Using infrastructure-provided context filters for efficient fine-grained activity sensingabstractWhile mobile and wearable sensing can capture unique insights into fine-grained activities (such as gestures and limb-based actions) at an individual level, their energy overheads are still prohibitive enough to prevent them from being executed continuously. In this paper, we explore practical alternatives to addressing this challenge-by exploring how cheap infrastructure sensors or information sources (e.g., BLE beacons) can be harnessed with such mobile/wearable sensors to provide an effective solution that reduces energy consumption without sacrificing accuracy. The key idea is that many fine-grained activities that we desire to capture are specific to certain location, movement or background context: infrastructure sensors and information sources (e.g., BLE beacons) offer practical and cheap ways to identify such context. In this paper, we first explore how various infrastructure, mobile & wearable sensors can be used to identify fine-grained location/movement context (e.g., transiting through a door). We then show, using a couple of illustrative examples (specifically, the detection of `switch pressing' before exiting a room and the identification of `water drinking' after approaching a water cooler) to show that such background context can be predicted, with sufficient accuracy, with sufficient lead time to enable a `triggered' model for mobile/wearable sensing of such microscopic, transient gestures and activities. Moreover, such `triggered' sensing also helps to improve the accuracy of such microscopic gesture recognition, by reducing the set of candidate activity labels. Empirical experiments show that we are able to identify 82.2% of switch-pressing and 91.73% of water-drinking activities in a campus lab setting, with a significant reduction in active sensing time (up to 92.9% compared to continuous sensing). Vigneshwaran Subbaraju, Sougata Sen, Archan Misra, Satyadip Chakraborti, Rajesh Krishna Balan |
PerCom | 2 |
| 2014 | Cloud-Based Query Evaluation for Energy-Efficient Mobile SensingabstractIn this paper, we reduce the energy overheads of continuous mobile sensing for context-aware applications that are interested in collective context or events. We propose a cloud-based query management and optimization framework, called CloQue, which can support concurrent queries, executing over thousands of individual smartphones. CloQue exploits correlation across context of different users to reduce energy overheads via two key innovations: i) Dynamically reordering the order of predicate processing to preferentially select predicates with not just lower sensing cost and higher selectivity, but that maximally reduce the uncertainty about other context predicates, and ii) intelligently propagating the query evaluation results to dynamically update the uncertainty of other correlated, but yet-to-be evaluated, context predicates. An evaluation, using real cell phone traces from a real world dataset shows significant energy savings (between 30 to 50% compared with traditional short-circuit systems) with little loss in accuracy (5% at most). Tianli Mo, Sougata Sen, Lipyeow Lim, Archan Misra, Rajesh Krishna Balan, Youngki Lee 0001 |
MDM (1) | 2 |