VLDB 2026 Research / reviewers in the wild / expert
Indrajeet Ghosh
dblp:272/1993
· DBLP profile ↗
17ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0003-2868-3766ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MorsEar: Toward Generalizable Low-Resource Covert Messaging via Earable based Inertial SensingabstractSilent, eyes-free text entry remains challenging when speech and conventional touch input are impractical. Prior wearable systems often required custom sensors or limited users to a small vocabulary. We present MorsEar, an IMU-only earable framework that maps near-ear micro-gestures such as taps for dot/dash; slide/pull/circle for space/delete/send into character-level Morse, enabling unrestricted character composition while using a compact lexicon solely for lightweight on-device autocorrect. The result is a low-bandwidth, reduced-exposure communication channel that works eyes-free and voice-free in accessibility scenarios, silent zones, and constrained environments. MorsEar infers words using a physics-aware preprocessing stack and compact CNN feed a tempo-adaptive segmentation with rolling buffers; an on-device decoder with lightweight autocorrect provides real-time feedback entirely on-phone. In a 24-participant study (with four accessibility users) across Silent, Cafe, and Metro, MorsEar achieved CER 7.3% and WER 12.5% → 7.8% (Autocorrect), with median 9.3/9.1/5.8 WPM, respectively. Similar to other accessibility-oriented encodings such as Braille, Morse requires a brief familiarization period to learn the timing and rhythm of dots and dashes; after which, MorsEar shows that commodity earable IMUs can support discreet, low-exposure text entry that scales beyond discrete commands to language-level interaction. Garvit Chugh, Indrajeet Ghosh, Nirmalya Roy, Sandip Chakraborty 0001, Suchetana Chakraborty |
CHI | 2 |
| 2026 | Fed-CASQ: Enhancing Class-Wise Accuracy in Pervasive Federated Learning with Class-Aware Scaling and QuantizationabstractFederated Learning (FL) enables collaborative machine learning across decentralized devices and data sources, but resource constraints on pervasive devices necessitate efficient model compression. Existing approaches, such as quantization for on-device training, often degrade accuracy, especially for classes that are difficult to learn due to imbalance, poor-quality samples, or inherent complexity. This results in persistent accuracy gaps across classes. We propose Fed-CASQ's a novel framework that couples class-aware strategies into the quantization process to jointly improve efficiency and accuracy in pervasive FL. Unlike prior works that address quantization and imbalance separately, Fed-CASQ adaptively selects quantization levels based on device resources and leverages Layer-wise Relevance Propagation (LRP) to assess class-relevant convolutional neural network (CNN) filters on the client side. An adaptive weight scaling mechanism is then applied to amplify critical information for low-accuracy classes before aggregation. At the server, a complementary novel aggregation strategy mitigates global imbalance across clients, ensuring that underperforming classes receive proportional attention during model updates. We theoretically establish that Fed-CASQ achieves a convergence rate of ${\mathcal{O}}\left({\frac{{\kappa *\hat \sigma *\hat \delta }}{{\sqrt T }}}\right)$ under non-convex settings. We empirically establish that quantization directly influences the performance of under sampled (minority) classes. Experimental results further show that Fed-CASQ substantially narrows the performance gap for low-accuracy classes, improving their accuracy by ≈30%, while reducing training latency by over 56% on resource-constrained pervasive devices. Emon Dey, Anuradha Ravi, Gaurav Shinde, Garvit Chugh, Indrajeet Ghosh, Archan Misra, Nirmalya Roy |
PerCom | 5 |
| 2026 | COHORT: Hybrid RL for Collaborative Large DNN Inference on Multi-Robot Systems Under Real-Time Constraints
Mohammad Saeid Anwar, Anuradha Ravi, Indrajeet Ghosh, Gaurav Shinde, Carl E. Busart, Nirmalya Roy |
WoWMoM | 3 |
| 2026 | Efficient personalized image memorability via gaze-guided semantic cloning distillationabstractAbstract Image memorability (IM) estimation typically relies on learning generic semantic features from large-scale datasets; however, memorability is intrinsically individual-dependent and personalized visual viewing behavior, but overlooking this variability can undermine model performance in downstream cognitive and human–machine interaction tasks, leading toward suboptimal performance. To address this, we propose PerMem , a unified framework that integrates knowledge distillation (KD) and imitation learning (IL) to jointly learn generic and personalized salient representations for memorability estimation by leveraging both image content and gaze-derived heatmaps. PerMem employs a teacher network built upon a pre-trained ResNet-50 backbone, followed by an encoder–decoder architecture coupled with spatial and channel attention mechanisms to generate generic saliency-aware memorability maps. A lightweight, attention-guided student encoder–decoder is then optimized through a composite imitation-guided distillation process, where knowledge is distilled from the teacher while simultaneously imitating user-specific gaze fixation heatmaps. Through this joint training process, the student network learns to produce personalized memorability estimates while achieving substantial reductions in computational complexity. We validate PerMem on two public IM estimation datasets (LaMem and SUN) and a in-house WoM dataset comprising of 45 participants (UMBC IRB #670) engaged in visual search and navigation tasks that reflect individualized visual attention patterns. PerMem outperforms nine state-of-the-art IM models by capturing coarse-to-fine saliency and adapting to individual attention, achieving $$\approx $$ ≈ 6% improvement in memorability prediction. Lastly, we evaluate PerMem on heterogeneous embedded edge devices, including Jetson Nano, Jetson Xavier NX, and Raspberry Pi, demonstrating efficient on-device inference with consistent reductions in memory usage (25.6–33.4%), power consumption (19.0–25.0%), and inference latency (34.9–39.8%) relative to the teacher model, highlighting its practical robustness for resource-constrained, human-centric applications. Indrajeet Ghosh, Mohammad Saeid Anwar, Kasthuri Jayarajah, Nirmalya Roy |
Knowl. Inf. Syst. | 1 |
| 2025 | High-Order Moments Conditional Domain Adaptation Networks for Wearable Human Activity RecognitionabstractDeveloping scalable wearable human activity recognition (wHAR) models is challenging due to domain shifts that substantially degrade performance across downstream tasks. Unsupervised domain adaptation (UDA) seeks to improve generalization by transferring knowledge from labeled source domains to unlabeled target domains. However, conventional UDA methods primarily align marginal feature distributions while neglecting feature-label dependencies, often leading to negative transfer and sub-optimal performance. Motivated by these limitations, we propose a novel optimization framework that tackles two key challenges: (i) generating reliable pseudo-labels for the unlabeled target domain and (ii) minimizing conditional discrepancies across domains. To address (i), we employ temperature-based entropy minimization (TEM), which calibrates prediction confidence by scaling logits with a temperature parameter to produce robust pseudo-labels. For (ii), we introduce a polynomial kernel-based cross-covariance (PkCC) loss, a high-order statistics-driven approach that maps features into a reproducing kernel hilbert space (RKHS) to capture richer feature-label dependencies and reduce conditional distribution gaps between domains. In addition, we demonstrate that CoDAN readily extends to partial UDA (pUDA), where the target label space is a subset of the source, and extensive evaluations on public wHAR datasets with diverse label spaces validate its superior performance over state-of-the-art methods in both UDA and pUDA scenarios. Indrajeet Ghosh, Garvit Chugh, Abu Zaher Md Faridee, Nirmalya Roy |
CIKM | 1 |
| 2025 | Imitation-Inspired Semantic-Guided Distillation for User-Conditioned Memorability PredictionabstractImage memorability (IM) estimation typically relies on learning generic semantic features from large-scale datasets; however, memorability is intrinsically individual-dependent and personalized visual viewing behavior, but overlooking this variability can undermine model performance in downstream cognitive and human-machine interaction tasks, leading towards suboptimal performance. To address this, we propose MemGaze, a unified framework that integrates knowledge distillation (KD) and imitation learning (IL) to jointly learn generic and personalized salient representations for memorability estimation by leveraging both image content and gaze-derived heatmaps. MemGaze employs a teacher network built upon a pretrained ResNet-50 backbone, followed by an encoder-decoder architecture coupled with spatial and channel attention mechanisms to generate generic saliency-aware memorability maps. A lightweight, attention-guided student encoder-decoder is then optimized through a composite imitation-guided distillation process, where knowledge is distilled from the teacher while simultaneously imitating user-specific gaze fixation heatmaps. Through this joint training process, the student network learns to produce personalized memorability estimates while achieving substantial reductions in computational complexity. We validate MemGaze on two public IM estimation datasets (LaMem and SUN) and a inhouse WoM dataset comprising of 45 participants (UMBC IRB #670) engaged in visual search and navigation tasks that reflect individualized visual attention patterns. MemGaze outperforms nine state-of-the-art IM models by capturing coarse-to-fine saliency and adapting to individual attention, achieving ≈6% improvement in memorability prediction. Indrajeet Ghosh, Mohammad Saeid Anwar, Kasthuri Jayarajah, Nirmalya Roy |
ICDM | 1 |
| 2025 | BiteSense: Earable-Based Inertial Sensing for Eating Behaviour AssessmentabstractAutomated dietary monitoring is essential for gaining insights into eating behaviors, especially for managing chronic conditions such as obesity, diabetes, and hypercholesterolemia. Earable-based inertial sensing has been found promising for detecting chewing and eating activities; however, further insights like what, when, and how much is being eaten are crucial information for effective dietary assessment. Therefore, we propose BiteSense, an earable-based system that leverages inertial sensors (IMU) to monitor food intake and classify various food types. Using a hierarchical classification model, the system analyzes masticatory kinematics to detect food states, textures, nutritional value, and cooking methods, ultimately identifying specific foods consumed, as well as estimating food intake amount and meal type. A semi-controlled user study involving 38 participants from diverse backgrounds demonstrated the system’s high accuracy, with an F1 score of 0.86 for detecting the masticatory process using a leave-one-subject-out (LOSO) approach, while exhibiting significant improvement over benchmark algorithms in extensive experiments by 8-12%. Garvit Chugh, Indrajeet Ghosh, Sandip Chakraborty 0001, Suchetana Chakraborty |
PerCom | 2 |
| 2025 | Augmenting Personalized Memory via Practical Multimodal Wearable Sensing in Visual Search and Wayfinding NavigationabstractACM UMAP June 16-19 2025, New York, USA Indrajeet Ghosh, Kasthuri Jayarajah, Nicholas R. Waytowich, Nirmalya Roy |
UMAP | 1 |
| 2024 | Unsupervised Domain Adaptation for Action Recognition via Self-Ensembling and Conditional Embedding AlignmentabstractRecent advancements in deep learning-based wearable human action recognition (wHAR) have improved the capture and classification of complex motions, but adoption remains limited due to the lack of expert annotations and domain discrepancies from user variations. Limited annotations hinder the model's ability to generalize to out-of-distribution samples. While data augmentation can improve generalizability, unsupervised augmentation techniques must be applied carefully to avoid introducing noise. Unsupervised domain adaptation (UDA) addresses domain discrepancies by aligning conditional distributions with labeled target samples, but vanilla pseudo-labeling can lead to error propagation. To address these challenges, we propose μDAR, a novel joint optimization architecture comprised of three functions: (i) consistency regularizer between augmented samples to improve model classification generalizability, (ii) temporal ensemble for robust pseudo-label generation and (iii) conditional distribution alignment to improve domain generalizability. The temporal ensemble works by aggregating predictions from past epochs to smooth out noisy pseudo-label predictions, which are then used in the conditional distribution alignment module to minimize kernel-based class-wise conditional maximum mean discrepancy (kCMMD) between the source and target feature space to learn a domain invariant embedding. The consistency-regularized augmentations ensure that multiple augmentations of the same sample share the same labels; this results in (a) strong generalization with limited source domain samples and (b) consistent pseudo-label generation in target samples. The novel integration of these three modules in μDAR results in a range of ~ 4-12% average macro-F1 score improvement over six state-of-the-art UDA methods in four benchmark wHAR datasets. Indrajeet Ghosh, Garvit Chugh, Abu Zaher Md Faridee, Nirmalya Roy |
ICDM | 1 |
| 2024 | EEGAmp+: Investigating the Efficacy of Functional Connectivity for Detecting Events in Low Resolution EEGabstractElectroencephalography (EEG) has found many applications cutting across many domains, such as digital health, affective computing, and human-machine interfaces. However, its widespread adoption in practice has been primarily inhibited by its susceptibility to noise artifacts and the low spatial resolution of electrodes on commercial EEG sensors. While several prior works have investigated techniques for detecting and extracting noise, our understanding of performance degradation due to electrode sparsity remains limited. In this work, we explore the feasibility of using Functional Connectivity (FC) for improving the EEG sensing-based accuracy using two exemplars downstream, working memory-related tasks: (a) cognitive task load (CTL) assessment and (b) high attentional event-evoked potential (EEP) episodes detection. This paper proposes an integrated approach, EEGAmp+ , that first utilizes channel-wise functional connectivity modules using independent component analysis (ICA) coupled with cosine distance for EEG signal reconstruction for cognitive task load assessment tasks. This is then coupled with a sliding window change point detection technique paired with continuous wavelet transformation (CWT) to extract high attentional EEP episodes. Our empirical results indicate that using independent component analysis (ICA) coupled with FC to improve spatial resolution increased cognitive load assessment accuracy by [5.6% ± 1.13] across four machine learning algorithms. Furthermore, after signal reconstruction, we introduce sliding window CPD coupled with CWT, which allows us to extract EEP segments legibly through decomposing the signals and the ability to capture both time and frequency representation from the reconstructed signal boosting detection accuracy by [11.1% ±1.31]. Indrajeet Ghosh, Kasthuri Jayarajah, Nicholas R. Waytowich, Nirmalya Roy |
MobiQuitous | 1 |
| 2023 | BeautyNet: A Makeup Activity Recognition Framework using Wrist-worn SensorabstractThe significance of enhancing facial features has grown increasingly in popularity among all groups of people bringing a surge in makeup activities. The makeup market is one of the most profitable and founding sectors in the fashion industry which involves product retailing and demands user training. Makeup activities imply exceptionally delicate hand movements and require much training and practice for perfection. However, the only available choices in learning makeup activities are hands-on workshops by professional instructors or, at most, video-based visual instructions. None of these exhibits immense benefits to beginners, or visually impaired people. One can consistently watch and listen to the best of their abilities, but to precisely practice, perform, and reach makeup satisfaction, recognition from an IoT (Internet-of-Things) device with results and feedback would be the utmost support. In this work, we propose a makeup activity recognition framework, BeautyNet which detects different makeup activities from wrist-worn sensor data collected from ten participants of different age groups in two experimental setups. Our framework employs a LSTM-autoencoder based classifier to extract features from the sensor data and classifies five makeup activities (i.e., applying cream, lipsticks, blusher, eyeshadow, and mascara) in controlled and uncontrolled environment. Empirical results indicate that BeautyNet achieves 95% and 93% accuracy for makeup activity detection in controlled and uncontrolled settings, respectively. In addition, we evaluate BeautyNet with various traditional machine learning algorithms using our in-house dataset and noted an increase in accuracy by ≈ 4-7%. Fatimah Albargi, Naima Khan, Indrajeet Ghosh, Ahana Roy |
SMARTCOMP | 3 |
| 2023 | HeteroSys: Heterogeneous and Collaborative Sensing in the WildabstractAdvances in Internet-of-Things, artificial intelligence, and ubiquitous computing technologies have contributed to building the next generation of context-aware heterogeneous systems with robust interoperability to control and monitor the environmental variables of smart environments. Motivated by this, we propose HeteroSys, an end-to-end multi-functional smart IoT-based system prototype for heterogeneous and collaborative sensing in a smart IoT-based environment. A unique characteristic of HeteroSys is that it relies on Home Assistant (HA) to collate heterogeneous sensors (e.g., passive infrared sensors (PIR), reed (door) switches, object tags, wearable wrist-mounted, water leak sensors, and internet protocol cameras), and uses a variety of networking protocols such as Zigbee open standard for mesh networking, WiFi, and Bluetooth Low Energy (BLE) for communication. The reliance on HA (and its broad community support) makes HeteroSys ideal for various applications such as object detection, human activity recognition and behavior patterns. We articulated the development phase, integration, testing challenges and evaluation of the HeteroSys. We conducted an extensive 24-hour longitudinal data collection from 5 participants performing 6 activities by deploying in an indoor home environment. Our assessment of the acquired dataset reveals that the representations learned using deep learning architecture aid in improving the detection of activities to 83.1% accuracy. Indrajeet Ghosh, Adam Goldstein, Avijoy Chakma, Jade Freeman, Timothy Gregory, Niranjan Suri, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy |
SMARTCOMP | 1 |
| 2022 | PerMTL: A Multi-Task Learning Framework for Skilled Human Performance AssessmentabstractIntelligent and complex human motion analysis can help design the next generation IoT and AR/VR systems for automated human performance assessment. Such an automated system can help advocate the interpretability and translatability of complex human motions, intelligent motion feedback, and fine-grained motion skill assessment to design next-generation interactive human-machine teaming systems. Motivated by this, we design a wearable sensing framework for assessing the players’ performance and consider a live badminton game as our use case. Generally, the players on the field try to improve their performance by focusing on fast and synchronous coordination of their limbs’ reflex actions to have the ideal body postures to perform the desired shot. Learning the minute dissimilarities and distinctive traits from each limb of the players simultaneously can help assess the players’ performance and specific skillsets during a game. This paper proposes a multi-task learning framework, PerMTL to learn the shared features from each player’s limb. The PerMTL comprises a task-specific regressor output layer that helps to determine the dissimilarities and distinctive traits between the player’s limbs for collective inference in a body sensor network (BSN) environment. We evaluate the PerMTL framework using publicly available Badminton Activity Recognition (BAR) and Daily and Sports Activities (DSA) datasets. Empirical results indicate that PerMTL achieves R2Score of ≈ 82% in predicting the players’ performance. Indrajeet Ghosh, Avijoy Chakma, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy, Nicholas R. Waytowich |
ICMLA | 1 |
| 2022 | SpecTextor: End-to-End Attention-based Mechanism for Dense Text Generation in Sports JournalismabstractLanguage-guided smart systems can help to design next-generation human-machine interactive applications. The dense text description is one of the research areas where systems learn the semantic knowledge and visual features of each video frame and map them to describe the video's most relevant subjects and events. In this paper, we consider untrimmed sports videos as our case study. Generating dense descriptions in the sports domain to supplement journalistic works without relying on commentators and experts requires more investigation. Motivated by this, we propose an end-to-end automated text-generator, SpecTextor, that learns the semantic features from untrimmed videos of sports games and generates associated descriptive texts. The proposed approach considers the video as a sequence of frames and sequentially generates words. After splitting videos into frames, we use a pre-trained VGG-16 model for feature extraction and encoding the video frames. With these encoded frames, we posit a Long Short-Term Memory (LSTM) based attention-decoder pipeline that leverages soft-attention mechanism to map the semantic features with relevant textual descriptions to generate the explanation of the game. Because developing a comprehensive description of the game warrants training on a set of dense time-stamped captions, we leverage two available public datasets: ActivityNet Captions and Microsoft Video Description. In addition, we utilized two different decoding algorithms: beam search and greedy search and computed two evaluation metrics: BLEU and METEOR scores. Indrajeet Ghosh, Matthew Ivler, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy |
SMARTCOMP | 1 |
| 2022 | DeCoach: Deep Learning-based Coaching for Badminton Player AssessmentabstractWearable devices have gained immense popularity among various pervasive computing and Internet-of-Things (IoT) applications in the past decade. Sports analytics researchers recently focused on improving a player’s performance to help devise a winning strategy based on the player’s gameplay. Especially in a racquet-based badminton sport, it is assumed that handling the racquet during the gameplay is one of the primary reasons to influence the players’ performance. On the contrary, we posit that the players’ stance, body movements, and posture are equally significant in evaluating a player’s performance during the game. A shot characterized by a recommended posture, stance, and body movements allows a player to play a stroke efficiently, thus aiding the player in guiding the shuttle to strategic spots and making it difficult for the opponent to return the shot and score a point. Relying on this hypothesis, we propose DeCoach, a data-driven framework that leverages the stance and posture of the players and ranks them based on their performances. In this effort, we first employ a deep learning-based algorithm to classify the strokes and stances of the players. Secondly, we propose a distance-based methodology to compare the obtained stance of a player with that of a professional player. Finally, we devise a deep learning-based regressor to predict the player’s performance which commences with ranking based on their performance. We evaluate DeCoach using our in-house dataset, Badminton Activity Recognition (BAR) Dataset that is collected using inertial measurement unit (IMU) sensors by placing them on the upper and lower limbs of the players. The BAR dataset is collected from 11 players in the controlled and uncontrolled environment settings for 12 frequently played shots in the game. Empirical results indicate that DeCoach achieves 89.09% accuracy for strokes detection and R2 score of 88.84% in estimating the players’ performance. Indrajeet Ghosh, Sreenivasan Ramasamy Ramamurthy, Avijoy Chakma, Nirmalya Roy |
Pervasive Mob. Comput. | 1 |
| 2022 | STAR-Lite: A light-weight scalable self-taught learning framework for older adults' activity recognition
Sreenivasan Ramasamy Ramamurthy, Indrajeet Ghosh, Aryya Gangopadhyay, Elizabeth Galik, Nirmalya Roy |
Pervasive Mob. Comput. | 2 |
| 2021 | STAR: A Scalable Self-taught Learning Framework for Older Adults' Activity RecognitionabstractActivity Recognition (AR) in older adults living with Neurocognitive disorders caused by diseases such as Alzheimer’s is still a challenging research problem. The inherent natural variation in performing an activity increases while repeating the same activity for an older adult, let alone the variation introduced when another older adult performs the same activity. Moreover, the challenges in acquiring the labeled data while preserving the privacy, availability of annotators with domain knowledge, aversion towards cameras even for a minimal amount of time for ground truth data collection, and psychological and mental health status make AR for older adults challenging. In this paper, we postulate a self-taught learning-based approach that helps recognize activities with variations that are not being directly seen during the training phase. We hypothesize that the features extracted using deep architectures from unlabeled data instances can learn general underlying representations of activities efficiently and help improve activity classification in a supervised setting, although the data instances in labeled data do not follow the generative distribution of that of unlabeled data. We posit real data from a retirement community center using our in-house SenseBox infrastructure and survey-based assessments concurrently done by a clinical evaluator to study the relationship between activities and functional/behavioral health of older adults. We evaluate our proposed self-taught learning-based approach, STAR, using the presented in-house Alzheimer’s Activity Recognition (AAR) dataset acquired in a real-world deployment in 25 homes which outperforms the state-of-the-art algorithm by about 20%. Sreenivasan Ramasamy Ramamurthy, Indrajeet Ghosh, Aryya Gangopadhyay, Elizabeth Galik, Nirmalya Roy |
SMARTCOMP | 2 |