Nirmalya Roy

dblp:79/1526 · DBLP profile ↗
← Back
107ranked-venue papers
18as first author
49since 2021 · last 2026
0000-0003-4827-3393ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 1 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 3 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 27 · 9 first-author · 7 since 2021Computer networks · 12 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 7 since 2021Systems, architecture and hardware · 8 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MorsEar: Toward Generalizable Low-Resource Covert Messaging via Earable based Inertial Sensing
abstract
Silent, eyes-free text entry remains challenging when speech and conventional touch input are impractical. Prior wearable systems often required custom sensors or limited users to a small vocabulary. We present MorsEar, an IMU-only earable framework that maps near-ear micro-gestures such as taps for dot/dash; slide/pull/circle for space/delete/send into character-level Morse, enabling unrestricted character composition while using a compact lexicon solely for lightweight on-device autocorrect. The result is a low-bandwidth, reduced-exposure communication channel that works eyes-free and voice-free in accessibility scenarios, silent zones, and constrained environments. MorsEar infers words using a physics-aware preprocessing stack and compact CNN feed a tempo-adaptive segmentation with rolling buffers; an on-device decoder with lightweight autocorrect provides real-time feedback entirely on-phone. In a 24-participant study (with four accessibility users) across Silent, Cafe, and Metro, MorsEar achieved CER 7.3% and WER 12.5% → 7.8% (Autocorrect), with median 9.3/9.1/5.8 WPM, respectively. Similar to other accessibility-oriented encodings such as Braille, Morse requires a brief familiarization period to learn the timing and rhythm of dots and dashes; after which, MorsEar shows that commodity earable IMUs can support discreet, low-exposure text entry that scales beyond discrete commands to language-level interaction.
Garvit Chugh, Indrajeet Ghosh, Nirmalya Roy, Sandip Chakraborty 0001, Suchetana Chakraborty
CHI3
2026 Fed-CASQ: Enhancing Class-Wise Accuracy in Pervasive Federated Learning with Class-Aware Scaling and Quantization
abstract
Federated Learning (FL) enables collaborative machine learning across decentralized devices and data sources, but resource constraints on pervasive devices necessitate efficient model compression. Existing approaches, such as quantization for on-device training, often degrade accuracy, especially for classes that are difficult to learn due to imbalance, poor-quality samples, or inherent complexity. This results in persistent accuracy gaps across classes. We propose Fed-CASQ's a novel framework that couples class-aware strategies into the quantization process to jointly improve efficiency and accuracy in pervasive FL. Unlike prior works that address quantization and imbalance separately, Fed-CASQ adaptively selects quantization levels based on device resources and leverages Layer-wise Relevance Propagation (LRP) to assess class-relevant convolutional neural network (CNN) filters on the client side. An adaptive weight scaling mechanism is then applied to amplify critical information for low-accuracy classes before aggregation. At the server, a complementary novel aggregation strategy mitigates global imbalance across clients, ensuring that underperforming classes receive proportional attention during model updates. We theoretically establish that Fed-CASQ achieves a convergence rate of ${\mathcal{O}}\left({\frac{{\kappa *\hat \sigma *\hat \delta }}{{\sqrt T }}}\right)$ under non-convex settings. We empirically establish that quantization directly influences the performance of under sampled (minority) classes. Experimental results further show that Fed-CASQ substantially narrows the performance gap for low-accuracy classes, improving their accuracy by ≈30%, while reducing training latency by over 56% on resource-constrained pervasive devices.
Emon Dey, Anuradha Ravi, Gaurav Shinde, Garvit Chugh, Indrajeet Ghosh, Archan Misra, Nirmalya Roy
PerCom7
2026 DataCenterGym: A Physics-Grounded Simulator for Multi-Objective Data Center Scheduling
abstract
Modern datacenters schedule heterogeneous workloads across geo-distributed sites with diverse compute capacities, electricity prices, and thermal conditions. Compute utilization, heat generation, cooling demand, and energy consumption are tightly coupled, yet most existing schedulers abstract these effects and treat them independently. We present \textit{DataCenterGym}, a physics-grounded simulation environment for job scheduling in geo-distributed data centers, designed as a reusable testbed for future research. The simulator integrates compute queueing, building thermal dynamics, localized HVAC behavior, and temperature-dependent service degradation within a Gymnasium-compatible interface. We also develop a Hierarchical Model Predictive Control (H-MPC) scheduling algorithm that performs distributed job placement while explicitly accounting for thermal and power dynamics. Through experiments on nominal operation and workload sensitivity, we demonstrate how H-MPC improves scheduling performance relative to baseline schedulers.
Nilavra Pathak, Samadrita Biswas, Nirmalya Roy
SmartComp3
2026 Lorentz Entailment Cone for Semantic Segmentation
abstract
Semantic segmentation in hyperbolic space can capture hierarchical structure in low dimensions with uncertainty quantification. Existing approaches choose the Poincaré ball model for hyperbolic geometry, which suffers from numerical instabilities, optimization, and computational challenges. We propose a novel, tractable, architecture-agnostic semantic segmentation framework in the hyperbolic Lorentz model. We employ text embeddings with semantic and visual cues to guide hierarchical pixel-level representations in Lorentz space. This enables stable and efficient optimization without requiring a Riemannian optimizer, and easily integrates with existing Euclidean architectures. Beyond segmentation, our approach yields free uncertainty estimation, confidence map, boundary delineation, hierarchical and text-based retrieval, and zero-shot performance, reaching generalized flatter minima. We further introduce a novel uncertainty and confidence indicator in Lorentz cone embeddings. Extensive experiments on ADE20K, COCO-Stuff-164k, Pascal-VOC, and Cityscapes with state-of-the-art models (DeepLabV3 and SegFormer) validate the effectiveness and generality of our approach. Our results demonstrate the potential of hyperbolic Lorentz embeddings for robust and uncertainty-aware semantic segmentation. Code is available at https://github.com/mxahan/Lorentz_semantic_segmentation.
Zahid Hasan 0001, Masud Ahmed, Nirmalya Roy
WACV3
2026 COHORT: Hybrid RL for Collaborative Large DNN Inference on Multi-Robot Systems Under Real-Time Constraints
Mohammad Saeid Anwar, Anuradha Ravi, Indrajeet Ghosh, Gaurav Shinde, Carl E. Busart, Nirmalya Roy
WoWMoM6
2026 CAViAR: Quality-Aware Vision-and-Radio Fusion for Relative Range Estimation Among Collaborative Autonomous Agents
Gaurav Shinde, Anuradha Ravi, Jared Lewis, Andre Harrison, Henry Gardiner, Mohammad Saeid Anwar, Shadman Sakib, Jade Freeman, Nirmalya Roy
WoWMoM9
2026 Efficient personalized image memorability via gaze-guided semantic cloning distillation
abstract
Abstract Image memorability (IM) estimation typically relies on learning generic semantic features from large-scale datasets; however, memorability is intrinsically individual-dependent and personalized visual viewing behavior, but overlooking this variability can undermine model performance in downstream cognitive and human–machine interaction tasks, leading toward suboptimal performance. To address this, we propose PerMem , a unified framework that integrates knowledge distillation (KD) and imitation learning (IL) to jointly learn generic and personalized salient representations for memorability estimation by leveraging both image content and gaze-derived heatmaps. PerMem employs a teacher network built upon a pre-trained ResNet-50 backbone, followed by an encoder–decoder architecture coupled with spatial and channel attention mechanisms to generate generic saliency-aware memorability maps. A lightweight, attention-guided student encoder–decoder is then optimized through a composite imitation-guided distillation process, where knowledge is distilled from the teacher while simultaneously imitating user-specific gaze fixation heatmaps. Through this joint training process, the student network learns to produce personalized memorability estimates while achieving substantial reductions in computational complexity. We validate PerMem on two public IM estimation datasets (LaMem and SUN) and a in-house WoM dataset comprising of 45 participants (UMBC IRB #670) engaged in visual search and navigation tasks that reflect individualized visual attention patterns. PerMem outperforms nine state-of-the-art IM models by capturing coarse-to-fine saliency and adapting to individual attention, achieving $$\approx $$ ≈ 6% improvement in memorability prediction. Lastly, we evaluate PerMem on heterogeneous embedded edge devices, including Jetson Nano, Jetson Xavier NX, and Raspberry Pi, demonstrating efficient on-device inference with consistent reductions in memory usage (25.6–33.4%), power consumption (19.0–25.0%), and inference latency (34.9–39.8%) relative to the teacher model, highlighting its practical robustness for resource-constrained, human-centric applications.
Indrajeet Ghosh, Mohammad Saeid Anwar, Kasthuri Jayarajah, Nirmalya Roy
Knowl. Inf. Syst.4
2025 High-Order Moments Conditional Domain Adaptation Networks for Wearable Human Activity Recognition
abstract
Developing scalable wearable human activity recognition (wHAR) models is challenging due to domain shifts that substantially degrade performance across downstream tasks. Unsupervised domain adaptation (UDA) seeks to improve generalization by transferring knowledge from labeled source domains to unlabeled target domains. However, conventional UDA methods primarily align marginal feature distributions while neglecting feature-label dependencies, often leading to negative transfer and sub-optimal performance. Motivated by these limitations, we propose a novel optimization framework that tackles two key challenges: (i) generating reliable pseudo-labels for the unlabeled target domain and (ii) minimizing conditional discrepancies across domains. To address (i), we employ temperature-based entropy minimization (TEM), which calibrates prediction confidence by scaling logits with a temperature parameter to produce robust pseudo-labels. For (ii), we introduce a polynomial kernel-based cross-covariance (PkCC) loss, a high-order statistics-driven approach that maps features into a reproducing kernel hilbert space (RKHS) to capture richer feature-label dependencies and reduce conditional distribution gaps between domains. In addition, we demonstrate that CoDAN readily extends to partial UDA (pUDA), where the target label space is a subset of the source, and extensive evaluations on public wHAR datasets with diverse label spaces validate its superior performance over state-of-the-art methods in both UDA and pUDA scenarios.
Indrajeet Ghosh, Garvit Chugh, Abu Zaher Md Faridee, Nirmalya Roy
CIKM4
2025 Imitation-Inspired Semantic-Guided Distillation for User-Conditioned Memorability Prediction
abstract
Image memorability (IM) estimation typically relies on learning generic semantic features from large-scale datasets; however, memorability is intrinsically individual-dependent and personalized visual viewing behavior, but overlooking this variability can undermine model performance in downstream cognitive and human-machine interaction tasks, leading towards suboptimal performance. To address this, we propose MemGaze, a unified framework that integrates knowledge distillation (KD) and imitation learning (IL) to jointly learn generic and personalized salient representations for memorability estimation by leveraging both image content and gaze-derived heatmaps. MemGaze employs a teacher network built upon a pretrained ResNet-50 backbone, followed by an encoder-decoder architecture coupled with spatial and channel attention mechanisms to generate generic saliency-aware memorability maps. A lightweight, attention-guided student encoder-decoder is then optimized through a composite imitation-guided distillation process, where knowledge is distilled from the teacher while simultaneously imitating user-specific gaze fixation heatmaps. Through this joint training process, the student network learns to produce personalized memorability estimates while achieving substantial reductions in computational complexity. We validate MemGaze on two public IM estimation datasets (LaMem and SUN) and a inhouse WoM dataset comprising of 45 participants (UMBC IRB #670) engaged in visual search and navigation tasks that reflect individualized visual attention patterns. MemGaze outperforms nine state-of-the-art IM models by capturing coarse-to-fine saliency and adapting to individual attention, achieving ≈6% improvement in memorability prediction.
Indrajeet Ghosh, Mohammad Saeid Anwar, Kasthuri Jayarajah, Nirmalya Roy
ICDM4
2025 QPRL : Learning Optimal Policies with Quasi-Potential Functions for Asymmetric Traversal
abstract
Reinforcement learning (RL) in real-world tasks such as robotic navigation often encounters environments with asymmetric traversal costs, where actions like climbing uphill versus moving downhill incur distinctly different penalties, or transitions may become irreversible. While recent quasimetric RL methods relax symmetry assumptions, they typically do not explicitly account for path-dependent costs or provide rigorous safety guarantees. We introduce Quasi-Potential Reinforcement Learning (QPRL), a novel framework that explicitly decomposes asymmetric traversal costs into a path-independent potential function ($\Phi$) and a path-dependent residual ($\Psi$). This decomposition allows efficient learning and stable policy optimization via a Lyapunov-based safety mechanism. Theoretically, we prove that QPRL achieves convergence with improved sample complexity of $\tilde{O}(\sqrt{T})$, surpassing prior quasimetric RL bounds of $\tilde{O}(T)$. Empirically, our experiments demonstrate that QPRL attains state-of-the-art performance across various navigation and control tasks, significantly reducing irreversible constraint violations by approximately $4\times$ compared to baselines.
Jumman Hossain, Nirmalya Roy
ICML2
2025 RespFormer: A Motion-Guided Temporal-Frequency Multimodal Fusion Transformer for Contactless Respiratory Monitoring
abstract
Monitoring respiratory rate (RR) is essential for early identification of respiratory and metabolic abnormalities. However, the limitations of contact-based sensors and the lack of reliability in many contactless methods make continuous and accurate monitoring difficult in non-clinical settings. To address these challenges, we introduce RespFormer, an edge-optimized, motion-guided temporal-frequency multimodal fusion transformer framework for real-time, contactless RR estimation and breathing pattern classification. RespFormer integrates dense optical flow analysis with temporal, statistical, and frequency-domain features derived from video sequences and enhances them through a multi-stage signal processing pipeline. These features are modeled using an ensemble of three time series transformer architectures (ETSformer, Temporal Fusion Transformer, and Informer) to capture distinct aspects of temporal dynamics. A shared attention-based refinement module enhances the feature representations, and final predictions are fused using a stacking-based meta-learner. We validate RespFormer on a multimodal dataset comprising synchronized RGB, NIR, and IR video data, including a custom in-house dataset captured under various conditions. Experimental results demonstrate that RespFormer achieves a mean absolute error (MAE) ≈ 0.98 bpm, improving prediction accuracy by ≈ 11% and reducing memory usage by ≈ 26%, while maintaining real-time inference (≈ 1.22 seconds) on resource-constrained devices. Furthermore, RespFormer accurately classifies breathing patterns (normal, bradypnea, tachypnea, and apnea) with 95% accuracy, underscoring it’s potential for practical application in telemedicine, clinical screening, and low-resource healthcare settings.
Shadman Sakib, Gaurav Shinde, Snehalraj Chugh, Mohammad Saeid Anwar, Nirmalya Roy
ICMLA5
2025 QuasiNav: Asymmetric Cost-Aware Navigation Planning with Constrained Quasimetric Reinforcement Learning
abstract
Autonomous navigation in unstructured outdoor environments is inherently challenging due to the presence of asymmetric traversal costs, such as varying energy expenditures for uphill versus downhill movement. Traditional reinforcement learning methods often assume symmetric costs, which can lead to suboptimal navigation paths and increased safety risks in realworld scenarios. In this paper, we introduce QuasiNav, a novel reinforcement learning framework that integrates quasimetric embeddings to explicitly model asymmetric costs and guide efficient, safe navigation. QuasiNav formulates the navigation problem as a constrained Markov decision process (CMDP) and employs quasimetric embeddings to capture directionally dependent costs, allowing for a more accurate representation of the terrain. We combine this approach with adaptive constraint tightening. This ensures that safety constraints are dynamically enforced during learning. We validate QuasiNav on a Clearpath Jackal robot in three challenging navigation scenarios-undulating terrains, asymmetric hill traversal, and directionally dependent terrain traversal-demonstrating its effectiveness in both simulated and real-world environments. Experimental results show that QuasiNav significantly outperforms conventional methods, achieving higher success rates, improved energy efficiency (13.6 % reduction in energy consumption compared to baseline methods), and better adherence to safety constraints.
Jumman Hossain, Abu Zaher Md Faridee, Derrik E. Asher, Jade Freeman, Theron Trout, Timothy Gregory, Nirmalya Roy
ICRA7
2025 Demo: InvisibleFence: Non-Lethal Edge-Optimized AI for Human Wildlife Coexistence and Crop Protection
abstract
Human-wildlife conflicts in residential/agricultural settings rely on ineffective deterrents like rodenticides or fences. We introduce InvisibleFence, a modular 3D-printed Vision Pod system with off-the-shelf deterrents. The Vision Pod fuses a 2K camera and 240° motion sensing with an edge-optimized pipeline trained on 44,000 wildlife images of eleven classes—achieving 86.7% mAP. Benchmarking YOLO variants (416p–2K) ensures performance. Upon detection, it sends MQTT commands to drive deterrent units—ultrasonic speakers or lighting/spray modules—that emit tones without affecting humans or pets. InvisibleFence creates adaptive zones that reduce false triggers and limit habituation.
Snehalraj Chugh, Elijah Polyakov, Milind Rampure, Bipendra Basnyat, Nirmalya Roy
MobiSys5
2025 Poster Abstract: Terrain Navigability Assessment of Autonomous Ground Robots Using mmWave Radar
abstract
We present a parameter evaluation of FMCW mmWave Radar to assess surface dampness and ruggedness and enhance the navigability of autonomous ground robots. We begin by designing and 3D-printing a mount for the mmWave Radar on a Rosmaster X3 platform. We then collect raw mmWave Radar data from various surfaces (grass, soil, puddles, and mulch) across different seasons (summer, winter, and rainy). Our findings demonstrate that the energy strength parameter is a reliable indicator for assessing the surface: dry surfaces (e.g., dry grass, dry mud) exhibit lower energy strength, whereas wet surfaces display higher values. This dampness and ruggedness factor can be leveraged to develop a cost-map navigability score for autonomous ground robots.
Anuradha Ravi, Eric Meza, Snehalraj Chugh, Andre Harrison, Timothy Gregory, Jade Freeman, Nirmalya Roy
SenSys7
2025 E2RespUNet: End-to-End Respiratory Signal Reconstruction and Rate Prediction Using a Unified Attention-Enhanced U-Net
abstract
Continuous, non-invasive respiratory rate (RR) monitoring is essential for the early diagnosis of many medi-cal problems. However, conventional contact-based sensors frequently under-perform in dynamic situations that can be un-comfortable and require human intervention. To overcome these limitations, we propose E2RespUNet, an end-to-end system that uses multimodal video data to estimate breathing rates and reconstruct respiratory signals using an attention-enhanced U-Net architecture. Our method combines optical flow analysis with preprocessing, detrending, and normalization to reliably extract chest motion features in a variety of settings. We validated our approach using a collection of RGB, near-infrared (NIR), and infrared (IR) videos from multiple sources, such as an in-house RGB dataset and a publicly available sleep dataset. According to our study in both the temporal and frequency domains, E2RespUNet surpasses existing baseline models by lowering the mean absolute error by up to 21 % in the sleep dataset and by up to 28% in the in-house dataset. Additionally, E2RespUNet addresses real-time deployment requirements on resource-constrained edge devices by processing each frame in less than 1 ms (0.99 ms in the Sleep dataset and 0.97 ms in the in-house dataset). These findings, together with a low KL divergence and a power spectral density (PSD) that corresponds to the ground truth, suggest that E2Resp UN et is accurate and fast enough for real-time use in emergency and clinical scenarios.
Shadman Sakib, Gaurav Shinde, Emon Dey, Nirmalya Roy
SMARTCOMP4
2025 Augmenting Personalized Memory via Practical Multimodal Wearable Sensing in Visual Search and Wayfinding Navigation
abstract
ACM UMAP June 16-19 2025, New York, USA
Indrajeet Ghosh, Kasthuri Jayarajah, Nicholas R. Waytowich, Nirmalya Roy
UMAP4
2025 Is Edge Computing Always Suitable for Image Analysis? An Experimental Analysis
abstract
In the era of Smart Cities, video surveillance stands as a pivotal tool for enhancing urban security, optimizing resource management, and improving the quality of urban life. When video surveillance is seamlessly integrated with image analysis systems, raw visual data are transformed into actionable insights, significantly enhancing the capability of Smart Cities to ensure public safety and optimize urban operations. Image analysis systems mainly rely on the cloud: images are offloaded to a cloud infrastructure to be processed, analyzed and segmented for inference. The analysis of images in external systems, however, is not always recommended, due to privacy/security concerns, e.g., human action recognition. In this paper, we investigate the opportunity to adopt edge computing to implement such systems, where images are analyzed directly on-premises. To investigate the suitability of this approach, we carried out an extensive experimentation using two large-scale Fed4Fire+ testbeds, namely, Grid’5000 and Virtual Wall. Specifically, we considered different cloud-edge configurations using different inference models, and evaluated the impact of those models on performance and resource utilization. Based on these results, we provide a set of guidelines for the adoption of different models depending on the requirements of the specific application.
Francesca Righetti, Carlo Vallati, Nirmalya Roy, Giuseppe Anastasi
WoWMoM3
2025 Recent Advancements in Machine Learning for Cybercrime Prediction
abstract
Cybercrime is a growing threat to organizations and individuals worldwide, with criminals using sophisticated techniques to breach security systems and steal sensitive data. This paper aims to comprehensively survey the latest advancements in cybercrime prediction, highlighting the relevant research. For this purpose, we reviewed more than 150 research articles and discussed 50 most recent and appropriate ones. We start the review with some standard methods cybercriminals use and then focus on the latest machine and deep learning techniques, which detect anomalous behavior and identify potential threats. We also discuss transfer learning, which allows models trained on one dataset to be adapted for use on another dataset. We then focus on active and reinforcement learning as part of early-stage algorithmic research in cybercrime prediction. Finally, we discuss critical innovations, research gaps, and future research opportunities in Cybercrime prediction. This paper presents a holistic view of cutting-edge developments and publicly available datasets.
Lavanya Elluri, Varun Mandalapu, Piyush Vyas, Nirmalya Roy
J. Comput. Inf. Syst.4
2024 Unsupervised Domain Adaptation for Action Recognition via Self-Ensembling and Conditional Embedding Alignment
abstract
Recent advancements in deep learning-based wearable human action recognition (wHAR) have improved the capture and classification of complex motions, but adoption remains limited due to the lack of expert annotations and domain discrepancies from user variations. Limited annotations hinder the model's ability to generalize to out-of-distribution samples. While data augmentation can improve generalizability, unsupervised augmentation techniques must be applied carefully to avoid introducing noise. Unsupervised domain adaptation (UDA) addresses domain discrepancies by aligning conditional distributions with labeled target samples, but vanilla pseudo-labeling can lead to error propagation. To address these challenges, we propose μDAR, a novel joint optimization architecture comprised of three functions: (i) consistency regularizer between augmented samples to improve model classification generalizability, (ii) temporal ensemble for robust pseudo-label generation and (iii) conditional distribution alignment to improve domain generalizability. The temporal ensemble works by aggregating predictions from past epochs to smooth out noisy pseudo-label predictions, which are then used in the conditional distribution alignment module to minimize kernel-based class-wise conditional maximum mean discrepancy (kCMMD) between the source and target feature space to learn a domain invariant embedding. The consistency-regularized augmentations ensure that multiple augmentations of the same sample share the same labels; this results in (a) strong generalization with limited source domain samples and (b) consistent pseudo-label generation in target samples. The novel integration of these three modules in μDAR results in a range of ~ 4-12% average macro-F1 score improvement over six state-of-the-art UDA methods in four benchmark wHAR datasets.
Indrajeet Ghosh, Garvit Chugh, Abu Zaher Md Faridee, Nirmalya Roy
ICDM4
2024 TopoNav: Topological Navigation for Efficient Exploration in Sparse Reward Environments
abstract
Autonomous robots exploring unknown environments face a significant challenge: navigating effectively without prior maps and with limited external feedback. This challenge intensifies in sparse reward environments, where traditional exploration techniques often fail. In this paper, we present TopoNav, a novel topological navigation framework that integrates active mapping, hierarchical reinforcement learning, and intrinsic motivation to enable efficient goal-oriented exploration and navigation in sparse-reward settings. TopoNav dynamically constructs a topological map of the environment, capturing key locations and pathways. A two-level hierarchical policy architecture, comprising a high-level graph traversal policy and low-level motion control policies, enables effective navigation and obstacle avoidance while maintaining focus on the overall goal. Additionally, TopoNav incorporates intrinsic motivation to guide exploration towards relevant regions and frontier nodes in the topological map, addressing the challenges of sparse extrinsic rewards. We evaluate TopoNav both in the simulated and real-world off-road environments using a Clearpath Jackal robot, across three challenging navigation scenarios: goal-reaching, feature-based navigation, and navigation in complex terrains. We observe an increase in exploration coverage by 7-20%, in success rates by 9-19%, and reductions in navigation times by 15-36% across various scenarios, compared to state-of-the-art methods.
Jumman Hossain, Abu Zaher Md Faridee, Nirmalya Roy, Jade Freeman, Timothy Gregory, Theron Trout
IROS3
2024 EEGAmp+: Investigating the Efficacy of Functional Connectivity for Detecting Events in Low Resolution EEG
abstract
Electroencephalography (EEG) has found many applications cutting across many domains, such as digital health, affective computing, and human-machine interfaces. However, its widespread adoption in practice has been primarily inhibited by its susceptibility to noise artifacts and the low spatial resolution of electrodes on commercial EEG sensors. While several prior works have investigated techniques for detecting and extracting noise, our understanding of performance degradation due to electrode sparsity remains limited. In this work, we explore the feasibility of using Functional Connectivity (FC) for improving the EEG sensing-based accuracy using two exemplars downstream, working memory-related tasks: (a) cognitive task load (CTL) assessment and (b) high attentional event-evoked potential (EEP) episodes detection. This paper proposes an integrated approach, EEGAmp+ , that first utilizes channel-wise functional connectivity modules using independent component analysis (ICA) coupled with cosine distance for EEG signal reconstruction for cognitive task load assessment tasks. This is then coupled with a sliding window change point detection technique paired with continuous wavelet transformation (CWT) to extract high attentional EEP episodes. Our empirical results indicate that using independent component analysis (ICA) coupled with FC to improve spatial resolution increased cognitive load assessment accuracy by [5.6% ± 1.13] across four machine learning algorithms. Furthermore, after signal reconstruction, we introduce sliding window CPD coupled with CWT, which allows us to extract EEP segments legibly through decomposing the signals and the ability to capture both time and frequency representation from the reconstructed signal boosting detection accuracy by [11.1% ±1.31].
Indrajeet Ghosh, Kasthuri Jayarajah, Nicholas R. Waytowich, Nirmalya Roy
MobiQuitous4
2024 AcouDL: Context-Aware Daily Activity Recognition from Natural Acoustic Signals
abstract
The ubiquitousness of smart and wearable devices with integrated acoustic sensors in modern human lives presents tremendous opportunities for recognizing human activities in our living spaces through ML-driven applications. However, their adoption is often hindered by the requirement of large amounts of labeled data during the model training phase. Integration of contextual metadata has the potential to alleviate this since the nature of these meta-data is often less dynamic (e.g. cleaning dishes, and cooking both can happen in the kitchen context) and can often be annotated in a less tedious manner (a sensor always placed in the kitchen). However, most models do not have good provisions for the integration of such meta-data information. Often, the additional metadata is leveraged in the form of multi-task learning with sub-optimal outcomes. On the other hand, reliably recognizing distinct in-home activities with similar acoustic patterns (e.g. chopping, hammering, knife sharpening) poses another set of challenges. To mitigate these challenges, we first show in our preliminary study that the room acoustics properties such as reverberation, room materials, and background noise leave a discernible fingerprint in the audio samples to recognize the room context and proposed AcouDL as a unified framework to exploit room context information to improve activity recognition performance. Our proposed self-supervision-based approach first learns the context features of the activities by leveraging a large amount of unlabeled data using a contrastive learning mechanism and then incorporates this feature induced with a novel attention mechanism into the activity classification pipeline to improve the activity recognition performance. Extensive evaluation of AcouDL on three datasets containing a wide range of activities shows that such an efficient feature fusion-mechanism enables the incorporation of metadata that helps to better recognition of the activities under challenging classification scenarios with 0.7-3.5% macro F1 score improvement over the baselines.
Avijoy Chakma, Anirban Das 0005, Abu Zaher Md Faridee, Suchetana Chakraborty, Sandip Chakraborty 0001, Nirmalya Roy
SMARTCOMP6
2024 Contactless Physiological Health Sensing: Challenges, Solutions & Opportunities
abstract
Contactless health monitoring techniques, such as Remote Photoplethysmography (rPPG) [1] and video-based Respiratory Rate (RR) [2] estimation, have emerged as promising methods utilizing regular camera sensors for capturing vital signs like Heart Rate (HR) and Respiratory Rate (RR). These approaches offer a cost-effective, widely applicable, and safe solution for regular and long-term health monitoring. However, the primary challenge lies in extracting accurate vital signals from the captured videos due to their low signal-to-noise ratios. Traditional signal processing analyzes temporal properties of spatial regions of interest across video frames to extract vital signs, while computer vision-based Deep Learning (DL) approaches leverage large-scale datasets and computational power for data-driven decision-making without heuristic reliance [3]. In this tutorial, we aim to provide a comprehensive background on rPPG and video-based heart rate estimation, covering signal acquisition and physics-inspired estimation principles. We will discuss signal-processing approaches and their limitations. We will delve into DL-based estimation systems, addressing the challenge of inherent aleatoric uncertainty (irreducible uncertainty from various stochastic factors like sensor variations and inter-subject differences) in the ground truth data annotation that hinders the development of generalized deep learning-based rPPG estimation system. To reinforce a robust DL model addressing these inherent uncertainties in rPPG data streams, we will discuss three novel deep-learning approaches. First, a multi-task learning method for rPPG estimation that learns the core rPPG features using shared embedding across noisy ground truth by separating individual targets. Second, a self-supervised learning method that leverages unlabeled rPPG data to learn the intrinsic rPPG properties and improve rPPG feature representation and signal reconstruction by incorporating the prior domain knowledge (HR frequency, phase, and temporal coherence). Third, a generative adversarial method that enables fine-grained rPPG learning without direct supervision using large-scale rPPG data. Finally, we will present the validation of the proposed methods across in-house MPSC-rPPG datasets, and multiple public datasets with open-source code for further exploration. Furthermore, we will demonstrate a real-time heart rate estimation system, RhythmEdge, using a low-cost camera sensor and describe the rPPG-specific pruning techniques to reduce rPPG model size for efficient edge implementation. We will conclude the tutorial with existing research gaps and potential directions in contactless rPPG and respiratory rate estimation research.
Nirmalya Roy, Zahid Hasan 0001
SMARTCOMP1
2023 A Novel ROS2 QoS Policy-Enabled Synchronizing Middleware for Co-Simulation of Heterogeneous Multi-Robot Systems
abstract
Recent Internet-of-Things (IoT) networks span across a multitude of stationary and robotic devices, namely unmanned ground vehicles, surface vessels, and aerial drones, to carry out mission-critical services such as search and rescue operations, wildfire monitoring, and flood/hurricane impact assessment. Achieving communication synchrony, reliability, and minimal communication jitter among these devices is a key challenge both at the simulation and system levels of implementation due to the underpinning differences between a physics-based robot operating system (ROS) simulator that is time-based and a network-based wireless simulator that is event-based, in addition to the complex dynamics of mobile and heterogeneous IoT devices deployed in a real environment. Nevertheless, synchronization between physics (robotics) and network simulators is one of the most difficult issues to address in simulating a heterogeneous multi-robot system before transitioning it into practice. The existing TCP/IP communication protocol-based synchronizing middleware mostly relied on Robot Operating System 1 (ROS1), which expends a significant portion of communication bandwidth and time due to its master-based architecture. To address these issues, we design a novel synchronizing middleware between robotics and traditional wireless network simulators, relying on the newly released real-time ROS2 architecture with a masterless packet discovery mechanism. Additionally, we propose a ground and aerial agents' velocity-aware customized QoS policy for Data Distribution Service (DDS) to minimize the packet loss and transmission latency between a diverse set of robotic agents, and we offer the theoretical guarantee of our proposed QoS policy. We performed extensive network performance evaluations both at the simulation and system levels in terms of packet loss probability and average latency with line-of-sight (LOS) and non-line-of-sight (NLOS) and TCP/UDP communication protocols over our proposed ROS2-based synchronization middleware. Moreover, for a comparative study, we presented a detailed ablation study replacing NS-3 with a real-time wireless network simulator, EMANE, and masterless ROS2 with master-based ROS1. Our proposed middleware attests to the promise of building a large-scale IoT infrastructure with a diverse set of stationary and robotic devices that achieve low-latency communications (12% and 11% reduction in simulation and reality, respectively) while satisfying the reliability (10% and 15% packet loss reduction in simulation and reality, respectively) and high-fidelity requirements of mission-critical applications.
Emon Dey, Mikolaj Walczak, Mohammad Saeid Anwar, Nirmalya Roy, Jade Freeman, Timothy Gregory, Niranjan Suri, Carl E. Busart
ICCCN4
2023 NEV-NCD: Negative Learning, Entropy, and Variance Regularization Based Novel Action Categories Discovery
abstract
Novel Categories Discovery (NCD) facilitates learning from a partially annotated label space and enables deep learning (DL) models to operate in an open-world setting by identifying and differentiating instances of novel classes based on the labeled data notions. One of the primary assumptions of NCD is that the novel label space is perfectly disjoint and can be equipartitioned, but it is rarely realized by most NCD approaches in practice. To better align with this assumption, we propose a novel single-stage joint optimization-based NCD method, Negative learning, Entropy, and Variance regularization NCD (NEV-NCD). We demonstrate the efficacy of NEV-NCD in previously unexplored NCD applications of video action recognition (VAR) with the public UCF101 dataset and a curated in-house partial action-space annotated multi-view video dataset. Further, we perform a thorough ablation study by varying the composition of final joint loss and associated hyper-parameters. During our experiments with UCF101 and multi-view action dataset, NEV-NCD achieves ≈ 83% classification accuracy in test instances of labeled data. NEV-NCD achieves ≈ 70% clustering accuracy over unlabeled data outperforming both naive baselines and state-of-the-art pseudo-labeling-based approaches by ≈ 40% and ≈ 3.5% over both datasets.
Zahid Hasan 0001, Masud Ahmed, Abu Zaher Md Faridee, Sanjay Purushotham, Heesung Kwon, Hyungtae Lee, Nirmalya Roy
ICIP7
2023 An Online Continuous Semantic Segmentation Framework With Minimal Labeling Efforts
abstract
The annotation load for a new dataset has been greatly decreased using domain adaptation based semantic segmentation, which iteratively constructs pseudo labels on unlabeled target data and retrains the network. However, realistic segmentation datasets are often imbalanced, with pseudo-labels tending to favor certain "head" classes while neglecting other "tail" classes. This can lead to an inaccurate and noisy mask. To address this issue, we propose a novel hard sample mining strategy for an active domain adaptation based semantic segmentation network, with the aim of automatically selecting a small subset of labeled target data to fine-tune the network. By calculating class-wise entropy, we are able to rank the difficulty level of different samples. We use a fusion of focal loss and regional mutual information loss instead of cross-entropy loss for the domain adaptation based semantic segmentation network. Our entire framework has been implemented in real-time using the Robotics Operating System (ROS) with a server PC and a small Unmanned Ground Vehicle (UGV) known as the ROSbot2.0 Pro. This implementation allows ROSbot2.0 Pro to access any type of data at any time, enabling it to perform a variety of tasks with ease. Our approach has been thoroughly evaluated through a series of extensive experiments, which demonstrate its superior performance compared to existing state-of-the-art methods. Remarkably, by using just 20% of hard samples for fine-tuning, our network has achieved a level of performance that is comparable (≈88%) to that of a fully supervised approach, with mIOU scores of 60.51% in the In-house dataset.
Masud Ahmed, Zahid Hasan 0001, Tim Yingling, Eric O'Leary, Sanjay Purushotham, Suya You, Nirmalya Roy
SMARTCOMP7
2023 A Systematic Study on Object Recognition Using Millimeter-wave Radar
abstract
Millimeter-wave (MMW) radar is becoming an essential sensing technology in smart environments due to its light and weather-independent sensing capability. Such capabilities have been widely explored and integrated with intelligent vehicle systems, often deployed in industry-grade MMW radars. However, industry-grade MMW radars are often expensive and difficult to attain for deployable community-purpose smart environment applications. On the other hand, commercially available MMW radars pose hidden underpinning challenges that are yet to be well investigated for tasks such as recognizing objects, and activities, real-time person tracking, object localization, etc. Such tasks are frequently accompanied by image and video data, which are relatively easy for an individual to obtain, interpret, and annotate. However, image and video data are light and weather-dependent, vulnerable to the occlusion effect, and inherently raise privacy concerns for individuals. It is crucial to investigate the performance of an alternative sensing mechanism where commercially available MMW radars can be a viable alternative to eradicate the dependencies and preserve privacy issues. Before championing MMW radar, several questions need to be answered regarding MMW radar’s practical feasibility and performance under different operating environments. To answer the concerns, we have collected a dataset using commercially available MMW radar, Automotive mmWave Radar (AWR2944) from Texas Instruments, and reported the optimum experimental settings for object recognition performance using several deep learning algorithms in this study. Moreover, our robust data collection procedure allows us to systematically study and identify potential challenges in the object recognition task under a cross-ambience scenario. We have explored the potential approaches to overcome the underlying challenges and reported extensive experimental results.
Maloy Kumar Devnath, Avijoy Chakma, Mohammad Saeid Anwar, Emon Dey, Zahid Hasan 0001, Marc Conn, Biplab Pal, Nirmalya Roy
SMARTCOMP8
2023 HeteroSys: Heterogeneous and Collaborative Sensing in the Wild
abstract
Advances in Internet-of-Things, artificial intelligence, and ubiquitous computing technologies have contributed to building the next generation of context-aware heterogeneous systems with robust interoperability to control and monitor the environmental variables of smart environments. Motivated by this, we propose HeteroSys, an end-to-end multi-functional smart IoT-based system prototype for heterogeneous and collaborative sensing in a smart IoT-based environment. A unique characteristic of HeteroSys is that it relies on Home Assistant (HA) to collate heterogeneous sensors (e.g., passive infrared sensors (PIR), reed (door) switches, object tags, wearable wrist-mounted, water leak sensors, and internet protocol cameras), and uses a variety of networking protocols such as Zigbee open standard for mesh networking, WiFi, and Bluetooth Low Energy (BLE) for communication. The reliance on HA (and its broad community support) makes HeteroSys ideal for various applications such as object detection, human activity recognition and behavior patterns. We articulated the development phase, integration, testing challenges and evaluation of the HeteroSys. We conducted an extensive 24-hour longitudinal data collection from 5 participants performing 6 activities by deploying in an indoor home environment. Our assessment of the acquired dataset reveals that the representations learned using deep learning architecture aid in improving the detection of activities to 83.1% accuracy.
Indrajeet Ghosh, Adam Goldstein, Avijoy Chakma, Jade Freeman, Timothy Gregory, Niranjan Suri, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy
SMARTCOMP8
2023 SrPPG: Semi-Supervised Adversarial Learning for Remote Photoplethysmography with Noisy Data
abstract
Remote Photoplethysmography (rPPG) systems offer contactless, low-cost, and ubiquitous heart rate (HR) monitoring by leveraging the skin-tissue blood volumetric variation-induced reflection. However, collecting large-scale time-synchronized rPPG data is costly and impedes the development of generalized end-to-end deep learning (DL) rPPG models to perform under diverse scenarios. We formulate the rPPG estimation as a generative task of recovering time-series PPG from facial videos and propose SrPPG, a novel semi-supervised adversarial learning framework using heterogeneous, asynchronous, and noisy rPPG data. More specifically, we develop a novel encoder-decoder architecture, where rPPG features are learned from video in a self-supervised manner (encoder) to reconstruct the time-series PPG (decoder/generator) with physics-inspired novel temporal consistency regularization. The generated PPG is scrutinized against the real rPPG signals by a frequency-class conditioned discriminator, forming a generative adversarial network. Thus, SrPPG generates samples without point-wise supervision, alleviating the need for time-synchronized data collection. We experiment and validate SrPPG by amassing three public datasets in heterogeneous settings. SrPPG outperforms both supervised and self-supervised state-of-the-art methods in HR estimation across all datasets without any time-synchronous rPPG data. We also perform extensive experiments to study the optimal generative setting (architecture, joint optimization) and provide insight into the SrPPG behavior.
Zahid Hasan 0001, Abu Zaher Md Faridee, Masud Ahmed, Shibi Ayyanar, Nirmalya Roy
SMARTCOMP5
2022 Simulated Forest Environment and Robot Control Framework for Integration with Cover Detection Algorithms
abstract
Simulated environments can be a quicker and more flexible alternative to training and testing machine learning models in the real world. Models also need to be able to efficiently communicate with the environment. In military-relevant environments, a trained model can play a valuable role in finding cover for an autonomous robot to prevent getting detected or attacked by adversaries. In this regard, we present a forest simulation and robot control framework that is ready for integration with machine learning or object recognition algorithms. Our framework includes an environment relevant to military situations and is capable of providing information about the environment to a machine learning model. A forest environment was designed with wooded areas, open paths, water, and bridges. A Clearpath Husky robot is simulated in the environment using Army Research Laboratory’s (ARL) Unity and ROS simulation framework. The Husky robot is equipped with a camera and lidar sensor. Data from these sensors can be read through ROS topics and RViz configuration windows. The robot can be moved using ROS velocity command topics. These communication methods can be employed by a machine learning algorithm for use in detecting trees to attain maximum cover. Our designed environment improves upon the default ARL framework environments by offering a more diverse terrain and more opportunities for cover. This makes the environment more relevant to a cover-seeking machine learning model. Code, videos, and integration process available at : https://github.com/avispector7/Forest-Simulation
Avi Spector, Wanying Zhu, Jumman Hossain, Nirmalya Roy
BDCAT4
2022 Cross-Domain Unseen Activity Recognition Using Transfer Learning
abstract
Activity Recognition (AR) models perform well with a large number of available training instances. However, in the presence of sensor heterogeneity, sensing biasness and variability of human behaviors and activities and unseen activity classes pose key challenges to adopting and scaling these pre-trained activity recognition models in the new environment. These challenging unseen activities recognition problems are addressed by applying transfer learning techniques that leverage a limited number of annotated samples and utilize the inherent structural patterns among activities within and across the source and target domains. This work proposes a novel AR framework that uses the pre-trained deep autoencoder model and generates features from source and target activity samples. Furthermore, this AR frame-work establishes correlations among activities between the source and target domain by exploiting intra- and inter-class knowledge transfer to mitigate the number of labeled samples and recognize unseen activities in the target domain. We validated the efficacy and effectiveness of our AR framework with three real-world data traces (Daily and Sports, Opportunistic, and Wisdm) that contain 41 users and 26 activities in total. Our AR framework achieves performance gains ≈ 5-6% with 111, 18, and 70 activity samples (20 % annotated samples) for Das, Opp, and Wisdm datasets. In addition, our proposed AR framework requires 56, 8, and 35 fewer activity samples (10% fewer annotated examples) for Das, Opp, and Wisdm, respectively, compared to the state-of-the-art Untran model.
Md Abdullah Al Hafiz Khan, Nirmalya Roy
COMPSAC2
2022 GADAN: Generative Adversarial Domain Adaptation Network For Debris Detection Using Drone
abstract
In maritime, coastal, and riverine environments debris has become an abundant pollutant, posing a significant threat to aquatic life. Existing debris detection systems are mostly designed to detect debris for specific type of environment. Therefore, we propose GADAN, an autonomous debris detection model fusing RGB and thermal images, and experiment on both over land and aquatic environment. We collected data using the Parrot Anafi Thermal drone from different water streams, and construction sites near the aquatic environment. We employ a generative domain adaptation based architecture to generate the thermal image compatible with the RGB image. We postulate a two-stream network on these thermal and RGB pairs of images separately, and pass the concatenated features to object detecting YOLO model. We also demonstrated the performance of debris detection using only RGB and only thermal images. Our proposed GADAN network incorporating RGB and thermal images outperforms (with a mean average precision of 0.97) RCNN and single modality based models.
Masud Ahmed, Naima Khan, Pretom Roy Ovi, Nirmalya Roy, Sanjay Purushotham, Aryya Gangopadhyay, Suya You
DCOSS4
2022 Semi-supervised Multi-source Domain Adaptation in Wearable Activity Recognition
abstract
The scarcity of labeled data has traditionally been the primary hindrance in building scalable supervised deep learning models that can retain adequate performance in the presence of various heterogeneities in sample distributions. Domain adaptation tries to address this issue by adapting features learned from a smaller set of labeled samples to that of the incoming unlabeled samples. The traditional domain adaptation approaches normally consider only a single source of labeled samples, but in real world use cases, labeled samples can originate from multiple-sources – providing motivation for multi-source domain adaptation (MSDA). Several MSDA approaches have been investigated for wearable sensor-based human activity recognition (HAR) in recent times, but their performance improvement compared to single source counterpart remained marginal. To remedy this performance gap that, we explore multiple avenues to align the conditional distributions in addition to the usual alignment of marginal ones. In our investigation, we extend an existing multi-source domain adaptation approach under semi-supervised settings. We assume the availability of partially labeled target domain data and further explore the pseudo labeling usage with a goal to achieve a performance similar to the former. In our experiments on three publicly available datasets, we find that a limited labeled target domain data and pseudo label data boost the performance over the unsupervised approach by 10-35% and 2-6%, respectively, in various domain adaptation scenarios.
Avijoy Chakma, Abu Zaher Md Faridee, Raghuveer M. Rao, Nirmalya Roy
DCOSS4
2022 SynchroSim: An Integrated Co-simulation Middleware for Heterogeneous Multi-robot System
abstract
With the advancement of modern robotics, autonomous agents are now capable of hosting sophisticated algorithms, which enables them to make intelligent decisions. But developing and testing such algorithms directly in real-world systems is tedious and may result in the wastage of valuable resources. Especially for heterogeneous multi-agent systems in battlefield environments where communication is critical in determining the system’s behavior and usability. Due to the necessity of simulators of separate paradigms (co-simulation) to simulate such scenarios before deploying, synchronization between those simulators is vital. Existing works aimed at resolving this issue fall short of addressing diversity among deployed agents. In this work, we propose SynchroSim, an integrated co-simulation middleware to simulate a heterogeneous multi-robot system. Here we propose a velocity difference-driven adjustable window size approach with a view to reducing packet loss probability. It takes into account the respective velocities of deployed agents to calculate a suitable window size before transmitting data between them. We consider our algorithm specific simulator agnostic but for the sake of implementation results, we have used Gazebo as a Physics simulator and NS-3 as a network simulator. Also, we design our algorithm considering the Perception-Action loop inside a closed communication channel, which is one of the essential factors in a contested scenario with the requirement of high fidelity in terms of data transmission. We validate our approach empirically at both the simulation and system level for both line-of-sight (LOS) and non-line-of-sight (NLOS) scenarios. Our approach achieves a noticeable improvement in terms of reducing packet loss probability (≈11%), and average packet delay (≈10%) compared to the fixed window size-based synchronization approach.
Emon Dey, Jumman Hossain, Nirmalya Roy, Carl E. Busart
DCOSS3
2022 PerMTL: A Multi-Task Learning Framework for Skilled Human Performance Assessment
abstract
Intelligent and complex human motion analysis can help design the next generation IoT and AR/VR systems for automated human performance assessment. Such an automated system can help advocate the interpretability and translatability of complex human motions, intelligent motion feedback, and fine-grained motion skill assessment to design next-generation interactive human-machine teaming systems. Motivated by this, we design a wearable sensing framework for assessing the players’ performance and consider a live badminton game as our use case. Generally, the players on the field try to improve their performance by focusing on fast and synchronous coordination of their limbs’ reflex actions to have the ideal body postures to perform the desired shot. Learning the minute dissimilarities and distinctive traits from each limb of the players simultaneously can help assess the players’ performance and specific skillsets during a game. This paper proposes a multi-task learning framework, PerMTL to learn the shared features from each player’s limb. The PerMTL comprises a task-specific regressor output layer that helps to determine the dissimilarities and distinctive traits between the player’s limbs for collective inference in a body sensor network (BSN) environment. We evaluate the PerMTL framework using publicly available Badminton Activity Recognition (BAR) and Daily and Sports Activities (DSA) datasets. Empirical results indicate that PerMTL achieves R2Score of ≈ 82% in predicting the players’ performance.
Indrajeet Ghosh, Avijoy Chakma, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy, Nicholas R. Waytowich
ICMLA4
2022 Environmental Sound Classification for Flood Event Detection
abstract
Flood is one of the common natural disasters that can severely affect human life and properties. Early detection, therefore, is of paramount importance to provide help through an emergency response team. Robust flood detection techniques so far have been based on computer vision using images either from cameras, satellite imagery, remote sensing, or radar-based images. However, sound signal-based flood event detection has not been widely explored. In this work, we design an end-to-end architecture for a deep learning-based flood-related sound event detection model. We employ Mel-Spectrogram-based auditory signal analysis and deep learning models for sound event detection (SED). We evaluated four deep learning models under the following two categories: (i) Binary classification Flood/No Flood, vs. Windy vs. Non-Windy, and (ii) Multi-classification for more granular flood and wind events. The experimental results performed in these settings on the datasets collected from real deployment showed an accuracy of around 78%.
Bipendra Basnyat, Nirmalya Roy, Aryya Gangopadhyay, Adrienne Raglin
Intelligent Environments2
2022 CogAx: Early Assessment of Cognitive and Functional Impairment from Accelerometry
abstract
An individual’s cognitive and functional abilities are commonly assessed through physical and mental status examination, observational performance measures, surveys and proxy reports of symptoms. These strategies are not ideal for early impairment detection as the individual needs to be present physically at the clinic to avail the assessments, especially for older adults who require assistance from a caregiver, and experience mobility, cognitive and functional disabilities from neurodegenerative disorders. Moreover, these strategies rely on self-reporting and proxy reports for evaluation which often leads to under-reporting of symptoms and decrease the validity of these measures. We argue that an early assessment of functional, and cognitive health impairment can be obtained from the individual’s daily activities captured through accelerometry. In this work, we postulate to learn high-level motion related representations from accelerometer data to better correlate with underlying functional and cognitive health parameters of older adults using a contrastive and multi-task learning framework. In particular, we posit a novel indicator, Impairment Indicator using the proposed multi-task learning framework that can indicate functional or cognitive decline as neurodegenerative disease progresses. An extensive 24-hour data collection from 25 older adults with the clinician in-the-loop was carried out in a retirement community center with IRB approval. We collected the activity patterns using wearables in their homes in addition to survey-based assessments and observational performance measures recorded by a clinical evaluator to infer their current cognitive and functional impairment status. Our evaluation on the acquired dataset reveals that the representations learned using contrastive learning aids in improving the detection of activities, activity performance score, and stage of dementia to 92%, 97%, and 98%, respectively.
Sreenivasan Ramasamy Ramamurthy, Soumyajit Chatterjee, Elizabeth Galik, Aryya Gangopadhyay, Nirmalya Roy, Bivas Mitra, Sandip Chakraborty 0001
PerCom5
2022 CoDEm: Conditional Domain Embeddings for Scalable Human Activity Recognition
abstract
We explore the effect of auxiliary labels in improving the classification accuracy of wearable sensor-based human activity recognition (HAR) systems, which are primarily trained with the supervision of the activity labels (e.g. running, walking, jumping). Supplemental meta-data are often available during the data collection process such as body positions of the wearable sensors, subjects' demographic information (e.g. gender, age), and the type of wearable used (e.g. smartphone, smart-watch). This information, while not directly related to the activity classification task, can nonetheless provide auxiliary supervision and has the potential to significantly improve the HAR accuracy by providing extra guidance on how to handle the introduced sample heterogeneity from the change in domains (i.e positions, persons, or sensors), especially in the presence of limited activity labels. However, integrating such meta-data information in the classification pipeline is non-trivial - (i) the complex interaction between the activity and domain label space is hard to capture with a simple multi-task and/or adversarial learning setup, (ii) meta-data and activity labels might not be simultaneously available for all collected samples. To address these issues, we propose a novel framework Conditional Domain Embeddings (CoDEm). From the available unlabeled raw samples and their domain meta-data, we first learn a set of domain embeddings using a contrastive learning methodology to handle inter-domain variability and inter-domain similarity. To classify the activities, CoDEm then learns the label embeddings in a contrastive fashion, conditioned on domain embeddings with a novel attention mechanism, enforcing the model to learn the complex domain-activity relationships. We extensively evaluate CoDEm in three benchmark datasets against a number of multi-task and adversarial learning baselines and achieve state-of-the-art nerformance in each avenue.
Abu Zaher Md Faridee, Avijoy Chakma, Zahid Hasan 0001, Nirmalya Roy, Archan Misra
SMARTCOMP4
2022 SpecTextor: End-to-End Attention-based Mechanism for Dense Text Generation in Sports Journalism
abstract
Language-guided smart systems can help to design next-generation human-machine interactive applications. The dense text description is one of the research areas where systems learn the semantic knowledge and visual features of each video frame and map them to describe the video's most relevant subjects and events. In this paper, we consider untrimmed sports videos as our case study. Generating dense descriptions in the sports domain to supplement journalistic works without relying on commentators and experts requires more investigation. Motivated by this, we propose an end-to-end automated text-generator, SpecTextor, that learns the semantic features from untrimmed videos of sports games and generates associated descriptive texts. The proposed approach considers the video as a sequence of frames and sequentially generates words. After splitting videos into frames, we use a pre-trained VGG-16 model for feature extraction and encoding the video frames. With these encoded frames, we posit a Long Short-Term Memory (LSTM) based attention-decoder pipeline that leverages soft-attention mechanism to map the semantic features with relevant textual descriptions to generate the explanation of the game. Because developing a comprehensive description of the game warrants training on a set of dense time-stamped captions, we leverage two available public datasets: ActivityNet Captions and Microsoft Video Description. In addition, we utilized two different decoding algorithms: beam search and greedy search and computed two evaluation metrics: BLEU and METEOR scores.
Indrajeet Ghosh, Matthew Ivler, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy
SMARTCOMP4
2022 RhythmEdge: Enabling Contactless Heart Rate Estimation on the Edge
abstract
The primary contribution of this paper is designing and prototyping a real-time edge computing system, RhythmEdge, that is capable of detecting changes in blood volume from facial videos (Remote Photoplethysmography; rPPG), enabling cardio-vascular health assessment instantly. The benefits of RhythmEdge include non-invasive measurement of cardiovascular activity, real-time system operation, inexpensive sensing components, and computing. RhythmEdge captures a short video of the skin using a camera and extracts rPPG features to estimate the Photoplethysmography (PPG) signal using a multi-task learning framework while offloading the edge computation. In addition, we intelligently apply a transfer learning approach to the multi-task learning framework to mitigate sensor heterogeneities to scale the RhythmEdge prototype to work with a range of commercially available sensing and computing devices. Besides, to further adapt the software stack for resource-constrained devices, we postulate novel pruning and quantization techniques (Quantization: FP32, FP16; Pruned-Quantized: FP32, FP16) that efficiently optimize the deep feature learning while minimizing the runtime, latency, memory, and power usage. We benchmark RhythmEdge prototype for three different cameras and edge computing platforms while evaluating it on three publicly available datasets and an in-house dataset collected under challenging environmental circumstances. Our analysis indicates that RhythmEdge performs on par with the existing contactless heart rate monitoring systems while utilizing only half of its available resources. Furthermore, we perform an ablation study with and without pruning and quantization to report the model size (87%) vs. inference time (70%) reduction. We attested the efficacy of RhythmEdge prototype with a maximum power of 8W and a memory usage of 290MB, with a minimal latency of 0.0625 seconds and a runtime of 0.64 seconds per 30 frames.
Zahid Hasan 0001, Emon Dey, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy, Archan Misra
SMARTCOMP4
2022 Demo: RhythmEdge: Enabling Contactless Heart Rate Estimation on the Edge
abstract
In this demo paper, we design and prototype RhythmEdge [1], a low-cost, deep-learning-based contact-less system for regular HR monitoring applications. RhythmEdge benefits over existing approaches by facilitating contact-less nature, real-time/offline operation, inexpensive and available sensing components, and computing devices. Our RhythmEdge system is portable and easily deployable for reliable HR estimation in moderately controlled indoor or outdoor environments. RhythmEdge measures HR via detecting changes in blood volume from facial videos (Remote Photoplethysmography; rPPG) and provides instant assessment using off-the-shelf commercially available resource-constrained edge platforms and video cameras. We demonstrate the scalability, flexibility, and compatibility of the RhythmEdge by deploying it on three resource-constrained platforms of differing architectures (NVIDIA Jetson Nano, Google Coral Development Board, Raspberry Pi) and three heterogeneous cameras of differing sensitivity, resolution, properties (web camera, action camera, and DSLR). RhythmEdge further stores longitudinal cardiovascular information and provides instant notification to the users. We thoroughly test the prototype stability, latency, and feasibility for three edge computing platforms by profiling their runtime, memory, and power usage.
Zahid Hasan 0001, Emon Dey, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy, Archan Misra
SMARTCOMP4
2022 Assessing the Feasibility of Exploiting Edge Computing for Real- Time Monitoring of Flash Floods
abstract
Monitoring flash floods and providing just-in-time notification to city officials for taking appropriate action and prompt intervention is crucial for any smart city located in flood-prone areas around the world. Flood monitoring systems that exploit image analysis via Machine Learning (ML) techniques have been already proposed in literature. Such systems, however, adopt a cloud-based approach that generates significant data traffic and could be susceptible to failures due to network outages. In such a framework, images are continuously offloaded from cameras deployed in flood-prone areas of the city towards a cloud infrastructure where a service is deployed to analyze the images and detect the rise of water in rivers or city canals in a timely way. In this paper, we present the activities of the project EdgeFlooding, which aims at investigating the opportunity of adopting a distributed approach based on edge computing for the implementation of more resilient and reliable flash flood monitoring systems, that helps mitigate the limitations of the cloud-based systems. We have developed a prototype of an edge computing flood monitoring system based on micro-services, and we run an extensive set of experiments exploiting one European Fed4Fire+ testbed, i.e., the Grid'5000 testbed. The aim of those experiments is to assess whether a distributed edge/cloud computing approach is feasible for the implementation of future flood or environmental monitoring systems.
Francesca Righetti, Carlo Vallati, Andrea Klaus Tubak, Nirmalya Roy, Bipendra Basnyat, Giuseppe Anastasi
SMARTCOMP4
2022 DeCoach: Deep Learning-based Coaching for Badminton Player Assessment
abstract
Wearable devices have gained immense popularity among various pervasive computing and Internet-of-Things (IoT) applications in the past decade. Sports analytics researchers recently focused on improving a player’s performance to help devise a winning strategy based on the player’s gameplay. Especially in a racquet-based badminton sport, it is assumed that handling the racquet during the gameplay is one of the primary reasons to influence the players’ performance. On the contrary, we posit that the players’ stance, body movements, and posture are equally significant in evaluating a player’s performance during the game. A shot characterized by a recommended posture, stance, and body movements allows a player to play a stroke efficiently, thus aiding the player in guiding the shuttle to strategic spots and making it difficult for the opponent to return the shot and score a point. Relying on this hypothesis, we propose DeCoach, a data-driven framework that leverages the stance and posture of the players and ranks them based on their performances. In this effort, we first employ a deep learning-based algorithm to classify the strokes and stances of the players. Secondly, we propose a distance-based methodology to compare the obtained stance of a player with that of a professional player. Finally, we devise a deep learning-based regressor to predict the player’s performance which commences with ranking based on their performance. We evaluate DeCoach using our in-house dataset, Badminton Activity Recognition (BAR) Dataset that is collected using inertial measurement unit (IMU) sensors by placing them on the upper and lower limbs of the players. The BAR dataset is collected from 11 players in the controlled and uncontrolled environment settings for 12 frequently played shots in the game. Empirical results indicate that DeCoach achieves 89.09% accuracy for strokes detection and R2 score of 88.84% in estimating the players’ performance.
Indrajeet Ghosh, Sreenivasan Ramasamy Ramamurthy, Avijoy Chakma, Nirmalya Roy
Pervasive Mob. Comput.4
2022 STAR-Lite: A light-weight scalable self-taught learning framework for older adults' activity recognition
Sreenivasan Ramasamy Ramamurthy, Indrajeet Ghosh, Aryya Gangopadhyay, Elizabeth Galik, Nirmalya Roy
Pervasive Mob. Comput.5
2021 A Scalable and Domain Adaptive Respiratory Symptoms Detection Framework using Earables
abstract
The COVID-19 pandemic has brought a devastating impact on human health across the globe, and people are still observing face-masking as a preventive measure to contain the spread of COVID-19. Coughing is one of the major transmission mediums of COVID-19, and early cough detection could play a significant r ole i n p reventing t he s pread o f t his life-threatening virus. Many approaches have been proposed for developing systems to detect coughing and other respiratory symptoms in literature, but earable devices are not well-studied and investigated for respiratory symptom detection. In this work, we posited an acoustic research prototype (earable device) - eSense that has acoustic and IMU sensors embedded into user-convenient earbuds to address the following issues: (i) feasibility of the earables in detecting respiratory symptoms, and (ii) scalability of trained machine learning models in the presence of unseen data samples. We performed experimentation with both shallow and deep learning models on the eSense collected data samples. We observed that the deep learning model outperforms the shallow learning models achieving 97% accuracy. Furthermore, we investigated the scalability of the deep learning model on unseen datasets and noticed that the performance of the deep learning model deteriorates when trained on a particular dataset and tested on an unseen dataset. To mitigate such challenges, we postulated an adversarial domain adaptation technique that helps improve the performance of our respiratory symptoms detection framework by a substantial margin.
Avijoy Chakma, Nirmalya Roy
IEEE BigData3
2021 BuiltNet: Graph based Spatio-Temporal Indoor Thermal Variation Detection
abstract
Monitoring thermal condition with thermal cameras is a potential non-intrusive way to supervise the structural well-being of buildings. Thermal variation can infer various structural damages or construction deficiencies including air leakages through inside and outside surfaces of buildings. Frequent monitoring with thermal images can track the thermal characteristics of different places of built environments which helps to prevent damages beforehand. Previous literature studied thermal conditions in buildings with thermal images are limited to specific regions with constrained environmental settings. In this work, we propose an automated scalable framework BuiltNet for analyzing spatial and temporal temperature variation over various building elements i.e., walls, windows, doors, etc. using longitudinal thermal images. We collected thermal images from a residential apartment home for 10 minutes in consecutive 4-5 hours on different days. The spatial and temporal relations among different spots in a region from sequential thermal images of the corresponding region are represented by graph. We propose an unsupervised deep clustering algorithm based on graph neural network, considering both spatial and temporal features from longitudinal thermal images. Our analysis on the spatial and temporal features of regions in the collected thermal images (from both day and night of different weather conditions) identifies the thermal variation and characterizes the spatiotemporal dynamics over different places in the built environment.
Naima Khan, Nirmalya Roy
ICMLA2
2021 ARIS: A Real Time Edge Computed Accident Risk Inference System
abstract
To deploy an intelligent transport system in urban environment, an effective and real-time accident risk prediction method is required that can help maintain road safety, provide adequate level of medical assistance and transport in case of an emergency. Reducing traffic accidents is an important problem for increasing public safety, so accident analysis and prediction have been a subject of extensive research in recent time. Even if a traffic hazard occurs, a readily deployable structure with an accurate prediction of accident can contribute to better management of rescue resources. But the significant shortcomings of current studies are the use of small-scale datasets with minimal scope, being based on extensive data sets, and not being applicable for real-time purposes. To overcome these challenges, we propose ARIS: a system for real-time traffic accident prediction built on a traffic accident dataset named ‘US-Accidents’ which covers 49 states of United States, collected from February 2016 to June 2020. Our approach is based on a deep neural network model that utilizes a variety of data characteristics, such as time-sensitive weather data, textual information, and discerning factors. We have tested ARIS against multiple baselines through a comprehensive series of experiments across several major cities of USA, and we have noticed significant improvement during inference especially in detecting accident classes. Additionally, to make our model edge-implementable we have compressed our model using a joint technique of magnitude-based weight pruning and model quantization. We have also demonstrated the inference results along with power consumption profiling after deploying the model on a resource constrained environment that consists of Intel Neural Compute Stick 2 (NCS2) with Raspberry Pi 4B (RPi4). Our investigation and observations indicate major improvements to predict unusual traffic accident event even after model compression and deployment. We have managed to reduce the model size and inference time by ≈ 6x, and ≈ 70 % respectively with insignificant drop in performance. Furthermore, to better understand the importance of each individual type of variables used in our analysis, we have showcased a comprehensive ablation study.
Pretom Roy Ovi, Emon Dey, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP3
2021 STAR: A Scalable Self-taught Learning Framework for Older Adults' Activity Recognition
abstract
Activity Recognition (AR) in older adults living with Neurocognitive disorders caused by diseases such as Alzheimer’s is still a challenging research problem. The inherent natural variation in performing an activity increases while repeating the same activity for an older adult, let alone the variation introduced when another older adult performs the same activity. Moreover, the challenges in acquiring the labeled data while preserving the privacy, availability of annotators with domain knowledge, aversion towards cameras even for a minimal amount of time for ground truth data collection, and psychological and mental health status make AR for older adults challenging. In this paper, we postulate a self-taught learning-based approach that helps recognize activities with variations that are not being directly seen during the training phase. We hypothesize that the features extracted using deep architectures from unlabeled data instances can learn general underlying representations of activities efficiently and help improve activity classification in a supervised setting, although the data instances in labeled data do not follow the generative distribution of that of unlabeled data. We posit real data from a retirement community center using our in-house SenseBox infrastructure and survey-based assessments concurrently done by a clinical evaluator to study the relationship between activities and functional/behavioral health of older adults. We evaluate our proposed self-taught learning-based approach, STAR, using the presented in-house Alzheimer’s Activity Recognition (AAR) dataset acquired in a real-world deployment in 25 homes which outperforms the state-of-the-art algorithm by about 20%.
Sreenivasan Ramasamy Ramamurthy, Indrajeet Ghosh, Aryya Gangopadhyay, Elizabeth Galik, Nirmalya Roy
SMARTCOMP5
2021 Tracking and Behavior Augmented Activity Recognition for Multiple Inhabitants
abstract
We develop CACE (Constraints And Correlations mining Engine), a framework that significantly improves the recognition accuracy of complex daily activities in multi-inhabitant smarthomes. CACE views the implicit relationships between the activities of multiple people as an asset, and exploits such constraints and correlations in a hierarchical fashion, taking advantage of both personspecific sensor data (generated by wearable devices) and person-independent ambient sensor data (generated by ambient sensors). To effectively utilize such couplings, CACE first uses a multi-target particle filtering approach over ambient sensors captured movement data, to identify the number of distinct users and infer individual-specific movement trajectories. We then utilize a Hierarchical Dynamic Bayesian Network (HDBN)-based model for activity recognition. This model utilizes the inter-and-intra individual correlations and constraints, at both micro-activity and macro-activity levels, to recognize individual activities accurately. These constraints are learnt automatically using data-mining techniques, and help to dramatically reduce the computational complexity of HDBN-based inferencing. Empirical studies using a real-world testbed of five multi-inhabitant smarthomes shows that CACE is able to achieve an activity recognition accuracy of 95%, with a 16-fold reduction in computational overhead compared to traditional hybrid classification approaches.
Mohammad Arif Ul Alam, Nirmalya Roy, Archan Misra
IEEE Trans. Mob. Comput.2
2020 OMAD: On-device Mental Anomaly Detection for Substance and Non-Substance Users
abstract
Stay at home order during the COVID-19 helps flatten the curve but ironically, instigate mental health problems among the people who have Substance Use Disorders. Measuring the electrical activity signals in brain using off-the-shelf consumer wearable devices such as smart wristwatch and mapping them in real time to underlying mood, behavioral and emotional changes play striking roles in postulating mental health anomalies. In this work, we propose to implement a wearable, On-device Mental Anomaly Detection (OMAD) system to detect anomalous behaviors and activities that render to mental health problems and help clinicians to design effective intervention strategies. We propose an intrinsic artifact removal model on Electroencephalogram (EEG) signal to better correlate the fine-grained behavioral changes. We design model compression technique on the artifact removal and activity recognition (main) modules. We implement a magnitude-based weight pruning technique both on convolutional neural network and Multilayer Perceptron to employ the inference phase on Nvidia Jetson Nano; one of the tightest resource-constrained devices for wearables. We experimented with three different combinations of feature extractions and artifact removal approaches. We evaluate the performance of OMAD in terms of accuracy, F1 score, memory usage and running time for both unpruned and compressed models using EEG data from both control and treatment (alcoholic) groups for different object recognition tasks. Our artifact removal model and main activity detection model achieved about ≈ 93% and 90% accuracy, respectively with significant reduction in model size (70%) and inference time (31%).
Emon Dey, Nirmalya Roy
BIBE2
2020 LASO: Exploiting Locomotive and Acoustic Signatures over the Edge to Annotate IMU Data for Human Activity Recognition
abstract
Annotated IMU sensor data from smart devices and wearables are essential for developing supervised models for fine-grained human activity recognition, albeit generating sufficient annotated data for diverse human activities under different environments is challenging. Existing approaches primarily use human-in-the-loop based techniques, including active learning; however, they are tedious, costly, and time-consuming. Leveraging the availability of acoustic data from embedded microphones over the data collection devices, in this paper, we propose LASO, a multimodal approach for automated data annotation from acoustic and locomotive information. LASO works over the edge device itself, ensuring that only the annotated IMU data is collected, discarding the acoustic data from the device itself, hence preserving the audio-privacy of the user. In the absence of any pre-existing labeling information, such an auto-annotation is challenging as the IMU data needs to be sessionized for different time-scaled activities in a completely unsupervised manner. We use a change-point detection technique while synchronizing the locomotive information from the IMU data with the acoustic data, and then use pre-trained audio-based activity recognition models for labeling the IMU data while handling the acoustic noises. LASO efficiently annotates IMU data, without any explicit human intervention, with a mean accuracy of $0.93$ ($\pm 0.04$) and $0.78$ ($\pm 0.05$) for two different real-life datasets from workshop and kitchen environments, respectively.
Soumyajit Chatterjee, Avijoy Chakma, Aryya Gangopadhyay, Nirmalya Roy, Bivas Mitra, Sandip Chakraborty 0001
ICMI4
2020 Vision Powered Conversational AI for Easy Human Dialogue Systems
abstract
In this paper, we propose an end to end goal-oriented conversational AI agent that can provide contextual information from a potential hazard site. We posit the conversational agent as a FloodBot capable of seeing, sensing, assessing hazard condition, and ultimately conversing about them. We present our domain-specific FloodBot design-solution and learning-experience from the real-time deployment in a flash flood devastated city that uses state-of-the-art deep learning models. We specifically used computer vision and pertinent natural language processing technologies to empower the conversation power of the FloodBot. To deliver such practical and usable AI, we chain multiple deep learning frameworks and create a human-friendly question-answer based dialogue system. We present our deployment details from the last five months and validate the results using ongoing COVID19's impact on the area as well.
Bipendra Basnyat, Neha Singh 0004, Nirmalya Roy, Aryya Gangopadhyay
MASS3
2020 AutoCogniSys: IoT Assisted Context-Aware Automatic Cognitive Health Assessment
abstract
Cognitive impairment has become epidemic in older adult population. The recent advent of tiny wearable and ambient devices, a.k.a Internet of Things (IoT) provides ample platforms for continuous functional and cognitive health assessment of older adults. In this paper, we design, implement and evaluate AutoCogniSys, a context-aware automated cognitive health assessment system, combining the sensing powers of wearable physiological (Electrodermal Activity, Photoplethysmography) and physical (Accelerometer, Object) sensors in conjunction with ambient sensors. We design appropriate signal processing and machine learning techniques, and develop an automatic cognitive health assessment system in a natural older adults living environment. We validate our approaches using two datasets: (i) a naturalistic sensor data streams related to Activities of Daily Living and mental arousal of 22 older adults recruited in a retirement community center, individually living in their own apartments using a customized inexpensive IoT system (IRB #HP-00064387) and (ii) a publicly available dataset for emotion detection. The performance of AutoCogniSys attests max. 93% of accuracy in assessing cognitive health of older adults.
Mohammad Arif Ul Alam, Nirmalya Roy, Sarah Holmes, Aryya Gangopadhyay, Elizabeth Galik
MobiQuitous2
2020 Flood Detection Framework Fusing The Physical Sensing & Social Sensing
abstract
We investigate the practical challenge of localized flood detection in real smart city environment using the fusion of physical sensor and social sensing models to depict a reliable and accurate flood monitoring and detection framework. Our proposed framework efficiently utilize the physical and social sensing models to provide the flood-related updates to the city officials. We deployed our flood monitoring system in Ellicott City, Maryland, USA and connect it to the social sensing module to perform the flood-related sensor and social data integration and analysis. Our ground-based sensor network model record and performs the predictive data analytic by forecasting the rise in water level (RMSE=0.2) that demonstrates the severity of upcoming flash floods whereas, our social sensing model helps collect and track the flood-related feeds from Twitter. We employ a pre-trained model and inductive transfer learning based approach to classify the flood-related tweets with 90% accuracy in the use of unseen target flood events. Finally our flood detection framework categorizes the flood relevant localized contextual details into more meaningful classes in order to help the emergency services and local authorities for effective decision making.
Neha Singh 0004, Bipendra Basnyat, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP3
2020 Design and Deployment of a Flash Flood Monitoring IoT: Challenges and Opportunities
abstract
Successful implementation of the Internet of Thing (IoT) is precursory to a thriving smart city. However, the technical, physical, and environmental conditions can often pose challenges in their successful deployments. The deployment is further complicated if the time and location of implementation are amidst a natural disaster. In this work, we use flash flood detection as a natural hazard testbed and describe various IoT deployment, our progression, and first-hand experience from those implementations. We compare and contrast three IoTs and their performance in real-time execution. Next, we discuss systems architecture and their end-to-end design and present lessons learned from these heterogeneous deployments. Additionally, we evaluate and outline our observations, challenges, and opportunities for further improvement. We also formulate standard evaluation metrics for their scoring and document our deployment journey.
Bipendra Basnyat, Neha Singh 0004, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP3
2020 Towards AI Conversing: FloodBot using Deep Learning Model Stacks
abstract
Talking to the electronic device and getting the required information at a minimal time has become today's norm. Although AI-powered conversational agents have percolated the commercial market, their use in a communal setting is still evolving. We postulate that the deployments of chatbots in disaster-prone areas can be beneficial to watch, monitor, and warn people during the crisis. Furthermore, the successful implementation of such technology can be life-saving. In this work, we discuss our deployment of a real-time flood monitoring chatbot called FloodBot. We collect, annotate and visually parse images from potentially hazardous areas. We detect the flood conditions and identify objects in harm's way by stacking deep learning models such as a convolutional neural network (CNN), single-shot multi-box object detection (SSD). We then feed the image contents to a knowledge base of our artificially intelligent FloodBot and explore its AI-Conversing power using end to end memory network. We also showcase the power of cross-domain transfer learning and model fusion techniques. In this work, we discuss our deployment of a real-time flood monitoring chatbot called FloodBot. We collect, annotate and visually parse images from potentially hazardous areas. We detect the flood conditions and identify objects in harm's way by stacking deep learning models such as a convolutional neural network (CNN), single-shot multi-box object detection (SSD). We then feed the image contents to a knowledge base of our artificially intelligent FloodBot and explore its AI-Conversing power using end to end memory network. We also showcase the power of cross-domain transfer learning and model fusion techniques.
Bipendra Basnyat, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP2
2020 Water Quality Assessment with Thermal Images
abstract
Water contamination has been a critical issue in many countries of the world including USA. Physical, chemical, biological, radio-logical substances can be the reason of this contamination. Drinking water systems are allowed to contain chlorine, calcium, lead, arsenic etc., at a certain level. However, there are expensive instruments and paper sensors to detect the quantity of minerals in water. But these instruments are not always convenient for easy determination of the quality of the sample as drinking water. Different minerals in the water reacts to heat heterogeneously. Some minerals (i.e., arsenic) stay in the water with noticeable amount even after reaching to boiling point. However, it requires cheaper and easier process to examine the quality of water samples for drinking from different sources. With this in mind, we experimented few water samples from different places of USA including artificially prepared samples by mixing different impurities. We investigated their heating property with the sample of marked safe drinking water. We collected thermal images with 10-seconds interval during cooling period of hot water samples from the boiling point to room temperature. We extracted features for each of the water samples with the combination of convolution and recurrent neural network based model and classified different water samples based on the added impurity types and sources from where the samples were collected. We also showed the feature distances of these water samples with the safe water sample. Our proposed framework can differentiate features for different impurities added in the water samples and detect different category of impurities with average accuracy of 70%.
Naima Khan, Nirmalya Roy
SMARTCOMP2
2019 Active Deep Learning for Activity Recognition with Context Aware Annotator Selection
abstract
Machine learning models are bounded by the credibility of ground truth data used for both training and testing. Regardless of the problem domain, this ground truth annotation is objectively manual and tedious as it needs considerable amount of human intervention. With the advent of Active Learning with multiple annotators, the burden can be somewhat mitigated by actively acquiring labels of most informative data instances. However, multiple annotators with varying degrees of expertise poses new set of challenges in terms of quality of the label received and availability of the annotator. Due to limited amount of ground truth information addressing the variabilities of Activity of Daily Living (ADLs), activity recognition models using wearable and mobile devices are still not robust enough for real-world deployment. In this paper, we first propose an active learning combined deep model which updates its network parameters based on the optimization of a joint loss function. We then propose a novel annotator selection model by exploiting the relationships among the users while considering their heterogeneity with respect to their expertise, physical and spatial context. Our proposed model leverages model-free deep reinforcement learning in a partially observable environment setting to capture the action-reward interaction among multiple annotators. Our experiments in real-world settings exhibit that our active deep model converges to optimal accuracy with fewer labeled instances and achieves ~8% improvement in accuracy in fewer iterations.
H. M. Sajjad Hossain, Nirmalya Roy
KDD2
2019 AugToAct: scaling complex human activity recognition with few labels
abstract
Human activity recognition (HAR) from wearable sensor data has recently gained widespread adoption in a number of fields. However, recognizing complex human activities, postural and rhythmic body movements (e.g. dance, sports) is challenging due to the lack of domain-specific labeling information, the perpetual variability in human movement kinematics profiles due to age, sex, dexterity and the level of professional training. In this paper, we propose a deep activity recognition model to work with limited labeled data, both for simple and complex human activities. To mitigate the intra and inter-user spatio-temporal variability of movements, we posit novel data augmentation and domain normalization techniques. We depict a semi-supervised technique that learns noise and transformation invariant feature representation from sparsely labeled data to accommodate intra-personal and inter-user variations of human movement kinematics. We also postulate a transfer learning approach to learn domain invariant feature representations by minimizing the feature distribution distance between the source and target domains. We showcase the improved performance of our proposed framework, AugToAct, using a public HAR dataset. We also design our own data collection, annotation and experimental setup on complex dance activity recognition steps and kinematics movements where we achieved higher performance metrics with limited label data compared to simple activity recognition tasks.
Abu Zaher Md Faridee, Md Abdullah Al Hafiz Khan, Nilavra Pathak, Nirmalya Roy
MobiQuitous4
2019 Identifying the Context of Hurricane Posts on Twitter using Wavelet Features
abstract
With the increase of natural disasters all over the world, we are in crucial need of innovative solutions with inexpensive implementations to assist the emergency response systems. Information collected through conventional sources (e.g., incident reports, 911 calls, physical volunteers, etc.) are proving to be insufficient [1]. Responsible organizations are now leaning towards research grounds that explore digital human connectivity and freely available sources of information. U.S. Geological Survey and Federal Emergency Management Agency (FEMA) introduced Critical Lifeline (CLL) s which identifies the most significant areas that require immediate attention in case of natural disasters. These organizations applied crowdsourcing by connecting digital volunteer networks to collect data on the critical lifelines from data sources including social media [3], [4], [5]. In the past couple of years, during some of the deadly hurricanes (e.g., Harvey, IRMA, Maria, Michael, Florence, etc.), people took on different social media platforms like never seen before, in search of help for rescue, shelter, and relief. Their posts reflect crisis updates and their real-time observations on the devastation that they witness. In this paper, we propose a methodology to build and analyze time-frequency features of words on social media to assist the volunteer networks in identifying the context before, during and after a natural disaster and distinguishing contexts connected to the critical lifelines. We employ Continuous Wavelet Transform to help create word features and propose two ways to reduce the dimensions which we use to create word clusters to identify themes of conversations associated with stages of a disaster and these lifelines. We compare two different methodologies of wavelet features and word clusters both qualitatively and quantitatively, to show that wavelet features can identify and separate context without using semantic information as inputs.
Amrita Anam, Aryya Gangopadhyay, Nirmalya Roy
SMARTCOMP3
2019 Developing Machine Learning Based Predictive Models for Smart Policing
abstract
Crimes are problematic where normal social issues are confronted and influence personal satisfaction, financial development, and quality-of-life of a region. There has been a surge in the crime rate over the past couple of years. To reduce the offense rate, law enforcement needs to embrace innovative preventive technological measures. Accurate crime forecasts help to decrease the crime rate. However, predicting criminal activities is difficult due to the high complexity associated with modeling numerous intricate elements. In this work, we employ statistical analysis methods and machine learning models for predicting different types of crimes in New York City, based on 2018 crime datasets. We combine weather, and its temporal attributes like cloud cover, lighting and time of day to identify relevance to crime data. We note that weather-related attributes play a negligible role in crime forecasting. We have evaluated the various performance metrics of crime prediction, with and without the consideration of weather datasets, on different types of crime committed. Our proposed methodology will enable law enforcement to make effective decisions on appropriate resource allocation, including backup officers related to crime type and location.
Lavanya Elluri, Varun Mandalapu, Nirmalya Roy
SMARTCOMP3
2019 Detecting Common Insulation Problems in Built Environments using Thermal Images
abstract
Proper thermal insulation yields optimum energy expenses in buildings by maintaining necessary heat gain or loss through the built envelope. However, improper thermal insulation causes significant energy wastage along with infusing various damages on indoor and outdoor walls of the buildings, for example, damp areas, black stains, cracks, paint bubbles etc. Therefore, it is important to inspect the temperature variations in different areas of the built environments in regular basis. We propose a method for identifying temperature variance in building thermal images based on Symbolic Aggregated Approximation (SAX). Our process helps detect the temperature variation over different image segments and infers the fault prone segments of leakages. We have collected about 50 thermal images associated with different types of wall specific insulation problems in indoor built environment and were able to identify the affected area with approximately 75% accuracy using our proposed method based on temperature variation detection approach.
Naima Khan, Nilavra Pathak, Nirmalya Roy
SMARTCOMP3
2019 Analyzing The Emotions of Crowd For Improving The Emergency Response Services
abstract
Twitter is an extremely popular micro-blogging social platform with millions of users, generating thousands of tweets per second. The huge amount of Twitter data inspire the researchers to explore the trending topics, event detection and event tracking which help to postulate the fine-grained details and situation awareness. Obtaining situational awareness of any event is crucial in various application domains such as natural calamities, man-made disaster and emergency responses. In this paper, we advocate that data analytics on Twitter feeds can help improve the planning and rescue operations and services as provided by the emergency personnel in the event of unusual circumstances. We take an emotional change detection approach and focus on the users’ emotions, concerns and feelings expressed in tweets during the emergency situations, and analyze those feelings and perceptions in the community involved during the events to provide appropriate feedback to emergency responders and local authorities. We employ improved emotion analysis and change point detection techniques to process, discover and infer the spatiotemporal sentiments of the users. We analyze the tweets from recent Las Vegas shooting (Oct. 2017) and note that the changes in the polarity of the sentiments and articulation of the emotional expressions, if captured successfully can be employed as an informative tool for providing feedback to EMS.
Neha Singh 0004, Nirmalya Roy, Aryya Gangopadhyay
Pervasive Mob. Comput.2
2018 Scaling Human Activity Recognition via Deep Learning-based Domain Adaptation
abstract
We investigate the problem of making human activity recognition (AR) scalable-i.e., allowing AR classifiers trained in one context to be readily adapted to a different contextual domain. This is important because AR technologies can achieve high accuracy if the classifiers are trained for a specific individual or device, but show significant degradation when the same classifier is applied context-e.g., to a different device located at a different on-body position. To allow such adaptation without requiring the onerous step of collecting large volumes of labeled training data in the target domain, we proposed a transductive transfer learning model that is specifically tuned to the properties of convolutional neural networks (CNNs). Our model, called HDCNN, assumes that the relative distribution of weights in the different CNN layers will remain invariant, as long as the set of activities being monitored does not change. Evaluation on real-world data shows that HDCNN is able to achieve high accuracy even without any labeled training data in the target domain, and offers even higher accuracy (significantly outperforming competitive shallow and deep classifiers) when even a modest amount of labeled training data is available.
Md Abdullah Al Hafiz Khan, Nirmalya Roy, Archan Misra
PerCom2
2018 Evaluating Disaster Time-Line from Social Media with Wavelet Analysis
abstract
For over a decade, social media has proved to be a functional and convenient data source in the Internet of things. Social platforms such as Facebook, Twitter, Instagram, and Reddit have their own styles and purposes. Twitter, among them, has become the most popular platform in the research community due to its nature of attracting people to write brief posts about current and unexpected events (e.g., natural disasters). The immense popularity of such sites has opened a new horizon in `social sensing' to manage disaster response. Sensing through social media platforms can be used to track and analyze natural disasters and evaluate the overall response (e.g., resource allocation, relief, cost and damage estimation). In this paper, we propose a two-step methodology: i) wavelet analysis and ii) predictive modeling to track the progression of a disaster aftermath and predict the time-line. We demonstrate that wavelet features can preserve text semantics and predict the total duration for localized small scale disasters. The experimental results and observations on two real data traces (flash flood in Cummins Falls state park and Arizona swimming hole) showcase that the wavelet features can predict disaster time-line with an error lower than 20% with less than 50% of the data when compared to ground truth.
Amrita Anam, Aryya Gangopadhyay, Nirmalya Roy
SMARTCOMP3
2018 A Flash Flood Categorization System Using Scene-Text Recognition
abstract
Detecting flash floods in real-time and taking rapid actions are of utmost importance to save human lives, loss of infrastructures, and personal properties in a smart city. In this paper, we develop a low-cost low-power cyber-physical System prototype using a Raspberry Pi camera to detect the rising water level. We deployed the system in the real word and collected data in different environmental conditions (early morning in the presence of fog, sunny afternoon, late afternoon with sunsetting). We employ image processing and text recognition techniques to detect the rising water level and articulate several challenges in deploying such a system in the real environment. We envision this prototype design will pave the way for mass deployment of the flash flood detection system with minimal human intervention.
Bipendra Basnyat, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP2
2018 Forecasting Gas Usage for Big Buildings Using Generalized Additive Models and Deep Learning
abstract
Time series behavior of gas consumption is highly irregular, non-stationary, and volatile due to its dependency on the weather, users' habits and lifestyle. This complicates the modeling and forecasting of gas consumption with most of the existing time series modeling techniques, specifically when missing values and outliers are present. To demonstrate and overcome these problems, we investigate two approaches to model the gas consumption, namely Generalized Additive Models (GAM) and Long Short-Term Memory (LSTM). We perform our evaluations on two building datasets from two different continents. We present each selected feature's influence, the tuning parameters, and the characteristics of the gas consumption on their forecasting abilities. We compare the performances of GAM and LSTM with other state-of-the-art forecasting approaches. We show that LSTM outperforms GAM and other existing approaches, however, GAM provides better interpretable results for building management systems (BMS).
Nilavra Pathak, Amadou Ba, Joern Ploennigs, Nirmalya Roy
SMARTCOMP4
2018 Analyzing the Sentiment of Crowd for Improving the Emergency Response Services
abstract
Twitter is an extremely popular micro-blogging social platform with millions of users, generating thousands of tweets per second. The huge amount of Twitter data inspire the researchers to explore the trending topics, event detection and event tracking which help to postulate the fine-grained details and situation awareness. Obtaining situational awareness of any event is crucial in various application domains such as natural calamities, man made disaster and emergency responses. In this paper, we advocate that data analytics on Twitter feeds can help improve the planning and rescue operations and services as provided by the emergency personnel in the event of unusual circumstances. We take a different approach and focus on the users' emotions, concerns and feelings expressed in tweets during the emergency situations, and analyze those feelings and perceptions in the community involved during the events to provide appropriate feedback to emergency responders and local authorities. We employ sentiment analysis and change point detection techniques to process, discover and infer the spatiotemporal sentiments of the users. We analyze the tweets from recent Las Vegas shooting (Oct. 2017) and note that the changes in the polarity of the sentiments and articulation of the emotional expressions, if captured successfully can be employed as an informative tool for providing feedback to EMS.
Neha Singh 0004, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP2
2018 An Active Sleep Monitoring Framework Using Wearables
abstract
Sleep is the most important aspect of healthy and active living. The right amount of sleep at the right time helps an individual to protect his or her physical, mental, and cognitive health and maintain his or her quality of life. The most durative of the Activities of Daily Living (ADL), sleep has a major synergic influence on a person’s fuctional, behavioral, and cognitive health. A deep understanding of sleep behavior and its relationship with its physiological signals, and contexts (such as eye or body movements), is necessary to design and develop a robust intelligent sleep monitoring system. In this article, we propose an intelligent algorithm to detect the microscopic states of sleep that fundamentally constitute the components of good and bad sleeping behaviors and thus help shape the formative assessment of sleep quality. Our initial analysis includes the investigation of several classification techniques to identify and correlate the relationship of microscopic sleep states with overall sleep behavior. Subsequently, we also propose an online algorithm based on change point detection to process and classify the microscopic sleep states. We also develop a lightweight version of the proposed algorithm for real-time sleep monitoring, recognition, and assessment at scale. For a larger deployment of our proposed model across a community of individuals, we propose an active-learning-based methodology to reduce the effort of ground-truth data collection and labeling. Finally, we evaluate the performance of our proposed algorithms on real data traces and demonstrate the efficacy of our models for detecting and assessing the fine-grained sleep states beyond an individual.
H. M. Sajjad Hossain, Sreenivasan Ramasamy Ramamurthy, Md Abdullah Al Hafiz Khan, Nirmalya Roy
ACM Trans. Interact. Intell. Syst.4
2017 COAR: Collaborative and Opportunistic Human Activity Recognition
abstract
The new era of consumer devices ranging from smartphones, smartwatches, and smart jewelries augmented with our everyday activities and lifestyle help postulate human behavior, activity, gesture, social interaction, and gaming experience. Intelligently tasking and sharing the sensing, processing, storing, and computing tasks among those emerging consumer-friendly commodity devices based on their proximities, advocate the development of resource-aware collaborative and opportunistic smart living applications. Motivated by this emerging subsets of phenomenal applications, we first propose a finite-state machine (FSM) based human activity recognition framework which opportunistically exploits the relevant data sources from multiple heterogeneous devices to help infer a variety of user contexts. We depict a lightweight maximum entropy based classifier and exploit the a-priori conditional dependences among the feature sets to opportunistically select the right set of sensors with the most appropriate devices. Experimental results on real data traces demonstrate that our proposed Collaborative Opportunistic Activity Recognition, COAR framework helps infer the activities of daily living with ≈ 90% accuracy.
Md Abdullah Al Hafiz Khan, Nirmalya Roy, H. M. Sajjad Hossain
DCOSS2
2017 Unseen Activity Recognitions: A Hierarchical Active Transfer Learning Approach
abstract
Human activity recognition (AR) is an essential element for user-centric and context-aware applications. While previous studies showed promising results using various machine learning algorithms, most of them can only recognize the activities that were previously seen in the training data. We investigate the challenges of improving the recognition of unseen daily activities in smart home environment, by better exploiting the hierarchical taxonomy of complex daily activities. We first (a) design a hierarchical representation of complex activity taxonomy in terms of human-readable semantic attributes, and (b) develop a hierarchy of classifiers which incorporates a cluster tree built on the domain knowledge from training samples. Though this model is rich in recognizing complex activities that are previously seen in training data, it is not well versed to recognize unseen complex activities without new training samples. To tackle this challenge, we extend Hierarchical Active Transfer Learning (HATL) approach that exploits semantic attribute cluster structure of complex activities shared between seen (source) and unseen (target) activity domains. Our approach employs transfer and active learning to help label target domain unlabeled data by spawning the most effective queries. We evaluated our approach with two real-time smart home systems (IRB #HP-00064387) which corroborates radical improvements in recognizing unseen complex activities.
Mohammad Arif Ul Alam, Nirmalya Roy
ICDCS2
2017 Analyzing Social Media Texts and Images to Assess the Impact of Flash Floods in Cities
abstract
Computer Vision and Image Processing are emerging research paradigms. The increasing popularity of social media, micro- blogging services and ubiquitous availability of high-resolution smartphone cameras with pervasive connectivity are propelling our digital footprints and cyber activities. Such online human footprints related with an event-of-interest, if mined appropriately, can provide meaningful information to analyze the current course and pre- and post- impact leading to the organizational planning of various real-time smart city applications. In this paper, we investigate the narrative (texts) and visual (images) components of Twitter feeds to improve the results of queries by exploiting the deep contexts of each data modality. We employ Latent Semantic Analysis (LSA)-based techniques to analyze the texts and Discrete Cosine Transformation (DCT) to analyze the images which help establish the cross-correlations between the textual and image dimensions of a query. While each of the data dimensions helps improve the results of a specific query on its own, the contributions from the dual modalities can potentially provide insights that are greater than what can be obtained from the individual modalities. We validate our proposed approach using real Twitter feeds from a recent devastating flash flood in Ellicott City near the University of Maryland campus. Our results show that the images and texts can be classified with 67\% and 94\% accuracies respectively
Bipendra Basnyat, Amrita Anam, Neha Singh 0004, Aryya Gangopadhyay, Nirmalya Roy
SMARTCOMP5
2017 A smart segmentation technique towards improved infrequent non-speech gestural activity recognition model
Mohammad Arif Ul Alam, Nirmalya Roy, Aryya Gangopadhyay, Elizabeth Galik
Pervasive Mob. Comput.2
2017 Active learning enabled activity recognition
H. M. Sajjad Hossain, Md Abdullah Al Hafiz Khan, Nirmalya Roy
Pervasive Mob. Comput.3
2016 CACE: Exploiting Behavioral Interactions for Improved Activity Recognition in Multi-inhabitant Smart Homes
abstract
We propose CACE (Constraints And Correlations mining Engine) which investigates the challenges of improving the recognition of complex daily activities in multi-inhabitant smart homes, by better exploiting the spatiotemporal relationships across the activities of different individuals. We first propose and develop a loosely-coupled Hierarchical Dynamic Bayesian Network (HDBN), which both (a) captures the hierarchical inference of complex (macro-activity) contexts from lower-layer microactivity context (postural and improved oral gestural context), and (b) embeds the various types of behavioral correlations and constraints (at both micro-and macro-activity contexts) across the individuals. While this model is rich in terms of accuracy, it is computationally prohibitive, due to the explosive increase in the number of jointly-defined states. To tackle this challenge, we employ data mining to learn behaviorally-driven context correlations in the form of association rules, we then use such rules to prune the state space dramatically. To evaluate our framework, we build a customized smart home system and collected naturalistic multi-inhabitant smart home activities data. The system performance is illustrated with results from real-time system deployment experiences in a smart home environment reveals a radical (max 16 fold) reduction in the computational overhead compared to traditional hybrid classification approaches, as well as an improved activity recognition accuracy of max 95%.
Mohammad Arif Ul Alam, Nirmalya Roy, Archan Misra, Joseph Taylor
ICDCS2
2016 RAM: Radar-based activity monitor
abstract
Activity recognition has applications in a variety of human-in-the-loop settings such as smart home health monitoring, green building energy and occupancy management, intelligent transportation, and participatory sensing. While fine-grained activity recognition systems and approaches help enable a multitude of novel applications, discovering them with non-intrusive ambient sensor systems pose challenging design, as well as data processing, mining, and activity recognition issues. In this paper, we develop a low-cost heterogeneous Radar based Activity Monitoring (RAM) system for recognizing fine-grained activities. We exploit the feasibility of using an array of heterogeneous micro-doppler radars to recognize low-level activities. We prototype a short-range and a long-range radar system and evaluate the feasibility of using the system for fine-grained activity recognition. In our evaluation, using real data traces, we show that our system can detect fine-grained user activities with 92.84% accuracy.
Md Abdullah Al Hafiz Khan, Ruthvik Kukkapalli, Piyush Waradpande, Sekar Kulandaivel, Nilanjan Banerjee, Nirmalya Roy, Ryan W. Robucci
INFOCOM6
2016 Active learning enabled activity recognition
abstract
Activity recognition in smart environment has been investigated rigorously in recent years. Researchers are enhancing the underlying activity discovery and recognition process by adding various dimensions and functionalities. But one significant barrier still persists which is collecting the ground truth information. Ground truth is very important to initialize a supervised learning of activities. Due to a large variety in number of Activities of Daily Living (ADLs), acknowledging them in a supervised way is a non-trivial research problem. Most of the previous researches have referenced a subset of ADLs and to initialize their model, they acquire a vast amount of informative labeled training data. On the other hand to collect ground truth and differentiate ADLs, human intervention is indispensable. As a result it takes an immense effort and raises privacy concerns to collect a reasonable amount of labeled data. In this paper, we propose to use active learning to alleviate the labeling effort and ground truth data collection in activity recognition pipeline. We investigate and analyze different active learning strategies to scale activity recognition and propose a dynamic k-means clustering based active learning approach. Experimental results on real data traces from a retirement community-(IRB #HP-00064387) help validate the early promise of our approach.
H. M. Sajjad Hossain, Nirmalya Roy, Md Abdullah Al Hafiz Khan
PerCom2
2016 Special Issue-Big Data for Healthcare
Sriram Chellappan, Nirmalya Roy, Sajal K. Das 0001
Pervasive Mob. Comput.2
2016 Fine-grained appliance usage and energy monitoring through mobile and power-line sensing
Nirmalya Roy, Nilavra Pathak, Archan Misra
Pervasive Mob. Comput.1
2016 Determining Quality- and Energy-Aware Multiple Contexts in Pervasive Computing Environments
abstract
In pervasive computing environments, understanding the context of an entity is essential for adapting the application behavior to changing situations. In our view, context is a high-level representation of a user or entity's state and can capture location, activities, social relationships, capabilities, etc. Inherently, however, these high-level context metrics are difficult to capture using uni-modal sensors only and must therefore be inferred using multi-modal sensors. A key challenge in supporting context-aware pervasive computing is how to determine multiple high-level context metrics simultaneously and energy-efficiently using low-level sensor data streams collected from the environment and the entities present therein. A key challenge is addressing the fact that the algorithms that determine different high-level context metrics may compete for access to low-level sensors. In this paper, we first highlight the complexities of determining multiple context metrics as compared to a single context and then develop a novel framework and practical implementation for this problem. The proposed framework captures the tradeoff between the accuracy of estimating multiple context metrics and the overhead incurred in acquiring the necessary sensor data streams. In particular, we develop two variants of a heuristic algorithm for multi-context search that compute the optimal set of sensors contributing to the multi-context determination as well as the associated parameters of the sensing tasks (e.g., the frequency of data acquisition). Our goal is to satisfy the application requirements for a specified accuracy at a minimum cost. We compare the performance of our heuristics with a brute-force based approach for multi-context determination. Experimental results with SunSPOT, Shimmer and Smartphone sensors in smart home environments demonstrate the potential impact of the proposed framework.
Nirmalya Roy, Archan Misra, Sajal K. Das 0001, Christine Julien 0001
IEEE/ACM Trans. Netw.1
2015 Sleep Well: A Sound Sleep Monitoring Framework for Community Scaling
abstract
Following healthy lifestyle is a key for active living. Regular exercise, controlled diet and sound sleep play an invisible role on the well being and independent living of the people. Sleep being the most durative activities of daily living (ADL) has a major synergistic influence on people's mental, physical and cognitive health. Understanding the sleep behavior longitudinally and its underpinning clausal relationships with physiological signals and contexts (such as eye or body movement etc.) horizontally responsible for a sound or disruptive sleep pattern help provide meaningful information for promoting healthy lifestyle and designing appropriate intervention strategy. In this paper we propose to detect the microscopic states of the sleep which fundamentally constitute the components of a good or bad sleeping behavior and help shape the formative assessment of sleep quality. We initially investigate several classification techniques to identify and correlate the relationship of microscopic sleep states with the overall sleep behavior. Subsequently we propose an online algorithm based on change point detection to better process and classify the microscopic sleep states and then test a lightweight version of this algorithm for real time sleep monitoring activity recognition and assessment at scale. For a larger deployment of our proposed model across a community of individuals we propose an active learning based methodology by reducing the effort of ground truth data collection. We evaluate the performance of our proposed algorithms on real data traces, and demonstrate the efficacy of our models for detecting and assessing fine-grained sleep states beyond an individual.
H. M. Sajjad Hossain, Nirmalya Roy, Md Abdullah Al Hafiz Khan
MDM (1)2
2015 SensePresence: Infrastructure-Less Occupancy Detection for Opportunistic Sensing Applications
abstract
Predicting the occupancy related information in an environment has been investigated to satisfy the myriad requirements of various evolving pervasive, ubiquitous, opportunistic and participatory sensing applications. Infrastructure and ambient sensors based techniques have been leveraged largely to determine the occupancy of an environment incurring a significant deployment and retrofitting costs. In this paper, we advocate an infrastructure-less zero-configuration multimodal smartphone sensor-based techniques to detect fine-grained occupancy information. We propose to exploit opportunistically smartphones' acoustic sensors in presence of human conversation and motion sensors in absence of any conversational data. We develop a novel speaker estimation algorithm based on unsupervised clustering of overlapped and non-overlapped conversational data to determine the number of occupants in a crowded environment. We also design a hybrid approach combining acoustic sensing opportunistically with locomotive model to further improve the occupancy detection accuracy. We evaluate our algorithms in different contexts, conversational, silence and mixed in presence of 10 domestic users. Our experimental results on real-life data traces collected from 10 occupants in natural setting show that using this hybrid approach we can achieve approximately 0.76 error count distance for occupancy detection accuracy on average.
Md Abdullah Al Hafiz Khan, H. M. Sajjad Hossain, Nirmalya Roy
MDM (2)3
2015 AARPA: Combining Mobile and Power-Line Sensing for Fine-Grained Appliance Usage and Energy Monitoring
abstract
To promote energy-efficient operations in residential and office buildings, non-intrusive load monitoring (NILM) techniques have been proposed to infer the fine-grained power consumption and usage patterns of appliances from power-line measurement data. Fine-grained monitoring of everyday appliances (such as toasters and coffee makers) can not only promote energy-efficient building operations, but also provide unique insights into the context and activities of individuals. Current building-level NILM techniques are unable to identify the consumption characteristics of relatively low-load appliances, whereas smart-plug based solutions incur significant deployment and maintenance costs. In this paper, we investigate an intermediate architecture, where smart circuit breakers provide measurements of aggregate power consumption at room (or section) level granularity. We then investigate techniques to identify the usage and energy consumption of individual appliances from such measurements. We first develop a novel correlation-based approach called CBPA to identify individual appliances based on both their unique transient and steady-state power signatures. While promising, CBPA fails when the set of candidate appliances is too large. To further improve the accuracy of appliance level usage estimation, we then propose a hybrid system called AARPA, which uses mobile sensing to first infer high-level activities of daily living (ADLs), and then uses knowledge of such ADLs to effectively reduce the set of candidate appliances that potentially contribute to the aggregate readings at any point. We evaluate two variants of this algorithm, and show, using real-life data traces gathered from 10 domestic users, that our fusion of mobile and power-line sensing is very promising: it identified all devices that were used in each data trace, and it identified the usage duration and energy consumption of low-load consumer appliances with 87% accuracy.
Nirmalya Roy, Nilavra Pathak, Archan Misra
MDM (1)1
2015 Acoustic based appliance state identifications for fine-grained energy analytics
abstract
Fine-grained monitoring of everyday appliances can provide better feedback to the consumers and motivate them to change behavior in order to reduce their energy usage. It also helps to detect abnormal power consumption events, long-term appliance malfunctions and potential safety concerns. Commercially available plug meters can be used for individual appliance monitoring but for an entire house, each such individual plug meters are expensive and tedious to setup. Alternative methods relying on Non-Intrusive Load Monitoring techniques help disaggregate electricity consumption data and learn about the individual appliance's power states and signatures. However fine-grained events (e.g., appliance malfunctions, abnormal power consumption, etc.) remain undetected and thus inferred contexts (such as safety hazards etc.) become invisible. In this work, we correlate an appliance's inherent acoustic noise with its energy consumption pattern individually and in presence of multiple appliances. We initially investigate classification techniques to establish the relationship between appliance power and acoustic states for efficient energy disaggregation and abnormal events detection. While promising, this approach fails when there are multiple appliances simultaneously in `ON' state. To further improve the accuracy of our energy disaggregation algorithm, we propose a probabilistic graphical model, based on a variation of Factorial Hidden Markov Model (FHMM) for multiple appliances energy disaggregation. We combine our probabilistic model with the appliances acoustic analytics and postulate a hybrid model for energy disaggregation. Our approach helps to improve the performance of energy disaggregation algorithms and provide critical insights on appliance longevity, abnormal power consumption, consumer behavior and their everyday lifestyle activities. We evaluate the performance of our proposed algorithms on real data traces and show that the fusion of acoustic and power signatures can successfully detect a number of appliances with 95% accuracy.
Nilavra Pathak, Md Abdullah Al Hafiz Khan, Nirmalya Roy
PerCom3
2014 Immersive Physiotherapy: Challenges for Smart Living Environments and Inclusive Communities
Nirmalya Roy, Christine Julien 0001
ICOST1
2014 Monitoring Patient Recovery Using Wireless Physiotherapy Devices
Nirmalya Roy, Brooks Reed Kindle
ICOST1
2014 GeSmart: A gestural activity recognition model for predicting behavioral health
abstract
To promote independent living for elderly population activity recognition based approaches have been investigated deeply to infer the activities of daily living (ADLs) and instrumental activities of daily living (I-ADLs). Deriving and integrating the gestural activities (such as talking, coughing, and deglutition etc.) along with activity recognition approaches can not only help identify the daily activities or social interaction of the older adults but also provide unique insights into their long-term health care, wellness management and ambulatory conditions. Gestural activities (GAs), in general, help identify fine-grained physiological symptoms and chronic psychological conditions which are not directly observable from traditional activities of daily living. In this paper, we propose GeSmart, an energy efficient wearable smart earring based GA recognition model for detecting a combination of speech and non-speech events. To capture the GAs we propose to use only the accelerometer sensor inside our smart earring due to its energy efficient operations and ubiquitous presence in everyday wearable devices. We present initial results and insights based on a C4.5 classification algorithm to infer the infrequent GAs. Subsequently, we propose a novel change-point detection based hybrid classification method exploiting the emerging patterns in a variety of GAs to detect and infer infrequent GAs. Experimental results based on real data traces collected from 10 users demonstrate that this approach improves the accuracy of GAs classification by over 23%, compared to previously proposed pure classification-based solutions. We also note that the accelerometer sensor based earrings are surprisingly informative and energy efficient (by 2.3 times) for identifying different types of GAs.
Mohammad Arif Ul Alam, Nirmalya Roy
SMARTCOMP2
2014 Special issue on data mining in pervasive environments
Nirmalya Roy, Parisa Rashidi, Lawrence B. Holder, Liming Chen 0001
Pervasive Mob. Comput.1
2013 Infrastructure-assisted smartphone-based ADL recognition in multi-inhabitant smart environments
abstract
We propose a hybrid approach for recognizing complex Activities of Daily Living that lie between the two extremes of intensive use of body-worn sensors and the use of infrastructural sensors. Our approach harnesses the power of infrastructural sensors (e.g., motion sensors) to provide additional `hidden' context (e.g., room-level location) of an individual and combines this context with smartphone-based sensing of micro-level postural/locomotive states. The major novelty is our focus on multi-inhabitant environments, where we show how spatiotemporal constraints can be used to significantly improve the accuracy and computational overhead of traditional coupled-HMM based approaches. Experimental results on a smart home dataset demonstrate that this approach improves the accuracy of complex ADL classification by over 30% compared to pure smartphone-based solutions.
Nirmalya Roy, Archan Misra, Diane J. Cook
PerCom1
2012 Auction-based task allocation with trust management for shared sensor networks
abstract
ABSTRACT Task allocation for wireless sensor networks with multiple concurrent applications (such as target tracking and event detection) requires sharing applications' tasks (such as sensing and computation) and available network resources. In this paper, we model the distributed task allocation problem for multiple concurrent applications by using a reverse combinatorial auction, in which the bidders (sensor nodes) are supposed to bid cost values (according to their available resources) for accomplishing the subset of the applications' tasks. Trust management schemes consist of a powerful tool for the detection of unexpected node behaviors (such as faulty or malicious). It is critical for participants (i.e., bidders and auctioneer) to estimate each other's trustworthiness before initiating the task allocation procedure. To address this issue, we introduce a real‐time trust management module for our auction system that is able to validate the reliable bid value and determine faulty nodes and malicious entities. The main objective of our task allocation scheme is to maximize the network lifetime by sharing tasks and network resources within applications, while enhancing the overall application quality of service (e.g., deadline). We also propose a heuristic two‐phase winner determination protocol to deal with the combinatorial reverse auction problem. Simulation results show that the proposed scheme offers the promising performance and efficiency. Copyright © 2012 John Wiley & Sons, Ltd.
Neda Edalat, Wendong Xiao, Mehul Motani, Nirmalya Roy, Sajal K. Das 0001
Secur. Commun. Networks4
2012 Resource-Optimized Quality-Assured Ambiguous Context Mediation Framework in Pervasive Environments
abstract
Pervasive computing applications often involve sensor-rich networking environments that capture various types of user contexts such as locations, activities, vital signs, and so on. Such context information is useful in a variety of applications, for example, monitoring health information to promote independent living in "aging-in-place” scenarios, or providing safety and security of people and infrastructures. In reality, both sensed and interpreted contexts are often ambiguous, thus leading to potentially dangerous decisions if not properly handled. Therefore, a significant challenge in the design and development of realistic and deployable context-aware services for pervasive computing applications lies in the ability to deal with ambiguous contexts. In this paper, we propose a resource-optimized, quality-assured context mediation framework for sensor networks. The underlying approach is based on efficient context-aware data fusion, information-theoretic reasoning, and selection of sensor parameters, leading to an optimal state estimation. In particular, we apply dynamic Bayesian networks to derive context and deal with context ambiguity or error in a probabilistic manner. Experimental results using SunSPOT sensors demonstrate the promise of this approach.
Nirmalya Roy, Sajal K. Das 0001, Christine Julien 0001
IEEE Trans. Mob. Comput.1
2011 Passive Network-Awareness for Dynamic Resource-Constrained Networks
Agoston Petz, Nirmalya Roy, Chien-Liang Fok, Christine Julien 0001
DAIS3
2011 Combinatorial Auction-Based Task Allocation in Multi-application Wireless Sensor Networks
abstract
Wireless sensor networks (WSNs) are usually assigned tasks for a single application. Recently, the concept of shared sensor networks, which support multiple concurrent applications, has emerged, reducing the deployment and administrative costs, and increasing the usability and efficiency of the network. Supporting task allocation for multiple concurrent applications in sensor networks (such as target tracking, event detection, etc.) requires sharing applications' tasks (such as sensing, computation, etc.) and available network resources. In this paper, we model the distributed task allocation problem for multiple concurrent applications using a reverse combinatorial auction, in which the bidders (sensor nodes) bid the cost value (in terms available resources) for accomplishing the subset of the applications' tasks. The main objective is to maximize the network lifetime by sharing tasks and network resources among applications, while enhancing the overall application QoS (e.g., deadline). We also propose a heuristic two-phase winner determination protocol to solve the combinatorial reverse auction problem. Simulation results show that the proposed scheme offers efficiency and network scalability.
Neda Edalat, Wendong Xiao, Nirmalya Roy, Sajal K. Das 0001, Mehul Motani
EUC3
2011 An energy-efficient quality adaptive framework for multi-modal sensor context recognition
abstract
In pervasive computing environments, understanding the context of an entity is essential for adapting the application behavior to changing situations. In our view, context is a high-level representation of a user or entity's state and can capture location, activities, social relationships, capabilities, etc. Inherently, however, these high-level context metrics are difficult to capture using uni-modal sensors only, and must therefore be inferred with the help of multi-modal sensors. However a key challenge in supporting context-aware pervasive computing environments, is how to determine in an energy-efficient manner multiple (potentially competing) high-level context metrics simultaneously using low-level sensor data streams about the environment and the entities present therein. In this paper, we first highlight the intricacies of determining multiple context metrics as compared to a single context, and then develop a novel framework and practical implementation for this problem. The proposed framework captures the tradeoff between the accuracy of estimating multiple context metrics and the overhead incurred in acquiring the necessary sensor data stream. In particular, we develop a multi-context search heuristic algorithm that computes the optimal set of sensors contributing to the multi-context determination as well as the associated parameters of the sensing tasks. Our goal is to satisfy the application requirements for a specified accuracy at a minimum cost. We compare the performance of our heuristic based framework with a brute-forced approach for multi-context determination. Experimental results with SunSPOT sensors demonstrate the potential impact of the proposed framework.
Nirmalya Roy, Archan Misra, Christine Julien 0001, Sajal K. Das 0001, Jit Biswas
PerCom1
2010 Usability of Semantic Web for Enhancing Digital Living Experience
abstract
The number of different types of devices on the home network is expanding rapidly. While this explosion of innovation provides compelling new devices to consumers, there are challenges ensuring compatibility among these devices and providing a comprehensive user interface that supports consumers managing their digital content across several devices. The Digital Living Network Alliance (DLNA) has specified an architecture that enables interoperability between the various devices and allows a user to enjoy the desired content across several devices. To complement DLNA, we investigate ontological representation from semantic web technology to model the interaction between multiple home devices and to provide an enriched and more comprehensive metadata based multimedia search. This removes the user burden of searching each device individually. Development of a prototype incorporating these ideas has begun, using a well known semantic web toolkit Jena.
Nirmalya Roy, Kevin Brooks, Christine Julien 0001
CCNC1
2010 Modeling Delivery Delay for Flooding in Mobile Ad Hoc Networks
abstract
Mobile ad hoc networks (MANETs) can have widely varying characteristics under different deployments, and previous studies show that the characteristics impact the behavior of routing protocols for MANETs. To deploy applications successfully in MANETs, application developers need to comprehend the potential behavior of any underlying protocol used. In mobile networks, a major component of many of these routing protocols is some form of flooding, which facilitates message delivery over an entire network in a relatively reliable way. Several MANET protocols use flooding to support the distribution of route request messages as well as delivering broadcast packets. Therefore, to developers of applications for MANETs, a major task in understanding MANET protocols is estimating the performance of flooding in a given operating environment. In this work, we develop an analytical model for the delay experienced in flooding a message. We model the one hop delay of a flooding message in MANETs in terms of parameters that can be acquired either from a system configuration or from application designers.
Nirmalya Roy, Christine Julien 0001
ICC2
2010 Supporting pervasive computing applications with active context fusion and semantic context delivery
Nirmalya Roy, Tao Gu 0001, Sajal K. Das 0001
Pervasive Mob. Comput.1
2009 Resolving and mediating ambiguous contexts for pervasive care environments
abstract
Ubiquitous (or smart) healthcare applications envision sensor rich computing and networking environments that can capture various types of contexts of patients (or inhabitants of the environment), such as their location, activities and vital signs. Such context information is useful in providing hea
Nirmalya Roy, Christine Julien 0001, Sajal K. Das 0001
MobiQuitous1
2009 Quality-of-Inference Aware Context Determination
abstract
Energy-efficient determination of an individual's context (both physiological and activity) is an important technical challenge for assisted living environments. Given the expected availability of multiple sensors, context determination may be viewed as an estimation problem over multiple sensor data streams. This paper develops a formal and practically applicable model to capture the tradeoff between the accuracy of context estimation and the communication overheads of sensing. In particular, we propose the use of tolerance ranges to reduce an individual sensor's reporting frequency, while ensuring acceptable accuracy of the derived context. In our vision, applications specify their minimally acceptable value for a quality-of-inference (QoINF) metric. We develop an optimization technique allowing the context service to compute both the best set of sensors and their associated tolerance values that satisfy the specified QoINF at a minimum communication cost. Experimental results with SunSPOT sensors demonstrate the potential impact of this approach.
Nirmalya Roy
PerCom1
2009 Enhancing Availability of Grid Computational Services to Ubiquitous Computing Applications
abstract
The grid is an integrated infrastructure that can play the dual roles of a coordinated resource consumer as well as a donator in distributed computing environments. The enormous growth in the use of mobile and embedded devices in ubiquitous computing environment and their interaction with human beings produces a huge amount of data that need to be processed efficiently anytime anywhere. However, such devices often have limited resources in terms of CPU, storage, battery power, and communication bandwidth. Thus, there is a need to transfer ubiquitous computing application services to more powerful computational resources. In this paper, we investigate the use of the grid as a candidate for provisioning computational services to applications in ubiquitous computing environments. In particular, we present a competitive model that describes the possible interaction between the competing resources in the grid infrastructure as service providers and ubiquitous applications as subscribers. The competition takes place in terms of quality of service (QoS) and cost offered by different grid service providers (GSPs). We also investigate the job allocation of different GSPs by exploiting the noncooperativeness among the strategies. We present the equilibrium behavior of our model facing global competition under stochastic demand and estimate guaranteed QoS assurance level by efficiently satisfying the requirement of ubiquitous application. We have also performed extensive experiments over distributed parallel computing cluster (DPCC) and studied overall job execution performance of different GSPs under a wide range of QoS parameters using different strategies. Our model and performance evaluation results can serve as a valuable reference for designing appropriate strategies in a practical grid environment.
Nirmalya Roy, Sajal K. Das 0001
IEEE Trans. Parallel Distributed Syst.1
2007 Mobility-Aware Efficient Job Scheduling in Mobile Grids
abstract
In this paper, we present a node mobility prediction framework based on a generic mobile grid architecture. We show how this framework can be used to formulate a cost effective job scheduling scheme based on a predetermined pricing strategy at the wireless access point. The proposed scheme is for distributing grid computing jobs to the mobile nodes and considers the bandwidth constraints along with any internal job (e.g., call processing) arrival rate at the nodes. The simulation results point to the efficacy of our algorithm.
Preetam Ghosh, Nirmalya Roy, Sajal K. Das 0001
CCGRID2
2007 A Middleware Framework for Ambiguous Context Mediation in Smart Healthcare Application
Nirmalya Roy, Gautham V. Pallapa, Sajal K. Das 0001
WiMob1
2006 Context-Aware Resource Management in Multi-Inhabitant Smart Homes: A Nash H-Learning based Approach
abstract
A smart home aims at building intelligence automation with a goal to provide its inhabitants with maximum possible comfort, minimize the resource consumption and thus overall cost of maintaining the home. 'Context awareness' is perhaps the most salient feature of such an intelligent environment. Clearly, an inhabitant's mobility and activities play a significant role in defining his contexts in and around the home. Although there exists an optimal algorithm for location and activity tracking of a single inhabitant, the correlation and dependence between multiple inhabitants' contexts within the same environment make the location and activity tracking more challenging. In this paper, we first prove that the optimal location prediction across multiple inhabitants in smart homes is an NP-hard problem. Next, to capture the correlation and interactions of different inhabitants' movements (and hence activities), we develop a novel framework based on a game theoretic, Nash H-learning approach that attempts to minimize the joint location uncertainty. The framework achieves a Nash equilibrium such that no inhabitant is given preference over others. This results in more accurate prediction of contexts and better adaptive control of automated devices, leading to a mobility-aware resource (say, energy) management scheme in multi-inhabitant smart homes. Experimental results demonstrate that the proposed framework is capable of adaptively controlling a smart environment, thus reducing energy consumption and enhancing the comfort of the inhabitants.
Nirmalya Roy, Abhishek Roy 0001, Sajal K. Das 0001
PerCom1
2006 Context-aware resource management in multi-inhabitant smart homes: A framework based on Nash H
Sajal K. Das 0001, Nirmalya Roy, Abhishek Roy 0001
Pervasive Mob. Comput.2
2005 A Cooperative Learning Framework for Mobility-Aware Resource Management in Multi-Inhabitant Smart Homes
abstract
The essence of pervasive (ubiquitous) computing lies in the creation of smart environments saturated with computing and communication capabilities, yet gracefully integrated with human users. 'Context Awareness' is perhaps the most important feature of such an intelligent computing paradigm. The mobility and activity of the inhabitants play significant roles in forming the context at any instance of time. In order to extract the best performance and efficacy of smart computing environments, one needs a technology-independent, context-aware platform spanning over multiple inhabitants. In this paper, we have developed a framework for mobility-aware resource (in particular, energy consumption) management in a multi-inhabitant smart home, based on a dynamic, cooperative reinforcement learning technique. The inhabitants' mobility creates uncertainty of his location and activity. Using the proposed cooperative game-theory based framework, all the inhabitants currently present in the house attempt to minimize this overall uncertainty in the form of utility functions associated with them. Joint optimization of the utility function corresponds to the convergence to Nash equilibrium and helps in accurate prediction of inhabitants' future locations and activities. This results in adaptive control of automated devices and temperature of the house, thus providing an amicable environment and sufficient comfort to the inhabitants. Simulation results point out that our framework can adaptively control the smart environment, while reducing the energy consumption and enhancing the comfort.
Nirmalya Roy, Abhishek Roy 0001, Sajal K. Das 0001, Kalyan Basu
MobiQuitous1
2005 A pricing strategy for job allocation in mobile grids using a non-cooperative bargaining theory framework
Preetam Ghosh, Nirmalya Roy, Sajal K. Das 0001, Kalyan Basu
J. Parallel Distributed Comput.2
2004 A Game Theory Based Pricing Strategy for Job Allocation in Mobile Grids
abstract
Summary form only given. This article realizes the vision of mobile grid computing by proposing a fair pricing strategy and an optimal, static job allocation scheme. Mobile devices has not yet been integrated into grid computing platforms mainly due to their inherent limitations in processing and storage capacity, power and bandwidth shortages. However, millions of laptops, PDAs and other mobile devices remain unused most of the time and this huge resource repository can be potentially utilized in the grid environment. Here, we propose a game theoretic pricing model, to address load balancing issues in mobile grids. In particular, by drawing upon the Nash bargaining solution (NBS), we show that we can obtain an unified framework for addressing such issues as network efficiency, fairness, utility maximization, and pricing. The advantage of this framework is that we have a precise mathematical characterization of the solutions and their properties. Our current endeavor characterizes a two-player alternating-offer bargaining game between the wireless access point (WAP) server and the mobile devices to determine the pricing strategy. This pricing strategy is then made use of to effectively distribute jobs to the mobile devices. Our job allocation scheme maximizes the revenue of the grid user, and yet is comparable to the overall system response time of other load balancing schemes.
Preetam Ghosh, Nirmalya Roy, Sajal K. Das 0001, Kalyan Basu
IPDPS2