Dong Ma 0001

dblp:68/7912-1 · DBLP profile ↗
← Back
35ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0003-3824-234XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 15 · 6 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 12 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BAIT: Visual-illusion-inspired Privacy Preservation for Mobile Data Visualization
abstract
With the prevalence of mobile data visualizations, there have been growing concerns about their privacy risks, especially shoulder surfing attacks. Inspired by prior research on visual illusion, we propose BAIT, a novel approach to automatically generate privacy-preserving visualizations by stacking a decoy visualization over a given visualization. It allows visualization owners at proximity to clearly discern the original visualization and makes shoulder surfers at a distance be misled by the decoy visualization, by adjusting different visual channels of a decoy visualization (e.g., shape, position, tilt, size, color and spatial frequency). We explicitly model human perception effect at different viewing distances to optimize the decoy visualization design. Privacy-preserving examples and two in-depth user studies demonstrate the effectiveness of BAIT in both controlled lab study and real-world scenarios.
Sizhe Cheng, Songheng Zhang, Dong Ma 0001, Yong Wang 0021
CHI3
2026 DeWater: Towards Efficient Underwater Communication via Fine-tuned Learning-enhanced Demodulation
Yuezhong Wu, Xing Chen 0002, Dong Ma 0001
INFOCOM6
2026 From Cheap to Chic: Enhancing Music Playback Quality of Budget Earphones via Hardware-Aware Learning
abstract
Low-end earphones are widely spread due to their affordability, but their limited speaker hardware often leads to poor music playback quality. This raises a key question: Can we compensate for hardware limitations to enhance the listening experience without modifying the device? Existing EQ-based approaches attempt this, but they rely on frequency response curves (FRCs) measured under ideal conditions, which fail to capture real-world distortions such as harmonic and intermodulation effects.
Changshuo Hu, Hung Manh Pham, Ting Dang, Jiannan Li, Rajesh Krishna Balan, Dong Ma 0001
SenSys6
2026 Dronaquatics: Real-time Swimming Analytics Using Drone Captured Imagery
abstract
Accurate swimming performance monitoring has traditionally relied on wearable sensors, which can disrupt natural technique and are impractical in competitive settings. In this paper, we present a fully vision-based system for automatic swimmer analysis using overhead drone footage, removing the need for any wearable device or underwater equipment. By fine-tuning pose estimation models for aerial aquatic conditions, our approach robustly extracts full-body swimmer skeletons even under challenging scenarios such as splashes and partial occlusions. From these poses, we classify swimming strokes, compute instantaneous speed, estimate lap times, and count individual strokes. Unlike existing methods, our system provides scalable, unobtrusive, and infrastructure-free tracking. Evaluated on real-world drone-captured swimming competition data, our method achieves a median speed estimation error below 4% (under 0.05 m/s), a median lap time error of just 0.03s, and stroke count errors typically under one stroke per lap.
Thu Tran, Harold Abraham Joseph, Kichang Lee, Kenny T. W. Choo, Dong Ma 0001, Shaohui Foong, Thivya Kandappu, JeongGil Ko, Rajesh Krishna Balan
WACV5
2026 A cascade framework for on-device uncertainty-aware event detection on microcontrollers
abstract
Pervasive sensing enables diverse wearable event detection (WED) applications, but deploying machine learning models on resource-constrained microcontrollers (MCUs) poses significant challenges, particularly in ensuring prediction reliability under data shifts or out-of-distribution (OOD) inputs. While Uncertainty quantification methods offer a way to assess this reliability, many are computationally prohibitive for MCUs, and detecting multiple events concurrently further exacerbates resource constraints. Addressing these combined challenges, this paper presents an uncertainty and resource-aware framework designed for reliable and efficient multi-event WED on MCUs, significantly extending our preliminary work. The proposed framework achieves this by integrating Evidential Deep Learning (EDL) for efficient, single-pass uncertainty estimation with a novel cascade learning architecture. This architecture promotes resource efficiency via: (i) intra-event sharing using uncertainty-aware early exits within a staged model (shallow, medium, deep), allowing simpler samples to terminate inference earlier; and (ii) inter-event sharing using a multi-head design where multiple event detectors share a common backbone, minimizing overhead. System efficiency is further enhanced through MCU-specific optimizations, including targeted architecture search, quantization, efficient uncertainty operator implementation using standard TensorFlow Lite Micro (TFLM) operations, and library footprint reduction. We conducted extensive experiments on four distinct wearable datasets (Oesense, KWS, ECG5000, and HHAR) and two MCU platforms (STM32F446ZE, STM32H747XI), comparing the proposed framework against strong baselines including Deep Ensembles and Vanilla EDL. Results demonstrate the proposed framework’s effectiveness, achieving competitive accuracy and uncertainty performance (e.g., up to 22% lower NLL than data augmentation) while drastically reducing resource consumption, offering up to 8.64 × faster inference, up to 8.57 × lower energy use, and 55% smaller memory footprint compared to ensemble methods. The proposed framework enables the deployment of reliable, uncertainty-aware multi-event detection on a wider range of low-power MCUs.
Hong Jia, Young D. Kwon, Dong Ma 0001, Nhat Pham, Lorena Qendro, Tam Vu 0001, Cecilia Mascolo
Pervasive Mob. Comput.3
2025 SmarTeeth: Augmenting Manual Toothbrushing with In-ear Microphones
Qiang Yang 0018, Yang Liu 0101, Jake Stuchbury-Wass, Kayla-Jade Butkow, Emeli Panariti, Dong Ma 0001, Cecilia Mascolo
CHI6
2025 Q-HEART: ECG Question Answering via Knowledge-Informed Multimodal LLMs
abstract
Electrocardiography (ECG) offers critical cardiovascular insights, such as identifying arrhythmias and myocardial ischemia, but enabling automated systems to answer complex clinical questions directly from ECG signals (ECG-QA) remains a significant challenge. Current approaches often lack robust multimodal reasoning capabilities or rely on generic architectures ill-suited for the nuances of physiological signals. We introduce Q-HEART, a novel multimodal framework designed to bridge this gap. Q-HEART leverages a powerful, adapted ECG encoder and integrates its representations with textual information via a specialized ECG-aware transformer-based mapping layer. Furthermore, Q-HEART leverages dynamic prompting and retrieval of relevant historical clinical reports to guide tuning the language model toward knowledge-aware ECG reasoning. Extensive evaluations on the benchmark ECG-QA dataset show Q-HEART achieves state-of-the-art performance, outperforming existing methods by a 4% improvement in exact match accuracy. Our work demonstrates the effectiveness of combining domain-specific architectural adaptations with knowledge-augmented LLM instruction tuning for complex physiological ECG analysis, paving the way for more capable and potentially interpretable clinical patient care systems. Our sample code and checkpoint are made available at https://github.com/manhph2211/Q-HEART.
Hung Manh Pham, Jialu Tang, Aaqib Saeed, Dong Ma 0001
ECAI4
2025 Boosting Masked ECG-Text Auto-Encoders as Discriminative Learners
abstract
The accurate interpretation of Electrocardiogram (ECG) signals is pivotal for diagnosing cardiovascular diseases. Integrating ECG signals with accompanying textual reports further holds immense potential to enhance clinical diagnostics by combining physiological data and qualitative insights. However, this integration faces significant challenges due to inherent modality disparities and the scarcity of labeled data for robust cross-modal learning. To address these obstacles, we propose D-BETA, a novel framework that pre-trains ECG and text data using a contrastive masked auto-encoder architecture, uniquely combining generative and boosted discriminative capabilities for robust cross-modal representations. This is accomplished through masked modality modeling, specialized loss functions, and an improved negative sampling strategy tailored for cross-modal alignment. Extensive experiments on five public datasets across diverse downstream tasks demonstrate that D-BETA significantly outperforms existing methods, achieving an average AUC improvement of 15% in linear probing with only one percent of training data and 2% in zero-shot performance without requiring training data over state-of-the-art models. These results highlight the effectiveness of D-BETA, underscoring its potential to advance automated clinical diagnostics through multi-modal representations.
Hung Manh Pham, Aaqib Saeed, Dong Ma 0001
ICML3
2025 ESPIRO: Natural Pulmonary Function Monitoring via Earphone-Acquired Speech
abstract
As a crucial tool for assessing health, spirometry provides valuable insights into pulmonary functions. Recent advancements have enabled more convenient measurements by shifting spirometry solutions from cumbersome clinical devices to portable devices. However, the forced maneuvers and burdensome procedures, which necessitate repeated maximal forced breathing, often lead to dizziness and discomfort, rendering them unsuitable for vulnerable populations. In this paper, we present ESPIRO (Earphone-enabled Speech sPIROmetry) system to furnish user-friendly pulmonary function monitoring for diverse populations. Basically, ESPIRO records normal speech using microphone-embedded earphones and characterizes pulmonary function-related glottal flow during speech production. ESPIRO advances existing spirometry solutions in i) leveraging phonetics to associate pulmonary function with glottal flow in normal speech, thereby eliminating the need for forced breathing; ii) identifying effective speech features according to physiological basis, ensuring reliable spirometry measurements; and iii) effectively addressing ambient noise, making it suitable for various real-world settings. Extensive experiments with 38 subjects on 18 commodity earphones confirm that ESPIRO accurately estimates pulmonary function indices in practice.
Yetong Cao, Dong Ma 0001, Wentao Xie 0001, Qian Zhang 0001, Jun Luo 0001
MobiCom2
2025 RespEar: Earable-Based Robust Respiratory Rate Monitoring
abstract
Continuous respiratory rate (RR) monitoring is essential for understanding physical and mental health, as well as tracking fitness. However, performing reliable and non-obtrusive RR monitoring across diverse daily routines and activities is still an open research problem. In this work, we present RespEar, a pipeline for robust RR monitoring across various sedentary and active scenarios using earphones. RespEar relies solely on in-ear microphones, repurposing them for continuous RR monitoring purposes. Specifically, leveraging the unique properties of in-ear audio, RespEar enables the use of respiratory sinus arrhythmia (RSA) and locomotor respiratory coupling (LRC), physiological couplings between cardiovascular activity, gait and respiration, to determine the RR. This effectively addresses the challenges posed by the almost imperceptible breathing signals encountered during common daily activities. Additionally, RespEar uniquely identifies and addresses three key practical issues for the RSA and LRC-based solutions and introduces a suite of meticulously crafted signal processing techniques to enhance the accuracy of RR measurements. With data collected from 18 subjects over 8 activities, RespEar measures RR with a mean absolute error (MAE) of 1.48 breaths per minute (BPM) and a mean absolute percent error (MAPE) of 9.12% in sedentary conditions, and a MAE of 2.28 BPM and a MAPE of 11.04% in active conditions, respectively. To the best of our knowledge, RespEar is the first earable-based system capable of accurately determining RR in a variety of realistic settings.
Yang Liu 0101, Kayla-Jade Butkow, Jake Stuchbury-Wass, Adam Pullin, Dong Ma 0001, Cecilia Mascolo
PerCom5
2025 WalkEar: Holistic Gait Monitoring using Earables
abstract
Gait behaviour is a key health metric. Temporal, spatial and kinetic walking gait parameters are valuable in enhancing sport performance and early health diagnostics Full gait assessment requires a gait clinic and existing wearable gait tracking systems typically measure isolated subsets of parameters tailored to specific applications. This is useful when the condition to be monitored is known, but fails to offer a comprehensive view of an individual’s gait traits when their pathology is unknown or changing, or a general assessment is required. To support holistic walking gait tracking, we introduce WalkEar, a novel sensing platform designed to simultaneously track gait parameters using commodity earbuds. WalkEar operates by detecting gait events to derive temporal gait parameters and segment the IMU data. WalkEar then progresses earable gait assessment by, for the first time, estimating kinetic gait parameters and reconstructing the vGRF curve using machine learning. Each parameter is calculated on a step-to-step basis for gait variability and asymmetry. We developed an earbud prototype and collected data from 13 participants using gold standard force plates and instrumented treadmill ground truth. Extensive experiments demonstrate the promising performance of WalkEar, achieving an overall MAPE of 5.1% in estimating gait, 2.0% MAPE on kinetic gait parameters, and an NRMSE of 5.3% for vGRF curve reconstruction.
Jake Stuchbury-Wass, Yang Liu 0101, Kayla-Jade Butkow, Joshua Carter, Qiang Yang 0018, Mathias Ciliberto, Ezio Preatoni, Dong Ma 0001, Cecilia Mascolo
PerCom8
2025 Toward Diverse Tiny-Model Selection for Microcontrollers
abstract
Enabling efficient and accurate deep neural network (DNN) inference on microcontrollers is challenging due to their constrained on-chip resources. Existing approaches mainly focus on compressing larger models, often compromising model accuracy as a trade-off. In this paper, we rethink the problem from the inverse perspective by directly constructing small/weak models, then enhancing their accuracy. Thus, we propose DiTMoS, a novel DNN training and inference framework featuring aselector-classifiersarchitecture, where the selector routes each input sample to the appropriate classifier for classification. DiTMoS is built on a key insight: a combination of weak models can exhibit high diversity and the union of them can significantly raise the upper bound of overall accuracy. To approach the upper bound, DiTMoS introduces three strategies including diverse training data splitting to enhance the classifiers' diversity, adversarial selector-classifiers training to ensure synergistic interactions thereby maximizing their complementarity, and heterogeneous feature aggregation to improve the capacity of classifiers. We further design a network slicing technique to eliminate the extra memory consumption incurred by feature aggregation. We deploy DiTMoS on the Nucleo STM32F767ZI board and evaluate its performance across three time-series datasets for human activity recognition, keyword spotting, and emotion recognition tasks. The experimental results show that: (a) DiTMoS improves accuracy by up to 13.4% compared to the best baseline; (b) network slicing successfully eliminates the memory overhead introduced by feature aggregation, with only a minimal increase in latency. The code of DiTMoS is released athttps://github.com/TheMaXiao/DiTMoS
Shengfeng He, Hezhe Qiao, Dong Ma 0001
IEEE Trans. Mob. Comput.4
2024 UR2M: Uncertainty and Resource-Aware Event Detection on Microcontrollers
abstract
Traditional machine learning techniques are prone to generating inaccurate predictions when confronted with shifts in the distribution of data between the training and testing phases. This vulnerability can lead to severe consequences, especially in applications such as mobile healthcare. Uncertainty estimation has the potential to mitigate this issue by assessing the reliability of a model's output. However, existing uncertainty estimation techniques often require substantial computational resources and memory, making them impractical for implementation on microcontrollers (MCUs). This limitation hinders the feasibility of many important on-device wearable event detection (WED) applications, such as heart attack detection. In this paper, we present UR2M, a novel Uncertainty and Resource-aware event detection framework for MCUs. Specifically, we (i) develop an uncertainty-aware WED based on evidential theory for accurate event detection and reliable uncertainty estimation; (ii) introduce a cascade ML framework to achieve efficient model inference via early exits, by sharing shallower model layers among different event models; (iii) optimize the deployment of the model and MCU library for system efficiency. We conducted extensive experiments and compared UR2M to traditional uncertainty baselines using three wearable datasets. Our results demonstrate that UR2M achieves up to 864% faster inference speed, 857% energy-saving for uncertainty estimation, 55% memory saving on two popular MCUs, and a 22% improvement in uncertainty quantification performance. UR2M can be deployed on a wide range of MCUs, significantly expanding real-time and reliable WED applications.
Hong Jia, Young D. Kwon, Dong Ma 0001, Nhat Pham, Lorena Qendro, Tam Vu 0001, Cecilia Mascolo
PerCom3
2024 DiTMoS: Delving into Diverse Tiny-Model Selection on Microcontrollers
abstract
Enabling efficient and accurate deep neural network (DNN) inference on microcontrollers is non-trivial due to the constrained on-chip resources. Current methodologies primarily focus on compressing larger models yet at the expense of model accuracy. In this paper, we rethink the problem from the inverse perspective by constructing small/weak models directly and improving their accuracy. Thus, we introduce DiTMoS, a novel DNN training and inference framework with a selector-classifiers architecture, where the selector routes each input sample to the appropriate classifier for classification. DiTMoS is grounded on a key insight: a composition of weak models can exhibit high diversity and the union of them can significantly boost the accuracy upper bound. To approach the upper bound, DiT-MoS introduces three strategies including diverse training data splitting to increase the classifiers' diversity, adversarial selector-classifiers training to ensure synergistic interactions thereby maximizing their complementarity, and heterogeneous feature aggregation to improve the capacity of classifiers. We further propose a network slicing technique to alleviate the extra memory overhead incurred by feature aggregation. We deploy DiTMoS on the Neucleo STM32F767ZI board and evaluate it based on three time-series datasets for human activity recognition, keywords spotting, and emotion recognition, respectively. The experiment results manifest that: (a) DiTMoS achieves up to 13.4% accuracy improvement compared to the best baseline; (b) network slicing almost completely eliminates the memory overhead incurred by feature aggregation with a marginal increase of latency. Code is released at https//github.com/TheMaXiao/DiTMoS
Shengfeng He, Hezhe Qiao, Dong Ma 0001
PerCom4
2024 VibMilk: Nonintrusive Milk Spoilage Detection via Smartphone Vibration
abstract
Quantifying the chemical process of milk spoilage is challenging due to the need for bulky, expensive equipment that is not user-friendly for milk producers or customers. This lack of a convenient and accurate milk spoilage detection system can cause two significant issues. First, people who consume spoiled milk may experience serious health problems. Secondly, milk manufacturers typically provide a “best before” date to indicate freshness, but this date only shows the highest quality of the milk, not the last day it can be safely consumed, leading to significant milk waste. A practical and efficient solution to this problem is proposed in this paper: a vibration-based milk spoilage detection method called VibMilk that utilizes the ubiquitous vibration motor and Inertial Measurement Unit (IMU) of off-the-shelf smartphones. The method detects spoilage based on the fact that the milk’s physical properties change, inducing different vibration responses at various stages of degradation. Using the InceptionTime deep learning model, VibMilk achieves 98.35% accuracy in detecting milk spoilage across 23 different stages, from fresh (pH = 6.6) to fully spoiled (pH = 4.4).
Yuezhong Wu, Dong Ma 0001, Weitao Xu, Mahbub Hassan, Wen Hu 0001
IEEE Internet Things J.4
2024 An evaluation of heart rate monitoring with in-ear microphones under motion
abstract
With the soaring adoption of in-ear wearables, the research community has started investigating suitable in-ear heart rate detection systems. Heart rate is a key physiological marker of cardiovascular health and physical fitness. Continuous and reliable heart rate monitoring with wearable devices has therefore gained increasing attention in recent years. Existing heart rate detection systems in wearables mainly rely on photoplethysmography (PPG) sensors, however, these are notorious for poor performance in the presence of human motion. In this work, leveraging the occlusion effect that enhances low-frequency bone-conducted sounds in the ear canal, we investigate for the first time in-ear audio-based motion-resilient heart rate monitoring. We first collected heart rate-induced sounds in the ear canal using an in-ear microphone under seven stationary activities and two full-body motion activities (i.e., walking, and running). Then, we devised a novel deep learning based motion artefact (MA) mitigation framework to denoise the in-ear audio signals, followed by a heart rate estimation algorithm to extract heart rate. With data collected from 15 subjects over nine activities, we demonstrate that hEARt, our end-to-end approach, achieves a mean absolute error (MAE) of 1.88 ± 2.89 BPM, 6.83 ± 5.05 BPM, and 13.19 ± 11.37 BPM for stationary, walking, and running, respectively, opening the door to a new non-invasive and affordable heart rate monitoring with usable performance for daily activities. Not only does hEARt outperform previous in-ear heart rate monitoring work, but it outperforms reported in-ear PPG performance.
Kayla-Jade Butkow, Ting Dang, Andrea Ferlini, Dong Ma 0001, Yang Liu 0101, Cecilia Mascolo
Pervasive Mob. Comput.4
2023 hEARt: Motion-resilient Heart Rate Monitoring with In-ear Microphones
abstract
With the soaring adoption of in-ear wearables, the research community has started investigating suitable in-ear heart rate (HR) detection systems. HR is a key physiological marker of cardiovascular health and physical fitness. Continuous and reliable HR monitoring with wearable devices has therefore gained increasing attention in recent years. Existing HR detection systems in wearables mainly rely on photoplethysmography (PPG) sensors, however, these are notorious for poor performance in the presence of human motion. In this work, leveraging the occlusion effect that enhances low-frequency bone-conducted sounds in the ear canal, we investigate for the first time in-ear audio-based motion-resilient HR monitoring. We first collected HR-induced sounds in the ear canal leveraging an in-ear microphone under stationary and three different activities (i.e., walking, running, and speaking). Then, we devised a novel deep learning based motion artefact (MA) mitigation framework to denoise the in-ear audio signals, followed by an HR estimation algorithm to extract HR. With data collected from 20 subjects over four activities, we demonstrate that hEARt, our end-to-end approach, achieves a mean absolute error (MAE) of 3.02$\pm\ \boldsymbol{ 2.97}$BPM, 8.12$\pm\ \boldsymbol{6.74}$BPM, 11.23$\pm\ \boldsymbol{9.20}$BPM and 9.39$\pm\ \boldsymbol{6.97}$BPM for stationary, walking, running and speaking, respectively, opening the door to a new non-invasive and affordable HR monitoring with usable performance for daily activities. Not only does hEARt outperform previous in-ear HR monitoring work, but it outperforms reported in-ear PPG performance.
Kayla-Jade Butkow, Ting Dang, Andrea Ferlini, Dong Ma 0001, Cecilia Mascolo
PERCOM4
2023 Don't Peek at My Chart: Privacy-preserving Visualization for Mobile Devices
abstract
Abstract Data visualizations have been widely used on mobile devices like smartphones for various tasks (e.g., visualizing personal health and financial data), making it convenient for people to view such data anytime and anywhere. However, others nearby can also easily peek at the visualizations, resulting in personal data disclosure. In this paper, we propose a perception‐driven approach to transform mobile data visualizations into privacy‐preserving ones. Specifically, based on human visual perception, we develop a masking scheme to adjust the spatial frequency and luminance contrast of colored visualizations. The resulting visualization retains its original information in close proximity but reduces visibility when viewed from a certain distance or farther away. We conducted two user studies to inform the design of our approach (N=16) and systematically evaluate its performance (N=18), respectively. The results demonstrate the effectiveness of our approach in terms of privacy preservation for mobile data visualizations.
Songheng Zhang, Dong Ma 0001, Yong Wang 0021
Comput. Graph. Forum2
2023 Recognizing Hand Gestures Using Solar Cells
abstract
We design a system, SolarGest, which can recognize hand gestures near a solar-powered device by analyzing the patterns of the photocurrent. SolarGest is based on the observation that each gesture interferes with incident light rays on the solar panel in a unique way, leaving its discernible signature in harvested photocurrent. Using solar energy harvesting laws, we develop a model to optimize design and usage of SolarGest. To further improve the robustness of SolarGest under non-deterministic operating conditions, we combine dynamic time warping with Z-score transformation in a signal processing pipeline to pre-process each gesture waveform before it is analyzed for classification. We evaluate SolarGest with both conventional opaque solar cells as well as emerging see-through transparent cells. Our experiments demonstrate that SolarGest achieves 99% for six gestures with a single cell and 95% for fifteen gesture with a$2\times 2$solar cell array. The power measuement study suggests that SolarGest consume 44% less power compared to light sensor based systems.
Dong Ma 0001, Guohao Lan, Changshuo Hu, Mahbub Hassan, Wen Hu 0001, Mushfika Baishakhi Upama, Ashraf Uddin 0002, Moustafa Youssef 0001
IEEE Trans. Mob. Comput.1
2022 Photovoltaic cells for energy harvesting and indoor positioning
abstract
We propose SoLoc, a lightweight probabilistic fingerprinting-based technique for energy-free device-free indoor localization. The system harnesses photovoltaic currents harvested by the photovoltaic cells in smart environments for simultaneously powering digital devices and user positioning. The basic principle is that the location of the human interferes with the lighting received by the photovoltaic cells, thus producing a location fingerprint on the generated photocurrents. To ensure resilience to noisy measurements, SoLoc constructs probability distributions as a photovoltaic fingerprint at each location. Then, we employ a probabilistic graphical model for estimating the user location in the continuous space. Results show that SoLoc can localize the user at sub-meter accuracy in a real indoor environment.
Hamada Rizk, Dong Ma 0001, Mahbub Hassan, Moustafa Youssef 0001
SIGSPATIAL/GIS2
2022 Improving Feature Generalizability with Multitask Learning in Class Incremental Learning
abstract
Many deep learning applications, like keyword spotting [1], [2], require the incorporation of new concepts (classes) over time, referred to as Class Incremental Learning (CIL). The major challenge in CIL is catastrophic forgetting, i.e., preserving as much of the old knowledge as possible while learning new tasks. Various techniques, such as regularization, knowledge distillation, and the use of exemplars, have been proposed to resolve this issue. However, prior works primarily focus on the incremental learning step, while ignoring the optimization during the base model training. We hypothesise that a more transferable and generalizable feature representation from the base model would be beneficial to incremental learning.In this work, we adopt multitask learning during base model training to improve the feature generalizability. Specifically, instead of training a single model with all the base classes, we decompose the base classes into multiple subsets and regard each of them as a task. These tasks are trained concurrently and a shared feature extractor is obtained for incremental learning. We evaluate our approach on two datasets under various configurations. The results show that our approach enhances the average incremental learning accuracy by up to 5.5%, which enables more reliable and accurate keyword spotting over time. Moreover, the proposed approach can be combined with many existing techniques and provides additional performance gain.
Dong Ma 0001, Chi Ian Tang, Cecilia Mascolo
ICASSP1
2022 PROS: an efficient pattern-driven compressive sensing framework for low-power biopotential-based wearables with on-chip intelligence
abstract
While the global healthcare market of wearable devices has been growing significantly in recent years and is predicted to reach $60 billion by 2028, many important healthcare applications such as seizure monitoring, drowsiness detection, etc. have not been deployed due to the limited battery lifetime, slow response rate, and inadequate biosignal quality.
Nhat Pham, Hong Jia, Tuan Dinh, Nam Bui, Young D. Kwon, Dong Ma 0001, Phuc Nguyen 0002, Cecilia Mascolo, Tam Vu 0001
MobiCom7
2022 Simultaneous Energy Harvesting and Gait Recognition Using Piezoelectric Energy Harvester
abstract
Piezoelectric energy harvester (PEH), which generates electricity from stress or vibrations, is attracting tremendous attention as a viable solution to extend battery life of wearable devices. More interestingly, besides the energy harvesting capability, recent research has demonstrated the feasibility of leveraging PEH as an power-free sensor for gait recognition as its stress or vibration patters are significantly influenced by the gait. However, as PEHs are not designed for precise motion sensing, the gait recognition accuracy remains low with conventional classification algorithms. The accuracy deteriorates further when the generated electricity is stored simultaneously. In this work, to achieve high performance gait recognition and efficient energy harvesting at the same time, we make two distinct contributions. First, we propose a preprocessing algorithm to filter out the effect of energy storage on PEH electricity signals. Second, we propose long short-term memory (LSTM) network-based classifiers to accurately capture temporal information in gait-induced electricity generation. We prototype the proposed gait recognition architecture in the form factor of an insole and evaluate its gait recognition as well as energy harvesting performance with 20 subjects. Our results show that the proposed architecture detects human gait with 12 percent higher recall and harvests up to 127 percent more energy while consuming 38 percent less power compared to the state-of-the-art.
Dong Ma 0001, Guohao Lan, Weitao Xu, Mahbub Hassan, Wen Hu 0001
IEEE Trans. Mob. Comput.1
2021 EarGate: gait-based user identification with in-ear microphones
abstract
Human gait is a widely used biometric trait for user identification and recognition. Given the wide-spreading, steady diffusion of ear-worn wearables (Earables) as the new frontier of wearable devices, we investigate the feasibility of earable-based gait identification. Specifically, we look at gait-based identification from the sounds induced by walking and propagated through the musculoskeletal system in the body. Our system, EarGate, leverages an in-ear facing microphone which exploits the earable's occlusion effect to reliably detect the user's gait from inside the ear canal, without impairing the general usage of earphones. With data collected from 31 subjects, we show that EarGate achieves up to 97.26% Balanced Accuracy (BAC) with very low False Acceptance Rate (FAR) and False Rejection Rate (FRR) of 3.23% and 2.25%, respectively. Further, our measurement of power consumption and latency investigates how this gait identification model could live both as a stand-alone or cloud-coupled earable system.
Andrea Ferlini, Dong Ma 0001, Robert K. Harle, Cecilia Mascolo
MobiCom2
2021 OESense: employing occlusion effect for in-ear human sensing
abstract
Smart earbuds are recognized as a new wearable platform for personal-scale human motion sensing. However, due to the interference from head movement or background noise, commonly-used modalities (e.g. accelerometer and microphone) fail to reliably detect both intense and light motions. To obviate this, we propose OESense, an acoustic-based in-ear system for general human motion sensing. The core idea behind OESense is the joint use of the occlusion effect (i.e., the enhancement of low-frequency components of bone-conducted sounds in an occluded ear canal) and inward-facing microphone, which naturally boosts the sensing signal and suppresses external interference. We prototype OESense as an earbud and evaluate its performance on three representative applications, i.e., step counting, activity recognition, and hand-to-face gesture interaction. With data collected from 31 subjects, we show that OESense achieves 99.3% step counting recall, 98.3% recognition recall for 5 activities, and 97.0% recall for five tapping gestures on human face, respectively. We also demonstrate that OESense is compatible with earbuds' fundamental functionalities (e.g. music playback and phone calls). In terms of energy, OESense consumes 746 mW during data recording and recognition and it has a response latency of 40.85 ms for gesture recognition. Our analysis indicates such overhead is acceptable and OESense is potential to be integrated into future earbuds.
Dong Ma 0001, Andrea Ferlini, Cecilia Mascolo
MobiSys1
2020 Skin-MIMO: Vibration-based MIMO Communication over Human Skin
abstract
We explore the feasibility of Multiple-Input-Multiple-Output (MIMO) communication through vibrations over human skin. Using off-the-shelf motors and piezo transducers as vibration transmitters and receivers, respectively, we build a 2x2 MIMO testbed to collect and analyze vibration signals from real subjects. Our analysis reveals that there exist multiple independent vibration channels between a pair of transmitter and receiver, confirming the feasibility of MIMO. Unfortunately, the slow ramping of mechanical motors and rapidly changing skin channels make it impractical for conventional channel sounding based channel state information (CSI) acquisition, which is critical for achieving MIMO capacity gains. To solve this problem, we propose Skin-MIMO, a deep learning based CSI acquisition technique to accurately predict CSI entirely based on inertial sensor (accelerometer and gyroscope) measurements at the transmitter, thus obviating the need for channel sounding. Based on experimental vibration data, we show that Skin-MIMO can improve MIMO capacity by a factor of 2.3 compared to Single-Input-Single-Output (SISO) or open-loop MIMO, which do not have access to CSI. A surprising finding is that gyroscope, which measures the angular velocity, is found to be superior in predicting skin vibrations than accelerometer, which measures linear acceleration and used widely in previous research for vibration communications over solid objects.
Dong Ma 0001, Yuezhong Wu, Ming Ding 0001, Mahbub Hassan, Wen Hu 0001
INFOCOM1
2020 Poster Abstract: Data Communication using Switchable Privacy Glass
abstract
Switchable privacy glass can electronically change its state between opaque and transparent. In this work, we propose to exploit the electronic configurability of switchable glass to modulate natural light, which can be demodulated by a nearby receiver with light sensing capability to realise data communication over natural light. A key advantage is that no energy is used to generate light, as it simply modulates the existing light in the nature. We demonstrate that the proposed data communication using switchable glass modulation can achieve 33.33 bits per second communication with a bit rate below 1% under a wide range of ambient luminance.
Changshuo Hu, Dong Ma 0001, Mahbub Hassan, Wen Hu 0001
IPSN2
2020 SolarSLAM: Battery-free Loop Closure for Indoor Localisation
abstract
In this paper, we propose SolarSLAM, a batteryfree loop closure method for indoor localisation. Inertial Measurement Unit (IMU) based indoor localisation method has been widely used due to its ubiquity in mobile devices, such as mobile phones, smartwatches and wearable bands. However, it suffers from the unavoidable long term drift. To mitigate the localisation error, many loop closure solutions have been proposed using sophisticated sensors, such as cameras, laser, etc. Despite achieving high-precision localisation performance, these sensors consume a huge amount of energy. Different from those solutions, the proposed SolarSLAM takes advantage of an energy harvesting solar cell as a sensor and achieves effective battery-free loop closure method. The proposed method suggests the key-point dynamic time warping for detecting loops and uses robust simultaneous localisation and mapping (SLAM) as the optimiser to remove falsely recognised loop closures. Extensive evaluations in the real environments have been conducted to demonstrate the advantageous photocurrent characteristics for indoor localisation and good localisation accuracy of the proposed method.
Bo Wei 0003, Weitao Xu, Chengwen Luo 0001, Guillaume Zoppi, Dong Ma 0001, Sen Wang 0002
IROS5
2020 Enhancing Cellular Communications for UAVs via Intelligent Reflective Surface
abstract
Intelligent reflective surfaces (IRSs) capable of reconfiguring their electromagnetic absorption and reflection properties in real-time are offering unprecedented opportunities to enhance wireless communication experience in challenging environments. In this paper, we analyze the potential of IRS in enhancing cellular communications for UAVs, which currently suffers from poor signal strength due to the down-tilt of base station antennas optimized to serve ground users. We consider deployment of IRS on building walls, which can be remotely configured by cellular base stations to coherently direct the reflected radio waves towards specific UAVs in order to increase their received signal strengths. Using the recently released 3GPP ground-to-air channel models, we analyze the signal gains at UAVs due to the IRS deployments as a function of UAV height as well as various IRS parameters including size, altitude, and distance from base station. Our analysis suggests that even with a small IRS, we can achieve significant signal gain for UAVs flying above the cellular base station. We also find that the maximum gain can be achieved by optimizing the location of IRS including its altitude and distance to BS.
Dong Ma 0001, Ming Ding 0001, Mahbub Hassan
WCNC1
2020 Capacitor-based Activity Sensing for Kinetic-powered Wearable IoTs
abstract
We propose the use of the conventional energy storage component, i.e., capacitor, in the kinetic-powered wearable IoTs as the sensor to detect human activities. Since activities accumulate energy in the capacitor at different rates, the charging rate of the capacitor can be used to detect the activities. The key advantage of the proposed capacitor-based activity sensing mechanism, called CapSense, is that it obviates the need for sampling the motion signal at a high rate, and thus, significantly reduces power consumption of the wearable device. The challenge we face is that capacitors are inherently non-linear energy accumulators, which leads to significant variations in the charging rates. We solve this problem by jointly configuring the parameters of the capacitor and the associated energy harvesting circuits, which allows us to operate in the charging cycles that are approximately linear. We design and implement a kinetic-powered shoe and conduct experiments with 10 subjects. Our results show that CapSense can classify five different daily activities with 95% accuracy while consuming 57% less system power compared to conventional motion-sensor-based approaches.
Guohao Lan, Dong Ma 0001, Weitao Xu, Mahbub Hassan, Wen Hu 0001
ACM Trans. Internet Things2
2020 EnTrans: Leveraging Kinetic Energy Harvesting Signal for Transportation Mode Detection
abstract
Monitoring the daily transportation modes of an individual provides useful information in many application domains, such as urban design, real-time journey recommendation, and providing location-based services. In existing systems, accelerometer and GPS are the dominantly used signal sources for transportation context monitoring which drain out the limited battery life of the wearable devices very quickly. To resolve the high energy consumption issue, in this paper, we present EnTrans, which enables transportation mode detection by using only the kinetic energy harvester as an energy-efficient signal source. The proposed idea is based on the intuition that the vibrations experienced by the passenger during traveling with different transportation modes are distinctive. Thus, voltage signal generated by the energy harvesting devices should contain sufficient features to distinguish different transportation modes. We evaluate our system using over 28 h of data, which is collected by eight individuals using a practical energy harvesting prototype. The evaluation results demonstrate that EnTrans is able to achieve an overall accuracy over 92% in classifying five different modes while saving more than 34% of the system power compared to conventional accelerometer-based approaches.
Guohao Lan, Weitao Xu, Dong Ma 0001, Sara Khalifa, Mahbub Hassan, Wen Hu 0001
IEEE Trans. Intell. Transp. Syst.3
2019 SolarGest: Ubiquitous and Battery-free Gesture Recognition using Solar Cells
abstract
We design a system, SolarGest, which can recognize hand gestures near a solar-powered device by analyzing the patterns of the photocurrent. SolarGest is based on the observation that each gesture interferes with incident light rays on the solar panel in a unique way, leaving its distinguishable signature in harvested photocurrent. Using solar energy harvesting laws, we develop a model to optimize design and usage of SolarGest. To further improve the robustness of SolarGest under non-deterministic operating conditions, we combine dynamic time warping with Z-score transformation in a signal processing pipeline to pre-process each gesture waveform before it is analyzed for classification. We evaluate SolarGest with both conventional opaque solar cells as well as emerging see-through transparent cells. Our experiments with 6,960 gesture samples for 6 different gestures reveal that even with transparent cells, SolarGest can detect 96% of the gestures while consuming 44% less power compared to light sensor based systems.
Dong Ma 0001, Guohao Lan, Mahbub Hassan, Wen Hu 0001, Mushfika Baishakhi Upama, Ashraf Uddin 0002, Moustafa Youssef 0001
MobiCom1
2018 HiddenCode: Hidden Acoustic Signal Capture with Vibration Energy Harvesting
abstract
The feasibility of using vibration energy harvesting (VEH) as an energy-efficient receiver for short-range acoustic data communication has been investigated recently. When data was encoded in acoustic signal within the energy harvesting frequency band and transmitted through a speaker, a VEH receiver was capable of decoding the data by processing the harvested energy signal. Although previous work created new opportunities for simultaneous energy harvesting and communication using the same hardware, the communication makes annoying sounds as the energy harvesting frequency band lies within the sensitive region of human auditory system. In this work, we present a novel modulation scheme to completely hide all communications within background music sound. The proposed modulation exploits sound masking theory to maximize signal to noise ratio of data communication without being audible to the music listener. We capitalize on the existence of repetitive sound patterns within popular music to realize synchronization between the transmitter and the receiver. We implement the proposed modulation within multiple hit songs and demonstrate its efficacy using a real VEH prototype made from off-the-shelf hardware. A user study involving 30 subjects confirms that the proposed modulation can completely hide VEH-based data communication from human perception while achieving up to 14 bps data rate, which is sufficient to transmit short codes or coupons of practical use.
Guohao Lan, Dong Ma 0001, Mahbub Hassan, Wen Hu 0001
PerCom2
2017 CapSense: Capacitor-based Activity Sensing for Kinetic Energy Harvesting Powered Wearable Devices
abstract
We propose a new activity sensing method, CapSense, which detects activities of daily living (ADL) by sampling the voltage of the kinetic energy harvesting (KEH) capacitor at an ultra low sampling rate. Unlike conventional sensors that generate only instantaneous motion information of the subject, KEH capacitors accumulate and store human generated energy over time. Given that humans produce kinetic energy at distinct rates for different ADL, the KEH capacitor can be sampled only once in a while to observe the energy generation rate and identify the current activity. Thus, with CapSense, it is possible to avoid collecting time series motion data at high frequency, which promises significant power saving for the sensing device. We prototype a shoe-mounted KEH-powered wearable device and conduct experiments with 10 subjects for detecting 5 different activities. Our results show that compared to the existing time-series-based activity recognition, CapSense reduces sampling-induced power consumption by 99% and the overall system power, after considering wireless transmissions, by 75%. CapSense recognizes activities with up to 90%.
Guohao Lan, Dong Ma 0001, Weitao Xu, Mahbub Hassan, Wen Hu 0001
MobiQuitous2
2017 Unobtrusive User Verification using Piezoelectric Energy Harvesting
abstract
With the capability to harvest energy from low frequency motions or vibrations, piezoelectric energy harvesting has become a promising solution to achieve self-powered wearable system. Apart from generating energy to power the wearable devices, the output electricity signal of the PEH can also be used as an information source as it reflects the activity or motion patterns of the user. In this paper, we have designed and built an insole-based user authentication system by leveraging the AC voltage generated by the PEH during human walking. Meanwhile, the generated power is also collected and stored, which could be later used as the power source of the mobile system. By using a dataset of 20 subjects, we have demonstrated that our system can achieve 89.76% of human recognition accuracy when using only one gait cycle signal, and the accuracy can be further increased to 95.86% when two gait cycles are utilized.
Dong Ma 0001, Guohao Lan, Weitao Xu, Mahbub Hassan, Wen Hu 0001
MobiQuitous1