Yongpan Zou

dblp:150/3243 · DBLP profile ↗
← Back
41ranked-venue papers
13as first author
23since 2021 · last 2026
0000-0002-4314-6259ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 26 · 9 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 8 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MentalCare: Contrastive Disentanglement and Positive Transfer for Psychiatric Disorder Detection with a Wristband-Smartphone System
abstract
Low-cost wrist-worn sensing for scalable, low-burden assessment of mental-health symptoms is a central topic in mobile health. In this paper, we propose MentalCare, a multi-class psychiatric disorder (PD) detection system that integrates ubiquitous multi-source sensors, including Pulse sensor, LMT70, Grove GSR, and Max30102. These sensors capture physiological signals such as photoplethysmography (PPG), skin temperature (SKT), galvanic skin response (GSR), and blood oxygen (OXY). To address physiological heterogeneity among individuals, we design a contrastive disentanglement with a positive transfer framework (CDPT), which consists of a stacked convolutional neural network module (SCNN), a cross-modal attention interaction module (CMAI), a subject-level contrastive disentanglement module (SCD), and an adaptive feature selector module (AFS). In particular, the SCD module employs supervised contrastive learning to promote the separation of heterogeneous features among different subjects in the latent subspace. Furthermore, the proposed AFS module selects features from the abundant source-domain heterogeneous representations that are most similar to the target domain, thereby facilitating feature alignment and enhancing detection performance. We collect a multi-class PD dataset comprising 74 participants diagnosed with depression, bipolar disorder, anxiety, or verified as healthy controls. The experiment result demonstrates that CDPT achieves 83.96% overall accuracy under the leave-one-subject-out cross-validation protocol, significantly outperforming existing state-of-the-art methods. At the system level, we further validate the effectiveness and rationality of the proposed system when deployed on edge devices.
Yufei Zhang 0005, Wenting Kuang, Zite Huang, Changhe Fan, Yongpan Zou
PerCom5
2026 M-Ring: A Lightweight Ring for Gesture Interaction and User Authentication With Multimodal Fusion
abstract
In extended reality (XR) virtual interaction environments, both flexible user interaction and reliable authentication are essential. However, current wearable systems relying on single-sensor data suffer from insufficient data dimensionality, often failing to simultaneously achieve fine-grained gesture recognition and continuous authentication. Multi-sensor systems, while more capable, tend to be bulky, comfort-compromising, and computationally demanding for edge deployment. Furthermore, inefficient parallel processing and the decoupled pipeline for gesture and authentication tasks degrade computational efficiency, ultimately hindering practical adoption. To overcome these limitations, we propose M-Ring, a lightweight ring-based system equipped with a compact, optimized sensor layout that synchronously captures impedance and IMU signals. By fusing complementary physiological and behavioral cues, M-Ring enables accurate hand-state perception. We design an adaptive-scale convolutional network with a parameter-separated and progressively routed architecture (PLE), which effectively captures intrinsic correlations within multimodal data and enhances synergy between gesture recognition and user authentication, all while minimizing computational overhead. For fine-grained authentication, we propose a dual-factor framework that integrates Siamese networks with contrastive learning to enhance inter-user discriminability. By jointly modeling physiological activations and kinematic patterns, our approach enables both seamless gesture-based interaction and non-intrusive continuous authentication. Experiments demonstrate that M-Ring achieves 97.8% gesture recognition accuracy and 97.5% identity authentication accuracy under complex conditions, while maintaining low enrollment cost, power efficiency, low latency, and strong edge deployment capability.
Yuling Tan, Chaonan Tang, Yongpan Zou, Victor C. M. Leung, Kaishun Wu
IEEE Internet Things J.6
2026 ArmPad: Transforming Forearms Into Interaction Interfaces With Smartwatches
abstract
With the rapid development of new smart devices, such as smart home appliances and VR/AR equipment, there is an increasing demand for novel interaction methods. However, many existing interaction methods require external devices, are unintuitive, and demand substantial user learning effort. To fill this gap, we propose ArmPad, a system that leverages the smartwatch's built-in IMU to enablemultidimensional inputon the user's forearm. Methodologically, ArmPad is explicitly designed to address three core research challenges in forearm-based interaction. First, to resolve the inherent feature conflicts between discrete gesture recognition and continuous distance estimation, we propose a multi-task learning framework with a dynamic gating mechanism for cross-task synergy. Second, to tackle the physical limitation of rapid vibration attenuation across the forearm, we introduce a cross-device guidance strategy that incorporates high-fidelity fingertip knowledge during the training phase. Finally, to ensure robust generalization across diverse populations, we develop task-specific data augmentation and a lightweight user registration mechanism to effectively mitigate physiological variances. Experiments on 20 subjects demonstrate that ArmPad achieves an accuracy of 92.51% on nine gestures and a Mean Absolute Error of 1.53$cm$for sliding distance in cross-user settings. Extensive robustness evaluations and case studies further confirm the system's stability and usability under diverse real-world conditions.
Qiang Yang 0018, Zhidan Liu 0001, Zhenjiang Li 0001, Yongpan Zou, Kaishun Wu
IEEE Trans. Mob. Comput.5
2026 DepGuard: Depression Recognition and Episode Monitoring System With a Ubiquitous Wrist-Worn Device
abstract
Depression significantly impacts mental health, severely disrupting patients' daily lives. During depressive episodes, individuals may experience symptoms such as excessive guilt, self-harm, and suicidal ideation. Compared to proprietary devices like brain electrode caps, wearable technologies for depression detection have gained attention due to their affordability and portability—enabling real-time monitoring of depressive states. However, challenges such as low-quality data from ubiquitous devices, individual variability, and the complexity of multimodal physiological signal analysis limit model generalizability. To address these issues, we present DepGuard, a novel ubiquitous wearable system for depression assessment based on multimodal physiological signals. DepGuard performs a two-stage detection process: depression recognition and real-time episode monitoring. For depression recognition, we propose an unsupervised domain adaptation method to reduce the domain gap between source and target subjects. For episode monitoring, we employ a few-shot learning strategy to enable personalized modeling. Both approaches enhance cross-subject generalization. Our system achieves 90.75% accuracy in cross-subject depression recognition using 30 unlabeled samples per target subject, and 93.52% accuracy in episode monitoring using 15 labeled samples per class.
Yufei Zhang 0005, Wenting Kuang, Yuda Zheng, Qifeng Song, Changhe Fan, Yongpan Zou, Victor C. M. Leung, Kaishun Wu
IEEE Trans. Mob. Comput.7
2026 VitalEar: An Earable Heartbeat and Respiratory Rate Monitoring System Under Aerobic Exercises
abstract
Heart rate (HR) and respiratory rate (RR) are essential physiological indicators of people's physical function and exercise performance. Advancement in sensor technology has rendered earable devices with in-ear microphones feasible for vital sign monitoring. However, it is rather challenging to monitor heart rate and respiration simultaneously with a single earable device especially when a person is doing exercises. This is because intense physical activities can lead to significant noise interference which can easily obscure physiological signals. To address this challenge, this paper presents VitalEar, an exercise physiological monitoring system based on in-ear microphones, designed to estimate HR and RR while addressing complex motion interference and variability in users and activities. VitalEar employs Empirical Wavelet Transform (EWT) to decompose heartbeats into periodic and harmonic coefficients, enhancing noise reduction in the ECG spectrogram reconstruction model. Additionally, VitalEar incorporates a DCN-LSTM-based breathing curve reconstruction model to mitigate background noise and variability in user and activity. The experiments show that VitalEar achieves an average MAE of 5.61 BPM and 2.31 RPM, MAPE of 4.16% and 10.58% for HR and RR estimation, respectively. Compared to related work, our approach offers significant advantages in robustness against intense physical activities
Yuzheng Zhu, Zhangxin Liang, Jie Zheng 0005, Yongpan Zou, Victor C. M. Leung, Kaishun Wu
IEEE Trans. Mob. Comput.4
2025 Generalizable Psychiatric Screening from Wearable: Causal Disentanglement of Multi-Channel Biosignals with Interpretability Analysis
abstract
Psychiatric disorders represent a growing global public health concern. Current screening methods primarily rely on subjective self-reports, which prone to bias and often impede timely intervention. In this paper, we propose G-PsychSW (Generalizable Psychiatric Screening from Wearable), a continuous and objective mental health monitoring based on multi-channel biosignals, including photoplethysmography (PPG), galvanic skin response (GSR), skin temperature (SKT), and blood oxygen (OXY). To address challenges such as individual variability and limited interpretability, we introduce causal disentanglement to learn subject-invariant representations to enable domain generalization (DG). We further incorporate multi-scale attention convolution (MAConv) and functional graph construction (FGC) to model temporal dynamic and inter-channel interaction, uncovering meaningful interpretability analysis corresponding to psychial biomarkers. Experiments on a customized dataset under leave-one-subject-out cross-validation show that G-PsychSW achieves 8 0. 1 6% accuracy and 8 6. 5 2% F1-score across four-class psychiatric screening, including healthy, depression, anxiety and bipolar disorders, outperforming state-of-the-art DG methods and validating its effectiveness.
Wenting Kuang, Yufei Zhang 0005, Kaishun Wu, Changhe Fan, Yongpan Zou
BIBM6
2025 CoMe: Contrastive Learning-based Mental Disorder Recognition with Wearable Multimodal Physiological Data
abstract
Mental disorders are increasingly prevalent, requiring timely and accurate recognition. Existing methods often depend on specialized devices and single-modality public datasets, limiting their ability to identify diverse mental disorders. Multiclass mental disorder recognition (MDR) faces challenges such as limited labeled data, overlapping symptoms among disorders, and significant individual variability. This paper introduces CoMe, a system that leverages low-cost, ubiquitous wearable devices to collect multimodal physiological signals, enabling real-time and daily monitoring to support the early detection of mental disorders. CoMe employs a multimodal residual fusion contrastive learning framework to extract both modality-specific and cross-modality knowledge, addressing challenges of sparse labeled data and symptom overlap. Through self-supervised and supervised fine-tuning, CoMe personalizes models for new users, facilitating accurate, real-time mental status monitoring in daily life. Validation with 74 participants demonstrates CoMe's superior performance under low-data conditions, outperforming both self-supervised and supervised baselines, and surpassing a state-of-the-art method. CoMe achieves average recognition accuracies of 84.5%, 85.36%, 83.25%, and 83.54% for depression, anxiety, bipolar disorders, and control group, respectively.
Yulan Li, Kaishun Wu, Changhe Fan, Yongpan Zou
BIBM5
2025 CHAR: Composite Head-Body Activities Recognition With a Single Earable Device
abstract
The increasing popularity of earable devices stimulates great academic interest to design novel head gesture-based interaction technologies. But existing works simply consider it as a singular activity recognition problem. This is not in line with practice since users may have different body movements such as walking and jogging along with head gestures. It is also beneficial to recognize body movements during human-device interaction since it provides useful context information. As a result, it is significant to recognize such composite activities in which actions of different body parts happen simultaneously. In this paper, we propose a system called CHAR to recognize composite head-body activities with a single IMU sensor. The key idea of our solution is to make use of the inter-correlation of different activities and design a multi-task learning network to extract shared and specific representations. We implement a real-time prototype and conduct extensive experiments to evaluate it. The results show that CHAR can recognize 60 kinds of composite activities (12 head gestures and 5 body movements) with high accuracies of 89.7% and 85.1% in sufficient data and insufficient data cases, respectively.
Peizhao Zhu, Yuzheng Zhu, Yanbo He, Yongpan Zou, Kaishun Wu, Victor C. M. Leung
IEEE Trans. Mob. Comput.5
2025 UltraWrite: A Lightweight Continuous Gesture Input System With Ultrasonic Signals on COTS Devices
abstract
Due to the advantages of device ubiquity, natural interaction and privacy preservation, acoustic-based gesture input has received widespread attention. Researchers have proposed various techniques for different applications. However, the existing work has shortcomings of heavy data-collection overhead, non-continuous input, and performance degradation in crossuser scenarios. To overcome these shortcomings, we propose UltraWrite, an acoustic-based gesture input system that only needs extremely low data-collection overhead, supports continuous input, and achieves high cross-user recognition accuracy. The key idea of our solution is to synthesize training data of continuous gestures from isolated ones, build a lightweight continuous gesture recognition model based on connectionist temporal classification (CTC) mechanism, and design a novel decoupled model training strategy to improve its cross-user recognition capability. We have implemented prototype systems on commercial devices and conducted comprehensive experiments to evaluate their performance. The results show that UltraWrite achieves an average top-1 word accuracy of 99.3% and top-1 word error rate of 0.34%. In addition, we have also evaluated UltraWrite's robustness to the sensing distance, angle, background noise, and device. The results reveal that UltraWrite possesses strong robustness to these factors.
Yongpan Zou, Yunshu Wang, Canlin Zheng, Wenfeng He, Kaishun Wu
IEEE Trans. Mob. Comput.1
2025 ${\sf Img2Acoustic}$Img2Acoustic: A Cross-Modal Gesture Recognition Method Based on Few-Shot Learning
abstract
Acoustic-based human gesture recognition (HGR) offers diverse applications due to the ubiquity of sensors and touch-free interaction. However, existing machine learning approaches require substantial training data, making the process time-consuming, costly, and labor-intensive. Recent studies have explored cross-modal methods to reduce the need for large training datasets in behavior recognition, but they typically rely on open-source datasets that closely align with the target domain, limiting flexibility and complicating data collection. In this paper, we propose${\sf Img2Acoustic}$, a novel cross-modal acoustic-based HGR approach that leverages models trained on open-source image datasets (i.e., EMNIST, Omniglot) to effectively recognize custom gestures detected via acoustic signals. Our model incorporates a task-aware attention layer (TAAL) and a task-aware local matching layer (TALML), enabling seamless transfer of knowledge from image datasets to acoustic gesture recognition. We implement${\sf Img2Acoustic}$on commercial devices and conduct comprehensive evaluations, demonstrating that our method not only delivers superior accuracy and robustness compared to existing approaches but also eliminates the need for extensive training data collection.
Yongpan Zou, Jianhao Weng, Wenting Kuang, Victor C. M. Leung, Kaishun Wu
IEEE Trans. Mob. Comput.1
2025 TimbreSense: Timbre Abnormality Detection for Bel Canto with Smart Devices
abstract
With the rise of mobile devices, bel canto practitioners increasingly utilize smart devices as auxiliary tools for improving their singing skills. However, they frequently encounter timbre abnormalities during practice, which, if left unaddressed, can potentially harm their vocal organs. Existing singing assessment systems primarily focus on pitch and melody and lack real-time detection of bel canto timbre abnormalities. Moreover, the diverse vocal habits and timbre compositions among individuals present significant challenges in cross-user recognition of such abnormalities. To address these limitations, we propose TimbreSense, a novel bel canto timbre abnormality detection system. TimbreSense enables real-time detection of the five major timbre abnormalities commonly observed in bel canto singing. We introduce an effective feature extraction pipeline that captures the acoustic characteristics of bel canto singing. By applying temporal average pooling to the Short-Time Fourier Transform spectrogram, we reduce redundancy while preserving essential frequency-domain information. Our system leverages a transformer model with self-attention mechanisms to extract correlation and semantic features of overtones in the frequency domain. Additionally, we employ a few-shot learning approach involving pre-training, meta-learning, and fine-tuning to enhance the system’s cross-domain recognition performance while minimizing user usage costs. Experimental results demonstrate the system’s strong cross-user domain recognition performance and real-time capabilities.
Yuzheng Zhu, Chengzhe Luo, Yongpan Zou, Dongping Chen, Kaishun Wu
ACM Trans. Sens. Networks3
2024 EmoTracer: A User-independent Wearable Emotion Tracer with Multi-source Physiological Sensors Based on Few-shot Learning
abstract
The rising prevalence of mood disorders, including depression and anxiety, underscores the critical state of physical and mental health issues in contemporary society. Automatic emotion recognition technology emerges as a potential tool for monitoring mental health disorders, and offering valuable guidance through human-computer interfaces. However, existing technologies grapple with limitations, such as emotional concealment, potential privacy leakage, and specific device location restrictions. To overcome these limitations, we propose a mobile wearable emotion recognition system called EmoTracer incorporating ubiquitous multi-source sensors to measure physiology signals, and address challenges such as the intricate signal-emotion relationship, sparse data processing, and notable disparities among different subjects. We implement a real-time prototype and carry out comprehensive experiments to evaluate its performance. For six basic emotions, the results show that EmoTracer can achieve 95.6% accuracy in intra-subject emotion classification and 80.7% accuracy with 5 shots in cross-subject emotion classification.
Wenting Kuang, Yuda Zheng, Yongpan Zou, Kaishun Wu
BIBM5
2024 Cop: Continuously Pairing of Heterogeneous Wearable Devices Based on Heartbeat
Wenfeng He, Yongpan Zou, Weipeng Cao
KSEM (3)3
2024 Ultra Write: A Lightweight Continuous Gesture Input System with Ultrasonic Signals on COTS Devices
abstract
Due to the advantages of acoustic sensing such as device ubiquity, hands-free interaction and privacy security, acoustic-based gesture input techniques have gained extensive at-tention and many excellent works have been proposed. However, these works have the following shortcomings: high cost of system construction, non-continuous input, and degraded performance in cross-user scenarios. To overcome the above shortcomings, we propose UltraWrite, a continuous gesture input system that needs rather low system construction cost, supports continuous input, and achieves high cross-user recognition performance. The key idea of our solution is to synthesize the data of continuous gestures from isolated ones, and build an end-to-end continuous gesture recognition model by introducing the connectionist temporal classification (CTC) mechanism widely used in natural language processing. We have implemented the system on a commercial tablet and conducted comprehensive experiments to evaluate its performance. The results demonstrate that UltraWrite achieves an average word accuracy of 99.3% and word error rate of 0.34% when considering only the first output candidate word. In addition, we have also evaluated the system's robustness to background noise, sensing distance and angle. The results reveal that UltraWrite displays strong robustness to these factors.
Canlin Zheng, Wenfeng He, Yongpan Zou, Kaishun Wu
PerCom4
2024 EchoGest: A Highly Scalable Unseen Gesture Recognition System Based on Feature-Wise Transformation
abstract
Recent research studies have made significant progress in acoustic-based gesture recognition. However, existing methods lack the capability to expand to customized gestures and adapt to different practical environments. We propose a highly scalable gesture recognition system called EchoGest which integrates a well-designed feature-wise transformation layer into prototypical network framework, and accomplishes unseen gesture recognition with a device’s built-in speaker and microphone. Our key insight involves gauging the similarity between query sample representations and class prototypes in the embedding space, and thus enabling the scalability to unseen gestures. Meanwhile, we introduce a feature transformation layer to linearly adjust feature maps and propose an efficient two-stage training strategy to obtain regularized parameters for this layer. Specifically, this layer employs affine transformation to enhance intermediate feature activations and yield more diverse feature distributions for cross-domain recognition, and it improves recognition accuracy by 10% in 1-shot cases. We train the system with a collected a letter gestures (i.e., writing ’A’ to ’Z’) dataset and test it on a digit gestures (i.e., writing ’0’ to ’9’) dataset with 10 volunteers. The results show that EchoGest can recognize unseen digit gestures with an accuracy of 93.7% in 2-shot cases, and 93.2% in the leave-one-user-out testing setting. We also explore a semi-supervised clustering approach in which each user’s data can be used to update his or her prototypes for personalized customization. The comprehensive experiments also verify that EchoGest remain good performance across various environments, age groups, and different devices.
Yunshu Wang, Weiwei Lu, Yanbo He, Yongpan Zou, Kaishun Wu, Victor C. M. Leung
IEEE Internet Things J.5
2024 EarPrint: Earphone-Based Implicit User Authentication With Behavioral and Physiological Acoustics
abstract
With the increasing pervasiveness of smart earphones, it is appealing to propose more unobtrusive and convenient wearable authentication methods. Researchers have designed earphone-based authentication systems which utilize high-frequency audio signals to scan ear canal structure. Nevertheless, they possess shortcomings of low unobtrusiveness and robustness. In this article, we put forward an earphone-based passive authentication system which makes use of physiological and behavioral acoustic signals caused by a user’s natural actions, including putting on earphones and inner organs’ activities, respectively. By introducing attention mechanism into the network design, our method adaptively weighs two channel signals, and extracts stable fingerprints for different people, which relieves model retraining for unseen users and improves its scalability. We have built a real-time prototype called EarPrint by designing the earphones and a mobile application, and conducted comprehensive experiments under diverse settings. Experimental results demonstrate that EarPrint has low false acceptance rate (FAR) and equal error rate (EER) less than 1% and 5% in most cases, respectively.
Yongpan Zou, Jianhao Weng, Haibo Lei, Dan Wang 0002, Victor C. M. Leung, Kaishun Wu
IEEE Internet Things J.1
2024 PreGesNet: Few-Shot Acoustic Gesture Recognition Based on Task-Adaptive Pretrained Networks
abstract
Acoustic-based human gesture recognition (HGR) applications have drawn increasing academic attention in order to overcome the shortcomings of conventional interaction methods on tiny devices. Existing techniques following a learning-based routine requires collecting massive application-specific training data. What is worse, the cross-domain problem induces additional retraining overhead to enable the systems recognize unseen gestures in different environments. This obviously decreases their scalability and prevent them from real-world deployment. Although some recent works propose different few-shot learning solutions to deal with the cross-domain problem in HGR, they possess shortcomings of being application-specific, high training overhead, and/or incapability to recognize unseen gestures. In this paper, we propose PreGesNet, a few-shot acoustic gesture recognition framework based on task-adaptive pretrained networks whose novelty lies in three aspects: i) leveraging pretrained feature extractor which captures generic knowledge of our collected and open-source large-scale gesture datasets; ii) designing task-specific parameter adaptation mechanism to efficiently update the feature extractor to adapt the pretrained feature extractor to each target task; iii) discovering suitable distance metric and task generation strategy which fit HGR application. According to the experiments, when the model is trained with 10 digit gestures, its recognition accuracies of 26 kinds of letter gestures and 8 kinds of other hand gestures can be up to 80.5% and 93.4% with only two shots, respectively. In addition, the average recognition latency of PreGesNet is less than 0.4 second.
Yongpan Zou, Yunshu Wang, Haozhi Dong, Yaqing Wang 0002, Yanbo He, Kaishun Wu
IEEE Trans. Mob. Comput.1
2023 CHAR: Composite Head-body Activities Recognition with A Single Earable Device
abstract
The increasing popularity of earable devices stimulates great academic interest to design novel head gesture-based interaction technologies. But existing works simply consider it as a singular activity recognition problem. This is not in line with practice since users may have different body movements such as walking and jogging along with head gestures. It is also beneficial to recognize body movements during human-device interaction since it provides useful context information. As a result, it is significant to recognize such composite activities in which actions of different body parts happen simultaneously. In this paper, we propose a system called CHAR to recognize composite head-body activities with a single IMU sensor. The key idea of our solution is to make use of the inter-correlation of different activities and design a multi-task learning network to extract shared and specific representations. We implement a real-time prototype and conduct extensive experiments to evaluate it. The results show that CHAR can recognize 60 kinds of composite activities (12 head gestures and 5 body movements) with high accuracies of 97.0% and 89.7% in user- dependent and independent cases, respectively.
Peizhao Zhu, Yongpan Zou, Kaishun Wu
PERCOM2
2023 Ubiquitous WiFi and Acoustic Sensing: Principles, Technologies, and Applications
Jia-Ling Huang, Yunshu Wang, Yongpan Zou, Kaishun Wu, Lionel M. Ni
J. Comput. Sci. Technol.3
2023 Beyond Legitimacy, Also With Identity: Your Smart Earphones Know Who You Are Quietly
abstract
User authentication and identification on smart devices has great significance in keeping data privacy and recommending personalized services. With the rising popularity of smart earphones recently, they open up a new world for users to enjoy music individually, but also bring about privacy concerns at the same time. Existing few research works propose positive sensing systems that emit and receive inaudible acoustic signals to authenticate users. However, they share shortcomings of intrusiveness to users, high power consumption, and purely focusing on authentication. Instead, in this paper, we propose a passive sensing system called${{\sf EarID}}$with low-cost customized earphones which attains user authentication and identification at once. It makes use of a embedded microphone to sense body sounds spread out through ear canals and extract ‘fingerprints’ as a novel biometric feature. With self-designed earphones, we design a deep learning-based real-time data processing pipeline and cope with external interference. Extensive experiments under different real-world settings show that${{\sf EarID}}$can achieve a rather low false acceptance rate of$3.4\%$for user authentication and a high$F1$score of$95.5\%$for legitimate user identification.
Yongpan Zou, Haibo Lei, Kaishun Wu
IEEE Trans. Mob. Comput.1
2022 EchoWrite 2.0: A Lightweight Zero-Shot Text-Entry System Based on Acoustics
abstract
Limited by size, shape, and other factors, it is rather inconvenient to interact with new smart devices by traditional methods. Acoustic-based methods following a machine learning approach have been put forward to resolve this problem in previous works. But they possess limitations of heavy training overhead, low performance for unseen users, and intensive computation cost. Following our previous work in this area, we further overcome shortcomings of existing work and propose a lightweight and zero-shot text-entry system for unseen users based on acoustic sensing. The key novelty of this work is proposing a new model training strategy including dataset construction and augmentation methods to effectively enhance generalization ability of a simple learning model with as few training data as possible, based on our insight into the problem. We design and implement a real-time Android application system called EchoWrite 2.0 to validate our idea with extensive experiments. Results show that EchoWrite 2.0 can recognize digits, English letters, and words with an accuracy of 85.3%, 73.2%, and 96.9%, respectively, for unseen users without providing any data to the learning model. The comparison with related work in different aspects shows overall superiority of EchoWrite 2.0 .
Yongpan Zou, Zhihong Xiao, Shicong Hong, Zishuo Guo, Kaishun Wu
IEEE Trans. Hum. Mach. Syst.1
2022 Beyond RSS: A PRR and SNR Aided Localization System for Transceiver-Free Target in Sparse Wireless Networks
abstract
Nowadays transceiver-free (also referred to as device-free) localization using Received Signal Strength (RSS) is a hot topic for researchers due to its widespread applicability. However, RSS is easily affected by the indoor environment, resulting in a dense deployment of reference nodes. Some hybrid systems have already been proposed to help RSS localization, but most of them require additional hardware support. In order to solve this problem, in this paper, we propose two algorithms, which leverage the Packet Received Rate (PRR) to help RSS localization without additional hardware support. Moreover, we take the environment noise information into consideration by utilizing the Signal-to-Noise Ratio (SNR) which is based on the RSS and Noise Floor (NF) information instead of pure RSS. Thus, we can alleviate the noise effect in the environment and make our system more sensitive to the target. Specifically, when reference nodes are sparsely deployed and RSS is very weak, PRR and SNR can help in performing localization more accurately. Our BEYOND RSS system is based on sparse wireless sensor networks, wherein the experimental results show that the average localization error of our approach outperforms the pure RSS based approach by about 15.19%.
Dian Zhang 0001, Wen Xie 0005, Zexiong Liao, Wenzhan Zhu, Landu Jiang, Yongpan Zou
IEEE Trans. Mob. Comput.6
2021 EchoWrite: An Acoustic-Based Finger Input System Without Training
abstract
Recently, wearable devices have become increasingly popular in our lives because of their neat features and stylish appearance. However, their tiny sizes bring about new challenges to human-device interaction such as texts input. Although some novel methods have been put forward, they possess different defects and are not applicable to deal with the problem. As a result, we propose an acoustic-based texts-entry system, i.e., EchoWrite, by which texts can be entered with a finger writing in the air without wearing any additional device. More importantly, different from many previous works, EchoWrite runs in a training-free style which reduces the training overhead and improves system scalability. We implement EchoWrite with commercial devices and conduct comprehensive experiments to evaluate its texts-entry performance. Experimental results show that EchoWrite enables users to enter texts at a speed of 7.5 WPM without practice, and 16.6 WPM after about 30-minute practice. This speed is better than touch screen-based method on smartwatches, and comparable with previous related works. Moreover, EchoWrite provides favorable user experience of entering texts.
Kaishun Wu, Qiang Yang 0018, Baojie Yuan, Yongpan Zou, Rukhsana Ruby, Mo Li 0001
IEEE Trans. Mob. Comput.4
2020 I am Smartglasses, and I Can Assist Your Reading
Baojie Yuan, Yetong Han, Jialu Dai, Yongpan Zou, Kaishun Wu
ICA3PP (2)4
2020 What you wear know how you feel: an emotion inference system with multi-modal wearable devices
abstract
Emotions show high significance on human health. Automatic emotion recognition is helpful for monitoring psychological disorders, mental problems and exploring behavioral mechanisms. Existing approaches adopt costly and bulky specialized hardware such as EEG/ECG helmet, possess privacy risks, or with low accuracy and user experience. With the increasing popularity of wearables, people tend to equip multiple smart devices, which provides potential opportunity for emotion perception. In this paper, we present a pervasive and portable system called MW-Emotion to recognize common emotional states with multi-modal wearable devices. However, ubiquitous wearable devices perceive shallow information which is not obviously related to human emotions. MW-Emotion excavates intrinsic mapping relationship between emotions and sensing data. Our experiments show that MW-Emotion can recognize different emotion states with a relatively high accuracy of 83.1%.
Dan Wang 0002, Haibo Lei, Haozhi Dong, Yunshu Wang, Yongpan Zou, Kaishun Wu
MobiCom5
2020 SilentSign: Device-free Handwritten Signature Verification through Acoustic Sensing
abstract
Signature is one of the most prevailing identity authorization approaches. It is yet inconvenient to use in real life in the sense that a majority of existing signature verification approaches rely on additional digital signing devices. In this paper, we propose a portable device-free signature verification system named SilentSign which makes use of acoustic sensors (i.e., microphone and speaker) embedded in smart devices to enable secure and convenient signature verification service. The basic idea is to leverage acoustic signals to measure the distance variation of the tip of the pen while signing. We carefully design the signal modulation scheme, develop a phase-based distance measurement technique, and train the verification model for high performance and robustness. Compared with conventional digital signing systems, SilentSign allows users to sign more invisibly and conveniently. We conduct extensive experiments involving 35 participants to evaluate SilentSign. Results show that SilentSign can achieve 98.2% AUC and 1.25% EER.
Yongpan Zou, Rukhsana Ruby, Kaishun Wu
PerCom3
2020 Smart earpieces that know who you are quietly: poster abstract
abstract
User authentication and identification on smart devices has great significance in keeping data privacy and recommending personalized services. Existing few research works propose active sensing systems that emit and receive inaudible acoustic signals to authenticate users. But they share shortcomings of intrusiveness to users, high power consumption, and purely focusing on authentication. Instead, in this paper, we propose a passive sensing system called EarID with low-cost customized earpieces which attains user authentication and identification simultaneously. It makes use of a embedded microphone to sense body sounds spread out through ear canals and extract 'fingerprints' as a novel biometric feature. With self-designed earpieces, we design a deep learning-based real-time data processing pipeline. Extensive experiments under different real-world settings show that EarID can achieve a rather low false acceptance rate less than 5% for user authentication and a high F1 score of 96% for user identification.
Haibo Lei, Yongpan Zou, Kaishun Wu
SenSys3
2020 Tap it and you know what it is: a surface identification system based on acoustic dispersion: poster abstract
abstract
Surface identification provides contextual services during humancomputer interaction, which is important for target detection and scene understanding. A robust and ubiquitous surface recognition system has a wide range of applications such as context awareness and robot operation. Existing methods have shortcomings of requiring specialized devices and limited usage scenarios. In this paper, we introduce Surtify, a surface identification system based on acoustic dispersion with a smartphone. By combining the intrinsic physical phenomenon (i.e., acoustic dispersion) with a deep learning model, Surtify can identify eleven kinds of surfaces with accuracies up to 96%, even in cross-person and cross-location scenarios.
Baojie Yuan, Shicong Hong, Yongpan Zou, Kaishun Wu
SenSys3
2020 A Low-Cost Smart Glove System for Real-Time Fitness Coaching
abstract
Strength training is becoming increasingly popular among all age groups, as it helps the participants increase muscle strength, improve body flexibility, reduce health risks, and reshape physical forms. However, strength training imposes strict regulations on gestures and requires professional instruction in real time for the sake of body-building efficiency and safety. For this purpose, in this article, we propose a novel low-cost system named iCoach, to provide real-time monitoring and coaching service for strength training participants. Specifically, we design and implement a smart fitness glove, which can be seamlessly equipped with a pervasive inertial unit. With this customized but low-cost device, we can recognize various training programs, detect nonstandard behaviors while exercising, and assess exercising qualities of a user. Our primary experimental results show that iCoach can recognize 15 sets of training programs, detect three common nonstandard behaviors, and assess the quality of training with high accuracy and reliability.
Yongpan Zou, Dan Wang 0002, Shicong Hong, Rukhsana Ruby, Dian Zhang 0001, Kaishun Wu
IEEE Internet Things J.1
2019 EchoWrite: An Acoustic-based Finger Input System Without Training
abstract
Recently, wearable devices have become increasingly popular in our lives because of their neat features and stylish appearance. However, their tiny sizes bring about new challenges to human-device interaction such as texts input. Although some novel methods have been put forward, they possess different defects and are not applicable to deal with the problem. As a result, we propose an acoustic-based texts-entry system, i.e., EchoWrite, by which texts can be entered with a finger writing in the air without wearing any additional device. More importantly, different from many previous works, EchoWrite runs in a training-free style which reduces the training overhead and improves system scalability. We implement EchoWrite with commercial devices and conduct comprehensive experiments to evaluate its texts-entry performance. Experimental results show that EchoWrite enables users to enter texts at a speed of 7.5 WPM without practice, and 16.6 WPM after about 30- minute practice. This speed is better than touch screen-based method on smartwatches, and comparable with previous related works.
Yongpan Zou, Qiang Yang 0018, Rukhsana Ruby, Yetong Han, Sicheng Wu, Mo Li 0001, Kaishun Wu
ICDCS1
2019 AcouDigits: Enabling Users to Input Digits in the Air
abstract
Recently, wearable devices have become increasingly popular in our lives because of their neat features and stylish appearance. However, due to the tiny size, it is inconvenient for users to interact with a device using conventional methods, especially for text entry. Although some methods have been proposed to handle this problem, they have different limitations and are not applicable to many existing mobile devices. As a result, we take the first step to propose a digits-entry system, i.e., AcouDigits, in which digits can be entered in the air using a finger without taking help from any additional hardware. We implement AcouDigits on two commercial devices and conduct experiments to evaluate its performance in recognizing ten basic digits. Experimental results show that AcouDigits can achieve average accuracies of 91.7% and 87.4% in recognizing basic digits and 26 English alphabets, respectively.
Yongpan Zou, Qiang Yang 0018, Yetong Han, Dan Wang 0002, Jiannong Cao 0001, Kaishun Wu
PerCom1
2018 ArmIn: Explore the Feasibility of Designing a Text-entry Application Using EMG Signals
abstract
EMG is becoming an emerging interface for human-computer interface and has been applied to gesture recognition in previous work. However, those existing EMG-based interfaces can only recognize gestures at a coarse-grained level such as hand and arm gestures, which constraints their usage in applications involving fine-grained activities such as text entry via keystrokes. As a result, in this paper, we attempt to push the limit of existing EMG-based interfaces and propose the first wearable text-entry system, named ArmIn, with EMG signals. ArmIn is designed to recognize keystroke gestures with the help of a finger on printed and physical keyboards. We implement ArmIn using commodity EMG sensors and custom hardware board, and conduct experiments to evaluate its performance. By carefully designing the data processing scheme, ArmIn can recognize keystrokes on both kinds of keyboard, with 89.5% and 87.5% accuracy respectively, when it is worn on a user's left arm.
Qiang Yang 0018, Yongpan Zou, Kaishun Wu
MobiQuitous2
2018 A Novel Finger-Assisted Touch-free Text Input System Without Training
abstract
Recently, tiny smart devices have become increasingly popular in our lives because of their neat features and stylish appearance. However, their small form factors, especially screens, make it inconvenient for users to enter texts with conventional methods such as soft keyboards, which need a fairly large screen. To address this problem, we propose a novel texts-input system, called EchoType, with which users can enter texts with a finger writing in the air. EchoType makes use of acoustic sensors (i.e., microphone and speaker) to sense finger gestures and infer texts based on mapping relation between gestures and basic letters. We take a step to enable users to input texts with acoustic signals. Compared with existing approaches, EchoType enjoys merits of low hardware requirements and high scalability to different mobile devices.
Qiang Yang 0018, Hongrui Fu, Yongpan Zou, Kaishun Wu
MobiSys3
2017 ABAid: Navigation Aid for Blind People Using Acoustic Signal
abstract
Blind mobility aid is a primary part in the daily life of blind people. Although plenty of systems or devices are invented to make the navigation of blind people easier, those are generally expensive and hardly affordable for them. To solve these issues, we introduce ABAid, a novel system designed for blind or visually impaired people to navigate, with commercial off-the-shelf (COTS) mobile devices. Based on in-depth acoustic localization and gyroscope techniques, this system is not only the means of huge convenience to carry, but also is capable of detecting obstacles before reaching them. In our experiments designed to detect the distance of wall, the proposed system achieves 3.24% average error rate. It can further measure the direction of wall, and the average error in this case is 2.73°. With high accuracy and stable measurement, ABAid is able to help blind people move independently in fairly uncomplicated scenarios.
Zehui Zheng, Rukhsana Ruby, Yongpan Zou, Kaishun Wu
MASS4
2017 TagFree: Passive object differentiation via physical layer radiometric signatures
abstract
Object differentiation plays a vital role in our daily life and such systems are widely deployed with RFID tags or bar codes attached on goods. In certain scenarios, however, attaching tags to objects may be impractical due to cost and protection issues. In this paper, we propose TagFree, a novel object differentiation scheme without attaching tags. Instead of relying on external tags, we exploit the inherent radiometric properties of different objects as their signatures. To improve the robustness and efficiency of TagFree, we empirically determine a spatial safe zone and harness successive cancellation to distinguish multiple objects simultaneously. We prototype TagFree on commercial WiFi infrastructure and evaluate its performance in various indoor scenarios. Experimental results demonstrate that TagFree achieves single object distinguishing accuracy of 96% measured at the same location, and over 80% within the safe zone range of up to 3m along a 7m link. TagFree can also differentiate up to 3 objects with acceptable accuracy.
Yongpan Zou, Shufeng Ye, Kaishun Wu, Lionel M. Ni
PerCom1
2017 GRfid: A Device-Free RFID-Based Gesture Recognition System
abstract
Gesture recognition has emerged recently as a promising application in our daily lives. Owing to low cost, prevalent availability, and structural simplicity, RFID shall become a popular technology for gesture recognition. However, the performance of existing RFID-based gesture recognition systems is constrained by unfavorable intrusiveness to users, requiring users to attach tags on their bodies. To overcome this, we propose GRfid, a novel device-free gesture recognition system based on phase information output by COTS RFID devices. Our work stems from the key insight that the RFID phase information is capable of capturing the spatial features of various gestures with low-cost commodity hardware. In GRfid, after data are collected by hardware, we process the data by a sequence of functional blocks, namely data preprocessing, gesture detection, profiles training, and gesture recognition, all of which are well-designed to achieve high performance in gesture recognition. We have implemented GRfid with a commercial RFID reader and multiple tags, and conducted extensive experiments in different scenarios to evaluate its performance. The results demonstrate that GRfid can achieve an average recognition accuracy of 96.5 and 92.8 percent in the identical-position and diverse-positions scenario, respectively. Moreover, experiment results show that GRfid is robust against environmental interference and tag orientations.
Yongpan Zou, Jiang Xiao 0001, Jinsong Han, Kaishun Wu, Yun Li 0002, Lionel M. Ni
IEEE Trans. Mob. Comput.1
2016 We Can Hear You with Wi-Fi!
abstract
Recent literature advances Wi-Fi signals to “see” people's motions and locations. This paper asks the following question: Can Wi-Fi “hear” our talks? We present WiHear, which enables Wi-Fi signals to “hear” our talks without deploying any devices. To achieve this, WiHear needs to detect and analyze fine-grained radio reflections from mouth movements. WiHear solves this micro-movement detection problem by introducing Mouth Motion Profile that leverages partial multipath effects and wavelet packet transformation. Since Wi-Fi signals do not require line-of-sight, WiHear can “hear” people talks within the radio range. Further, WiHear can simultaneously “hear” multiple people's talks leveraging MIMO technology. We implement WiHear on both USRP N210 platform and commercial Wi-Fi infrastructure. Results show that within our pre-defined vocabulary, WiHear can achieve detection accuracy of 91 percent on average for single individual speaking no more than six words and up to 74 percent for no more than three people talking simultaneously. Moreover, the detection accuracy can be further improved by deploying multiple receivers from different angles.
Yongpan Zou, Zimu Zhou, Kaishun Wu, Lionel M. Ni
IEEE Trans. Mob. Comput.2
2016 SmartScanner: Know More in Walls with Your Smartphone!
abstract
Seeing through walls and knowing clearly what exist inside just like a superman are not only fantastic wishes for humans, but also of much practical significance. For example, you would like to know whether there are pipes, or rebars inside a wall before drilling into it. Moreover, knowing how pipes are configured in a wall before attempting to fix defects would definitely prevent unnecessary damages. Existing methods that intend to address this issue are either costly due to the use of high-end technology, or restrictive for reasons of some strong assumptions. However, in this paper, we present a novel system, SmartScanner, which is based on off-the-shelf sensors embedded in a smartphone. SmartScanner makes full use of in-built sensors, namely, the accelerometer, gyroscope, and magnetometer to achieve this goal inexpensively and conveniently. Specifically, by combining these sensors, we are able to clearly distinguish certain objects inside a wall and map out the layout of an in-wall pipeline system. We implement SmartScanner on two smartphone platforms, namely iPhone 4 and Xiaomi Mi2S, and conduct extensive experiments to evaluate its performance. Experiments show that SmartScanner can achieve high accuracies in distinguishing objects in various scenarios. Meanwhile, as for layout mapping, 90 percent of length errors are limited to several centimeters for horizontal and vertical pipeline segments, respectively. Also, SmartScanner can achieve centimeter-level position errors of turning points in horizontal and vertical directions in the testbed.
Yongpan Zou, Kaishun Wu, Lionel M. Ni
IEEE Trans. Mob. Comput.1
2015 WiG: WiFi-Based Gesture Recognition System
abstract
Most recently, gesture recognition has increasingly attracted intense academic and industrial interest due to its various applications in daily life, such as home automation, mobile games. Present approaches for gesture recognition, mainly including vision-based, sensor-based and RF-based, all have certain limitations which hinder their practical use in some scenarios. For example, the vision-based approaches fail to work well in poor light conditions and the sensor-based ones require users to wear devices. To address these, we propose WiG in this paper, a device-free gesture recognition system based solely on Commercial Off-The-Shelf (COTS) WiFi infrastructures and devices. Compared with existing Radio Frequency (RF)-based systems, WiG stands out for its systematic simplicity, extremely low cost and high practicability. We implemented WiG in indoor environment and conducted experiments to evaluate its performance in two typical scenarios. The results demonstrate that WiG can achieve an average recognition accuracy of 92% in line-of-sight scenario and average accuracy of 88% in the none-line-of sight scenario.
Wenfeng He, Kaishun Wu, Yongpan Zou, Zhong Ming 0001
ICCCN3
2014 SmartSensing: Sensing Through Walls with Your Smartphone!
abstract
Seeing through walls and knowing clearly what exist inside just like a superman are not only fantastic wishes for humans, but also of much practical significance. For example, you would like to know whether there are pipes, or rebars inside a wall before drilling into it. Moreover, knowing how pipes are configured in a wall before attempting to fix defects would definitely prevent unnecessary damages. Existing methods that intend to address this issue are either costly due to the use of high-end technology, or too restrictive for reasons of some strong assumptions. However, in this paper, we present a novel system, SmartSening, which is based on off-the-shelf sensors embedded in smartphones. SmartSensing makes full use of in-built sensors, namely, the accelerometer, the gyroscope, and the magnetometer to achieve this goal inexpensively and conveniently. Specifically, by combining these sensors, we are able to clearly distinguish certain objects inside a wall. In addition, the layout of a pipeline system can be mapped out automatically in an economical and laborsaving way. We implement this system on two different kinds of smartphone platforms, namely iPhone4 and Xiaomi Mi2S. We conduct experiments in a proof-of-concept testbed of size 1.8m×1.0m. Experimental results show that SmartSensing can achieve no less than an average accuracy of 96%, 89% and 77% in distinguishing objects under three different depths, respectively. Also, as for layout mapping, it can achieve less than 32cm and 28cm length error with 90% probability on average for whole horizontal and vertical pipeline segments, with a 6.8m and 4.0m total length, respectively.
Yongpan Zou, Kaishun Wu, Lionel M. Ni
MASS1
2014 We can hear you with Wi-Fi!
abstract
Recent literature advances Wi-Fi signals to "see" people's motions and locations. This paper asks the following question: Can Wi-Fi "hear" our talks? We present WiHear, which enables Wi-Fi signals to "hear" our talks without deploying any devices. To achieve this, WiHear needs to detect and analyze fine-grained radio reflections from mouth movements. WiHear solves this micro-movement detection problem by introducing Mouth Motion Profile that leverages partial multipath effects and wavelet packet transformation. Since Wi-Fi signals do not require line-of-sight, WiHear can "hear" people talks within the radio range. Further, WiHear can simultaneously "hear" multiple people's talks leveraging MIMO technology. We implement WiHear on both USRP N210 platform and commercial Wi-Fi infrastructure. Results show that within our pre-defined vocabulary, WiHear can achieve detection accuracy of 91% on average for single individual speaking no more than 6 words and up to 74% for no more than 3 people talking simultaneously. Moreover, the detection accuracy can be further improved by deploying multiple receivers from different angles.
Yongpan Zou, Zimu Zhou, Kaishun Wu, Lionel M. Ni
MobiCom2