Tao Chen 0033

dblp:69/510-33 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0003-4565-5548ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ZA-SLAM: Leveraging Vision-Language Model for Zero-Shot Acoustic SLAM
abstract
Existing acoustic indoor location sensing systems are limited by the need for extensive data collection and model retraining in unseen environments. This paper introduces ZA-SLAM, a novel zero-shot acoustic Simultaneous Localization and Mapping (SLAM) system that can be deployed in unseen environments without model retraining. Our core idea is to train an acoustic encoder that inherits the generalization capabilities of pre-trained Vision-Language Models (VLMs), which show superiority in tasks like zero-shot visual SLAM. To achieve this goal, we perform Acoustic-Visual Feature Alignment to enable the acoustic encoder to generate features aligned with visual features from VLMs. To select high-quality images for effective alignment, we design a Semantic-Guided Image Selection that filters out low-quality collected images caused by factors like abrupt view changes, occlusions, and uninformative views. Furthermore, we address the challenge of false positive loop closures in structurally similar locations with the Learning-Based Trajectory Reachability Matching that validates loop closures leveraging IMU trajectory features. Extensive real-world experiments demonstrate that our system achieves comparable SLAM performance to retraining-based acoustic SLAM, and much improved performance compared to existing zero-shot Wi-Fi and geomagnetic SLAM systems. Our system achieves a mean mapping error of 0.56 m and a localization error of 0.78 m across multiple unseen environments.
Zhuochen Yu, David K. Y. Yau, Yijie Shen, Xiaoran Fan, Tao Chen 0033, Qun Song 0001
MobiSys5
2026 Longan: Ultra-Low-Power, Long-Range LoRa Receiver for LPWANs
abstract
Low-Power Wide-Area Networks (LPWANs) are essential for IoT connectivity, but they face a longstanding fundamental challenge: achieving both continuous receiver availability and long-term battery operation without compromising network performance. In low-duty-cycle modes like LoRaWAN’s Class A or B, data timeliness and network throughput are limited, and devices are deaf to peer transmissions, restricting topologies to a single hop, hindering scalability in practical applications. While Class C enables always-on operation for real-time communication and multi-hop networking, commercial LoRa transceivers consume up to 42.4 mW in this mode, rendering sustained battery-powered deployments infeasible. Recent low-power designs cut consumption to sub-milliwatt levels but limit communication ranges to mere hundreds of meters, undermining LPWAN’s “wide-area’’ vision. This paper introduces Longan, the first LoRa receiver to overcome this power-range tradeoff boundary, enabling kilometer-scale links with a power of about 1 mW only. Longan achieves this via two breakthroughs: (1) We design a novel LoRa receiver analog radio frequency (RF) front-end that exploits negative differential resistors for high-efficiency signal amplification while simultaneously reducing power consumption by 2–3 orders of magnitude; (2) Building on this, we further introduce an architectural decoupling of detection and demodulation, and design a lightweight analog dechirping circuit using our front-end. This circuit enables always-on preamble detection at just one-tenth the power of digital counterparts, triggering COTS demodulation only on demand. Evaluation shows that Longan sustains continuous detection at 1.16 mW, a 36 × reduction from COTS LoRa, while preserving sensitivity within 3–13 dB of commercial devices and enabling long-range communication.
Tao Chen 0033, Zhenjiang Li 0001
SenSys2
2026 Neural-Enhanced Modulation for Spatial Selective Transmission on Low-End IoT Devices
abstract
This paper tries to answer a question: “Can we achieve spatial-selective transmission on IoT devices?” A positive answer would enable more secure data transmission among IoT devices. The challenge, however, is how to manipulate signal propagation without relying on beamforming antenna arrays which are usually unavailable on low-end IoT devices. We give an affirmative answer by introducing SpotSound, a novel acoustic communication system that exploits the diversity of multi-path indoors as a naturalbeamformer. By judiciously controlling the way how the information is embedded into the signal, SpotSound can make the signal decodable only when the signal propagates along a certain multipath channel. Since the multipath channel decorrelates rapidly over the distance between receivers, SpotSound can ensure the signal is decodable only at the target position, achieving precise physical isolation. SpotSound is a purely software-based solution that can run on most IoT devices where speakers and microphones are widely used. We implement SpotSound on Raspberry Pi connected with COTS microphone and speaker. Experimental results show that SpotSound could precisely focus its signal on spots with customized sizes ranging from 0.04m2to 0.5m2.
Huangwei Wu, Tingchao Fan, Meng Jin 0002, Tao Chen 0033, Xinbing Wang, Chenghu Zhou
IEEE Trans. Netw.4
2025 Heart Rate Monitoring Through ANC Headphones in Unconstrained Environments
abstract
This paper introduces CLEAR-APG, a novel acoustic sensing approach that enables reliable heart rate monitoring in unconstrained environments using off-the-shelf active noise cancellation (ANC) headphones. By emitting ultrasonic signals into the user's ear canal via the headphone speaker and analyzing their echoes, which can detect the frequency of a pulsating vein along the canal wall. However, everyday activities such as exercising, speaking, or eating cause jaw movements that deform the ear canal, overwhelming the subtle deformation caused by blood flowing. To overcome this challenge, we employ the ANC headphone's built-in gyroscope to capture body motion and identify how various motion patterns influence the heartbeat waveform. Building on this insight, we propose a multi-modal method that effectively denoises the heartbeat waveform measurements and further accurately extracts heart rate. We implement CLEARAPG on ANC earbuds and conduct comprehensive field studies on 14 users. The results show that CLEAR-APG achieves an average heart rate error of 4.01% across seven different activities, satisfying industry-required margin of 10% heart rate error.
Maanya Shanker, Tao Chen 0033, Xiaoran Fan, Longfei Shangguan
BSN3
2025 LeakyFeeder: In-Air Gesture Control Through Leaky Acoustic Waves
abstract
We present LeakyFeeder, a mobile application that explores the acoustic signals leaked from headphones to reconstruct gesture motions around the ear for fine-grained gesture control. To achieve this goal, LeakyFeeder repurposes the speaker and a single feedforward microphone on active noise cancellation (ANC) headphones as a SONAR system, using inaudible frequency-modulated continuous-wave (FMCW) signals to track gesture reflections for accurate sensing. Since this single-receiver SONAR system is unable to differentiate reflection angles and further disentangle signal reflections from different gesture parts, we draw on principles of multi-modal learning to frame gesture motion reconstruction as a multi-modal translation task and propose a deep learning-based approach to fill the information gap between low-dimensional FMCW ranging readings and high-dimensional 3D hand movements. We implement LeakyFeeder on a pair of Google Pixel Buds and conduct experiments to examine the efficacy and robustness of LeakyFeeder in various conditions. Experiments based on six gesture types inspired by Apple Vision Pro demonstrate that LeakyFeeder achieves a PCK performance of 89% at 3cm across ten users, with an average MPJPE and MPJRPE error of 2.71cm and 1.88cm, respectively.
Yongjie Yang 0008, Tao Chen 0033, Zhenlin An, Shirui Cao, Xiaoran Fan, Longfei Shangguan
SenSys2
2024 MAF: Exploring Mobile Acoustic Field for Hand-to-Face Gesture Interactions
abstract
We present MAF, a novel acoustic sensing approach that leverages the commodity hardware in bone conduction earphones for hand-to-face gesture interactions. Briefly, by shining audio signals with bone conduction earphones, we observe that these signals not only propagate along the surface of the human face but also dissipate into the air, creating an acoustic field that envelops the individual’s head. We conduct benchmark studies to understand how various hand-to-face gestures and human factors influence this acoustic field. Building on the insights gained from these initial studies, we then propose a deep neural network combined with signal preprocessing techniques. This combination empowers MAF to effectively detect, segment, and subsequently recognize a variety of hand-to-face gestures, whether in close contact with the face or above it. Our comprehensive evaluation based on 22 participants demonstrates that MAF achieves an average gesture recognition accuracy of 92% across ten different gestures tailored to users’ preferences.
Yongjie Yang 0008, Tao Chen 0033, Yujing Huang, Xiuzhen Guo, Longfei Shangguan
CHI2
2024 Leveraging Foundation Models for Zero-Shot IoT Sensing
abstract
Deep learning models are increasingly deployed on edge Internet of Things (IoT) devices. However, these models typically operate under supervised conditions and fail to recognize unseen classes different from training. To address this, zero-shot learning (ZSL) aims to classify data of unseen classes with the help of semantic information. Foundation models (FMs) trained on web-scale data have shown impressive ZSL capability in natural language processing and visual understanding. However, leveraging FMs’ generalized knowledge for zero-shot IoT sensing using signals such as mmWave, IMU, and Wi-Fi has not been fully investigated. In this work, we align the IoT data embeddings with the semantic embeddings generated by an FM’s text encoder for zero-shot IoT sensing. To utilize the physics principles governing the generation of IoT sensor signals to derive more effective prompts for semantic embedding extraction, we propose to use cross-attention to combine a learnable soft prompt that is optimized automatically on training data and an auxiliary hard prompt that encodes domain knowledge of the IoT sensing task. To address the problem of IoT embeddings biasing to seen classes due to the lack of unseen class data during training, we propose using data augmentation to synthesize unseen class IoT data for fine-tuning the IoT feature extractor and embedding projector. We evaluate our approach on multiple IoT sensing tasks. Results show that our approach achieves superior open-set detection and generalized zero-shot learning performance compared with various baselines. Our code is available at https://github.com/schrodingho/FM_ZSL_IoT.
Dinghao Xue, Xiaoran Fan, Tao Chen 0033, Guohao Lan, Qun Song 0001
ECAI3
2024 Exploring the Feasibility of Remote Cardiac Auscultation Using Earphones
abstract
The elderly over 65 accounts for 80% of COVID deaths in the United States. In response to the pandemic, the federal, state governments, and commercial insurers are promoting video visits, through which the elderly can access specialists at home over the Internet, without the risk of COVID exposure. However, the current video visit practice barely relies on video observation and talking. The specialist could not assess the patient's health conditions by performing auscultations.
Tao Chen 0033, Yongjie Yang 0008, Xiaoran Fan, Xiuzhen Guo, Jie Xiong 0001, Longfei Shangguan
MobiCom1
2024 Exploring Biomagnetism for Inclusive Vital Sign Monitoring: Modeling and Implementation
abstract
This paper presents the design, implementation, and evaluation of MagWear, a novel biomagnetism-based system that can accurately and inclusively monitor the heart rate and respiration rate of mobile users with diverse skin tones. MagWear's contributions are twofold. Firstly, we build a mathematical model that characterizes the magnetic coupling effect of blood flow under the influence of an external magnetic field. This model uncovers the variations in accuracy when monitoring vital signs among individuals. Secondly, leveraging insights derived from this mathematical model, we present a softwarehardware co-design that effectively handles the impact of human diversity on the performance of vital sign monitoring, pushing this generic solution one big step closer to real adoptions. We have implemented a prototype of MagWear on a two-layer PCB board and followed IRB protocols to conduct system evaluations. Our extensive experiments involving 30 volunteers demonstrate that MagWear achieves high monitoring accuracy with a mean percentage error (MPE) of 1.55% for heart rate and 1.79% for respiration rate. The head-to-head comparison with Apple Watch 8 further demonstrates MagWear's consistently high performance in different user conditions.
Xiuzhen Guo, Long Tan, Tao Chen 0033, Chaojie Gu, Yuanchao Shu, Shibo He, Yuan He 0004, Jiming Chen 0001, Longfei Shangguan
MobiCom3
2024 Enabling Hands-Free Voice Assistant Activation on Earphones
abstract
We present the design and implementation of EarVoice, a lightweight mobile service that enables hands-free voice assistant activation on commodity earphones. EarVoice comprises two design modules: one for joint speech detection and primary user identification that explores the attributes of the air channel and in-body audio pathway to differentiate between the primary user and others nearby; and another for accurate wakeup word enhancement, which employs a "copy, paste, and adapt" approach to reconstruct the missing high-frequency component in speech recordings. To minimize false positives, enhance agility, and preserve privacy, we deploy EarVoice on a dongle where the proposed signal processing algorithms are streamlined with a gating mechanism to permit only the primary user's speech to enter the pairing device (e.g., a smartphone) for wakeup word recognition, preventing unintended disclosure of ambient conversations. We implemented the dongle on a 4-layer PCB board and conducted extensive experiments with 23 participants in both controlled and uncontrolled scenarios. The experiment results show that EarVoice achieves around 90% wakeup word recognition accuracy in stationary scenarios, which is on par with the high-end, multi-sensor fusion-based Airpods Pro earbud. EarVoice's performance drops to 84% on mobile cases, slightly worse than Airpods (around 90%).
Tao Chen 0033, Yongjie Yang 0008, Chonghao Qiu, Xiaoran Fan, Xiuzhen Guo, Longfei Shangguan
MobiSys1
2023 Towards Spatial Selection Transmission for Low-end IoT devices with SpotSound
abstract
This paper tries to answer a question: "Can we achieve spatial-selective transmission on IoT devices?" A positive answer would enable more secure data transmission among IoT devices. The challenge, however, is how to manipulate signal propagation without relying on beamforming antenna arrays which are usually unavailable on low-end IoT devices.
Tingchao Fan, Huangwei Wu, Meng Jin 0002, Tao Chen 0033, Longfei Shangguan, Xinbing Wang, Chenghu Zhou
MobiCom4
2023 Keystroke Recognition With the Tapping Sound Recorded by Mobile Phone Microphones
abstract
Mobile phones nowadays are equipped with at least dual microphones. We find when a user is typing on a phone, the sounds generated from the vibration caused by finger’s tapping on the screen surface can be captured by both microphones, and these recorded sounds alone are informative enough to localize the user’s keystrokes. This ability can be leveraged to enable useful application designs, while it also raises a crucial privacy risk that the private information typed by users on mobile phones has a great potential to be leaked through such a recognition ability. In this paper, we address two key design issues and demonstrate, more importantly alarm people, that this risk is possible, which could be related to many of us when we use our mobile phones. We implement our proposed techniques in a prototype system and conduct extensive experiments. The evaluation results indicate promising successful rates for more than 4000 keystrokes from different users on various types of mobile phones.
Tao Chen 0033, Yang Liu 0101, Jiao Li 0002, Zhenjiang Li 0001
IEEE Trans. Mob. Comput.2
2023 The Design and Implementation of a Steganographic Communication System over In-Band Acoustical Channels
abstract
This article presents SoundSticker, a system for steganographic, in-band data communication over an acoustic channel. In contrast with recent works that hide bits in inaudible frequency bands, SoundSticker embeds hidden bits in the audible sounds, making them more reliably survive audio codecs and bandpass filtering, while achieving a higher data rate and remaining imperceptible to a listener. The key observation behind SoundSticker is that the human ear is less sensitive to the audio phase changes than the frequency and amplitude changes, which leaves us an opportunity to alter the phase of an audio clip to convey hidden information. We take advantage of this opportunity and build an OFDM-based physical layer. To make this PHY-layer design work for a variety of end devices with heterogeneous computation resources, SoundSticker addresses multiple technical challenges including perceivable waveform artifacts caused by the phase-based modulation, bit rate adaptation without channel sounding and real-time preamble detection. Our prototype on both smartphones and ESP32 platforms demonstrates SoundSticker’s superior performance against the state of the arts, while preserving excellent sound quality and remaining unaffected by common audio codecs like MP3 and AAC. Audio clips produced by SoundSticker can be found at https://soundsticker.github.io/ .
Tao Chen 0033, Longfei Shangguan, Zhenjiang Li 0001, Kyle Jamieson
ACM Trans. Sens. Networks1
2022 Towards Remote Auscultation with Commodity Earphones
abstract
Virtual visits (a.k.a., telehealth) have been promoted in response to the COVID pandemic since early 2020. Despite its convenience, the current virtual visit practice barely relies on video observation and talking. The specialist, however, cannot accurately assess the patient's health condition by listening to acoustic cardiopulmonary signals emanating from the patient's heart with a stethoscope. In this poster, we explore the feasibility of remote auscultation in virtual visits settings by reusing the patient's earphones as a stethoscope. The proposed hardware-software system captures the minute heartbeats from the patient's ear canal. It then offloads these noisy cardiac signals to the pairing device (e.g., a smartphone or a laptop) to reconstruct fine-grained Phonocardiogram (PCG) signals. By listening to the reconstructed PCG signals, the specialist can easily assess the patient's health condition and make the most informed diagnosis. We describe the design challenges and explain our technical roadmap.
Tao Chen 0033, Xiaoran Fan, Yongjie Yang 0008, Longfei Shangguan
SenSys1
2020 Mobile Phones Know Your Keystrokes through the Sounds from Finger's Tapping on the Screen
abstract
Mobile phones nowadays are equipped with at least dual microphones. We find when a user is typing on a phone, the sounds generated from the vibration caused by finger's tapping on the screen surface can be captured by both microphones, and these recorded sounds alone are informative enough to infer the user's keystrokes. This ability can be leveraged to enable useful application designs, while it also raises a crucial privacy risk that the private information typed by users on mobile phones has a great potential to be leaked through such a recognition ability. In this paper, we address two key design issues and demonstrate, more importantly alarm people, that this risk is possible, which could be related to many of us when we use our mobile phones. We implement our proposed techniques in a prototype system and conduct extensive experiments. The evaluation results indicate promising successful rates for more than 4000 keystrokes from different users on various types of mobile phones.
Tao Chen 0033, Yang Liu 0101, Zhenjiang Li 0001
ICDCS2
2020 Metamorph: Injecting Inaudible Commands into Over-the-air Voice Controlled Systems
Tao Chen 0033, Longfei Shangguan, Zhenjiang Li 0001, Kyle Jamieson
NDSS1
2020 Adversarial Attacks and Defenses on Cyber-Physical Systems: A Survey
abstract
Cyber-security issues on adversarial attacks are actively studied in the field of computer vision with the camera as the main sensor source to obtain the input image or video data. However, in modern cyber-physical systems (CPSs), many other types of sensors are becoming popularly used, such as surveillance sensors, microphones, and textual interfaces. A series of recent works investigates the adversarial attacks and the potential defenses in these noncamera sensor-based CPSs. Therefore, this article provides a systematic discussion on these existing works and serves as a complimentary summary of the adversarial attacks and defenses for CPSs beyond the field of computer vision. We first introduce a general working flow for adversarial attacks on CPSs. On this basis, a clear taxonomy is provided to organize existing attacks effectively and indicate where the defenses can be potentially performed in CPSs as well. Then, we discuss these existing attacks and defenses with detailed comparison studies. Finally, we point out concrete research opportunities to be further explored along this research direction.
Jiao Li 0002, Yang Liu 0101, Tao Chen 0033, Zhenjiang Li 0001, Jianping Wang 0001
IEEE Internet Things J.3