Tianfang Zhang

dblp:231/7517 · DBLP profile ↗
← Back
32ranked-venue papers
11as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 4 first-author · 12 since 2021Security and privacy · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CAS-ViT: Convolutional Additive Self-Attention Vision Transformers for Efficient Mobile Applications
abstract
Vision Transformers (ViTs) mark a revolutionary advance in neural networks with their token mixer's powerful global context capability. However, the pairwise token affinity and complex matrix operations limit its deployment on resource-constrained scenarios and real-time applications, such as mobile devices, although considerable efforts have been made in previous works. In this paper, we introduce CAS-ViT: Convolutional Additive Self-attention Vision Transformers, to achieve a balance between efficiency and performance in mobile applications. Firstly, we argue that the capability of token mixers to obtain global contextual information hinges on multiple information interactions, such as spatial and channel domains. Subsequently, we propose Convolutional Additive Token Mixer (CATM) employing underlying spatial and channel attention as novel interaction forms. This module eliminates troublesome complex operations such as matrix multiplication and Softmax. We introduce Convolutional Additive Self-attention(CAS) block hybrid architecture and utilize CATM for each block. And further, we build a family of lightweight networks, which can be easily extended to various downstream tasks. Finally, we evaluate CAS-ViT across a variety of vision tasks, including image classification, object detection, instance segmentation, and semantic segmentation. Our M and T model achieves 83.0%/84.1% top-1 with only 12M/21M parameters on ImageNet-1K. Meanwhile, throughput evaluations on GPUs, ONNX, and iPhones also demonstrate superior results compared to other state-of-the-art backbones. Extensive experiments demonstrate that our approach achieves a better balance of performance, efficient inference and easy-to-deploy. Our code and model are available at: https://github.com/Tianfang-Zhang/CAS-ViT.
Tianfang Zhang, Wentao Liu 0002, Chen Qian 0006, Jenq-Neng Hwang, Xiangyang Ji
IEEE Trans. Image Process.1
2025 The Role of Deductive and Inductive Reasoning in Large Language Models
abstract
Chengkun Cai, Xu Zhao, Haoliang Liu, Zhongyu Jiang, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang, Lei Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chengkun Cai, Haoliang Liu, Zhongyu Jiang, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang, Lei Li 0050
ACL (1)5
2025 Harnessing Vital Sign Vibration Harmonics for Effortless and Inbuilt XR User Authentication
abstract
Extended Reality (XR) headsets are increasingly serving as repositories for substantial volumes of sensitive data and gateways to web applications. This transition highlights the need for convenient and secure user authentication solutions. Traditional password/PIN-based schemes are ill-suited to the XR's gesture- and voice-based interfaces and are prone to shoulder-surfing attacks. Some recent XR systems incorporate two-factor authentication, but it requires additional operations on a second device (e.g., a smartphone or wearable). In this work, we introduce the first effortless and inbuilt XR user authentication system by leveraging the harmonics of vibrations excited by users' vital signs. The system is transparent to users (no efforts during enrollment and authentication) and requires no additional hardware. The key idea is that vital signs (i.e., breathing and heart beating) naturally generate low-frequency mechanical vibrations, causing human skull to vibrate and produces harmonic signals. When the harmonics pass the human head, they carry rich biometrics associated with the wearer's skull structure and soft tissues, which can be captured by the XR motion sensors. Instead of directly utilizing the vibrations, we extract more reliable biometrics from the ratios among different harmonic frequencies, which capture wearers' unique head and facial attenuation properties and are non-volatile when the periodicity and amplitude of vital signs fluctuate. We further design an adaptive filter to mitigate the body motion distortions in common XR interactions. By adopting advanced deep learning models with the attention mechanism, our system realizes effective and robust authentication across XR scenarios. Evaluations across 10 months, with 52 users and two popular XR headsets, show that our system can accurately authenticate users with over 95% true positive rates and rejects unauthorized users with over 98% true negative rates under various XR scenarios, with biometrics remaining consistent over long-term periods.
Tianfang Zhang, Qiufan Ji, Md Mojibur Rahman Redoy Akanda, Zhengkun Ye, Ahmed Tanvir Mahdad, Cong Shi 0004, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001
CCS1
2025 Human Motion Instruction Tuning
abstract
This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model’s ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction.
Lei Li 0050, Sen Jia 0003, Zhongyu Jiang, Feng Zhou 0007, Ju Dai, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang
CVPR7
2025 Exploring Cross-Environment modeling and Robustness in Palm-based User Authentication using mmWave Testbed
abstract
Reliable and ubiquitous user authentication has become essential in smart cities, connected vehicles, and smart homes where users interact with multiple devices in their daily lives. However, existing biometric approaches, such as fingerprint, facial, or voice recognition, often require expensive hardware intrusive interaction, or raise privacy concerns, limiting their scalability in everyday settings [1–3]. To address these limitations, we explore a millimeter-wave (mmWave) testbed that enables palm-based user authentication through fine-grained sensing of palm geometry, skin thickness, and surface texture. By leveraging the widespread integration of mmWave technology in WiGig and 5G, this approach provides a low-cost, contactless, and privacy-preserving alternative to conventional biometrics. This work presents how the mmWave testbed is utilized to investigate cross-environment modeling and robustness in palm-based user authentication. Our system, named mmPalm, captures the reflections of Frequency-Modulated Continuous Wave (FMCW) signals from a user's palm to construct a distinctive palm profile that represents both structural and material characteristics of the hand. These reflections contain rich information about the three-dimensional geometry of the palm, sub-surface tissue variations, and fine surface textures, allowing unique identification without visual or physical contact. The mmWave testbed allows us to systematically collect palm data under varied distances, angles, and environments, providing a consistent platform for model development and evaluation.
Yucheng Xie, Xiaonan Guo 0003, Yan Wang 0003, Jerry Q. Cheng, Tianfang Zhang, Yingying Chen 0001
SEC5
2025 Passive Vital Sign Monitoring via Facial Vibrations Extracted from AR/VR Vibration Sensing Based Testbed
abstract
The adoption of augmented reality/virtual reality (AR/VR) has dramatically risen over the past few years across various application sectors, including immersive gaming, social communication, education, and tourism. The emerging use of AR/VR headsets has also created an excellent opportunity to promote pervasive health monitoring service as most AR/VR devices are already equipped with enriched sensing paradigm and will interact with users for a long time. In this talk, we aim to explore innovative technologies that enable fine-grained and personalized health status monitoring (e.g. vital signs and user identities) leveraging facial vibrations captured by the in-built motion sensor testbed on commodity AR/VR headsets. On one hand, it provides real-time health information required in virtual healthcare applications. For instance, a doctor can continuously monitor a patient's vital signs during the tele-medicine session at home, which helps the doctor to realize timely and precise diagnoses [2]. On the other hand, as people are spending increasing time in cyberspace (e.g., Metaverse), exposure to virtual and immersive contents requires high concentration on users' mind. Such usage cases may significantly increase the visual and psychological burden and induce potential health issues (e.g., anxiety, hypertension, sleep disorders) [1, 3, 5].
Tianfang Zhang, Cong Shi 0004, Payton Walker, Zhengkun Ye, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001
SEC1
2025 Saliency at the Helm: Steering Infrared Small Target Detection With Learnable Kernels
abstract
Infrared small target detection (ISTD) boasts extensive applications across civil and military domains, owing to its exceptional all-day performance. Neural network innovations have led to deep ISTD models that achieve heightened accuracy through extensive datasets. However, these general networks often fail to perceive the sensitivity of small targets and adopt heavy constructions to preserve potential target features, neglecting domain-specific insights and suffering from poor explainability. Our work seeks to rectify this by revisiting the saliency principles inherent to ISTD and developing a learnable local saliency kernel network (L2SKNet). This approach implements a learnable local saliency kernel module (LLSKM) that embodies the concept of “Center subtracts Neighbors,” guiding the network to capture the saliency features (points or edges). We enhance LLSKM by incorporating strategic dilation and structuring it hierarchically, which boosts its capability to capture multiscale infrared features while avoiding parameter explosion. In pursuit of efficiency, we also refine LLSKM into a more compact form by factorizing it into two orthogonal 1-D kernels, yielding a lightweight version. Heatmap visualizations and rigorous quantitative analyses corroborate the effectiveness of our local saliency-guided networks. Comprehensive testing reveals that L2SKNet variants outperform established baselines, demonstrating significant improvements in both visual and numerical assessments. The code is available athttps://github.com/fengyiwu98/L2SKNet.
Fengyi Wu, Tianfang Zhang, Junhai Luo, Zhenming Peng
IEEE Trans. Geosci. Remote. Sens.3
2024 SAFARI: Speech-Associated Facial Authentication for AR/VR Settings via Robust VIbration Signatures
abstract
In AR/VR devices, the voice interface, serving as one of the primary AR/VR control mechanisms, enables users to interact naturally using speeches (voice commands) for accessing data, controlling applications, and engaging in remote communication/meetings. Voice authentication can be adopted to protect against unauthorized speech inputs. However, existing voice authentication mechanisms are usually susceptible to voice spoofing attacks and are unreliable under the variations of phonetic content. In this work, we propose SAFARI, a spoofing-resistant and text-independent speech authentication system that can be seamlessly integrated into AR/VR voice interfaces. The key idea is to elicit phonetic-invariant biometrics from the facial muscle vibrations upon the headset. During speech production, a user's facial muscles are deformed for articulating phoneme sounds. The facial deformations associated with the phonemes are referred to as visemes. They carry rich biometrics of the wearer's muscles, tissue, and bones, which can propagate through the head and vibrate the headset. SAFARI aims to derive reliable facial biometrics from the viseme-associated facial vibrations captured by the AR/VR motion sensors. Particularly, it identifies the vibration data segments that contain rich viseme patterns (prominent visemes) less susceptible to phonetic variations. Based on the prominent visemes, SAFARI learns on the correlations among facial vibrations of different frequencies to extract biometric representations invariant to the phonetic context. The key advantages of SAFARI are that it is suitable for commodity AR/VR headsets (no additional sensors) and is resistant to voice spoofing attacks as the conductive property of the facial vibrations prevents biometric disclosure via the air media or the audio channel. To mitigate the impacts of body motions in AR/VR scenarios, we also design a generative diffusion model trained to reconstruct the viseme patterns from the data distorted by motion artifacts. We conduct extensive experiments with two representative AR/VR headsets and 35 users under various usage and attack settings. We demonstrate that SAFARI can achieve over 96% true positive rate on verifying legitimate users while successfully rejecting different kinds of spoofing attacks with over 97% true negative rates.
Tianfang Zhang, Qiufan Ji, Zhengkun Ye, Md Mojibur Rahman Redoy Akanda, Ahmed Tanvir Mahdad, Cong Shi 0004, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001
CCS1
2024 Clean and Compact: Efficient Data-Free Backdoor Defense with Model Compactness
Huy Phan, Jinqi Xiao, Yang Sui 0001, Tianfang Zhang, Zijie Tang, Cong Shi 0004, Yan Wang 0003, Yingying Chen 0001, Bo Yuan 0001
ECCV (60)4
2024 Palm-Based User Authentication Through mmWave
abstract
Biometric authentication systems are increasingly needed across a broad range of applications including in smart city environments (e.g., entering hotels, high-rise buildings, train stations, hospitals, and personalizing vehicles settings), and in smart home environments (e.g., controlling smart devices, en-hancing VR/AR experience). Traditional methods, such as face-based and fingerprint-based authentication, usually incur high cost to be installed in all this kind of environments, making them hard to become a ubiquitous authentication approach. In this paper, we develop a ubiquitous low-effort user authentication approach based on palm recognition using millimeter wave (mmWave) signals. Extensive experiments demonstrate that our system achieves 99% authentication accuracy.
Yucheng Xie, Tianfang Zhang, Xiaonan Guo 0003, Yan Wang 0003, Jerry Q. Cheng, Yingying Chen 0001
ICDCS2
2024 Inaudible Backdoor Attack via Stealthy Frequency Trigger Injection in Audio Spectrogram
abstract
Deep learning-enabled Voice User Interfaces (VUIs) have surpassed human-level performance in acoustic perception tasks. However, the significant cost associated with training these models compels users to rely on third-party data or outsource training services. Such emerging trends have drawn substantial attention to training-phase attacks, particularly backdoor attacks. Such attacks implant hidden trigger patterns (e.g., tones, environmental sounds) into the model during training, thereby manipulating the model's predictions in the inference phase. However, existing backdoor attacks can be easily undermined in practice as the inserted triggers are audible. Users may notice such attacks when listening to the training data and remaining alert for suspicious sounds. In this work, we present a novel audio backdoor attack that exploits completely inaudible triggers in the frequency domain of the audio spectrograms. Specifically, we optimize the trigger to be a frequency-domain pattern with the energy below the noise floor (e.g., background and hardware noises) at any given frequency, thereby rendering the trigger inaudible. To realize such attacks, we design a strategy that automatically generates inaudible triggers in the spectrum supported by commodity playback devices (e.g., smartphones and laptops). We further develop optimization techniques to enhance the trigger's robustness against speech content and onset variations. Experiments on hotword and speaker recognition indicate that our attack can achieve attack success rates of more than 98.2% and 81.0% under digital and physical attack scenarios. The results also demonstrate the trigger's inaudibility with a Signal-to-Noise Ratio (SNR) less than -3.54 dB against background noises. We further verify that our attack can successfully bypass state-of-the-art backdoor defense strategies based on learning and audio processing.
Tianfang Zhang, Huy Phan, Zijie Tang, Cong Shi 0004, Yan Wang 0003, Bo Yuan 0001, Yingying Chen 0001
MobiCom1
2024 RPCANet: Deep Unfolding RPCA Based Infrared Small Target Detection
abstract
Deep learning (DL) networks have achieved remarkable performance in infrared small target detection (ISTD). However, these structures exhibit a deficiency in interpretability and are widely regarded as black boxes, as they disregard domain knowledge in ISTD. To alleviate this issue, this work proposes an interpretable deep network for detecting infrared dim targets, dubbed RPCANet. Specifically, our approach formulates the ISTD task as sparse target extraction, low-rank background estimation, and image reconstruction in a relaxed Robust Principle Component Analysis (RPCA) model. By unfolding the iterative optimization updating steps into a deep-learning framework, time-consuming and complex matrix calculations are replaced by theory-guided neural networks. RPCANet detects targets with clear interpretability and preserves the intrinsic image feature, instead of directly transforming the detection task into a matrix decomposition problem. Extensive experiments substantiate the effectiveness of our deep unfolding framework and demonstrate its trustworthy results, surpassing baseline methods in both qualitative and quantitative evaluations. Our source code is available at https://github.com/fengyiwu98/RPCANet.
Fengyi Wu, Tianfang Zhang, Lei Li 0050, Yian Huang, Zhenming Peng
WACV2
2024 RF Domain Backdoor Attack on Signal Classification via Stealthy Trigger
abstract
Deep learning (DL) has recently become a key technology supporting radio frequency (RF) signal classification applications. Given the heavy DL training requirement, adopting outsourced training is a practical option for RF application developers. However, the outsourcing process exposes a security vulnerability that enables a backdoor attack. While backdoor attacks have been explored in the vision domain, it is rarely explored in the RF domain. In this work, we present a stealthy backdoor attack that targets DL-based RF signal classification. To realize such an attack, we extensively explore the characteristics of the RF data in different applications, which include RF modulation classification and RF fingerprint-based device identification. Then, we design a training-based backdoor trigger generation approach with different optimization procedures for two backdoor attack scenarios (i.e., poison-label and clean-label). Extensive experiments on two RF signal classification datasets show that the attack success rate is over 99.2%, while its classification accuracy for the clean data remains high (i.e., less than a 0.6% drop compared to the clean model). The low NMSE (less than 0.091) indicates the stealthiness of the attack. Additionally, we demonstrate that our attack can bypass existing defense strategies, such as Neural Cleanse and STRIP.
Zijie Tang, Tianming Zhao 0001, Tianfang Zhang, Huy Phan, Yan Wang 0003, Cong Shi 0004, Bo Yuan 0001, Yingying Chen 0001
IEEE Trans. Mob. Comput.3
2023 Privacy Leakage via Speech-induced Vibrations on Room Objects through Remote Sensing based on Phased-MIMO
abstract
Speech eavesdropping has long been an important threat to the privacy of individuals and enterprises. Recent research has shown the possibility of deriving private speech information from sound-induced vibrations. Acoustic signals transmitted through a solid medium or air may induce vibrations upon solid surfaces, which can be picked up by various sensors (e.g., motion sensors, high-speed cameras and lasers), without using a microphone. To date, these threats are limited to scenarios where the sensor is in contact with the vibration surface or at least in the visual line-of-sight.
Cong Shi 0004, Tianfang Zhang, Donglin Gao, Athina P. Petropulu, Chung-Tse Michael Wu, Yingying Chen 0001
CCS2
2023 FaceReader: Unobtrusively Mining Vital Signs and Vital Sign Embedded Sensitive Info via AR/VR Motion Sensors
abstract
The market size of augmented reality and virtual reality (AR/VR) has been expanding rapidly in recent years, with the use of face-mounted headsets extending beyond gaming to various application sectors, such as education, healthcare, and the military. Despite the rapid growth, the understanding of information leakage through sensor-rich headsets remains in its infancy. Some of the headset's built-in sensors do not require users' permission to access, and any apps and websites can acquire their readings. While theseunrestricted sensors are generally considered free of privacy risks, we find that an adversary could uncover private information by scrutinizing sensor readings, making existing AR/VR apps and websites potential eavesdroppers. In this work, we investigate a novel, unobtrusive privacy attack called FaceReader, which reconstructs high-quality vital sign signals (breathing and heartbeat patterns) based on unrestricted AR/VR motion sensors. FaceReader is built on the key insight that the headset is closely mounted on the user's face, allowing the motion sensors to detect subtle facial vibrations produced by users' breathing and heartbeats. Based on the reconstructed vital signs, we further investigate three more advanced attacks, including gender recognition, user re-identification, and body fat ratio estimation. Such attacks pose severe privacy concerns, as an adversary may obtain users' sensitive demographic/physiological traits and potentially uncover their real-world identities. Compared to prior privacy attacks relying on speeches and activities, FaceReader targets spontaneous breathing and heartbeat activities that are naturally produced by the human body and are unobtrusive to victims. In particular, we design an adaptive filter to dynamically mitigate the impacts of body motions. We further employ advanced deep-learning techniques to reconstruct vital sign signals, achieving signal qualities comparable to those of dedicated medical instruments, as well as deriving sensitive gender, identity, and body fat information. We conduct extensive experiments involving 35 users on three types of mainstream AR/VR headsets across 3 months. The results reveal that FaceReader can reconstruct vital signs with low mean errors and accurately detect gender (over 93.33%). The attack can also link/re-identify users across different apps, websites, and longitudinal sessions with over 97.83% accuracy. Furthermore, we present the first successful attempt at revealing body fat information from motion sensor data, achieving a remarkably low estimation error of 4.43%.
Tianfang Zhang, Zhengkun Ye, Ahmed Tanvir Mahdad, Md Mojibur Rahman Redoy Akanda, Cong Shi 0004, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001
CCS1
2023 Stealthy Backdoor Attack on RF Signal Classification
abstract
Recently, deep learning (DL) has become one of the key technologies supporting radio frequency (RF) signal classification applications. Given the heavy DL training requirement, adopting outsourced training is a practical option for RF application developers. However, the outsourcing process exposes a security vulnerability that enables a backdoor attack. While backdoor attacks have been explored in the computer vision domain, it is rarely explored in the RF domain. In this work, we present a stealthy backdoor attack that targets DL-based RF signal classification. To realize such an attack, we extensively explore the characteristics of the RF data in different applications, which include RF modulation classification and RF fingerprint-based device identification. Particularly, we design a training-based backdoor trigger generation approach with an optimization procedure that not only accommodates dynamic application inputs but also is stealthy to RF receivers. Extensive experiments on two RF signal classification datasets show that the average attack success rate of our backdoor attack is over 99.2%, while its classification accuracy for the clean data remains high (i.e., less than a 0.6% drop compared to the clean model). Additionally, we demonstrate that our attack can bypass existing defense strategies, such as Neural Cleanse and STRIP.
Tianming Zhao 0001, Zijie Tang, Tianfang Zhang, Huy Phan, Yan Wang 0003, Cong Shi 0004, Bo Yuan 0001, Yingying Chen 0001
ICCCN3
2023 Poster: Extracting Speech from Subtle Room Object Vibrations Using Remote mmWave Sensing
abstract
Speech privacy leakage has long been a public concern. Existing non-microphone-based eavesdropping attacks rely on physical contact or line-of-sight between the sensor (e.g., a motion sensor or a radar) and the victim sound source. In this poster, we investigate a new form of attack that remotely elicits speech from minute surface vibrations upon common room objects (e.g., paper bags, plastic storage bin) via mmWave sensing. We design and implement a highresolution software-defined phased-MIMO radar that integrates transmit beamforming, virtual array, and receive beamforming. The proposed system enhances sensing directivity by focusing all the mmWave beams toward a target room object. We successfully demonstrate such an attack by developing a deep speech recognition scheme grounded on unsupervised domain adaptation. Without prior training on the victim's data, our attack can achieve a high success rate of over 90% in recognizing simple digits.
Cong Shi 0004, Tianfang Zhang, Donglin Gao, Athina P. Petropulu, Chung-Tse Michael Wu, Yingying Chen 0001
MobiHoc2
2023 Poster: Unobtrusively Mining Vital Sign and Embedded Sensitive Info via AR/VR Motion Sensors
abstract
Despite the rapid growth of augmented reality and virtual reality (AR/VR) in various applications, the understanding of information leakage through sensor-rich headsets remains in its infancy. In this poster, we investigate an unobtrusive privacy attack, which exposes users' vital signs and embedded sensitive information (e.g., gender, identity, body fat ratio), based on unrestricted AR/VR motion sensors. The key insight is that the headset is closely mounted on the user's face, allowing the motion sensors to detect facial vibrations produced by users' breathing and heartbeats. Specifically, we employ deep-learning techniques to reconstruct vital signs, achieving signal qualities comparable to dedicated medical instruments, as well as deriving users' gender, identity, and body fat information. Experiments on three types of commodity AR/VR headsets reveal that our attack can successfully reconstruct high-quality vital signs, detect gender (accuracy over 93.33%), re-identify users (accuracy over 97.83%), and derive body fat ratio (error less than 4.43%).
Tianfang Zhang, Zhengkun Ye, Ahmed Tanvir Mahdad, Md Mojibur Rahman Redoy Akanda, Cong Shi 0004, Nitesh Saxena, Yan Wang 0003, Yingying Chen 0001
MobiHoc1
2023 Passive Vital Sign Monitoring via Facial Vibrations Leveraging AR/VR Headsets
abstract
Vital signs (e.g., breathing and heart rates) and personal identities are essential information for personalized medicine and healthcare. The popularity of augmented reality/virtual reality (AR/VR) provides an excellent opportunity for enabling long-term health monitoring in a broad range of scenarios, including virtual entertainment, education, and telemedicine. However, commercial-off-the-shelf AR/VR devices do not have dedicated biosensors for providing vital signs and personal identities. In this work, we propose a novel framework that can generate fine-grained vital sign signals and other personalized health information of an AR/VR user through passive sensing on AR/VR devices. In particular, we find that the user's minute facial vibrations induced by breathing and heart beating can impact the readily available motion sensors on AR/VR headsets, which encode rich vital sign patterns and unique biometrics. The proposed framework further estimates the breathing and heartbeat rates, detects the gender and identity, and derives the body fat percentage of the user. To mitigate the impacts of body movement, we design an adaptive filtering scheme to cancel the spontaneous and non-spontaneous motion artifacts. We also develop unique facial vibration features and deep learning techniques to facilitate vital sign signal reconstruction and user identification. Extensive experiments demonstrate that our framework can achieve a low error of vital sign signal reconstruction and rate measurement, along with 95.51% and 93.33% accuracy on identity and gender recognition.
Tianfang Zhang, Cong Shi 0004, Payton Walker, Zhengkun Ye, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001
MobiSys1
2023 Privacy Leakage via Unrestricted Motion-Position Sensors in the Age of Virtual Reality: A Study of Snooping Typed Input on Virtual Keyboards
abstract
Virtual Reality (VR) has gained popularity in numerous fields, including gaming, social interactions, shopping, and education. In this paper, we conduct a comprehensive study to assess the trustworthiness of the embedded sensors on VR, which embed various forms of sensitive data that may put users’ privacy at risk. We find that accessing most on-board sensors (e.g., motion, position, and button sensors) on VR SDKs/APIs, such as OpenVR, Oculus Platform, and WebXR, requires no security permission, exposing a huge attack surface for an adversary to steal the user’s privacy. We validate this vulnerability through developing malware programs and malicious websites and specifically explore to what extent it exposes the user’s information in the context of keystroke snooping. To examine its actual threat in practice, the adversary in the considered attack model doesn’t possess any labeled data from the user nor knowledge about the user’s VR settings. Extensive experiments, involving two mainstream VR systems and four keyboards with different typing mechanisms, demonstrate that our proof-of-concept attack can recognize the user’s virtual typing with over 89.7% accuracy. The attack can recover the user’s passwords with up to 84.9% recognition accuracy if three attempts are allowed and achieve an average of 87.1% word recognition rate for paragraph inference. We hope this study will help the community gain awareness of the vulnerability in the sensor management of current VR systems and provide insights to facilitate the future design of more comprehensive and restricted sensor access control mechanisms.
Yi Wu 0020, Cong Shi 0004, Tianfang Zhang, Payton Walker, Jian Liu 0001, Nitesh Saxena, Yingying Chen 0001
SP3
2023 BarrierBypass: Out-of-Sight Clean Voice Command Injection Attacks through Physical Barriers
abstract
The growing adoption of voice-enabled devices (e.g., smart speakers), particularly in smart home environments, has introduced many security vulnerabilities that pose significant threats to users' privacy and safety. When multiple devices are connected to a voice assistant, an attacker can cause serious damage if they can gain control of these devices. We ask where and how can an attacker issue clean voice commands stealthily across a physical barrier, and perform the first academic measurement study of this nature on the command injection attack. We present the BarrierBypass attack that can be launched against three different barrier-based scenarios termed across-door, across-window, and across-wall. We conduct a broad set of experiments to observe the command injection attack success rates for multiple speaker samples (TTS and live human recorded) at different command audio volumes (65, 75, 85 dB), and smart speaker locations (0.1-4.0m from barrier). Against Amazon Echo Dot 2, BarrierBypass is able to achieve 100% wake word and command injection success for the across-wall and across-window attacks, and for the across-door attack (up to 2 meters). At 4 meters for the across-door attack, BarrierBypass can achieve 90% and 80% injection accuracy for the wake word and command, respectively. Against Google Home mini BarrierBypass is able to achieve 100% wake word injection accuracy for all attack scenarios. For command injection BarrierBypass can achieve 100% accuracy for all the three barrier settings (up to 2 meters). For the across-door attack at 4 meters, BarrierBypass can achieve 80% command injection accuracy. Further, our demonstration using drones yielded high command injection success, up to 100%. Overall, our results demonstrate the potentially devastating nature of this vulnerability to control a user's device from outside of the device's physical space, and its limitations, without the need for complex and error-prone command injection.
Payton Walker, Tianfang Zhang, Cong Shi 0004, Nitesh Saxena, Yingying Chen 0001
WISEC2
2023 Mask-FPAN: Semi-supervised face parsing in the wild with de-occlusion and UV GAN
Lei Li 0050, Tianfang Zhang, Zhongfeng Kang, Xikun Jiang
Comput. Graph.2
2023 Optimization-inspired Cumulative Transmission Network for image compressive sensing
Tianfang Zhang, Lei Li 0050, Zhenming Peng
Knowl. Based Syst.1
2022 RIBAC: Towards Robust and Imperceptible Backdoor Attack against Compact DNN
Huy Phan, Cong Shi 0004, Yi Xie 0001, Tianfang Zhang, Tianming Zhao 0001, Jian Liu 0001, Yan Wang 0003, Yingying Chen 0001, Bo Yuan 0001
ECCV (4)4
2022 Audio-domain position-independent backdoor attack via unnoticeable triggers
abstract
Deep learning models have become key enablers of voice user interfaces. With the growing trend of adopting outsourced training of these models, backdoor attacks, stealthy yet effective training-phase attacks, have gained increasing attention. They inject hidden trigger patterns through training set poisoning and overwrite the model's predictions in the inference phase. Research in backdoor attacks has been focusing on image classification tasks, while there have been few studies in the audio domain. In this work, we explore the severity of audio-domain backdoor attacks and demonstrate their feasibility under practical scenarios of voice user interfaces, where an adversary injects (plays) an unnoticeable audio trigger into live speech to launch the attack. To realize such attacks, we consider jointly optimizing the audio trigger and the target model in the training phase, deriving a position-independent, unnoticeable, and robust audio trigger. We design new data poisoning techniques and penalty-based algorithms that inject the trigger into randomly generated temporal positions in the audio input during training, rendering the trigger resilient to any temporal position variations. We further design an environmental sound mimicking technique to make the trigger resemble unnoticeable situational sounds and simulate played over-the-air distortions to improve the trigger's robustness during the joint optimization process. Extensive experiments on two important applications (i.e., speech command recognition and speaker recognition) demonstrate that our attack can achieve an average success rate of over 99% under both digital and physical attack settings.
Cong Shi 0004, Tianfang Zhang, Huy Phan, Tianming Zhao 0001, Yan Wang 0003, Jian Liu 0001, Bo Yuan 0001, Yingying Chen 0001
MobiCom2
2022 Continuous blood pressure monitoring using low-cost motion sensors on AR/VR headsets
abstract
The Augmented reality/Virtual reality (AR/VR) industry has ushered in a period of rapid development. The next decade leaves a massive imagination for AR/VR in terms of end product form, software, content, applications, and user increment. The AR & VR technology offers a gazillion of possibilities for smart healthcare. In this poster, we develop an innovative continuous blood pressure (CBP) estimation system leveraging the built-in motion sensors of AR/VR headsets for users. We design a deep learning-based PPG construction scheme using the motion sensor-based cardiac signal and estimate the continuous blood pressure using the regression model. Our experimental results show that our system can continuously estimate both systolic blood pressure (SBP) and diastolic blood pressure (DBP) with a mean error of less than 4 mmHg and 0.9 mmHg respectively within a day.
Tianming Zhao 0001, Zhengkun Ye, Tianfang Zhang, Cong Shi 0004, Ahmed Tanvir Mahdad, Yan Wang 0003, Yingying Chen 0001, Nitesh Saxena
MobiSys3
2022 Speech privacy attack via vibrations from room objects leveraging a phased-MIMO radar
abstract
Speech privacy leakage has long been a public concern. Through speech eavesdropping, an adversary may steal a user's private information or an enterprise's financial/intellectual properties, leading to catastrophic consequences. Existing non-microphone-based eavesdropping attacks rely on physical contact or line-of-sight between the sensor (e.g., a motion sensor or a radar) and the victim sound source. In this poster, we discover a new form of speech eavesdropping attack that senses minor speech-induced vibrations upon common room objects using mmWave. By integrating phasedarray and multiple-input and multiple-output (MIMO) on a single mmWave transceiver, our attack can capture and fuse micrometerlevel vibrations upon the surfaces of multiple objects to reveal speech content in a remote and non-line-of-sight fashion. We successfully demonstrate such an attack by developing a deep speech recognition scheme grounded on unsupervised domain adaptation. Without prior training on the victim's data, our attack can achieve a high success rate of over 90% in recognizing simple speech content.
Cong Shi 0004, Tianfang Zhang, Yichao Yuan, Athina P. Petropulu, Chung-Tse Michael Wu, Yingying Chen 0001
MobiSys2
2022 Personalized health monitoring via vital sign measurements leveraging motion sensors on AR/VR headsets
abstract
Augmented reality/virtual reality (AR/VR) headsets have attracted millions of users and gained predictable popularity. However, long-period usage of immersive technology may lead to health issues (e.g., cybersickness, anxiety). In this poster, we design a low-cost and personalized healthcare monitoring system grounded on vital sign tracking (i.e., breathing and heartbeat rate tracking), by exploiting built-in AR/VR motion sensors. The key insight is that the conductive vibrations induced by chest and heart movements can propagate through the user's cranial bones, thereby vibrating the AR/VR headset mounted on the user's head. To realize this system, we design signal processing techniques to cancel the human motions and derive the periods of breathing and heartbeat through frequency-domain analyses. We further design a user identification scheme based on respiratory and cardiac biometrics, which works with vital sign monitoring to provide personalized healthcare recommendations. Our experiment shows that the proposed scheme can achieve less than 5.7% error rate on breathing/heartbeat rate estimation and 95% accuracy on user identification.
Tianfang Zhang, Cong Shi 0004, Tianming Zhao 0001, Zhengkun Ye, Payton Walker, Nitesh Saxena, Yan Wang 0003, Yingying Chen 0001
MobiSys1
2021 Robust Detection of Machine-induced Audio Attacks in Intelligent Audio Systems with Microphone Array
abstract
With the popularity of intelligent audio systems in recent years, their vulnerabilities have become an increasing public concern. Existing studies have designed a set of machine-induced audio attacks, such as replay attacks, synthesis attacks, hidden voice commands, inaudible attacks, and audio adversarial examples, which could expose users to serious security and privacy threats. To defend against these attacks, existing efforts have been treating them individually. While they have yielded reasonably good performance in certain cases, they can hardly be combined into an all-in-one solution to be deployed on the audio systems in practice. Additionally, modern intelligent audio devices, such as Amazon Echo and Apple HomePod, usually come equipped with microphone arrays for far-field voice recognition and noise reduction. Existing defense strategies have been focusing on single- and dual-channel audio, while only few studies have explored using multi-channel microphone array for defending specific types of audio attack. Motivated by the lack of systematic research on defending miscellaneous audio attacks and the potential benefits of multi-channel audio, this paper builds a holistic solution for detecting machine-induced audio attacks leveraging multi-channel microphone arrays on modern intelligent audio systems. Specifically, we utilize magnitude and phase spectrograms of multi-channel audio to extract spatial information and leverage a deep learning model to detect the fundamental difference between human speech and adversarial audio generated by the playback machines. Moreover, we adopt an unsupervised domain adaptation training framework to further improve the model's generalizability in new acoustic environments. Evaluation is conducted under various settings on a public multi-channel replay attack dataset and a self-collected multi-channel audio attack dataset involving 5 types of advanced audio attacks. The results show that our method can achieve an equal error rate (EER) as low as 6.6% in detecting a variety of machine-induced attacks. Even in new acoustic environments, our method can still achieve an EER as low as 8.8%.
Cong Shi 0004, Tianfang Zhang, Yi Xie 0001, Jian Liu 0001, Bo Yuan 0001, Yingying Chen 0001
CCS3
2021 Environment-independent In-baggage Object Identification Using WiFi Signals
abstract
Low-cost in-baggage object identification is highly demanded in enhancing public safety and smart manufacturing. Existing approaches usually require specialized equipment and heavy deployment overhead, making them hard to scale for wide deployment. The recent WiFi-based approach is unsuitable for practical deployment as it did not address dynamic environmental impacts. In this work, we propose an environment-independent in-baggage object identification system by leveraging low-cost WiFi. We exploit the channel state information (CSI) to capture material and shape characteristics to facilitate fine-grained inbaggage object identification. A major challenge of building such a system is that CSI measurements are sensitive to real-world dynamics, such as different types of baggage, time-varying ambient noises and interferences, and different deployment environments. To tackle these problems, we develop WiFi features based on polarized directional antennas that can capture objects’ material and shape characteristics. A convolutional neural network-based model is developed to constructively integrate the WiFi features and perform accurate in-baggage object identification. We also develop a material-based domain adaptation using adversarial learning to facilitate fast deployments in different environments. We conduct extensive experiments involving 14 representation objects, 4 types of bags in 3 different room environments. The results show that our system can achieve over 97% in the same environment, and our domain adaptation method can improve the object identification accuracy by 42% when the system is deployed in a new environment with little training.
Cong Shi 0004, Tianming Zhao 0001, Yucheng Xie, Tianfang Zhang, Yan Wang 0003, Xiaonan Guo 0003, Yingying Chen 0001
MASS4
2021 Face-Mic: inferring live speech and speaker identity via subtle facial dynamics captured by AR/VR motion sensors
abstract
Augmented reality/virtual reality (AR/VR) has extended beyond 3D immersive gaming to a broader array of applications, such as shopping, tourism, education. And recently there has been a large shift from handheld-controller dominated interactions to headset-dominated interactions via voice interfaces. In this work, we show a serious privacy risk of using voice interfaces while the user is wearing the face-mounted AR/VR devices. Specifically, we design an eavesdropping attack, Face-Mic, which leverages speech-associated subtle facial dynamics captured by zero-permission motion sensors in AR/VR headsets to infer highly sensitive information from live human speech, including speaker gender, identity, and speech content. Face-Mic is grounded on a key insight that AR/VR headsets are closely mounted on the user's face, allowing a potentially malicious app on the headset to capture underlying facial dynamics as the wearer speaks, including movements of facial muscles and bone-borne vibrations, which encode private biometrics and speech characteristics. To mitigate the impacts of body movements, we develop a signal source separation technique to identify and separate the speech-associated facial dynamics from other types of body movements. We further extract representative features with respect to the two types of facial dynamics. We successfully demonstrate the privacy leakage through AR/VR headsets by deriving the user's gender/identity and extracting speech information via the development of a deep learning-based framework. Extensive experiments using four mainstream VR headsets validate the generalizability, effectiveness, and high accuracy of Face-Mic.
Cong Shi 0004, Xiangyu Xu 0001, Tianfang Zhang, Payton Walker, Yi Wu 0020, Jian Liu 0001, Nitesh Saxena, Yingying Chen 0001, Jiadi Yu
MobiCom3
2021 Infrared small target detection via self-regularized weighted sparse model
Tianfang Zhang, Zhenming Peng, Hao Wu 0043, Yanmin He, Chaohai Li, Chunping Yang
Neurocomputing1