VLDB 2026 Research / reviewers in the wild / expert
Cong Shi 0004
dblp:07/6946-4
· DBLP profile ↗
44ranked-venue papers
12as first author
37since 2021 · last 2026
0000-0002-7599-202XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 23 · 9 first-author · 18 since 2021Security and privacy · 9 · 2 first-author · 8 since 2021Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Solving Scarce Wireless Signal Dilemma in Model Training using Cross-Modal Learning Leveraging Limited Video Data
Qiufan Ji, Honglu Li, Cong Shi 0004, Yan Wang 0003, Jerry Q. Cheng, Yingying Chen 0001 |
MobiSys | 3 |
| 2026 | Solving the WiFi Sensing Dilemma in Reality Leveraging Conformal PredictionabstractWith the extensive deployment of smart environments and IoT devices, WiFi sensing has proven its significant convenience and contact-free sensing capabilities in supporting a wide range of applications. However, designing a ubiquitous WiFi sensing system for diverse real-world scenarios presents a substantial dilemma, as the system performance deteriorates when the testing data diverges significantly from the training data due to domain variations. To address this dilemma, existing studies need extra efforts to develop new features or even retrain the original model under environmental variations. However, these approaches have not efficiently resolved the dilemma. In this study, we conduct a comprehensive study on the domain variation problem to make WiFi sensing robust and accurate in practical applications. Our definition of domains is comprehensive and includes environmental conditions, surrounding settings, user differences, user orientations, user's positions relative to WiFi sensors, and user participation time frames. We design a novel conformal prediction framework that quantifies the conformity (i.e., similarity) between the testing and training WiFi samples, then labels the testing samples with the most probable class(es). Unlike traditional conformal prediction which relies on data from a single domain, we develop a new statistical (Type I) approach to assess the conformity of the testing WiFi samples to individual training domains and aggregate the outcomes. To further improve the framework's generalization, we design a (Type II) fusion approach that utilizes the inter-relationships among domains for more accurate conformity quantification. Built upon these two methods, kernel density estimation-based and SVM-based methods are developed to compute the conformity scores for new testing samples to make conformal predictions. Extensive experiments, utilizing both self-collected and publicly available datasets show that our framework can improve prediction accuracies ranging from 20.1% to 77.1% in three of the most representative WiFi-based applications across six types of domain variations. Honglu Li, Qiufan Ji, Cong Shi 0004, Yan Wang 0003, Jerry Q. Cheng, Kailong Wang 0003, Min-ge Xie, Yingying Chen 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Harnessing Vital Sign Vibration Harmonics for Effortless and Inbuilt XR User AuthenticationabstractExtended Reality (XR) headsets are increasingly serving as repositories for substantial volumes of sensitive data and gateways to web applications. This transition highlights the need for convenient and secure user authentication solutions. Traditional password/PIN-based schemes are ill-suited to the XR's gesture- and voice-based interfaces and are prone to shoulder-surfing attacks. Some recent XR systems incorporate two-factor authentication, but it requires additional operations on a second device (e.g., a smartphone or wearable). In this work, we introduce the first effortless and inbuilt XR user authentication system by leveraging the harmonics of vibrations excited by users' vital signs. The system is transparent to users (no efforts during enrollment and authentication) and requires no additional hardware. The key idea is that vital signs (i.e., breathing and heart beating) naturally generate low-frequency mechanical vibrations, causing human skull to vibrate and produces harmonic signals. When the harmonics pass the human head, they carry rich biometrics associated with the wearer's skull structure and soft tissues, which can be captured by the XR motion sensors. Instead of directly utilizing the vibrations, we extract more reliable biometrics from the ratios among different harmonic frequencies, which capture wearers' unique head and facial attenuation properties and are non-volatile when the periodicity and amplitude of vital signs fluctuate. We further design an adaptive filter to mitigate the body motion distortions in common XR interactions. By adopting advanced deep learning models with the attention mechanism, our system realizes effective and robust authentication across XR scenarios. Evaluations across 10 months, with 52 users and two popular XR headsets, show that our system can accurately authenticate users with over 95% true positive rates and rejects unauthorized users with over 98% true negative rates under various XR scenarios, with biometrics remaining consistent over long-term periods. Tianfang Zhang, Qiufan Ji, Md Mojibur Rahman Redoy Akanda, Zhengkun Ye, Ahmed Tanvir Mahdad, Cong Shi 0004, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001 |
CCS | 6 |
| 2025 | Fine-grained Vital Sign Reconstruction through Machine Learning on Multi-channel Radar SignalsabstractMonitoring vital signs such as breathing rate (BR) and heart rate (HR) is crucial for early detection of health issues and supports a wide range of health-related applications. Traditional monitoring methods often involve body-attached medical devices, which can be intrusive and inconvenient for continuous use in daily life. Contactless monitoring using radio frequency (RF) signals has emerged as a promising alternative, but acquiring precise vital sign measurements remains challenging due to the limited sensing resolution of RF devices. In this paper, we design a high-resolution contactless vital sensing system by leveraging advanced beamforming in combination with machine learning (ML) methods. The key idea of our system is to reconstruct fine-grained vital sign measurements from RF signals, achieving low estimation error, comparable to that of dedicated medical devices such as photoplethysmography sensors, respiration monitoring belts. To enhance the reconstruction performance, we integrate an antenna array with double phase shifters to acquire RF data that captures precise chest displacement of human subjects. An encoder-decoder model based on a 1D convolutional neural network is then developed to map the RF signals into vital sign measurements. Extensive evaluations show that our system has low errors of 0.3 beat per minute (BPM) for BR estimation and 2.7 BPM for HR estimation. Cong Shi 0004, Athina P. Petropulu, Yingying Chen 0001 |
ICASSP | 2 |
| 2025 | Robust Point Cloud Recognition Model SharingabstractWith the rapid development of mobile and edge-integrated sensing technologies, 3D point clouds have emerged as a fundamental data modality for understanding and interacting with the physical world. They provide rich spatial and geometric information that enables autonomous driving, mobile robotics, and intelligent IoT devices to perceive and reason about their surroundings in real time. In particular, the development of edge-deployed sensors, such as LiDAR, depth cameras, and structured-light sensors, has made it feasible to capture and process 3D point clouds directly at the network edge, empowering low-latency perception and decision-making for safety-critical systems. Qiufan Ji, Lin Wang 0025, Cong Shi 0004, Shengshan Hu, Yingying Chen 0001, Lichao Sun 0001 |
SEC | 3 |
| 2025 | Mobile Edge Testbed for Driving Behavior Data Collection and Cognitive Impairment AnalysisabstractStudying cognitive impairment and its impact on driving behaviors is crucial for enhancing public safety. To facilitate cognitive impairment studies, we devise a testbed for realworld driving data collection using ubiquitous mobile edge devices (i.e., smartphones) [4]. Toward this end, we develop an application for autonomous data collection using smartphones. To enable robust data collection in real-world driving scenarios, we design a coordinate alignment method that automatically aligns the smartphone's coordinate system with the vehicle's by continuously detecting stationary and straight-line acceleration periods. We also design a two-step segmentation algorithm that first utilizes gyroscope readings to segment rotation-based behaviors (e.g., turning) and then employs accelerometer data to segment non-rotation-based behaviors (e.g., braking). The processed data is then uploaded to a cloud server through WiFi connections for further analysis. Honglu Li, Cong Shi 0004, Yan Wang 0003, Tammy Chung, Yingying Chen 0001 |
SEC | 3 |
| 2025 | VR Testbed-based Blood Pressure Privacy Leakage AnalysisabstractBlood pressure (BP) is one of the most essential biomarkers for human health, widely used to diagnose cardiovascular diseases [3] and assess mental states [2, 5]. It is considered Protected Health Information (PHI) under HIPAA, and access to it typically requires explicit user consent. In this work, we uncover a novel privacy breach in the metaverse usage: a user's private BP information can be covertly and continuously surveilled using the unrestricted in-built motion sensors present in commodity VR headsets. Zhengkun Ye, Ahmed Tanvir Mahdad, Yan Wang 0003, Cong Shi 0004, Yingying Chen 0001, Nitesh Saxena |
SEC | 4 |
| 2025 | Passive Vital Sign Monitoring via Facial Vibrations Extracted from AR/VR Vibration Sensing Based TestbedabstractThe adoption of augmented reality/virtual reality (AR/VR) has dramatically risen over the past few years across various application sectors, including immersive gaming, social communication, education, and tourism. The emerging use of AR/VR headsets has also created an excellent opportunity to promote pervasive health monitoring service as most AR/VR devices are already equipped with enriched sensing paradigm and will interact with users for a long time. In this talk, we aim to explore innovative technologies that enable fine-grained and personalized health status monitoring (e.g. vital signs and user identities) leveraging facial vibrations captured by the in-built motion sensor testbed on commodity AR/VR headsets. On one hand, it provides real-time health information required in virtual healthcare applications. For instance, a doctor can continuously monitor a patient's vital signs during the tele-medicine session at home, which helps the doctor to realize timely and precise diagnoses [2]. On the other hand, as people are spending increasing time in cyberspace (e.g., Metaverse), exposure to virtual and immersive contents requires high concentration on users' mind. Such usage cases may significantly increase the visual and psychological burden and induce potential health issues (e.g., anxiety, hypertension, sleep disorders) [1, 3, 5]. Tianfang Zhang, Cong Shi 0004, Payton Walker, Zhengkun Ye, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001 |
SEC | 2 |
| 2025 | BPSniff: Continuously Surveilling Private Blood Pressure Information in the Metaverse via Unrestricted Inbuilt Motion SensorsabstractBlood pressure (BP) is one of the most essential biomarkers for various diseases. It is considered protected health information under HIPAA and usually needs the user's consent for access. In this work, we uncover an insidious privacy breach in metaverse usage: private BP information can be covertly obtained from unrestricted motion sensors in virtual reality (VR) headsets. The insight is that the motion sensors can capture the subtle vibrations induced by the blood waves in the major arteries. Such vibrations are highly correlated with users' cardiac cycles and BP. As adversaries can continuously obtain motion sensor data from VR headsets without users' consent, they can derive and collect users' BP information in metaverse apps or websites, leading to more severe consequences, such as discrimination, exploitation, and targeted harassment. To demonstrate this severe privacy leakage in the meta-verse, we develop a practical attack, BPSniff, which can reconstruct fine-grained blood flow patterns and derive BP based on motion sensor data from users' VR headsets. BP-Sniff is the first practical attack revealing the BP leakage in the metaverse without using dedicated equipment. Unlike previous mobile sensing approaches that require user-specific calibration, BPSniff bypasses this constraint, enabling truly stealthy passive BP attacks at scale. Our attack first employs a variational autoencoder to reconstruct high-fidelity blood flow patterns from VR headset motion sensor data. We then develop an Adam-optimized long short-term memory (LSTM) regression model that leverages BP-related fiducial features from successive blood flow patterns to continuously estimate the user's BP. We evaluate BPSniff through extensive experiments and a longitudinal study of 8 weeks, involving 37 participants and two VR headset models. The results show that BPSniff can achieve low mean errors of 1.75 mmHg for systolic blood pressure (SBP) and 1.34 mmHg for diastolic blood pressure (DBP), which are comparable to commercial BP monitors and satisfy the standard (i.e., mean error ≤ 5.0 mmHg) specified by FDA's AAMI protocol. Zhengkun Ye, Ahmed Tanvir Mahdad, Yan Wang 0003, Cong Shi 0004, Yingying Chen 0001, Nitesh Saxena |
SP | 4 |
| 2024 | SAFARI: Speech-Associated Facial Authentication for AR/VR Settings via Robust VIbration SignaturesabstractIn AR/VR devices, the voice interface, serving as one of the primary AR/VR control mechanisms, enables users to interact naturally using speeches (voice commands) for accessing data, controlling applications, and engaging in remote communication/meetings. Voice authentication can be adopted to protect against unauthorized speech inputs. However, existing voice authentication mechanisms are usually susceptible to voice spoofing attacks and are unreliable under the variations of phonetic content. In this work, we propose SAFARI, a spoofing-resistant and text-independent speech authentication system that can be seamlessly integrated into AR/VR voice interfaces. The key idea is to elicit phonetic-invariant biometrics from the facial muscle vibrations upon the headset. During speech production, a user's facial muscles are deformed for articulating phoneme sounds. The facial deformations associated with the phonemes are referred to as visemes. They carry rich biometrics of the wearer's muscles, tissue, and bones, which can propagate through the head and vibrate the headset. SAFARI aims to derive reliable facial biometrics from the viseme-associated facial vibrations captured by the AR/VR motion sensors. Particularly, it identifies the vibration data segments that contain rich viseme patterns (prominent visemes) less susceptible to phonetic variations. Based on the prominent visemes, SAFARI learns on the correlations among facial vibrations of different frequencies to extract biometric representations invariant to the phonetic context. The key advantages of SAFARI are that it is suitable for commodity AR/VR headsets (no additional sensors) and is resistant to voice spoofing attacks as the conductive property of the facial vibrations prevents biometric disclosure via the air media or the audio channel. To mitigate the impacts of body motions in AR/VR scenarios, we also design a generative diffusion model trained to reconstruct the viseme patterns from the data distorted by motion artifacts. We conduct extensive experiments with two representative AR/VR headsets and 35 users under various usage and attack settings. We demonstrate that SAFARI can achieve over 96% true positive rate on verifying legitimate users while successfully rejecting different kinds of spoofing attacks with over 97% true negative rates. Tianfang Zhang, Qiufan Ji, Zhengkun Ye, Md Mojibur Rahman Redoy Akanda, Ahmed Tanvir Mahdad, Cong Shi 0004, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001 |
CCS | 6 |
| 2024 | Clean and Compact: Efficient Data-Free Backdoor Defense with Model Compactness
Huy Phan, Jinqi Xiao, Yang Sui 0001, Tianfang Zhang, Zijie Tang, Cong Shi 0004, Yan Wang 0003, Yingying Chen 0001, Bo Yuan 0001 |
ECCV (60) | 6 |
| 2024 | Practical Adversarial Attack on WiFi Sensing Through Unnoticeable Communication Packet PerturbationabstractThe pervasive use of WiFi has driven the recent research in WiFi sensing, converting communication tech into sensing for applications such as activity recognition, user authentication, and vital sign monitoring. Despite the integration of deep learning into WiFi sensing systems, potential security vulnerabilities to adversarial attacks remain unexplored. This paper introduces the first physical attack focusing on deep learning-based WiFi sensing systems, demonstrating how adversaries can subtly manipulate WiFi packet preambles to affect channel state information (CSI), a critical feature in such systems, and thereby influence underlying deep learning models without disrupting regular communication. To realize the proposed attack in practical scenarios, we rigorously analyze and derive the intricate relationship between the pilot symbol and CSI. A novel mechanism is proposed to facilitate quantitive control of receiver-side CSI through minimal modifications to the pilot symbols of WiFi packets at the transmitter. We further develop a perturbation optimization method based on the Carlini & Wagner (CW) attack and a penalty-based training process to ensure the attack's universal efficacy across various CSI responses and noise. The physical attack is implemented and evaluated in two representative WiFi sensing systems (i.e., activity recognition and user authentication) with 35 participants over 3 months. Extensive experiments demonstrate the remarkable attack success rates of 90.47% and 83.83% for activity recognition and user authentication, respectively. Mingjing Xu, Yicong Du, Cong Shi 0004, Yan Wang 0003, Hongbo Liu 0002, Yingying Chen 0001 |
MobiCom | 5 |
| 2024 | Inaudible Backdoor Attack via Stealthy Frequency Trigger Injection in Audio SpectrogramabstractDeep learning-enabled Voice User Interfaces (VUIs) have surpassed human-level performance in acoustic perception tasks. However, the significant cost associated with training these models compels users to rely on third-party data or outsource training services. Such emerging trends have drawn substantial attention to training-phase attacks, particularly backdoor attacks. Such attacks implant hidden trigger patterns (e.g., tones, environmental sounds) into the model during training, thereby manipulating the model's predictions in the inference phase. However, existing backdoor attacks can be easily undermined in practice as the inserted triggers are audible. Users may notice such attacks when listening to the training data and remaining alert for suspicious sounds. In this work, we present a novel audio backdoor attack that exploits completely inaudible triggers in the frequency domain of the audio spectrograms. Specifically, we optimize the trigger to be a frequency-domain pattern with the energy below the noise floor (e.g., background and hardware noises) at any given frequency, thereby rendering the trigger inaudible. To realize such attacks, we design a strategy that automatically generates inaudible triggers in the spectrum supported by commodity playback devices (e.g., smartphones and laptops). We further develop optimization techniques to enhance the trigger's robustness against speech content and onset variations. Experiments on hotword and speaker recognition indicate that our attack can achieve attack success rates of more than 98.2% and 81.0% under digital and physical attack scenarios. The results also demonstrate the trigger's inaudibility with a Signal-to-Noise Ratio (SNR) less than -3.54 dB against background noises. We further verify that our attack can successfully bypass state-of-the-art backdoor defense strategies based on learning and audio processing. Tianfang Zhang, Huy Phan, Zijie Tang, Cong Shi 0004, Yan Wang 0003, Bo Yuan 0001, Yingying Chen 0001 |
MobiCom | 4 |
| 2024 | RF Domain Backdoor Attack on Signal Classification via Stealthy TriggerabstractDeep learning (DL) has recently become a key technology supporting radio frequency (RF) signal classification applications. Given the heavy DL training requirement, adopting outsourced training is a practical option for RF application developers. However, the outsourcing process exposes a security vulnerability that enables a backdoor attack. While backdoor attacks have been explored in the vision domain, it is rarely explored in the RF domain. In this work, we present a stealthy backdoor attack that targets DL-based RF signal classification. To realize such an attack, we extensively explore the characteristics of the RF data in different applications, which include RF modulation classification and RF fingerprint-based device identification. Then, we design a training-based backdoor trigger generation approach with different optimization procedures for two backdoor attack scenarios (i.e., poison-label and clean-label). Extensive experiments on two RF signal classification datasets show that the attack success rate is over 99.2%, while its classification accuracy for the clean data remains high (i.e., less than a 0.6% drop compared to the clean model). The low NMSE (less than 0.091) indicates the stealthiness of the attack. Additionally, we demonstrate that our attack can bypass existing defense strategies, such as Neural Cleanse and STRIP. Zijie Tang, Tianming Zhao 0001, Tianfang Zhang, Huy Phan, Yan Wang 0003, Cong Shi 0004, Bo Yuan 0001, Yingying Chen 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | Privacy Leakage via Speech-induced Vibrations on Room Objects through Remote Sensing based on Phased-MIMOabstractSpeech eavesdropping has long been an important threat to the privacy of individuals and enterprises. Recent research has shown the possibility of deriving private speech information from sound-induced vibrations. Acoustic signals transmitted through a solid medium or air may induce vibrations upon solid surfaces, which can be picked up by various sensors (e.g., motion sensors, high-speed cameras and lasers), without using a microphone. To date, these threats are limited to scenarios where the sensor is in contact with the vibration surface or at least in the visual line-of-sight. Cong Shi 0004, Tianfang Zhang, Donglin Gao, Athina P. Petropulu, Chung-Tse Michael Wu, Yingying Chen 0001 |
CCS | 1 |
| 2023 | FaceReader: Unobtrusively Mining Vital Signs and Vital Sign Embedded Sensitive Info via AR/VR Motion SensorsabstractThe market size of augmented reality and virtual reality (AR/VR) has been expanding rapidly in recent years, with the use of face-mounted headsets extending beyond gaming to various application sectors, such as education, healthcare, and the military. Despite the rapid growth, the understanding of information leakage through sensor-rich headsets remains in its infancy. Some of the headset's built-in sensors do not require users' permission to access, and any apps and websites can acquire their readings. While theseunrestricted sensors are generally considered free of privacy risks, we find that an adversary could uncover private information by scrutinizing sensor readings, making existing AR/VR apps and websites potential eavesdroppers. In this work, we investigate a novel, unobtrusive privacy attack called FaceReader, which reconstructs high-quality vital sign signals (breathing and heartbeat patterns) based on unrestricted AR/VR motion sensors. FaceReader is built on the key insight that the headset is closely mounted on the user's face, allowing the motion sensors to detect subtle facial vibrations produced by users' breathing and heartbeats. Based on the reconstructed vital signs, we further investigate three more advanced attacks, including gender recognition, user re-identification, and body fat ratio estimation. Such attacks pose severe privacy concerns, as an adversary may obtain users' sensitive demographic/physiological traits and potentially uncover their real-world identities. Compared to prior privacy attacks relying on speeches and activities, FaceReader targets spontaneous breathing and heartbeat activities that are naturally produced by the human body and are unobtrusive to victims. In particular, we design an adaptive filter to dynamically mitigate the impacts of body motions. We further employ advanced deep-learning techniques to reconstruct vital sign signals, achieving signal qualities comparable to those of dedicated medical instruments, as well as deriving sensitive gender, identity, and body fat information. We conduct extensive experiments involving 35 users on three types of mainstream AR/VR headsets across 3 months. The results reveal that FaceReader can reconstruct vital signs with low mean errors and accurately detect gender (over 93.33%). The attack can also link/re-identify users across different apps, websites, and longitudinal sessions with over 97.83% accuracy. Furthermore, we present the first successful attempt at revealing body fat information from motion sensor data, achieving a remarkably low estimation error of 4.43%. Tianfang Zhang, Zhengkun Ye, Ahmed Tanvir Mahdad, Md Mojibur Rahman Redoy Akanda, Cong Shi 0004, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001 |
CCS | 5 |
| 2023 | Stealthy Backdoor Attack on RF Signal ClassificationabstractRecently, deep learning (DL) has become one of the key technologies supporting radio frequency (RF) signal classification applications. Given the heavy DL training requirement, adopting outsourced training is a practical option for RF application developers. However, the outsourcing process exposes a security vulnerability that enables a backdoor attack. While backdoor attacks have been explored in the computer vision domain, it is rarely explored in the RF domain. In this work, we present a stealthy backdoor attack that targets DL-based RF signal classification. To realize such an attack, we extensively explore the characteristics of the RF data in different applications, which include RF modulation classification and RF fingerprint-based device identification. Particularly, we design a training-based backdoor trigger generation approach with an optimization procedure that not only accommodates dynamic application inputs but also is stealthy to RF receivers. Extensive experiments on two RF signal classification datasets show that the average attack success rate of our backdoor attack is over 99.2%, while its classification accuracy for the clean data remains high (i.e., less than a 0.6% drop compared to the clean model). Additionally, we demonstrate that our attack can bypass existing defense strategies, such as Neural Cleanse and STRIP. Tianming Zhao 0001, Zijie Tang, Tianfang Zhang, Huy Phan, Yan Wang 0003, Cong Shi 0004, Bo Yuan 0001, Yingying Chen 0001 |
ICCCN | 6 |
| 2023 | Benchmarking and Analyzing Robust Point Cloud Recognition: Bag of Tricks for Defending Adversarial ExamplesabstractDeep Neural Networks (DNNs) for 3D point cloud recognition are vulnerable to adversarial examples, threatening their practical deployment. Despite the many research endeavors have been made to tackle this issue in recent years, the diversity of adversarial examples on 3D point clouds makes them more challenging to defend against than those on 2D images. For examples, attackers can generate adversarial examples by adding, shifting, or removing points. Consequently, existing defense strategies are hard to counter unseen point cloud adversarial examples. In this paper, we first establish a comprehensive, and rigorous point cloud adversarial robustness benchmark to evaluate adversarial robustness, which can provide a detailed understanding of the effects of the defense and attack methods. We then collect existing defense tricks in point cloud adversarial defenses and then perform extensive and systematic experiments to identify an effective combination of these tricks. Furthermore, we propose a hybrid training augmentation methods that consider various types of point cloud adversarial examples to adversarial training, significantly improving the adversarial robustness. By combining these tricks, we construct a more robust defense framework achieving an average accuracy of 83.45% against various attacks, demonstrating its capability to enabling robust learners. Our codebase are open-sourced on: https://github.com/qiufan319/benchmark_pc_attack.git. Qiufan Ji, Lin Wang 0025, Cong Shi 0004, Shengshan Hu, Yingying Chen 0001, Lichao Sun 0001 |
ICCV | 3 |
| 2023 | EmoLeak: Smartphone Motions Reveal EmotionsabstractEmotional state leakage attracts increasing concerns as it reveals rich sensitive information, such as intent, demo graphic, personality, and health information. Existing emotion recognition techniques rely on vision and audio data, which have limited threat due to the requirements of accessing restricted sensors (e.g., cameras and microphones). In this work, we first investigate the feasibility of detecting the emotional state of people in the vibration domain via zero-permission motion sensors. We find that when voice is being played through a smartphone's loudspeaker or ear speaker, it generates vibration signals on the smartphone surface, which encodes rich emotional information. As the smartphone is the go-to device for almost everyone nowadays, our attack based only on motion sensors raises severe concerns about emotion state leakage. We comprehensively study the relationship between vibration data and human emotion based on several publicly available emotion datasets (e.g., SAVEE, TESS). Time-frequency features and machine learning techniques are developed to determine the emotion of the victim based on speech vibrations. We evaluate our attack on both the ear speakers and loudspeakers on a diverse set of smartphones. The results demonstrate our attack can achieve a high accuracy, with around 95.3% (random guess 14.3%) accuracy for the loudspeaker setting and 60.52% (random guess 14.3%) accuracy for the ear speaker setting. Ahmed Tanvir Mahdad, Cong Shi 0004, Zhengkun Ye, Tianming Zhao 0001, Yan Wang 0003, Yingying Chen 0001, Nitesh Saxena |
ICDCS | 2 |
| 2023 | Poster: Extracting Speech from Subtle Room Object Vibrations Using Remote mmWave SensingabstractSpeech privacy leakage has long been a public concern. Existing non-microphone-based eavesdropping attacks rely on physical contact or line-of-sight between the sensor (e.g., a motion sensor or a radar) and the victim sound source. In this poster, we investigate a new form of attack that remotely elicits speech from minute surface vibrations upon common room objects (e.g., paper bags, plastic storage bin) via mmWave sensing. We design and implement a highresolution software-defined phased-MIMO radar that integrates transmit beamforming, virtual array, and receive beamforming. The proposed system enhances sensing directivity by focusing all the mmWave beams toward a target room object. We successfully demonstrate such an attack by developing a deep speech recognition scheme grounded on unsupervised domain adaptation. Without prior training on the victim's data, our attack can achieve a high success rate of over 90% in recognizing simple digits. Cong Shi 0004, Tianfang Zhang, Donglin Gao, Athina P. Petropulu, Chung-Tse Michael Wu, Yingying Chen 0001 |
MobiHoc | 1 |
| 2023 | Poster: Unobtrusively Mining Vital Sign and Embedded Sensitive Info via AR/VR Motion SensorsabstractDespite the rapid growth of augmented reality and virtual reality (AR/VR) in various applications, the understanding of information leakage through sensor-rich headsets remains in its infancy. In this poster, we investigate an unobtrusive privacy attack, which exposes users' vital signs and embedded sensitive information (e.g., gender, identity, body fat ratio), based on unrestricted AR/VR motion sensors. The key insight is that the headset is closely mounted on the user's face, allowing the motion sensors to detect facial vibrations produced by users' breathing and heartbeats. Specifically, we employ deep-learning techniques to reconstruct vital signs, achieving signal qualities comparable to dedicated medical instruments, as well as deriving users' gender, identity, and body fat information. Experiments on three types of commodity AR/VR headsets reveal that our attack can successfully reconstruct high-quality vital signs, detect gender (accuracy over 93.33%), re-identify users (accuracy over 97.83%), and derive body fat ratio (error less than 4.43%). Tianfang Zhang, Zhengkun Ye, Ahmed Tanvir Mahdad, Md Mojibur Rahman Redoy Akanda, Cong Shi 0004, Nitesh Saxena, Yan Wang 0003, Yingying Chen 0001 |
MobiHoc | 5 |
| 2023 | Passive Vital Sign Monitoring via Facial Vibrations Leveraging AR/VR HeadsetsabstractVital signs (e.g., breathing and heart rates) and personal identities are essential information for personalized medicine and healthcare. The popularity of augmented reality/virtual reality (AR/VR) provides an excellent opportunity for enabling long-term health monitoring in a broad range of scenarios, including virtual entertainment, education, and telemedicine. However, commercial-off-the-shelf AR/VR devices do not have dedicated biosensors for providing vital signs and personal identities. In this work, we propose a novel framework that can generate fine-grained vital sign signals and other personalized health information of an AR/VR user through passive sensing on AR/VR devices. In particular, we find that the user's minute facial vibrations induced by breathing and heart beating can impact the readily available motion sensors on AR/VR headsets, which encode rich vital sign patterns and unique biometrics. The proposed framework further estimates the breathing and heartbeat rates, detects the gender and identity, and derives the body fat percentage of the user. To mitigate the impacts of body movement, we design an adaptive filtering scheme to cancel the spontaneous and non-spontaneous motion artifacts. We also develop unique facial vibration features and deep learning techniques to facilitate vital sign signal reconstruction and user identification. Extensive experiments demonstrate that our framework can achieve a low error of vital sign signal reconstruction and rate measurement, along with 95.51% and 93.33% accuracy on identity and gender recognition. Tianfang Zhang, Cong Shi 0004, Payton Walker, Zhengkun Ye, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001 |
MobiSys | 2 |
| 2023 | Privacy Leakage via Unrestricted Motion-Position Sensors in the Age of Virtual Reality: A Study of Snooping Typed Input on Virtual KeyboardsabstractVirtual Reality (VR) has gained popularity in numerous fields, including gaming, social interactions, shopping, and education. In this paper, we conduct a comprehensive study to assess the trustworthiness of the embedded sensors on VR, which embed various forms of sensitive data that may put users’ privacy at risk. We find that accessing most on-board sensors (e.g., motion, position, and button sensors) on VR SDKs/APIs, such as OpenVR, Oculus Platform, and WebXR, requires no security permission, exposing a huge attack surface for an adversary to steal the user’s privacy. We validate this vulnerability through developing malware programs and malicious websites and specifically explore to what extent it exposes the user’s information in the context of keystroke snooping. To examine its actual threat in practice, the adversary in the considered attack model doesn’t possess any labeled data from the user nor knowledge about the user’s VR settings. Extensive experiments, involving two mainstream VR systems and four keyboards with different typing mechanisms, demonstrate that our proof-of-concept attack can recognize the user’s virtual typing with over 89.7% accuracy. The attack can recover the user’s passwords with up to 84.9% recognition accuracy if three attempts are allowed and achieve an average of 87.1% word recognition rate for paragraph inference. We hope this study will help the community gain awareness of the vulnerability in the sensor management of current VR systems and provide insights to facilitate the future design of more comprehensive and restricted sensor access control mechanisms. Yi Wu 0020, Cong Shi 0004, Tianfang Zhang, Payton Walker, Jian Liu 0001, Nitesh Saxena, Yingying Chen 0001 |
SP | 2 |
| 2023 | BarrierBypass: Out-of-Sight Clean Voice Command Injection Attacks through Physical BarriersabstractThe growing adoption of voice-enabled devices (e.g., smart speakers), particularly in smart home environments, has introduced many security vulnerabilities that pose significant threats to users' privacy and safety. When multiple devices are connected to a voice assistant, an attacker can cause serious damage if they can gain control of these devices. We ask where and how can an attacker issue clean voice commands stealthily across a physical barrier, and perform the first academic measurement study of this nature on the command injection attack. We present the BarrierBypass attack that can be launched against three different barrier-based scenarios termed across-door, across-window, and across-wall. We conduct a broad set of experiments to observe the command injection attack success rates for multiple speaker samples (TTS and live human recorded) at different command audio volumes (65, 75, 85 dB), and smart speaker locations (0.1-4.0m from barrier). Against Amazon Echo Dot 2, BarrierBypass is able to achieve 100% wake word and command injection success for the across-wall and across-window attacks, and for the across-door attack (up to 2 meters). At 4 meters for the across-door attack, BarrierBypass can achieve 90% and 80% injection accuracy for the wake word and command, respectively. Against Google Home mini BarrierBypass is able to achieve 100% wake word injection accuracy for all attack scenarios. For command injection BarrierBypass can achieve 100% accuracy for all the three barrier settings (up to 2 meters). For the across-door attack at 4 meters, BarrierBypass can achieve 80% command injection accuracy. Further, our demonstration using drones yielded high command injection success, up to 100%. Overall, our results demonstrate the potentially devastating nature of this vulnerability to control a user's device from outside of the device's physical space, and its limitations, without the need for complex and error-prone command injection. Payton Walker, Tianfang Zhang, Cong Shi 0004, Nitesh Saxena, Yingying Chen 0001 |
WISEC | 3 |
| 2022 | RIBAC: Towards Robust and Imperceptible Backdoor Attack against Compact DNN
Huy Phan, Cong Shi 0004, Yi Xie 0001, Tianfang Zhang, Tianming Zhao 0001, Jian Liu 0001, Yan Wang 0003, Yingying Chen 0001, Bo Yuan 0001 |
ECCV (4) | 2 |
| 2022 | Defending against Thru-barrier Stealthy Voice Attacks via Cross-Domain Sensing on Phoneme SoundsabstractThe open nature of voice input makes voice assistant (VA) systems vulnerable to various acoustic attacks (e.g., replay and voice synthesis attacks). A simple yet effective way for adversaries to launch these attacks is to hide behind barriers (e.g., a wall, a window, or a door) and give unauthorized voice commands without being observed by legitimate users. In this work, we develop an automated, training-free defense system that can protect VA systems from such thru-barrier acoustic attacks. Our study finds that acoustic signals passing through the barriers generally present a unique frequency-selective effect in the vibration domain. Thus, we propose to devise a system to capture this unique effect of barriers by leveraging low-cost, cross-domain sensing available in users’ wearables. The system replays the audio-domain signals with the wearable’s speaker and captures the conductive vibrations caused by the audio sounds in the vibration domain via the built-in accelerometer. To improve the proposed system’s reliability, we develop a unique vibration-domain enhancement method to extract the phonemes most sensitive to the frequency-selective effect of barriers. We identify effective vibration-domain features that capture the barriers’ effects in the vibration domain. A 2D-correlation-based method is developed to examine the speech similarity between the recordings from the VA system and the user’s wearable and detect thru-barrier attacks. Extensive experiments with various barriers and environments demonstrate that the proposed defense system can effectively defend random, replay, synthesis, and hidden voice attacks with less than 4% equal error rates. Cong Shi 0004, Tianming Zhao 0001, Ahmed Tanvir Mahdad, Zhengkun Ye, Yan Wang 0003, Nitesh Saxena, Yingying Chen 0001 |
ICDCS | 1 |
| 2022 | Audio-domain position-independent backdoor attack via unnoticeable triggersabstractDeep learning models have become key enablers of voice user interfaces. With the growing trend of adopting outsourced training of these models, backdoor attacks, stealthy yet effective training-phase attacks, have gained increasing attention. They inject hidden trigger patterns through training set poisoning and overwrite the model's predictions in the inference phase. Research in backdoor attacks has been focusing on image classification tasks, while there have been few studies in the audio domain. In this work, we explore the severity of audio-domain backdoor attacks and demonstrate their feasibility under practical scenarios of voice user interfaces, where an adversary injects (plays) an unnoticeable audio trigger into live speech to launch the attack. To realize such attacks, we consider jointly optimizing the audio trigger and the target model in the training phase, deriving a position-independent, unnoticeable, and robust audio trigger. We design new data poisoning techniques and penalty-based algorithms that inject the trigger into randomly generated temporal positions in the audio input during training, rendering the trigger resilient to any temporal position variations. We further design an environmental sound mimicking technique to make the trigger resemble unnoticeable situational sounds and simulate played over-the-air distortions to improve the trigger's robustness during the joint optimization process. Extensive experiments on two important applications (i.e., speech command recognition and speaker recognition) demonstrate that our attack can achieve an average success rate of over 99% under both digital and physical attack settings. Cong Shi 0004, Tianfang Zhang, Huy Phan, Tianming Zhao 0001, Yan Wang 0003, Jian Liu 0001, Bo Yuan 0001, Yingying Chen 0001 |
MobiCom | 1 |
| 2022 | Continuous blood pressure monitoring using low-cost motion sensors on AR/VR headsetsabstractThe Augmented reality/Virtual reality (AR/VR) industry has ushered in a period of rapid development. The next decade leaves a massive imagination for AR/VR in terms of end product form, software, content, applications, and user increment. The AR & VR technology offers a gazillion of possibilities for smart healthcare. In this poster, we develop an innovative continuous blood pressure (CBP) estimation system leveraging the built-in motion sensors of AR/VR headsets for users. We design a deep learning-based PPG construction scheme using the motion sensor-based cardiac signal and estimate the continuous blood pressure using the regression model. Our experimental results show that our system can continuously estimate both systolic blood pressure (SBP) and diastolic blood pressure (DBP) with a mean error of less than 4 mmHg and 0.9 mmHg respectively within a day. Tianming Zhao 0001, Zhengkun Ye, Tianfang Zhang, Cong Shi 0004, Ahmed Tanvir Mahdad, Yan Wang 0003, Yingying Chen 0001, Nitesh Saxena |
MobiSys | 4 |
| 2022 | Speech privacy attack via vibrations from room objects leveraging a phased-MIMO radarabstractSpeech privacy leakage has long been a public concern. Through speech eavesdropping, an adversary may steal a user's private information or an enterprise's financial/intellectual properties, leading to catastrophic consequences. Existing non-microphone-based eavesdropping attacks rely on physical contact or line-of-sight between the sensor (e.g., a motion sensor or a radar) and the victim sound source. In this poster, we discover a new form of speech eavesdropping attack that senses minor speech-induced vibrations upon common room objects using mmWave. By integrating phasedarray and multiple-input and multiple-output (MIMO) on a single mmWave transceiver, our attack can capture and fuse micrometerlevel vibrations upon the surfaces of multiple objects to reveal speech content in a remote and non-line-of-sight fashion. We successfully demonstrate such an attack by developing a deep speech recognition scheme grounded on unsupervised domain adaptation. Without prior training on the victim's data, our attack can achieve a high success rate of over 90% in recognizing simple speech content. Cong Shi 0004, Tianfang Zhang, Yichao Yuan, Athina P. Petropulu, Chung-Tse Michael Wu, Yingying Chen 0001 |
MobiSys | 1 |
| 2022 | Personalized health monitoring via vital sign measurements leveraging motion sensors on AR/VR headsetsabstractAugmented reality/virtual reality (AR/VR) headsets have attracted millions of users and gained predictable popularity. However, long-period usage of immersive technology may lead to health issues (e.g., cybersickness, anxiety). In this poster, we design a low-cost and personalized healthcare monitoring system grounded on vital sign tracking (i.e., breathing and heartbeat rate tracking), by exploiting built-in AR/VR motion sensors. The key insight is that the conductive vibrations induced by chest and heart movements can propagate through the user's cranial bones, thereby vibrating the AR/VR headset mounted on the user's head. To realize this system, we design signal processing techniques to cancel the human motions and derive the periods of breathing and heartbeat through frequency-domain analyses. We further design a user identification scheme based on respiratory and cardiac biometrics, which works with vital sign monitoring to provide personalized healthcare recommendations. Our experiment shows that the proposed scheme can achieve less than 5.7% error rate on breathing/heartbeat rate estimation and 95% accuracy on user identification. Tianfang Zhang, Cong Shi 0004, Tianming Zhao 0001, Zhengkun Ye, Payton Walker, Nitesh Saxena, Yan Wang 0003, Yingying Chen 0001 |
MobiSys | 2 |
| 2022 | Solving the WiFi Sensing Dilemma in Reality Leveraging Conformal PredictionabstractWith the wide deployment of smart environments and IoT devices, WiFi sensing has demonstrated its great convenience and contactless sensing capabilities in supporting a broad array of applications. However, designing a ubiquitous WiFi sensing system for heterogeneous scenarios in practice is still a big dilemma as the system performs poorly when the testing data is significantly different from the training data caused by domain variations. To address this dilemma, existing studies involve extra efforts to develop new features or even to retrain the original model under environmental variations. However, none of them can resolve the dilemma completely. In this work, we conduct a comprehensive study on the domain variation problem to make WiFi sensing robust and accurate in reality. Our definition of domains is comprehensive and includes environments, surrounding settings, user differences, user's facing directions, user's positions relative to WiFi sensors, and user participating time frames. Our innovation is to achieve reliable WiFi sensing across all the domains based on the conformal prediction framework. Our approach quantifies the conformity (i.e., similarity) between the testing WiFi samples and the training samples, then labels the testing samples with the most probable class(es). We develop a novel cross-domain transformal prediction scheme based on the multivariate kernel density estimation to effectively assess and learn the conformity of each domain in the training data. To meet various application-specific requirements, we further develop two approaches to fuse the knowledge of conformity derived from the training domains to perform predictions. Extensive experiments with both self-collected and public datasets show that our framework can improve prediction accuracies from 30% to 74% improvements in three most representative WiFi-based applications across six types of domain variations. Kailong Wang 0003, Cong Shi 0004, Jerry Q. Cheng, Yan Wang 0003, Min-ge Xie, Yingying Chen 0001 |
SenSys | 2 |
| 2022 | Enabling Secret Key Distribution Over Screen-to-Camera Channel Leveraging Color Shift PropertyabstractRecent years witnessed the emergence of visible light communication (VLC) over screen-to-camera channel, such as barcode and unobtrusive optical pattern, due to the widely adoption of screen and camera in plenty of electronic devices. The prevalence of wide viewing angle screen and high standard cameras also imposes great threat for visible light communication, and the information leakage over screen-to-camera channel has been rarely explored. In this paper, we propose a secret key distribution system leveraging the unique color shift property over screen-to-camera channel. To facilitate such design, two practical secret key distribution methods, key matching-based and nearest next hop-based, are developed to map the secret key into a unique optical pattern on screen, which can only be correctly decoded by the legitimate user situated at an accessible region. We also provide theoretical analysis on the security of both methods. The performance of the proposed system is implemented with off-the-shelf devices and validated under various experimental scenarios. The results demonstrate that our system can achieve high bit-decoding accuracy for the legitimate users while maintaining low recovery accuracy for the attackers. Hongbo Liu 0002, Cong Shi 0004, Yingying Chen 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2021 | Enabling Fast and Universal Audio Adversarial Attack Using Generative ModelabstractRecently, the vulnerability of deep neural network (DNN)-based audio systems to adversarial attacks has obtained increasing attention. However, the existing audio adversarial attacks allow the adversary to possess the entire user's audio input as well as granting sufficient time budget to generate the adversarial perturbations. These idealized assumptions, however, make the existing audio adversarial attacks mostly impossible to be launched in a timely fashion in practice (e.g., playing unnoticeable adversarial perturbations along with user's streaming input). To overcome these limitations, in this paper we propose fast audio adversarial perturbation generator (FAPG), which uses generative model to generate adversarial perturbations for the audio input in a single forward pass, thereby drastically improving the perturbation generation speed. Built on the top of FAPG, we further propose universal audio adversarial perturbation generator (UAPG), a scheme to craft universal adversarial perturbation that can be imposed on arbitrary benign audio input to cause misclassification. Extensive experiments on DNN-based audio systems show that our proposed FAPG can achieve high success rate with up to 214X speedup over the existing audio adversarial attack methods. Also our proposed UAPG generates universal adversarial perturbations that can achieve much better attack performance than the state-of-the-art solutions. Yi Xie 0001, Cong Shi 0004, Jian Liu 0001, Yingying Chen 0001, Bo Yuan 0001 |
AAAI | 3 |
| 2021 | Robust Detection of Machine-induced Audio Attacks in Intelligent Audio Systems with Microphone ArrayabstractWith the popularity of intelligent audio systems in recent years, their vulnerabilities have become an increasing public concern. Existing studies have designed a set of machine-induced audio attacks, such as replay attacks, synthesis attacks, hidden voice commands, inaudible attacks, and audio adversarial examples, which could expose users to serious security and privacy threats. To defend against these attacks, existing efforts have been treating them individually. While they have yielded reasonably good performance in certain cases, they can hardly be combined into an all-in-one solution to be deployed on the audio systems in practice. Additionally, modern intelligent audio devices, such as Amazon Echo and Apple HomePod, usually come equipped with microphone arrays for far-field voice recognition and noise reduction. Existing defense strategies have been focusing on single- and dual-channel audio, while only few studies have explored using multi-channel microphone array for defending specific types of audio attack. Motivated by the lack of systematic research on defending miscellaneous audio attacks and the potential benefits of multi-channel audio, this paper builds a holistic solution for detecting machine-induced audio attacks leveraging multi-channel microphone arrays on modern intelligent audio systems. Specifically, we utilize magnitude and phase spectrograms of multi-channel audio to extract spatial information and leverage a deep learning model to detect the fundamental difference between human speech and adversarial audio generated by the playback machines. Moreover, we adopt an unsupervised domain adaptation training framework to further improve the model's generalizability in new acoustic environments. Evaluation is conducted under various settings on a public multi-channel replay attack dataset and a self-collected multi-channel audio attack dataset involving 5 types of advanced audio attacks. The results show that our method can achieve an equal error rate (EER) as low as 6.6% in detecting a variety of machine-induced attacks. Even in new acoustic environments, our method can still achieve an EER as low as 8.8%. Cong Shi 0004, Tianfang Zhang, Yi Xie 0001, Jian Liu 0001, Bo Yuan 0001, Yingying Chen 0001 |
CCS | 2 |
| 2021 | Environment-independent In-baggage Object Identification Using WiFi SignalsabstractLow-cost in-baggage object identification is highly demanded in enhancing public safety and smart manufacturing. Existing approaches usually require specialized equipment and heavy deployment overhead, making them hard to scale for wide deployment. The recent WiFi-based approach is unsuitable for practical deployment as it did not address dynamic environmental impacts. In this work, we propose an environment-independent in-baggage object identification system by leveraging low-cost WiFi. We exploit the channel state information (CSI) to capture material and shape characteristics to facilitate fine-grained inbaggage object identification. A major challenge of building such a system is that CSI measurements are sensitive to real-world dynamics, such as different types of baggage, time-varying ambient noises and interferences, and different deployment environments. To tackle these problems, we develop WiFi features based on polarized directional antennas that can capture objects’ material and shape characteristics. A convolutional neural network-based model is developed to constructively integrate the WiFi features and perform accurate in-baggage object identification. We also develop a material-based domain adaptation using adversarial learning to facilitate fast deployments in different environments. We conduct extensive experiments involving 14 representation objects, 4 types of bags in 3 different room environments. The results show that our system can achieve over 97% in the same environment, and our domain adaptation method can improve the object identification accuracy by 42% when the system is deployed in a new environment with little training. Cong Shi 0004, Tianming Zhao 0001, Yucheng Xie, Tianfang Zhang, Yan Wang 0003, Xiaonan Guo 0003, Yingying Chen 0001 |
MASS | 1 |
| 2021 | Face-Mic: inferring live speech and speaker identity via subtle facial dynamics captured by AR/VR motion sensorsabstractAugmented reality/virtual reality (AR/VR) has extended beyond 3D immersive gaming to a broader array of applications, such as shopping, tourism, education. And recently there has been a large shift from handheld-controller dominated interactions to headset-dominated interactions via voice interfaces. In this work, we show a serious privacy risk of using voice interfaces while the user is wearing the face-mounted AR/VR devices. Specifically, we design an eavesdropping attack, Face-Mic, which leverages speech-associated subtle facial dynamics captured by zero-permission motion sensors in AR/VR headsets to infer highly sensitive information from live human speech, including speaker gender, identity, and speech content. Face-Mic is grounded on a key insight that AR/VR headsets are closely mounted on the user's face, allowing a potentially malicious app on the headset to capture underlying facial dynamics as the wearer speaks, including movements of facial muscles and bone-borne vibrations, which encode private biometrics and speech characteristics. To mitigate the impacts of body movements, we develop a signal source separation technique to identify and separate the speech-associated facial dynamics from other types of body movements. We further extract representative features with respect to the two types of facial dynamics. We successfully demonstrate the privacy leakage through AR/VR headsets by deriving the user's gender/identity and extracting speech information via the development of a deep learning-based framework. Extensive experiments using four mainstream VR headsets validate the generalizability, effectiveness, and high accuracy of Face-Mic. Cong Shi 0004, Xiangyu Xu 0001, Tianfang Zhang, Payton Walker, Yi Wu 0020, Jian Liu 0001, Nitesh Saxena, Yingying Chen 0001, Jiadi Yu |
MobiCom | 1 |
| 2021 | WiFi-Enabled User Authentication through Deep Learning in Daily ActivitiesabstractUser authentication is a critical process in both corporate and home environments due to the ever-growing security and privacy concerns. With the advancement of smart cities and home environments, the concept of user authentication is evolved with a broader implication by not only preventing unauthorized users from accessing confidential information but also providing the opportunities for customized services corresponding to a specific user. Traditional approaches of user authentication either require specialized device installation or inconvenient wearable sensor attachment. This article supports the extended concept of user authentication with a device-free approach by leveraging the prevalent WiFi signals made available by IoT devices, such as smart refrigerator, smart TV, and smart thermostat, and so on. The proposed system utilizes the WiFi signals to capture unique human physiological and behavioral characteristics inherited from their daily activities, including both walking and stationary ones. Particularly, we extract representative features from channel state information (CSI) measurements of WiFi signals, and develop a deep-learning-based user authentication scheme to accurately identify each individual user. To mitigate the signal distortion caused by surrounding people’s movements, our deep learning model exploits a CNN-based architecture that constructively combines features from multiple receiving antennas and derives more reliable feature abstractions. Furthermore, a transfer-learning-based mechanism is developed to reduce the training cost for new users and environments. Extensive experiments in various indoor environments are conducted to demonstrate the effectiveness of the proposed authentication system. In particular, our system can achieve over 94% authentication accuracy with 11 subjects through different activities. Cong Shi 0004, Jian Liu 0001, Hongbo Liu 0002, Yingying Chen 0001 |
ACM Trans. Internet Things | 1 |
| 2020 | WearID: Low-Effort Wearable-Assisted Authentication of Voice Commands via Cross-Domain Comparison without TrainingabstractDue to the open nature of voice input, voice assistant (VA) systems (e.g., Google Home and Amazon Alexa) are vulnerable to various security and privacy leakages (e.g., credit card numbers, passwords), especially when issuing critical user commands involving large purchases, critical calls, etc. Though the existing VA systems may employ voice features to identify users, they are still vulnerable to various acoustic-based attacks (e.g., impersonation, replay, and hidden command attacks). In this work, we propose a training-free voice authentication system, WearID, leveraging the cross-domain speech similarity between the audio domain and the vibration domain to provide enhanced security to the ever-growing deployment of VA systems. In particular, when a user gives a critical command, WearID exploits motion sensors on the user’s wearable device to capture the aerial speech in the vibration domain and verify it with the speech captured in the audio domain via the VA device’s microphone. Compared to existing approaches, our solution is low-effort and privacy-preserving, as it neither requires users’ active inputs (e.g., replying messages/calls) nor to store users’ privacy-sensitive voice samples for training. In addition, our solution exploits the distinct vibration sensing interface and its short sensing range to sound (e.g., 25cm) to verify voice commands. Examining the similarity of the two domains’ data is not trivial. The huge sampling rate gap (e.g., 8000Hz vs. 200Hz) between the audio and vibration domains makes it hard to compare the two domains’ data directly, and even tiny data noises could be magnified and cause authentication failures. To address the challenges, we investigate the complex relationship between the two sensing domains and develop a spectrogram-based algorithm to convert the microphone data into the lower-frequency “ motion sensor data” to facilitate cross-domain comparisons. We further develop a user authentication scheme to verify that the received voice command originates from the legitimate user based on the cross-domain speech similarity of the received voice commands. We report on extensive experiments to evaluate the WearID under various audible and inaudible attacks. The results show WearID can verify voice commands with 99.8% accuracy in the normal situation and detect 97.2% fake voice commands from various attacks, including impersonation/replay attacks and hidden voice/ultrasound attacks. Cong Shi 0004, Yan Wang 0003, Yingying Chen 0001, Nitesh Saxena, Chen Wang 0009 |
ACSAC | 1 |
| 2020 | Real-Time, Universal, and Robust Adversarial Attacks Against Speaker Recognition SystemsabstractAs the popularity of voice user interface (VUI) exploded in recent years, speaker recognition system has emerged as an important medium of identifying a speaker in many security-required applications and services. In this paper, we propose the first real-time, universal, and robust adversarial attack against the state-of-the-art deep neural network (DNN) based speaker recognition system. Through adding an audio-agnostic universal perturbation on arbitrary enrolled speaker's voice input, the DNN-based speaker recognition system would identify the speaker as any target (i.e., adversary-desired) speaker label. In addition, we improve the robustness of our attack by modeling the sound distortions caused by the physical over-the-air propagation through estimating room impulse response (RIR). Experiment using a public dataset of 109 English speakers demonstrates the effectiveness and robustness of our proposed attack with a high attack success rate of over 90%. The attack launching time also achieves a 100× speedup over contemporary non-universal attacks. Yi Xie 0001, Cong Shi 0004, Jian Liu 0001, Yingying Chen 0001, Bo Yuan 0001 |
ICASSP | 2 |
| 2020 | Mobile Device Usage Recommendation based on User Context Inference Using Embedded SensorsabstractThe proliferation of mobile devices along with their rich functionalities/applications have made people form addictive and potentially harmful usage behaviors. Though this problem has drawn considerable attention, existing solutions (e.g., text notification or setting usage limits) are insufficient and cannot provide timely recommendations or control of inappropriate usage of mobile devices. This paper proposes a generalized context inference framework, which supports timely usage recommendations using low-power sensors in mobile devices Comparing to existing schemes that rely on detection of single type user contexts (e.g., merely on location or activity), our framework derives a much larger-scale of user contexts that characterize the phone usages, especially those causing distraction or leading to dangerous situations. We propose to uniformly describe the general user context with context fundamentals, i.e., physical environments, social situations, and human motions, which are the underlying constituent units of diverse general user contexts. To mitigate the profiling efforts across different environments, devices, and individuals, we develop a deep learning-based architecture to learn transferable representations derived from sensor readings associated with the context fundamentals. Based on the derived context fundamentals, our framework quantifies how likely an inferred user context would lead to distractions/dangerous situations, and provides timely recommendations for mobile device access/usage. Extensive experiments during a period of 7 months demonstrate that the system can achieve 95% accuracy on user context inference while offering the transferability among different environments, devices, and users. Cong Shi 0004, Xiaonan Guo 0003, Ting Yu 0001, Yingying Chen 0001, Yucheng Xie, Jian Liu 0001 |
ICCCN | 1 |
| 2020 | Towards Environment-independent Behavior-based User Authentication Using WiFiabstractWith the increasing prevalence of smart mobile and Internet of things (IoT) environments, user authentication has become a critical component for not only preventing unauthorized access to security-sensitive systems but also providing customized services for individual users. Unlike traditional approaches relying on tedious passwords or specialized biometric/wearable sensors, this paper presents a device-free user authentication via daily human behavioral patterns captured by existing WiFi infrastructures. Specifically, our system exploits readily available channel state information (CSI) in WiFi signals to capture unique behavioral biometrics residing in the user’s daily activities, without requiring any dedicated sensors or wearable device attachment. To build such a system, one major challenge is that wireless signals always carry substantial information that is specific to the user’s location and surrounding environment, rendering the trained model less effective when being applied to the data collected in a new location or environment. This issue could lead to significant authentication errors and may quickly ruin the whole system in practice. To disentangle the behavioral biometrics for practical environment-independent user authentication, we propose an end-to-end deep-learning based approach with domain adaptation techniques to remove the environment-and location-specific information contained in the collected WiFi measurements. Extensive experiments in a residential apartment and an office with various scales of user location variations and environmental changes demonstrate the effectiveness and generalizability of the proposed authentication system. Cong Shi 0004, Jian Liu 0001, Nick Borodinov, Bruno Leão, Yingying Chen 0001 |
MASS | 1 |
| 2019 | CardioCam: Leveraging Camera on Mobile Devices to Verify Users While Their Heart is PumpingabstractWith the increasing prevalence of mobile and IoT devices (e.g., smartphones, tablets, smart-home appliances), massive private and sensitive information are stored on these devices. To prevent unauthorized access on these devices, existing user verification solutions either rely on the complexity of user-defined secrets (e.g., password) or resort to specialized biometric sensors (e.g., fingerprint reader), but the users may still suffer from various attacks, such as password theft, shoulder surfing, smudge, and forged biometrics attacks. In this paper, we propose, CardioCam, a low-cost, general, hard-to-forge user verification system leveraging the unique cardiac biometrics extracted from the readily available built-in cameras in mobile and IoT devices. We demonstrate that the unique cardiac features can be extracted from the cardiac motion patterns in fingertips, by pressing on the built-in camera. To mitigate the impacts of various ambient lighting conditions and human movements under practical scenarios, CardioCam develops a gradient-based technique to optimize the camera configuration, and dynamically selects the most sensitive pixels in a camera frame to extract reliable cardiac motion patterns. Furthermore, the morphological characteristic analysis is deployed to derive user-specific cardiac features, and a feature transformation scheme grounded on Principle Component Analysis (PCA) is developed to enhance the robustness of cardiac biometrics for effective user verification. With the prototyped system, extensive experiments involving $25$ subjects are conducted to demonstrate that CardioCam can achieve effective and reliable user verification with over $99%$ average true positive rate (TPR) while maintaining the false positive rate (FPR) as low as $4%$. Jian Liu 0001, Cong Shi 0004, Yingying Chen 0001, Hongbo Liu 0002, Marco Gruteser |
MobiSys | 2 |
| 2017 | Smart User Authentication through Actuation of Daily Activities Leveraging WiFi-enabled IoTabstractUser authentication is a critical process in both corporate and home environments due to the ever-growing security and privacy concerns. With the advancement of smart cities and home environments, the concept of user authentication is evolved with a broader implication by not only preventing unauthorized users from accessing confidential information but also providing the opportunities for customized services corresponding to a specific user. Traditional approaches of user authentication either require specialized device installation or inconvenient wearable sensor attachment. This paper supports the extended concept of user authentication with a device-free approach by leveraging the prevalent WiFi signals made available by IoT devices, such as smart refrigerator, smart TV and thermostat, etc. The proposed system utilizes the WiFi signals to capture unique human physiological and behavioral characteristics inherited from their daily activities, including both walking and stationary ones. Particularly, we extract representative features from channel state information (CSI) measurements of WiFi signals, and develop a deep learning based user authentication scheme to accurately identify each individual user. Extensive experiments in two typical indoor environments, a university office and an apartment, are conducted to demonstrate the effectiveness of the proposed authentication system. In particular, our system can achieve over 94% and 91% authentication accuracy with 11 subjects through walking and stationary activities, respectively. Cong Shi 0004, Jian Liu 0001, Hongbo Liu 0002, Yingying Chen 0001 |
MobiHoc | 1 |
| 2017 | WiFi-Enabled Smart Human Dynamics MonitoringabstractThe rapid pace of urbanization and socioeconomic development encourage people to spend more time together and therefore monitoring of human dynamics is of great importance, especially for facilities of elder care and involving multiple activities. Traditional approaches are limited due to their high deployment costs and privacy concerns (e.g., camera-based surveillance or sensor-attachment-based solutions). In this work, we propose to provide a fine-grained comprehensive view of human dynamics using existing WiFi infrastructures often available in many indoor venues. Our approach is low-cost and device-free, which does not require any active human participation. Our system aims to provide smart human dynamics monitoring through participant number estimation, human density estimation and walking speed and direction derivation. A semi-supervised learning approach leveraging the non-linear regression model is developed to significantly reduce training efforts and accommodate different monitoring environments. We further derive participant number and density estimation based on the statistical distribution of Channel State Information (CSI) measurements. In addition, people's walking speed and direction are estimated by using a frequency-based mechanism. Extensive experiments over 12 months demonstrate that our system can perform fine-grained effective human dynamic monitoring with over 90% accuracy in estimating participants number, density, and walking speed and direction at various indoor environments. Xiaonan Guo 0003, Bo Liu 0058, Cong Shi 0004, Hongbo Liu 0002, Yingying Chen 0001, Mooi Choo Chuah |
SenSys | 3 |