Xiangyu Xu 0001

dblp:172/1282-1 · DBLP profile ↗
← Back
32ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0002-1628-906XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 22 · 11 first-author · 14 since 2021Security and privacy · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs
abstract
Deep Neural Networks (DNNs), as valuable intellectual property, face unauthorized use. Existing protections, such as digital watermarking, are largely passive; they provide only post-hoc ownership verification and cannot actively prevent the illicit use of a stolen model. This work proposes a proactive protection scheme, dubbed ``Authority Backdoor," which embeds access constraints directly into the model. In particular, the scheme utilizes a backdoor learning framework to intrinsically lock a model's utility, such that it performs normally only in the presence of a specific trigger (e.g., a hardware fingerprint). But in its absence, the DNN's performance degrades to be useless. To further enhance the security of the proposed authority scheme, the certifiable robustness is integrated to prevent an adaptive attacker from removing the implanted backdoor. The resulting framework establishes a secure authority mechanism for DNNs, combining access control with certifiable robustness against adversarial attacks. Extensive experiments on diverse architectures and datasets validate the effectiveness and certifiable robustness of the proposed framework.
Shaofeng Li 0001, Tian Dong 0003, Xiangyu Xu 0001, Guangchi Liu, Zhen Ling 0001
AAAI4
2026 A Needle in a Haystack: Defending Federated Learning Backdoor Attacks via Orthogonal Subnetwork Pruning
Zihan Ma 0008, Guangchi Liu, Xiangyu Xu 0001, Shaofeng Li 0001, Zhen Ling 0001, Junzhou Luo
INFOCOM3
2026 Time will Tell: Large-scale De-anonymization of Hidden I2P Services via Live Behavior Alignment
Hongze Wang, Zhen Ling 0001, Xiangyu Xu 0001, Yumingzhi Pan, Guangchi Liu, Junzhou Luo, Xinwen Fu
NDSS3
2026 Toward Practical Headphones Eavesdropping Leveraging COTS mmWave Radar
abstract
Headphones have become ubiquitous in daily work and communication, leading users to assume a sense of privacy and security during confidential conversations while overlooking the potential risk of eavesdropping. In this paper, we present mmEar, an end-to-end eavesdropping system that demonstrates the feasibility of compromising headphones using a commercial off-the-shelf (COTS) mmWave radar. Unlike previous approaches that rely on relatively strong vibrations, mmEar targets extremely faint, low-SNR speech-induced vibrations on headphone surfaces. To address this challenge, we introduce a Faint Vibration Emphasis (FVE) technique that amplifies phase variations on the IQ plane, followed by a deep denoising network for enhanced signal quality. Furthermore, we design a diffusion-based generative model within a pretrain–finetune framework, leveraging large-scale synthetic data to significantly improve generalization and robustness across diverse scenarios. Extensive experiments on multiple headphone and earphone models validate the practicality and effectiveness of the proposed attack, revealing that most tested devices can be compromised to recover intelligible speech.
Xiangyu Xu 0001, Hao Kong 0004, Zhen Ling 0001, Jiadi Yu, Junzhou Luo, Xinwen Fu
IEEE Trans. Mob. Comput.1
2025 FlexEmu: Towards Flexible MCU Peripheral Emulation
abstract
Microcontroller units (MCUs) are widely used in embedded devices due to their low power consumption and cost-effectiveness. MCU firmware controls these devices and is vital to the security of embedded systems. However, performing dynamic security analyses for MCU firmware has remained challenging due to the lack of usable execution environments -- existing dynamic analyses cannot run on physical devices (e.g., insufficient computational resources), while building emulators is costly due to the massive amount of heterogeneous hardware, especially peripherals. Recent advances in automated peripheral emulation have made MCU emulation more scalable. However, these efforts only support limited peripherals and are hard to extend because they require ad-hoc adaptations.
Chongqing Lei, Zhen Ling 0001, Xiangyu Xu 0001, Shaofeng Li 0001, Guangchi Liu, Kai Dong 0001, Junzhou Luo
CCS3
2025 Bilateral Virtual Companions: The Impact of Virtual Humans' Movement and Voice Realism on User Perception and Experience in Multi-user VR Cinemas
abstract
Despite the increasing prevalence of online social interaction, challenges such as insufficient immersion and lack of interactivity still persist. To overcome these limitations, this study developed a multi-user virtual reality (VR) cinema system with motion capture (Mocap) and multi-user VR technology. The system demonstrated strengths in overcoming physical space restrictions and saving travel costs, providing a more enriched interactive experience and fulfilling social needs under special circumstances, and facilitating metaverse applications. Furthermore, to give insight into the further design of multi-user VR cinemas, this study investigated the impact of virtual humans' (VHs') characteristics on user perception and experience of bilateral virtual companions, by evaluating the impact of movement and voice realism on immersion, social presence and intimacy. The results show that while neither movement nor voice realism significantly influences immersion, voice realism rather than movement realism significantly affects social presence and intimacy.
Jingfeng Hu, Ding Ding 0002, Xiangyu Xu 0001, Jinghui Zhang 0001, Jiahui Jin 0001, Fang Dong 0001
CSCWD3
2025 Poster: KeyRadar: Contactless Touchscreen Keystroke Inference via mmWave Sensing and Language Models
abstract
We present KeyRadar, a contactless keystroke inference system that leverages mmWave radar to capture fine-grained 2D motion signals from virtual keypresses on touchscreen devices. Unlike existing visual or acoustic side-channel methods, KeyRadar uses a MIMO radar array to sense subtle back-surface vibrations and employs a hybrid CNN-Transformer model for accurate single-key recognition. To reconstruct full text input, it combines a Tire-based decoding algorithm with a large language model for semantic correction. Evaluated on a nine-key touchscreen in various user conditions, KeyRadar achieves more than 77% single-key precision and 0.8 + semantic similarity, demonstrating a practical and stealthy threat to input privacy.
Haixin Zhang, Xiangyu Xu 0001, Zhen Ling 0001
MobiCom2
2025 Poster: Object-Aware Vibration Fusion: Leveraging Frequency Response Diversity for Through-Wall Eavesdropping
abstract
Traditional mmWave-based acoustic eavesdropping systems primarily focus on sensing techniques, overlooking the physical characteristics of the objects being sensed. In this work, we propose Object-Aware Vibration Fusion, a through-wall eavesdropping system that leverages the frequency response diversity of everyday objects to enhance speech reconstruction. Instead of treating environmental surfaces as generic reflectors, we model and exploit their distinct resonance patterns, which respond selectively to different frequency bands of speech. By capturing and fusing vibrations from multiple objects through a frequency-aware fusion framework, our system constructs a richer and more intelligible representation of the original audio. Experimental results show that this object-centric approach significantly improves speech intelligibility and quality across varying conditions, highlighting the power of frequency response diversity in passive acoustic sensing.
Xiangyu Xu 0001, Zhen Ling 0001
MobiCom2
2025 Pivot: Panoramic-Image-Based VR User Authentication against Side-Channel Attacks
abstract
With metaverse attracting increasing attention from both academic and industry, the application of virtual reality (VR) has extended beyond 3D immersive viewing/gaming to a broader range of areas, such as banking, shopping, tourism, education, and so on, which involves a growing amount of sensitive and private user data into VR systems. However, with current password-based user authentication schemes in mainstream VR devices, studies demonstrate that side-channel attacks can pose a severe threat to VR user privacy. To mitigate the threat, we propose a novel panoramic-image-based VR user authentication system, i.e., Pivot , to defend against such attacks, yet maintain high usability. Specifically, in Pivot , we design an image-random-pivoting-based user interaction mechanism to assist users in quickly and securely selecting memorable points of interest in a panoramic image. Then an image region segmentation algorithm is designed to automatically scatter the points to regions to form the customized graphic password for the user, which could ensure a sufficiently large password space and also reduce the near-region point misclicks. Afterward, the region indexes are used to generate the hashed password for authentication. Both theoretical security analysis and extensive user studies demonstrate that Pivot is secure and user-friendly in practice.
Gui Xiao, Zhen Ling 0001, Qunqun Fan, Xiangyu Xu 0001, Wenjia Wu, Ding Ding 0002, Chen Chen 0147, Xinwen Fu
ACM Trans. Multim. Comput. Commun. Appl.4
2024 WFGuard: an Effective Fuzzing-testing-based Traffic Morphing Defense against Website Fingerprinting
abstract
Website fingerprinting (WF) attack is a type of traffic analysis attack. It enables a local and passive eavesdropper situated between the Tor client and the Tor entry node to deduce which websites the client is visiting. Currently, deep learning (DL) based WF attacks have overcome a number of proposed WF defenses, demonstrating superior performance compared to traditional machine learning (ML) based WF attacks. To mitigate this threat, we present WFGuard, a fuzzing-testing-based traffic morphing WF defense technique. WFGuard employs fine-grained neuron information within WF classifiers to design a joint optimization function and then applies gradient ascent to maximize both neurons value and misclassification possibility in DL-based WF classifiers. During each traffic mutation cycle, we propose a gradient based dummy traffic injection pattern generation approach, continuously mutating the traffic until a pattern emerges that can successfully deceive the classifier. Finally, the pattern present in successful variant traces are extracted and applied as defense strategies to Tor traffic. Extensive evaluations reveal that WFGuard can effectively decrease the accuracy of DL-based WF classifiers (e.g., DF and Var-CNN) to a mere 4.43%, while only incurring an 11.04% bandwidth overhead. This highlights the potential efficacy of our approach in mitigating WF attacks.
Zhen Ling 0001, Gui Xiao, Xiangyu Xu 0001, Guangchi Liu
INFOCOM5
2024 mmEar: Push the Limit of COTS mmWave Eavesdropping on Headphones
abstract
Recent years have witnessed a surge of headphones (including in-ear headphones) usage in works and communications. Because of the privacy-preserve property, people feel comfortable having confidential communication wearing headphones and pay little attention to speech leakage. In this paper, we present an end-to-end eavesdropping system, mmEar, which shows the feasibility of launching an eavesdropping attack on headphones leveraging a commercial mmWave radar. Different from previous works that realize eavesdropping by sensing speech-induced vibrations with reasonable amplitude, mmEar focuses on capturing the extremely faint vibrations with a low signal-to-noise ratio (SNR) on the surface of headphones. Toward this end, we propose a faint vibration emphasis (FVE) method that models and amplifies the mmWave responses to speech-induced vibrations on the In-phase and Quadrature (IQ) plane, followed by a deep denoising network to further improve the SNR. To achieve practical eavesdropping on various headphones and setups, we propose a cGAN model with a pretrain-finetune scheme, boosting the generalization ability and robustness of the attack by generating high-quality synthesis data. We evaluate mmEar with extensive experiments on different headphones and earphones and find that most of them can be compromised by the proposed attack for speech recovery.
Xiangyu Xu 0001, Zhen Ling 0001, Li Lu 0008, Junzhou Luo, Xinwen Fu
INFOCOM1
2024 Conan's Bow Tie: A Streaming Voice Conversion for Real-Time VTuber Livestreaming
abstract
Recent years have witnessed a dramatic growing trend of Virtual YouTubers (VTubers) as a new business on social media, such as YouTube, Twitch, and TikTok. However, a significant challenge arises when VTuber voice actors face health issues or retire, jeopardizing the continuity of their avatar’s recognizable voices. A potential solution reminiscent of Conan’s Bow Tie voice changer in the popular animation Case Closed (i.e., Detective Conan) has inspired our work. To make this a reality, we introduce VTuberBowTie, a user-friendly streaming voice conversion system for real-time VTuber livestreaming. We propose an innovative streaming voice conversion approach that tackles the challenges of limited context modeling and bidirectional context dependence inherent to conventional real-time voice conversion. Rather than individually processing the voice stream in data chunks, our approach adopts a fully sequential structure that leverages contextual information preceding the input chunk, thereby expanding the perceptual range and enabling seamless concatenation. Moreover, we developed a ready-to-use interaction interface for VTuberBowTie and deployed it on various computing platforms. The experimental results show that VTuberBowTie can achieve high-quality voice conversion in a streaming manner with a latency of 179.1ms on CPU and 70.8ms on GPU while providing users a friendly interactive experience.
Qianniu Chen, Zhehan Gu, Li Lu 0008, Xiangyu Xu 0001, Zhongjie Ba, Feng Lin 0004, Zhenguang Liu, Kui Ren 0001
IUI4
2024 FraudWhistler: A Resilient, Robust and Plug-and-play Adversarial Example Detection Method for Speaker Recognition
Kun Wang 0025, Xiangyu Xu 0001, Li Lu 0008, Zhongjie Ba, Feng Lin 0004, Kui Ren 0001
USENIX Security Symposium2
2024 Devil in the Room: Triggering Audio Backdoors in the Physical World
Meng Chen 0011, Xiangyu Xu 0001, Li Lu 0008, Zhongjie Ba, Feng Lin 0004, Kui Ren 0001
USENIX Security Symposium2
2024 iStrayPaws: Immersing in a Stray Animal's World through First-Person VR to Bridge Human-Animal Empathy
abstract
While Virtual Reality Perspective-Taking (VRPT) demonstrates its efficiency in inducing empathy, its application primarily focuses on vulnerable humans, not animals. Existing animal-related works mainly targets farm animals and wildlife. In this work, we focus on stray animals and introduce iStrayPaws, a VRPT system that simulates stray animals’ challenging lives. The system offers users an immersive first-person journey into the world of stray animals encountering different difficulties like inclement weather, hunger, and illnesses. Enriched with audio-visual and kinesthetic design, the system seeks to deepen users’ understanding of stray animals’ life and foster profound emotional connections. To evaluate the system, a user study was conducted, which showed that VRPT recipients exhibited significant improvement in both state and trait empathy compared to traditional method. Our research not only delivers a novel, accessible, and interactive animal empathy experience but also provides innovative solutions for addressing stray animal issues and advancing broader animal welfare work.
Ding Ding 0002, Yongxin Chen 0004, Zhuying Li 0001, Xiangyu Xu 0001
VRST5
2023 FlyingLoRa: Towards energy efficient data collection in UAV-assisted LoRa networks
Runqun Xiong, Chuan Liang, Xiangyu Xu 0001, Junzhou Luo
Comput. Networks4
2023 Toward Multi-User Authentication Using WiFi Signals
abstract
User authentication nowadays has become an important support for not only security guarantees but also emerging novel applications. Although WiFi signal-based user authentication has achieved initial success, it works in single-user scenarios while multi-user authentication remains a challenging task. In this paper, we present MultiAuth, a multi-user authentication system that can authenticate multiple users with a single pair of commodity WiFi devices. The basic idea is to profile multipath components of WiFi signals, and leverage the multipath components to characterize each user individually for multi-user authentication. MultiAuth first profiles multipath components of WiFi signals through a proposed MUltipath Time-of-Arrival estimation algorithm (MUTA). Then, after matching corresponding multipath components to each user in complex multi-user scenarios, MultiAuth constructs individual CSI based on the multipath components to characterize each user individually. An AoA-based approach is exploited to further separate individual CSI constructed by the users with same ToA. To identify users through their activities, MultiAuth extracts user behavior profiles based on the individual CSI, and leverages a dual-task neural network for robust user authentication. Extensive experiments involving 3 simultaneously present users demonstrate that MultiAuth is effective in multi-user authentication with 86.2% average accuracy and 9.5% average false accept rate.
Hao Kong 0004, Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Xiangyu Xu 0001, Feng Lyu 0001
IEEE/ACM Trans. Netw.5
2022 mmECG: Monitoring Human Cardiac Cycle in Driving Environments Leveraging Millimeter Wave
abstract
The continuously increasing time spent on car trips in recent years brings growing attention to the physical and mental health of drivers on roads. As one of the key vital signs, the heartbeat is a critical indicator of drivers' health states. Most existing studies on heartbeat monitoring either require sensor attachment or could only provide sketchy heart rates. Moreover, most approaches require the subject to remain stationary or a quiet measuring environment, which is hard to apply to dynamic driving environments. In this paper, we propose a contactless cardiac cycle monitoring system, mmECG, which leverages Commercial-Off-The-Shelf mmWave radar to estimate the fine-grained heart movements of drivers in moving vehicles. By exploring the principle of mmWave signal-based sensing, we first perform studies in static environments and find the fine-grained heart movements, represented as stages of atria and ventricles in repetitive cardiac cycles, can be captured by the FMCW-based mmWave radar as phase changes in signals. Whereas in driving environments, such phase changes are caused and influenced by not only the heartbeat of drivers but also driving operations and vehicle dynamics. To further extract the minute heart movements of drivers and eliminate other influences in phase changes, we construct a movement mixture model to represent the phase changes caused by different movements, and further design a hierarchy variational mode decomposition (VMD) approach to extract and estimate the essential heart movement in mmWave signals. Finally, based on the extracted phase changes, mmECG reconstructs the cardiac cycle by estimating fine-grained movements of atria and ventricles leveraging a template-based optimization method. Experimental results involving 25 drivers in real driving scenarios demonstrate that mmECG can accurately estimate not only heart rates but also cardiac cycles of drivers in real driving environments.
Xiangyu Xu 0001, Jiadi Yu, Chengguang Ma, Yanzhi Ren, Hongbo Liu 0002, Yanmin Zhu 0006, Yingying Chen 0001, Feilong Tang 0001
INFOCOM1
2022 m3Track: mmwave-based multi-user 3D posture tracking
abstract
Nowadays, the market of 3D human posture tracking has extended to a broad range of application scenarios. As current mainstream solutions, vision-based posture tracking systems suffer from privacy leakage concerns and depend on lighting conditions. Towards more privacy-preserving and robust tracking manner, recent works have exploited commodity radio frequency signals to realize 3D human posture tracking. However, these studies cannot handle the case where multiple users are in the same space. In this paper, we present a mmWave-based multi-user 3D posture tracking system, m3Track, which leverages a single commercial off-the-shelf (COTS) mmWave radar to track multiple users' postures simultaneously as they move, walk, or sit. Based on the sensing signals from a mmWave radar in multi-user scenarios, m3Track first separates all the users on mmWave signals. Then, m3Track extracts shape and motion features of each user, and reconstructs 3D human posture for each user through a designed deep learning model. Furthermore. m3Track maps the reconstructed 3D postures of all users into 3D space, and tracks users' positions through a coordinate-corrected tracking method, realizing practical multi-user 3D posture tracking with a COTS mmWave radar. Experiments conducted in real-world multi-user scenarios validate the accuracy and robustness of m3Track on multi-user 3D posture tracking.
Hao Kong 0004, Xiangyu Xu 0001, Jiadi Yu, Qilin Chen, Chenguang Ma, Yingying Chen 0001, Yi-Chao Chen 0001, Linghe Kong
MobiSys2
2022 Leveraging Acoustic Signals for Fine-Grained Breathing Monitoring in Driving Environments
abstract
Given the increasing amount of time people spent on driving, the physical and mental health of drivers is essential to road safety. Breathing patterns are critical indicators of the wellbeing of drivers on the road. Existing studies on breathing monitoring require active user participation of wearing special sensors or relatively quiet environments during sleep, which are hardly applicable to noisy driving environments. In this work, we propose a fine-grained breathing monitoring system,BreathListener, which leverages audio devices on smartphones to estimate the fine-grained breathing waveform in driving environments. By investigating the data collected from real driving environments, we find that energy spectrum density (ESD) of acoustic signals can be utilized to capture breathing procedures in driving environments. To extract breathing pattern in ESD signals,BreathListenereliminates interference from driving environments in ESD signals utilizing background subtraction and variational mode decomposition (VMD). After that, the extracted breathing pattern is transformed into Hilbert spectrum, and we further design a deep learning architecture based on generative adversarial network (GAN) to generate fine-grained breathing waveform from the Hilbert spectrum of extracted breathing patterns in ESD signals. Experiments with ten drivers in real driving environments show thatBreathListenercan accurately capture breathing patterns of drivers in driving environments.
Xiangyu Xu 0001, Jiadi Yu, Yingying Chen 0001
IEEE Trans. Mob. Comput.1
2021 HVAC: Evading Classifier-based Defenses in Hidden Voice Attacks
abstract
Recent years have witnessed the rapid development of automatic speech recognition (ASR) systems, providing a practical voice-user interface for widely deployed smart devices. With the ever-growing deployment of such an interface, several voice-based attack schemes have been proposed towards current ASR systems to exploit certain vulnerabilities. Posing one of the more serious threats,hidden voice attack uses the human-machine perception gap to generate obfuscated/hidden voice commands that are unintelligible to human listeners but can be interpreted as commands by machines. However, due to the nature of hidden voice commands (i.e., normal and obfuscated samples exhibit a significant difference in their acoustic features), recent studies show that they can be easily detected and defended by a pre-trained classifier, thereby making it less threatening. In this paper, we validate that such a defense strategy can be circumvented with a more advanced type of hidden voice attack calledHVAC. Our proposed HVAC attack can easily bypass the existing learning-based defense classifiers while preserving all the essential characteristics of hidden voice attacks (i.e., unintelligible to humans and recognizable to machines). Specifically, we find that all classifier-based defenses build on top of classification models that are trained with acoustic features extracted from the entire audio of normal and obfuscated samples. However, only speech parts (i.e., human voice parts) of these samples contain the useful linguistic information needed for machine transcription. We thus propose a fusion-based method to combine the normal sample and corresponding obfuscated sample as a hybrid HVAC command, which can effectively cheat the defense classifiers. Moreover, to make the command more unintelligible to humans, we tune the speed and pitch of the sample and make it even more distorted in the time domain while ensuring it can still be recognized by machines. Extensive physical over-the-air experiments demonstrate the robustness and generalizability of our HVAC attack under different realistic attack scenarios. Results show that our HVAC commands can achieve an average 94.1% success rate of bypassing machine-learning-based defense approaches under various realistic settings.
Yi Wu 0020, Xiangyu Xu 0001, Payton Walker, Jian Liu 0001, Nitesh Saxena, Yingying Chen 0001, Jiadi Yu
AsiaCCS2
2021 Face-Mic: inferring live speech and speaker identity via subtle facial dynamics captured by AR/VR motion sensors
abstract
Augmented reality/virtual reality (AR/VR) has extended beyond 3D immersive gaming to a broader array of applications, such as shopping, tourism, education. And recently there has been a large shift from handheld-controller dominated interactions to headset-dominated interactions via voice interfaces. In this work, we show a serious privacy risk of using voice interfaces while the user is wearing the face-mounted AR/VR devices. Specifically, we design an eavesdropping attack, Face-Mic, which leverages speech-associated subtle facial dynamics captured by zero-permission motion sensors in AR/VR headsets to infer highly sensitive information from live human speech, including speaker gender, identity, and speech content. Face-Mic is grounded on a key insight that AR/VR headsets are closely mounted on the user's face, allowing a potentially malicious app on the headset to capture underlying facial dynamics as the wearer speaks, including movements of facial muscles and bone-borne vibrations, which encode private biometrics and speech characteristics. To mitigate the impacts of body movements, we develop a signal source separation technique to identify and separate the speech-associated facial dynamics from other types of body movements. We further extract representative features with respect to the two types of facial dynamics. We successfully demonstrate the privacy leakage through AR/VR headsets by deriving the user's gender/identity and extracting speech information via the development of a deep learning-based framework. Extensive experiments using four mainstream VR headsets validate the generalizability, effectiveness, and high accuracy of Face-Mic.
Cong Shi 0004, Xiangyu Xu 0001, Tianfang Zhang, Payton Walker, Yi Wu 0020, Jian Liu 0001, Nitesh Saxena, Yingying Chen 0001, Jiadi Yu
MobiCom2
2021 MultiAuth: Enable Multi-User Authentication with Single Commodity WiFi Device
abstract
With the increasing integration of humans and the cyber world, user authentication becomes critical to support various emerging application scenarios requiring security guarantees. Existing works utilize Channel State Information (CSI) of WiFi signals to capture single human activities for non-intrusive and device-free user authentication, but multi-user authentication remains a challenging task. In this paper, we present a multi-user authentication system, MultiAuth, which can authenticate multiple users with a single commodity WiFi device. The key idea is to profile multipath components of WiFi signals induced by multiple users, and construct individual CSI from the multipath components to solely characterize each user for user authentication. Specifically, we propose a MUltipath Time-of-Arrival measurement algorithm (MUTA) to profile multipath components of WiFi signals in high resolution. Then, after aggregating and separating the multipath components related to users, MultiAuth constructs individual CSI based on the multipath components to solely characterize each user. To identify users, MultiAuth further extracts user behavior profiles based on the individual CSI of each user through time-frequency analysis, and leverages a dual-task neural network for robust user authentication. Extensive experiments involving 3 simultaneously present users demonstrate that MultiAuth is accurate and reliable for multi-user authentication with 87.6% average accuracy and 8.8% average false accept rate.
Hao Kong 0004, Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Xiangyu Xu 0001, Feilong Tang 0001, Yi-Chao Chen 0001
MobiHoc5
2020 TouchPass: towards behavior-irrelevant on-touch user authentication on smartphones leveraging vibrations
abstract
With increasing private and sensitive data stored in mobile devices, secure and effective mobile-based user authentication schemes are desired. As the most natural way to contact with mobile devices, finger touches have shown potentials for user authentication. Most existing approaches utilize finger touches as behavioral biometrics for identifying individuals, which are vulnerable to spoofer attacks. To resist attacks for on-touch user authentication on mobile devices, this paper exploits physical characters of touching fingers by investigating active vibration signal transmission through fingers, and we find that physical characters of touching fingers present unique patterns on active vibration signals for different individuals. Based on the observation, we propose a behavior-irrelevant on-touch user authentication system, TouchPass, which leverages active vibration signals on smartphones to extract only physical characters of touching fingers for user identification. TouchPass first extracts features that mix physical characters of touching fingers and behavior biometrics of touching behaviors from vibration signals generated and received by smartphones. Then, we design a Siamese network-based architecture with a specific training sample selection strategy to reconstruct the extracted signal features to behavior-irrelevant features and further build a behavior-irrelevant on-touch user authentication scheme leveraging knowledge distillation. Our extensive experiments validate that TouchPass can accurately authenticate users and defend various attacks.
Xiangyu Xu 0001, Jiadi Yu, Yingying Chen 0001, Qin Hua, Yanmin Zhu 0006, Yi-Chao Chen 0001, Minglu Li 0001
MobiCom1
2020 Leveraging Acoustic Signals for Vehicle Steering Tracking with Smartphones
abstract
Given the increasing popularity, mobile devices are exploited to enhance active driving safety nowadays. Among all safety services provided for vehicles, tracking the rotation angle of steering wheel in real time can monitor the vehicles' dynamics and drivers' behaviors at the same time. In this paper, we propose a steering tracking system, SteerTrack, which tracks the rotation angle of the steering wheel in real time leveraging audio devices on smartphones. SteerTrack seeks a device-free approach for steering tracking without requiring installation of specialized sensors on the steering wheels nor asking drivers to wear sensors on their wrists. Since the steering wheel is operated by a driver's hands, the rotation angle of the steering wheel can be tracked based on movements of the driver's hands. SteerTrack first builds an acoustic signal field inside of a vehicle and then analyzes the echoes reflected from the driver's hands with relative correlation coefficient (RCC) and reference frame to track the movement trajectory of hands under different steering maneuvers. Given the tracked movement trajectory, SteerTrackfurther develops a geometrical transformation-based method for estimating the rotation angle of the steering wheel in 3D driving environments by projecting the steering wheel to a 2D ellipse. Through extensive experiments in real driving environments with five volunteers for several weeks, SteerTrack can achieve an average steering wheel estimation error of 1.48 degree during driving, and 4.61 degree for turns.
Xiangyu Xu 0001, Jiadi Yu, Yingying Chen 0001, Yanmin Zhu 0006, Minglu Li 0001
IEEE Trans. Mob. Comput.1
2019 KeyListener: Inferring Keystrokes on QWERTY Keyboard of Touch Screen through Acoustic Signals
abstract
This paper demonstrates the feasibility of a side-channel attack to infer keystrokes on touch screen leveraging an off-the-shelf smartphone. Although there exist some studies on keystroke eavesdropping attacks on touch screen, they are mainly direct eavesdropping attacks, i.e., require the device of victims compromised to provide side-channel information for the adversary, which are hardly launched in practical scenarios. In this work, we show the practicability of an indirect eavesdropping attack, KeyListener, which infers keystrokes on QWERTY keyboards of touch screen leveraging audio devices on a smartphone. We investigate the attenuation of acoustic signals, and find that a user's keystroke fingers can be localized through the attenuation of acoustic signals received by the microphones in the smartphone. We then utilize the attenuation of acoustic signals to localize each keystroke, and further analyze errors induced by ambient noises. To improve the accuracy of keystroke localization, KeyListener further tracks finger movements during inputs through phase change and Doppler effect to reduce errors of acoustic signal attenuation-based keystroke localization. In addition, a binary tree-based search approach is employed to infer keystrokes in a context-aware manner. The proposed keystroke eavesdropping attack is robust to various environments without the assistance of additional infrastructures. Extensive experiments demonstrate that the accuracy of keystroke inference in top-5 candidates can approach 90% with a top-5 error rate of around 6%, which is a strong indication of the possible user privacy leakage of inputs on QWERTY keyboard.
Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Yanmin Zhu 0006, Xiangyu Xu 0001, Guangtao Xue, Minglu Li 0001
INFOCOM5
2019 BreathListener: Fine-grained Breathing Monitoring in Driving Environments Utilizing Acoustic Signals
abstract
Given the increasing amount of time people spent on driving, the physical and mental health of drivers is essential to road safety. Breathing patterns are critical indicators of the well-being of drivers on the road. Existing studies on breathing monitoring require active user participation of wearing special sensors or relatively quiet environments during sleep, which are hardly applicable to noisy driving environments. In this work, we propose a fine-grained breathing monitoring system, BreathListener, which leverages audio devices on smartphones to estimate the fine-grained breathing waveform in driving environments. By investigating the data collected from real driving environments, we find that Energy Spectrum Density (ESD) of acoustic signals can be utilized to capture breathing procedures in driving environments. To extract breathing pattern in ESD signals, BreathListener eliminates interference from driving environments in ESD signals utilizing background subtraction and Ensemble Empirical Mode Decomposition (EEMD). After that, the extracted breathing pattern is transformed into Hilbert spectrum, and we further design a deep learning architecture based on Generative Adversarial Network (GAN) to generate fine-grained breathing waveform from the Hilbert spectrum of extracted breathing patterns in ESD signals. Experiments with 10 drivers in real driving environments show that BreathListener can accurately capture breathing patterns of drivers in driving environments.
Xiangyu Xu 0001, Jiadi Yu, Yingying Chen 0001, Yanmin Zhu 0006, Linghe Kong, Minglu Li 0001
MobiSys1
2018 VPad: Virtual Writing Tablet for Laptops Leveraging Acoustic Signals
abstract
Human-computer interaction based on touch screens plays an increasing role in our daily lives. Besides smartphones and tablets, laptops are the most popular mobile devices used in both work and leisure. To satisfy requirements of many emerging applications, it becomes desirable to equip both writing and drawing functions directly on laptop screens. In this paper, we design a virtual writing tablet system, VPad, for traditional laptops without touch screens. VPad leverages two speakers and one microphone, which are available in most commodity laptops, for trajectory tracking without additional hardware. It employs acoustic signals to accurately track hand movements and recognize characters user writes in the air. Specifically, VPad emits inaudible acoustic signals from two speakers in a laptop. Then VPad applies Sliding-window Overlap Fourier Transformation technique to find Doppler frequency shift with higher resolution and accuracy in real time. Furthermore, we analyze frequency shifts and energy features of acoustic signals received by the microphone to track the trajectory of hand movements. Finally, we employ a stroke direction sequence model based on possibility estimation to recognize characters users write in the air. Our experimental results show that VPad achieves the average trajectory tracking error of only 1.55cm and the character recognition accuracy of above 90% merely through two speakers and one microphone on a laptop.
Li Lu 0008, Jian Liu 0001, Jiadi Yu, Yingying Chen 0001, Yanmin Zhu 0006, Xiangyu Xu 0001, Minglu Li 0001
ICPADS6
2018 SteerTrack: Acoustic-Based Device-Free Steering Tracking Leveraging Smartphones
abstract
Given the increasing popularity, mobile devices are exploited to enhance active driving safety nowadays. Among all safety services provided for vehicles, tracking the rotation angle of steering wheel in real time can monitor the vehicles' dynamics and drivers' behaviors at the same time. In this paper, we propose a steering tracking system, SteerTrack, which tracks the rotation angle of steering wheel in real time leveraging audio devices on smartphones. SteerTrack seeks a device-free approach for steering tracking without requiring installation of specialized sensors on steering wheels nor asking drivers to wear sensors on their wrists. Since the steering wheel is operated by a driver's hands, the rotation angle of steering wheel can be tracked based on movements of the driver's hands. SteerTrack first builds an acoustic signal field inside of a vehicle and then analyzes the echoes reflected from the driver's hands with relative correlation coefficient(RCC) and reference frame to track the movement trajectory of hands under different steering maneuvers. Given the tracked movement trajectory, SteerTrack further develops a geometrical transformation-based method for estimating the rotation angle of steering wheel in 3D driving environments by projecting the steering wheel to a 2D ellipse. Through extensive experiments in real driving environments with 5 volunteers for several weeks, SteerTrack can achieve an average error of 4.61 degree for estimating the rotation angle of steering wheel.
Xiangyu Xu 0001, Jiadi Yu, Yingying Chen 0001, Yanmin Zhu 0006, Minglu Li 0001
SECON1
2018 Leveraging Audio Signals for Early Recognition of Inattentive Driving with Smartphones
abstract
Real-time driving behavior monitoring is a corner stone to improve driving safety. Most of the existing studies on driving behavior monitoring using smartphones only provide detection results after an abnormal driving behavior is finished, not sufficient for driver alerting and avoiding car accidents. In this paper, we leverage built-in audio devices on smartphones to realize early recognition of inattentive driving events including Fetching Forward, Picking up Drops, Turning Back, and Eating or Drinking. Through empirical studies of driving traces collected in real driving environments, we find that each type of inattentive driving event exhibits unique patterns on Doppler profiles of audio signals. This enables us to develop an Early Recognition system, ER, which can recognize inattentive driving events at an early stage and alert drivers timely. ER employs machine learning methods to first generate binary classifiers for every pair of inattentive driving events, and then develops a modified vote mechanism to form a multi-classifier for all four types of inattentive driving events, for atypical inattentive driving events along with other driving behaviors. It next turns the multi-classifier into a gradient model forestto achieve early recognition of inattentive driving. Through extensive experiments with eight volunteers driving for about two months, ER can achieve an average total accuracy of 94.80 percent for inattentive driving recognition and recognize over 80 percent inattentive driving events before the event is 50 percent finished.
Xiangyu Xu 0001, Jiadi Yu, Yingying Chen 0001, Yanmin Zhu 0006, Shiyou Qian, Minglu Li 0001
IEEE Trans. Mob. Comput.1
2018 Leveraging Smartphones for Vehicle Lane-Level Localization on Highways
abstract
When vehicle road-level localization cannot satisfy people's need for convenience and safety driving, lane-level localization becomes a corner stone in Intelligent Transportation System. Existing works of tracking vehicles on lane-level mostly depend on pre-deployed infrastructures and additional hardwares. In this paper, we utilize smartphones to sense driving conditions for vehicle lane-level localization on highways. By analyzing driving traces collected from real driving environments, we find that each type of lane-change has its unique pattern on the vehicle's lateral acceleration. Based on this observation, we propose a Lane-Level Localization (L3) system, which can perform real-time vehicle localization on lane-level only using smartphones when vehicles are driving on highways. Our system first uses embedded sensors in smartphones to capture the patterns of lane-change behaviors. Then, a Finite State Machine is employed to track vehicles on lane-level leveraging the patterns. Extensive experiments demonstrate that L3is accurate and robust in real driving environments. The experimental results show that, on average, L3achieves the accuracy of 91.49 percent on lane change detection and 90.31 percent on lane-level localization.
Xiangyu Xu 0001, Jiadi Yu, Yanmin Zhu 0006, Zhichen Wu, Jianda Li, Minglu Li 0001
IEEE Trans. Mob. Comput.1
2017 ER: Early recognition of inattentive driving leveraging audio devices on smartphones
abstract
Real-time driving behavior monitoring is a corner stone to improve driving safety. Most of the existing studies on driving behavior monitoring using smartphones only provide detection results after an abnormal driving behavior is finished, not sufficient for driver alert and avoiding car accidents. In this paper, we leverage existing audio devices on smartphones to realize early recognition of inattentive driving events including Fetching Forward, Picking up Drops, Turning Back and Eating or Drinking. Through empirical studies of driving traces collected in real driving environments, we find that each type of inattentive driving event exhibits unique patterns on Doppler profiles of audio signals. This enables us to develop an Early Recognition system, ER, which can recognize inattentive driving events at an early stage and alert drivers timely. ER employs machine learning methods to first generate binary classifiers for every pair of inattentive driving events, and then develops a modified vote mechanism to form a multi-classifier for all inattentive driving events along with other driving behaviors. It next turns the multi-classifier into a gradient model forest to achieve early recognition of inattentive driving. Through extensive experiments with 8 volunteers driving for about half a year, ER can achieve an average total accuracy of 94.80% for inattentive driving recognition and recognize over 80% inattentive driving events before the event is 50% finished.
Xiangyu Xu 0001, Jiadi Yu, Yingying Chen 0001, Yanmin Zhu 0006, Guangtao Xue, Minglu Li 0001
INFOCOM1