EDBT 2026 Demo / reviewers in the wild / expert
Li Lu 0008
dblp:49/2793-8
· DBLP profile ↗
70ranked-venue papers
9as first author
55since 2021 · last 2026
0000-0001-5230-3749ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 30 · 5 first-author · 21 since 2021Security and privacy · 22 · 22 since 2021Systems, architecture and hardware · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic CuesabstractToxic speech detection has become a crucial challenge in maintaining safe online communication environments. However, existing approaches to toxic speech detection often neglect the contribution of paralinguistic cues, such as emotion, intonation, and speech rate, which are key to detecting speech toxicity. Moreover, current toxic speech datasets are predominantly text-based, limiting the development of models that can capture paralinguistic cues. To address these challenges, we present ToxiAlert-Bench, a large-scale audio dataset comprising over 30,000 audio clips annotated with seven major toxic categories and twenty fine-grained toxic labels. Uniquely, our dataset annotates toxicity sources—distinguishing between textual content and paralinguistic origins—for comprehensive toxic speech analysis. Furthermore, we propose a dual-head neural network with a multi-stage training strategy tailored for toxic speech detection. This architecture features two task-specific classification headers: one for identifying the source of sensitivity (textual or paralinguistic), and the other for categorizing the specific toxic type. The training process involves independent head training followed by joint fine-tuning to reduce task interference. To mitigate data class imbalance, we incorporate class-balanced sampling and weighted loss functions. Our experimental results show that leveraging paralinguistic features significantly improves detection performance. Our method consistently outperforms existing baselines across multiple evaluation metrics, with a 21.1% relative improvement in Macro-F1 score and a 13.0% relative gain in accuracy over the strongest baseline, highlighting its enhanced effectiveness and practical applicability. Zhongjie Ba, Liang Yi, Peng Cheng 0007, Qingcao Li, Qinglong Wang 0003, Li Lu 0008 |
AAAI | 6 |
| 2026 | Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
Meng Chen 0011, Kun Wang 0025, Li Lu 0008, Jiaheng Zhang, Tianwei Zhang 0004 |
SP | 3 |
| 2026 | EF-IDS: An efficient intrusion detection system with enriched features for CAN bus in modern vehicles
Aya El-Fatyany, Xiaohang Wang 0001, Li Lu 0008, Kui Ren 0001 |
J. Syst. Archit. | 3 |
| 2026 | FBRE: Fuzzing Based Bit-Level Reverse Engineering of Vehicular CAN BusabstractThe Controller Area Network (CAN) bus serves as a foundational communication architecture in modern vehicles, supporting a wide range of functions, from engine control to auxiliary systems. However, lacking built-in security mechanisms makes CAN vulnerable to cyberattacks. Accurately mapping CAN signals to specific car-control actions becomes critical as it allows the detection of security breaches by pinpointing potential vulnerabilities exploited to compromise vehicular functions. Existing mapping techniques rely on CAN reverse engineering, which struggle to achieve bit-level resolution due to the huge search space of IDs and payload combinations. To address this challenge, we propose a systematic framework that includes signal boundary identification, targeted fuzz testing, and control bit analysis. Our method achieves high efficiency and precision in mapping control bits in CAN frames to car-control actions. Additionally, we developed a compact and user-friendly reverse engineering toolkit, incorporating a graphical interface to facilitate practical vehicle function testing and CAN message monitoring. Experiments on Tesla Model 3 and Leapmotor C11/C10 demonstrate that our framework is validated across different vehicle models and capable of identifying a wide range of car-control actions. Compared with previous works, our method significantly improves the resolution and automation of CAN reverse engineering. Hanxue Shi, Yunlang Cai, Xiaohang Wang 0001, Haoting Shen, Li Lu 0008, Kui Ren 0001, Kaiwei Wu, Yinhe Shen |
IEEE Trans. Computers | 5 |
| 2026 | A Passive Defense Against Out-of-Band Injection Threats to Microphone-Based DevicesabstractThe integration of microphones into a broad array of devices, from consumer electronics to industrial sensors, introduces vulnerabilities to out-of-band injection attacks, including ultrasound, laser, electromagnetic, and magnetic field attacks. These attacks enable adversaries to inject inaudible or imperceptible commands, compromising systems without direct physical access. This paper presents a robust, passive detection framework designed to address the full spectrum of out-of-band attacks on microphone-equipped devices. Unlike prior approaches, our system leverages advanced speech disentanglement to separate semantic and acoustic features from recorded audio, enabling a refined analysis of injection artifacts within each feature domain. By quantifying entropy-based chaos within the disen tangled representations, we detect subtle spectral and structural irregularities indicative of injected signals. The system further incorporates a preliminary stage to identify carrier traces where applicable, expediting detection in cases such as ultrasound and laser attacks. Extensive evaluations across various device types, including smartphones, tablets, and microphones, demonstrate the system's high accuracy and stability, achieving an AUC of 98% under diverse conditions and attack configurations. Feng Lin 0004, Tiantian Liu 0002, Teshi Meng, Zhongjie Ba, Li Lu 0008, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | AE-IPP: Adversarial Example Enabled Identity Privacy Preserving With mmWave SignalsabstractDespite convenience and reliability of mmWave-based action recognition, it still raises privacy concerns on identity leakage threat since human behaviors could meanwhile expose massive user information in real-world applications. Existing solutions attempt to send anonymized features extracted from mmWave signals; however, features not only reduce the application flexibility but also increase the privacy disclosure risk due to original data reconstruction. Instead, in this paper we propose a de-identification system, AE-IPP, which customizes learned noises into the raw data to generate adversarial examples for identity privacy and action utility balance. In other words, the noises are sample-specific perturbations that are automatically learned for each sample through our presented network. To achieve the performance balance and ensure robustness to other models, we are faced with two challenges, including the decoupling of action and identity information and the transferability of models. To this end, AE-IPP focuses on respective attention areas by leveraging task-specific gradients and designs a dynamic attention mechanism to update the attention weights according to the final optimization objective. Moreover, we present a multidirectional perturbation strategy to improve the model generalization capabilities, enabling robust de-identification. Extensive experiments on mmWave datasets demonstrate the superiority of our method over state-of-the-art approaches. Biyun Sheng, Wangquan Qin, Jun Li 0033, Li Lu 0008, Tie Qiu 0001, Fu Xiao 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | On Bit-level Reverse Engineering of Vehicular CAN BusabstractThe Controller Area Network (CAN) bus is a cornerstone of modern vehicles, orchestrating functions from engine control to auxiliary systems. However, its lack of inherent security measures makes it vulnerable to cyberattacks. Accurately mapping CAN signals with car-control actions is critical for detecting security breaches, as it allows pinpointing potential vulnerabilities exploited to compromise vehicular functions. Despite this, existing CAN reverse engineering methods struggle to achieve bit-level resolution due to the huge search space of unique IDs and payload combinations. To address this challenge, we propose a systematic framework for reverse engineering CAN bus messages, achieving precise mapping of control bits in CAN frames to car-control actions. The framework was validated on Tesla Model 3, Leapmotor C10 and C11, demonstrating its versatility across different vehicle platforms. In particular, it successfully identified 43 car-control actions on the Tesla Model 3, showcasing its extensive coverage. Furthermore, its low resource consumption enables seamless integration into compact platforms like the Raspberry Pi, supporting practical deployment in real-world automotive systems. Yunlang Cai, Hanxue Shi, Xiaohang Wang 0001, Haoting Shen, Li Lu 0008, Kui Ren 0001 |
DAC | 5 |
| 2025 | GFuzz4CAN: A Generative Model-based Fuzzing Method for In-vehicle Controller Area NetworkabstractController Area Network (CAN) bus has been the fundamental communication method to interconnect the critical powertrain and body domains in vehicles for decades. Typical in-vehicle CAN buses employ rule-based Intrusion Detection Systems (IDS) to prevent non-compliant CAN messages for security enhancement. However, as the progress of connected autonomous vehicular techniques, more undiscovered vulnerabilities behind CAN IDS are gradually exposed, causing severe threats to vehicle safety, even threatening human life. To address it, fuzzing has become an efficient security testing method to automatically explore unknown vulnerabilities underlying the in-vehicle CAN bus. Existing fuzzing methods either conduct time-consuming random testing that is easily detected by IDS, causing inefficient vulnerability exploration, or rely on precious expert experience for higher efficiency. Instead, this paper turns to seek the power of deep generative models that automatically learn the data communication protocol for compliant fuzzing sample generation. Such a method both releases the experts' knowledge of reverse engineering, and can generate compliant samples to bypass typical rule-based IDS. To further improve the efficiency of exploring exceptions based on the generated samples, we optimize the reward function of SeqGAN with Extreme Value Theory (EVT), to fit the distribution of exception samples. With the trained model, we design a fuzzing system GFuzz4CAN, and conduct extensive experiments for evaluation. Experimental results on open-source datasets and a CAN bus device demonstrate that GFuzz4CAN can achieve$2 \sim 11 \mathrm{x}$higher efficiency compared to byte-level random fuzzing, and is feasible to identify 2 vulnerabilities on the device. Li Lu 0008, Yuli Wu 0004, Shuguo Zhuo, Zhan Qin, Kui Ren 0001 |
ICC | 2 |
| 2025 | Evaluating Robustness of Voice Conversion Systems under Multi-source Channel InterferenceabstractVoice Conversion (VC) technology holds significant potential for enhancing communication across varied application scenarios, such as voice chats, video conferencing, and VTuber live streaming. During the VC systems’ use, there is inevitable and complex channel interference, which becomes a key factor in downgrading the performance of VC systems. However, few studies comprehensively reveal the specific impact of these interferences on VC systems, which gradually become a significant gap between VC design and its landing application requirements. Toward this end, this paper proposes a comprehensive evaluation framework to systematically analyze the effects of multi-source channel interference on VC systems. We investigate multi-source channel interference across physical and digital domains, including noise, reverberation, device distortion, codec compression and transmission loss, and integrate all of them into our proposed evaluation framework. The framework assesses VC systems in terms of three complementary dimensions, i.e., the speaker timbre, speech semantics, and signal consistency, by introducing three core metrics: timbre similarity, semantic similarity, and acoustic similarity. Through extensive experiments on 44,000 samples from 109 speakers across six representative VC systems, we find that VC systems: 1) handle digital interference better than physical interference, 2) maintain timbre features well but struggle with semantics and sound quality, and 3) could achieve robust conversion through feature disentanglement. Qianniu Chen, Xiaodi Zhao, Zhehan Gu, Li Lu 0008 |
IJCNN | 5 |
| 2025 | SecHeadset: A Practical Privacy Protection System for Real-time Voice CommunicationabstractVoice communication is convenient while also poses risks of privacy leakage, due to potential interception or eavesdropping during voice transmission. Current protections of voice privacy are almost entirely controlled by communication service providers (CSPs), which operate as a black-box to users thus hard to fully trust. To take back the control of user privacy, in this paper, we introduce SecHeadset, an end-to-end solution for secure voice communication based on voice obfuscation, which is plug-and-play and compatible with various CSPs. Our solution involves two parts. First, we design a voice-like noise masking scheme for voice obfuscation. The noise, mimicking voice characteristics, could effectively obscure users' voices while demonstrating resilience against noise reduction methods. Second, we develop a protocol that enables efficient channel state estimation and secure information exchange between two communication entities. Based on this information, we propose a lightweight algorithm for voice retrieval during communication. We develop a prototype of SecHeadset and evaluate its performance with 8 widely-used applications, including Telegram and Skype. It reduces the voice recognition accuracy of various adversaries to below 15% while maintaining communication quality. We also integrate SecHeadset with off-the-shelf portable devices and verify its real-world effectiveness. Kun Pan, Qinglong Wang 0003, Peng Cheng 0007, Li Lu 0008, Zhongjie Ba, Kui Ren 0001 |
MobiSys | 5 |
| 2025 | From One Stolen Utterance: Assessing the Risks of Voice Cloning in the AIGC EraabstractThe advent of voice cloning has fundamentally threatened the role of voice as a unique biometric. Many criminal incidents have already been reported to demonstrate its significant risks of identity forgery. Previous works explored the risks of voice cloning in constrained settings, which require victim speakers to either be already seen in the training data of voice cloning models, or leak dozens of minutes of their speech samples to adversaries. However, with the rapid progress of voice cloning in AIGC (Artificial Intelligence Generated Content) era, these requirements have largely been released, leaving the exact risks of state-of-the-art (SOTA) voice cloning techniques shrouded in a dense fog. To uncover it, this paper conducts a large-scale study in real-world scenarios to assess the risks of advanced voice cloning techniques. This study involves 5 SOTA voice cloning techniques (open-source and commercial), across 8 SOTA voice authentication systems (open-source and real-world) and 30 human listeners, using voice data of over 7,000 speakers (public and custom). By experimental and theoretical analysis, this study reveals that 1) state-of-the-art voice cloning techniques pose severe threats in spoofing voice authentication systems and human listeners; 2) demographic factors such as age and gender of victim speakers have a subtle impact on voice cloning attacks; 3) human listeners' subjective opinions and background about voice cloning play an important role in their susceptibility to attacks; 4) advanced detection methods still fail to identify voice cloning samples as expected. Kun Wang 0025, Meng Chen 0011, Li Lu 0008, Jingwen Feng, Qianniu Chen, Zhongjie Ba, Kui Ren 0001, Chun Chen 0001 |
SP | 3 |
| 2025 | LUFT-CAN: A lightweight unsupervised learning based intrusion detection system with frequency-time analysis for vehicular CAN bus
Xiaohang Wang 0001, Li Lu 0008, Shuguo Zhuo, Yingtao Jiang, Amit Kumar Singh 0002, Kui Ren 0001, Mei Yang 0001, Kaiwei Wu |
J. Syst. Archit. | 3 |
| 2025 | Phoneme-Based Proactive Anti-Eavesdropping With Controlled Recording PrivilegeabstractThe widespread smart devices raise people’s concerns of being eavesdropped on. To enhance voice privacy, recent studies exploit the nonlinearity in microphone to jam audio recorders with inaudible ultrasound. However, existing solutions solely rely on energetic masking. Their simple-form noise leads to several problems, such as high energy requirements and being easily removed by speech enhancement techniques. Besides, most of these solutions do not support authorized recording, which restricts their usage scenarios. In this paper, we design an efficient yet robust system that can jam microphones while preserving authorized recording. Specifically, we propose a novel phoneme-based noise with the idea of informational masking, which can distract both machines and humans and is resistant to denoising techniques. Besides, we optimize the noise transmission strategy for broader coverage and implement a hardware prototype of our system. Experimental results show that our system can reduce the recognition accuracy of recordings to below 50% under all tested speech recognition systems, which is much better than existing solutions. Yao Wei 0002, Peng Cheng 0007, Zhongjie Ba, Li Lu 0008, Feng Lin 0004, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | ACL: Account Linking in Online Social Networks With Robust Camera Fingerprint MatchingabstractPseudonyms used in Online social networks (OSNs) post a great challenge to fighting against cyber crimes. To build a strong case, law enforcement may want to link multiple user accounts with pseudonyms to a physical suspect. Images based camera fingerprinting has been used for account linking when a suspect takes pictures and videos with his camera and posts them online. However, image post-processing software may introduce noise into images. This noise is hard to eliminate by conventional strategies, is partly resident in the estimated photo-response non-uniformity (PRNU) fingerprints, and interferes with matching fingerprints. We define this noise as software noise, which pollutes PRNU fingerprints and affects accounts linking in online social networks. In this article, we propose new approaches for camera fingerprint matching given software noise. The key idea is to determine the PRNU hardware noise correlation component with our new test statistic–fingerprinttosoftware noise ratio (FITS). We performed extensive experiments and 10,000+ images taken by 90+ smartphones were used to validate our robust camera fingerprint matching system. The experiment results show FITS outperforms the state-of-the-art approaches for polluted fingerprints. This is the first work studying camera fingerprint matching with the presence of software noise. Xinwen Fu, Zhongjie Ba, Feng Lin 0004, Li Lu 0008, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2025 | MagShadow: Physical Adversarial Example Attacks via Electromagnetic InjectionabstractPhysical adversarial examples (AEs) have become an increasing threat to deploying deep neural network (DNN) models in the real world. Popular approaches adopt sticking-based or projecting-based strategies that stick the printed adversarial patches to objects or directly project the AE onto objects. Although effective, these methods require access to target objects and generate visible artifacts, which reduces the attack's stealthiness. In this article, we propose MagShadow, a new attack vector that leverages imperceptible electromagnetic (EM) signals to realize physical AEs. MagShadow utilizes the CCD camera sensor's susceptibility to EM injection attacks and induces fine-grained adversarial perturbations on the camera's captured image by injecting carefully-crafted signals with a low-cost portable device. As MagShadow directly manipulates the image sensor's output with EM signals, the attack requires no access to the target object and can keep stealthy. We study the feasibility of MagShadow in two typical DNN application scenarios (image classification and object detection) and design a framework to implement four different types of attacks, i.e., untargeted, targeted, hiding, and appearing attacks. Extensive real-world experiments on five different cameras are conducted, which demonstrate MagShadow's effectiveness against different popular DNN models (Inception v3, ResNet101, YOLO v3/v4). Ziwei Liu 0007, Feng Lin 0004, Zhongjie Ba, Li Lu 0008, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Liquid Crystal Mimics Your Heart: A Physical Spoofing Attack Against PPG-Based SystemsabstractPhotoplethysmography (PPG) has been extensively employed in commercial and medical products to assess human cardiac activities. However, despite PPG’s active role in improving people’s daily lives, research on the vulnerabilities of PPG systems is still in its infancy. This paper investigates the feasibility of deceiving PPG sensors in the physical domain. We propose FakePPG, which utilizes a low-cost Liquid Crystal Modulator (LCM) device to mimic the PPG signals of a legitimate user, thus deceiving both the PPG-based health assessment and potential authentication applications. To implement FakePPG in practical scenarios, we build the attack prototype using commercial off-the-shelf electronic components and further design an automated optimization and attack framework. By leveraging the modified multi-Gaussian model for parameterization, the evolutionary strategy for optimization, and the reference heart rate model for heartbeat variability alignment, FakePPG can achieve efficient, flexible, and automated PPG forgery against arbitrary users and heart states. Extensive experimental results show that FakePPG can achieve a success rate of 96.7% for Atrial Fibrillation (AFib) spoofing and 91.2% for identity spoofing, respectively, revealing a realistic threat to PPG systems. Li Lu 0008, Hao Kong 0004, Feng Lin 0004, Zhongjie Ba, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | DroneAudioID: A Lightweight Acoustic Fingerprint-Based Drone Authentication System for Secure Drone DeliveryabstractWith the increasing accessibility of drones, they have been warmly embraced across various sectors, especially in low-altitude logistics transportation. However, during drone delivery, legal drones dispatched by logistics companies are susceptible to malicious attacks, resulting in package theft or substitution. To address this, existing works focus on designing drone authentication to secure drone delivery. However, most of these methods require expensive specialized equipment, such as high-quality microphones and professional recording devices, resulting in high real-world application costs. In this paper, we propose DroneAudioID, a lightweight acoustic fingerprint-based drone authentication system that relies solely on common mobile devices. The basic idea is to employ acoustic fingerprints to authenticate different drones of the same model based on differences in fundamental frequency and harmonic components of drone audio. Specifically, the drone audio is recorded by a mobile device instead of sophisticated equipment. We apply wavelet transform to remove high-frequency noise during data preprocessing. Then, specialized filter banks are designed for feature extraction, leveraging the frequency characteristics of drone audio. Finally, we construct a Bi-Long Short-Term Memory (Bi-LSTM) with an Open-Max model for open-set classification. Extensive experiments are conducted on eight crafts drones of$DJI Mini2$, showing an authentication accuracy of 99.6%. A series of comprehensive experiments further validate DroneAudioID’s capability to defend against various attacks. Meng Zhang 0022, Li Lu 0008, Zheng Yan 0002, Feng Lin 0004, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | mmZeAR: Zero-Effort Cross-Category Action Recognition With mmWave RadarabstractDespite the widespread application of radio frequency (RF) signal-based human action recognition, traditional solutions can only recognize seen categories and the perception scope is restrained by the limited activity classes. When a novel category emerges, the model needs to be optimized again on additionally collected samples at the cost of computation and labor burden. To address this challenge, we develop the mmZeAR system, which learns semantic knowledge from available vision data as class attributes and then transforms the classification into a matching problem. Specifically, we build the attribute space by fusing the coarse-grained video classification features and fine-grained angle change features of 3D joint skeletons. Then we design an efficient feature extraction backbone named TriSqN, which integrates triple radar heatmaps into the final representations by sufficiently exploring the heterogeneous and complementary characteristics. Finally, a projection network is developed between semantic attributes and radar features to construct indirect relationships between samples and labels. By implementing mmZeAR on millimeter wave (mmWave) radar signal datasets, our extensive experiments have demonstrated its remarkable recognition accuracy in novel category recognition with zero effort and achieved state-of-the-art performance. Biyun Sheng, Jiabin Li, Yiping Zuo, Li Lu 0008, Fu Xiao 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Exposing the Deception: Uncovering More Forgery Clues for Deepfake DetectionabstractDeepfake technology has given rise to a spectrum of novel and compelling applications. Unfortunately, the widespread proliferation of high-fidelity fake videos has led to pervasive confusion and deception, shattering our faith that seeing is believing. One aspect that has been overlooked so far is that current deepfake detection approaches may easily fall into the trap of overfitting, focusing only on forgery clues within one or a few local regions. Moreover, existing works heavily rely on neural networks to extract forgery features, lacking theoretical constraints guaranteeing that sufficient forgery clues are extracted and superfluous features are eliminated. These deficiencies culminate in unsatisfactory accuracy and limited generalizability in real-life scenarios. In this paper, we try to tackle these challenges through three designs: (1) We present a novel framework to capture broader forgery clues by extracting multiple non-overlapping local representations and fusing them into a global semantic-rich feature. (2) Based on the information bottleneck theory, we derive Local Information Loss to guarantee the orthogonality of local representations while preserving comprehensive task-relevant information. (3) Further, to fuse the local representations and remove task-irrelevant information, we arrive at a Global Information Loss through the theoretical analysis of mutual information. Empirically, our method achieves state-of-the-art performance on five benchmark datasets. Our code is available at https://github.com/QingyuLiu/Exposing-the-Deception, hoping to inspire researchers. Zhongjie Ba, Zhenguang Liu, Shuang Wu 0002, Feng Lin 0004, Li Lu 0008, Kui Ren 0001 |
AAAI | 6 |
| 2024 | mmEar: Push the Limit of COTS mmWave Eavesdropping on HeadphonesabstractRecent years have witnessed a surge of headphones (including in-ear headphones) usage in works and communications. Because of the privacy-preserve property, people feel comfortable having confidential communication wearing headphones and pay little attention to speech leakage. In this paper, we present an end-to-end eavesdropping system, mmEar, which shows the feasibility of launching an eavesdropping attack on headphones leveraging a commercial mmWave radar. Different from previous works that realize eavesdropping by sensing speech-induced vibrations with reasonable amplitude, mmEar focuses on capturing the extremely faint vibrations with a low signal-to-noise ratio (SNR) on the surface of headphones. Toward this end, we propose a faint vibration emphasis (FVE) method that models and amplifies the mmWave responses to speech-induced vibrations on the In-phase and Quadrature (IQ) plane, followed by a deep denoising network to further improve the SNR. To achieve practical eavesdropping on various headphones and setups, we propose a cGAN model with a pretrain-finetune scheme, boosting the generalization ability and robustness of the attack by generating high-quality synthesis data. We evaluate mmEar with extensive experiments on different headphones and earphones and find that most of them can be compromised by the proposed attack for speech recovery. Xiangyu Xu 0001, Zhen Ling 0001, Li Lu 0008, Junzhou Luo, Xinwen Fu |
INFOCOM | 4 |
| 2024 | Conan's Bow Tie: A Streaming Voice Conversion for Real-Time VTuber LivestreamingabstractRecent years have witnessed a dramatic growing trend of Virtual YouTubers (VTubers) as a new business on social media, such as YouTube, Twitch, and TikTok. However, a significant challenge arises when VTuber voice actors face health issues or retire, jeopardizing the continuity of their avatar’s recognizable voices. A potential solution reminiscent of Conan’s Bow Tie voice changer in the popular animation Case Closed (i.e., Detective Conan) has inspired our work. To make this a reality, we introduce VTuberBowTie, a user-friendly streaming voice conversion system for real-time VTuber livestreaming. We propose an innovative streaming voice conversion approach that tackles the challenges of limited context modeling and bidirectional context dependence inherent to conventional real-time voice conversion. Rather than individually processing the voice stream in data chunks, our approach adopts a fully sequential structure that leverages contextual information preceding the input chunk, thereby expanding the perceptual range and enabling seamless concatenation. Moreover, we developed a ready-to-use interaction interface for VTuberBowTie and deployed it on various computing platforms. The experimental results show that VTuberBowTie can achieve high-quality voice conversion in a streaming manner with a latency of 179.1ms on CPU and 70.8ms on GPU while providing users a friendly interactive experience. Qianniu Chen, Zhehan Gu, Li Lu 0008, Xiangyu Xu 0001, Zhongjie Ba, Feng Lin 0004, Zhenguang Liu, Kui Ren 0001 |
IUI | 3 |
| 2024 | EMTrig: Physical Adversarial Examples Triggered by Electromagnetic Injection towards LiDAR PerceptionabstractLiDAR sensors measure the environment by emitting lasers and, when combined with deep neural networks (DNNs), can effectively identify surrounding obstacles such as vehicles and pedestrians. Given its crucial role in autonomous driving perception, the security of LiDAR is closely tied to driving safety. Some studies have explored its vulnerabilities to physical-world attacks, such as laser-based attacks or adversarial objects. However, these methods are either extremely difficult to execute or lack stealth and flexibility. In this paper, we propose a novel attack method called EMTrig, which leverages common roadside objects combined with controlled intentional electromagnetic interference (IEMI) targeting specific LiDARs to create flexible and covert adversarial attacks against designated vehicles. This causes the victim vehicle to misidentify roadside objects as obstacles, such as other vehicles, leading to dangerous driving behaviors like sudden stops and lane changes. Unlike conventional adversarial examples, our deployed objects are common items (e.g., signboards) that are harmless without the IEMI trigger but pose a threat only under IEMI attacks, providing better stealthiness and flexibility. Extensive experiments in both digital and physical domains validate the effectiveness of EMTrig, demonstrating its significant threat to LiDAR perception. Ziwei Liu 0007, Feng Lin 0004, Teshi Meng, Benaouda Chouaib Baha-eddine, Li Lu 0008, Kui Ren 0001 |
SenSys | 5 |
| 2024 | ALIF: Low-Cost Adversarial Audio Attacks on Black-Box Speech Platforms using Linguistic FeaturesabstractExtensive research has revealed that adversarial examples (AE) pose a significant threat to voice-controllable smart devices. Recent studies have proposed black-box adversarial attacks that require only the final transcription from an automatic speech recognition (ASR) system. However, these attacks typically involve many queries to the ASR, resulting in substantial costs. Moreover, AE-based adversarial audio samples are susceptible to ASR updates. In this paper, we identify the root cause of these limitations, namely the inability to construct AE attack samples directly around the decision boundary of deep learning (DL) models. Building on this observation, we propose ALIF, the first black-box adversarial linguistic feature-based attack pipeline. We leverage the reciprocal process of text-to-speech (TTS) and ASR models to generate perturbations in the linguistic embedding space where the decision boundary resides. Based on the ALIF pipeline, we present the ALIF-OTL and ALIF-OTA schemes for launching attacks in both the digital domain and the physical playback environment on four commercial ASRs and voice assistants. Extensive evaluations demonstrate that ALIF-OTL and -OTA significantly improve query efficiency by 97.7% and 73.3%, respectively, while achieving competitive performance compared to existing methods. Notably, ALIF-OTL can generate an attack sample with only one query. Furthermore, our test-of-time experiment validates the robustness of our approach against ASR updates. Peng Cheng 0007, Yuwei Wang 0009, Zhongjie Ba, Xiaodong Lin 0001, Feng Lin 0004, Li Lu 0008, Kui Ren 0001 |
SP | 7 |
| 2024 | MicGuard: A Comprehensive Detection System against Out-of-band Injection Attacks for Different Level Microphone-based Devices
Tiantian Liu 0002, Feng Lin 0004, Zhongjie Ba, Li Lu 0008, Zhan Qin, Kui Ren 0001 |
USENIX Security Symposium | 4 |
| 2024 | FraudWhistler: A Resilient, Robust and Plug-and-play Adversarial Example Detection Method for Speaker Recognition
Kun Wang 0025, Xiangyu Xu 0001, Li Lu 0008, Zhongjie Ba, Feng Lin 0004, Kui Ren 0001 |
USENIX Security Symposium | 3 |
| 2024 | Devil in the Room: Triggering Audio Backdoors in the Physical World
Meng Chen 0011, Xiangyu Xu 0001, Li Lu 0008, Zhongjie Ba, Feng Lin 0004, Kui Ren 0001 |
USENIX Security Symposium | 3 |
| 2024 | UniAP: Protecting Speech Privacy With Non-Targeted Universal Adversarial PerturbationsabstractUbiquitous microphones on smart devices considerably raise users’ concerns about speech privacy. Since the microphones are primarily controlled by hardware/software developers, profit-driven organizations can easily collect and analyze individuals’ daily conversations on a large scale with deep learning models, and users have no means to stop such privacy-violating behavior. In this article, we propose UniAP to empower users with the capability of protecting their speech privacy from the large-scale analysis without affecting their routine voice activities. Based on our observation of the recognition model, we utilize adversarial learning to generate quasi-imperceptible perturbations to disturb speech signals captured by nearby microphones, thus obfuscating the recognition results of recordings into meaningless contents. As validated in experiments, our perturbations can protect user privacy regardless of what users speak and when they speak. The jamming performance stability is further improved by training optimization. Additionally, the perturbations are robust against noise removal techniques. Extensive evaluations show that our perturbations achieve successful jamming rates of more than 87% in the digital domain and at least 90% and 70% for common and challenging settings, respectively, in the real-life chatting scenario. Moreover, our perturbations, solely trained on DeepSpeech, exhibit good transferability over other models based on similar architecture. Peng Cheng 0007, Yuexin Wu, Yuan Hong 0001, Zhongjie Ba, Feng Lin 0004, Li Lu 0008, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | High-Quality Speech Recovery Through Soundproof Protections via mmWave SensingabstractOnline voice communications are widely used nowadays. To protect speech from leakage, people tend to initiate the talk in sound-isolated environments. In this paper, we reveal a novel attack that recovers high-quality speech from outside soundproof zones. The rationale of the attack is to leverage sound-sensitive characteristics of piezoelectric materials, i.e., a piezo film that can change the phase of reflected mmWaves when placed in a sound field. If the attacker transmits mmWaves and analyzes reflected signals from the piezo film, the speech information can be compromised. More importantly, the piezo film is paper-like and works without a power supply. We propose a new speech recovery methodology to transform sound waves into wireless signals and build an end-to-end eavesdropping system working as a through-wall “microphone” to recover high-quality speech stealthily. To combat signal attenuation and improve speech quality, we develop a speech-enhancement scheme based on generative adversarial networks and propose to use multi-antenna information for intelligible speech reconstruction. We conduct extensive experiments to evaluate the system. The results indicate that the system achieves over 98% accuracy for digit recognition and works well over 5m away through the wall. We also test the system under complex scenarios and give countermeasures. Feng Lin 0004, Chao Wang 0097, Tiantian Liu 0002, Ziwei Liu 0007, Yijie Shen, Zhongjie Ba, Li Lu 0008, Wenyao Xu, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2024 | MotoPrint: Reconfigurable Vibration Motor Fingerprint via Homologous Signals LearningabstractDevice fingerprints can satisfy the high-security requirement of modern mobile applications (e.g., mobile payments) by guaranteeing the operation is performed on a trusted device. However, existing works on device fingerprints are weak to leakage, which leads to an irreversible failure of the device fingerprint authentication system after suffering from fingerprint theft attacks. The vulnerability drives us to propose a reconfigurable device fingerprint, i.e.,MotoPrint, that can recover the system after suffering from such attacks.MotoPrintstems from the motor vibration that can represent in both signals of the accelerometer and the gyroscope (i.e., they are homologous motion signals). Therefore, we designed a two-path feature extracting network and a sensor-independent training strategy to eliminate sensor noise that can decline authentication performance. In addition,MotoPrinthas a complete reconfiguration mechanism to cope with fingerprint leakage, which brings the damaged authentication system back to health. The evaluation of 80 stand-alone vibration motors and 20 in-built ones shows thatMotoPrintcan achieve high authentication accuracy of 98.5%. Meanwhile, we also demonstrate the reconfiguredMotoPrint, which can also effectively indicate the device's uniqueness with over 98% accuracy, is independent ofMotoPrints under other stimulating codes. Yijie Shen, Feng Lin 0004, Chao Wang 0097, Tiantian Liu 0002, Zhongjie Ba, Li Lu 0008, Wenyao Xu, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | Indelible "Footprints" of Inaudible Command InjectionabstractInaudible command injection transmits inaudible ultrasounds to inject adversarial speech commands into a voice assistant, therefore manipulating voice control systems (e.g., a garage door or a security camera) for illegitimate purposes. Although the attack is inaudible, we find it does leave visible “footprints”. Such attack “footprints” are the side product due to the interaction between the attack signal (i.e., input) and the acoustic components (i.e., transfer function), so they reflect the hardware characteristics of the sound capture system, including the microphone diaphragm, the low-pass filter, and the analog-to-digital converter. Moreover, unlike the non-linearity distortion that is erasable with signal-shaping techniques, the “footprints” are indelible because they are unrelated to the content of injected commands. We discover two types of indelible “footprints” embedded in the recording spectrogram, namely abnormal interfering noise and abnormal demodulation. A software-based detection method and a portable detector, DolphinTag, are further designed to identify these “footprints”. The software-based method achieves a detection accuracy of 99.8% on the phone models exhibiting abnormal interfering noise, and our DolphinTag achieves 100% detection accuracy which detects the ultrasound attack by actively facilitating the abnormal demodulation. Zhongjie Ba, Bin Gong 0001, Yuwei Wang 0009, Peng Cheng 0007, Feng Lin 0004, Li Lu 0008, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | AdvReverb: Rethinking the Stealthiness of Audio Adversarial Examples to Human PerceptionabstractAs one of the most representative applications built on deep learning, audio systems, including keyword spotting, automatic speech recognition, and speaker identification, have recently been demonstrated to be vulnerable to adversarial examples, which have already raised general concerns in both academia and industry. Existing attacks follow the same adversarial example generation paradigm from computer vision, i.e., overlaying the optimized additive perturbations on original voices. However, due to the additive perturbations’ nature on human audibility, balancing the stealthiness and attack capability remains a challenging problem. In this paper, we rethink the stealthiness of audio adversarial examples and turn to introduce another kind of audio distortion, i.e., reverberation, as a new perturbation format for stealthy adversarial example generation. Such convolutional adversarial perturbations are crafted as real-world impulse responses and behave as a natural reverberation for deceiving humans. Based on this idea, we propose AdvReverb to construct, optimize, and deliver phoneme-level convolutional adversarial perturbations on both speech and music carriers with a well-designed objective. Experimental results demonstrate that AdvReverb could realize high attack success rates over 95% on three audio-domain tasks while achieving superior perceptual quality and keeping stealthy from human perception in over-the-air and over-the-line delivery scenarios. Meng Chen 0011, Li Lu 0008, Jiadi Yu, Zhongjie Ba, Feng Lin 0004, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | PhaDe: Practical Phantom Spoofing Attack Detection for Autonomous VehiclesabstractDespite their prevalence and indispensability in the perception modules of autonomous vehicles, cameras have shown susceptibility to numerous attacks. Among them, the phantom spoofing attack is of significant concern. In such attacks, malefactors employ electronic display devices like projectors and display monitors to generate deceptive objects, thereby duping the object detectors in autonomous vehicles. However, existing detection methodologies are narrowly focused on a single device category, ignoring the multitude of devices that could be leveraged for attacks. Furthermore, the artificial modality-based solution presently in use lacks efficacious fusion mechanisms. In response to these limitations, we propose PhaDe, a practical deep learning-based system adept at detecting phantom spoofing attacks from a variety of and even unfamiliar attack devices. Our approach introduces two image processing techniques to construct artificial modalities and further advances a multi-head self-attention MSA-based fusion module for more versatile integration of disparate modalities. To boost the generalization capacity of our system against novel, unseen attacks, we incorporate two representation-level losses to align feature distributions from various domains. Evaluations conducted on our own dataset, encompassing fake objects from several device types, attest to the efficacy of our system. Our results indicate an accuracy of 98.80% on familiar domains and a detection success rate of 94.03% on unfamiliar domains. Additionally, PhaDe demonstrates a swift response time, fulfilling the practicality requisites. Feng Lin 0004, Jin Li 0033, Ziwei Liu 0007, Li Lu 0008, Zhongjie Ba, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | FLTracer: Accurate Poisoning Attack Provenance in Federated LearningabstractFederated Learning (FL) is a promising distributed learning approach that enables multiple clients to collaboratively train a shared global model. However, recent studies show that FL is vulnerable to various poisoning attacks, which can degrade the performance of global models or introduce backdoors into them. In this paper, we first conduct a comprehensive study on prior FL attacks and detection methods. The results show that all existing detection methods are only effective against limited and specific attacks. Most detection methods suffer from high false positives, which lead to significant performance degradation, especially in not independent and identically distributed (non-IID) settings. To address these issues, we propose FLTracer, the first FL attack provenance framework to accurately detect various attacks and trace the attack time, objective, type, and poisoned location of updates. Different from existing methodologies that rely solely on cross-client anomaly detection, we propose a Kalman filter-based cross-round detection to identify adversaries by seeking the behavior changes before and after the attack. Thus, this makes it resilient to data heterogeneity and is effective even in non-IID settings. To further improve the accuracy of our detection method, we employ four novel features and capture their anomalies with the joint decisions. Extensive evaluations show that FLTracer achieves an average true positive rate of over 96.88% at an average false positive rate of less than 2.67%, significantly outperforming SOTA detection methods (https://github.com/Eyr3/FLTracer). Xinyu Zhang 0016, Zhongjie Ba, Yuan Hong 0001, Tianhang Zheng, Feng Lin 0004, Li Lu 0008, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | A Resilience Evaluation Framework on Ultrasonic Microphone JammersabstractCovert eavesdropping via microphones has always been a major threat to user privacy. Benefiting from the acoustic non-linearity property, the ultrasonic microphone jammer (UMJ) is effective in resisting this long-standing attack. However, prior UMJ researches underestimate adversary's attacking capability in reality and miss critical metrics for a thorough evaluation. The strong assumptions of adversary unable to retrieve information under low word recognition rate, and adversary's weak denoising abilities in the threat model make these works overlook the vulnerability of existing UMJs. As a result, their UMJs' resilience is overestimated. In this paper, we refine the adversary model and completely investigate potential eavesdropping threats. Correspondingly, we define a total of 12 metrics that are necessary for evaluating UMJs' resilience. Using these metrics, we propose a comprehensive framework to quantify UMJs' practical resilience. It fully covers three perspectives that prior works ignored to some degree, i.e., ambient information, semantic comprehension, and collaborative recognition. Guided by this framework, we can thoroughly and quantitatively evaluate the resilience of existing UMJs towards eavesdroppers. Our extensive assessment results reveal that most existing UMJs are vulnerable to sophisticated adverse approaches. We further outline the key factors influencing jammers' performance and present constructive suggestions for UMJs' future designs. Ming Gao 0023, Yike Chen, Lingfeng Zhang 0004, Jianwei Liu 0008, Li Lu 0008, Feng Lin 0004, Jinsong Han, Kui Ren 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | An Imperceptible Eavesdropping Attack on WiFi Sensing SystemsabstractRecent years have witnessed enormous research efforts on WiFi sensing to enable intelligent services of Internet of Things. However, due to the omni-directional broadcasting manner of WiFi signals, the activity semantic underlying the signals can be leaked to adversaries for surveillance, as demonstrated by our previous work. In this paper, we further extend the attack capability ofActListenerto impersonation attack, which could eavesdrop on users’ behavioral uniqueness imperceptibly using a WiFi infrastructure in any location of user sensing area. In particular,ActListener detects each human activityand converts the eavesdropped signals to that by legitimate devices based on our proposed signal propagation models. To extract noise-resilient individual behavioral uniqueness from converted CSI of WiFi signals, we further add user identification models into the substitute model set for training the signal pattern calibration generative model. Experimental results demonstrate thatActListenercould achieve over 80% accuracy in activity semantics retrieval and impersonation by using the converted signals. Li Lu 0008, Meng Chen 0011, Jiadi Yu, Zhongjie Ba, Feng Lin 0004, Jinsong Han, Yanmin Zhu 0006, Kui Ren 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | FITS: Matching Camera Fingerprints Subject to Software Noise PollutionabstractPhysically unclonable hardware fingerprints can be used for device authentication. The photo-response non-uniformity (PRNU) is the most reliable hardware fingerprint of digital cameras and can be conveniently extracted from images. However, we find image post-processing software may introduce extra noise into images. Part of this noise remains in the extracted PRNU fingerprints and is hard to be eliminated by traditional approaches, such as denoising filters. We define this noise as software noise, which pollutes PRNU fingerprints and interferes with authenticating a camera armed device. In this paper, we propose novel approaches for fingerprint matching, a critical step in device authentication, in the presence of software noise. We calculate the cross correlation between PRNU fingerprints of different cameras using a test statistic such as the Peak to Correlation Energy (PCE) so as to estimate software noise correlation. During fingerprint matching, we derive the ratio of the test statistic on two PRNU fingerprints of interest over the estimated software noise correlation. We denote this ratio as the fingerprint to software noise ratio (FITS), which allows us to detect the PRNU hardware noise correlation component in the test statistic for fingerprint matching. Extensive experiments over 10,000 images taken by more than 90 smartphones are conducted to validate our approaches, which outperform the state-of-the-art approaches significantly for polluted fingerprints. We are the first to study fingerprint matching with the existence of software noise. Xinwen Fu, Zhongjie Ba, Feng Lin 0004, Li Lu 0008, Kui Ren 0001 |
CCS | 7 |
| 2023 | Shift to Your Device: Data Augmentation for Device-Independent Speaker Verification Anti-SpoofingabstractThis paper proposes a novel Deconvolution-enhanced data Augmentation method, DeAug, for ultrasonic-based speaker verification anti-spoofing systems to detect the liveness of voice sources in physical access, which aims to improve the performance of liveness detection on unseen devices where no data is collected yet. Specifically, DeAug first employs a wiener deconvolution pre-processing on available collected data to generate enhanced clean signal samples. Then, the generated samples are convolved with different device impulse responses, to enable the signal with the unseen devices' channel characteristics. Experiments on cross-domain datasets show that our proposed augmentation method can improve the performance of ultrasonic-based anti-spoofing systems by 97.8% relatively, and a further improvement of up to 43.4% can be obtained after applying domain adversarial training on multi-device augmented data. Li Lu 0008, Zhongjie Ba, Feng Lin 0004, Kui Ren 0001 |
ICASSP | 2 |
| 2023 | BypTalker: An Adaptive Adversarial Example Attack to Bypass Prefilter-enabled Speaker RecognitionabstractWith the broad integration of deep learning in Speaker Recognition (SR) systems, adversarial example attacks have been a significant threat raising user security concerns. Nevertheless, recent studies demonstrate that using input transformations (e.g., re-quantization, resampling, bandpass filtering) as a low-cost prefilter can efficiently mitigate such adversarial example attacks. These prefilters constrain the injection space of adversarial perturbations in both time and frequency domains, leading to either degraded attack performance or amplified perturbation noise. This paper proposes a new adversarial example attack, BypTalker, which could bypass these prefilter-enabled SR systems while remaining imperceptible to human listeners. BypTalker employs ensemble learning with diverse substitute pre-filters in the training phase to enhance the adversarial example’s adaptiveness to different prefilters. Furthermore, it incorporates an Acoustic Masker to cloak adversarial perturbations based on psychoacoustics effectively. This masker is well selected from a proposed metric M-Sup for minimizing the perturbation’s auditory to human perception. Experimental results show that BypTalker can achieve an Attack Success Rate of 99.1% and a Perceptual Evaluation of Speech Quality of 4.32, respectively. Qianniu Chen, Li Lu 0008, Meng Chen 0011, Zhongjie Ba, Feng Lin 0004, Kui Ren 0001 |
MSN | 3 |
| 2023 | InfoMasker: Preventing Eavesdropping Using Phoneme-Based Noise
Yao Wei 0002, Peng Cheng 0007, Zhongjie Ba, Li Lu 0008, Feng Lin 0004, Fan Zhang 0010, Kui Ren 0001 |
NDSS | 5 |
| 2023 | FingerFaker: Spoofing Attack on COTS Fingerprint Recognition Without Victim's KnowledgeabstractFingerprint recognition has been a vital security guard for various applications whose vulnerability has been explored by different works. However, previous works on spoofing fingerprint recognition rely on prior knowledge (e.g., photos and minutiae) of the target fingerprint, which fails to implement in practical scenarios. In this paper, we design a fingerprint spoofing attack, namely FingerFaker, to explore the vulnerability of fingerprint recognition, which can spoof automated fingerprint recognition systems (AFRSs) without prior knowledge of target fingerprints. Specifically, we propose a novel concept of "pseudo-minutiae-set" as an effective optimization object and design a two-stage scheme to optimize "pseudo-minutiaeset" leveraging a two-factor evolutionary strategy. In addition, we use a GAN-based training strategy with a minutiae loss function to pre-train a fingerprint generator to map a "pseudo-minutiae-set" into a fingerprint. We use 6342 fingerprint images to verify the performance of FingerFaker on spoofing the open-source AFRS, which shows a high attack success rate (ASR) of 97.78%. Meanwhile, we conduct a realistic case study on commercial off-the-shelf (COTS) AFRS, where FingerFaker also shows 94.22% ASR. Finally, we explore the impact of different conditions to guide the attack and propose countermeasures to mitigate the harm. Yijie Shen, Feng Lin 0004, Zhongjie Ba, Li Lu 0008, Wenyao Xu, Kui Ren 0001 |
SenSys | 6 |
| 2023 | MagBackdoor: Beware of Your Loudspeaker as A Backdoor For Magnetic Injection AttacksabstractAn audio system containing loudspeakers and microphones is the fundamental hardware for voice-enabled devices, enabling voice interaction with mobile applications and smart homes. This paper presents MagBackdoor, the first magnetic field attack that injects malicious commands via a loudspeaker-based backdoor of the audio system, compromising the linked voice interaction system. MagBackdoor focuses on the magnetic threat on loudspeakers and manipulates their sound production stealthily. Consequently, the microphone will inevitably pick up malicious sound generated by the attacked speaker, due to the closely packed arrangement of internal audio systems. To prove the feasibility of MagBackdoor, we conduct comprehensive simulations and experiments. This study further models the mechanism by which an external magnetic field excites the sound production of loudspeakers, giving theoretical guidance to MagBackdoor. Aiming at stealthy magnetic attacks in real-world scenarios, we self-design a prototype that can emit magnetic fields modulated by voice commands. We implement MagBackdoor and evaluate it across a wide range of smart devices involving 16 smartphones, four laptops, two tablets, and three smart speakers, achieving an average 95% injection success rate with high-quality injected acoustic signals. Tiantian Liu 0002, Feng Lin 0004, Zhangsen Wang, Chao Wang 0097, Zhongjie Ba, Li Lu 0008, Wenyao Xu, Kui Ren 0001 |
SP | 6 |
| 2023 | Transferring Audio Deepfake Detection Capability across LanguagesabstractThe proliferation of deepfake content has motivated a surge of detection studies. However, existing detection methods in the audio area exclusively work in English, and there is a lack of data resources in other languages. Cross-lingual deepfake detection, a critical but rarely explored area, urges more study. This paper conducts the first comprehensive study on the cross-lingual perspective of deepfake detection. We observe that English data enriched in deepfake algorithms can teach a detector the knowledge of various spoofing artifacts, contributing to performing detection across language domains. Based on the observation, we first construct a first-of-its-kind cross-lingual evaluation dataset including heterogeneous spoofed speech uttered in the two most widely spoken languages, then explored domain adaptation (DA) techniques to transfer the artifacts detection capability and propose effective and practical DA strategies fitting the cross-lingual scenario. Our adversarial-based DA paradigm teaches the model to learn real/fake knowledge while losing language dependency. Extensive experiments over 137-hour audio clips validate the adapted models can detect fake audio generated by unseen algorithms in the new domain. Zhongjie Ba, Qing Wen, Peng Cheng 0007, Yuwei Wang 0009, Feng Lin 0004, Li Lu 0008, Zhenguang Liu |
WWW | 6 |
| 2023 | Toward Multi-User Authentication Using WiFi SignalsabstractUser authentication nowadays has become an important support for not only security guarantees but also emerging novel applications. Although WiFi signal-based user authentication has achieved initial success, it works in single-user scenarios while multi-user authentication remains a challenging task. In this paper, we present MultiAuth, a multi-user authentication system that can authenticate multiple users with a single pair of commodity WiFi devices. The basic idea is to profile multipath components of WiFi signals, and leverage the multipath components to characterize each user individually for multi-user authentication. MultiAuth first profiles multipath components of WiFi signals through a proposed MUltipath Time-of-Arrival estimation algorithm (MUTA). Then, after matching corresponding multipath components to each user in complex multi-user scenarios, MultiAuth constructs individual CSI based on the multipath components to characterize each user individually. An AoA-based approach is exploited to further separate individual CSI constructed by the users with same ToA. To identify users through their activities, MultiAuth extracts user behavior profiles based on the individual CSI, and leverages a dual-task neural network for robust user authentication. Extensive experiments involving 3 simultaneously present users demonstrate that MultiAuth is effective in multi-user authentication with 86.2% average accuracy and 9.5% average false accept rate. Hao Kong 0004, Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Xiangyu Xu 0001, Feng Lyu 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | ActListener: Imperceptible Activity Surveillance by Pervasive Wireless InfrastructuresabstractRecent years have witnessed enormous research efforts on WiFi sensing to enable intelligent services of Internet of Things. However, due to the omni-directional broadcasting manner of WiFi signals, the activity semantic underlying the signals is leaked to adversaries for surveillance in all probability. To reveal the threat, this paper demonstrates ActListener, which could eavesdrop on user activities imperceptibly using a WiFi infrastructure in any location of user sensing area. The proposed attack requires no direct physical access to the victim user’s devices and prior knowledge of activity recognition model details and device locations. In particular, ActListener first detects the signal segment induced by each human activity, and estimates the locations of legitimate devices and the victim users relative to the adversary’s device for further signal modeling. Then, ActListener models propagating WiFi signals to construct the relationship between physical locations and received signals, and converts the eavesdropped signals to that by legitimate devices based on the models. Furthermore, a neural network-based generative model is designed to calibrate the converted signals for resisting noises in over-the-air WiFi signals. Experiments show ActListener achieves 88.4% average α-similarity on recovering originally signals from eavesdropped ones, and over 90% accuracy in activity recognition. Li Lu 0008, Zhongjie Ba, Feng Lin 0004, Jinsong Han, Kui Ren 0001 |
ICDCS | 1 |
| 2022 | Big Brother is Listening: An Evaluation Framework on Ultrasonic Microphone JammersabstractCovert eavesdropping via microphones has always been a major threat to user privacy. Benefiting from the acoustic non-linearity property, the ultrasonic microphone jammer (UMJ) is effective in resisting this long-standing attack. However, prior UMJ researches underestimate adversary’s attacking capability in reality and miss critical metrics for a thorough evaluation. The strong assumptions of adversary unable to retrieve information under low word recognition rate, and adversary’s weak denoising abilities in the threat model make these works overlook the vulnerability of existing UMJs. As a result, their UMJs’ resilience is overestimated. In this paper, we refine the adversary model and completely investigate potential eavesdropping threats. Correspondingly, we define a total of 12 metrics that are necessary for evaluating UMJs’ resilience. Using these metrics, we propose a comprehensive framework to quantify UMJs’ practical resilience. It fully covers three perspectives that prior works ignored in some degree, i.e., ambient information, semantic comprehension, and collaborative recognition. Guided by this framework, we can thoroughly and quantitatively evaluate the resilience of existing UMJs towards eavesdroppers. Our extensive assessment results reveal that most existing UMJs are vulnerable to sophisticated adverse approaches. We further outline the key factors influencing jammers’ performance and present constructive suggestions for UMJs’ future designs. Yike Chen, Ming Gao 0023, Lingfeng Zhang 0004, Li Lu 0008, Feng Lin 0004, Jinsong Han, Kui Ren 0001 |
INFOCOM | 5 |
| 2022 | PhoneyTalker: An Out-of-the-Box Toolkit for Adversarial Example Attack on Speaker RecognitionabstractVoice has become a fundamental method for human-computer interactions and person identification these days. Benefit from the rapid development of deep learning, speaker recognition exploiting voice biometrics has achieved great success in various applications. However, the shadow of adversarial example attacks on deep neural network-based speaker recognition recently raised extensive public concerns and enormous research interests. Although existing studies propose to generate adversarial examples by iterative optimization to deceive speaker recognition, these methods require multiple iterations to construct specific perturbations for a single voice, which is input-specific, time-consuming, and non-transferable, hindering the deployment and application for non-professional adversaries. In this paper, we propose PhoneyTalker, an out-of-the-box toolkit for any adversary to generate universal and transferable adversarial examples with low complexity, releasing the requirement for professional background and specialized equipment. PhoneyTalker decomposes an arbitrary voice into phone combinations and generates phone-level perturbations using a generative model, which are reusable for voices from different persons with various texts. Experiments on mainstream speaker recognition systems with large-scale corpus show that PhoneyTalker outperforms state-of-the-art methods with overall attack success rates of 99.9% and 84.0% under white-box and black-box settings respectively. Meng Chen 0011, Li Lu 0008, Zhongjie Ba, Kui Ren 0001 |
INFOCOM | 2 |
| 2022 | Push the Limit of WiFi-based User Authentication towards Undefined GesturesabstractWith the development of smart indoor environments, user authentication becomes an essential mechanism to support various secure accesses. Although recent studies have shown initial success on authenticating users with human activities or gestures using WiFi, they rely on predefined body gestures and perform poorly when meeting undefined body gestures. This work aims to enable WiFi-based user authentication with undefined body gestures rather than only predefined body gestures, i.e., realizing a gesture-independent user authentication. In this paper, we first explore physiological characteristics underlying body gestures, and find that statistical distributions under WiFi signals induced by body gestures can exhibit invariant individual uniqueness unrelated to specific body gestures. Inspired by this observation, we propose a user authentication system, which utilizes WiFi signals to identify individuals in a gesture-independent manner. Specifically, we design an adversarial learning-based model, which suppresses specific gesture characteristics, and extracts invariant individual uniqueness unrelated to specific body gestures, to authenticate users in a gesture-independent manner. Extensive experiments in indoor environments show that the proposed system is feasible and effective in gesture-independent user authentication. Hao Kong 0004, Li Lu 0008, Jiadi Yu, Yanmin Zhu 0006, Feilong Tang 0001, Yingying Chen 0001, Linghe Kong, Feng Lyu 0001 |
INFOCOM | 2 |
| 2022 | mmPhone: Acoustic Eavesdropping on Loudspeakers via mmWave-characterized Piezoelectric EffectabstractMore and more people turn to online voice communication with loudspeaker-equipped devices due to its convenience. To prevent speech leakage, soundproof rooms are often adopted. This paper presents mmPhone, a novel acoustic eavesdropping system that recovers loudspeaker speech protected by soundproof environments. The key idea is that properties of piezoelectric films in mmWave band can change with sound pressure due to the piezoelectric effect. If the property changes are acquired by an adversary (i.e., characterizing the piezoelectric effect with mmWaves), speech leakage can happen. More importantly, the piezoelectric film can work without a power supply. Base on this, we proposed a methodology using mmWaves to sense the film and decoding the speech from mmWaves, which turns the film into a passive "microphone". To recover intelligible speech, we further develop an enhancement scheme based on a denoising neural network, multi-channel augmentation, and speech synthesis, to compensate for the propagation and penetration loss of mmWaves. We perform extensive experiments to evaluate mmPhone and conduct digit recognition with over 93% accuracy. The results indicate mmPhone can recover high-quality and intelligible speech from a distance over 5m and is resilient to incident angles of sound waves (within 55 degrees) and different types of loudspeakers. Chao Wang 0097, Feng Lin 0004, Tiantian Liu 0002, Ziwei Liu 0007, Yijie Shen, Zhongjie Ba, Li Lu 0008, Wenyao Xu, Kui Ren 0001 |
INFOCOM | 7 |
| 2022 | A non-intrusive and adaptive speaker de-identification scheme using adversarial examplesabstractFaced with the threat of identity leakage during voice data publishing, users are engaged in a privacy-utility dilemma while enjoying convenient voice services. Existing studies employ direct modification or text-based re-synthesis to de-identify users' voices, but resulting in inconsistent audibility for human participants and not adaptive to informed attacks. In this poster, we propose a non-intrusive and adaptive speaker de-identification scheme to balance the privacy and utility of voice services. We generate adversarial examples to conceal user identity from exposure by Automatic Speaker Identification (ASI). By learning a compact distribution with a conditional variational auto-encoder, our system enables on-demand target sampling and diverse identity transformation. We also introduce the acoustic masking effect to construct inaudible perturbations, thus preserving the speech content and perceptual quality. Experiments on 50 speakers show our system could achieve 98.2% successful de-identification on 4 mainstream ASIs with an objective perceptual quality of 4.38 and a subjective mean opinion score of 4.56. Meng Chen 0011, Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Zhongjie Ba, Feng Lin 0004, Kui Ren 0001 |
MobiCom | 2 |
| 2022 | Push the Limit of Adversarial Example Attack on Speaker Recognition in Physical DomainabstractThe integration of deep learning on Speaker Recognition (SR) advances its development and wide deployment, but also introduces the emerging threat of adversarial examples. However, only a few existing studies investigate its practical threat in physical domain, which either evaluate its feasibility only by directly replaying generated adversarial examples, or explore the partial channel interference for robustness improvement. In this paper, we propose a physical adversarial example attack, PhyTalker, which could generate and inject perturbations on voices in a live-streaming manner on attacking various SR models in different physical channels. Compared with the typical adversarial example for digital attacks, PhyTalker generates a subphoneme-level perturbation dictionary to decouple the perturbation optimization and injection. Moreover, we introduce the channel augmentation to compensate both device and environmental distortions, as well as model ensemble to improve the perturbation transferability. Finally, PhyTalker recognizes and localizes the latest recorded phoneme to determine the corresponding perturbations for real-time broadcasting. Extensive experiments are conducted with a large-scale corpus in real physical scenarios, and results show that PhyTalker achieves an overall Attack Success Rate (ASR) of 85.5% in attacking mainstream SR systems and Mel Cepstral Distortion (MCD) of 2.45dB in human audibility. Qianniu Chen, Meng Chen 0011, Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Zhibo Wang 0001, Zhongjie Ba, Feng Lin 0004, Kui Ren 0001 |
SenSys | 3 |
| 2021 | PassFace: Enabling Practical Anti-Spoofing Facial Recognition with Camera FingerprintingabstractFacial recognition has become the surge on mobile authentication scenarios and makes up a huge market share for various apps, such as MasterCard, Google Wallet, and AliPay. However, existing solutions suffer from various impersonation attacks, including photo-spoofing attack, video-replay attack, and 3D facial mask attack. State-of-the-art countermeasures either require additional user intervention or introduce specialized high-end sensors. Even introducing these extra efforts, these approaches still hardly defend the latest 3D facial mask attacks, which gradually become accessible due to the prevalence of low-cost 3D printing. In this paper, we propose an anti-spoofing facial recognition system, PassFace, which verifies the smartphone for authentication as the second factor merely using raw facial videos without any user intervention, to defeat impersonation attacks. In particular, when receiving a user’s selfie video, PassFace identifies the user’s face from the video, and meanwhile extracts the highly unique and physically irreproducible camera fingerprint, i.e., Photo Response Non-Uniformity (PRNU), built in the smartphone from key frames of the video. After that, the system compares the Peak to Correlation Energy (PCE) calculated by the estimated PRNU and the reference profile with a threshold for authentication. Experiment results demonstrate PassFace can achieve satisfactory performance in authentication and attack resistance. Hanlin Yu, Zhongjie Ba, Li Lu 0008, Feng Lin 0004, Kui Ren 0001 |
ICC | 4 |
| 2021 | MultiAuth: Enable Multi-User Authentication with Single Commodity WiFi DeviceabstractWith the increasing integration of humans and the cyber world, user authentication becomes critical to support various emerging application scenarios requiring security guarantees. Existing works utilize Channel State Information (CSI) of WiFi signals to capture single human activities for non-intrusive and device-free user authentication, but multi-user authentication remains a challenging task. In this paper, we present a multi-user authentication system, MultiAuth, which can authenticate multiple users with a single commodity WiFi device. The key idea is to profile multipath components of WiFi signals induced by multiple users, and construct individual CSI from the multipath components to solely characterize each user for user authentication. Specifically, we propose a MUltipath Time-of-Arrival measurement algorithm (MUTA) to profile multipath components of WiFi signals in high resolution. Then, after aggregating and separating the multipath components related to users, MultiAuth constructs individual CSI based on the multipath components to solely characterize each user. To identify users, MultiAuth further extracts user behavior profiles based on the individual CSI of each user through time-frequency analysis, and leverages a dual-task neural network for robust user authentication. Extensive experiments involving 3 simultaneously present users demonstrate that MultiAuth is accurate and reliable for multi-user authentication with 87.6% average accuracy and 8.8% average false accept rate. Hao Kong 0004, Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Xiangyu Xu 0001, Feilong Tang 0001, Yi-Chao Chen 0001 |
MobiHoc | 2 |
| 2021 | Enable Traditional Laptops with Virtual Writing Capability Leveraging Acoustic SignalsabstractAbstract Human–computer interaction through touch screens plays an increasingly important role in our daily lives. Besides smartphones and tablets, laptops are the most prevalent mobile devices for both work and leisure. To satisfy the requirements of some applications, it is desirable to re-equip a typical laptop with both handwriting and drawing capability. In this paper, we design a virtual writing tablet system, VPad, for traditional laptops without touch screens. VPad leverages two speakers and one microphone, which are available in most commodity laptops, to accurately track hand movements and recognize writing characters in the air without additional hardware. Specifically, VPad emits inaudible acoustic signals from two speakers in a laptop and then analyzes energy features and Doppler shifts of acoustic signals received by the microphone to track the trajectory of hand movements. Furthermore, we propose a state machine-based trajectory optimization method to correct the unexpected trajectory and employ a stroke direction sequence model based on probability estimation to recognize characters users write in the air. Experimental results show that VPad achieves the average error of 1.55 cm for trajectory tracking and the accuracy over 90% of character recognition merely through built-in audio devices on a laptop. Li Lu 0008, Jian Liu 0001, Jiadi Yu, Yingying Chen 0001, Yanmin Zhu 0006, Linghe Kong, Minglu Li 0001 |
Comput. J. | 1 |
| 2021 | Continuous Authentication Through Finger Gesture Interaction for Smart Homes Using WiFiabstractThe development of smart homes has advanced the concept of user authentication to not only protecting user privacy but also facilitating personalized services to users. Along this direction, we propose to integrate user authentication with human-computer interactions between users and smart household appliances through widely-deployed WiFi infrastructures, which is non-intrusive and device-free. In this paper, we propose$FingerPass$which leverages channel state information (CSI) of surrounding WiFi signals to continuously authenticate users through finger gestures in smart homes.$FingerPass$separates the user authentication process into two stages, login and interaction, to achieve high authentication accuracy and low response latency simultaneously. In the login stage, we develop a deep learning-based approach to extract behavioral characteristics of finger gestures for highly accurate user identification. For the interaction stage, to provide continuous authentication in real time for satisfactory user experience, we design a verification mechanism with lightweight classifiers to continuously authenticate the user’s identity during each interaction of finger gestures. Experiments in real environments show that$FingerPass$can achieve the authentication accuracies of 90.6 percent under in-domain scenarios and 87.6 percent under cross-domain scenarios, as well as$186.6\;ms$response time during interactions. Hao Kong 0004, Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Feilong Tang 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2021 | An Indirect Eavesdropping Attack of Keystrokes on Touch Screen through Acoustic SensingabstractThis paper demonstrates the feasibility of a side-channel attack to infer keystrokes on touch screen leveraging an off-the-shelf smartphone. Although there exist some studies on keystroke eavesdropping attacks on touch screen, they are mainly direct eavesdropping attacks, i.e., require the device of victims compromised to provide side-channel information for the adversary, which are hardly launched in practical scenarios. In this work, we show the practicability of an indirect eavesdropping attack, KeyListener, which infers keystrokes on QWERTY keyboards of touch screen leveraging audio devices on a smartphone. We investigate the attenuation of acoustic signals, and find that a user's keystroke fingers can be localized through the attenuation of acoustic signals received by the microphones in the smartphone. We then utilize the attenuation of acoustic signals to localize each keystroke, and further analyze errors induced by ambient noises. To improve the accuracy of keystroke localization, KeyListener further tracks finger movements during inputs through phase change and Doppler effect to reduce errors of acoustic signal attenuation-based keystroke localization. In addition, a binary tree-based search approach is employed to infer keystrokes in a context-aware manner. The proposed keystroke eavesdropping attack is robust to various environments without the assistance of additional infrastructures. Extensive experiments demonstrate that the accuracy of keystroke inference in top-5 candidates can approach 90 percent with a top-5 error rate of around 6 percent, which is a strong indication of the possible user privacy leakage of inputs on QWERTY keyboard. Jiadi Yu, Li Lu 0008, Yingying Chen 0001, Yanmin Zhu 0006, Linghe Kong |
IEEE Trans. Mob. Comput. | 2 |
| 2020 | BatComm: enabling inaudible acoustic communication with high-throughput for mobile devicesabstractAcoustic communication is an increasingly popular alternative to existing short-range wireless communication technologies for mobile devices, such as NFC and QR codes. Unlike the current standards, there are no requirements for extra hardware, lighting conditions, or Internet connection. However, the audibility and limited throughput of existing studies hinder their deployment on a wide range of applications. In this paper, we aim to redesign acoustic communication mechanism to push the boundary of potential throughput while keeping the inaudibility. Specifically, we propose BatComm, a high-throughput and inaudible acoustic communication system for mobile devices capable of throughput rates 12X higher than contemporary state-of-the-art acoustic communication for mobile devices. We theoretically model the non-linearity of microphone and use orthogonal frequency division multiplexing (OFDM) to transmit data bits over multiple orthogonal channels with an ultrasound frequency carrier. We also design a series of techniques to mitigate interference caused by sources such as the signal's unbalanced frequency response, ambient noise, and unrelated residual signals created through OFDM, amplitude modulation (AM), and related processes. Extensive evaluations under multiple realistic settings demonstrate that our inaudible acoustic communication system can achieve over 47kbps within a 10cm communication range. We also show the possibility of increasing the communication range to room scale (i.e., around 2m) while maintaining high-throughput and inaudibility. Our findings offer a new direction for future inaudible acoustic communication techniques to pursue in emerging mobile and IoT applications. Yang Bai 0009, Jian Liu 0001, Li Lu 0008, Yingying Chen 0001, Jiadi Yu |
SenSys | 3 |
| 2020 | Acoustic-based sensing and applications: A survey
Yang Bai 0009, Li Lu 0008, Jerry Q. Cheng, Jian Liu 0001, Yingying Chen 0001, Jiadi Yu |
Comput. Networks | 2 |
| 2019 | Dynamically Adjusting Scale of a Kubernetes Cluster under QoS GuaranteeabstractNowadays, the container-based virtualization technologies have become very popular due to lightweight nature, scalability, flexibility and others. Kubernetes is one of the most popular container cluster management systems, which enables users to deploy applications on the container easily, so more and more web applications are deployed in a Kubernetes clusters. However, a Kubernetes cluster is generally designed to handle the peak of workloads, so that most of resources are idle in usual time, which results in an huge waste of resource. Hence, it is necessary to design a system to improve the cluster resource utilization and promise Quality of Service (QoS) in a Kubernetes cluster. In this paper, we propose a generic system to dynamically adjust the scale of a Kubernetes cluster, which is able to reduce the waste of resource on the premise of QoS guarantee. The proposed system contains four modules: monitor module, QoS module, scaling module, and executing module. First, the monitor module uses two open-source tools, Heapster and InfluxDB, to monitor and store real-time status of a Kubernetes cluster. Then, to guarantee QoS in the Kubernetes cluster, the QoS module presents a method to automatically decide a threshold of CPU utilization that is able to meet requirements of a specific application. Next, the scaling module provides a cluster scaling algorithm to get an ideal number of nodes in the Kubernetes cluster, which is used to allocate resources in a cluster-level allocation. Finally, according to the ideal number of nodes, the executing module adjusts the scale of the Kubernetes cluster to carry out the application. Extensive experiments in real environments show that, on average, the proposed system improves CPU utilization of a Kubernetes cluster by 28.99%. Jiadi Yu, Li Lu 0008, Shiyou Qian, Guangtao Xue |
ICPADS | 3 |
| 2019 | KeyListener: Inferring Keystrokes on QWERTY Keyboard of Touch Screen through Acoustic SignalsabstractThis paper demonstrates the feasibility of a side-channel attack to infer keystrokes on touch screen leveraging an off-the-shelf smartphone. Although there exist some studies on keystroke eavesdropping attacks on touch screen, they are mainly direct eavesdropping attacks, i.e., require the device of victims compromised to provide side-channel information for the adversary, which are hardly launched in practical scenarios. In this work, we show the practicability of an indirect eavesdropping attack, KeyListener, which infers keystrokes on QWERTY keyboards of touch screen leveraging audio devices on a smartphone. We investigate the attenuation of acoustic signals, and find that a user's keystroke fingers can be localized through the attenuation of acoustic signals received by the microphones in the smartphone. We then utilize the attenuation of acoustic signals to localize each keystroke, and further analyze errors induced by ambient noises. To improve the accuracy of keystroke localization, KeyListener further tracks finger movements during inputs through phase change and Doppler effect to reduce errors of acoustic signal attenuation-based keystroke localization. In addition, a binary tree-based search approach is employed to infer keystrokes in a context-aware manner. The proposed keystroke eavesdropping attack is robust to various environments without the assistance of additional infrastructures. Extensive experiments demonstrate that the accuracy of keystroke inference in top-5 candidates can approach 90% with a top-5 error rate of around 6%, which is a strong indication of the possible user privacy leakage of inputs on QWERTY keyboard. Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Yanmin Zhu 0006, Xiangyu Xu 0001, Guangtao Xue, Minglu Li 0001 |
INFOCOM | 1 |
| 2019 | Poster: Inaudible High-throughput Communication Through Acoustic SignalsabstractIn recent decades, countless efforts have been put into the research and development of short-range wireless communication, which offers a convenient way for numerous applications (e.g., mobile payments, mobile advertisement). Regarding the design of acoustic communication, throughput and inaudibility are the most vital aspects, which greatly affect available applications that can be supported and their user experience. Existing studies on acoustic communication either use audible frequency band (e.g., <20kHz) to achieve a relatively high throughput or realize inaudibility using near-ultrasonic frequency band (e.g., 18-20kHz) which however can only achieve limited throughput. Leveraging the non-linearity of microphones, voice commands can be demodulated from the ultrasound signals, and further recognized by the speech recognition systems. In this poster, we design an acoustic communication system, which achieves high-throughput and inaudibility at the same time, and the highest throughput we achieve is over 17x higher than the state-of-the-art acoustic communication systems. Yang Bai 0009, Jian Liu 0001, Yingying Chen 0001, Li Lu 0008, Jiadi Yu |
MobiCom | 4 |
| 2019 | FingerPass: Finger Gesture-based Continuous User Authentication for Smart Homes Using Commodity WiFiabstractThe development of smart homes has advanced the concept of user authentication to not only protecting user privacy but also facilitating personalized services to users. Along this direction, we propose to integrate user authentication with human-computer interactions between users and smart household appliances through widely-deployed WiFi infrastructures, which is non-intrusive and device-free. In this paper, we propose FingerPass which leverages channel state information (CSI) of surrounding WiFi signals to continuously authenticate users through finger gestures in smart homes. We investigate CSI of WiFi signals in depth and find CSI phase can be used to capture and distinguish the unique behavioral characteristics from different users. FingerPass separates the user authentication process into two stages, login and interaction, to achieve high authentication accuracy and low response latency simultaneously. In the login stage, we develop a deep learning-based approach to extract behavioral characteristics of finger gestures for highly accurate user identification. For the interaction stage, to provide continuous authentication in real time for satisfactory user experience, we design a verification mechanism with lightweight classifiers to continuously authenticate the user's identity during each interaction of finger gestures. Experiments in real environments show that FingerPass can achieve 91.4% authentication accuracy, and 186.6ms response time during interactions. Hao Kong 0004, Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Linghe Kong, Minglu Li 0001 |
MobiHoc | 2 |
| 2019 | WiZoom: Accurate Multipath Profiling using Commodity WiFi Devices with Limited BandwidthabstractMultipath profiling is to characterize multipath components of wireless channels, which can be done using Channel State Information (CSI) from WiFi devices. To do so with satisfactory accuracy, recent studies rely on either a large number of receiving antennas or large bandwidth. However, it is difficult for commodity WiFi devices to meet these requirements. In this paper, we propose a scheme, WiZoom, that can perform accurate multipath profiling using single-band CSI from commodity WiFi devices. In order to achieve accurate multipath profiling with limited bandwidth, WiZoom first incorporates the MUltiple SIgnal Classification (MUSIC) algorithm with CSI to estimate ToAs of multipath components, and then combines multiple antennas to improve the resolution of ToA estimation. WiZoom further estimates attenuations and phase shifts for multipath components using the ToAs. So far, multipath components are fully characterized, and all these estimated parameters form the multipath profile. We evaluate the performance of WiZoom using commodity WiFi devices in real environment, and results show that WiZoom achieves high accuracy in multipath profiling. Jiadi Yu, Yanmin Zhu 0006, Li Lu 0008, Shiyou Qian, Minglu Li 0001 |
SECON | 4 |
| 2019 | Online cost-rejection rate scheduling for resource requests in hybrid clouds
Yanhua Cao, Li Lu 0008, Jiadi Yu, Shiyou Qian, Yanmin Zhu 0006, Minglu Li 0001 |
Parallel Comput. | 2 |
| 2019 | Lip Reading-Based User Authentication Through Acoustic Sensing on SmartphonesabstractTo prevent users privacy from leakage, more and more mobile devices employ biometric-based authentication approaches, such as fingerprint, face recognition, voiceprint authentications, and so on, to enhance the privacy protection. However, these approaches are vulnerable to replay attacks. Although the state-of-art solutions utilize liveness verification to combat the attacks, existing approaches are sensitive to ambient environments, such as ambient lights and surrounding audible noises. Toward this end, we explore liveness verification of user authentication leveraging users mouth movements, which are robust to noisy environments. In this paper, we propose a lip reading-based user authentication system, LipPass, which extracts unique behavioral characteristics of users speaking mouths through acoustic sensing on smartphones for user authentication. We first investigate Doppler profiles of acoustic signals caused by users' speaking mouths and find that there are unique mouth movement patterns for different individuals. To characterize the mouth movements, we propose a deep learning-based method to extract efficient features from Doppler profiles and employ softmax function, support vector machine and support vector domain description to construct multi-class identifier, binary classifiers and spoofer detectors for mouth state identification, user identification and spoofer detection, respectively. Afterward, we develop a balanced binary tree-based authentication approach to accurately identify each individual leveraging these binary classifiers and spoofer detectors with respect to registered users. Through extensive experiments involving 48 volunteers in four real environments, LipPass can achieve 90.2% accuracy in user identification and 93.1% accuracy in spoofer detection. Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Hongbo Liu 0002, Yanmin Zhu 0006, Linghe Kong, Minglu Li 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2018 | VPad: Virtual Writing Tablet for Laptops Leveraging Acoustic SignalsabstractHuman-computer interaction based on touch screens plays an increasing role in our daily lives. Besides smartphones and tablets, laptops are the most popular mobile devices used in both work and leisure. To satisfy requirements of many emerging applications, it becomes desirable to equip both writing and drawing functions directly on laptop screens. In this paper, we design a virtual writing tablet system, VPad, for traditional laptops without touch screens. VPad leverages two speakers and one microphone, which are available in most commodity laptops, for trajectory tracking without additional hardware. It employs acoustic signals to accurately track hand movements and recognize characters user writes in the air. Specifically, VPad emits inaudible acoustic signals from two speakers in a laptop. Then VPad applies Sliding-window Overlap Fourier Transformation technique to find Doppler frequency shift with higher resolution and accuracy in real time. Furthermore, we analyze frequency shifts and energy features of acoustic signals received by the microphone to track the trajectory of hand movements. Finally, we employ a stroke direction sequence model based on possibility estimation to recognize characters users write in the air. Our experimental results show that VPad achieves the average trajectory tracking error of only 1.55cm and the character recognition accuracy of above 90% merely through two speakers and one microphone on a laptop. Li Lu 0008, Jian Liu 0001, Jiadi Yu, Yingying Chen 0001, Yanmin Zhu 0006, Xiangyu Xu 0001, Minglu Li 0001 |
ICPADS | 1 |
| 2018 | LipPass: Lip Reading-based User Authentication on Smartphones Leveraging Acoustic SignalsabstractTo prevent users' privacy from leakage, more and more mobile devices employ biometric-based authentication approaches, such as fingerprint, face recognition, voiceprint authentications, etc., to enhance the privacy protection. However, these approaches are vulnerable to replay attacks. Although state-of-art solutions utilize liveness verification to combat the attacks, existing approaches are sensitive to ambient environments, such as ambient lights and surrounding audible noises. Towards this end, we explore liveness verification of user authentication leveraging users' lip movements, which are robust to noisy environments. In this paper, we propose a lip reading-based user authentication system, LipPass, which extracts unique behavioral characteristics of users' speaking lips leveraging build-in audio devices on smartphones for user authentication. We first investigate Doppler profiles of acoustic signals caused by users' speaking lips, and find that there are unique lip movement patterns for different individuals. To characterize the lip movements, we propose a deep learning-based method to extract efficient features from Doppler profiles, and employ Support Vector Machine and Support Vector Domain Description to construct binary classifiers and spoofer detectors for user identification and spoofer detection, respectively. Afterwards, we develop a binary tree-based authentication approach to accurately identify each individual leveraging these binary classifiers and spoofer detectors with respect to registered users. Through extensive experiments involving 48 volunteers in four real environments, LipPass can achieve 90.21% accuracy in user identification and 93.1% accuracy in spoofer detection. Li Lu 0008, Jiadi Yu, Yingying Chen 0001, Hongbo Liu 0002, Yanmin Zhu 0006, Minglu Li 0001 |
INFOCOM | 1 |
| 2018 | A Double Auction Mechanism to Bridge Users' Task Requirements and Providers' Resources in Two-Sided Cloud MarketsabstractDouble auction-based pricing model is an efficient pricing model to balance users' and providers' benefits. Existing double auction mechanisms usually require both users and providers to bid with the unit price and the number of VMs. However, in practice users seldom know the exact number of VMs that meets their task requirements, which leads to users' task requirements inconsistent with providers' resource. In this paper, we propose a truthful double auction mechanism, including a matching process as well as a pricing and VM allocation scheme, to bridge users' task requirements and providers' resources in two-sided cloud markets. In the matching process, we design a cost-aware resource algorithm based on Lyapunov optimization techniques to precisely obtain the number of VMs that meets users' task requirements. In the pricing and VM allocation scheme, we apply the idea of second-price auction to determine the final price and the number of provisioned VMs in the double auction. We theoretically prove our proposed mechanism is individual-rational, truthful and budget-balanced, and analyze the optimality of proposed algorithm. Through simulation experiments, the results show that the individual profits achieved by our algorithm are 12.35 and 11.02 percent larger than that of scaleout and greedy scale-up algorithms respectively for 90 percent of users, and the social welfare of our mechanism is only 7.01 percent smaller than that of the optimum mechanism in the worst case. Li Lu 0008, Jiadi Yu, Yanmin Zhu 0006, Minglu Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | An Efficient Sampling and Classification Approach for Flow Detection in SDN-Based Big Data CentersabstractSoftware defined networking (SDN) provides flexible management for datacenter networks with the flow-level control. Such the fine-grained management, however, consumes large amount of bandwidth between data and control planes, which results in the bottleneck in the scalability of SDN-based datacenters. "The elephant and mouse phenomenon" suggests that there are only very few elephant flows that carry the majority of bytes in datacenters so that it can improve management efficiency to detect and reroute elephant flows while leaving mice flows in data plane leveraging wildcard flow table in OpenFlow. Unfortunately, existing mechanisms for elephant flow detection suffer from high bandwidth consumption and long detection time. In this paper, we propose an efficient sampling and classification approach (ESCA) with the two-phase elephant flow detection. In the first phase, ESCA improves sampling efficiency by estimating the arrival interval of elephant flows and filtering out redundant samples using a filtering flow table. In the second phase, ESCA classifies samples with a new supervised classification algorithm based on correlation among data flows. The mathematical analysis proofs our ESCA outperforms related schemes. Extensive experiment results on real public datacenter traces further demonstrate that our ESCA can provide accurate detection with less sampled packets and shorter detection time. Feilong Tang 0001, Li Lu 0008, Leonard Barolli, Can Tang |
AINA | 2 |
| 2017 | Cost-efficient VM configuration algorithm in the cloud using mix scaling strategyabstractBenefiting from the pay-per-use pricing model of cloud computing, many companies migrate their services and applications from typical expensive infrastructures to the cloud. However, due to fluctuations in the workload of services and applications, making a cost-efficient VM configuration decision in the cloud remains a critical challenge. Even experienced administrators cannot accurately predict the workload in the future. Since the pricing model of cloud provider is convex other than linear that often assumed in past research, instead of typical scaling out strategy. In this paper, we adopt mix scale strategy. Based on this observation, we model an optimization problem aiming to minimize the VM configuration cost under the constraint of migration delay. Taking advantages of Lyapunov optimization techniques, we propose a mix scale online algorithm which achieves more cost-efficiency than that of scale out strategy. Experimental results shows that the mix scale algorithm saves 30.8% and 31.1% cost where controlling migration delay in a tolerable range under different workload respectively. Li Lu 0008, Jiadi Yu, Yanmin Zhu 0006, Guangtao Xue, Shiyou Qian, Minglu Li 0001 |
ICC | 1 |
| 2017 | Online Cost-Aware Service Requests Scheduling in Hybrid Clouds for Cloud Bursting
Yanhua Cao, Li Lu 0008, Jiadi Yu, Shiyou Qian, Yanmin Zhu 0006, Minglu Li 0001, Jian Cao 0001, Zhong Wang 0013, Juan Li 0011, Guangtao Xue |
WISE (1) | 2 |