Hanqing Guo

dblp:206/6471 · DBLP profile ↗
← Back
36ranked-venue papers
11as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 4 first-author · 7 since 2021Security and privacy · 12 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Yours or Mine? Overwriting Attacks Against Neural Audio Watermarking
abstract
As generative audio models are rapidly evolving, AI-generated audios increasingly raise concerns about copyright infringement and misinformation spread. Audio watermarking, as a proactive defense, can embed secret messages into audio for copyright protection and source verification. However, current neural audio watermarking methods focus primarily on the imperceptibility and robustness of watermarking, while ignoring its vulnerability to security attacks. In this paper, we develop a simple yet powerful attack: the overwriting attack that overwrites the legitimate audio watermark with a forged one and makes the original legitimate watermark undetectable. Based on the audio watermarking information that the adversary has, we propose three categories of overwriting attacks, i.e., white-box, gray-box, and black-box attacks. We also thoroughly evaluate the proposed attacks on state-of-the-art neural audio watermarking methods. Experimental results demonstrate that the proposed overwriting attacks can effectively compromise existing watermarking schemes across various settings and achieve a nearly 100% attack success rate. The practicality and effectiveness of the proposed overwriting attacks expose security flaws in existing neural audio watermarking systems, underscoring the need to enhance security in future audio watermarking designs.
Lingfeng Yao, Chenpei Huang, Shengyao Wang, Junpei Xue, Hanqing Guo, Phone Lin, Tomoaki Ohtsuki, Miao Pan
AAAI5
2026 Sensing-Communication Tradeoffs in Closed-Loop ISAC via Covariance-Aware Optimization
Anindya Bal, Thomas Yang 0003, Minmei Wang, Hanqing Guo, Haofan Cai
INFOCOM4
2026 WebCloak: Characterizing and Mitigating Threats From LLM-Driven Web Agents as Intelligent Scrapers
Xinfeng Li, Tianze Qiu, Yingbin Jin, Lixu Wang, Hanqing Guo, Xiaojun Jia, Xiaofeng Wang 0001, Wei Dong 0007
SP5
2025 ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks
Yuanda Wang, Bocheng Chen, Hanqing Guo, Guangjing Wang 0001, Weikang Ding, Qiben Yan 0001
AsiaCCS3
2025 Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language Models
abstract
The protection of cyber Intellectual Property (IP) such as web content is an increasingly critical concern.The rise of large language models (LLMs) with online retrieval capabilities enables convenient access to information but often undermines the rights of original content creators.As users increasingly rely on LLM-generated responses, they gradually diminish direct engagement with original information sources, which will significantly reduce the incentives for IP creators to contribute, and lead to a saturating cyberspace with more AIgenerated content.In response, we propose a novel defense framework that empowers web content creators to safeguard their web-based IP from unauthorized LLM real-time extraction and redistribution by leveraging the semantic understanding capability of LLMs themselves.Our method follows principled motivations and effectively addresses an intractable black-box optimization problem.Real-world experiments demonstrated that our methods improve defense success rates from 2.5% to 88.6% on different LLMs, outperforming traditional defenses such as configuration-based restrictions.
Yisheng Zhong, Yizhu Wen, Mehran Kafai, Heng Huang 0001, Hanqing Guo, Zhuangdi Zhu
EMNLP6
2025 A Cooperative Bearing-Rate Approach for Observability-Enhanced Target Motion Estimation
abstract
Vision-based target motion estimation is a fundamental problem in many robotic tasks. The existing methods have the limitation of low observability and, hence, face challenges in tracking highly maneuverable targets. Motivated by the aerial target pursuit task where a target may maneuver in 3D space, this paper studies how to further enhance observability by incorporating the bearing rate information that has not been well explored in the literature. The main contribution of this paper is to propose a new cooperative estimator called STT-R (Spatial-Temporal Triangulation with bearing Rate), which is designed under the framework of distributed recursive least squares. This theoretical result is further verified by numerical simulation and real-world experiments. It is shown that the proposed STT-R algorithm can effectively generate more accurate estimations and effectively reduce the lag in velocity estimation, enabling tracking of more maneuverable targets.
Canlun Zheng, Hanqing Guo, Shiyu Zhao 0002
ICRA2
2025 Vision-Based Cooperative MAV-Capturing-MAV
abstract
MAV-capturing-MAV (MCM) is one of the few effective methods for physically countering misused or malicious MAVs. This paper presents a vision-based cooperative MCM system, where multiple pursuer MAVs equipped with onboard vision systems detect, localize, and pursue a target MAV. To enhance robustness, a distributed state estimation and control framework enables the pursuer MAVs to autonomously coordinate their actions. Pursuer trajectories are optimized using Model Predictive Control (MPC) and executed via a low-level SO(3) controller, ensuring smooth and stable pursuit. Once the capture conditions are satisfied, the pursuer MAVs automatically deploy a flying net to intercept the target. These capture conditions are determined based on the predicted motion of the net. To enable real-time decision-making, we propose a lightweight computational method to approximate the net’s motion, avoiding the prohibitive cost of solving the full net dynamics. The effectiveness of the proposed system is validated through simulations and real-world experiments. In real-world tests, our approach successfully captures a moving target traveling at 4 m/s with an acceleration of 1 m/s2, achieving a success rate of 64.7%.
Canlun Zheng, Yize Mi, Hanqing Guo, Huaben Chen, Shiyu Zhao 0002
IROS3
2025 Enabling Joint Sensing and Communication via STBC Assisted NOMA in ISAC Systems
abstract
Integrated Sensing and Communication (ISAC) is a key enabler for Sixth-Generation (6G) and future wireless networks, which seamlessly combines ambient sensing with data communication. In this paper, we propose a novel ISAC-enabled Non-Orthogonal Multiple Access (NOMA) scheme named ISAC-Space-Time Block Coding (STBC) NOMA. We present a thorough performance comparison of our proposed scheme against three previously studied ISAC-NOMA variants: Conventional ISAC-NOMA, ISAC-Unmanned Aerial Vehicle (UAV) NOMA, and ISAC-Generalized Space Shift Keying (GSSK) NOMA, in a multi-user scenario. The comparison specifically focuses on three critical performance metrics: spectral efficiency, Bit Error Rate (BER), and Successive Interference Cancellation (SIC) decoding complexity. The evaluation results show that ISAC-STBC NOMA consistently outperforms the other schemes across all metrics. Specifically, ISAC-STBC NOMA achieves approximately 24% higher spectral efficiency at 30 dB Signal-to-Noise Ratio (SNR), reduces the BER by approximately 30%, and lowers SIC decoding complexity by up to 25%. These findings position ISAC-STBC NOMA as a strong candidate for next-generation networks, offering a well-balanced solution that enhances communication robustness, spectrum utilization, and computational efficiency.
Anindya Bal, Haofan Cai, Hanqing Guo, Yao Zheng 0004, Xiaoxue Zhang 0001
MASS3
2025 AUDIO WATERMARK: Dynamic and Harmless Watermark for Black-box Voice Dataset Copyright Protection
Hanqing Guo, Bocheng Chen, Yuanda Wang, Heng Huang 0001, Qiben Yan 0001, Li Xiao 0001
USENIX Security Symposium1
2024 SwitchTab: Switched Autoencoders Are Effective Tabular Learners
abstract
Self-supervised representation learning methods have achieved significant success in computer vision and natural language processing (NLP), where data samples exhibit explicit spatial or semantic dependencies. However, applying these methods to tabular data is challenging due to the less pronounced dependencies among data samples. In this paper, we address this limitation by introducing SwitchTab, a novel self-supervised method specifically designed to capture latent dependencies in tabular data. SwitchTab leverages an asymmetric encoder-decoder framework to decouple mutual and salient features among data pairs, resulting in more representative embeddings. These embeddings, in turn, contribute to better decision boundaries and lead to improved results in downstream tasks. To validate the effectiveness of SwitchTab, we conduct extensive experiments across various domains involving tabular data. The results showcase superior performance in end-to-end prediction tasks with fine-tuning. Moreover, we demonstrate that pre-trained salient embeddings can be utilized as plug-and-play features to enhance the performance of various traditional classification methods (e.g., Logistic Regression, XGBoost, etc.). Lastly, we highlight the capability of SwitchTab to create explainable representations through visualization of decoupled mutual and salient features in the latent space.
Suiyao Chen, Renat Sergazinov, Chongchao Zhao, Tianpei Xie, Hanqing Guo, Cheng Ji 0005, Daniel Cociorva, Hakan Brunzell
AAAI9
2024 WavePurifier: Purifying Audio Adversarial Examples via Hierarchical Diffusion Models
abstract
In this paper, we propose WavePurifier, an audio purification framework to defend against audio adversarial attacks. Audio adversarial attacks craft adversarial examples or perturbations to attack the automated speech recognition (ASR) models. Although existing defense mechanisms can detect such attacks and raise alarms, they fail to recover or maintain benign commands. Consequently, this leads to the denial of users' benign commands. Different than existing defenses, WavePurifier aims to purify adversarial examples, thereby rectifying the user's benign commands. We find that the forward diffusion process of the diffusion model effectively eliminates perturbations, whereas the reverse diffusion process restores benign speech. Based on this, we develop a hierarchical diffusion model to defend against audio adversarial examples. This model is capable of purifying different spectrogram bands to varying degrees. To validate the performance of WavePurifier, we purify the adversarial examples from 3 different adversarial attacks in 140 distinct settings. In total, we collect 78,864 diffused spectrograms and 21,000 purified audios. Then, we evaluate WavePurifier on 2 different ASR models, 4 commercial speech-to-text APIs, 2 real-world attack scenarios, and compare them against 7 existing defense approaches. Our result shows that WavePurifier is a universal framework, demonstrating adaptability across diverse attacks with the same hyperparameters. Notably, WavePurifier outperforms existing methods with the lowest character error rate (CER), word error rate (WER), and a high purification success rate against different attacks.
Hanqing Guo, Guangjing Wang 0001, Bocheng Chen, Yuanda Wang, Xiao Zhang 0037, Qiben Yan 0001, Li Xiao 0001
MobiCom1
2024 PiezoBud: A Piezo-Aided Secure Earbud with Practical Speaker Authentication
abstract
With the advancement of AI-powered personal voice assistants, speaker authentication via earbuds has become increasingly vital, serving as a critical interface between users and mobile devices. However, existing audio-based speaker authentication methods fail to defend against voice spoofing threats such as replay and deep-fake attacks. To counteract these risks, we introduce PiezoBud, a pioneering multi-modal user authentication system that is truly practical and lightweight for earbuds. PiezoBud uses miniature piezoelectric sensors to detect micro-vibrations on the skin, extracting user-specific biometric data to authenticate legitimate access on the local smartphone and protect against malicious attacks. Our exploratory study, involving 85 participants, demonstrates the effectiveness of PiezoBud in various everyday scenarios, including ambient noise, body movement, and in-ear media playing. Using only 15 seconds of enrollment data, PiezoBud achieves an Equal Error Rate (EER) of 1.05% and attain a mean authentication latency of 0.06 seconds on mobile devices. We also evaluate PiezoBud's effectiveness in countering challenging adaptive attack scenarios and its overall performance in various real-world situations. Our evaluation highlights that PiezoBud stands out as a practical, resilient, responsive, and secure option for earbuds users.
Huaili Zeng, Hanqing Guo, Yidong Ren, Aiden Dixon, Zhichao Cao 0001, Tianxing Li 0001
SenSys3
2024 Motion-guided small MAV detection in complex and non-planar scenes
abstract
In recent years, there has been a growing interest in the visual detection of micro aerial vehicles (MAVs) due to its importance in numerous applications. However, the existing methods based on either appearance or motion features encounter difficulties when the background is complex or the MAV is too small. In this paper, we propose a novel motion-guided MAV detector that can accurately identify small MAVs in complex and non-planar scenes. This detector first exploits a motion feature enhancement module to capture the motion features of small MAVs. Then it uses multi-object tracking and trajectory filtering to eliminate false positives caused by motion parallax . Finally, an appearance-based classifier and an appearance-based detector that operates on the cropped regions are used to achieve precise detection results. Our proposed method can effectively and efficiently detect extremely small MAVs from dynamic and complex backgrounds because it aggregates pixel-level motion features and eliminates false positives based on the motion and appearance features of MAVs. Experiments on the ARD-MAV dataset demonstrate that the proposed method could achieve high performance in small MAV detection under challenging conditions and outperform other state-of-the-art methods across various metrics.
Hanqing Guo, Canlun Zheng, Shiyu Zhao 0002
Pattern Recognit. Lett.1
2024 Global-Local MAV Detection Under Challenging Conditions Based on Appearance and Motion
abstract
Visual detection of micro aerial vehicles (MAVs) has received increasing research attention in recent years due to its importance in many applications. However, the existing approaches based on either appearance or motion features of MAVs still face challenges when the background is complex, the MAV target is small, or the computation resource is limited. In this paper, we propose a global-local MAV detector that can fuse both motion and appearance features for MAV detection under challenging conditions. This detector first searches MAV targets using a global detector and then switches to a local detector which works in an adaptive search region to enhance accuracy and efficiency. Additionally, a detector switcher is applied to coordinate the global and local detectors. A new dataset is created to train and verify the effectiveness of the proposed detector. This dataset contains more challenging scenarios that can occur in practice. Extensive experiments on three challenging datasets show that the proposed detector outperforms the state-of-the-art ones in terms of detection accuracy and computational efficiency. In particular, this detector can run with near real-time frame rate on NVIDIA Jetson NX Xavier, which demonstrates the usefulness of our approach for real-world applications. The dataset is available at https://github.com/WestlakeIntelligentRobotics/GLAD. In addition, A video summarizing this work is available at https://youtu.be/Tv473mAzHbU.
Hanqing Guo, Zhi Gao 0005, Shiyu Zhao 0002
IEEE Trans. Intell. Transp. Syst.1
2023 Federated IoT Interaction Vulnerability Analysis
abstract
IoT devices provide users with great convenience in smart homes. However, the interdependent behaviors across devices may yield unexpected interactions. To analyze the potential IoT interaction vulnerabilities, in this paper, we propose a federated and explicable IoT interaction data management system FexIoT. To address the lack of information in the closed-source platforms, FexIoT captures causality information by fusing multi-domain data, including the descriptions of apps and real-time event logs, into interaction graphs. The interaction graph representation is encoded by graph neural networks (GNNs). To collaboratively train the GNN model without sharing the raw data, we design a layer-wise clustering-based federated GNN framework for learning intrinsic clustering relationships among GNN model weights, which copes with the statistical heterogeneity and the concept drift problem of graph data. In addition, we propose the Monte Carlo beam search with the SHAP method to search and measure the risk of subgraphs, in order to explain the potential vulnerability causes. We evaluate our prototype on datasets collected from five IoT automation platforms. The results show that FexIoT achieves more than 90% average accuracy for interaction vulnerability detection, outperforming the existing methods. Moreover, FexIoT offers an explainable result for the detected vulnerabilities.
Guangjing Wang 0001, Hanqing Guo, Anran Li 0001, Qiben Yan 0001
ICDE2
2023 SCALE-UP: An Efficient Black-box Input-level Backdoor Detection via Analyzing Scaled Prediction Consistency
Yiming Li 0004, Hanqing Guo, Lichao Sun 0001, Cong Liu 0005
ICLR4
2023 MASTERKEY: Practical Backdoor Attack Against Speaker Verification Systems
abstract
Speaker Verification (SV) is widely deployed in mobile systems to authenticate legitimate users by using their voice traits. In this work, we propose a backdoor attack MasterKey, to compromise the SV models. Different from previous attacks, we focus on a real-world practical setting where the attacker possesses no knowledge of the intended victim. To design MasterKey, we investigate the limitation of existing poisoning attacks against unseen targets. Then, we optimize a universal backdoor that is capable of attacking arbitrary targets. Next, we embed the speaker's characteristics and semantics information into the backdoor, making it imperceptible. Finally, we estimate the channel distortion and integrate it into the backdoor. We validate our attack on 6 popular SV models. Specifically, we poison a total of 53 models and use our trigger to attack 16,430 enrolled speakers, composed of 310 target speakers enrolled in 53 poisoned models. Our attack achieves 100% attack success rate with a 15% poison rate. By decreasing the poison rate to 3%, the attack success rate remains around 50%. We validate our attack in 3 real-world scenarios, and successfully demonstrate the attack through both over-the-air and over-the-telephony-line scenarios.
Hanqing Guo, Li Xiao 0001, Qiben Yan 0001
MobiCom1
2023 Understanding Multi-Turn Toxic Behaviors in Open-Domain Chatbots
abstract
Recent advances in natural language processing and machine learning have led to the development of chatbot models, such as ChatGPT, that can engage in conversational dialogue with human users. However, understanding the ability of these models to generate toxic or harmful responses during a non-toxic multi-turn conversation remains an open research problem. Existing research focuses on single-turn sentence testing, while we find that 82% of the individual non-toxic sentences that elicit toxic behaviors in a conversation are considered safe by existing tools. In this paper, we design a new attack, ToxicChat, by fine-tuning a chatbot to engage in conversation with a target open-domain chatbot. The chatbot is fine-tuned with a collection of crafted conversation sequences. Particularly, each conversation begins with a sentence from a crafted prompt sentences dataset. Our extensive evaluation shows that open-domain chatbot models can be triggered to generate toxic responses in a multi-turn conversation. In the best scenario, ToxicChat achieves a 67% toxicity activation rate. The conversation sequences in the fine-tuning stage help trigger the toxicity in a conversation, which allows the attack to bypass two defense methods. Our findings suggest that further research is needed to address chatbot toxicity in a dynamic interactive environment. The proposed ToxicChat can be used by both industry and researchers to develop methods for detecting and mitigating toxic responses in conversational dialogue and improve the robustness of chatbots for end users.
Bocheng Chen, Guangjing Wang 0001, Hanqing Guo, Yuanda Wang, Qiben Yan 0001
RAID3
2023 PhantomSound: Black-Box, Query-Efficient Audio Adversarial Attack via Split-Second Phoneme Injection
abstract
In this paper, we propose PhantomSound, a query-efficient black-box attack toward voice assistants. Existing black-box adversarial attacks on voice assistants either apply substitution models or leverage the intermediate model output to estimate the gradients for crafting adversarial audio samples. However, these attack approaches require a significant amount of queries with a lengthy training stage. PhantomSound leverages the decision-based attack to produce effective adversarial audios, and reduces the number of queries by optimizing the gradient estimation. In the experiments, we perform our attack against 4 different speech-to-text APIs under 3 real-world scenarios to demonstrate the real-time attack impact. The results show that PhantomSound is practical and robust in attacking 5 popular commercial voice controllable devices over the air, and is able to bypass 3 liveness detection mechanisms with success rate. The benchmark result shows that PhantomSound can generate adversarial examples and launch the attack in a few minutes. We significantly enhance the query efficiency and reduce the cost of a successful untargeted and targeted adversarial attack by 93.1% and 65.5% compared with the state-of-the-art black-box attacks, using merely ∼ 300 queries (∼ 5 minutes) and ∼ 1,500 queries (∼ 25 minutes), respectively.
Hanqing Guo, Guangjing Wang 0001, Yuanda Wang, Bocheng Chen, Qiben Yan 0001, Li Xiao 0001
RAID1
2023 VSMask: Defending Against Voice Synthesis Attack via Real-Time Predictive Perturbation
abstract
Deep learning based voice synthesis technology generates artificial human-like speeches, which has been used in deepfakes or identity theft attacks. Existing defense mechanisms inject subtle adversarial perturbations into the raw speech audios to mislead the voice synthesis models. However, optimizing the adversarial perturbation not only consumes substantial computation time, but it also requires the availability of entire speech. Therefore, they are not suitable for protecting live speech streams, such as voice messages or online meetings. In this paper, we propose VSMask, a real-time protection mechanism against voice synthesis attacks. Different from offline protection schemes, VSMask leverages a predictive neural network to forecast the most effective perturbation for the upcoming streaming speech. VSMask introduces a universal perturbation tailored for arbitrary speech input to shield a real-time speech in its entirety. To minimize the audio distortion within the protected speech, we implement a weight-based perturbation constraint to reduce the perceptibility of the added perturbation. We comprehensively evaluate VSMask protection performance under different scenarios. The experimental results indicate that VSMask can effectively defend against 3 popular voice synthesis models. None of the synthetic voice could deceive the speaker verification models or human ears with VSMask protection. In a physical world experiment, we demonstrate that VSMask successfully safeguards the real-time speech by injecting the perturbation over the air.
Yuanda Wang, Hanqing Guo, Guangjing Wang 0001, Bocheng Chen, Qiben Yan 0001
WISEC2
2022 SUPERVOICE: Text-Independent Speaker Verification Using Ultrasound Energy in Human Speech
abstract
Voice-activated systems are integrated into a variety of desktop, mobile, and Internet-of-Things (IoT) devices. However, voice spoofing attacks, such as impersonation and replay attacks, in which malicious attackers synthesize the voice of a victim or simply replay it, have brought growing security concerns. Existing speaker verification techniques distinguish individual speakers via the spectrographic features extracted from an audible frequency range of voice commands. However, they often have high error rates and/or long delays. In this paper, we explore a new direction of human voice research by scrutinizing the unique characteristics of human speech at the ultrasound frequency band. Our research indicates that the high-frequency ultrasound components (e.g. speech fricatives) from 20 to 48 kHz can significantly enhance the security and accuracy of speaker verification. We propose a speaker verification system, SUPERVOICE that uses a two-stream DNN architecture with a feature fusion mechanism to generate distinctive speaker models. To test the system, we create a speech dataset with 12 hours of audio (8,950 voice samples) from 127 participants. In addition, we create a second spoofed voice dataset to evaluate its security. In order to balance between controlled recordings and real-world applications, the audio recordings are collected from two quiet rooms by 8 different recording devices, including 7 smartphones and an ultrasound microphone. Our evaluation shows that SUPERVOICE achieves 0.58% equal error rate in the speaker verification task, which reduces the best equal error rate of the existing systems by 86.1%. SUPERVOICE only takes 120 ms for testing an incoming utterance, outperforming all existing speaker verification systems. Moreover, within 91 ms processing time, SUPERVOICE achieves 0% equal error rate in detecting replay attacks launched by 5 different loudspeakers. Finally, we demonstrate that SUPERVOICE can be used in retail smartphones by integrating an off-the-shelf ultrasound microphone.
Hanqing Guo, Qiben Yan 0001, Li Xiao 0001, Eric J. Hunter
AsiaCCS1
2022 SPECPATCH: Human-In-The-Loop Adversarial Audio Spectrogram Patch Attack on Speech Recognition
abstract
In this paper, we propose SpecPatch, a human-in-the loop adversarial audio attack on automated speech recognition (ASR) systems. Existing audio adversarial attacker assumes that the users cannot notice the adversarial audios, and hence allows the successful delivery of the crafted adversarial examples or perturbations. However, in a practical attack scenario, the users of intelligent voice-controlled systems (e.g., smartwatches, smart speakers, smartphones) have constant vigilance for suspicious voice, especially when they are delivering their voice commands. Once the user is alerted by a suspicious audio, they intend to correct the falsely-recognized commands by interrupting the adversarial audios and giving more powerful voice commands to overshadow the malicious voice. This makes the existing attacks ineffective in the typical scenario when the user's interaction and the delivery of adversarial audio coincide. To truly enable the imperceptible and robust adversarial attack and handle the possible arrival of user interruption, we design SpecPatch, a practical voice attack that uses a sub-second audio patch signal to deliver an attack command and utilize periodical noises to break down the communication between the user and ASR systems. We analyze the CTC (Connectionist Temporal Classification) loss forwarding and backwarding process and exploit the weakness of CTC to achieve our attack goal. Compared with the existing attacks, we extend the attack impact length (i.e., the length of attack target command) by 287%. Furthermore, we show that our attack achieves 100% success rate in both over-the-line and over-the-air scenarios amid user intervention.
Hanqing Guo, Yuanda Wang, Li Xiao 0001, Qiben Yan 0001
CCS1
2022 NEC: Speaker Selective Cancellation via Neural Enhanced Ultrasound Shadowing
abstract
In this paper, we propose NEC (Neural Enhanced Cancellation), a defense mechanism, which prevents unautho-rized microphones from capturing a target speaker’s voice. Compared with the existing scrambling-based audio cancellation approaches, NEC can selectively remove a target speaker’s voice from a mixed speech without causing interference to others. Specifically, for a target speaker, we design a Deep Neural Network (DNN) model to extract high-level speaker-specific but utterance-independent vocal features from his/her reference audios. When the microphone is recording, the DNN generates a shadow sound to cancel the target voice in real-time. Moreover, we modulate the audible shadow sound onto an ultrasound frequency, making it inaudible for humans. By leveraging the non-linearity of the microphone circuit, the microphone can accurately decode the shadow sound for target voice cancellation. We implement and evaluate NEC comprehensively with 8 smartphone microphones in different settings. The results show that NEC effectively mutes the target speaker at a microphone without interfering with other users’ normal conversations.
Hanqing Guo, Chenning Li, Lingkun Li, Zhichao Cao 0001, Qiben Yan 0001, Li Xiao 0001
DSN1
2022 U-star: an underwater navigation system based on passive 3D optical identification tags
abstract
Underwater optical wireless communication techniques are promising due to a broad bandwidth with a long communication range compared with existing expensive acoustic and RF-based underwater communication techniques. For underwater navigation assistance during dive and rescue, it is more practical to adopt passive optical tags for objects/human identification and location-based services. However, existing optical tags (bar/QR codes) employ one/two dimensional designs, which lack significant element/symbol distance for robust decoding and full-directional localization capabilities for underwater navigation tasks. This paper investigates opportunities to increase the element distance in passive low-order optical tags by exploiting 3D spatial diversity. Specifically, we design U-Star, a system that consists of Underwater Optical Identification (UOID) tags and commercial camera-based tag readers for underwater navigation. Our UOID tags embed rich location and guidance information. Additionally, because our UOID tags employ a three-dimensional design, they can also determine the relative location of a user in real-time based on the perspective principles. We design AI based mobile algorithms for underwater denoising, relative positioning, and robust data parsing for tag readers. Finally, we evaluate U-Star on real UOID tag prototypes under different underwater scenarios. Results show that our 3-order UOID tag can embed 21 bits with a BER of 0.003 at 1m and less than 0.05 at up to 3m, which is sufficient for underwater navigation guidance with backup database.
Xiao Zhang 0037, Hanqing Guo, James Mariani, Li Xiao 0001
MobiCom2
2022 GhostTalk: Interactive Attack on Smartphone Voice System Through Power Line
Yuanda Wang, Hanqing Guo
NDSS2
2021 Rectifying Administrated ERC20 Tokens
Hanqing Guo
ICICS (1)2
2021 NELoRa: Towards Ultra-low SNR LoRa Communication with Neural-enhanced Demodulation
abstract
Low-Power Wide-Area Networks (LPWANs) are an emerging Internet-of-Things (IoT) paradigm marked by low-power and long-distance communication. Among them, LoRa is widely deployed for its unique characteristics and open-source technology. By adopting the Chirp Spread Spectrum (CSS) modulation, LoRa enables low signal-to-noise ratio (SNR) communication. However, the standard demodulation method does not fully exploit the properties of chirp signals, thus yields a sub-optimal SNR threshold under which the decoding fails. Consequently, the communication range and energy consumption have to be compromised for robust transmission. This paper presents NELoRa, a neural-enhanced LoRa demodulation method, exploiting the feature abstraction ability of deep learning to support ultra-low SNR LoRa communication. Taking the spectrogram of both amplitude and phase as input, we first design a mask-enabled Deep Neural Network (DNN) filter that extracts multi-dimension features to capture clean chirp symbols. Second, we develop a spectrogram-based DNN decoder to decode these chirp symbols accurately. Finally, we propose a generic packet demodulation system by incorporating a method that generates high-quality chirp symbols from received signals. We implement and evaluate NELoRa on both indoor and campus-scale outdoor testbeds. The results show that NELoRa achieves 1.84-2.35 dB SNR gains and extends the battery life up to 272% (~0.38-1.51 years) in average for various LoRa configurations.
Chenning Li, Hanqing Guo, Shuai Tong, Zhichao Cao 0001, Mi Zhang 0002, Qiben Yan 0001, Li Xiao 0001, Jiliang Wang, Yunhao Liu 0001
SenSys2
2021 Automated Labeling for Robotic Autonomous Navigation Through Multi-Sensory Semi-Supervised Learning on Big Data
abstract
Imitation learning holds the promise to address challenging robotic tasks such as autonomous navigation. It however requires a human supervisor to oversee the training process and send correct control commands to robots without feedback, which is always prone to error and expensive. To minimize human involvement and avoid manual labeling of data in the robotic autonomous navigation with imitation learning, this paper proposes a novel semi-supervised imitation learning solution based on a multi-sensory design. This solution includes a suboptimalsensor policybased on sensor fusion to automatically label states encountered by a robot to avoid human supervision during training. In addition, arecording policyis developed to throttle the adversarial affect of learning too much from the suboptimal sensor policy. As a result, this solution allows the robot to learn a navigation policy in a self-supervised manner without human intervention after the initial data collection. With extensive experiments in indoor environments, this solution can achieve near human performance in most of the tasks and even surpasses human performance in case of unexpected events such as hardware failures or human operation errors. To best of our knowledge, this is the first work that synthesizes sensor fusion and imitation learning to enable robotic autonomous navigation in the real world without human supervision.
Junhong Xu, Shangyue Zhu, Hanqing Guo, Shaoen Wu
IEEE Trans. Big Data3
2020 Deep Learning Driven Wireless Real-time Human Activity Recognition
abstract
Human activity recognition based on wireless sensing is advantageous at various features such as privacy preservation, but also very challenging due to the instability of wireless signals. This paper proposes a deep learning driven wireless human activity recognition solution based on Multiple-Input-Multiple-Output (MIMO) radar sensing. User activities are first sensed by a low-power Frequency-Modulated Continuous Wave (FMCW) MIMO radar. Then a sequence of 3D images are generated out of the reflected signal strength. Next, deep neural networks (DNNs) are designed to analyze the correlation among the sequential 3D images to recognize various types of human activities. This work has developed: 1) a large dataset containing over 1,500 training videos of six different types of indoor activities, 2) a customized deep learning video data-loader to select proper training data in each training epoch, 3) a deep recurrent neural network (RNN) model to recognize human activities based on radar imaging results. This solution has been extensively evaluated in a research lab room. The results show that the solution is able to generate wireless imaging frame-by-frame, and it can achieve over 86.7% accuracy in recognizing six different types of human activities.
Hanqing Guo, Nan Zhang 0021, Shaoen Wu
ICC1
2020 SurfingAttack: Interactive Hidden Attack on Voice Assistants Using Ultrasonic Guided Waves
Qiben Yan 0001, Kehai Liu, Hanqing Guo, Ning Zhang 0017
NDSS4
2019 DSIC: Deep Learning Based Self-Interference Cancellation for In-Band Full Duplex Wireless
abstract
In-band full duplex (IBFD) wireless is of utmost interest to future wireless communication and networking due to great potentials of spectrum efficiency. IBFD wireless, how- ever, is throttled by its key challenge, namely self-interference. Therefore, effective self- interference cancellation is the key to enable IBFD wireless. This paper proposes a real-time non- linear self-interference cancellation solution: Deep learning based Self-Interference Cancellation (DSIC) to enable IBFD wireless. In this solution, a self-interference channel is modeled by a deep neural network (DNN). Synchronized self- interference channel data is first collected to train the DNN of the self-interference channel. Afterwards, the trained DNN is used to cancel the self-interference at a wireless node. This solution has been implemented on a USRP SDR testbed and evaluated in real world in multiple scenarios with various modulations in transmitting information including numbers, texts as well as images. It results in the performance of 17dB in digital cancellation, which is very close to the self-interference power and nearly cancels the self- interference at a SDR node in the testbed. The solution yields an average of 8.5% bit error rate (BER) over many scenarios and different modulation schemes.
Hanqing Guo, Shaoen Wu, Honggang Wang 0001, Mahmoud Daneshmand
GLOBECOM1
2019 In-band full duplex wireless communications and networking for IoT devices: Progress, challenges and opportunities
Shaoen Wu, Hanqing Guo, Junhong Xu, Shangyue Zhu, Honggang Wang 0001
Future Gener. Comput. Syst.2
2018 Non-Contact Non-Invasive Heart and Respiration Rates Monitoring with MIMO Radar Sensing
abstract
Smart health calls for novel approaches to detect vital signs in non- contact, non-invasive and non-intrusive matters. In this work, we design a solution that monitors the rates of heartbeats and respiration simultaneously by using a Frequency Modulated Continuous Wave (FMCW) radar with multiple antennas. This solution measures the reflections from heartbeats and respiration at a high frequency of 4 K H z to capture fine dynamics of motions with big data. It employs multiple antennas and superposition to reduce the interference noises from unwanted motions in the background and any detection defects. The heart and respiration rates are detected in the frequency domains after a chain of preprocessing techniques on the sensed big data. With extensive experiments in a lab office, this system demonstrates high accuracies in various cases: 98% in the still case, 95% with finger motions and 96% with body motions. The tests also confirm that multiple antennas and signal superposition improve the detection accuracy and reliability.
Hanqing Guo, Junhong Xu, Honggang Wang 0001, Aaron Kageza, Saeed AlQarni, Shaoen Wu
GLOBECOM2
2018 Shared Multi-Task Imitation Learning for Indoor Self-Navigation
abstract
Deep imitation learning enables robots to learn from expert demonstrations to perform tasks such as lane following or obstacle avoidance. However, in the traditional imitation learning framework, one model only learns one task, and thus it lacks of the capability to support a robot to perform various different navigation tasks with one model in indoor environments. This paper proposes a new framework, Shared Multi-headed Imitation Learning (SMIL), that allows a robot to perform multiple tasks with one model without switching among different models. We model each task as a sub-policy and design a multi-headed policy to learn the shared information among related tasks by summing up activations from all sub-policies. Compared to single or non-shared multi-headed policies, this framework is able to leverage correlated information among tasks to increase performance. We have implemented this framework using a robot based on NVIDIA TX2 and performed extensive experiments in indoor environments with different baseline solutions. The results demonstrate that SMIL has doubled the performance over non-shared multi-headed policy.
Junhong Xu, Hanqing Guo, Aaron Kageza, Saeed AlQarni, Shaoen Wu
GLOBECOM3
2018 Indoor Human Activity Recognition Based on Ambient Radar with Signal Processing and Machine Learning
abstract
Indoor human activity recognition has been extensively investigated. However, most of the solutions require sensors e.g. 9-axis IMU be equipped on human body or use image processing that presents privacy issues. This work proposes an ambient radar sensor based a solution to recognize the activities that humans normally perform in indoor environments. This solution uses a 7.8 GHz radar to emit 16 pulse signals every second and samples the reflected signals at 128 KHz to capture the fine dynamics of human activities. This solution designs a set of data preprocessing algorithms, including a data refining algorithm to filter outlier data, a contrastive divergence algorithm to remove background static reflection, and a transformation algorithm to convert the signal data into feature- rich spatial location changes. This solution also develops schemes to separate a collection of various activities into individuals. A lowpass frequency filter is designed to remove unwanted noisy data and the motion intensity is used to classify the activities into two high-level groups. It uses a slope-based approach and a k- means clustering to further finely recognize each activity. This solution has been extensively evaluated in a spacious research lab room and shows outstanding accuracy.
Shangyue Zhu, Junhong Xu, Hanqing Guo, Shaoen Wu, Honggang Wang 0001
ICC3
2018 A Deep Residual convolutional neural network for facial keypoint detection with missing labels
Shaoen Wu, Junhong Xu, Shangyue Zhu, Hanqing Guo
Signal Process.4