Zheng Fang 0014

dblp:77/4730-14 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0005-9308-7452ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Selective Masking Adversarial Attack on Automatic Speech Recognition Systems
abstract
Extensive research has shown that Automatic Speech Recognition (ASR) systems are vulnerable to audio adversarial attacks. Current attacks mainly focus on single-source scenarios, ignoring dual-source scenarios where two people are speaking simultaneously. To bridge the gap, we propose a Selective Masking Adversarial attack, namely SMA attack, which ensures that one audio source is selected for recognition while the other audio source is muted in dual-source scenarios. To better adapt to the dual-source scenario, our SMA attack constructs the normal dual-source audio from the muted audio and selected audio. SMA attack initializes the adversarial perturbation with a small Gaus-sian noise and iteratively optimizes it using a selective masking optimization algorithm. Extensive experiments demonstrate that the SMA attack can generate effective and imperceptible audio adversarial examples in the dual-source scenario, achieving an average success rate of attack of 100% and signal-to-noise ratio of 37.15dB on Conformer-CTC, outperforming the baselines.
Zheng Fang 0014, Shenyi Zhang, Tao Wang 0081, Bowen Li 0016, Lingchen Zhao, Zhangyi Wang
ICME1
2025 JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
Shenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu, Zheng Fang 0014, Lingchen Zhao, Chao Shen 0001, Cong Wang 0001, Qian Wang 0002
USENIX Security Symposium6
2025 CuckooAttack: Towards Practical Backdoor Attack against Automatic Speech Recognition Systems
abstract
Deep learning-based automatic speech recognition (ASR) systems are capable of transcribing input audio of arbitrary duration into character sequences, which are widely used in daily life. However, recent research has found that deep learning models are vulnerable to backdoor attacks. A malicious adversary can embed a backdoor functionality into the model during the training phase and manipulate the output of the backdoored model by adding a specific trigger to the input during the inference phase. Unfortunately, TrojanModel, the existing state-of-the-art backdoor attack against ASR systems (Zong et al. S&P’23), relies on an overly strong assumption that requires the adversary to modify the model structure beyond data poisoning, which significantly limits its practicability. In this paper, we propose CuckooAttack, a more practical backdoor attack against ASR systems that only requires poisoning a small portion of the training data. We first construct a phoneme-level auxiliary dataset to generate effective, robust, and unnoticeable triggers, while substantially lowering computational expenses. Considering the real-world ASR application scenarios, we propose an adaptive trigger injection mechanism to ensure that the backdoor can be activated on variable-duration input audio under asynchronous temporal conditions. To further enhance the efficacy of CuckooAttack, we design a character-filling strategy tailored for ASR to construct poisoned samples, which facilitates the model in establishing backdoor connections. Extensive experiments show that CuckooAttack achieves comparable performance with TrojanModel under a weaker assumption. Specifically, CuckooAttack achieves an attack success rate of about 99% in the digital domain and over 90% in the physical domain, with a poison rate of only 1%.
Bowen Li 0016, Yunjie Ge, Zheng Fang 0014, Tao Wang 0081, Lingchen Zhao, Ning Jiang 0001, Qian Wang 0002
IEEE Trans. Dependable Secur. Comput.3
2024 Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition Systems
abstract
In recent years, extensive research has been conducted on the vulnerability of ASR systems, revealing that black-box adversarial example attacks pose significant threats to real-world ASR systems. However, most existing black-box attacks rely on queries to the target ASRs, which is impractical when queries are not permitted. In this paper, we propose ZQ-Attack, a transfer-based adversarial attack on ASR systems in the zero-query black-box setting. Through a comprehensive review and categorization of modern ASR technologies, we first meticulously select surrogate ASRs of diverse types to generate adversarial examples. Following this, ZQ-Attack initializes the adversarial perturbation with a scaled target command audio, rendering it relatively imperceptible while maintaining effectiveness. Subsequently, to achieve high transferability of adversarial perturbations, we propose a sequential ensemble optimization algorithm, which iteratively optimizes the adversarial perturbation on each surrogate model, leveraging collaborative information from other models. We conduct extensive experiments to evaluate ZQ-Attack. In the over-the-line setting, ZQ-Attack achieves a 100% success rate of attack (SRoA) with an average signal-to-noise ratio (SNR) of 21.91dB on 4 online speech recognition services, and attains an average SRoA of 100% and SNR of 19.67dB on 16 open-source ASRs. In the over-the-air setting, ZQ-Attack also achieves a 100% SRoA with an average SNR of 15.77dB on 2 commercial intelligent voice control devices.
Zheng Fang 0014, Tao Wang 0081, Lingchen Zhao, Shenyi Zhang, Bowen Li 0016, Yunjie Ge, Qi Li 0002, Chao Shen 0001, Qian Wang 0002
CCS1
2024 Hijacking Attacks against Neural Network by Analyzing Training Data
Yunjie Ge, Qian Wang 0002, Huayang Huang, Qi Li 0002, Cong Wang 0001, Chao Shen 0001, Lingchen Zhao, Peipei Jiang 0002, Zheng Fang 0014, Shenyi Zhang
USENIX Security Symposium9
2024 Palette: Physically-Realizable Backdoor Attacks Against Video Recognition Models
abstract
Backdoor attacks have been widely studied for image classification tasks, but rarely investigated for video recognition tasks. In this paper, we explore the possibility of physically-realizable backdoor attacks against video recognition models. Different from existing works that directly apply image backdoor attacks to videos, i.e., patch a visible trigger to each frame of a video, we carefully take into consideration the temporal interactions among frames in a video. Our proposed video backdoor attack, namedPalette, features two special design choices. The first is to utilize natural-light-alike RGB offset as triggers rather than traditional patch triggers. Such triggers may be applied in the physical world through lighting without the need to modify video files. The second is to make the backdoored model more robust to temporal asynchronization between the trigger and the video samples by performing rolling operations during sample poisoning. Extensive experiments show thatPaletteoutperforms existing video backdoor attacks, especially in the physical world. It is shown thatPaletteis also resistant to backdoor defense methods. We will open-source our codes upon publication.
Xueluan Gong, Zheng Fang 0014, Bowen Li 0016, Tao Wang 0081, Yanjiao Chen, Qian Wang 0002
IEEE Trans. Dependable Secur. Comput.2