Hanbo Cai

dblp:285/4032 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-3701-6383ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Modulation-Based Backdoors: Leveraging Amplitude and Frequency Patterns to Attack Speaker Recognition
abstract
Deep neural networks (DNNs) are widely and successfully applied in the field of speaker recognition. However, recent studies reveal that these models are vulnerable to backdoor attacks, where adversaries inject malicious behaviors into victim models by poisoning the training process. Existing attack methods often rely on environmental noise or complex voice transformations, which are typically difficult to implement and exhibit poor stealthiness. To address these issues, this paper proposes two modulation-based backdoor attacks that leverage frequency modulation (FM) and amplitude modulation (AM) to construct audio triggers. In real-world scenarios, regular variations in frequency and amplitude are often imperceptible to human listeners, making the proposed attacks more covert. Experimental results show that our methods achieve high attack success rates in both digital and physical settings, while also demonstrating strong resistance to various state-of-the-art backdoor defenses.
Hanbo Cai, Pengcheng Zhang 0001, Yan Xiao 0002, Hanting Chu
AAAI1
2026 Automated robustness testing for LLM-based natural language processing software
Mingxuan Xiao, Yan Xiao 0002, Shunhui Ji, Hanbo Cai, Lei Xue 0001, Pengcheng Zhang 0001
Expert Syst. Appl.4
2025 Clean-label backdoor attack based on robust feature attenuation for speech recognition
Hanbo Cai, Pengcheng Zhang 0001, Yan Xiao 0002, Shunhui Ji, Mingxuan Xiao, Letian Cheng
Expert Syst. Appl.1
2024 Audio Steganography Based Backdoor Attack for Speech Recognition Software
abstract
With the growing prevalence of deep learning in the speech area, speech recognition, voice control, and related applications have become integral parts of people's lives. However, the rise of malicious third-party platforms has introduced significant security concerns, particularly through backdoor attacks. These attacks implant triggers that manipulate speech recognition models to produce specific labels, thereby compromising the system's integrity. Studying speech backdoor attacks is crucial for evaluating the security of speech recognition software, and iden-tifying and addressing potential vulnerabilities. Existing methods for speech backdoor attacks usually employ fixed perturbations as triggers. However, these perturbations may be discernible to the human ear, making them easily detectable. To address this issue, we propose a frequency domain-embedded backdoor attack method based on echo hiding. Echo hiding is a steganography technique based on audio. This method embeds hidden information into the frequency spectrum of the echo signal, leveraging the masking property of the human auditory system. It is difficult to arouse suspicion or detect the presence of hidden information since echo is perceived as a natural phenomenon in auditory perception. Furthermore, it does not cause a significant decrease in audio quality. Experimental results show the effectiveness of our method in different settings.
Shunhui Ji, Hanbo Cai, Hai Dong 0001, Pengcheng Zhang 0001
COMPSAC3
2024 Toward Stealthy Backdoor Attacks Against Speech Recognition via Elements of Sound
abstract
Deep neural networks (DNNs) have been widely and successfully adopted and deployed in various applications of speech recognition. Recently, a few works revealed that these models are vulnerable to backdoor attacks, where the adversaries can implant malicious prediction behaviors into victim models by poisoning their training process. In this paper, we revisit poison-only backdoor attacks against speech recognition. We reveal that existing methods are not stealthy since their trigger patterns are perceptible to humans or machine detection. This limitation is mostly because their trigger patterns are simple noises or separable and distinctive clips. Motivated by these findings, we propose to exploit elements of sound (e.g., pitch and timbre) to design more stealthy yet effective poison-only backdoor attacks. Specifically, we insert a short-duration high-pitched signal as the trigger and increase the pitch of remaining audio clips to ‘mask’ it for designing stealthy pitch-based triggers. We manipulate timbre features of victim audio to design the stealthy timbre-based attack and design a voiceprint selection module to facilitate the multi-backdoor attack. Our attacks can generate more ‘natural’ poisoned samples and therefore are more stealthy. Extensive experiments are conducted on benchmark datasets, which verify the effectiveness of our attacks under different settings (e.g., all-to-one, all-to-all, clean-label, physical, and multi-backdoor settings) and their stealthiness. Our methods achieve attack success rates of over 95% in most cases and are nearly undetectable. The code for reproducing main experiments are available at https://github.com/HanboCai/BadSpeech_SoE.
Hanbo Cai, Pengcheng Zhang 0001, Hai Dong 0001, Yan Xiao 0002, Stefanos Koffas, Yiming Li 0004
IEEE Trans. Inf. Forensics Secur.1
2023 Adversarial example-based test case generation for black-box speech recognition systems
abstract
Abstract Test case generation techniques based on adversarial examples are commonly used to enhance the reliability and robustness of image‐based and text‐based machine learning applications. However, efficient techniques for speech recognition systems are still absent. This paper proposes a family of methods that generate targeted adversarial examples for speech recognition systems. All are based on thefirefly algorithm (F), and are enhanced withgaussmutations and / orgradientestimation (F‐GM, F‐GE, F‐GMGE) to fit the specific problem of targeted adversarial test case generation. We conduct an experimental evaluation on three different types of speech datasets, includingGoogle Command,Common VoiceandLibriSpeech. In addition, we recruit volunteers to evaluate the performance of the adversarial examples. The experimental results show that, compared with existing approaches, these approaches can effectively improve the success rate of the targeted adversarial example generation. The code is publicly available at https://github.com/HanboCai/FGMGE .
Hanbo Cai, Pengcheng Zhang 0001, Hai Dong 0001, Lars Grunske, Shunhui Ji, Tianhao Yuan
Softw. Test. Verification Reliab.1
2021 Differential Privacy Preservation in Adaptive K-Nets Clustering
abstract
K-Nets is a deterministic clustering algorithm based on the network structure. It can automatically detect the sym-metric structure in the data and can be used to process clusters of different sizes, shapes or a specific number. However, K-Nets has the following shortcomings: (1) the clustering result is more sensitive to the manually input parameter K, so the accuracy will be affected; (2) the algorithm only considers the average distance of K-nearest neighbors, which may lead to some wrong distribution center points in the dataset with large density difference or the same score values during calculation; (3) it does not consider the privacy leakage during the clustering process. To solve the above problems, we propose a differential privacy protection method in adaptive K-Nets clustering, called ADP-K-Nets. Firstly, for reducing the influence of the parameters on the result, the natural eigenvalues are adaptively obtained through the characteristic of the natural neighbors and used as parameter values to find data points. Then we define a new method for calculating the score, which can solve the problem of incorrectly selecting cluster centers when there are large density differences or conflicts in the calculation process. Also, the Laplace noise is added in calculating the local density of every data point to protect data privacy. Experimental results show that our method ensures the performance of clustering compared with some existing algorithms.
Hanbo Cai, Xianxian Li
TrustCom2