EDBT 2026 Demo / reviewers in the wild / expert
Xuejing Yuan
dblp:213/8010
· DBLP profile ↗
10ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0003-1866-5828ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedWM: Data-Free Watermarking for Model Ownership Protection in Federated LearningabstractThe widespread adoption of federated learning has been driven by growing demands for privacy protection in model training. Federated learning enables multiple clients to collaboratively train a global model coordinated by a central server without sharing their raw data. However, when distributing the global model to clients, the central server faces significant security risks from malicious clients who may steal and misuse the model, thereby compromising its ownership. While existing watermarking techniques typically rely on main task data for ownership protection, their application in federated learning is limited since the server lacks access to this data, which remains with the clients. To address this challenge, we propose a novel data-free watermarking method. We utilize substitute data unrelated to the main task and improve efficiency by filtering out redundant samples. To optimize the watermarking process, we introduce a logits alignment-based optimization strategy that uses the substitute dataset with watermark triggers for effective embedding. Additionally, we propose a dynamic optimization algorithm to balance the trade-off between watermark embedding and main task. We comprehensively evaluate our approach across four datasets, four model architectures, and three mainstream deep learning tasks. Our experimental results demonstrate nearly perfect watermark performance while maintaining minimal impact on the main task. Notably, our watermarking method proves resistant to existing backdoor detection techniques, establishing its effectiveness, robustness and stealthiness. Congyi Li, Peizhuo Lv, Xuejing Yuan, Shengzhi Zhang, Kai Chen 0012, Yingjiu Li |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | EvilHarmony: Stealthy Adversarial Attacks Against Black-Box Speech Recognition SystemsabstractAutomatic Speech Recognition (ASR) systems are vulnerable to adversarial examples (AEs), where small, carefully designed perturbations are added to original audio to mislead the systems into generating target commands. Existing adversarial attacks typically initialize perturbations either as zero or as Text-to-Speech clips of the target command. The former accumulates the features of the command in the perturbed audio, while the latter constantly reduces the features of the command, resulting in the generation of AEs. Although most target commands in the AEs are imperceptible to humans, the audio often exhibits noticeable distortions or disruptions, making it apparent that the sound has been tampered with. This work aims to retain only the essential features of adversarial audio, minimizing distortions from unnecessary elements to improve quality and make the attack less detectable. Our findings highlight the importance of formants as critical features for black-box adversarial attacks, motivating the development of a novel Formant Filter Bank (FFB) tailored to the target command. By inputting musical audio into the FFB, we utilize the filtered output as the perturbation seed, which retains the formant features of the target command and blends in certain features of the original music. Then we search for a minimum enhancement factor for the perturbation seed to generate high-quality AEs. Our perturbation can be regarded as local amplitude modulation of the music, so we define the AE as EvilHarmony. Experimental results demonstrate that our method successfully attacks commercial black-box ASR models, including Microsoft, Google, Amazon, Tencentyun, Aliyun, and OpenAI Whisper-V3. Compared to existing approaches, our AEs achieve significantly greater stealth, with 53% to 77% of participants perceiving them as indistinguishable from normal audio across the six ASR API services. Additionally, our approach successfully attacks Google Assistant and voice assistants on Surface Pro 9 in the real world. Demos are uploaded at https://sites.google.com/view/evilharmony. Xuejing Yuan, Jiangshan Zhang, Kai Chen 0012, XiaoFeng Wang 0001, Shengzhi Zhang, Dun Liu, Runnan Zhu |
SP | 1 |
| 2025 | Adversarial Attack and Defense for Commercial Black-box Chinese-English Speech Recognition SystemsabstractThe attacker can generate adversarial examples (AEs) to stealthily mislead automatic speech recognition (ASR) models, raising significant concerns about the security of intelligent voice control (IVC) devices. Existing adversarial attacks mainly generate AEs to mislead ASR models to output specific target English commands (e.g., open the door). However, it remains unknown whether AEs can be used to issue commands in other languages to attack commercial black-box ASR models. In this article, taking Chinese phrases (e.g., 支付宝付款) and “Chinese–English code-switching” phrases (e.g., 关闭GPS) as the target commands, we propose adversarial attacks for commercial multilingual ASR models. In particular, if a multilingual speech recognition model can recognize Chinese and English, we call it a Chinese–English speech recognition model. In English, the meaning of “支付宝付款” and “关闭GPS” are “Alipay payment” and “turn off GPS”, respectively. In detail, we generate transferable AEs based on the open-sourced conventional DataTang Mandarin ASR model. Given 55 target commands, the success rate for generating AEs of them is up to 96% and 80% for Aliyun ASR API and Tencentyun ASR API, respectively. Our AEs can trigger actual attack actions on voice assistants (e.g., Apple Siri, Xiaomi Xiaoaitongxue) or spread malicious messages through ASR API services, while the target commands in the AEs are inaudible to human beings. 1 Finally, by analyzing the spectrum differences between benign audio clips and AEs, we propose a general defense against adversarial audio attacks. Xuejing Yuan, Jiangshan Zhang, Kai Chen 0012, Cheng'an Wei, Zhenkun Ma, Xinqi Ling |
ACM Trans. Priv. Secur. | 1 |
| 2024 | EvilPromptFuzzer: generating inappropriate content based on text-to-image modelsabstractAbstract Text-to-image (TTI) models provide huge innovation ability for many industries, while the content security triggered by them has also attracted wide attention. Considerable research has focused on content security threats of large language models (LLMs), yet comprehensive studies on the content security of TTI models are notably scarce. This paper introduces a systematic tool, named EvilPromptFuzzer, designed to fuzz evil prompts in TTI models. For 15 kinds of fine-grained risks, EvilPromptFuzzer employs the strong knowledge-mining ability of LLMs to construct seed banks, in which the seeds cover various types of characters, interrelations, actions, objects, expressions, body parts, locations, surroundings, etc. Subsequently, these seeds are fed into the LLMs to build scene-diverse prompts, which can weaken the semantic sensitivity related to the fine-grained risks. Hence, the prompts can bypass the content audit mechanism of the TTI model, and ultimately help to generate images with inappropriate content. For the risks of violence, horrible, disgusting, animal cruelty, religious bias, political symbol, and extremism, the efficiency of EvilPromptFuzzer for generating inappropriate images based on DALL.E 3 are greater than 30%, namely, more than 30 generated images are malicious among 100 prompts. Specifically, the efficiency of horrible, disgusting, political symbols, and extremism up to 58%, 64%, 71%, and 50%, respectively. Additionally, we analyzed the vulnerability of existing popular content audit platforms, including Amazon, Google, Azure, and Baidu. Even the most effective Google SafeSearch cloud platform identifies only 33.85% of malicious images across three distinct categories. Juntao He, Runqi Sui, Xuejing Yuan, Dun Liu, Wenchuan Yang, Baojiang Cui, Kedan Li |
Cybersecur. | 4 |
| 2023 | Enhancing Privacy Preservation in Federated Learning via Learning Rate PerturbationabstractFederated learning (FL) is a privacy-enhanced distributed machine learning framework, in which multiple clients collaboratively train a global model by exchanging their model updates without sharing local private data. However, the adversary can use gradient inversion attacks to reveal the clients’ privacy from the shared model updates. Previous attacks assume the adversary can infer the local learning rate of each client, while we observe that: (1) using the uniformly distributed random local learning rates does not incur much accuracy loss of the global model, and (2) personalizing local learning rates can mitigate the drift issue which is caused by non-IID (identically and in-dependently distributed) data. Moreover, we theoretically derive a convergence guarantee to FedAvg with uniformly perturbed local learning rates. Therefore, by perturbing the learning rate of each client with random noise, we propose a learning rate perturbation (LRP) defense against gradient inversion attacks. Specifically, for classification tasks, we adapt LPR to ada-LPR by personalizing the expectation of each local learning rate. The experiments show that our defenses can well enhance privacy preservation against existing gradient inversion attacks, and LRP outperforms 5 baseline defenses against a state-of-the-art gradient inversion attack. In addition, our defenses only incur minor ac-curacy reductions (less than 0.5%) of the global model. So they are effective in real applications. Guangnian Wan, Haitao Du, Xuejing Yuan, Meiling Chen, Jie Xu 0038 |
ICCV | 3 |
| 2022 | Dimensionality reduction algorithm of tensor data based on orthogonal tucker decomposition and local discrimination difference
Wenxu Gao, Zhengming Ma, Xuejing Yuan |
Appl. Intell. | 3 |
| 2022 | SoK: A Modularized Approach to Study the Security of Automatic Speech Recognition SystemsabstractWith the wide use of Automatic Speech Recognition (ASR) in applications such as human machine interaction, simultaneous interpretation, audio transcription, and so on, its security protection becomes increasingly important. Although recent studies have brought to light the weaknesses of popular ASR systems that enable out-of-band signal attack, adversarial attack, and so on, and further proposed various remedies (signal smoothing, adversarial training, etc.), a systematic understanding of ASR security (both attacks and defenses) is still missing, especially on how realistic such threats are and how general existing protection could be. In this article, we present our systematization of knowledge for ASR security and provide a comprehensive taxonomy for existing work based on a modularized workflow. More importantly, we align the research in this domain with that on security in Image Recognition System (IRS), which has been extensively studied, using the domain knowledge in the latter to help understand where we stand in the former. Generally, both IRS and ASR are perceptual systems. Their similarities allow us to systematically study existing literature in ASR security based on the spectrum of attacks and defense solutions proposed for IRS, and pinpoint the directions of more advanced attacks and the directions potentially leading to more effective protection in ASR. In contrast, their differences, especially the complexity of ASR compared with IRS, help us learn unique challenges and opportunities in ASR security. Particularly, our experimental study shows that transfer attacks across ASR models are feasible, even in the absence of knowledge about models (even their types) and training data. Jiangshan Zhang, Xuejing Yuan, Shengzhi Zhang, Kai Chen 0012, XiaoFeng Wang 0001, Shanqing Guo |
ACM Trans. Priv. Secur. | 3 |
| 2020 | Devil's Whisper: A General Approach for Physical Adversarial Attacks against Commercial Black-box Speech Recognition Devices
Xuejing Yuan, Jiangshan Zhang, Yue Zhao 0018, Shengzhi Zhang, Kai Chen 0012, XiaoFeng Wang 0001 |
USENIX Security Symposium | 2 |
| 2018 | All Your Alexa Are Belong to Us: A Remote Voice Control Attack against EchoabstractVoice controlled system becomes increasingly popular these days due to the convenient and natural control over lots of functionalities and smart devices. Amazon Echo, designed around Alexa, is capable of controlling smart devices such as locks, sending emails, making phone calls, and even bridging the gap between online services such as Twitter, Facebook, etc. Previously, researchers demonstrated that by carefully crafting obfuscated commands or transmitting commands over ultrasound carrier, voice controlled systems can be compromised without people's awareness. However, those researches require the target voice controlled systems to be close enough to their speaker or ultrasound transducer. In this paper, we proposed REEVE (REmotE VoicE control) attack that can manipulate Amazon Alexa remotely, e.g., via signal broadcasting to compromise radio, TV, speaker, etc. It works on behalf of the attackers to operate various commands beneficial to them. By analyzing more than 15,000 Alexa skills and 600 IFTTT Applets related to Alexa, we found that more than 100 of them can be used to attack Echo. We also thoroughly scrutinized the attack surface of Echo's voice control and conducted security analysis based on different consequences. Xuejing Yuan, Aohui Wang, Kai Chen 0012, Shengzhi Zhang, Heqing Huang 0001, Ian M. Molloy |
GLOBECOM | 1 |
| 2018 | CommanderSong: A Systematic Approach for Practical Adversarial Voice Recognition
Xuejing Yuan, Yue Zhao 0018, Yunhui Long, Kai Chen 0012, Shengzhi Zhang, Heqing Huang 0001, XiaoFeng Wang 0001, Carl A. Gunter |
USENIX Security Symposium | 1 |