Yue Zhao 0018

dblp:48/76-18 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
10since 2021 · last 2025
0009-0007-4708-8061ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 11 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 SCOPE: Expanding Client-Side Post-Processing for Efficient Privacy-Preserving Model Inference
abstract
Privacy-Preserving Inference (PPI) enables users to leverage powerful machine learning models without revealing sensitive input data. However, existing state-of-the-art solutions remain impractical due to significant computation and communication overheads.
Shenchen Zhu, Kai Chen 0012, Yue Zhao 0018, Cheng'an Wei
CCS3
2025 DEO: Jailbreak a Black-box Multimodal Large Language Model with Dual-Embedding Alignment
abstract
Multimodal Large Language Models (MLLMs), which integrate textual and visual modalities, have demonstrated unparalleled capabilities in diverse multimodal tasks. However, the inclusion of visual inputs exposes MLLMs to security risks, one of which is jailbreak attacks. Although various methods have been proposed to jailbreak MLLMs via the visual modality, attacks in black-box settings have some limitations. Existing black-box attacks either fail to generate precise harmful outputs in practical scenarios or require substantial preparatory work in constructing adversarial images. In this work, we propose a novel dual-embedding optimization (DEO) attack approach to generate visual adversarial perturbations that induce the MLLMs to produce harmful responses that violate common AI safety policies. Specifically, DEO iteratively optimizes the visual input by enforcing alignment objectives across both the input and output embedding spaces: the image embedding of the input and the text embedding generated by the MLLM are both required to align with a harmful target text within a shared embedding space, which is defined by a frozen pretrained encoder. This alignment is conducted entirely under a black-box setting using a query-based strategy, where the attacker issues queries and observes only the model’s outputs, without access to its internal parameters or gradients. By optimizing in the dual-embedding space, our method can generate an adversarial perturbation to elicit more harmful and precise responses, overcoming the limitations of existing approaches. Experimental results demonstrate that our method significantly improves attack success rates of existing black-box attack methods by up to 30% against two MLLM families, including MiniGPT4 and LLaVa, achieving an average attack success rate of 87% across different models and eight scenarios, demonstrating its superior attack effectiveness. These findings highlight the urgent need for systematic robustness evaluations and improved safety mechanisms in MLLMs.1Content Warning: This paper contains harmful model responses.
Mingsi Wang, Yue Zhao 0018, Zijin Lin, Kai Chen 0012
IJCNN3
2025 PrivacyXray: Detecting Privacy Breaches in LLMs through Semantic Consistency and Probability Certainty
Jinwen He, Zijin Lin, Kai Chen 0012, Yue Zhao 0018
USENIX Security Symposium5
2025 EGRTE: adversarially training a self-explaining smoothed classifier for certified robustness
abstract
Abstract Deep learning has transformed fields such as computer vision, natural language processing, and audio analysis through its powerful pattern recognition and predictive capabilities. However, the robustness of these models remains a major concern, as they are highly vulnerable to adversarial attacks-subtle, intentional perturbations that lead to incorrect predictions. While recent defenses like adversarial training and defensive distillation aim to improve robustness, they have notable drawbacks, including overfitting and degraded performance under strong attacks. Certified defenses, such as robust training and Randomized Smoothing, offer theoretical guarantees within a specific perturbation radius, yet struggle to reflect real-world robustness due to efficiency bottlenecks and the unpredictable nature of actual adversarial attacks. These challenges reveal a critical gap between current defenses and real-world attack scenarios, highlighting the need for more practical and resilient solutions. To address the challenges of defense-attack gaps and the inefficiency in robust training, we introduce the Explanation-Guided Robust Training Enhancer (EGRTE). EGRTE combines a self-explaining mechanism, which guides adversarial training to focus on generalized features for improved robustness and accuracy, with a masking mechanism that transforms noised data for easier model learning. This approach not only mitigates noise effects, including adversarial perturbations, but also eliminates the need for time-intensive gradient calculations, greatly enhancing training efficiency. Comprehensive experiments on several datasets show EGRTE’s superior certified accuracy and robustness against adversarial attacks, with a 6.24-fold efficiency increase over comparable methods, positioning EGRTE as a highly effective solution for robust and efficient deep learning.
Zijin Lin, Jinwen He, Yue Zhao 0018, Ruigang Liang, Zhendong Wu
Cybersecur.3
2024 UMA: Facilitating Backdoor Scanning via Unlearning-Based Model Ablation
abstract
Recent advances in backdoor attacks, like leveraging complex triggers or stealthy implanting techniques, have introduced new challenges in backdoor scanning, limiting the usability of Deep Neural Networks (DNNs) in various scenarios. In this paper, we propose Unlearning-based Model Ablation (UMA), a novel approach to facilitate backdoor scanning and defend against advanced backdoor attacks. UMA filters out backdoor-irrelevant features by ablating the inherent features of the target class within the model and subsequently reveals the backdoor through dynamic trigger optimization. We evaluate our method on 1700 models (700 benign and 1000 trojaned) with 6 model structures, 7 different backdoor attacks and 4 datasets. Our results demonstrate that the proposed methodology effectively detect these advanced backdoors. Specifically, our method can achieve 91% AUC-ROC and 86.6% detection accuracy on average, which outperforms the baselines, including Neural Cleanse, ABS, K-Arm and MNTD.
Yue Zhao 0018, Congyi Li, Kai Chen 0012
AAAI1
2024 I Don't Know You, But I Can Catch You: Real-Time Defense against Diverse Adversarial Patches for Object Detectors
abstract
Deep neural networks (DNNs) have revolutionized the field of computer vision like object detection with their unparalleled performance. However, existing research has shown that DNNs are vulnerable to adversarial attacks. In the physical world, an adversary could exploit adversarial patches to implement a Hiding Attack (HA) which patches the target object to make it disappear from the detector, and an Appearing Attack (AA) which fools the detector into misclassifying the patch as a specific object. Recently, many defense methods for detectors have been proposed to mitigate the potential threats of adversarial patches. However, such methods still have limitations in generalization, robustness and efficiency. Most defenses are only effective against the HA, leaving the detector vulnerable to the AA.
Zijin Lin, Yue Zhao 0018, Kai Chen 0012, Jinwen He
CCS2
2024 AE-Morpher: Improve Physical Robustness of Adversarial Objects against LiDAR-based Detectors via Object Reconstruction
Shenchen Zhu, Yue Zhao 0018, Kai Chen 0012, Hualong Ma, Cheng'an Wei
USENIX Security Symposium2
2024 NeuralSanitizer: Detecting Backdoors in Neural Networks
abstract
Deep neural networks (DNNs) have been pervasively used in many areas, e.g., computer vision, speech recognition, natural language processing, etc. However, recent works show that they are vulnerable to backdoor/Trojan attacks, severely restricting their usage in various scenarios. In this paper, we proposeNeuralSanitizer, a novel approach to detect and remove backdoors in DNNs, capable of capturing various triggers with better accuracy and higher efficiency. In particular, we identify two fundamental properties of triggers, i.e., their effectiveness in the backdoored model and ineffectiveness in other clean models, and design a novel objective function to reconstruct triggers based on them. Then we present a new approach that leverages transferability to identify adversarial patches that could be generated during trigger reconstruction, thus detecting backdoors more accurately. We evaluate NeuralSanitizer on real-world backdoored DNNs and achieve 2.1% FNR and 0.9% FPR on average, significantly outperforming the state-of-the-art works by 1~14 times. In addition, NeuralSanitizer can reconstruct triggers up to 25% of the size of the original inputs on average, compared to only 6~10% by existing works. Finally, NeuralSanitizer is also 1~25 times faster than existing works.
Yue Zhao 0018, Shengzhi Zhang, Kai Chen 0012
IEEE Trans. Inf. Forensics Secur.2
2023 A Robustness-Assured White-Box Watermark in Neural Networks
abstract
Recently, stealing highly-valuable and large-scale deep neural network (DNN) models becomes pervasive. The stolen models may be re-commercialized, e.g., deployed in embedded devices, released in model markets, utilized in competitions, etc, which infringes the Intellectual Property (IP) of the original owner. Detecting IP infringement of the stolen models is quite challenging, even with the white-box access to them in the above scenarios, since they may have experienced fine-tuning, pruning, functionality-equivalent adjustment to destruct any embedded watermark. Furthermore, the adversaries may also attempt to extract the embedded watermark or forge a similar watermark to falsely claim ownership. In this article, we propose a novel DNN watermarking solution, named$HufuNet$, to detect IP infringement of DNN models against the above mentioned attacks. Furthermore, HufuNet is the first one theoretically proved to guarantee robustness against fine-tuning attacks. We evaluate HufuNet rigorously on four benchmark datasets with five popular DNN models, including convolutional neural network (CNN) and recurrent neural network (RNN). The experiments and analysis demonstrate that HufuNet is highly robust against model fine-tuning/pruning, transfer learning, kernels cutoff/supplement, functionality-equivalent attacks and fraudulent ownership claims, thus highly promising to protect large-scale DNN models in the real world.
Peizhuo Lv, Shengzhi Zhang, Kai Chen 0012, Ruigang Liang, Hualong Ma, Yue Zhao 0018, Yingjiu Li
IEEE Trans. Dependable Secur. Comput.7
2021 AI-Lancet: Locating Error-inducing Neurons to Optimize Neural Networks
abstract
Deep neural network (DNN) has been widely utilized in many areas due to its increasingly high accuracy. However, DNN models could also produce wrong outputs due to internal errors, which may lead to severe security issues. Unlike fixing bugs in traditional computer software, tracing the errors in DNN models and fixing them are much more difficult due to the uninterpretability of DNN. In this paper, we present a novel and systematic approach to trace and fix the errors in deep learning models. In particular, we locate the error-inducing neurons that play a leading role in the erroneous output. With the knowledge of error-inducing neurons, we propose two methods to fix the errors: the neuron-flip and the neuron-fine-tuning. We evaluate our approach using five different training datasets and seven different model architectures. The experimental results demonstrate its efficacy in different application scenarios, including backdoor removal and general defects fixing.
Yue Zhao 0018, Kai Chen 0012, Shengzhi Zhang
CCS1
2020 Devil's Whisper: A General Approach for Physical Adversarial Attacks against Commercial Black-box Speech Recognition Devices
Xuejing Yuan, Jiangshan Zhang, Yue Zhao 0018, Shengzhi Zhang, Kai Chen 0012, XiaoFeng Wang 0001
USENIX Security Symposium4
2019 Seeing isn't Believing: Towards More Robust Adversarial Attack Against Real World Object Detectors
abstract
Recently Adversarial Examples (AEs) that deceive deep learning models have been a topic of intense research interest. Compared with the AEs in the digital space, the physical adversarial attack is considered as a more severe threat to the applications like face recognition in authentication, objection detection in autonomous driving cars, etc. In particular, deceiving the object detectors practically, is more challenging since the relative position between the object and the detector may keep changing. Existing works attacking object detectors are still very limited in various scenarios, e.g., varying distance and angles, etc. In this paper, we presented systematic solutions to build robust and practical AEs against real world object detectors. Particularly, for Hiding Attack (HA), we proposed thefeature-interference reinforcement (FIR) method and theenhanced realistic constraints generation (ERG) to enhance robustness, and for Appearing Attack (AA), we proposed thenested-AE, which combines two AEs together to attack object detectors in both long and short distance. We also designed diverse styles of AEs to make AA more surreptitious. Evaluation results show that our AEs can attack the state-of-the-art real-time object detectors (i.e., YOLO V3 and faster-RCNN) at the success rate up to 92.4% with varying distance from 1m to 25m and angles from -60º to 60º. Our AEs are also demonstrated to be highly transferable, capable of attacking another three state-of-the-art black-box models with high success rate.
Yue Zhao 0018, Ruigang Liang, Qintao Shen, Shengzhi Zhang, Kai Chen 0012
CCS1
2018 CommanderSong: A Systematic Approach for Practical Adversarial Voice Recognition
Xuejing Yuan, Yue Zhao 0018, Yunhui Long, Kai Chen 0012, Shengzhi Zhang, Heqing Huang 0001, XiaoFeng Wang 0001, Carl A. Gunter
USENIX Security Symposium3