EDBT 2026 Demo / reviewers in the wild / expert
Zhaohan Xi
dblp:224/9296
· DBLP profile ↗
12ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Data to Defense: The Role of Curation in Aligning Large Language Models Against Safety CompromiseabstractLarge language models (LLMs) are widely adapted for downstream applications through fine-tuning, a process named customization.However, recent studies have identified a vulnerability during this process, where malicious samples can compromise the robustness of LLMs and amplify harmful behaviors.To address this challenge, we propose an adaptive data curation approach allowing any text to be curated to enhance its effectiveness in counteracting harmful samples during customization.To avoid the need for additional defensive modules, we further introduce a comprehensive mitigation framework spanning the lifecycle of the customization process: before customization to immunize LLMs against future compromise attempts, during customization to neutralize risks, and after customization to restore compromised models.Experimental results demonstrate a significant reduction in compromising effects, achieving up to a 100% success rate in generating safe responses.By combining adaptive data curation with lifecycle-based mitigation strategies, this work represents a solid step forward in mitigating compromising risks and ensuring the secure adaptation of LLMs. Xiaoqun Liu, Jiacheng Liang, Luoxi Tang, Muchao Ye, Zhaohan Xi |
EMNLP | 6 |
| 2024 | PromptFix: Few-shot Backdoor Removal via Adversarial Prompt TuningabstractPre-trained language models (PLMs) have attracted enormous attention over the past few years with their unparalleled performances. Meanwhile, the soaring cost to train PLMs as well as their amazing generalizability have jointly contributed to few-shot fine-tuning and prompting as the most popular training paradigms for natural language processing (NLP) models. Nevertheless, existing studies have shown that these NLP models can be backdoored such that model behavior is manipulated when trigger tokens are presented. In this paper, we propose PromptFix, a novel backdoor mitigation strategy for NLP models via adversarial prompt-tuning in few-shot settings. Unlike existing NLP backdoor removal methods, which rely on accurate trigger inversion and subsequent model fine-tuning, PromptFix keeps the model parameters intact and only utilizes two extra sets of soft tokens which approximate the trigger and counteract it respectively. The use of soft tokens and adversarial optimization eliminates the need to enumerate possible backdoor configurations and enables an adaptive balance between trigger finding and preservation of performance. Experiments with various backdoor attacks validate the effectiveness of the proposed method and the performances when domain shift is present further shows PromptFix's applicability to models pre-trained on unknown data source which is the common case in prompt tuning scenarios. Tianrong Zhang, Zhaohan Xi, Ting Wang 0006, Prasenjit Mitra 0001 |
NAACL-HLT | 2 |
| 2024 | On the Difficulty of Defending Contrastive Learning against Backdoor Attacks
Changjiang Li, Ren Pang, Bochuan Cao, Zhaohan Xi, Shouling Ji, Ting Wang 0006 |
USENIX Security Symposium | 4 |
| 2023 | An Embarrassingly Simple Backdoor Attack on Self-supervised LearningabstractAs a new paradigm in machine learning, self-supervised learning (SSL) is capable of learning high-quality representations of complex data without relying on labels. In addition to eliminating the need for labeled data, research has found that SSL improves the adversarial robustness over supervised learning since lacking labels makes it more challenging for adversaries to manipulate model predictions. However, the extent to which this robustness superiority generalizes to other types of attacks remains an open question.We explore this question in the context of backdoor attacks. Specifically, we design and evaluate Ctrl, an embarrassingly simple yet highly effective self-supervised backdoor attack. By only polluting a tiny fraction of training data (≤ 1%) with indistinguishable poisoning samples, Ctrl causes any trigger-embedded input to be misclassified to the adversary's designated class with a high probability (≥ 99%) at inference time. Our findings suggest that SSL and supervised learning are comparably vulnerable to backdoor attacks. More importantly, through the lens of Ctrl, we study the inherent vulnerability of SSL to backdoor attacks. With both empirical and analytical evidence, we reveal that the representation invariance property of SSL, which benefits adversarial robustness, may also be the very reason making SSL highly susceptible to backdoor attacks. Our findings also imply that the existing defenses against supervised backdoor attacks are not easily retrofitted to the unique vulnerability of SSL. Code is available at: https://github.com/meet-cjli/CTRL Changjiang Li, Ren Pang, Zhaohan Xi, Tianyu Du, Shouling Ji, Yuan Yao 0001, Ting Wang 0006 |
ICCV | 3 |
| 2023 | The Dark Side of AutoML: Towards Architectural Backdoor Search
Ren Pang, Changjiang Li, Zhaohan Xi, Shouling Ji, Ting Wang 0006 |
ICLR | 3 |
| 2023 | Defending Pre-trained Language Models as Few-shot Learners against Backdoor AttacksabstractPre-trained language models (PLMs) have demonstrated remarkable performance as few-shot learners. However, their security risks under such settings are largely unexplored. In this work, we conduct a pilot study showing that PLMs as few-shot learners are highly vulnerable to backdoor attacks while existing defenses are inadequate due to the unique challenges of few-shot scenarios. To address such challenges, we advocate MDP, a novel lightweight, pluggable, and effective defense for PLMs as few-shot learners. Specifically, MDP leverages the gap between the masking-sensitivity of poisoned and clean samples: with reference to the limited few-shot data as distributional anchors, it compares the representations of given samples under varying masking and identifies poisoned samples as ones with significant variations. We show analytically that MDP creates an interesting dilemma for the attacker to choose between attack effectiveness and detection evasiveness. The empirical evaluation using benchmark datasets and representative attacks validates the efficacy of MDP. The code of MDP is publicly available. Zhaohan Xi, Tianyu Du, Changjiang Li, Ren Pang, Shouling Ji, Fenglong Ma, Ting Wang 0006 |
NeurIPS | 1 |
| 2023 | On the Security Risks of Knowledge Graph Reasoning
Zhaohan Xi, Tianyu Du, Changjiang Li, Ren Pang, Shouling Ji, Xiapu Luo, Xusheng Xiao, Fenglong Ma, Ting Wang 0006 |
USENIX Security Symposium | 1 |
| 2022 | TrojanZoo: Towards Unified, Holistic, and Practical Evaluation of Neural BackdoorsabstractNeural backdoors represent one primary threat to the security of deep learning systems. The intensive research has produced a plethora of backdoor attacks/defenses, resulting in a constant arms race. However, due to the lack of evaluation benchmarks, many critical questions remain under-explored: (i) what are the strengths and limitations of different attacks/defenses? (ii) what are the best practices to operate them? and (iii) how can the existing attacks/defenses be further improved? To bridge this gap, we design and implement TROJAN-ZOO, the first open-source platform for evaluating neural backdoor attacks/defenses in a unified, holistic, and practical manner. Thus far, focusing on the computer vision domain, it has incorporated 8 representative attacks, 14 state-of-the-art defenses, 6 attack performance metrics, 10 defense utility metrics, as well as rich tools for in-depth analysis of the attack-defense interactions. Leveraging TROJANZOO, we conduct a systematic study on the existing attacks/defenses, unveiling their complex design spectrum: both manifest intricate trade-offs among multiple desiderata (e.g., the effectiveness, evasiveness, and transferability of attacks). We further explore improving the existing attacks/defenses, leading to a number of interesting findings: (i) one-pixel triggers often suffice; (ii) training from scratch often outperforms perturbing benign models to craft trojan models; (iii) optimizing triggers and trojan models jointly greatly improves both attack effectiveness and evasiveness; (iv) individual defenses can often be evaded by adaptive attacks; and (v) exploiting model interpretability significantly improves defense robustness. We envision that TROJANZOO will serve as a valuable platform to facilitate future research on neural backdoors. Ren Pang, Xiangshan Gao, Zhaohan Xi, Shouling Ji, Peng Cheng 0001, Xiapu Luo, Ting Wang 0006 |
EuroS&P | 4 |
| 2022 | Seeing is Living? Rethinking the Security of Facial Liveness Verification in the Deepfake Era
Changjiang Li, Li Wang 0120, Shouling Ji, Xuhong Zhang 0002, Zhaohan Xi, Shanqing Guo, Ting Wang 0006 |
USENIX Security Symposium | 5 |
| 2022 | On the Security Risks of AutoML
Ren Pang, Zhaohan Xi, Shouling Ji, Xiapu Luo, Ting Wang 0006 |
USENIX Security Symposium | 2 |
| 2021 | Graph Backdoor
Zhaohan Xi, Ren Pang, Shouling Ji, Ting Wang 0006 |
USENIX Security Symposium | 1 |
| 2018 | Towards a Secure Zero-rating Framework with Three Parties
Yinzhi Cao, Zhaohan Xi, Shihao Jing, Humberto J. La Roche |
USENIX Security Symposium | 4 |