VLDB 2026 Research / reviewers in the wild / expert
Tinghao Xie
dblp:307/5298
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Trustworthy machine learning · 52% Language models and text generation · 24% Generative modeling · 10% | |
| Network and information security
6 papers |
Security and privacy of machine learning · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Security and privacy of machine learning › poisoning attack
fine-tuning attack |
1.6 | 2 | 2025 | On Evaluating the Durability of Safeguards for Open-Weight LLMs · ICLR 2025 Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! · ICLR 2024 |
Security and privacy of machine learning
large language model safety |
1.6 | 2 | 2025 | On Evaluating the Durability of Safeguards for Open-Weight LLMs · ICLR 2025 Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! · ICLR 2024 |
Security and privacy of machine learning › adversarial attack › backdoor attack
backdoor defense |
1.4 | 2 | 2024 | BaDExpert: Extracting Backdoor Functionality for Accurate Backdoor Input Detection · ICLR 2024 Revisiting the Assumption of Latent Separability for Backdoor Defenses · ICLR 2023 |
Security and privacy of machine learning › adversarial attack
backdoor attack |
1.2 | 2 | 2023 | Towards A Proactive ML Approach for Detecting Backdoor Poison Samples · USENIX Security Symposium 2023 Towards Practical Deployment-Stage Backdoor Attack on Deep Neural Networks · CVPR 2022 |
Natural language and speech › Language models and text generation › text generation › large language model generation
prompt-based generation |
0.9 | 1 | 2025 | Fantastic Copyrighted Beasts and How (Not) to Generate Them · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | Fantastic Copyrighted Beasts and How (Not) to Generate Them · ICLR 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 2 | 2023 | Revisiting the Assumption of Latent Separability for Backdoor Defenses · ICLR 2023 Towards A Proactive ML Approach for Detecting Backdoor Poison Samples · USENIX Security Symposium 2023 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.8 | 1 | 2024 | Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications · ICML 2024 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.8 | 1 | 2024 | Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications · ICML 2024 |
Security and privacy of machine learning › large language model alignment
safety alignment |
0.8 | 1 | 2024 | Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! · ICLR 2024 |
Machine learning › Trustworthy machine learning › robustness
backdoor defense |
0.7 | 1 | 2023 | Revisiting the Assumption of Latent Separability for Backdoor Defenses · ICLR 2023 |
Security and privacy of machine learning › poisoning attack
model weight attack |
0.6 | 1 | 2022 | Towards Practical Deployment-Stage Backdoor Attack on Deep Neural Networks · CVPR 2022 |
Natural language and speech › Language models and text generation
large language model |
0.5 | 2 | 2025 | On Evaluating the Durability of Safeguards for Open-Weight LLMs · ICLR 2025 Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! · ICLR 2024 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.2 | 1 | 2024 | Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
fine-tuning · 2.4red-teaming · 1.6red teaming · 1.6ensemble strategy · 1.5latent separability analysis · 1.3prompt rewriting · 0.9negative prompting · 0.9human-in-the-loop · 0.9LLM-as-a-judge · 0.9pruning · 0.8low-rank modification · 0.8machine learning · 0.7subnet replacement attack · 0.6gray-box attack · 0.6adversarial weight attack · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fantastic Copyrighted Beasts and How (Not) to Generate ThemabstractRecent studies show that image and video generation models can be prompted to reproduce copyrighted content from their training data, raising serious legal con- cerns about copyright infringement. Copyrighted characters (e.g., Mario, Batman) present a significant challenge: at least one lawsuit has already awarded damages based on the generation of such characters. Consequently, commercial services like DALL·E have started deploying interventions. However, little research has systematically examined these problems: (1) Can users easily prompt models to generate copyrighted characters, even if it is unintentional?; (2) How effective are the existing mitigation strategies? To address these questions, we introduce a novel evaluation framework with metrics that assess both the generated image’s similarity to copyrighted characters and its consistency with user intent, grounded in a set of popular copyrighted characters from diverse studios and regions. We show that state-of-the-art image and video generation models can still generate characters even if characters’ names are not explicitly mentioned, sometimes with only two generic keywords (e.g., prompting with “videogame, plumber” consistently gener- ates Nintendo’s Mario character). We also introduce semi-automatic techniques to identify such keywords or descriptions that trigger character generation. Using this framework, we evaluate mitigation strategies, including prompt rewriting and new approaches we propose. Our findings reveal that common methods, such as DALL·E’s prompt rewriting, are insufficient alone and require supplementary strategies like negative prompting. Our work provides empirical grounding for discussions on copyright mitigation strategies and offers actionable insights for model deployers implementing these safeguards. Luxi He, Yangsibo Huang, Tinghao Xie, Luke Zettlemoyer, Chiyuan Zhang, Danqi Chen 0001, Peter Henderson 0002 |
ICLR | 4 |
| 2025 | On Evaluating the Durability of Safeguards for Open-Weight LLMsabstractMany stakeholders---from model developers to policymakers---seek to minimize the risks of large language models (LLMs). Key to this goal is whether technical safeguards can impede the misuse of LLMs, even when models are customizable via fine-tuning or when model weights are openly available. Several recent studies have proposed methods to produce durable LLM safeguards for open-weight LLMs that can withstand adversarial modifications of the model's weights via fine-tuning. This holds the promise of raising adversaries' costs even under strong threat models where adversaries can directly fine-tune parameters. However, we caution against over-reliance on such methods in their current state. Through several case studies, we demonstrate that even the evaluation of these defenses is exceedingly difficult and can easily mislead audiences into thinking that safeguards are more durable than they really are. We draw lessons from the failure modes that we identify and suggest that future research carefully cabin claims to more constrained, well-defined, and rigorously examined threat models, which can provide useful and candid assessments to stakeholders. Xiangyu Qi, Boyi Wei, Nicholas Carlini, Yangsibo Huang, Tinghao Xie, Luxi He, Matthew Jagielski, Milad Nasr, Prateek Mittal, Peter Henderson 0002 |
ICLR | 5 |
| 2025 | SORRY-Bench: Systematically Evaluating Large Language Model Safety RefusalabstractEvaluating aligned large language models' (LLMs) ability to recognize and reject unsafe user requests is crucial for safe, policy-compliant deployments. Existing evaluation efforts, however, face three limitations that we address with **SORRY-Bench**, our proposed benchmark. **First**, existing methods often use coarse-grained taxonomies of unsafe topics, and are over-representing some fine-grained topics. For example, among the ten existing datasets that we evaluated, tests for refusals of self-harm instructions are over 3x less represented than tests for fraudulent activities. SORRY-Bench improves on this by using a fine-grained taxonomy of 44 potentially unsafe topics, and 440 class-balanced unsafe instructions, compiled through human-in-the-loop methods. **Second**, evaluations often overlook the linguistic formatting of prompts, like different languages, dialects, and more --- which are only implicitly considered in many evaluations. We supplement SORRY-bench with 20 diverse linguistic augmentations to systematically examine these effects. **Third**, existing evaluations rely on large LLMs (e.g., GPT-4) for evaluation, which can be computationally expensive. We investigate design choices for creating a fast, accurate automated safety evaluator. By collecting 7K+ human annotations and conducting a meta-evaluation of diverse LLM-as-a-judge designs, we show that fine-tuned 7B LLMs can achieve accuracy comparable to GPT-4 scale LLMs, with lower computational cost. Putting these together, we evaluate over 50 proprietary and open-weight LLMs on SORRY-Bench, analyzing their distinctive safety refusal behaviors. We hope our effort provides a building block for systematic evaluations of LLMs' safety refusal capabilities, in a balanced, granular, and efficient manner. Benchmark demo, data, code, and models are available through [https://sorry-bench.github.io](https://sorry-bench.github.io). Tinghao Xie, Xiangyu Qi, Yi Zeng 0005, Yangsibo Huang, T. W. U. Madhushani, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng 0007, Ruoxi Jia 0001, Bo Li 0026, Kai Li 0001, Danqi Chen 0001, Peter Henderson 0002, Prateek Mittal |
ICLR | 1 |
| 2024 | Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!abstractOptimizing large language models (LLMs) for downstream use cases often involves the customization of pre-trained LLMs through further fine-tuning. Meta's open-source release of Llama models and OpenAI's APIs for fine-tuning GPT-3.5 Turbo on customized datasets accelerate this trend. But, what are the safety costs associated with such customized fine-tuning? While existing safety alignment techniques restrict harmful behaviors of LLMs at inference time, they do not cover safety risks when fine-tuning privileges are extended to end-users. Our red teaming studies find that the safety alignment of LLMs can be compromised by fine-tuning with only a few adversarially designed training examples. For instance, we jailbreak GPT-3.5 Turbo's safety guardrails by fine-tuning it on only 10 such examples at a cost of less than $0.20 via OpenAI's APIs, making the model responsive to nearly any harmful instructions. Disconcertingly, our research also reveals that, even without malicious intent, simply fine-tuning with benign and commonly used datasets can also inadvertently degrade the safety alignment of LLMs, though to a lesser extent. These findings suggest that fine-tuning aligned LLMs introduces new safety risks that current safety infrastructures fall short of addressing --- even if a model's initial safety alignment is impeccable, how can it be maintained after customized fine-tuning? We outline and critically analyze potential mitigations and advocate for further research efforts toward reinforcing safety protocols for the customized fine-tuning of aligned LLMs. (This paper contains red-teaming data and model-generated content that can be offensive in nature.) Xiangyu Qi, Yi Zeng 0005, Tinghao Xie, Ruoxi Jia 0001, Prateek Mittal, Peter Henderson 0002 |
ICLR | 3 |
| 2024 | BaDExpert: Extracting Backdoor Functionality for Accurate Backdoor Input DetectionabstractWe present a novel defense, against backdoor attacks on Deep Neural Networks (DNNs), wherein adversaries covertly implant malicious behaviors (backdoors) into DNNs. Our defense falls within the category of post-development defenses that operate independently of how the model was generated. The proposed defense is built upon a novel reverse engineering approach that can directly extract **backdoor functionality** of a given backdoored model to a *backdoor expert* model. The approach is straightforward --- finetuning the backdoored model over a small set of intentionally mislabeled clean samples, such that it unlearns the normal functionality while still preserving the backdoor functionality, and thus resulting in a model~(dubbed a backdoor expert model) that can only recognize backdoor inputs. Based on the extracted backdoor expert model, we show the feasibility of devising highly accurate backdoor input detectors that filter out the backdoor inputs during model inference. Further augmented by an ensemble strategy with a finetuned auxiliary model, our defense, **BaDExpert** (**Ba**ckdoor Input **D**etection with Backdoor **Expert**), effectively mitigates 17 SOTA backdoor attacks while minimally impacting clean utility. The effectiveness of BaDExpert has been verified on multiple datasets (CIFAR10, GTSRB and ImageNet) across various model architectures (ResNet, VGG, MobileNetV2 and Vision Transformer). Our code is integrated into our research toolbox: [https://github.com/vtu81/backdoor-toolbox](https://github.com/vtu81/backdoor-toolbox). Tinghao Xie, Xiangyu Qi, Jiachen T. Wang, Prateek Mittal |
ICLR | 1 |
| 2024 | Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsabstractLarge language models (LLMs) show inherent brittleness in their safety mechanisms, as evidenced by their susceptibility to jailbreaking and even non-malicious fine-tuning. This study explores this brittleness of safety alignment by leveraging pruning and low-rank modifications. We develop methods to identify critical regions that are vital for safety guardrails, and that are disentangled from utility-relevant regions at both the neuron and rank levels. Surprisingly, the isolated regions we find are sparse, comprising about $3$ % at the parameter level and $2.5$ % at the rank level. Removing these regions compromises safety without significantly impacting utility, corroborating the inherent brittleness of the model's safety mechanisms. Moreover, we show that LLMs remain vulnerable to low-cost fine-tuning attacks even when modifications to the safety-critical regions are restricted. These findings underscore the urgent need for more robust safety strategies in LLMs. Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang 0001, Peter Henderson 0002 |
ICML | 4 |
| 2023 | Revisiting the Assumption of Latent Separability for Backdoor Defenses
Xiangyu Qi, Tinghao Xie, Saeed Mahloujifar, Prateek Mittal |
ICLR | 2 |
| 2023 | Towards A Proactive ML Approach for Detecting Backdoor Poison Samples
Xiangyu Qi, Tinghao Xie, Jiachen T. Wang, Saeed Mahloujifar, Prateek Mittal |
USENIX Security Symposium | 2 |
| 2022 | Towards Practical Deployment-Stage Backdoor Attack on Deep Neural NetworksabstractOne major goal of the AI security community is to securely and reliably produce and deploy deep learning models for real-world applications. To this end, data poisoning based backdoor attacks on deep neural networks (DNNs) in the production stage (or training stage) and corresponding defenses are extensively explored in recent years. Ironically, backdoor attacks in the deployment stage, which can often happen in unprofessional users' devices and are thus arguably far more threatening in real-world scenarios, draw much less attention of the community. We attribute this imbalance of vigilance to the weak practicality of existing deployment-stage backdoor attack algorithms and the insufficiency of real-world attack demonstrations. To fill the blank, in this work, we study the realistic threat of deployment-stage backdoor attacks on DNNs. We base our study on a commonly used deployment-stage attack paradigm - adversarial weight attack, where adversaries selectively modify model weights to embed backdoor into deployed DNNs. To approach realistic practicality, we propose the first gray-box and physically realizable weights attack algorithm for backdoor injection, namely subnet replacement attack (SRA), which only requires architecture information of the victim model and can support physical triggers in the real world. Extensive experimental simulations and system-level real- world attack demonstrations are conducted. Our results not only suggest the effectiveness and practicality of the proposed attack algorithm, but also reveal the practical risk of a novel type of computer virus that may widely spread and stealthily inject backdoor into DNN models in user devices. By our study, we call for more attention to the vulnerability of DNNs in the deployment stage. Xiangyu Qi, Tinghao Xie, Ruizhe Pan, Jifeng Zhu, Kai Bu |
CVPR | 2 |