Yupei Liu

dblp:204/1178 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Critical Evaluation of Defenses against Prompt Injection Attacks: [Dataset/Tool Paper]
abstract
Large Language Models (LLMs) are vulnerable to prompt injection attacks, and several defenses have recently been proposed, often claiming to mitigate these attacks successfully. However, we argue that existing studies lack a principled approach to evaluating these defenses. In this paper, we argue the need to assess defenses across two critical dimensions: (1) effectiveness, measured against both existing and adaptive prompt injection attacks involving diverse target and injected prompts, and (2) general-purpose utility, ensuring that the defense does not compromise the foundational capabilities of the LLM. Our critical evaluation reveals that prior studies have not followed such a comprehensive evaluation methodology. When assessed using this principled approach, we show that existing defenses are not as successful as previously reported. This work provides a foundation for evaluating future defenses and guiding their development. Our anonymous code and data are available at https://github.com/PIEval123/PIEval.
Yuqi Jia 0001, Zedian Shao, Yupei Liu, Jinyuan Jia 0001, Dawn Song, Neil Zhenqiang Gong
SACMAT3
2026 PromptLocate: Localizing Prompt Injection Attacks
abstract
Prompt injection attacks deceive a large language model into completing an attacker-specified task instead of its intended task by contaminating its input data with an injected prompt, which consists of injected instruction(s) and data. Localizing the injected prompt within contaminated data is crucial for post-attack forensic analysis and data recovery. Despite its growing importance, prompt injection localization remains largely unexplored. In this work, we bridge this gap by proposing PromptLocate, the first method for localizing injected prompts. PromptLocate comprises three steps: (1) splitting the contaminated data into semantically coherent segments, (2) identifying segments contaminated by injected instructions, and (3) pinpointing segments contaminated by injected data. We show PromptLocate accurately localizes injected prompts across eight existing and eight adaptive attacks.
Yuqi Jia 0001, Yupei Liu, Zedian Shao, Jinyuan Jia 0001, Neil Zhenqiang Gong
SP2
2025 TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
abstract
An image encoder pre-trained by self-supervised learning can be used as a general-purpose feature extractor to build downstream classifiers for various downstream tasks. However, many studies showed that an attacker can embed a trojan into an encoder such that multiple downstream classifiers built based on the trojaned encoder simultaneously inherit the trojan behavior. In this work, we propose TrojanDec, the first data-free method to identify and recover a test input embedded with a trigger. Given a (trojaned or clean) encoder and a test input, TrojanDec first predicts whether the test input is trojaned. If not, the test input is processed in a normal way to maintain the utility. Otherwise, the test input will be further restored to remove the trigger. Our extensive evaluation shows that TrojanDec can effectively identify the trojan (if any) from a given test input and recover it under state-of-the-art trojan attacks. We further demonstrate by experiments that our TrojanDec outperforms the state-of-the-art defenses.
Yupei Liu, Yanting Wang 0001, Jinyuan Jia 0001
AAAI1
2025 SecureGaze: Defending Gaze Estimation Against Backdoor Attacks
abstract
Gaze estimation models are widely used in applications such as driver attention monitoring and human-computer interaction. While many methods for gaze estimation exist, they rely heavily on data-hungry deep learning to achieve high performance. This reliance often forces practitioners to harvest training data from unverified public datasets, outsource model training, or rely on pre-trained models. However, such practices expose gaze estimation models to backdoor attacks. In such attacks, adversaries inject backdoor triggers by poisoning the training data, creating a backdoor vulnerability: the model performs normally with benign inputs, but produces manipulated gaze directions when a specific trigger is present. This compromises the security of many gaze-based applications, such as causing the model to fail in tracking the driver's attention. To date, there is no defense that addresses backdoor attacks on gaze estimation models. In response, we introduce SecureGaze, the first solution designed to protect gaze estimation models from such attacks. Unlike classification models, defending gaze estimation poses unique challenges due to its continuous output space and globally activated backdoor behavior. By identifying distinctive characteristics of backdoored gaze estimation models, we develop a novel and effective approach to reverse-engineer the trigger function for reliable backdoor detection. Extensive evaluations in both digital and physical worlds demonstrate that SecureGaze effectively counters a range of backdoor attacks and outperforms seven state-of-the-art defenses adapted from classification models.
Lingyu Du, Yupei Liu, Jinyuan Jia 0001, Guohao Lan
SenSys2
2025 DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
abstract
LLM-integrated applications and agents are vulnerable to prompt injection attacks, where an attacker injects prompts into their inputs to induce attacker-desired outputs. A detection method aims to determine whether a given input is contaminated by an injected prompt. However, existing detection methods have limited effectiveness against state-of-the-art attacks, let alone adaptive ones. In this work, we propose DataSentinel, a game-theoretic method to detect prompt injection attacks. Specifically, DataSentinel fine-tunes an LLM to detect inputs contaminated with injected prompts that are strategically adapted to evade detection. We formulate this as a minimax optimization problem, with the objective of fine-tuning the LLM to detect strong adaptive attacks. Furthermore, we propose a gradient-based method to solve the minimax optimization problem by alternating between the inner max and outer min problems. Our evaluation results on multiple benchmark datasets and LLMs show that DataSentinel effectively detects both existing and adaptive prompt injection attacks. Our code and data are available at: https://github.com/liu00222/Open-Prompt-Injection.
Yupei Liu, Yuqi Jia 0001, Jinyuan Jia 0001, Dawn Song, Neil Zhenqiang Gong
SP1
2025 Evaluating LLM-based Personal Information Extraction and Countermeasures
Yupei Liu, Yuqi Jia 0001, Jinyuan Jia 0001, Neil Zhenqiang Gong
USENIX Security Symposium1
2025 Policy-Oriented Cognitive Risk Map Modeling for Lane Change via Deep Successor Representation
abstract
Risk assessment plays an essential role in the improvement of driving safety for intelligent vehicles. Current methods ignoring the predictive and personalized impact of driving policies weaken the effectiveness of risk assessment and lead to human-machine conflicts. By combining subjective cognition of drivers and objective risk metrics, a policy-oriented cognitive risk map (POCRM) is proposed in this paper to encode different driving policies in risk assessment for lane-changing scenarios. To obtain the objective safety metrics, insecurity quantification is built based on the fuzzy theory and fault tree analysis. The subjective cognition of drivers for different driving policies is modeled by deep successor representation and encoded in POCRM using deep reinforcement learning. Driving data collected from the public dataset for realistic traffic environment are used to evaluate the proposed POCRM. The experimental results show that the risk map can take into account future risks and provide driving advice that balances human-machine conflicts with safety in scenarios where drivers can or cannot correctly perceive risk.
Danni Chen, Chao Lu 0006, Yupei Liu, Xianghao Meng, Jianwei Gong
IEEE Trans. Intell. Transp. Syst.3
2025 Risk Assessment of Cyclists in the Mixed Traffic Based on Multilevel Graph Representation
abstract
Accurate assessment of the cyclist risk is a crucial task for the safety system of autonomous vehicles (AVs). This paper proposes a framework for defining and evaluating cyclist risk levels, considering behavioral cues. The framework comprises three modules: the cyclist graph construction (CGC) module, the risk label generation (RLG) module, and the risk assessment (RA) module. The CGC module constructs a spatiotemporal graph model of the cyclist with both the behavioral and risk information. The RLG module leverages the graph representation method (GRM) to extract features and assigns risk labels using unsupervised learning. The RA module employs spatiotemporal graph convolutional networks (ST-GCN) to extract features from the cyclist graph. Additionally, it facilitates feature fusion through interactions between the human body and the two-wheeler and between hierarchical levels. The fused features, along with the risk labels, are used to train a classifier for the risk assessment of cyclists. The proposed framework is validated using real-world data, and the comparative results with state-of-the-art methods demonstrate the effectiveness and accuracy of the proposed approach in cyclist risk assessment in mixed traffic.
Gege Cui, Chao Lu 0006, Yupei Liu, Xianghao Meng, Jianwei Gong
IEEE Trans. Intell. Transp. Syst.3
2024 Formalizing and Benchmarking Prompt Injection Attacks and Defenses
Yupei Liu, Yuqi Jia 0001, Runpeng Geng, Jinyuan Jia 0001, Neil Zhenqiang Gong
USENIX Security Symposium1
2023 PORE: Provably Robust Recommender Systems against Data Poisoning Attacks
Jinyuan Jia 0001, Yupei Liu, Yuepeng Hu, Neil Zhenqiang Gong
USENIX Security Symposium2
2022 Certified Robustness of Nearest Neighbors against Data Poisoning and Backdoor Attacks
abstract
Data poisoning attacks and backdoor attacks aim to corrupt a machine learning classifier via modifying, adding, and/or removing some carefully selected training examples, such that the corrupted classifier makes incorrect predictions as the attacker desires. The key idea of state-of-the-art certified defenses against data poisoning attacks and backdoor attacks is to create a majority vote mechanism to predict the label of a testing example. Moreover, each voter is a base classifier trained on a subset of the training dataset. Classical simple learning algorithms such as k nearest neighbors (kNN) and radius nearest neighbors (rNN) have intrinsic majority vote mechanisms. In this work, we show that the intrinsic majority vote mechanisms in kNN and rNN already provide certified robustness guarantees against data poisoning attacks and backdoor attacks. Moreover, our evaluation results on MNIST and CIFAR10 show that the intrinsic certified robustness guarantees of kNN and rNN outperform those provided by state-of-the-art certified defenses. Our results serve as standard baselines for future certified defenses against data poisoning attacks and backdoor attacks.
Jinyuan Jia 0001, Yupei Liu, Neil Zhenqiang Gong
AAAI2
2022 StolenEncoder: Stealing Pre-trained Encoders in Self-supervised Learning
abstract
Pre-trained encoders are general-purpose feature extractors that can be used for many downstream tasks. Recent progress in self-supervised learning can pre-train highly effective encoders using a large volume of unlabeled data, leading to the emerging encoder as a service (EaaS). A pre-trained encoder may be deemed confidential because its training often requires lots of data and computation resources as well as its public release may facilitate misuse of AI, e.g., for deepfakes generation. In this paper, we propose the first attack called StolenEncoder to steal pre-trained image encoders. We evaluate StolenEncoder on multiple target encoders pre-trained by ourselves and three real-world target encoders including the ImageNet encoder pre-trained by Google, CLIP encoder pre-trained by OpenAI, and Clarifai's General Embedding encoder deployed as a paid EaaS. Our results show that the encoders stolen by StolenEncoder have similar functionality with the target encoders. In particular, the downstream classifiers built upon a target encoder and a stolen encoder have similar accuracy. Moreover, stealing a target encoder using StolenEncoder requires much less data and computation resources than pre-training it from scratch. We also explore three defenses that perturb feature vectors produced by a target encoder. Our evaluation shows that these defenses are not enough to mitigate StolenEncoder.
Yupei Liu, Jinyuan Jia 0001, Hongbin Liu 0005, Neil Zhenqiang Gong
CCS1
2022 BadEncoder: Backdoor Attacks to Pre-trained Encoders in Self-Supervised Learning
abstract
Self-supervised learning in computer vision aims to pre-train an image encoder using a large amount of unlabeled images or (image, text) pairs. The pre-trained image encoder can then be used as a feature extractor to build downstream classifiers for many downstream tasks with a small amount of or no labeled training data. In this work, we propose BadEncoder, the first backdoor attack to self-supervised learning. In particular, our BadEncoder injects backdoors into a pre-trained image encoder such that the downstream classifiers built based on the backdoored image encoder for different downstream tasks simultaneously inherit the backdoor behavior. We formulate our BadEncoder as an optimization problem and we propose a gradient descent based method to solve it, which produces a backdoored image encoder from a clean one. Our extensive empirical evaluation results on multiple datasets show that our BadEncoder achieves high attack success rates while preserving the accuracy of the downstream classifiers. We also show the effectiveness of BadEncoder using two publicly available, real-world image encoders, i.e., Google’s image encoder pre-trained on ImageNet and OpenAI’s Contrastive Language-Image Pre-training (CLIP) image encoder pre-trained on 400 million (image, text) pairs collected from the Internet. Moreover, we consider defenses including Neural Cleanse and MNTD (empirical defenses) as well as PatchGuard (a provable defense). Our results show that these defenses are insufficient to defend against BadEncoder, highlighting the needs for new defenses against our BadEncoder. Our code is publicly available at: https://github.com/jjy1994/BadEncoder.
Jinyuan Jia 0001, Yupei Liu, Neil Zhenqiang Gong
SP2
2022 Security Analysis of Camera-LiDAR Fusion Against Black-Box Attacks on Autonomous Vehicles
Spencer Hallyburton, Yupei Liu, Z. Morley Mao, Miroslav Pajic
USENIX Security Symposium2