EDBT 2026 Demo / reviewers in the wild / expert
Yuqi Jia 0001
dblp:359/3184-1
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0009-0006-2954-3609ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
Reachal Wang, Yuqi Jia 0001, Neil Zhenqiang Gong |
NDSS | 2 |
| 2026 | A Critical Evaluation of Defenses against Prompt Injection Attacks: [Dataset/Tool Paper]abstractLarge Language Models (LLMs) are vulnerable to prompt injection attacks, and several defenses have recently been proposed, often claiming to mitigate these attacks successfully. However, we argue that existing studies lack a principled approach to evaluating these defenses. In this paper, we argue the need to assess defenses across two critical dimensions: (1) effectiveness, measured against both existing and adaptive prompt injection attacks involving diverse target and injected prompts, and (2) general-purpose utility, ensuring that the defense does not compromise the foundational capabilities of the LLM. Our critical evaluation reveals that prior studies have not followed such a comprehensive evaluation methodology. When assessed using this principled approach, we show that existing defenses are not as successful as previously reported. This work provides a foundation for evaluating future defenses and guiding their development. Our anonymous code and data are available at https://github.com/PIEval123/PIEval. Yuqi Jia 0001, Zedian Shao, Yupei Liu, Jinyuan Jia 0001, Dawn Song, Neil Zhenqiang Gong |
SACMAT | 1 |
| 2026 | PromptLocate: Localizing Prompt Injection AttacksabstractPrompt injection attacks deceive a large language model into completing an attacker-specified task instead of its intended task by contaminating its input data with an injected prompt, which consists of injected instruction(s) and data. Localizing the injected prompt within contaminated data is crucial for post-attack forensic analysis and data recovery. Despite its growing importance, prompt injection localization remains largely unexplored. In this work, we bridge this gap by proposing PromptLocate, the first method for localizing injected prompts. PromptLocate comprises three steps: (1) splitting the contaminated data into semantically coherent segments, (2) identifying segments contaminated by injected instructions, and (3) pinpointing segments contaminated by injected data. We show PromptLocate accurately localizes injected prompts across eight existing and eight adaptive attacks. Yuqi Jia 0001, Yupei Liu, Zedian Shao, Jinyuan Jia 0001, Neil Zhenqiang Gong |
SP | 1 |
| 2025 | Competitive Advantage Attacks to Decentralized Federated LearningabstractDecentralized federated learning (DFL) enables clients (e.g., hospitals and banks) to jointly train machine learning models without a central orchestration server. In each global training round, each client trains a local model on its own training data and then they exchange local models for aggregation. In this work, we propose SelfishAttack, a new family of attacks to DFL. In SelfishAttack, a set of selfish clients aim to achieve competitive advantages over the remaining non-selfish ones, i.e., the final learnt local models of the selfish clients are more accurate than those of the non-selfish ones. Towards this goal, the selfish clients send carefully crafted local models to each remaining non-selfish one in each global training round. We formulate finding such local models as an optimization problem and propose methods to solve it when DFL uses different aggregation rules. Theoretically, we show that our methods find the optimal solutions to the optimization problem. Empirically, we show that SelfishAttack successfully increases the accuracy gap (i.e., competitive advantage) between the final learnt local models of selfish clients and those of non-selfish ones. Moreover, SelfishAttack achieves larger accuracy gaps than poisoning attacks when extended to increase competitive advantages. Yuqi Jia 0001, Minghong Fang, Neil Zhenqiang Gong |
NeurIPS | 1 |
| 2025 | Tracing Back the Malicious Clients in Poisoning Attacks to Federated LearningabstractPoisoning attacks compromise the training phase of federated learning (FL) such that the learned global model misclassifies attacker-chosen inputs called target inputs. Existing defenses mainly focus on protecting the training phase of FL such that the learnt global model is poison free. However, these defenses often achieve limited effectiveness when the clients' local training data is highly non-iid or the number of malicious clients is large, as confirmed in our experiments. In this work, we propose FLForensics, the first poison-forensics method for FL. FLForensics complements existing training-phase defenses. In particular, when training-phase defenses fail and a poisoned global model is deployed, FLForensics aims to trace back the malicious clients that performed the poisoning attack after a misclassified target input is identified. We theoretically show that FLForensics can accurately distinguish between benign and malicious clients under a formal definition of poisoning attack. Moreover, we empirically show the effectiveness of FLForensics at tracing back both existing and adaptive poisoning attacks on five benchmark datasets. Yuqi Jia 0001, Minghong Fang, Hongbin Liu 0005, Jinghuai Zhang, Neil Zhenqiang Gong |
NeurIPS | 1 |
| 2025 | DataSentinel: A Game-Theoretic Detection of Prompt Injection AttacksabstractLLM-integrated applications and agents are vulnerable to prompt injection attacks, where an attacker injects prompts into their inputs to induce attacker-desired outputs. A detection method aims to determine whether a given input is contaminated by an injected prompt. However, existing detection methods have limited effectiveness against state-of-the-art attacks, let alone adaptive ones. In this work, we propose DataSentinel, a game-theoretic method to detect prompt injection attacks. Specifically, DataSentinel fine-tunes an LLM to detect inputs contaminated with injected prompts that are strategically adapted to evade detection. We formulate this as a minimax optimization problem, with the objective of fine-tuning the LLM to detect strong adaptive attacks. Furthermore, we propose a gradient-based method to solve the minimax optimization problem by alternating between the inner max and outer min problems. Our evaluation results on multiple benchmark datasets and LLMs show that DataSentinel effectively detects both existing and adaptive prompt injection attacks. Our code and data are available at: https://github.com/liu00222/Open-Prompt-Injection. Yupei Liu, Yuqi Jia 0001, Jinyuan Jia 0001, Dawn Song, Neil Zhenqiang Gong |
SP | 2 |
| 2025 | Evaluating LLM-based Personal Information Extraction and Countermeasures
Yupei Liu, Yuqi Jia 0001, Jinyuan Jia 0001, Neil Zhenqiang Gong |
USENIX Security Symposium | 2 |
| 2025 | Periodic Recovery From Poisoning Attacks in Machine LearningabstractRecovery from poisoning attacks aims to eliminate the influence of a given set of deleted poisoned training data on a model. In practice, model recovery often happensperiodicallysince data deletion occurs repeatedly after a model has been trained. Existing efficient model recovery methods are designed forsingle-shotmodel recovery. When applied to periodic model recovery, they treat the instances of recovery independently, leading to a large total overhead over time. In this work, we propose PeriRecover, an efficient periodic model recovery method. Our key idea is to extract some common information during the original model training, which can be used to accelerate all instances of model recovery. In particular, we propose to compute and store the diagonals of the Hessian matrix of the loss function during the original model training. Given such information, each instance of model recovery can efficiently estimate the gradients to update the model instead of exactly computing them. Theoretically, we show that the model recovered by PeriRecover is close to the one recovered by training-from-scratch under some assumptions, achievingcertified recovery. Empirically, we apply PeriRecover to supervised learning and recommender systems, and we consider targeted attacks and untargeted attacks. Our results show that PeriRecover is much more efficient and/or accurate than existing model recovery methods. Yuepeng Hu, Minghong Fang, Yuqi Jia 0001, Hongbin Liu 0005, Neil Zhenqiang Gong |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | Unlocking the Potential of Federated Learning: The Symphony of Dataset Distillation via Deep Generative Latents
Yuqi Jia 0001, Saeed Vahidian, Jingwei Sun 0002, Vyacheslav Kungurtsev, Neil Zhenqiang Gong, Yiran Chen 0001 |
ECCV (78) | 1 |
| 2024 | Formalizing and Benchmarking Prompt Injection Attacks and Defenses
Yupei Liu, Yuqi Jia 0001, Runpeng Geng, Jinyuan Jia 0001, Neil Zhenqiang Gong |
USENIX Security Symposium | 2 |