VLDB 2026 Research / reviewers in the wild / expert
Yanting Wang 0001
dblp:72/7715-1
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Security and privacy · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PIArena: A Platform for Prompt Injection EvaluationabstractPrompt injection attacks pose serious security risks across a wide range of real-world applications.While receiving increasing attention, the community faces a critical gap: the lack of a unified platform for prompt injection evaluation.This makes it challenging to reliably compare defenses, understand their true robustness under diverse attacks, or assess how well they generalize across tasks and benchmarks.For instance, many defenses initially reported as effective were later found to exhibit limited robustness on diverse datasets and attacks.To bridge this gap, we introduce PIArena, a unified and extensible platform for prompt injection evaluation that enables users to easily integrate state-of-the-art attacks and defenses and evaluate them across a variety of existing and new benchmarks.We also design a dynamic strategy-based attack that adaptively optimizes injected prompts based on defense feedback.Through comprehensive evaluation using PIArena, we uncover critical limitations of state-of-the-art defenses: limited generalizability across tasks, vulnerability to adaptive attacks, and fundamental challenges when an injected task aligns with the target task.The code and datasets are available at https://github.com/sleeepeer/PIArena. Runpeng Geng, Chenlong Yin, Yanting Wang 0001, Jinyuan Jia 0001 |
ACL (1) | 3 |
| 2026 | AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
Yanting Wang 0001, Runpeng Geng, Jinyuan Jia 0001 |
SP | 1 |
| 2025 | TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised LearningabstractAn image encoder pre-trained by self-supervised learning can be used as a general-purpose feature extractor to build downstream classifiers for various downstream tasks. However, many studies showed that an attacker can embed a trojan into an encoder such that multiple downstream classifiers built based on the trojaned encoder simultaneously inherit the trojan behavior. In this work, we propose TrojanDec, the first data-free method to identify and recover a test input embedded with a trigger. Given a (trojaned or clean) encoder and a test input, TrojanDec first predicts whether the test input is trojaned. If not, the test input is processed in a normal way to maintain the utility. Otherwise, the test input will be further restored to remove the trigger. Our extensive evaluation shows that TrojanDec can effectively identify the trojan (if any) from a given test input and recover it under state-of-the-art trojan attacks. We further demonstrate by experiments that our TrojanDec outperforms the state-of-the-art defenses. Yupei Liu, Yanting Wang 0001, Jinyuan Jia 0001 |
AAAI | 2 |
| 2025 | TracLLM: A Generic Framework for Attributing Long Context LLMs
Yanting Wang 0001, Runpeng Geng, Jinyuan Jia 0001 |
USENIX Security Symposium | 1 |
| 2024 | MMCert: Provable Defense Against Adversarial Attacks to Multi-Modal ModelsabstractDifferent from a unimodal model whose input is from a single modality, the input (called multi-modal input) of a multi-modal model is from multiple modalities such as image, 3D points, audio, text, etc. Similar to unimodal models, many existing studies show that a multi-modal model is also vulnerable to adversarial perturbation, where an attacker could add small perturbation to all modalities of a multi-modal input such that the multi-modal model makes incorrect predictions for it. Existing certified defenses are mostly designed for unimodal models, which achieve sub optimal certified robustness guarantees when extended to multi-modal models as shown in our experimental results. In our work, we propose MMCert, the first certified defense against adversarial attacks to a multi-modal model. We derive a lower bound on the performance of our MMCert under arbitrary adversarial attacks with bounded perturbations to both modalities (e.g., in the context of auto-driving, we bound the number of changed pixels in both RGB image and depth image). We evaluate our MMCert using two benchmark datasets: one for the multi-modal road segmentation task and the other for the multi-modal emotion recognition task. Moreover, we compare our MMCert with a state-of-the-art certified defense extended from unimodal models. Our experimental results show that our MMCert outperforms the baseline. Yanting Wang 0001, Hongye Fu, Jinyuan Jia 0001 |
CVPR | 1 |
| 2024 | FCert: Certifiably Robust Few-Shot Classification in the Era of Foundation ModelsabstractFew-shot classification with foundation models (e.g., CLIP, DINOv2, PaLM-2) enables users to build an accurate classifier with a few labeled training samples (called support samples) for a classification task. However, an attacker could perform data poisoning attacks by manipulating some support samples such that the classifier makes the attacker-desired, arbitrary prediction for a testing input. Empirical defenses cannot provide formal robustness guarantees, leading to a cat-and-mouse game between the attacker and defender. Existing certified defenses are designed for traditional supervised learning, resulting in sub-optimal performance when extended to few-shot classification. In our work, we propose FCert, the first certified defense against data poisoning attacks to few-shot classification. We show our FCert provably predicts the same label for a testing input under arbitrary data poisoning attacks when the total number of poisoned support samples is bounded. We perform extensive experiments on benchmark few-shot classification datasets with foundation models released by OpenAI, Meta, and Google in both vision and text domains. Our experimental results show our FCert: 1) maintains classification accuracy without attacks, 2) outperforms existing state-of-the-art certified defenses for data poisoning attacks, and 3) is efficient and general. Yanting Wang 0001, Jinyuan Jia 0001 |
SP | 1 |