VLDB 2026 Research / reviewers in the wild / expert
Xinyue Shen 0001
dblp:148/9731-1 · also Xinyue (Vera) Shen
· DBLP profile ↗
18ranked-venue papers
6as first author
18since 2021 · last 2026
0009-0006-9954-587XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Open Schrödinger's Closed Box: Identifying Retrieval Augmented Generation in API-Accessible Large Language Model ServicesabstractLarge language models (LLMs) are powerful at question-answering but prone to hallucinations due to limited domain-specific or up-todate knowledge.Retrieval augmented generation (RAG) mitigates this by adding an external retriever and knowledge database, yet RAG remains vulnerable to targeted attacks that degrade outputs or manipulate opinions.Prior attacks typically assume adversaries know the service is RAG-enhanced and may even know deployment details, an assumption often invalid for real-world commercial LLMs that expose only black-box APIs.This opacity also risks misleading users about system capabilities.This work aims to bridge this gap by proposing RAG-ID, a framework for IDentifying RAG properties in LLM services.We classify adversaries into three knowledge levels and design six attack methods.Experiments show these attacks reliably detect RAG -up to 99.97% accuracy with partial or no optional knowledge, and nearly 100% when the LLM and database are known.After detection, RAG-ID can infer finer RAG properties (e.g., deployed LLM and knowledge database).We consider RAG-ID a reconnaissance tool for attackers, a way to facilitate users' transparent selection of LLM services, and a guide for RAG developers in refining security measures. Yukun Jiang 0001, Xinyue Shen 0001, Michael Backes 0001, Zheng Li 0023, Yang Zhang 0016 |
ACL (1) | 2 |
| 2025 | When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTsabstractKnowledge files have been widely used in large language model (LLM) agents, such as GPTs, to improve response quality.However, concerns about the potential leakage of knowledge files have grown significantly.Existing studies demonstrate that adversarial prompts can induce GPTs to leak knowledge file content.Yet, it remains uncertain whether additional leakage vectors exist, particularly given the complex data flows across clients, servers, and databases in GPTs.In this paper, we present a comprehensive risk assessment of knowledge file leakage, leveraging a novel workflow inspired by Data Security Posture Management (DSPM).Through the analysis of 651,022 GPT metadata, 11,820 flows, and 1,466 responses, we identify five leakage vectors: metadata, GPT initialization, retrieval, sandboxed execution environments, and prompts.These vectors enable adversaries to extract sensitive knowledge file data such as titles, content, types, and sizes.Notably, the activation of the built-in tool Code Interpreter leads to a privilege escalation vulnerability, enabling adversaries to directly download original knowledge files with a 95.95% success rate.Further analysis reveals that 28.80% of leaked files are copyrighted, including digital copies from major publishers and internal materials from a listed company.In the end, we provide actionable solutions for GPT builders and platform providers to secure the GPT data supply chain. Xinyue Shen 0001, Michael Backes 0001, Yang Zhang 0016 |
ACL (1) | 1 |
| 2025 | Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social MediaabstractSocial media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs). However, the misuse of AIGTs could have profound implications for public opinion, such as spreading misinformation and manipulating narratives. Despite its importance, it remains unclear how prevalent AIGTs are on social media. To address this gap, this paper aims to quantify and monitor the AIGTs on online social media platforms. We first collect a dataset (SM-D) with around 2.4M posts from 3 major social media platforms: Medium, Quora, and Reddit. Then, we construct a diverse dataset (AIGTBench) to train and evaluate AIGT detectors. AIGTBench combines popular open-source datasets and our AIGT datasets generated from social media texts by 12 LLMs, serving as a benchmark for evaluating mainstream detectors. With this setup, we identify the best-performing detector (OSM-Det). We then apply OSM-Det to SM-D to track AIGTs across social media platforms from January 2022 to October 2024, using the AI Attribution Rate (AAR) as the metric. Specifically, Medium and Quora exhibit marked increases in AAR, rising from 1.77% to 37.03% and 2.06% to 38.95%, respectively. In contrast, Reddit shows slower growth, with AAR increasing from 1.31% to 2.45% over the same period. Our further analysis indicates that AIGTs on social media differ from human-written texts across several dimensions, including linguistic patterns, topic distributions, engagement levels, and the follower distribution of authors. We envision our analysis and findings on AIGTs in social media can shed light on future research in this domain. Zhen Sun 0001, Zongmin Zhang, Xinyue Shen 0001, Yule Liu, Michael Backes 0001, Yang Zhang 0016, Xinlei He 0001 |
ACL (1) | 3 |
| 2025 | JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMsabstractJailbreak attacks aim to bypass the LLMs' safeguards.While researchers have proposed different jailbreak attacks in depth, they have done so in isolation-either with unaligned settings or comparing a limited range of methods.To fill this gap, we present a large-scale evaluation of various jailbreak attacks.We collect 17 representative jailbreak attacks, summarize their features, and establish a novel jailbreak attack taxonomy.Then we conduct comprehensive measurement and ablation studies across nine aligned LLMs on 160 forbidden questions from 16 violation categories.Also, we test jailbreak attacks under eight advanced defenses.Based on our taxonomy and experiments, we identify some important patterns, such as heuristicbased attacks could achieve high attack success rates but are easy to mitigate by defenses, causing low practicality.Our study offers valuable insights for future research on jailbreak attacks and defenses.We hope our work could help the community avoid incremental work and serve as an effective benchmark tool for practitioners. Junjie Chu 0002, Yugeng Liu, Ziqing Yang 0002, Xinyue Shen 0001, Michael Backes 0001, Yang Zhang 0016 |
ACL (1) | 4 |
| 2025 | UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated ImagesabstractWith the advent of text-to-image models and concerns about their misuse, developers are increasingly relying on image safety classifiers to moderate their generated unsafe images. Yet, the performance of current image safety classifiers remains unknown for both real-world and AI-generated images. In this work, we propose UnsafeBench, a benchmarking framework that evaluates the effectiveness and robustness of image safety classifiers, with a particular focus on the impact of AI-generated images on their performance. First, we curate a large dataset of 10K real-world and AI-generated images that are annotated as safe or unsafe based on a set of 11 unsafe categories of images (sexual, violent, hateful, etc.). Then, we evaluate the effectiveness and robustness of five popular image safety classifiers, as well as three classifiers that are powered by general-purpose visual language models. Our assessment indicates that existing image safety classifiers are not comprehensive and effective enough to mitigate the multifaceted problem of unsafe images. Also, there exists a distribution shift between real-world and AI-generated images in image qualities, styles, and layouts, leading to degraded effectiveness and robustness. Motivated by these findings, we build a comprehensive image moderation tool called PerspectiveVision, which improves the effectiveness and robustness of existing classifiers, especially on AI-generated images. UnsafeBench and PerspectiveVision can aid the research community in better understanding the landscape of image safety classification in the era of generative AI. Yiting Qu, Xinyue Shen 0001, Yixin Wu 0001, Michael Backes 0001, Savvas Zannettou, Yang Zhang 0016 |
CCS | 2 |
| 2025 | GPTracker: A Large-Scale Measurement of Misused GPTsabstractLarge language model (LLM)-powered agents, particularly GPTs by OpenAI, have revolutionized how AI is customized, deployed, and used. However, misuse of GPTs has emerged as a critical, yet largely underexplored, issue within OpenAI's GPT Store. In this paper, we present the first large-scale measurement study on misused GPTs. We introduce GPTRACKER, a framework designed to continuously collect GPTs from the official GPT Store and automate the interaction with them. As of the submission of this paper, GPTRACKER has collected 755,297 GPTs and 28,464 GPT conversation flows over eight months. Using an LLM-driven scoring system combined with human review, we identify 2,051 misused GPTs across ten forbidden scenarios. Through both static and dynamic analyses, we explore the landscape of these misused GPTs, including the trends, builders, operation mechanisms, and effectiveness. We find that builders of misused GPTs employ various tactics to bypass OpenAI's review system, such as integrating external APIs, hiding intention in descriptions, and URL redirection. Notably, GPTs activating external APIs are more likely to provide answers to inappropriate queries than other misused GPTs, showing an average 22.81% increase in answer rate in the Illegal Activity scenario. Leveraging VirusTotal, we identify 50 malicious domains shown on 446 GPTs, where 33 are labeled as phishing, 28 as malware, and 2 as spam, with some domains receiving multiple labels. We responsibly disclosed our findings to OpenAI on September 11, 2024, and November 12, 2024. 1,316 out of 1,804 GPTs reported in the first disclosure were removed by September 25. Our study sheds light on the alarming misuse of GPTs in the emerging GPT marketplace and offers actionable recommendations for stakeholders to mitigate future misuse.11Our code is available at https://github.com/TrustAIRLab/GPTracker. Disclaimer. This paper includes examples of hateful and disturbing content. Reader discretion is advised. Xinyue Shen 0001, Michael Backes 0001, Yang Zhang 0016 |
SP | 1 |
| 2025 | On the Effectiveness of Prompt Stealing Attacks on In-the-Wild PromptsabstractLarge Language Models (LLMs) have increased demand for high-quality prompts, which are now considered valuable commodities in prompt marketplaces. However, this demand has also led to the emergence of prompt stealing attacks, where the adversary attempts to infer prompts from generated outputs, threatening the intellectual property and business models of these marketplaces. Previous research primarily examines prompt stealing on academic datasets. The key question remains unanswered: Do these attacks genuinely threaten in-the-wild prompts curated by real-world users? In this paper, we provide the first systematic study on the efficacy of prompt stealing attacks against in-the-wild prompts. Our analysis shows that in-the-wild prompts differ significantly from academic ones in length, semantics, and topics. Our evaluation subsequently reveals that current prompt stealing attacks perform poorly in this context. To improve attack efficacy, we employ a Text Gradient based method to iteratively refine prompts to better reproduce outputs. This leads to enhanced attack performance, as evidenced by improvements in METEOR score from 0.207 to 0.253 for prompt recovery and from 0.323 to 0.440 for output recovery. Despite these improvements, we showcase that the fundamental challenges persist, highlighting the necessity for further research to improve and evaluate the effectiveness of prompt stealing attacks in practical scenarios. Yicong Tan, Xinyue Shen 0001, Michael Backes 0001, Yang Zhang 0016 |
SP | 2 |
| 2025 | From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language Models
Yihan Ma 0001, Xinyue Shen 0001, Yiting Qu, Ning Yu 0006, Michael Backes 0001, Savvas Zannettou, Yang Zhang 0016 |
USENIX Security Symposium | 2 |
| 2025 | HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
Xinyue Shen 0001, Yixin Wu 0001, Yiting Qu, Michael Backes 0001, Savvas Zannettou, Yang Zhang 0016 |
USENIX Security Symposium | 1 |
| 2024 | MGTBench: Benchmarking Machine-Generated Text DetectionabstractNowadays, powerful large language models (LLMs) such as ChatGPT have demonstrated revolutionary power in a variety of natural language processing (NLP) tasks such as text classification, sentiment analysis, language translation, and question-answering. Consequently, the detection of machine-generated texts (MGTs) is becoming increasingly crucial as LLMs become more advanced and prevalent. These models have the ability to generate human-like language, making it challenging to discern whether a text is authored by a human or a machine. This raises concerns regarding authenticity, accountability, and potential bias. However, existing methods for detecting MGTs are evaluated using different model architectures, datasets, and experimental settings, resulting in a lack of a comprehensive evaluation framework that encompasses various methodologies. Furthermore, it remains unclear how existing detection methods would perform against powerful LLMs. Xinlei He 0001, Xinyue Shen 0001, Zeyuan Chen 0002, Michael Backes 0001, Yang Zhang 0016 |
CCS | 2 |
| 2024 | "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsabstractThe misuse of large language models (LLMs) has drawn significant attention from the general public and LLM vendors. One particular type of adversarial prompt, known as jailbreak prompt, has emerged as the main attack vector to bypass the safeguards and elicit harmful content from LLMs. In this paper, employing our new framework JailbreakHub, we conduct a comprehensive analysis of 1,405 jailbreak prompts spanning from December 2022 to December 2023. We identify 131 jailbreak communities and discover unique characteristics of jailbreak prompts and their major attack strategies, such as prompt injection and privilege escalation. We also observe that jailbreak prompts increasingly shift from online Web communities to prompt-aggregation websites and 28 user accounts have consistently optimized jailbreak prompts over 100 days. To assess the potential harm caused by jailbreak prompts, we create a question set comprising 107,250 samples across 13 forbidden scenarios. Leveraging this dataset, our experiments on six popular LLMs show that their safeguards cannot adequately defend jailbreak prompts in all scenarios. Particularly, we identify five highly effective jailbreak prompts that achieve 0.95 attack success rates on ChatGPT (GPT-3.5) and GPT-4, and the earliest one has persisted online for over 240 days. We hope that our study can facilitate the research community and LLM vendors in promoting safer and regulated LLMs. Xinyue Shen 0001, Zeyuan Chen 0002, Michael Backes 0001, Yang Zhang 0016 |
CCS | 1 |
| 2024 | ModSCAN: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language ModalitiesabstractLarge vision-language models (LVLMs) have been rapidly developed and widely used in various fields, but the (potential) stereotypical bias in the model is largely unexplored.In this study, we present a pioneering measurement framework, ModSCAN, to SCAN the stereotypical bias within LVLMs from both vision and language Modalities.ModSCAN examines stereotypical biases with respect to two typical stereotypical attributes (gender and race) across three kinds of scenarios: occupations, descriptors, and persona traits.Our findings suggest that 1) the currently popular LVLMs show significant stereotype biases, with CogVLM emerging as the most biased model; 2) these stereotypical biases may stem from the inherent biases in the training dataset and pre-trained models; 3) the utilization of specific prompt prefixes (from both vision and language modalities) performs well in reducing stereotypical biases.We believe our work can serve as the foundation for understanding and addressing stereotypical bias in LVLMs. Yukun Jiang 0001, Zheng Li 0023, Xinyue Shen 0001, Yugeng Liu, Michael Backes 0001, Yang Zhang 0016 |
EMNLP | 3 |
| 2024 | The Death and Life of Great Prompts: Analyzing the Evolution of LLM Prompts from the Structural PerspectiveabstractEffective utilization of large language models (LLMs), such as ChatGPT, relies on the quality of input prompts.This paper explores prompt engineering, specifically focusing on the disparity between experimentally designed prompts and real-world "in-the-wild" prompts.We analyze 10,538 in-the-wild prompts collected from various platforms and develop a framework that decomposes the prompts into eight key components.Our analysis shows that Role and Requirement are the most prevalent two components.Roles specified in the prompts, along with their capabilities, have become increasingly varied over time, signifying a broader range of application scenarios for LLMs.However, from the response of GPT-4, there is a marginal improvement with a specified role, whereas leveraging less prevalent components such as Capability and Demonstration can result in a more satisfying response.Overall, our work sheds light on the essential components of in-the-wild prompts and the effectiveness of these components on the broader landscape of LLM prompt engineering, providing valuable guidelines for the LLM community to optimize high-quality prompts. Yihan Ma 0001, Xinyue Shen 0001, Yixin Wu 0001, Boyang Zhang 0008, Michael Backes 0001, Yang Zhang 0016 |
EMNLP | 2 |
| 2024 | Games and Beyond: Analyzing the Bullet Chats of Esports LivestreamingabstractEsports, short for electronic sports, is a form of competition using video games and has attracted more than 530 million audiences worldwide. To watch esports, people utilize online livestreaming platforms. Recently, a novel interaction method, namely "bullet chats," has been introduced on these platforms. Different from conventional comments, bullet chats are scrolling comments posted by audiences that are synchronized to the livestreaming timeline, enabling audiences to share and communicate their immediate perspectives. The real-time nature of bullet chats, therefore, brings a new perspective to esports analysis. In this paper, we conduct the first empirical study on the bullet chats for esports, focusing on one of the most popular video games, i.e., League of Legends (LoL). Specifically, we collect 21 million bullet chats of LoL from Jan. 2023 to Mar. 2023 across two mainstream platforms (Bilibili and Huya). By performing quantitative analysis, we reveal how the quantity and toxicity of bullet chats distribute (and change) w.r.t. three aspects, i.e., the season, the team, and the match. Our findings show that teams with higher rankings tend to attract a greater quantity of bullet chats, and these chats are often characterized by a higher degree of toxicity. We then utilize topic modeling to identify topics among bullet chats. Interestingly, we find that a considerable portion of topics (14.14% on Bilibili and 22.94% on Huya) discuss themes beyond the game, including genders, entertainment stars, non-esports athletes, and so on. Besides, by further modeling topics on toxic bullet chats, we find hateful speech targeting different social groups, ranging from professions, regions, etc. To the best of our knowledge, this work is the first measurement of bullet chats on esports livestreaming. We believe our study can shed light on esports research from the perspective of bullet chats. Yukun Jiang 0001, Xinyue Shen 0001, Rui Wen 0002, Zeyang Sha, Junjie Chu 0002, Yugeng Liu, Michael Backes 0001, Yang Zhang 0016 |
ICWSM | 2 |
| 2024 | Prompt Stealing Attacks Against Text-to-Image Generation Models
Xinyue Shen 0001, Yiting Qu, Michael Backes 0001, Yang Zhang 0016 |
USENIX Security Symposium | 1 |
| 2023 | Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image ModelsabstractState-of-the-art Text-to-Image models like Stable Diffusion and DALLE\cdot2 are revolutionizing how people generate visual content. At the same time, society has serious concerns about how adversaries can exploit such models to generate problematic or unsafe images. In this work, we focus on demystifying the generation of unsafe images and hateful memes from Text-to-Image models. We first construct a typology of unsafe images consisting of five categories (sexually explicit, violent, disturbing, hateful, and political). Then, we assess the proportion of unsafe images generated by four advanced Text-to-Image models using four prompt datasets. We find that Text-to-Image models can generate a substantial percentage of unsafe images; across four models and four prompt datasets, 14.56% of all generated images are unsafe. When comparing the four Text-to-Image models, we find different risk levels, with Stable Diffusion being the most prone to generating unsafe content (18.92% of all generated images are unsafe). Given Stable Diffusion's tendency to generate more unsafe content, we evaluate its potential to generate hateful meme variants if exploited by an adversary to attack a specific individual or community. We employ three image editing methods, DreamBooth, Textual Inversion, and SDEdit, which are supported by Stable Diffusion to generate variants. Our evaluation result shows that 24% of the generated images using DreamBooth are hateful meme variants that present the features of the original hateful meme and the target individual/community; these generated images are comparable to hateful meme variants collected from the real world. Overall, our results demonstrate that the danger of large-scale generation of unsafe images is imminent. We discuss several mitigating measures, such as curating training data, regulating prompts, and implementing safety filters, and encourage better safeguard tools to be developed to prevent unsafe generation.1 Our code is available at https://github.com/YitingQu/unsafe-diffusion. Yiting Qu, Xinyue Shen 0001, Xinlei He 0001, Michael Backes 0001, Savvas Zannettou, Yang Zhang 0016 |
CCS | 2 |
| 2022 | On Xing Tian and the Perseverance of Anti-China Sentiment Online
Xinyue Shen 0001, Xinlei He 0001, Michael Backes 0001, Jeremy Blackburn, Savvas Zannettou, Yang Zhang 0016 |
ICWSM | 1 |
| 2021 | Evil Under the Sun: Understanding and Discovering Attacks on Ethereum Decentralized Applications
Liya Su, Xinyue Shen 0001, Xiangyu Du, Xiaojing Liao, XiaoFeng Wang 0001, Luyi Xing, Baoxu Liu |
USENIX Security Symposium | 2 |