VLDB 2026 Research / reviewers in the wild / expert
Yao Dou
dblp:262/0556
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2025
0009-0001-2157-5482ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CROSSNEWS: A Cross-Genre Authorship Verification and Attribution BenchmarkabstractAuthorship models have historically generalized poorly to new domains because of the wide distribution of author-identifying signals across domains. In particular, the effects of topic and genre are highly domain-dependent and impact authorship analysis performance greatly. This paper addresses the existing data gap in authorship for these resources by introducing CROSSNEWS, a novel cross-genre dataset that connects formal journalistic articles and casual social media posts. CROSSNEWS is the largest authorship dataset of its kind for supporting both verification and attribution tasks, with comprehensive topic and genre annotations. We use CROSSNEWS to demonstrate that current models exhibit poor performance in genre transfer scenarios, underscoring the need for authorship models robust to genre-specific effects. We also explore SELMA, a new LLM embedding approach for large-scale authorship setups that outperforms existing models in both same-genre and cross-genre settings. Marcus Ma, Duong Minh Le, Junmo Kang, Yao Dou, John Cadigan, Dayne Freitag, Alan Ritter, Wei Xu 0004 |
AAAI | 4 |
| 2025 | SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?abstractYao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, Jianfeng Gao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu 0004, Jianfeng Gao 0001 |
EMNLP | 1 |
| 2025 | CollabLLM: From Passive Responders to Active CollaboratorsabstractLarge Language Models are typically trained with next-turn rewards, limiting their ability to optimize for long-term interaction. As a result, they often respond passively to ambiguous or open-ended user requests, failing to help users reach their ultimate intents and leading to inefficient conversations. To address these limitations, we introduce CollabLLM, a novel and general training framework that enhances multiturn human-LLM collaboration. Its key innovation is a collaborative simulation that estimates the long-term contribution of responses
using Multiturn-aware Rewards. By reinforcement fine-tuning these rewards, CollabLLM goes beyond responding to user requests, and actively uncovers user intent and offers insightful suggestions—a key step towards more human-centered AI. We also devise a multiturn interaction benchmark with three challenging tasks such as document creation. CollabLLM significantly outperforms our baselines with averages of 18.5% higher task performance and 46.3% improved interactivity by LLM judges. Finally, we conduct a large user study with 201 judges, where CollabLLM increases user satisfaction by 17.6% and reduces user spent time by 10.4%. Shirley Wu, Michel Galley, Baolin Peng, Hao Cheng 0002, Gavin Li, Yao Dou, Weixin Cai, James Zou 0001, Jure Leskovec, Jianfeng Gao 0001 |
ICML | 6 |
| 2025 | Measuring, Modeling, and Helping People Account for Privacy Risks in Online Self-Disclosures with AIabstractIn pseudonymous online fora like Reddit, the benefits of self-disclosure are often apparent to users (e.g., I can vent about my in-laws to understanding strangers), but the privacy risks are more abstract (e.g., will my partner be able to tell that this is me?). Prior work has sought to develop natural language processing (NLP) tools that help users identify potentially risky self-disclosures in their text, but none have been designed for or evaluated with the users they hope to protect. Absent this assessment, these tools will be limited by the social-technical gap: users need assistive tools that help them make informed decisions, not paternalistic tools that tell them to avoid self-disclosure altogether. To bridge this gap, we conducted a study with N =21 Reddit users; we had them use a state-of-the-art NLP disclosure detection model on two of their authored posts and asked them questions to understand if and how the model helped, where it fell short, and how it could be improved to help them make more informed decisions. Despite its imperfections, users responded positively to the model and highlighted its use as a tool that can help them catch mistakes, inform them of risks they were unaware of, and encourage self-reflection. However, our work also shows how, to be useful and usable, AI for supporting privacy decision-making must account for posting context, disclosure norms, and users' lived threat models, and provide explanations that help contextualize detected risks. Isadora Krsek, Anubha Kabra, Yao Dou, Tarek Naous, Laura A. Dabbish, Alan Ritter, Wei Xu 0004, Sauvik Das |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | Reducing Privacy Risks in Online Self-Disclosures with Language ModelsabstractYao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, Wei Xu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, Wei Xu 0004 |
ACL (1) | 1 |
| 2024 | Improving Minimum Bayes Risk Decoding with Multi-PromptabstractWhile instruction fine-tuned LLMs are effective text generators, sensitivity to prompt construction makes performance unstable and suboptimal in practice.Relying on a single 'best' prompt cannot capture all differing approaches to a generation problem.Using this observation, we propose multi-prompt decoding, where many candidate generations are decoded from a prompt bank at inference-time.To ensemble candidates, we use Minimum Bayes Risk (MBR) decoding, which selects a final output using a trained value metric.We show multiprompt improves MBR across a comprehensive set of conditional generation tasks (Figure 1), and show this is a result of estimating a more diverse and higher quality candidate space than that of a single prompt.Further experiments confirm multi-prompt improves generation across tasks, models and metrics. 1 David Heineman, Yao Dou, Wei Xu 0004 |
EMNLP | 2 |
| 2024 | GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-ExplanationabstractResearch on jailbreaking has been valuable for testing and understanding the safety and security issues of large language models (LLMs).In this paper, we introduce Iterative Refinement Induced Self-Jailbreak (IRIS), a novel approach that leverages the reflective capabilities of LLMs for jailbreaking with only blackbox access.Unlike previous methods, IRIS simplifies the jailbreaking process by using a single model as both the attacker and target.This method first iteratively refines adversarial prompts through self-explanation, which is crucial for ensuring that even well-aligned LLMs obey adversarial instructions.IRIS then rates and enhances the output given the refined prompt to increase its harmfulness.We find that IRIS achieves jailbreak success rates of 98% on GPT-4, 92% on GPT-4 Turbo, and 94% on Llama-3.1-70B in under 7 queries.It significantly outperforms prior approaches in automatic, black-box, and interpretable jailbreaking, while requiring substantially fewer queries, thereby establishing a new standard for interpretable jailbreaking methods. Govind Ramesh, Yao Dou, Wei Xu 0004 |
EMNLP | 2 |
| 2023 | LENS: A Learnable Evaluation Metric for Text SimplificationabstractTraining learnable metrics using modern language models has recently emerged as a promising method for the automatic evaluation of machine translation.However, existing human evaluation datasets for text simplification have limited annotations that are based on unitary or outdated models, making them unsuitable for this approach.To address these issues, we introduce the SIMPEVAL corpus that contains: SIMPEVAL PAST , comprising 12K human ratings on 2.4K simplifications of 24 past systems, and SIMPEVAL 2022 , a challenging simplification benchmark consisting of over 1K human ratings of 360 simplifications including GPT-3.5 generated text.Training on SIMPEVAL, we present LENS, a Learnable Evaluation Metric for Text Simplification.Extensive empirical results show that LENS correlates much better with human judgment than existing metrics, paving the way for future progress in the evaluation of text simplification.We also introduce RANK & RATE, a human evaluation framework that rates simplifications from several models in a list-wise manner using an interactive interface, which ensures both consistency and accuracy in the evaluation process and is used to create the SIMPEVAL datasets. Mounica Maddela, Yao Dou, David Heineman, Wei Xu 0004 |
ACL (1) | 2 |
| 2023 | Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSAabstractLarge language models (e.g., GPT-4) are uniquely capable of producing highly rated text simplification, yet current human evaluation methods fail to provide a clear understanding of systems' specific strengths and weaknesses.To address this limitation, we introduce SALSA, an edit-based human annotation framework that enables holistic and fine-grained text simplification evaluation.We develop twenty one linguistically grounded edit types, covering the full spectrum of success and failure across dimensions of conceptual, syntactic and lexical simplicity.Using SALSA, we collect 19K edit annotations on 840 simplifications, revealing discrepancies in the distribution of simplification strategies performed by fine-tuned models, prompted LLMs and humans, and find GPT-3.5 performs more quality edits than humans, but still exhibits frequent errors.Using our finegrained annotations, we develop LENS-SALSA, a reference-free automatic simplification metric, trained to predict sentence-and word-level quality simultaneously.Additionally, we introduce word-level quality estimation for simplification and report promising baseline results.Our data, new metric, and annotation toolkit are available at https://salsa-eval.com. EXAMPLEZero-shot GPT-3. David Heineman, Yao Dou, Mounica Maddela, Wei Xu 0004 |
EMNLP | 2 |
| 2022 | Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine TextabstractModern neural language models can produce remarkably fluent and grammatical text.So much, in fact, that recent work by Clark et al. (2021) has reported that conventional crowdsourcing can no longer reliably distinguish between machine-authored (GPT-3) and humanauthored writing.As errors in machine generations become ever subtler and harder to spot, it poses a new challenge to the research community for robust machine text evaluation.We propose a new framework called SCARE-CROW for scrutinizing machine text via crowd annotation.To support the broad range of real machine errors that can be identified by laypeople, the ten error categories of SCARECROWsuch as redundancy , commonsense errors , and incoherence -are identified through several rounds of crowd annotation experiments without a predefined ontology.We then use SCARECROW to collect over 41k error spans in human-written and machinegenerated paragraphs of English language news text.We isolate factors for detailed analysis, including parameter count, training data, and various decoding-time configurations.Our approach successfully quantifies measurable gaps between human authored text and generations from models of several sizes, including fourteen configurations of GPT-3.In addition, our analysis unveils new insights, with detailed rationales provided by laypeople, e.g., that the commonsense capabilities have been improving with larger models while math capabilities have not, and that the choices of simple decoding hyperparameters can make remarkable differences on the perceived quality of machine text.We release our training material, annotation toolkit and dataset at Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith, Yejin Choi 0001 |
ACL (1) | 1 |
| 2022 | Improving Large-scale Paraphrase Acquisition and GenerationabstractThis paper addresses the quality issues in existing Twitter-based paraphrase datasets, and discusses the necessity of using two separate definitions of paraphrase for identification and generation tasks.We present a new Multi-Topic Paraphrase in Twitter (MULTIPIT) corpus that consists of a total of 130k sentence pairs with crowdsoursing (MULTIPIT CROWD ) and expert (MULTIPIT EXPERT ) annotations using two different paraphrase definitions for paraphrase identification, in addition to a multi-reference test set (MULTIPIT NMR ) and a large automatically constructed training set (MULTIPIT AUTO ) for paraphrase generation.With improved data annotation quality and task-specific paraphrase definition, the best pre-trained language model fine-tuned on our dataset achieves the stateof-the-art performance of 84.2 F 1 for automatic paraphrase identification.Furthermore, our empirical results also demonstrate that the paraphrase generation models trained on MUL-TIPIT AUTO generate more diverse and highquality paraphrases compared to their counterparts fine-tuned on other corpora such as Quora, MSCOCO, and ParaNMT.Topic Domains #Train #Dev #Test Sent/Tweet Len %Paraphrase #Trends/URLs #Uniq Sent %Multi-Ref Our Multi-Topic Paraphrase in Twitter (MULTIPIT CROWD ) Dataset Trends Sports 25,255 3,157 3,157 10.24 / 13.79 40.52% 1,201 34,786 17.89% Entertainment 11,547 1,443 1,444 10.44 / 13.80 62.33% 610 15,784 18.11% Event 8,624 1,078 1,079 10.86 / 15.32 82.83% 359 11,746 17.75% Others 17,751 2,219 2,219 10.41 / 14.56 67. Yao Dou, Wei Xu 0004 |
EMNLP | 1 |
| 2021 | MultiTalk: A Highly-Branching Dialog Testbed for Diverse ConversationsabstractWe study conversational dialog in which there are many possible responses to a given history. We present the MultiTalk Dataset, a corpus of over 320,000 sentences of written conversational dialog that balances a high branching factor (10) with several conversation turns (6) through selective branch continuation. We make multiple contributions to study dialog generation in the highly branching setting. In order to evaluate a diverse set of generations, we propose a simple scoring algorithm, based on bipartite graph matching, to optimally incorporate a set of diverse references. We study multiple language generation tasks at different levels of predictive conversation depth, using textual attributes induced automatically from pretrained classifiers. Our culminating task is a challenging theory of mind problem, a controllable generation task which requires reasoning about the expected reaction of the listener. Yao Dou, Maxwell Forbes, Ari Holtzman, Yejin Choi 0001 |
AAAI | 1 |