VLDB 2026 Research / reviewers in the wild / expert
Elizabeth Clark
dblp:148/6935
· DBLP profile ↗
15ranked-venue papers
6as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing FeedbackabstractCan LLMs provide support to creative writers by giving meaningful writing feedback?In this paper, we explore the challenges and limitations of model-generated writing feedback by defining a new task, dataset, and evaluation frameworks.To study model performance in a controlled manner, we present a novel test set of 1,300 stories that we corrupted to intentionally introduce writing issues.We study the performance of commonly used LLMs in this task with both automatic and human evaluation metrics.Our analysis shows that current models have strong out-of-the-box behavior in many respects-providing specific and mostly accurate writing feedback.However, models often fail to identify the biggest writing issue in the story and to correctly decide when to offer critical vs. positive feedback. Hannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella Lapata |
ACL (1) | 2 |
| 2025 | Agents' Room: Narrative Generation through Multi-step CollaborationabstractWriting compelling fiction is a multifaceted process combining elements such as crafting a plot, developing interesting characters, and using evocative language. While large language models (LLMs) show promise for story writing, they currently rely heavily on intricate prompting, which limits their use. We propose Agents' Room, a generation framework inspired by narrative theory, that decomposes narrative writing into subtasks tackled by specialized agents. To illustrate our method, we introduce Tell Me A Story, a high-quality dataset of complex writing prompts and human-written stories, and a novel evaluation framework designed specifically for assessing long narratives. We show that Agents' Room generates stories that are preferred by expert evaluators over those produced by baseline systems by leveraging collaboration and specialization to decompose the complex story writing task into tractable components. We provide extensive analysis with automated and human-based metrics of the generated output. Fantine Huot, Reinald Kim Amplayo, Jennimaria Palomaki, Alice Shoshana Jakobovits, Elizabeth Clark, Mirella Lapata |
ICLR | 5 |
| 2024 | Evaluating LLMs for Targeted Concept Simplification for Domain-Specific TextsabstractOne useful application of NLP models is to support people in reading complex text from unfamiliar domains (e.g., scientific articles).Simplifying the entire text makes it understandable but sometimes removes important details.On the contrary, helping adult readers understand difficult concepts in context can enhance their vocabulary and knowledge.In a preliminary human study, we first identify that lack of context and unfamiliarity with difficult concepts is a major reason for adult readers' difficulty with domain-specific text.We then introduce targeted concept simplification, a simplification task for rewriting text to help readers comprehend text containing unfamiliar concepts.We also introduce WIKIDOMAINS 1 , a new dataset of 22k definitions from 13 academic domains paired with a difficult concept within each definition.We benchmark the performance of open-source and commercial LLMs, and a simple dictionary baseline on this task across human judgments of ease of understanding and meaning preservation.Interestingly, our human judges preferred explanations about the difficult concept more than simplification of the concept phrase.Further, no single model achieved superior performance across all quality dimensions, and automated metrics also show low correlations with human evaluations of concept simplification (∼ 0.2), opening up rich avenues for research on personalized human reading comprehension support.* Work done as student researcher at Google DeepMind. 1 https://github.com/google-deepmind/wikidomains Sumit Asthana, Hannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella Lapata |
EMNLP | 3 |
| 2023 | Dialect-robust Evaluation of Generated TextabstractJiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann |
ACL (1) | 3 |
| 2023 | A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for SummarizationabstractLining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, João Sedoc. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Lining Zhang, Simon Mille, Yufang Hou 0001, Daniel Deutsch, Elizabeth Clark, Yixin Liu 0003, Saad Mahamood, Sebastian Gehrmann, Miruna-Adriana Clinciu, Khyathi Raghavi Chandu, João Sedoc |
ACL (1) | 5 |
| 2023 | SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization EvaluationabstractElizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das, Ankur Parikh. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Elizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das 0001, Ankur P. Parikh |
EMNLP | 1 |
| 2023 | Don't Take This Out of Context!: On the Need for Contextual Models and Evaluations for Stylistic RewritingabstractMost existing stylistic text rewriting methods and evaluation metrics operate on a sentence level, but ignoring the broader context of the text can lead to preferring generic, ambiguous, and incoherent rewrites.In this paper, we investigate integrating the preceding textual context into both the rewriting and evaluation stages of stylistic text rewriting, and introduce a new composite contextual evaluation metric CtxSimFit that combines similarity to the original sentence with contextual cohesiveness.We comparatively evaluate non-contextual and contextual rewrites in formality, toxicity, and sentiment transfer tasks.Our experiments show that humans significantly prefer contextual rewrites as more fitting and natural over non-contextual ones, yet existing sentence-level automatic metrics (e.g., ROUGE, SBERT) correlate poorly with human preferences (ρ=0-0.3).In contrast, human preferences are much better reflected by both our novel CtxSimFit (ρ=0.7-0.9) as well as proposed context-infused versions of common metrics (ρ=0.4-0.7).Overall, our findings highlight the importance of integrating context into the generation and especially the evaluation stages of stylistic text rewriting. Akhila Yerukola, Elizabeth Clark, Maarten Sap |
EMNLP | 3 |
| 2023 | Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated TextabstractEvaluation practices in natural language generation (NLG) have many known flaws, but improved evaluation approaches are rarely widely adopted. This issue has become more urgent, since neural generation models have improved to the point where their outputs can often no longer be distinguished based on the surface-level features that older metrics rely on. This paper surveys the issues with human and automatic model evaluations and with commonly used datasets in NLG that have been pointed out over the past 20 years. We summarize, categorize, and discuss how researchers have been addressing these issues and what their findings mean for the current state of model evaluations. Building on those insights, we lay out a long-term vision for evaluation research and propose concrete steps for researchers to improve their evaluation processes. Finally, we analyze 66 generation papers from recent NLP conferences in how well they already follow these suggestions and identify which areas require more drastic changes to the status quo. Sebastian Gehrmann, Elizabeth Clark, Thibault Sellam |
J. Artif. Intell. Res. | 2 |
| 2021 | All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated TextabstractElizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, Noah A. Smith. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, Noah A. Smith |
ACL/IJCNLP (1) | 1 |
| 2021 | Choose Your Own Adventure: Paired Suggestions in Collaborative Writing for Evaluating Story Generation ModelsabstractStory generation is an open-ended and subjective task, which poses a challenge for evaluating story generation models.We present CHOOSE YOUR OWN ADVENTURE, a collaborative writing setup for pairwise model evaluation.Two models generate suggestions to people as they write a short story; we ask writers to choose one of the two suggestions, and we observe which model's suggestions they prefer.The setup also allows further analysis based on the revisions people make to the suggestions.We show that these measures, combined with automatic metrics, provide an informative picture of the models' performance, both in cases where the differences in generation methods are small (nucleus vs. top-k sampling) and large (GPT2 vs. Fusion models). Elizabeth Clark, Noah A. Smith |
NAACL-HLT | 1 |
| 2021 | TuringAdvice: A Generative and Dynamic Evaluation of Language UseabstractRowan Zellers, Ari Holtzman, Elizabeth Clark, Lianhui Qin, Ali Farhadi, Yejin Choi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Rowan Zellers, Ari Holtzman, Elizabeth Clark, Lianhui Qin, Ali Farhadi, Yejin Choi 0001 |
NAACL-HLT | 3 |
| 2019 | Sentence Mover's Similarity: Automatic Evaluation for Multi-Sentence TextsabstractFor evaluating machine-generated texts, automatic methods hold the promise of avoiding collection of human judgments, which can be expensive and time-consuming.The most common automatic metrics, like BLEU and ROUGE, depend on exact word matching, an inflexible approach for measuring semantic similarity.We introduce methods based on sentence mover's similarity; our automatic metrics evaluate text in a continuous space using word and sentence embeddings.We find that sentence-based metrics correlate with human judgments significantly better than ROUGE, both on machine-generated summaries (average length of 3.4 sentences) and human-authored essays (average length of 7.5).We also show that sentence mover's similarity can be used as a reward when learning a generation model via reinforcement learning; we present both automatic and human evaluations of summaries learned in this way, finding that our approach outperforms ROUGE. Elizabeth Clark, Asli Celikyilmaz, Noah A. Smith |
ACL (1) | 1 |
| 2019 | Counterfactual Story Reasoning and GenerationabstractLianhui Qin, Antoine Bosselut, Ari Holtzman, Chandra Bhagavatula, Elizabeth Clark, Yejin Choi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lianhui Qin, Antoine Bosselut, Ari Holtzman, Chandra Bhagavatula, Elizabeth Clark, Yejin Choi 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2018 | Creative Writing with a Machine in the Loop: Case Studies on Slogans and StoriesabstractAs the quality of natural language generated by artificial intelligence systems improves, writing interfaces can support interventions beyond grammar-checking and spell-checking, such as suggesting content to spark new ideas. To explore the possibility of machine-in-the-loop creative writing, we performed two case studies using two system prototypes, one for short story writing and one for slogan writing. Participants in our studies were asked to write with a machine in the loop or alone (control condition). They assessed their writing and experience through surveys and an open-ended interview. We collected additional assessments of the writing from Amazon Mechanical Turk crowdworkers. Our findings indicate that participants found the process fun and helpful and could envision use cases for future systems. At the same time, machine suggestions do not necessarily lead to better written artifacts. We therefore suggest novel natural language models and design choices that may better support creative writing. Elizabeth Clark, Anne Spencer Ross, Chenhao Tan, Yangfeng Ji, Noah A. Smith |
IUI | 1 |
| 2018 | Neural Text Generation in Stories Using Entity Representations as ContextabstractElizabeth Clark, Yangfeng Ji, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Elizabeth Clark, Yangfeng Ji, Noah A. Smith |
NAACL-HLT | 1 |