VLDB 2026 Research / reviewers in the wild / expert
Maria Antoniak
dblp:162/6913
· DBLP profile ↗
14ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-4807-2850ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Social Story Frames: Contextual Reasoning about Narrative Intent and ReceptionabstractJoel Mire, Maria Antoniak, Steven R Wilson, Zexin Ma, Achyutarama R Ganti, Andrew Piper, Maarten Sap. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Joel Mire, Maria Antoniak, Steven R. Wilson 0001, Zexin Ma, Achyutarama R. Ganti, Andrew Piper, Maarten Sap |
ACL (1) | 2 |
| 2025 | Research Borderlands: Analysing Writing Across Research CulturesabstractImproving cultural competence of language technologies is important.However most recent works rarely engage with the communities they study, and instead rely on synthetic setups and imperfect proxies of culture.In this work, we take a human-centered approach to discover and measure language-based cultural norms, and cultural competence of LLMs.We focus on a single kind of culture, research cultures, and a single task, adapting writing across research cultures.Through a set of interviews with interdisciplinary researchers, who are experts at moving between cultures, we create a framework of structural, stylistic, rhetorical, and citational norms that vary across research cultures.We operationalise these features with a suite of computational metrics and use them for (a) surfacing latent cultural norms in human-written research papers at scale; and (b) highlighting the lack of cultural competence of LLMs, and their tendency to homogenise writing.Overall, our work illustrates the efficacy of a humancentered approach to measuring cultural norms in human-written and LLM-generated texts. Shaily Bhatt, Tal August, Maria Antoniak |
ACL (1) | 3 |
| 2025 | CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-TeamingabstractYu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yu Ying Chiu, Bill Y. Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi 0001 |
ACL (1) | 8 |
| 2025 | Multi-Modal Framing Analysis of NewsabstractAutomated frame analysis of political communication is a popular task in computational social science that is used to study how authors select aspects of a topic to frame its reception.So far, such studies have been narrow, in that they use a fixed set of pre-defined frames and focus only on the text, ignoring the visual contexts in which those texts appear.Especially for framing in the news, this leaves out valuable information about editorial choices, which include not just the written article but also accompanying photographs.To overcome such limitations, we present a method for conducting multi-modal, multi-label framing analysis at scale using large (vision-) language models.Grounding our work in framing theory, we extract latent meaning embedded in images used to convey a certain point and contrast that to the text by comparing the respective frames used.We also identify highly partisan framing of topics with issue-specific frame analysis found in prior qualitative work.We demonstrate a method for doing scalable integrative framing analysis of both text and image in news, providing a more complete picture for understanding media bias. Arnav Arora, Srishti Yadav, Maria Antoniak, Serge J. Belongie, Isabelle Augenstein |
EMNLP | 3 |
| 2025 | so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMsabstractWhitespace is a critical component of poetic form, reflecting both adherence to standardized forms and rebellion against those forms.Each poem's whitespace distribution reflects the artistic choices of the poet and is an integral semantic and spatial feature of the poem.Yet, despite the popularity of poetry as both a longstanding art form and as a generation task for large language models (LLMs), whitespace has not received sufficient attention from the NLP community.Using a corpus of 19k Englishlanguage published poems from Poetry Foundation, we investigate how 4k poets have used whitespace in their works.We release a subset of 2.8k public-domain poems with preserved formatting to facilitate further research in this area.We compare whitespace usage in the published poems to (1) 51k LLM-generated poems, and (2) 12k unpublished poems posted in an online community.We also explore whitespace usage across time periods, poetic forms, and data sources.Additionally, we find that different text processing methods can result in significantly different representations of whitespace in poetry data, motivating us to use these poems and whitespace patterns to discuss implications for the processing strategies used to assemble pretraining datasets for LLMs. Sriharsh Bhyravajjula, Melanie Walsh, Anna Preus, Maria Antoniak |
EMNLP | 4 |
| 2025 | Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language ModelsabstractAbhilasha Ravichander, Jillian Fisher, Taylor Sorensen, Ximing Lu, Maria Antoniak, Bill Yuchen Lin, Niloofar Mireshghallah, Chandra Bhagavatula, Yejin Choi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Abhilasha Ravichander, Jillian Fisher, Taylor Sorensen, Ximing Lu, Maria Antoniak, Bill Y. Lin, Niloofar Mireshghallah, Chandra Bhagavatula, Yejin Choi 0001 |
NAACL (Long Papers) | 5 |
| 2024 | Where Do People Tell Stories Online? Story Detection Across Online CommunitiesabstractStory detection in online communities is a challenging task as stories are scattered across communities and interwoven with non-storytelling spans within a single text.We address this challenge by building and releasing the StorySeeker toolkit, including a richly annotated dataset of 502 Reddit posts and comments, a detailed codebook adapted to the social media context, and models to predict storytelling at the document and span levels.Our dataset is sampled from hundreds of popular Englishlanguage Reddit communities ranging across 33 topic categories, and it contains fine-grained expert annotations, including binary story labels, story spans, and event spans.We evaluate a range of detection methods using our data, and we identify the distinctive textual features of online storytelling, focusing on storytelling spans.We illuminate distributional characteristics of storytelling on a large communitycentric social media platform, and we also conduct a case study on r/ChangeMyView, where storytelling is used as one of many persuasive strategies, illustrating that our data and models can be used for both inter-and intra-community research.Finally, we discuss implications of our tools and analyses for narratology and the study of online communities. Maria Antoniak, Joel Mire, Maarten Sap, Elliott Ash, Andrew Piper |
ACL (1) | 1 |
| 2024 | The Empirical Variability of Narrative Perceptions of Social Media TextsabstractMost NLP work on narrative detection has focused on prescriptive definitions of stories crafted by researchers, leaving open the questions: how do crowd workers perceive texts to be a story, and why?We investigate this by building STORYPERCEPTIONS, a dataset of 2,496 perceptions of storytelling in 502 social media texts from 255 crowd workers, including categorical labels along with free-text storytelling rationales, authorial intent, and more.We construct a fine-grained bottom-up taxonomy of crowd workers' varied and nuanced perceptions of storytelling by open-coding their free-text rationales.Through comparative analyses at the label and code level, we illuminate patterns of disagreement among crowd workers and across other annotation contexts, including prescriptive labeling from researchers and LLM-based predictions.Notably, plot complexity, references to generalized or abstract actions, and holistic aesthetic judgments (such as a sense of cohesion) are especially important in disagreements.Our empirical findings broaden understanding of the types, relative importance, and contentiousness of features relevant to narrative detection, highlighting opportunities for future work on reader-contextualized models of narrative reception. Joel Mire, Maria Antoniak, Elliott Ash, Andrew Piper, Maarten Sap |
EMNLP | 2 |
| 2024 | Sensemaking about Contraceptive Methods across Online PlatformsabstractSelecting a birth control method is a complex healthcare decision. While birth control methods provide important benefits, they can also cause unpredictable side effects and be stigmatized, leading many people to seek additional information online, where they can privately find reviews, advice, hypotheses, and experiences of other birth control users. However, the relationships between their healthcare concerns, sensemaking activities, and online settings are not well understood. We gather texts about birth control shared on Twitter and Reddit—popular communities with different affordances, moderation, and audiences—to study where and how birth control is discussed online. Using a combination of topic modeling and hand annotation, we identify and characterize the dominant sensemaking practices across these platforms, and we create lexica to draw comparisons across birth control methods and side effects. We use these to measure variations from survey reports of side effect experiences, highlighting topics that social media users discuss more than expected online. Our findings characterize how online platforms are used to make sense of difficult healthcare choices, including analyzing risks, calculating timing and dosages, hypothesizing about causes of side effects, and storytelling about painful experiences. We contribute both to understanding unmet needs of birth control users and to exploring context-specific patterns in social media discussions. LeAnn McDowall, Maria Antoniak, David M. Mimno |
ICWSM | 2 |
| 2024 | Personalized Jargon Identification for Enhanced Interdisciplinary CommunicationabstractScientific jargon can confuse researchers when they read materials from other domains. Identifying and translating jargon for individual researchers could speed up research, but current methods of jargon identification mainly use corpus-level familiarity indicators rather than modeling researcher-specific needs, which can vary greatly based on each researcher's background. We collect a dataset of over 10K term familiarity annotations from 11 computer science researchers for terms drawn from 100 paper abstracts. Analysis of this data reveals that jargon familiarity and information needs vary widely across annotators, even within the same sub-domain (e.g., NLP). We investigate features representing domain, subdomain, and individual knowledge to predict individual jargon familiarity. We compare supervised and prompt-based approaches, finding that prompt-based methods using information about the individual researcher (e.g., personal publications, self-defined subfield of research) yield the highest accuracy, though the task remains difficult and supervised approaches have lower false positive rates. This research offers insights into features and methods for the novel task of integrating personal data into scientific jargon identification. Yue Guo 0007, Joseph Chee Chang, Maria Antoniak, Erin Bransom, Trevor Cohen, Lucy Lu Wang, Tal August |
NAACL-HLT | 3 |
| 2021 | Bad Seeds: Evaluating Lexical Methods for Bias MeasurementabstractMaria Antoniak, David Mimno. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Maria Antoniak, David M. Mimno |
ACL/IJCNLP (1) | 1 |
| 2021 | Tags, Borders, and Catalogs: Social Re-Working of Genre on LibraryThingabstractThrough a computational reading of the online book reviewing community LibraryThing, we examine the dynamics of a collaborative tagging system and learn how its users refine and redefine literary genres. LibraryThing tags are overlapping and multi-dimensional, created in a shared space by thousands of users, including readers, bookstore owners, and librarians. A common understanding of genre is that it relates to the content of books, but this resource allows us to view genre as an intersection of user communities and reader values and interests. We explore different methods of computational genre measurement within the open space of user-created tags. We measure overlap between books, tags, and users, and we also measure the homogeneity of communities associated with genre tags and correlate this homogeneity with reviewing behavior.Finally, by analyzing the text of reviews, we identify the thematic signatures of genres on LibraryThing, revealing similarities and differences between them. These measurements are intended to elucidate the genre conceptions of the users, not, as in prior work, to normalize the tags or enforce a hierarchy. We find that LibraryThing users make sense of genre through a variety of values and expectations, many of which fall outside common definitions and understandings of genre. Maria Antoniak, Melanie Walsh, David M. Mimno |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2019 | Narrative Paths and Negotiation of Power in Birth StoriesabstractBirth stories have become increasingly common on the internet, but they have received little attention as a computational dataset. These unsolicited, publicly posted stories provide rich descriptions of decisions, emotions, and relationships during a common but sometimes traumatic medical experience. These personal details can be illuminating for medical practitioners, and due to their shared structures, birth stories are also an ideal testing ground for narrative analysis techniques. We present an analysis of 2,847 birth stories from an online forum and demonstrate the utility of these stories for computational work. We discover clear sentiment, topic and persona-based patterns that both model the expected narrative event sequences of birth stories and highlight diverging pathways and exceptions to narrative norms. The authors' motivation to publicly post these personal stories can be a way to regain power after a surveilled and disempowering experience, and we explore power relationships between the personas in the stories, showing that these dynamics can vary with the type of birth (e.g., medicated vs unmedicated). Finally, birth stories exist in a space that is both public and deeply personal. This liminality poses a challenge for analysis and presentation, and we discuss tradeoffs and ethical practices for this collection. WARNING: This paper includes detailed narratives of pregnancy and birth. Maria Antoniak, David M. Mimno, Karen Levy |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2018 | Evaluating the Stability of Embedding-based Word SimilaritiesabstractWord embeddings are increasingly being used as a tool to study word associations in specific corpora. However, it is unclear whether such embeddings reflect enduring properties of language or if they are sensitive to inconsequential variations in the source documents. We find that nearest-neighbor distances are highly sensitive to small changes in the training corpus for a variety of algorithms. For all methods, including specific documents in the training set can result in substantial variations. We show that these effects are more prominent for smaller training corpora. We recommend that users never rely on single embedding models for distance calculations, but rather average over multiple bootstrap samples, especially for small corpora. Maria Antoniak, David M. Mimno |
Trans. Assoc. Comput. Linguistics | 1 |