VLDB 2026 Research / reviewers in the wild / expert
Su Lin Blodgett
dblp:182/2034
· DBLP profile ↗
21ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-9861-3483ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text GenerationabstractKatelyn X. Mei, Yi-Li Hsu, Minjoon Choi, Zongwan Cao, Chenjun Xu, Bingbing Wen, Su Lin Blodgett, Lucy Lu Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Katelyn Mei, Yi-Li Hsu, Minjoon Choi, Zongwan Cao, Chenjun Xu, Bingbing Wen, Su Lin Blodgett, Lucy Lu Wang |
ACL (1) | 7 |
| 2026 | A Framework to Characterize Reporting on Generative AI UseabstractUnlike with traditional predictive AI models, today’s generative AI models are increasingly designed to be general-purpose, able to perform a wide range of tasks. This makes it challenging to develop a reliable and useful understanding of the ways in which this technology is and could be used. As a result, academic and policy researchers and generative AI providers have started to publish the results of their own investigations about the use of generative AI. This information is, however, fragmented, potentially incomplete, sometimes ambiguous, and often lacking in methodological specificity. In this paper, we conducted an integrative review to build a multi-dimensional framework that specifies what kind of information about generative AI use could be reported and how, and illustrated its analytical utility by applying the framework to a collection of over 110 industry documents. Our analysis reveals systematic patterns and omissions in current industry reporting and reflects on the narratives this reporting collectively advance about generative AI use. Agathe Balayn, Varun Nagaraj Rao, Su Lin Blodgett, Aylin Caliskan, Solon Barocas |
CHI | 3 |
| 2026 | From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing AssistantsabstractAI-based writing assistants are ubiquitous, yet little is known about how users’ mental models shape their use. We examine two types of mental models—functional or related to what the system does, and structural or related to how the system works—and how they affect control behavior—how users request, accept, or edit AI suggestions as they write—and writing outcomes. We primed participants (N = 48) with different system descriptions to induce these mental models before asking them to complete a cover letter writing task using a writing assistant that occasionally offered preconfigured ungrammatical suggestions to test whether the mental models affected participants’ critical oversight. We find that while participants in the structural mental model condition demonstrate a better understanding of the system, this can have a backfiring effect: while these participants judged the system as more usable, they also produced letters with more grammatical errors, highlighting a complex relationship between system understanding, trust, and control in contexts that require user oversight of error-prone AI outputs. Shalaleh Rismani, Su Lin Blodgett, Qingzi Vera Liao, Alexandra Olteanu, AJung Moon |
CHI | 2 |
| 2025 | Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation SystemsabstractAs text generation systems' outputs are increasingly anthropomorphic-perceived as humanlike-scholars have also increasingly raised concerns about how such outputs can lead to harmful outcomes, such as users over-relying or developing emotional dependence on these systems.How to intervene on such system outputs to mitigate anthropomorphic behaviors and their attendant harmful outcomes, however, remains understudied.With this work, we aim to provide empirical and theoretical grounding for developing such interventions.To do so, we compile an inventory of interventions grounded both in prior literature and a crowdsourcing study where participants edited system outputs to make them less human-like.Drawing on this inventory, we also develop a conceptual framework to help characterize the landscape of possible interventions, articulate distinctions between different types of interventions, and provide a theoretical basis for evaluating the effectiveness of different interventions. Myra Cheng, Su Lin Blodgett, Alicia DeVrio, Lisa Egede, Alexandra Olteanu |
ACL (1) | 2 |
| 2025 | A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language TechnologiesabstractRecent attention to anthropomorphism -- the attribution of human-like qualities to non-human objects or entities -- of language technologies like LLMs has sparked renewed discussions about potential negative impacts of anthropomorphism. To productively discuss the impacts of this anthropomorphism and in what contexts it is appropriate, we need a shared vocabulary for the vast variety of ways that language can be anthropomorphic. In this work, we draw on existing literature and analyze empirical cases of user interactions with language technologies to develop a taxonomy of textual expressions that can contribute to anthropomorphism. We highlight challenges and tensions involved in understanding linguistic anthropomorphism, such as how all language is fundamentally human and how efforts to characterize and shift perceptions of humanness in machines can also dehumanize certain humans. We discuss ways that our taxonomy supports more precise and effective discussions of and decisions about anthropomorphism of language technologies. Alicia DeVrio, Myra Cheng, Lisa Egede, Alexandra Olteanu, Su Lin Blodgett |
CHI | 5 |
| 2025 | Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of RigorabstractIn AI research and practice, rigor remains largely understood in terms of methodological rigor---such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI community, including overblown claims about the capabilities of AI systems. Our position is that a broader conception of what rigorous AI research and practice should entail is needed. We believe such a conception---in addition to a more expansive understanding of 1) methodological rigor---should include aspects related to 2) what background knowledge informs what to work on (epistemic rigor); 3) how disciplinary, community, or personal norms, standards, or beliefs influence the work (normative rigor); 4) how clearly articulated the theoretical constructs under use are (conceptual rigor); 5) what is reported and how (reporting rigor); and 6) how well-supported the inferences from existing evidence are (interpretative rigor). In doing so, we also provide useful language and a framework for much needed dialogue about the AI community's work by researchers, policymakers, journalists, and other stakeholders. Alexandra Olteanu, Su Lin Blodgett, Agathe Balayn, Angelina Wang, Fernando Diaz 0001, Flávio P. Calmon, Margaret Mitchell, Michael D. Ekstrand, Reuben Binns, Solon Barocas |
NeurIPS | 2 |
| 2025 | 'It was 80% me, 20% AI': Seeking Authenticity in Co-Writing with Large Language ModelsabstractGiven the rising proliferation and diversity of AI writing assistance tools, especially those powered by large language models (LLMs), both writers and readers may have concerns about the impact of these tools on the authenticity of writing work. We examine whether and how writers want to preserve their authentic voice when co-writing with AI tools and whether personalization of AI writing support could help achieve this goal. We conducted semi-structured interviews with 19 professional writers, during which they co-wrote with both personalized and non-personalized AI writing-support tools. We supplemented writers' perspectives with opinions from 30 avid readers about the written work co-produced with AI collected through an online survey. Our findings illuminate conceptions of authenticity in human-AI co-creation, which focus more on the process and experience of constructing creators' authentic selves. While writers reacted positively to personalized AI writing tools, they believed the form of personalization needs to target writers' growth and go beyond the phase of text production. Overall, readers' responses showed less concern about human-AI co-writing. Readers could not distinguish AI-assisted work, personalized or not, from writers' solo-written work and showed positive attitudes toward writers experimenting with new technology for creative writing. Angel Hwang, Qingzi Vera Liao, Su Lin Blodgett, Alexandra Olteanu, Adam Trischler |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | ECBD: Evidence-Centered Benchmark Design for NLPabstractYu Lu Liu, Su Lin Blodgett, Jackie Cheung, Q. Vera Liao, Alexandra Olteanu, Ziang Xiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung, Qingzi Vera Liao, Alexandra Olteanu, Ziang Xiao |
ACL (1) | 2 |
| 2024 | Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPabstractThis paper introduces the concept of actionability in the context of bias measures in natural language processing (NLP).We define actionability as the degree to which a measurement's results enable informed action and propose a set of desiderata for assessing it.Building on existing frameworks such as measurement modeling, we argue that actionability is a crucial aspect of bias measures that has been largely overlooked in the literature.We conduct a comprehensive review of 146 papers proposing bias measures in NLP, examining whether and how they provide the information required for actionable results.Our findings reveal that many key elements of actionability, including a measure's intended use and reliability assessment, are often unclear or absent.This study highlights a significant gap in the current approach to developing and reporting bias measures in NLP.We argue that this lack of clarity may impede the effective implementation and utilization of these measures.To address this issue, we offer recommendations for more comprehensive and actionable metric development and reporting practices in NLP bias research. Pieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett, Zeerak Talat |
EMNLP | 4 |
| 2024 | The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human LabelsabstractEve Fleisig, Su Lin Blodgett, Dan Klein, Zeerak Talat. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Eve Fleisig, Su Lin Blodgett, Daniel Klein 0001, Zeerak Talat |
NAACL-HLT | 2 |
| 2024 | "One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System BehaviorsabstractLi Lucy, Su Lin Blodgett, Milad Shokouhi, Hanna Wallach, Alexandra Olteanu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Li Lucy, Su Lin Blodgett, Milad Shokouhi, Hanna M. Wallach, Alexandra Olteanu |
NAACL-HLT | 2 |
| 2023 | Taxonomizing and Measuring Representational Harms: A Look at Image TaggingabstractIn this paper, we examine computational approaches for measuring the "fairness" of image tagging systems, finding that they cluster into five distinct categories, each with its own analytic foundation. We also identify a range of normative concerns that are often collapsed under the terms "unfairness," "bias," or even "discrimination" when discussing problematic cases of image tagging. Specifically, we identify four types of representational harms that can be caused by image tagging systems, providing concrete examples of each. We then consider how different computational measurement approaches map to each of these types, demonstrating that there is not a one-to-one mapping. Our findings emphasize that no single measurement approach will be definitive and that it is not possible to infer from the use of a particular measurement approach which type of harm was intended to be measured. Lastly, equipped with this more granular understanding of the types of representational harms that can be caused by image tagging systems, we show that attempts to mitigate some of these types of harms may be in tension with one another. Jared Katzman, Angelina Wang, Morgan Klaus Scheuerman, Su Lin Blodgett, Kristen Laird, Hanna M. Wallach, Solon Barocas |
AAAI | 4 |
| 2023 | FairPrism: Evaluating Fairness-Related Harms in Text GenerationabstractEve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, Hanna Wallach. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, Hanna M. Wallach |
ACL (1) | 4 |
| 2022 | Examining Responsibility and Deliberation in AI Impact Statements and Ethics ReviewsabstractThe artificial intelligence research community is continuing to grapple with the ethics of its work by encouraging researchers to discuss potential positive and negative consequences. Neural Information Processing Systems (NeurIPS), a top-tier conference for machine learning and artificial intelligence research, first required a statement of broader impact in 2020. In 2021, NeurIPS updated their call for papers such that 1) the impact statement focused on negative societal impacts and was not required but encouraged, 2) a paper checklist and ethics guidelines were provided to authors, and 3) papers underwent ethics reviews and could be rejected on ethical grounds. In light of these changes, we contribute a qualitative analysis of 231 impact statements and all publicly-available ethics reviews. We describe themes arising around the ways in which authors express agency (or lack thereof) in identifying or mitigating negative consequences and assign responsibility for mitigating negative societal impacts. We also characterize ethics reviews in terms of the types of issues raised by ethics reviewers (falling into categories of policy-oriented and non-policy-oriented), recommendations ethics reviewers make to authors (e.g., in terms of adding or removing content), and interaction between authors, ethics reviewers, and original reviewers (e.g., consistency between issues flagged by original reviewers and those discussed by ethics reviewers). Finally, based on our analysis we make recommendations for how authors can be further supported in engaging with the ethical implications of their work. David Liu 0006, Priyanka Nanayakkara, Sarah Ariyan Sakha, Grace Abuhamad, Su Lin Blodgett, Nicholas Diakopoulos, Jessica Hullman, Tina Eliassi-Rad |
AIES | 5 |
| 2022 | Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their ImplicationsabstractKaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, Alexandra Olteanu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, Alexandra Olteanu |
NAACL-HLT | 2 |
| 2021 | Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark DatasetsabstractSu Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, Hanna Wallach. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, Hanna M. Wallach |
ACL/IJCNLP (1) | 1 |
| 2021 | A Survey of Race, Racism, and Anti-Racism in NLPabstractAnjalie Field, Su Lin Blodgett, Zeerak Waseem, Yulia Tsvetkov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Anjalie Field, Su Lin Blodgett, Zeerak Talat, Yulia Tsvetkov |
ACL/IJCNLP (1) | 2 |
| 2020 | Language (Technology) is Power: A Critical Survey of "Bias" in NLPabstractWe survey 146 papers analyzing "bias" in NLP systems, fnding that their motivations are often vague, inconsistent, and lacking in normative reasoning, despite the fact that analyzing "bias" is an inherently normative process.We further fnd that these papers' proposed quantitative techniques for measuring or mitigating "bias" are poorly matched to their motivations and do not engage with the relevant literature outside of NLP.Based on these fndings, we describe the beginnings of a path forward by proposing three recommendations that should guide work analyzing "bias" in NLP systems.These recommendations rest on a greater recognition of the relationships between language and social hierarchies, encouraging researchers and practitioners to articulate their conceptualizations of "bias"-i.e., what kinds of system behaviors are harmful, in what ways, to whom, and why, as well as the normative reasoning underlying these statements-and to center work around the lived experiences of members of communities affected by NLP systems, while interrogating and reimagining the power relations between technologists and such communities.NLP task Papers Embeddings (type-level or contextualized) 54 Coreference resolution 20 Language modeling or dialogue generation 17 Hate-speech detection 17 Sentiment analysis 15 Machine translation 8 Tagging or parsing 5 Surveys, frameworks, and meta-analyses 20 Other 22 Su Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. Wallach |
ACL | 1 |
| 2018 | Twitter Universal Dependency Parsing for African-American and Mainstream American EnglishabstractDue to the presence of both Twitterspecific conventions and non-standard and dialectal language, Twitter presents a significant parsing challenge to current dependency parsing tools.We broaden English dependency parsing to handle social media English, particularly social media African-American English (AAE), by developing and annotating a new dataset of 500 tweets, 250 of which are in AAE, within the Universal Dependencies 2.0 framework.We describe our standards for handling Twitter-and AAE-specific features and evaluate a variety of crossdomain strategies for improving parsing with no, or very little, in-domain labeled data, including a new data synthesis approach.We analyze these methods' impact on performance disparities between AAE and Mainstream American English tweets, and assess parsing accuracy for specific AAE lexical and syntactic features.Our Su Lin Blodgett, Johnny Wei, Brendan T. O'Connor 0001 |
ACL (1) | 1 |
| 2018 | Monte Carlo Syntax Marginals for Exploring and Using Dependency ParsesabstractKatherine Keith, Su Lin Blodgett, Brendan O’Connor. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Katherine A. Keith, Su Lin Blodgett, Brendan T. O'Connor 0001 |
NAACL-HLT | 2 |
| 2016 | Demographic Dialectal Variation in Social Media: A Case Study of African-American EnglishabstractThough dialectal language is increasingly abundant on social media, few resources exist for developing NLP tools to handle such language.We conduct a case study of dialectal language in online conversational text by investigating African-American English (AAE) on Twitter.We propose a distantly supervised model to identify AAE-like language from demographics associated with geo-located messages, and we verify that this language follows well-known AAE linguistic phenomena.In addition, we analyze the quality of existing language identification and dependency parsing tools on AAE-like text, demonstrating that they perform poorly on such text compared to text associated with white speakers.We also provide an ensemble classifier for language identification which eliminates this disparity and release a new corpus of tweets containing AAE-like language. Su Lin Blodgett, Lisa Green, Brendan T. O'Connor 0001 |
EMNLP | 1 |