Elisa Kreiss

dblp:212/3956 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
12since 2021 · last 2025
0009-0003-7574-6480ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
abstract
As LLMs are increasingly applied in socially impactful settings, concerns about gender bias have prompted growing efforts both to measure and mitigate such bias.These efforts often rely on evaluation tasks that differ from natural language distributions, as they typically involve carefully constructed task prompts that overtly or covertly signal the presence of gender bias-related content.In this paper, we examine how signaling the evaluative purpose of a task impacts measured gender bias in LLMs.Concretely, we test models under prompt conditions that (1) make the testing context salient, and (2) make gender-focused content salient.We then assess prompt sensitivity across four task formats with both token-probability and discrete-choice metrics.We find that prompts that more clearly align with (gender bias) evaluation framing elicit distinct gender output distributions compared to less evaluation-framed prompts.Discrete-choice metrics further tend to amplify bias relative to probabilistic measures.These findings do not only highlight the brittleness of LLM gender bias evaluations but open a new puzzle for the NLP benchmarking and development community: To what extent can well-controlled testing designs trigger LLM "testing mode" performance, and what does this mean for the ecological validity of future benchmarks.
Bufan Gao, Elisa Kreiss
EMNLP2
2025 MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations
abstract
We present a novel, open-source social network simulation framework MOSAIC where generative language agents predict user behaviors such as liking, sharing, and flagging content.This simulation combines LLM agents with a directed social graph to analyze emergent deception behaviors and gain a better understanding of how users determine the veracity of online social content.By constructing user representations from diverse fine-grained actual user personas, our system enables multi-agent simulations that model content dissemination and engagement dynamics at scale.Within this framework, we evaluate three different content moderation strategies with simulated misinformation dissemination, and we find that they not only mitigate the spread of non-factual content but also increase user engagement.In addition, we analyze the trajectories of popular content in our simulations, and explore whether simulation agents' articulated reasoning for their social interactions truly aligns with their collective engagement patterns.We opensource our simulation software to encourage further research within AI and social sciences:
Genglin Liu, Vivian T. Le, Salman Rahman, Elisa Kreiss, Marzyeh Ghassemi, Saadia Gabriel
EMNLP4
2024 Evaluating human and machine understanding of data visualizations
Arnav Verma, Kushin Mukherjee, Christopher Potts, Elisa Kreiss, Judith E. Fan
CogSci4
2024 CommVQA: Situating Visual Question Answering in Communicative Contexts
abstract
Current visual question answering (VQA) models tend to be trained and evaluated on imagequestion pairs in isolation.However, the questions people ask are dependent on their informational needs and prior knowledge about the image content.To evaluate how situating images within naturalistic contexts shapes visual questions, we introduce CommVQA, a VQA dataset consisting of images, image descriptions, real-world communicative scenarios where the image might appear (e.g., a travel website), and follow-up questions and answers conditioned on the scenario and description.CommVQA, which contains 1000 images and 8,949 question-answer pairs, poses a challenge for current models.Error analyses and a humansubjects study suggest that generated answers still contain high rates of hallucinations, fail to fittingly address unanswerable questions, and don't suitably reflect contextual information.Overall, we show that access to contextual information is essential for solving CommVQA, leading to the highest performing VQA model and highlighting the relevance of situating systems within communicative scenarios.
Nandita Naik, Christopher Potts, Elisa Kreiss
EMNLP3
2024 Updating CLIP to Prefer Descriptions Over Captions
abstract
Although CLIPScore is a powerful generic metric that captures the similarity between a text and an image, it fails to distinguish between a caption that is meant to complement the information in an image and a description that is meant to replace an image entirely, e.g., for accessibility.We address this shortcoming by updating the CLIP model with the Concadia dataset to assign higher scores to descriptions than captions using parameter efficient finetuning and a loss objective derived from work on causal interpretability.This model correlates with the judgements of blind and low-vision people while preserving transfer capabilities and has interpretable structure that sheds light on the caption-description distinction.1
Amir Zur, Elisa Kreiss, Karel D'Oosterlinck, Christopher Potts, Atticus Geiger
EMNLP2
2024 ContextRef: Evaluating Referenceless Metrics for Image Description Generation
abstract
Referenceless metrics (e.g., CLIPScore) use pretrained vision--language models to assess image descriptions directly without costly ground-truth reference texts. Such methods can facilitate rapid progress, but only if they truly align with human preference judgments. In this paper, we introduce ContextRef, a benchmark for assessing referenceless metrics for such alignment. ContextRef has two components: human ratings along a variety of established quality dimensions, and ten diverse robustness checks designed to uncover fundamental weaknesses. A crucial aspect of ContextRef is that images and descriptions are presented in context, reflecting prior work showing that context is important for description quality. Using ContextRef, we assess a variety of pretrained models, scoring functions, and techniques for incorporating context. None of the methods is successful with ContextRef, but we show that careful fine-tuning yields substantial improvements. ContextRef remains a challenging benchmark though, in large part due to the challenge of context dependence.
Elisa Kreiss, Eric Zelikman, Christopher Potts, Nick Haber
ICLR1
2023 A Semantics for Causing, Enabling, and Preventing Verbs Using Structural Causal Models
Angela Cao, Atticus Geiger, Elisa Kreiss, Thomas Icard, Tobias Gerstenberg
CogSci3
2022 Color Overmodification Emerges from Data-Driven Learning and Pragmatic Reasoning
Fei Fang 0005, Kunal Sinha, Noah D. Goodman, Christopher Potts, Elisa Kreiss
CogSci5
2022 Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation Metrics
abstract
Few images on the Web receive alt-text descriptions that would make them accessible to blind and low vision (BLV) users.Imagebased NLG systems have progressed to the point where they can begin to address this persistent societal problem, but these systems will not be fully successful unless we evaluate them on metrics that guide their development correctly.Here, we argue against current referenceless metrics -those that don't rely on human-generated ground-truth descriptions -on the grounds that they do not align with the needs of BLV users.The fundamental shortcoming of these metrics is that they do not take context into account, whereas contextual information is highly valued by BLV users.To substantiate these claims, we present a study with BLV participants who rated descriptions along a variety of dimensions.An in-depth analysis reveals that the lack of context-awareness makes current referenceless metrics inadequate for advancing image accessibility.As a proof-of-concept, we provide a contextual version of the referenceless metric CLIPScore which begins to address the disconnect to the BLV data.
Elisa Kreiss, Cynthia L. Bennett, Shayan Hooshmand, Eric Zelikman, Meredith Ringel Morris, Christopher Potts
EMNLP1
2022 Concadia: Towards Image-Based Text Generation with a Purpose
abstract
Current deep learning models often achieve excellent results on benchmark image-to-text datasets but fail to generate texts that are useful in practice.We argue that to close this gap, it is vital to distinguish descriptions from captions based on their distinct communicative roles.Descriptions focus on visual features and are meant to replace an image (often to increase accessibility), whereas captions appear alongside an image to supply additional information.To motivate this distinction and help people put it into practice, we introduce the publicly available Wikipedia-based dataset Concadia consisting of 96,918 images with corresponding English-language descriptions, captions, and surrounding context.Using insights from Concadia, models trained on it, and a preregistered human-subjects experiment with human-and model-generated texts, we characterize the commonalities and differences between descriptions and captions.In addition, we show that, for generating both descriptions and captions, it is useful to augment image-totext models with representations of the textual context in which the image appeared.split datapoints unique articles avg length (words) avg word length vocab size train 77,534 31,240 caption: 12.79
Elisa Kreiss, Fei Fang 0005, Noah D. Goodman, Christopher Potts
EMNLP1
2022 Inducing Causal Structure for Interpretable Neural Networks
abstract
In many areas, we have well-founded insights about causal structure that would be useful to bring into our trained models while still allowing them to learn in a data-driven fashion. To achieve this, we present the new method of interchange intervention training (IIT). In IIT, we (1) align variables in a causal model (e.g., a deterministic program or Bayesian network) with representations in a neural model and (2) train the neural model to match the counterfactual behavior of the causal model on a base input when aligned representations in both models are set to be the value they would be for a source input. IIT is fully differentiable, flexibly combines with other objectives, and guarantees that the target causal model is a causal abstraction of the neural model when its loss is zero. We evaluate IIT on a structural vision task (MNIST-PVR), a navigational language task (ReaSCAN), and a natural language inference task (MQNLI). We compare IIT against multi-task training objectives and data augmentation. In all our experiments, IIT achieves the best results and produces neural models that are more interpretable in the sense that they more successfully realize the target causal model.
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah D. Goodman, Christopher Potts
ICML5
2022 Causal Distillation for Language Models
abstract
Zhengxuan Wu, Atticus Geiger, Joshua Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, Noah Goodman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Zhengxuan Wu, Atticus Geiger, Josh Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, Noah D. Goodman
NAACL-HLT4
2020 Production expectations modulate contrastive inference
Elisa Kreiss, Judith Degen
CogSci1
2020 Modeling Subjective Assessments of Guilt in Newspaper Crime Narratives
abstract
Crime reporting is a prevalent form of journalism with the power to shape public perceptions and social policies.How does the language of these reports act on readers?We seek to address this question with the SuspectGuilt Corpus of annotated crime stories from Englishlanguage newspapers in the U.S. For Suspect-Guilt, annotators read short crime articles and provided text-level ratings concerning the guilt of the main suspect as well as span-level annotations indicating which parts of the story they felt most influenced their ratings.Sus-pectGuilt thus provides a rich picture of how linguistic choices affect subjective guilt judgments.We use SuspectGuilt to train and assess predictive models which validate the usefulness of the corpus, and show that these models benefit from genre pretraining and joint supervision from the text-level ratings and spanlevel annotations.Such models might be used as tools for understanding the societal effects of crime reporting.
Elisa Kreiss, Zijian Wang 0002, Christopher Potts
CoNLL1
2019 Uncertain evidence statements and guilt perception in iterative reproductions of crime stories
Elisa Kreiss, Michael Franke, Judith Degen
CogSci1
2017 Mentioning atypical properties of objects is communicatively efficient
Elisa Kreiss, Robert D. Hawkins, Judith Degen, Noah D. Goodman
CogSci1