Eve Fleisig

dblp:276/0223 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2025
0009-0004-6831-7319ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021
YearPublicationVenuePosition
2025 GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
abstract
Language models are often miscalibrated, leading to confidently incorrect answers. We introduce GRACE, a benchmark for language model calibration that incorporates comparison with human calibration. GRACE consists of question-answer pairs, in which each question contains a series of clues that gradually become easier, all leading to the same answer; models must answer correctly as early as possible as the clues are revealed. This setting permits granular measurement of model calibration based on how early, accurately, and confidently a model answers. After collecting these questions, we host live human vs. model competitions to gather 1,749 data points on human and model teams’ timing, accuracy, and confidence. We propose a metric, CalScore, that uses GRACE to analyze model calibration errors and identify types of model miscalibration that differ from human behavior. We find that although humans are less accurate than models, humans are generally better calibrated. Since state-of-the-art models struggle on GRACE, it effectively evaluates progress on improving model calibration.
Yoo Yeon Sung, Eve Fleisig, Ishan Upadhyay, Jordan L. Boyd-Graber
ACL (1)2
2025 Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
abstract
Yoo Yeon Sung, Maharshi Gor, Eve Fleisig, Ishani Mondal, Jordan Lee Boyd-Graber. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yoo Yeon Sung, Maharshi Gor, Eve Fleisig, Ishani Mondal, Jordan L. Boyd-Graber
NAACL (Long Papers)3
2024 Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
abstract
We present a large-scale study of linguistic bias exhibited by ChatGPT covering ten dialects of English (Standard American English, Standard British English, and eight widely spoken non-"standard" varieties from around the world).We prompted GPT-3.5 Turbo and GPT-4 with text by native speakers of each variety and analyzed the responses via detailed linguistic feature annotation and native speaker evaluation.We find that the models default to "standard" varieties of English; based on evaluation by native speakers, we also find that model responses to non-"standard" varieties consistently exhibit a range of issues: stereotyping (19% worse than for "standard" varieties), demeaning content (25% worse), lack of comprehension (9% worse), and condescending responses (15% worse).Moreover, if these models are asked to imitate the writing style of prompts in non-"standard" varieties, they produce text that exhibits lower comprehension of the input and is especially prone to stereotyping.GPT-4 improves on GPT-3.5 in terms of comprehension, warmth, and friendliness, but also exhibits a marked increase in stereotyping (+18%).The results indicate that GPT-3.5 Turbo and GPT-4 can perpetuate linguistic discrimination toward speakers of non-"standard" varieties.
Eve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi, Xavier Yin, Daniel Klein 0001
EMNLP1
2024 Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
abstract
When annotators disagree, predicting the labels given by individual annotators can capture nuances overlooked by traditional label aggregation.We introduce three approaches to predict individual annotator ratings on the toxicity of text by incorporating individual annotator-specific information: a neural collaborative filtering (NCF) approach, an in-context learning (ICL) approach, and an intermediate embedding-based architecture.We also study the utility of demographic information for rating prediction.NCF showed limited utility; however, integrating annotator history, demographics, and survey information permits both the embedding-based architecture and ICL to substantially improve prediction accuracy, with the embedding-based architecture outperforming the other methods.We also find that, if demographics are predicted from survey information, using these imputed demographics as features performs comparably to using true demographic data.This suggests that demographics may not provide substantial information for modeling ratings beyond what is captured in survey responses.Our findings raise considerations about the relative utility of different types of annotator information and provide new approaches for modeling annotators in subjective NLP tasks.
Harbani Jaggi, Kashyap Coimbatore Murali, Eve Fleisig, Erdem Biyik
EMNLP3
2024 The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
abstract
Eve Fleisig, Su Lin Blodgett, Dan Klein, Zeerak Talat. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Eve Fleisig, Su Lin Blodgett, Daniel Klein 0001, Zeerak Talat
NAACL-HLT1
2024 First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models
abstract
Naomi Saphra, Eve Fleisig, Kyunghyun Cho, Adam Lopez. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Naomi Saphra, Eve Fleisig, Kyunghyun Cho, Adam Lopez
NAACL-HLT2
2024 Ghostbuster: Detecting Text Ghostwritten by Large Language Models
abstract
Vivek Verma, Eve Fleisig, Nicholas Tomlin, Dan Klein. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Eve Fleisig, Nicholas Tomlin, Daniel Klein 0001
NAACL-HLT2
2023 FairPrism: Evaluating Fairness-Related Harms in Text Generation
abstract
Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, Hanna Wallach. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, Hanna M. Wallach
ACL (1)1
2023 When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
abstract
Though majority vote among annotators is typically used for ground truth labels in machine learning, annotator disagreement in tasks such as hate speech detection may reflect systematic differences in opinion across groups, not noise.Thus, a crucial problem in hate speech detection is determining if a statement is offensive to the demographic group that it targets, when that group may be a small fraction of the annotator pool.We construct a model that predicts individual annotator ratings on potentially offensive text and combines this information with the predicted target group of the text to predict the ratings of target group members.We show gains across a range of metrics, including raising performance over the baseline by 22% at predicting individual annotators' ratings and by 33% at predicting variance among annotators, which provides a metric for model uncertainty downstream.We find that annotators' ratings can be predicted using their demographic information as well as opinions on online content, and that non-invasive questions on annotators' online experiences minimize the need to collect demographic information when predicting annotators' opinions.
Eve Fleisig, Rediet Abebe, Daniel Klein 0001
EMNLP1
2023 Incorporating Worker Perspectives into MTurk Annotation Practices for NLP
abstract
Current practices regarding data collection for natural language processing on Amazon Mechanical Turk (MTurk) often rely on a combination of studies on data quality and heuristics shared among NLP researchers.However, without considering the perspectives of MTurk workers, these approaches are susceptible to issues regarding workers' rights and poor response quality.We conducted a critical literature review and a survey of MTurk workers aimed at addressing open questions regarding best practices for fair payment, worker privacy, data quality, and considering worker incentives.We found that worker preferences are often at odds with received wisdom among NLP researchers.Surveyed workers preferred reliable, reasonable payments over uncertain, very high payments; reported frequently lying on demographic questions; and expressed frustration at having work rejected with no explanation.We also found that workers view some quality control methods, such as requiring minimum response times or Master's qualifications, as biased and largely ineffective.Based on the survey results, we provide recommendations on how future NLP studies may better account for MTurk workers' experiences in order to respect workers' rights and improve data quality.
Olivia Huang, Eve Fleisig, Daniel Klein 0001
EMNLP2
2023 Centering the Margins: Outlier-Based Identification of Harmed Populations in Toxicity Detection
abstract
The impact of AI models on marginalized communities has traditionally been measured by identifying performance differences between specified demographic subgroups.Though this approach aims to center vulnerable groups, it risks obscuring patterns of harm faced by intersectional subgroups or shared across multiple groups.To address this, we draw on theories of marginalization from disability studies and related disciplines, which state that people farther from the norm face greater adversity, to consider the "margins" in the domain of toxicity detection.We operationalize the "margins" of a dataset by employing outlier detection to identify text about people with demographic attributes distant from the "norm".We find that model performance is consistently worse for demographic outliers, with mean squared error (MSE) between outliers and non-outliers up to 70.4% worse across toxicity types.It is also worse for text outliers, with a MSE up to 68.4% higher for outliers than non-outliers.We also find text and demographic outliers to be particularly susceptible to errors in the classification of severe toxicity and identity attacks.Compared to analysis of disparities using traditional demographic breakdowns, we find that our outlier analysis frequently surfaces greater harms faced by a larger, more intersectional group, which suggests that outlier analysis is particularly beneficial for identifying harms against those groups.* Eve and Vyoma co-created the theoretical framework for this paper, and Vyoma implemented it.
Vyoma Raman, Eve Fleisig, Daniel Klein 0001
EMNLP2