Maarten Sap

dblp:153/9519 · DBLP profile ↗
← Back
64ranked-venue papers
9as first author
50since 2021 · last 2026
0000-0002-0701-4654ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 61 · 9 first-author · 47 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Social Story Frames: Contextual Reasoning about Narrative Intent and Reception
abstract
Joel Mire, Maria Antoniak, Steven R Wilson, Zexin Ma, Achyutarama R Ganti, Andrew Piper, Maarten Sap. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Joel Mire, Maria Antoniak, Steven R. Wilson 0001, Zexin Ma, Achyutarama R. Ganti, Andrew Piper, Maarten Sap
ACL (1)7
2026 Black LLMirror: User (Self) Perceptions in Black American English Interactions with LLMs
abstract
LLMs becoming increasingly personalized to users’ language style raises both excitement and concerns for minority users such as Black American English (BAE) speakers. Yet, previous work has predominantly focused on user perceptions of out-of-context BAE statements by LLMs rather than naturalistic multi-turn interactions, and has ignored such systems’ effects on users’ self-perception. In this work, we examine the effects that multi-turn interactions with speech and text BAE-producing LLMs have on BAE speakers’ perceptions of the LLM and of themselves. We observe a significant change in participant self-esteem following the interactions, and notable qualitative differences between BAE-LLM and Standard American English (SAE) LLM interactions. We also observe significant effects of BAE-usage on user perception of the model within speech-based interactions. Our findings suggest that the effects of BAE-usage by an LLM agent on model- and self-perception among BAE-speaking users are complex and widely varied.
Mikayla Campbell, Joel Mire, Mark Diaz, Maarten Sap
CHI4
2025 BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
abstract
In this work, we tackle the challenge of embedding realistic human personality traits into LLMs.Previous approaches have primarily focused on prompt-based methods that describe the behavior associated with the desired personality traits, suffering from realism and validity issues.To address these limitations, we introduce BIG5-CHAT, a large-scale dataset containing 100,000 dialogues designed to ground models in how humans express their personality in language.Leveraging this dataset, we explore Supervised Fine-Tuning and Direct Preference Optimization as training-based methods to align LLMs more naturally with human personality patterns.Our methods outperform prompting on personality assessments such as BFI and IPIP-NEO, with trait correlations more closely matching human data.Furthermore, our experiments reveal that models trained to exhibit higher conscientiousness, higher agreeableness, lower extraversion, and lower neuroticism display better performance on reasoning tasks, aligning with psychological findings on how these traits impact human cognitive performance.To our knowledge, this work is the first comprehensive study to demonstrate how training-based methods can shape LLM personalities through learning from real human behaviors. Whenever I lay on my bed I get so tired.
Jiarui Liu 0004, Andy Liu, Mona T. Diab, Maarten Sap
ACL (1)6
2025 Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
abstract
Gestures are an integral part of non-verbal communication, with meanings that vary across cultures, and misinterpretations that can have serious social and diplomatic consequences.As AI systems become more integrated into global applications, ensuring they do not inadvertently perpetuate cultural offenses is critical.To this end, we introduce Multi-Cultural Set of Inappropriate Gestures and Nonverbal Signs (MC-SIGNS), a dataset of 288 gesture-country pairs annotated for offensiveness, cultural significance, and contextual factors across 25 gestures and 85 countries.Through systematic evaluation using MC-SIGNS, we uncover critical limitations: text-to-image (T2I) systems exhibit strong US-centric biases, performing better at detecting offensive gestures in US contexts than in non-US ones; large language models (LLMs) tend to over-flag gestures as offensive; and vision-language models (VLMs) default to US-based interpretations when responding to universal concepts like wishing someone luck, frequently suggesting culturally inappropriate gestures.These findings highlight the urgent need for culturally-aware AI safety mechanisms to ensure equitable global deployment of AI technologies.
Akhila Yerukola, Saadia Gabriel, Nanyun Peng 0001, Maarten Sap
ACL (1)4
2025 User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
abstract
Peer Reviewed
Xianzhe Fan, Qing Xiao 0002, Jiaxin Pei, Maarten Sap, Zhicong Lu, Hong Shen 0004
CHI5
2025 AutoPresent: Designing Structured Visuals from Scratch
abstract
Designing structured visuals such as presentation slides is essential for communicative needs, necessitating both content creation and visual planning skills. In this work, we tackle the challenge of automated slide generation, where models produce slide presentations from natural language (NL) instructions. We first introduce the SlidesBench benchmark, the first benchmark for slide generation with 7k training and 585 testing examples derived from 310 slide decks across 10 domains. SlidesBench supports evaluations that are (i) reference-based to measure similarity to a target slide, and (ii) reference-free to measure the design quality of generated slides alone. We benchmark end-to-end image generation and program generation methods with a variety of models, and find that programmatic methods produce higher-quality slides in user-interactable formats. Built on the success of program generation, we create AutoPresent, an 8B LlaMa-based model trained on 7k pairs of instructions paired with code for slide generation, and achieve results comparable to the closed-source model GPT-4O. We further explore iterative design refinement where the model is tasked to self-refine its own output, and we found that this process improves the slide’s quality. We hope that our work will provide a basis for future work on generating structured visuals. Our code, data, demo, and video demonstrations are publicly available at https: //github.com/para-lost/AutoPresent
Jiaxin Ge, Zhiruo Wang 0001, Yi-Hao Peng, Sanjay Subramanian, Qinyue Tan, Maarten Sap, Alane Suhr, Daniel Fried, Graham Neubig, Trevor Darrell
CVPR7
2025 SOCIAL SCAFFOLDS: A Generalization Framework for Social Understanding Tasks
abstract
Effective human communication in social settings is contingent on recognizing subtle cues, such as intentions or implications.Without such cues, NLP models risk missing social signals, instead relying on surface patterns.We introduce SOCIAL SCAFFOLDS, an automated framework for facilitating generalization across social reasoning tasks by generating rationales that make these social cues explicit.Grounded in narrative modeling principles, we generate task-agnostic rationales that capture different perspectives, i.e., that of the speaker, the listener, and the general world-view.Our experimental suite showcases that providing rationales as augmentations aids task performance for both supervised fine-tuning and incontext learning paradigms.Notably, providing all three rationale types significantly improves cross-task performance in 44% of cases, and inferred speaker intent in 31.3% of cases.We conduct statistical and ablation analyses that show how rationales complement the input text and are used effectively by models.
Ritam Dutt, Carolyn P. Rosé, Maarten Sap
EMNLP3
2025 Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
abstract
Jiarui Liu, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona T. Diab, Maarten Sap. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jiarui Liu 0004, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona T. Diab, Maarten Sap
EMNLP8
2025 Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
abstract
Conversational breakdowns in close relationships are deeply shaped by personal histories and emotional context, yet most NLP research treats conflict detection as a general task, overlooking the relational dynamics that influence how messages are perceived.In this work, we leverage nonviolent communication (NVC) theory to evaluate LLMs in detecting conversational breakdowns and assessing how relationship backstory influences both human and model perception of conflicts.Given the sensitivity and scarcity of real-world datasets featuring conflict between familiar social partners with rich personal backstories, we contribute the PERSONACONFLICTS CORPUS 1 , a dataset of N = 5, 772 naturalistic simulated dialogues spanning diverse conflict scenarios between friends, family members, and romantic partners.Through a controlled human study, we annotate a subset of dialogues and obtain finegrained labels of communication breakdown types on individual turns, and assess the impact of backstory on human and model perception of conflict in conversation.We find that the polarity of relationship backstories significantly shifted human perception of communication breakdowns and impressions of the social partners, yet models struggle to meaningfully leverage those backstories in the detection task.Additionally, we find that models consistently overestimate how positively a message will make a listener feel.Our findings underscore the critical role of personalization to relationship contexts in enabling LLMs to serve as effective mediators in human communication for authentic connection.
Jocelyn Shen, Akhila Yerukola, Cynthia Breazeal, Maarten Sap, Hae Won Park 0001
EMNLP5
2025 On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
abstract
Large language model-based multi-agent systems have shown great abilities across various tasks due to the collaboration of expert agents, each focusing on a specific domain. However, the impact of clumsy or even malicious agents—those who frequently make errors in their tasks—on the overall performance of the system remains underexplored. This paper investigates: (1) What is the resilience of various system structures (e.g., A$\rightarrow$B$\rightarrow$C, A$\leftrightarrow$B$\leftrightarrow$C) under faulty agents, on different downstream tasks? (2) How can we increase system resilience to defend against these agents? To simulate faulty agents, we propose two approaches—AutoTransform and AutoInject—which introduce mistakes into the agents’ responses. Experiments on four downstream tasks using six systems show that the "hierarchical" structure, i.e., A$\rightarrow$(B$\leftrightarrow$C), exhibits superior resilience with the lowest performance drop of 5.5%, compared to 10.5% and 23.7% of other two structures. To further improve resilience, we introduce (1) Challenger, that introduces a mechanism for each agent to challenge others’ outputs, and (2) Inspector, an additional agent to review and correct messages, recovering up to 96.4% errors made by faulty agents. Our code and data are available at https://github.com/CUHK-ARISE/MAS-Resilience.
Jen-tse Huang 0001, Jiaxu Zhou, Tailin Jin, Wenxuan Wang 0001, Youliang Yuan, Michael R. Lyu, Maarten Sap
ICML9
2025 SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
abstract
The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect a community’s values), which current systems fall short on. To address this gap, we present SafetyAnalyst, a novel AI safety moderation framework. Given an AI behavior, SafetyAnalyst uses chain-of-thought reasoning to analyze its potential consequences by creating a structured "harm-benefit tree," which enumerates harmful and beneficial actions and effects the AI behavior may lead to, along with likelihood, severity, and immediacy labels that describe potential impacts on stakeholders. SafetyAnalyst then aggregates all effects into a harmfulness score using 28 fully interpretable weight parameters, which can be aligned to particular safety preferences. We applied this framework to develop an open-source LLM prompt safety classification system, distilled from 18.5 million harm-benefit features generated by frontier LLMs on 19k prompts. On comprehensive benchmarks, we show that SafetyAnalyst (average F1=0.81) outperforms existing moderation systems (average F1$<$0.72) on prompt safety classification, while offering the additional advantages of interpretability, transparency, and steerability.
Valentina Pyatkin, Max Kleiman-Weiner, Nouha Dziri, Anne Gabrielle Eva Collins, Jana Schaich Borg, Maarten Sap, Yejin Choi 0001, Sydney Levine
ICML8
2025 NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
abstract
Abhinav Sukumar Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, Maarten Sap. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, Maarten Sap
NAACL (Long Papers)5
2025 AI-LieDar : Examine the Trade-off Between Utility and Truthfulness in LLM Agents
abstract
Zhe Su, Xuhui Zhou, Sanketh Rangreji, Anubha Kabra, Julia Mendelsohn, Faeze Brahman, Maarten Sap. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sanketh Rangreji, Anubha Kabra, Julia Mendelsohn, Faeze Brahman, Maarten Sap
NAACL (Long Papers)7
2025 REL-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
abstract
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, Nouha Dziri, Dan Jurafsky, Maarten Sap. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren 0001, Nouha Dziri, Daniel Jurafsky, Maarten Sap
NAACL (Long Papers)6
2025 SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
abstract
Humans continuously infer the states, goals, and behaviors of others by perceiving their surroundings in dynamic, real-world social interactions. However, most Theory of Mind (ToM) benchmarks only evaluate static, text-based scenarios, which have a significant gap compared to real interactions. We propose the SoMi-ToM benchmark, designed to evaluate multi-perspective ToM in embodied multi-agent complex social interactions. This benchmark is based on rich multimodal interaction data generated by the interaction environment SoMi, covering diverse crafting goals and social relationships. Our framework supports multi-level evaluation: (1) first-person evaluation provides multimodal (visual, dialogue, action, etc.) input from a first-person perspective during a task for real-time state inference, (2) third-person evaluation provides complete third-person perspective video and text records after a task for goal and behavior inference. This evaluation method allows for a more comprehensive examination of a model's ToM capabilities from both the subjective immediate experience and the objective global observation. We constructed a challenging dataset containing 35 third-person perspective videos, 363 first-person perspective images, and 1225 expert-annotated multiple-choice questions (three options). On this dataset, we systematically evaluated the performance of human subjects and several state-of-the-art large vision-language models (LVLMs). The results show that LVLMs perform significantly worse than humans on SoMi-ToM: the average accuracy gap between humans and models is 40.1% in first-person evaluation and 26.4% in third-person evaluation. This indicates that future LVLMs need to further improve their ToM capabilities in embodied, complex social interactions.
Xianzhe Fan, Chuanyang Jin, Kolby Nottingham, Hao Zhu 0011, Maarten Sap
NeurIPS6
2025 Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
abstract
Recent advances in reasoning techniques have substantially improved the performance of large language models (LLMs), raising expectations for their ability to provide accurate, truthful, and reliable information. However, emerging evidence suggests that iterative reasoning may foster belief entrenchment, rather than enhancing truth-seeking behavior. In this study, we propose a systematic evaluation framework for *belief entrenchment* in LLM reasoning by leveraging the Martingale property from Bayesian statistics. This property implies that, under rational belief updating, the expected value of future beliefs should remain equal to the current belief, i.e., belief updates cannot be predicted from solely the current belief. We propose the unsupervised, regression-based *Martingale Score* to measure violations of this property, signaling a deviation from the Bayesian ability of updating on new evidence. In open-ended problem domains, including event forecasting, value-laden questions, and academic paper review, we found such violations to be widespread across models, reasoning paradigms, problem domains, and system prompts, where the future beliefs are consistently predictable from the model's current belief, a phenomenon which we term *belief entrenchment*. Through comprehensive experiments, we identify the models (e.g., GPT-4o), reasoning techniques (e.g., chain of thought), and domains (e.g., forecasting) more prone to belief entrenchment. Finally, we validate the Martingale Score by showing that it predicts ground-truth accuracy on problem domains where ground truth labels are available. This indicates that, while designed as an unsupervised metric that operates even in domains without access to ground truth, the Martingale Score is a useful proxy of the truth-seeking ability of the LLM reasoning process.
Zhonghao He, Tianyi Qiu, Hirokazu Shirado, Maarten Sap
NeurIPS4
2025 Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
abstract
Large language models (LMs) often struggle to generate diverse, human-like creative content, raising concerns about the long-term homogenization of human thought through repeated exposure to similar outputs. Yet scalable methods for evaluating LM output diversity remain limited, especially beyond narrow tasks such as random number or name generation, or beyond repeated sampling from a single model. To address this gap, we introduce Infinity-Chat, a large-scale dataset of 26K diverse, real-world, open-ended user queries that admit a wide range of plausible answers with no single ground truth. We introduce the first comprehensive taxonomy for characterizing the full spectrum of open-ended prompts posed to LMs, comprising 6 top-level categories (e.g., creative content generation, brainstorm & ideation) that further breaks down to 17 subcategories. Using Infinity-Chat, we present a large-scale study of mode collapse in LMs, revealing a pronounced Artificial Hivemind effect in open-ended generation of LMs, characterized by (1) intra-model repetition, where a single model consistently generates similar responses, and more so (2) inter-model homogeneity, where different models produce strikingly similar outputs. Infinity-Chat also includes 31,250 human annotations, across absolute ratings and pairwise preferences, with 25 independent human annotations per example. This enables studying collective and individual-specific human preferences in response to open-ended queries. Our findings show that state-of-the-art LMs, reward models, and LM judges are less well calibrated to human ratings on model generations that elicit differing idiosyncratic annotator preferences, despite maintaining comparable overall quality. Overall, INFINITY-CHAT presents the first large-scale resource for systematically studying real-world open-ended queries to LMs, revealing critical insights to guide future research for mitigating long-term AI safety risks posed by the Artificial Hivemind.
Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, Yejin Choi 0001
NeurIPS8
2024 Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
abstract
Human values are crucial to human decision-making. Value pluralism is the view that multiple correct values may be held in tension with one another (e.g., when considering lying to a friend to protect their feelings, how does one balance honesty with friendship?). As statistical learners, AI systems fit to averages by default, washing out these potentially irreducible value conflicts. To improve AI systems to better reflect value pluralism, the first-order challenge is to explore the extent to which AI systems can model pluralistic human values, rights, and duties as well as their interaction. We introduce ValuePrism, a large-scale dataset of 218k values, rights, and duties connected to 31k human-written situations. ValuePrism’s contextualized values are generated by GPT-4 and deemed high-quality by human annotators 91% of the time. We conduct a large-scale study with annotators across diverse social and demographic backgrounds to try to understand whose values are represented. With ValuePrism, we build Value Kaleidoscope (or Kaleido), an open, light-weight, and structured language-based multi-task model that generates, explains, and assesses the relevance and valence (i.e., support or oppose) of human values, rights, and duties within a specific context. Humans prefer the sets of values output by our system over the teacher GPT- 4, finding them more accurate and with broader coverage. In addition, we demonstrate that Kaleido can help explain variability in human decision-making by outputting contrasting values. Finally, we show that Kaleido’s representations transfer to other philosophical frameworks and datasets, confirming the benefit of an explicit, modular, and interpretable approach to value pluralism. We hope that our work will serve as a step to making more explicit the implicit values behind human decision-making and to steering AI systems to make decisions that are more in accordance with them.
Taylor Sorensen, Jena D. Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, Maarten Sap, John Tasioulas, Yejin Choi 0001
AAAI11
2024 Where Do People Tell Stories Online? Story Detection Across Online Communities
abstract
Story detection in online communities is a challenging task as stories are scattered across communities and interwoven with non-storytelling spans within a single text.We address this challenge by building and releasing the StorySeeker toolkit, including a richly annotated dataset of 502 Reddit posts and comments, a detailed codebook adapted to the social media context, and models to predict storytelling at the document and span levels.Our dataset is sampled from hundreds of popular Englishlanguage Reddit communities ranging across 33 topic categories, and it contains fine-grained expert annotations, including binary story labels, story spans, and event spans.We evaluate a range of detection methods using our data, and we identify the distinctive textual features of online storytelling, focusing on storytelling spans.We illuminate distributional characteristics of storytelling on a large communitycentric social media platform, and we also conduct a case study on r/ChangeMyView, where storytelling is used as one of many persuasive strategies, illustrating that our data and models can be used for both inter-and intra-community research.Finally, we discuss implications of our tools and analyses for narratology and the study of online communities.
Maria Antoniak, Joel Mire, Maarten Sap, Elliott Ash, Andrew Piper
ACL (1)3
2024 SOTOPIA-π: Interactive Learning of Socially Intelligent Language Agents
abstract
Ruiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi, Maarten Sap, Yonatan Bisk, Graham Neubig, Hao Zhu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Ruiyi Wang, Haofei Yu, Wenxin Zhang 0004, Zhengyang Qi, Maarten Sap, Yonatan Bisk, Graham Neubig, Hao Zhu 0011
ACL (1)5
2024 Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
abstract
As natural language becomes the default interface for human-AI interaction, there is a need for LMs to appropriately communicate uncertainties in downstream applications.In this work, we investigate how LMs incorporate confidence in responses via natural language and how downstream users behave in response to LM-articulated uncertainties.We examine publicly deployed models and find that LMs are reluctant to express uncertainties when answering questions even when they produce incorrect responses.LMs can be explicitly prompted to express confidences, but tend to be overconfident, resulting in high error rates (an average of 47%) among confident responses.We test the risks of LM overconfidence by conducting human experiments and show that users rely heavily on LM generations, whether or not they are marked by certainty.Lastly, we investigate the preference-annotated datasets used in post training alignment and find that humans are biased against texts with uncertainty.Our work highlights new safety harms facing human-LM interactions and proposes design recommendations and mitigating strategies moving forward.
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren 0001, Maarten Sap
ACL (1)4
2024 Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
abstract
General purpose AI, such as ChatGPT, seems to have lowered the barriers for the public to use AI and harness its power. However, the governance and development of AI still remain in the hands of a few, and the pace of development is accelerating without a comprehensive assessment of risks. As a first step towards democratic risk assessment and design of general purpose AI, we introduce PARTICIP-AI, a carefully designed framework for laypeople to speculate and assess AI use cases and their impacts. Our framework allows us to study more nuanced and detailed public opinions on AI through collecting use cases, surfacing diverse harms through risk assessment under alternate scenarios (i.e., developing and not developing a use case), and illuminating tensions over AI devel- opment through making a concluding choice on its development. To showcase the promise of our framework towards informing democratic AI development, we run a medium-scale study with inputs from 295 demographically diverse participants. Our analyses show that participants’ responses emphasize applications for personal life and society, contrasting with most current AI development’s business focus. We also surface diverse set of envisioned harms such as distrust in AI and institutions, complementary to those defined by experts. Furthermore, we found that perceived impact of not developing use cases significantly predicted participants’ judgements of whether AI use cases should be developed, and highlighted lay users’ concerns of techno-solutionism. We conclude with a discussion on how frameworks like PARTICIP-AI can further guide democratic AI development and governance.
Jimin Mun, Jenny T. Liang, Inyoung Cheong, Nicole DeCario, Yejin Choi 0001, Tadayoshi Kohno, Maarten Sap
AIES (1)8
2024 Counterspeakers' Perspectives: Unveiling Barriers and AI Needs in the Fight against Online Hate
abstract
Counterspeech, i.e., direct responses against hate speech, has become an important tool to address the increasing amount of hate online while avoiding censorship. Although AI has been proposed to help scale up counterspeech efforts, this raises questions of how exactly AI could assist in this process, since counterspeech is a deeply empathetic and agentic process for those involved. In this work, we aim to answer this question, by conducting in-depth interviews with 10 extensively experienced counterspeakers and a large scale public survey with 342 everyday social media users. In participant responses, we identified four main types of barriers and AI needs related to resources, training, impact, and personal harms. However, our results also revealed overarching concerns of authenticity, agency, and functionality in using AI tools for counterspeech. To conclude, we discuss considerations for designing AI assistants that lower counterspeaking barriers without jeopardizing its meaning and purpose.
Jimin Mun, Cathy Buerger, Jenny T. Liang, Joshua Garland, Maarten Sap
CHI5
2024 Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
abstract
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, Vered Shwartz. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Yejin Choi 0001, Yoav Goldberg, Maarten Sap, Vered Shwartz
EACL (1)7
2024 The Empirical Variability of Narrative Perceptions of Social Media Texts
abstract
Most NLP work on narrative detection has focused on prescriptive definitions of stories crafted by researchers, leaving open the questions: how do crowd workers perceive texts to be a story, and why?We investigate this by building STORYPERCEPTIONS, a dataset of 2,496 perceptions of storytelling in 502 social media texts from 255 crowd workers, including categorical labels along with free-text storytelling rationales, authorial intent, and more.We construct a fine-grained bottom-up taxonomy of crowd workers' varied and nuanced perceptions of storytelling by open-coding their free-text rationales.Through comparative analyses at the label and code level, we illuminate patterns of disagreement among crowd workers and across other annotation contexts, including prescriptive labeling from researchers and LLM-based predictions.Notably, plot complexity, references to generalized or abstract actions, and holistic aesthetic judgments (such as a sense of cohesion) are especially important in disagreements.Our empirical findings broaden understanding of the types, relative importance, and contentiousness of features relevant to narrative detection, highlighting opportunities for future work on reader-contextualized models of narrative reception.
Joel Mire, Maria Antoniak, Elliott Ash, Andrew Piper, Maarten Sap
EMNLP5
2024 HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs
abstract
Empathy serves as a cornerstone in enabling prosocial behaviors, and can be evoked through sharing of personal experiences in stories.While empathy is influenced by narrative content, intuitively, people respond to the way a story is told as well, through narrative style.Yet the relationship between empathy and narrative style is not fully understood.In this work, we empirically examine and quantify this relationship between style and empathy using LLMs and large-scale crowdsourcing studies.We introduce a novel, theory-based taxonomy, HEART (Human Empathy and Narrative Taxonomy) that delineates elements of narrative style that can lead to empathy with the narrator of a story.We establish the performance of LLMs in extracting narrative elements from HEART, showing that prompting with our taxonomy leads to reasonable, human-level annotations beyond what prior lexicon-based methods can do.To show empirical use of our taxonomy, we collect a dataset of empathy judgments of stories via a large-scale crowdsourcing study with N = 2, 624 participants.1 We show that narrative elements extracted via LLMs, in particular, vividness of emotions and plot volume, can elucidate the pathways by which narrative style cultivates empathy towards personal stories.Our work suggests that such models can be used for narrative analyses that lead to human-centered social and behavioral insights.1. Flatness/roundness (Keen, 2006) of the charac-
Jocelyn Shen, Joel Mire, Hae Park, Cynthia Breazeal, Maarten Sap
EMNLP5
2024 Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
abstract
Recent advances in large language models (LLM) have enabled richer social simulations, allowing for the study of various social phenomena.However, most recent work has used a more omniscient perspective on these simulations (e.g., single LLM to generate all interlocutors), which is fundamentally at odds with the non-omniscient, information asymmetric interactions that involve humans and AI agents in the real world.To examine these differences, we develop an evaluation framework to simulate social interactions with LLMs in various settings (omniscient, non-omniscient).Our experiments show that LLMs perform better in unrealistic, omniscient simulation settings but struggle in ones that more accurately reflect real-world conditions with information asymmetry.Our findings indicate that addressing information asymmetry remains a fundamental challenge for LLM-based agents.
Tiwalayo Eisape, Hyunwoo Kim 0002, Maarten Sap
EMNLP5
2024 Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
abstract
Reinforcement Learning with Human Feedback (RLHF) is the most prominent method for Language Model (LM) alignment. However, RLHF is an unstable and data-hungry process that continually requires new high-quality LM-generated data for finetuning. We introduce Advantage-Leftover Lunch RL (A-LoL), a new class of offline policy gradient algorithms that enable RL training on any pre-existing data. By assuming the entire LM output sequence as a single action, A-LoL allows incorporating sequence-level classifiers or human-designed scoring functions as rewards. Subsequently, by using LM’s value estimate, A-LoL only trains on positive advantage (leftover) data points, making it resilient to noise. Overall, A-LoL is an easy-to-implement, sample-efficient, and stable LM training recipe. We demonstrate the effectiveness of A-LoL and its variants with a set of four different language generation tasks. We compare against both online RL (PPO) and recent preference-based (DPO, PRO) and reward-based (GOLD) offline RL baselines. On the commonly-used RLHF benchmark, Helpful and Harmless Assistant (HHA), LMs trained with A-LoL methods achieve the highest diversity while also being rated more safe and helpful than the baselines according to humans. Additionally, in the remaining three tasks, A-LoL could optimize multiple distinct reward functions even when using noisy or suboptimal training data.
Ashutosh Baheti, Ximing Lu, Faeze Brahman, Ronan Le Bras 0001, Maarten Sap, Mark O. Riedl
ICLR5
2024 Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
abstract
Existing efforts on quantifying privacy implications for large language models (LLMs) solely focus on measuring leakage of training data. In this work, we shed light on the often-overlooked interactive settings where an LLM receives information from multiple sources and generates an output to be shared with other entities, creating the potential of exposing sensitive input data in inappropriate contexts. In these scenarios, humans nat- urally uphold privacy by choosing whether or not to disclose information depending on the context. We ask the question “Can LLMs demonstrate an equivalent discernment and reasoning capability when considering privacy in context?” We propose CONFAIDE, a benchmark grounded in the theory of contextual integrity and designed to identify critical weaknesses in the privacy reasoning capabilities of instruction-tuned LLMs. CONFAIDE consists of four tiers, gradually increasing in complexity, with the final tier evaluating contextual privacy reasoning and theory of mind capabilities. Our experiments show that even commercial models such as GPT-4 and ChatGPT reveal private information in contexts that humans would not, 39% and 57% of the time, respectively, highlighting the urgent need for a new direction of privacy-preserving approaches as we demonstrate a larger underlying problem stemmed in the models’ lack of reasoning capabilities.
Niloofar Mireshghallah, Hyunwoo Kim 0002, Yulia Tsvetkov, Maarten Sap, Reza Shokri, Yejin Choi 0001
ICLR5
2024 SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
abstract
*Humans are social beings*; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and *interact* under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents.
Hao Zhu 0011, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, Maarten Sap
ICLR11
2024 WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
abstract
We introduce WildTeaming, an automatic red-teaming framework that mines in-the-wild user-chatbot interactions to discover 5.7K unique clusters of novel jailbreak tactics, and then composes selections of multiple mined tactics for systematic exploration of novel and even more challenging jailbreaks. Compared to prior work that performed red-teaming via recruited human workers, gradient-based optimization, or iterative revision with large language models (LLMs), our work investigates jailbreaks from chatbot users in-the-wild who were not specifically instructed to break the system. WildTeaming reveals previously unidentified vulnerabilities of frontier LLMs, resulting in more diverse and successful adversarial attacks compared to state-of-the-art jailbreaking methods. While there exist many datasets for jailbreak evaluation, very few open-source datasets exist for jailbreak training, as safety training data has been closed among all frontier models even when their weights are open. Therefore, with WildTeaming we create WildJailbreak, a large-scale open-source synthetic safety dataset with 262K vanilla (direct request) and adversarial (complex jailbreak) prompt-response pairs. In order to mitigate exaggerated safety behaviors, WildJailbreak provides two contrastive types of queries: 1) harmful queries (both vanilla and adversarial) and 2) benign queries that resemble harmful queries in form but contain no harmful intent. As WildJailbreak considerably upgrades the quality and scale of existing safety resources, it uniquely enables us to examine the scaling effects of data and the interplay of data properties and model capabilities during safety training. Through extensive model training and evaluations, we identify the training properties that enable an ideal balance of safety behaviors: appropriate safeguarding without over-refusal, effective handling of both vanilla and adversarial queries, and minimal, if any, decrease in general capabilities. All the components of WildJailbreak contribute to achieving balanced safety behaviors of models
Kavel Rao, Seungju Han 0002, Allyson Ettinger, Faeze Brahman, Sachin Kumar 0009, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi 0001, Nouha Dziri
NeurIPS9
2023 From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
abstract
Warning: content in this paper may be upsetting or offensive to some readers.Dogwhistles are coded expressions that simultaneously convey one meaning to a broad audience and a second one, often hateful or provocative, to a narrow in-group; they are deployed to evade both political repercussions and algorithmic content moderation.For example, in the sentence "we need to end the cosmopolitan experiment," the word "cosmopolitan" likely means "worldly" to many, but secretly means "Jewish" to a select few.We present the first large-scale computational investigation of dogwhistles.We develop a typology of dogwhistles, curate the largest-to-date glossary of over 300 dogwhistles with rich contextual information and examples, and analyze their usage in historical U.S. politicians' speeches.We then assess whether a large language model (GPT-3) can identify dogwhistles and their meanings, and find that GPT-3's performance varies widely across types of dogwhistles and targeted groups.Finally, we show that harmful content containing dogwhistles avoids toxicity detection, highlighting online risks of such coded language.This work sheds light on the theoretical and applied importance of dogwhistles in both NLP and computational social science, and provides resources for future research in modeling dogwhistles and mitigating their online harms.
Julia Mendelsohn, Ronan Le Bras 0001, Yejin Choi 0001, Maarten Sap
ACL (1)4
2023 NLPositionality: Characterizing Design Biases of Datasets and Models
abstract
Design biases in NLP systems, such as performance differences for different populations, often stem from their creator's positionality, i.e., views and lived experiences shaped by identity and background.Despite the prevalence and risks of design biases, they are hard to quantify because researcher, system, and dataset positionality is often unobserved.We introduce NLPositionality, a framework for characterizing design biases and quantifying the positionality of NLP datasets and models.Our framework continuously collects annotations from a diverse pool of volunteer participants on LabintheWild, and statistically quantifies alignment with dataset labels and model predictions.We apply NLPositionality to existing datasets and models for two tasks-social acceptability and hate speech detection.To date, we have collected 16, 299 annotations in over a year for 600 instances from 1, 096 annotators across 87 countries.We find that datasets and models align predominantly with Western, White, college-educated, and younger populations.Additionally, certain groups, such as nonbinary people and non-native English speakers, are further marginalized by datasets and models as they rank least in alignment across all tasks.Finally, we draw from prior literature to discuss how researchers can examine their own positionality and that of their datasets and models, opening the door for more inclusive NLP systems.
Sebastin Santy, Jenny T. Liang, Ronan Le Bras 0001, Katharina Reinecke, Maarten Sap
ACL (1)5
2023 SODA: Million-scale Dialogue Distillation with Social Commonsense Contextualization
abstract
Hyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Hyunwoo Kim 0002, Jack Hessel, Peter West, Ximing Lu, Youngjae Yu, Ronan Le Bras 0001, Malihe Alikhani, Gunhee Kim, Maarten Sap, Yejin Choi 0001
EMNLP11
2023 FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
abstract
Theory of mind (ToM) evaluations currently focus on testing models using passive narratives that inherently lack interactivity.We introduce FANTOM, a new benchmark designed to stress-test ToM within information-asymmetric conversational contexts via question answering.Our benchmark draws upon important theoretical requisites from psychology and necessary empirical considerations when evaluating large language models (LLMs).In particular, we formulate multiple types of questions that demand the same underlying reasoning to identify illusory or false sense of ToM capabilities in LLMs.We show that FANTOM is challenging for state-of-the-art LLMs, which perform significantly worse than humans even with chainof-thought reasoning or fine-tuning.1 Linda: Yeah, I got a golden retriever.She's so adorable.David: What's her favorite food?Kailey: Hey guys, I'
Hyunwoo Kim 0002, Melanie Sclar, Ronan Le Bras 0001, Gunhee Kim, Yejin Choi 0001, Maarten Sap
EMNLP7
2023 Modeling Empathic Similarity in Personal Narratives
abstract
The most meaningful connections between people are often fostered through expression of shared vulnerability and emotional experiences in personal narratives.We introduce a new task of identifying similarity in personal stories based on empathic resonance, i.e., the extent to which two people empathize with each others' experiences, as opposed to raw semantic or lexical similarity, as has predominantly been studied in NLP.Using insights from social psychology, we craft a framework that operationalizes empathic similarity in terms of three key features of stories: main events, emotional trajectories, and overall morals or takeaways.We create EM-PATHICSTORIES, a dataset of 1,500 personal stories annotated with our empathic similarity features, and 2,000 pairs of stories annotated with empathic similarity scores.Using our dataset, we finetune a model to compute empathic similarity of story pairs, and show that this outperforms semantic similarity models on automated correlation and retrieval metrics.Through a user study with 150 participants, we also assess the effect our model has on retrieving stories that users empathize with, compared to naive semantic similarity-based retrieval, and find that participants empathized significantly more with stories retrieved by our model.Our work has strong implications for the use of empathy-aware models to foster human connection and empathy between people.
Jocelyn Shen, Maarten Sap, Pedro Colon-Hernandez, Hae Park, Cynthia Breazeal
EMNLP2
2023 Don't Take This Out of Context!: On the Need for Contextual Models and Evaluations for Stylistic Rewriting
abstract
Most existing stylistic text rewriting methods and evaluation metrics operate on a sentence level, but ignoring the broader context of the text can lead to preferring generic, ambiguous, and incoherent rewrites.In this paper, we investigate integrating the preceding textual context into both the rewriting and evaluation stages of stylistic text rewriting, and introduce a new composite contextual evaluation metric CtxSimFit that combines similarity to the original sentence with contextual cohesiveness.We comparatively evaluate non-contextual and contextual rewrites in formality, toxicity, and sentiment transfer tasks.Our experiments show that humans significantly prefer contextual rewrites as more fitting and natural over non-contextual ones, yet existing sentence-level automatic metrics (e.g., ROUGE, SBERT) correlate poorly with human preferences (ρ=0-0.3).In contrast, human preferences are much better reflected by both our novel CtxSimFit (ρ=0.7-0.9) as well as proposed context-infused versions of common metrics (ρ=0.4-0.7).Overall, our findings highlight the importance of integrating context into the generation and especially the evaluation stages of stylistic text rewriting.
Akhila Yerukola, Elizabeth Clark, Maarten Sap
EMNLP4
2023 BiasX: "Thinking Slow" in Toxic Content Moderation with Explanations of Implied Social Biases
abstract
Toxicity annotators and content moderators often default to mental shortcuts when making decisions.This can lead to subtle toxicity being missed, and seemingly toxic but harmless content being over-detected.We introduce BIASX, a framework that assists content moderators with free-text explanations of statements' implied social biases, and explore its effectiveness through a large-scale user study.We show that participants indeed benefit substantially from explanations for correctly moderating subtly (non-)toxic content.The quality of explanations is critical: imperfect machine-generated explanations (+2.4% on hard toxic examples) help less compared to expert-written human explanations (+7.2%).Our results showcase the promise of using free-text explanations to encourage more thoughtful toxicity moderation. 1
Yiming Zhang 0022, Sravani Nanduri, Sherry Tongshuang Wu, Maarten Sap
EMNLP5
2022 Misinfo Reaction Frames: Reasoning about Readers' Reactions to News Headlines
abstract
Saadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen, Franziska Roesner, Eunsol Choi, Yejin Choi. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Saadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen, Franziska Roesner, Eunsol Choi, Yejin Choi 0001
ACL (1)3
2022 ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
abstract
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, Ece Kamar. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, Ece Kamar
ACL (1)4
2022 ProsocialDialog: A Prosocial Backbone for Conversational Agents
abstract
Most existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them.To address this issue, we introduce PROSOCIALDIALOG, the first large-scale multi-turn dialogue dataset to teach conversational agents to respond to problematic content following social norms.Covering diverse unethical, problematic, biased, and toxic situations, PROSOCIALDIALOG contains responses that encourage prosocial behavior, grounded in commonsense social rules (i.e., rules-ofthumb, RoTs).Created via a human-AI collaborative framework, PROSOCIALDIALOG consists of 58K dialogues, with 331K utterances, 160K unique RoTs, and 497K dialogue safety labels accompanied by free-form rationales.With this dataset, we introduce a dialogue safety detection module, Canary, capable of generating RoTs given conversational context, and a socially-informed dialogue agent, Prost.Empirical results show that Prost generates more socially acceptable dialogues compared to other state-of-the-art language and dialogue models in both in-domain and out-of-domain settings.Additionally, Canary effectively guides off-the-shelf language models to generate significantly more prosocial responses.Our work highlights the promise and importance of creating and steering conversational AI to be socially responsible.
Hyunwoo Kim 0002, Youngjae Yu, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi 0001, Maarten Sap
EMNLP8
2022 Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs
abstract
Social intelligence and Theory of Mind (TOM), i.e., the ability to reason about the different mental states, intents, and reactions of all people involved, allow humans to effectively navigate and understand everyday social interactions.As NLP systems are used in increasingly complex social situations, their ability to grasp social dynamics becomes crucial.
Maarten Sap, Ronan Le Bras 0001, Daniel Fried, Yejin Choi 0001
EMNLP1
2022 Aligning to Social Norms and Values in Interactive Narratives
abstract
Prithviraj Ammanabrolu, Liwei Jiang, Maarten Sap, Hannaneh Hajishirzi, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Prithviraj Ammanabrolu, Maarten Sap, Hannaneh Hajishirzi, Yejin Choi 0001
NAACL-HLT3
2022 Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection
abstract
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Yejin Choi 0001, Noah A. Smith
NAACL-HLT1
2022 When to Make Exceptions: Exploring Language Models as Accounts of Human Moral Judgment
abstract
AI systems are becoming increasingly intertwined with human life. In order to effectively collaborate with humans and ensure safety, AI systems need to be able to understand, interpret and predict human moral judgments and decisions. Human moral judgments are often guided by rules, but not always. A central challenge for AI safety is capturing the flexibility of the human moral mind — the ability to determine when a rule should be broken, especially in novel or unusual situations. In this paper, we present a novel challenge set consisting of moral exception question answering (MoralExceptQA) of cases that involve potentially permissible moral exceptions – inspired by recent moral psychology studies. Using a state-of-the-art large language model (LLM) as a basis, we propose a novel moral chain of thought (MoralCoT) prompting strategy that combines the strengths of LLMs with theories of moral reasoning developed in cognitive science to predict human moral judgments. MoralCoT outperforms seven existing LLMs by 6.2% F1, suggesting that modeling human reasoning might be necessary to capture the flexibility of the human moral mind. We also conduct a detailed error analysis to suggest directions for future work to improve AI safety using MoralExceptQA. Our data is open-sourced at https://huggingface.co/datasets/feradauto/MoralExceptQA and code at https://github.com/feradauto/MoralCoT.
Zhijing Jin 0001, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, Bernhard Schölkopf
NeurIPS5
2021 DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts
abstract
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, Yejin Choi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, Yejin Choi 0001
ACL/IJCNLP (1)2
2021 Challenges in Automated Debiasing for Toxic Language Detection
abstract
Warning: this paper contains content that may be offensive or upsetting.Biased associations have been a challenge in the development of classifiers for detecting toxic language, hindering both fairness and accuracy.As potential solutions, we investigate recently introduced debiasing methods for text classification datasets and models, as applied to toxic language detection.Our focus is on lexical (e.g., swear words, slurs, identity mentions) and dialectal markers (specifically African American English).Our comprehensive experiments establish that existing methods are limited in their ability to prevent biased behavior in current toxicity detectors.We then propose an automatic, dialect-aware data correction method, as a proof-of-concept study.Despite the use of synthetic labels, this method reduces dialectal associations with toxicity.Overall, our findings show that debiasing a model trained on biased toxic language data is not as effective as simply relabeling the data to remove existing biases.
Maarten Sap, Swabha Swayamdipta, Yejin Choi 0001, Noah A. Smith
EACL2
2021 Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive Contexts
abstract
Dialogue models trained on human conversations inadvertently learn to generate toxic responses.In addition to producing explicitly offensive utterances, these models can also implicitly insult a group or individual by aligning themselves with an offensive statement.To better understand the dynamics of contextually offensive language, we investigate the stance of dialogue model responses in offensive Reddit conversations.Specifically, we create TOXICHAT, a crowd-annotated dataset of 2,000 Reddit threads and model responses labeled with offensive language and stance.Our analysis reveals that 42% of human responses agree with toxic comments, whereas only 13% agree with safe comments.This undesirable behavior is learned by neural dialogue models, such as DialoGPT, which we show are two times more likely to agree with offensive comments.To enable automatic detection of offensive language, we fine-tuned transformerbased classifiers on TOXICHAT that achieve 0.71 F 1 for offensive labels and 0.53 Macro-F 1 for stance labels.Finally, we quantify the effectiveness of controllable text generation (CTG) methods to mitigate the tendency of neural dialogue models to agree with offensive comments.Compared to the baseline, our best CTG model achieves a 19% reduction in agreement with offensive comments and produces 29% fewer offensive replies.Our work highlights the need for further efforts to characterize and analyze inappropriate behavior in dialogue models, in order to help make them safer. 1
Ashutosh Baheti, Maarten Sap, Alan Ritter, Mark O. Riedl
EMNLP (1)2
2021 Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
abstract
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner 0001
EMNLP (1)2
2021 Detoxifying Language Models Risks Marginalizing Minority Voices
abstract
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, Dan Klein. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, Daniel Klein 0001
NAACL-HLT5
2020 Social Bias Frames: Reasoning about Social and Power Implications of Language
abstract
contains content that may be offensive or upsetting.
Maarten Sap, Saadia Gabriel, Lianhui Qin, Daniel Jurafsky, Noah A. Smith, Yejin Choi 0001
ACL1
2020 Recollection versus Imagination: Exploring Human Memory and Cognition via Neural Language Models
abstract
We investigate the use of NLP as a measure of the cognitive processes involved in storytelling, contrasting imagination and recollection of events. To facilitate this, we collect and release Hippocorpus, a dataset of 7,000 stories about imagined and recalled events. We introduce a measure of narrative flow and use this to examine the narratives for imagined and recalled events. Additionally, we measure the differential recruitment of knowledge attributed to semantic memory versus episodic memory (Tulving, 1972) for imagined and recalled storytelling by comparing the frequency of descriptions of general commonsense events with more specific realis events. Our analyses show that imagined stories have a substantially more linear narrative flow, compared to recalled stories in which adjacent sentences are more disconnected. In addition, while recalled stories rely more on autobiographical events based on episodic memory, imagined stories express more commonsense knowledge based on semantic memory. Finally, our measures reveal the effect of narrativization of memories in stories (e.g., stories about frequently recalled memories flow more linearly; Bartlett, 1932). Our findings highlight the potential of using NLP tools to study the traces of human cognition in language.
Maarten Sap, Eric Horvitz, Yejin Choi 0001, Noah A. Smith, James W. Pennebaker
ACL1
2020 Social Chemistry 101: Learning to Reason about Social and Moral Norms
abstract
Social norms-the unspoken commonsense rules about acceptable social behavior-are crucial in understanding the underlying causes and intents of people's actions in narratives.For example, underlying an action such as "wanting to call cops on my neighbor" are social norms that inform our conduct, such as "It is expected that you report crimes."We present SOCIAL CHEMISTRY, a new conceptual formalism to study people's everyday social norms and moral judgments over a rich spectrum of real life situations described in natural language.We introduce SOCIAL-CHEM-101, a large-scale corpus that catalogs 292k rules-of-thumb such as "It is rude to run a blender at 5am" as the basic conceptual units.Each rule-of-thumb is further broken down with 12 different dimensions of people's judgments, including social judgments of good and bad, moral foundations, expected cultural pressure, and assumed legality, which together amount to over 4.5 million annotations of categorical labels and free-text descriptions.Comprehensive empirical results based on state-of-the-art neural models demonstrate that computational modeling of social norms is a promising research direction.Our model framework, NEURAL NORM TRANSFORMER, learns and generalizes SOCIAL-CHEM-101 to successfully reason about previously unseen situations, generating relevant (and potentially novel) attribute-aware social rules-of-thumb.Punching a friend who stole from me.RoT 1: It is unacceptable to injure a person.RoT 2: People should not steal from others.RoT 3: It is bad to betray a friend.RoT 4: It is OK to want to take revenge.
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, Yejin Choi 0001
EMNLP (1)4
2020 PowerTransformer: Unsupervised Controllable Revision for Biased Language Correction
abstract
Unconscious biases continue to be prevalent in modern text and media, calling for algorithms that can assist writers with bias correction.For example, a female character in a story is often portrayed as passive and powerless ("She daydreams about being a doctor") while a man is portrayed as more proactive and powerful ("He pursues his dream of being a doctor").We formulate Controllable Debiasing, a new revision task that aims to rewrite a given text to correct the implicit and potentially undesirable bias in character portrayals.We then introduce POWERTRANSFORMER as an approach that debiases text through the lens of connotation frames (Sap et al., 2017), which encode pragmatic knowledge of implied power dynamics with respect to verb predicates.One key challenge of our task is the lack of parallel corpora.To address this challenge, we adopt an unsupervised approach using auxiliary supervision with related tasks such as paraphrasing and self-supervision based on a reconstruction loss, building on pretrained language models.Through comprehensive experiments based on automatic and human evaluations, we demonstrate that our approach outperforms ablations and existing methods from related tasks.Furthermore, we demonstrate the use of POWER-TRANSFORMER as a step toward mitigating the well-documented gender bias in character portrayal in movie scripts.
Xinyao Ma, Maarten Sap, Hannah Rashkin, Yejin Choi 0001
EMNLP (1)2
2019 ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning
abstract
We present ATOMIC, an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge. Compared to existing resources that center around taxonomic knowledge, ATOMIC focuses on inferential knowledge organized as typed if-then relations with variables (e.g., “if X pays Y a compliment, then Y will likely return the compliment”). We propose nine if-then relation types to distinguish causes vs. effects, agents vs. themes, voluntary vs. involuntary events, and actions vs. mental states. By generatively training on the rich inferential knowledge described in ATOMIC, we show that neural models can acquire simple commonsense capabilities and reason about previously unseen events. Experimental results demonstrate that multitask models that incorporate the hierarchical structure of if-then relation types lead to more accurate inference compared to models trained in isolation, as measured by both automatic and human evaluation.
Maarten Sap, Ronan Le Bras 0001, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A. Smith, Yejin Choi 0001
AAAI1
2019 COMET: Commonsense Transformers for Automatic Knowledge Graph Construction
abstract
We present the first comprehensive study on automatic knowledge base construction for two prevalent commonsense knowledge graphs: ATOMIC (Sap et al., 2019) and Con-ceptNet (Speer et al., 2017).Contrary to many conventional KBs that store knowledge with canonical templates, commonsense KBs only store loosely structured open-text descriptions of knowledge.We posit that an important step toward automatic commonsense completion is the development of generative models of commonsense knowledge, and propose COMmonsEnse Transformers (COMET ) that learn to generate rich and diverse commonsense descriptions in natural language.Despite the challenges of commonsense modeling, our investigation reveals promising results when implicit knowledge from deep pre-trained language models is transferred to generate explicit knowledge in commonsense knowledge graphs.Empirical results demonstrate that COMET is able to generate novel knowledge that humans rate as high quality, with up to 77.5% (ATOMIC) and 91.7% (ConceptNet) precision at top 1, which approaches human performance for these resources.Our findings suggest that using generative commonsense models for automatic commonsense KB completion could soon be a plausible alternative to extractive methods.
Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Celikyilmaz, Yejin Choi 0001
ACL (1)3
2019 The Risk of Racial Bias in Hate Speech Detection
abstract
We investigate how annotators' insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations.We first uncover unexpected correlations between surface markers of African American English (AAE) and ratings of toxicity in several widely-used hate speech datasets.Then, we show that models trained on these corpora acquire and propagate these biases, such that AAE tweets and tweets by self-identified African Americans are up to two times more likely to be labelled as offensive compared to others.Finally, we propose dialect and race priming as ways to reduce the racial bias in annotation, showing that when annotators are made explicitly aware of an AAE tweet's dialect they are significantly less likely to label the tweet as offensive.
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi 0001, Noah A. Smith
ACL (1)1
2019 Social IQa: Commonsense Reasoning about Social Interactions
abstract
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, Yejin Choi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras 0001, Yejin Choi 0001
EMNLP/IJCNLP (1)1
2018 Modeling Naive Psychology of Characters in Simple Commonsense Stories
abstract
Understanding a narrative requires reading between the lines and reasoning about the unspoken but obvious implications about events and people's mental states -a capability that is trivial for humans but remarkably hard for machines.To facilitate research addressing this challenge, we introduce a new annotation framework to explain naive psychology of story characters as fully-specified chains of mental states with respect to motivations and emotional reactions.Our work presents a new largescale dataset with rich low-level annotations and establishes baseline performance on several new tasks, suggesting avenues for future research.
Hannah Rashkin, Antoine Bosselut, Maarten Sap, Kevin Knight, Yejin Choi 0001
ACL (1)3
2018 Event2Mind: Commonsense Inference on Events, Intents, and Reactions
abstract
We investigate a new commonsense inference task: given an event described in a short free-form text ("X drinks coffee in the morning"), a system reasons about the likely intents ("X wants to stay awake") and reactions ("X feels alert") of the event's participants.To support this study, we construct a new crowdsourced corpus of 25,000 event phrases covering a diverse range of everyday events and situations.We report baseline performance on this task, demonstrating that neural encoder-decoder models can successfully compose embedding representations of previously unseen events and reason about the likely intents and reactions of the event participants.In addition, we demonstrate how commonsense inference on people's intents and reactions can help unveil the implicit gender inequality prevalent in modern movie scripts. 1 https://tinyurl.com/event2mind
Hannah Rashkin, Maarten Sap, Emily Allaway, Noah A. Smith, Yejin Choi 0001
ACL (1)2
2017 The Effect of Different Writing Tasks on Linguistic Style: A Case Study of the ROC Story Cloze Task
abstract
A writer's style depends not just on personal traits but also on her intent and mental state.In this paper, we show how variants of the same writing task can lead to measurable differences in writing style.We present a case study based on the story cloze task (Mostafazadeh et al., 2016a), where annotators were assigned similar writing tasks with different constraints: (1) writing an entire story, (2) adding a story ending for a given story context, and (3) adding an incoherent ending to a story.We show that a simple linear classifier informed by stylistic features is able to successfully distinguish among the three cases, without even looking at the story context.In addition, combining our stylistic features with language model predictions reaches state of the art performance on the story cloze challenge.Our results demonstrate that different task framings can dramatically affect the way people write. 1 1 This paper extends our LSDSem 2017 shared task submission (Schwartz et al., 2017).
Roy Schwartz 0001, Maarten Sap, Ioannis Konstas, Leila Zilles, Yejin Choi 0001, Noah A. Smith
CoNLL2
2017 Connotation Frames of Power and Agency in Modern Films
abstract
The framing of an action influences how we perceive its actor.We introduce connotation frames of power and agency, a pragmatic formalism organized using frame semantic representations, to model how different levels of power and agency are implicitly projected on actors through their actions.We use the new power and agency frames to measure the subtle, but prevalent, gender bias in the portrayal of modern film characters and provide insights that deviate from the well-known Bechdel test.Our contributions include an extended lexicon of connotation frames along with a web interface that provides a comprehensive analysis through the lens of connotation frames.
Maarten Sap, Marcella Cindy Prasettio, Ari Holtzman, Hannah Rashkin, Yejin Choi 0001
EMNLP1
2015 Extracting Human Temporal Orientation from Facebook Language
abstract
H. Andrew Schwartz, Gregory Park, Maarten Sap, Evan Weingarten, Johannes Eichstaedt, Margaret Kern, David Stillwell, Michal Kosinski, Jonah Berger, Martin Seligman, Lyle Ungar. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
H. Andrew Schwartz, Gregory J. Park, Maarten Sap, Evan Weingarten, Johannes C. Eichstaedt, Margaret L. Kern, David Stillwell, Michal Kosinski, Jonah Berger, Martin E. P. Seligman, Lyle H. Ungar
HLT-NAACL3
2014 Developing Age and Gender Predictive Lexica over Social Media
abstract
Maarten Sap, Gregory Park, Johannes Eichstaedt, Margaret Kern, David Stillwell, Michal Kosinski, Lyle Ungar, Hansen Andrew Schwartz. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Maarten Sap, Gregory J. Park, Johannes C. Eichstaedt, Margaret L. Kern, David Stillwell, Michal Kosinski, Lyle H. Ungar, H. Andrew Schwartz
EMNLP1