Myra Cheng

dblp:226/7067 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0002-5052-2929ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 8 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 78% Trustworthy machine learning · 22%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
1.012026
Thinking beyond the anthropomorphic paradigm benefits LLM research · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model safety
1.012026
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs · ACL (1) 2026
Machine learning › Trustworthy machine learning
fairness
0.712023
CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations · EMNLP 2023
Machine learning › Trustworthy machine learning › fairness › bias in language models
stereotype measurement in language models
0.712023
Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models · ACL (1) 2023
Computational social science and digital humanities › agent-based simulation
LLM-based social simulation
0.712023
CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations · EMNLP 2023
Natural language and speech › Language models and text generation › alignment
sycophancy
0.312026
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs · ACL (1) 2026
Machine learning › Trustworthy machine learning
interpretability
0.312025
HumT DumT: Measuring and controlling human-like language in LLMs · ACL (1) 2025
Machine learning › Trustworthy machine learning › fairness › social bias
representational harm
0.212023
Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

taxonomy development · 1.7empirical interaction analysis · 1.7crowdsourcing study · 1.7conceptual framework · 1.7prompting interventions · 1.0pragmatic analysis · 1.0bibliometric analysis · 1.0activation steering · 0.9LLM-based probability metrics · 0.9persona simulation · 0.7markedness analysis · 0.7evaluation framework · 0.7
YearPublicationVenuePosition
2026 Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
abstract
Recent evaluations show that large language models (LLMs) frequently fail to challenge users' harmful beliefs in domains ranging from medical advice to social reasoning.We present a unifying analysis through the lens of pragmatics: these safety failures can be understood and addressed as LLMs exhibiting excessive accommodation and insufficient epistemic vigilance.We show that the pragmatic factors affecting accommodation and epistemic vigilance in humans (at-issueness, linguistic encoding, and source reliability) influence LLM behaviors in similar ways.We demonstrate how these factors explain performance differences across three safety benchmarks that test models' ability to challenge harmful beliefs, spanning misinformation (Cancer-Myth, SAGE-Eval) and sycophancy (ELEPHANT).This pragmatic lens further motivates prompting interventions, such as adding the phrase "wait a minute", that drastically improve performance on these difficult benchmarks by shifting pragmatic cues.Our results have practical implications for benchmark design and underscore the importance of pragmatics for understanding model behavior and improving performance.
Myra Cheng, Robert D. Hawkins, Daniel Jurafsky
ACL (1)1
2026 Thinking beyond the anthropomorphic paradigm benefits LLM research
abstract
Anthropomorphism, or the attribution of human traits to technology, is an automatic and unconscious response that occurs even in those with advanced technical expertise.In this position paper, we analyze hundreds of thousands of research articles to present empirical evidence of the prevalence and growth of anthropomorphic terminology in research on large language models (LLMs).We argue for challenging the deeper assumptions reflected in this terminology -which, though often useful, may inadvertently constrain LLM development -and broadening beyond them to open new pathways for understanding and improving LLMs.Specifically, we identify and examine five anthropomorphic assumptions that shape research across the LLM development lifecycle.For each assumption (e.g., that LLMs must use natural language for reasoning, or that they should be evaluated on benchmarks originally meant for humans), we demonstrate empirical, non-anthropomorphic alternatives that remain under-explored yet offer promising directions for LLM research and development.
Lujain Ibrahim, Myra Cheng
ACL (1)2
2025 Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems
abstract
As text generation systems' outputs are increasingly anthropomorphic-perceived as humanlike-scholars have also increasingly raised concerns about how such outputs can lead to harmful outcomes, such as users over-relying or developing emotional dependence on these systems.How to intervene on such system outputs to mitigate anthropomorphic behaviors and their attendant harmful outcomes, however, remains understudied.With this work, we aim to provide empirical and theoretical grounding for developing such interventions.To do so, we compile an inventory of interventions grounded both in prior literature and a crowdsourcing study where participants edited system outputs to make them less human-like.Drawing on this inventory, we also develop a conceptual framework to help characterize the landscape of possible interventions, articulate distinctions between different types of interventions, and provide a theoretical basis for evaluating the effectiveness of different interventions.
Myra Cheng, Su Lin Blodgett, Alicia DeVrio, Lisa Egede, Alexandra Olteanu
ACL (1)1
2025 HumT DumT: Measuring and controlling human-like language in LLMs
abstract
Should LLMs generate language that makes them seem human?Human-like language might improve user experience, but might also lead to deception, overreliance, and stereotyping.Assessing these potential impacts requires a systematic way to measure human-like tone in LLM outputs.We introduce HUMT and SO-CIOT, metrics for human-like tone and other dimensions of social perceptions in text data based on relative probabilities from an LLM.By measuring HUMT across preference and usage datasets, we find that users prefer less human-like outputs from LLMs in many contexts.HUMT also offers insights into the perceptions and impacts of anthropomorphism: human-like LLM outputs are highly correlated with warmth, social closeness, femininity, and low status, which are closely linked to the aforementioned harms.We introduce DUMT, a method using HUMT to systematically control and reduce the degree of human-like tone while preserving model performance.DUMT offers a practical approach for mitigating risks associated with anthropomorphic language generation.
Myra Cheng, Sunny Yu, Daniel Jurafsky
ACL (1)1
2025 A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies
abstract
Recent attention to anthropomorphism -- the attribution of human-like qualities to non-human objects or entities -- of language technologies like LLMs has sparked renewed discussions about potential negative impacts of anthropomorphism. To productively discuss the impacts of this anthropomorphism and in what contexts it is appropriate, we need a shared vocabulary for the vast variety of ways that language can be anthropomorphic. In this work, we draw on existing literature and analyze empirical cases of user interactions with language technologies to develop a taxonomy of textual expressions that can contribute to anthropomorphism. We highlight challenges and tensions involved in understanding linguistic anthropomorphism, such as how all language is fundamentally human and how efforts to characterize and shift perceptions of humanness in machines can also dehumanize certain humans. We discuss ways that our taxonomy supports more precise and effective discussions of and decisions about anthropomorphism of language technologies.
Alicia DeVrio, Myra Cheng, Lisa Egede, Alexandra Olteanu, Su Lin Blodgett
CHI2
2024 AnthroScore: A Computational Linguistic Measure of Anthropomorphism
abstract
Myra Cheng, Kristina Gligoric, Tiziano Piccardi, Dan Jurafsky. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Myra Cheng, Kristina Gligoric, Tiziano Piccardi, Daniel Jurafsky
EACL (1)1
2024 NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
abstract
Kristina Gligoric, Myra Cheng, Lucia Zheng, Esin Durmus, Dan Jurafsky. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Kristina Gligoric, Myra Cheng, Lucia Zheng, Esin Durmus, Daniel Jurafsky
NAACL-HLT2
2023 Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models
abstract
To recognize and mitigate harms from large language models (LLMs), we need to understand the prevalence and nuances of stereotypes in LLM outputs.Toward this end, we present Marked Personas, a prompt-based method to measure stereotypes in LLMs for intersectional demographic groups without any lexicon or data labeling.Grounded in the sociolinguistic concept of markedness (which characterizes explicitly linguistically marked categories versus unmarked defaults), our proposed method is twofold: 1) prompting an LLM to generate personas, i.e., natural language descriptions, of the target demographic group alongside personas of unmarked, default groups; 2) identifying the words that significantly distinguish personas of the target group from corresponding unmarked ones.We find that the portrayals generated by GPT-3.5 and GPT-4 contain higher rates of racial stereotypes than human-written portrayals using the same prompts.The words distinguishing personas of marked (non-white, non-male) groups reflect patterns of othering and exoticizing these demographics.An intersectional lens further reveals tropes that dominate portrayals of marginalized groups, such as tropicalism and the hypersexualization of minoritized women.These representational harms have concerning implications for downstream applications like story generation.
Myra Cheng, Esin Durmus, Daniel Jurafsky
ACL (1)1
2023 CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations
abstract
Recent work has aimed to capture nuances of human behavior by using LLMs to simulate responses from particular demographics in settings like social science experiments and public opinion surveys.However, there are currently no established ways to discuss or evaluate the quality of such LLM simulations.Moreover, there is growing concern that these simulations are flattened caricatures of the personas that they aim to simulate, failing to capture the multidimensionality of people and perpetuating stereotypes.To bridge these gaps, we present CoMPosT, a framework to characterize LLM simulations using four dimensions: Context, Model, Persona, and Topic.We use this framework to measure open-ended LLM simulations' susceptibility to caricature, defined via two criteria: individuation and exaggeration.We evaluate the level of caricature in scenarios from existing work on LLM simulations.We find that for GPT-4, simulations of certain demographics (political and marginalized groups) and topics (general, uncontroversial) are highly susceptible to caricature.
Myra Cheng, Tiziano Piccardi, Diyi Yang
EMNLP1
2023 Social norm bias: residual harms of fairness-aware algorithms
Myra Cheng, Maria De-Arteaga, Lester Mackey, Adam Tauman Kalai
Data Min. Knowl. Discov.1
2020 Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
abstract
Optimizing lower-body exoskeleton walking gaits for user comfort requires understanding users' preferences over a high-dimensional gait parameter space. However, existing preference-based learning methods have only explored low-dimensional domains due to computational limitations. To learn user preferences in high dimensions, this work presents LINECOSPAR, a human-in-the-loop preference-based framework that enables optimization over many parameters by iteratively exploring one-dimensional subspaces. Additionally, this work identifies gait attributes that characterize broader preferences across users. In simulations and human trials, we empirically verify that LINECOSPAR is a sample-efficient approach for high-dimensional preference optimization. Our analysis of the experimental data reveals a correspondence between human preferences and objective measures of dynamicity, while also highlighting differences in the utility functions underlying individual users' gait preferences. This result has implications for exoskeleton gait synthesis, an active field with applications to clinical use and patient rehabilitation.
Maegan Tucker, Myra Cheng, Ellen R. Novoseller, Richard Cheng, Yisong Yue, Joel W. Burdick, Aaron D. Ames
IROS2