Amanda Cercas Curry

dblp:185/0457 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-5576-2550ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 10 since 2021
YearPublicationVenuePosition
2026 ACID: On the Perception of Online Classism
abstract
Socioeconomic status (SES) structures social inequality and underlies class-based discrimination that is often rationalised through stereotypes expressed in public discourse. However, despite extensive research on hate speech detection in Natural Language Processing, classism detection remains an underexplored phenomenon. We introduce ACID, a cross-cultural corpus with over 1.15 million instances, to investigate classism across YouTube and Twitter from 14 English-speaking countries. We examine (i) which stereotypes are invoked towards lower-SES, (ii) whether blame for lower-SES is attributed to individuals or structural factors, and (iii) whether these people are portrayed offensively. Across platforms, explanations are predominantly framed in terms of individual responsibility. Across countries, class stereotypes consistently revolve around moralized notions of dependency, laziness, and ignorance, revealing a shared global structure of class-based stigma. Our dataset and analysis are a foundation to advance research on class-based discrimination and its representation in online discourse.
Arianna Muti, Elisa Bassignana, Amanda Cercas Curry, Federica Durante, Dirk Hovy, Debora Nozza
LREC3
2025 The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
abstract
Socioeconomic status (SES) fundamentally influences how people interact with each other and more recently, with digital technologies like Large Language Models (LLMs).While previous research has highlighted the interaction between SES and language technology, it was limited by reliance on proxy metrics and synthetic data.We survey 1,000 individuals from diverse socioeconomic backgrounds about their use of language technologies and generative AI, and collect 6,482 prompts from their previous interactions with LLMs.We find systematic differences across SES groups in language technology usage (i.e., frequency, performed tasks), interaction styles, and topics.Higher SES entails a higher level of abstraction, convey requests more concisely, and topics like 'inclusivity' and 'travel'.Lower SES correlates with higher anthropomorphization of LLMs (using "hello" and "thank you") and more concrete language.Our findings suggest that while generative language technologies are becoming more accessible to everyone, socioeconomic linguistic differences still stratify their use to exacerbate the digital divide.These differences underscore the importance of considering SES in developing language technologies to accommodate varying linguistic needs rooted in socioeconomic factors and limit the AI Gap across SES groups.
Elisa Bassignana, Amanda Cercas Curry, Dirk Hovy
ACL (1)2
2025 Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks
abstract
Justin Zhao, Flor Miriam Plaza-del-Arco, Benjamin Genchel, Amanda Cercas Curry. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Justin Zhao, Flor Miriam Plaza del Arco, Amanda Cercas Curry
NAACL (Long Papers)3
2024 Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution
abstract
Flor Miriam Plaza-del-Arco, Amanda Cercas Curry, Alba Curry, Gavin Abercrombie, Dirk Hovy. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Flor Miriam Plaza del Arco, Amanda Cercas Curry, Alba Curry, Gavin Abercrombie, Dirk Hovy
ACL (1)2
2024 Classist Tools: Social Class Correlates with Performance in NLP
abstract
The field of sociolinguistics has studied factors affecting language use for the last century.Labov (1964) and Bernstein (1960) showed that socioeconomic class strongly influences our accents, syntax and lexicon.However, despite growing concerns surrounding fairness and bias in Natural Language Processing (NLP), there is a dearth of studies delving into the effects it may have on NLP systems.We show empirically that NLP systems' performance is affected by speakers' SES, potentially disadvantaging less-privileged socioeconomic groups.We annotate a corpus of 95K utterances from movies with social class, ethnicity and geographical language variety and measure the performance of NLP systems on three tasks: language modelling, automatic speech recognition, and grammar error correction.We find significant performance disparities that can be attributed to socioeconomic status as well as ethnicity and geographical differences.1 With NLP technologies becoming ever more ubiquitous and quotidian, they must accommodate all language varieties to avoid disadvantaging already marginalised groups.We argue for the inclusion of socioeconomic class in future language technologies.
Amanda Cercas Curry, Giuseppe Attanasio, Zeerak Talat, Dirk Hovy
ACL (1)1
2024 Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions
abstract
Emotions are a central aspect of communication. Consequently, emotion analysis (EA) is a rapidly growing field in natural language processing (NLP). However, there is no consensus on scope, direction, or methods. In this paper, we conduct a thorough review of 154 relevant NLP publications from the last decade. Based on this review, we address four different questions: (1) How are EA tasks defined in NLP? (2) What are the most prominent emotion frameworks and which emotions are modeled? (3) Is the subjectivity of emotions considered in terms of demographics and cultural factors? and (4) What are the primary NLP applications for EA? We take stock of trends in EA and tasks, emotion frameworks used, existing datasets, methods, and applications. We then discuss four lacunae: (1) the absence of demographic and cultural aspects does not account for the variation in how emotions are perceived, but instead assumes they are universally experienced in the same manner; (2) the poor fit of emotion categories from the two main emotion theories to the task; (3) the lack of standardized EA terminology hinders gap identification, comparison, and future goals; and (4) the absence of interdisciplinary research isolates EA from insights in other fields. Our work will enable more focused research into EA and a more holistic approach to modeling emotions in NLP.
Flor Miriam Plaza del Arco, Alba Curry, Amanda Cercas Curry, Dirk Hovy
LREC/COLING3
2024 Impoverished Language Technology: The Lack of (Social) Class in NLP
abstract
Since Labov’s foundational 1964 work on the social stratification of language, linguistics has dedicated concerted efforts towards understanding the relationships between socio-demographic factors and language production and perception. Despite the large body of evidence identifying significant relationships between socio-demographic factors and language production, relatively few of these factors have been investigated in the context of NLP technology. While age and gender are well covered, Labov’s initial target, socio-economic class, is largely absent. We survey the existing Natural Language Processing (NLP) literature and find that only 20 papers even mention socio-economic status. However, the majority of those papers do not engage with class beyond collecting information of annotator-demographics. Given this research lacuna, we provide a definition of class that can be operationalised by NLP researchers, and argue for including socio-economic class in future language technologies.
Amanda Cercas Curry, Zeerak Talat, Dirk Hovy
LREC/COLING1
2023 Mirages. On Anthropomorphism in Dialogue Systems
abstract
Automated dialogue or conversational systems are anthropomorphised by developers and personified by users.While a degree of anthropomorphism may be inevitable due to the choice of medium, conscious and unconscious design choices can guide users to personify such systems to varying degrees.Encouraging users to relate to automated systems as if they were human can lead to high risk scenarios caused by over-reliance on their outputs.As a result, natural language processing researchers have investigated the factors that induce personification and develop resources to mitigate such effects.However, these efforts are fragmented, and many aspects of anthropomorphism have yet to be explored.In this paper, we discuss the linguistic factors that contribute to the anthropomorphism of dialogue systems and the harms that can arise, including reinforcing gender stereotypes and notions of acceptable language.We recommend that future efforts towards developing dialogue systems take particular care in their design, development, release, and description; and attend to the many linguistic cues that can elicit personification by users.
Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser, Zeerak Talat
EMNLP2
2023 Viewpoint: Artificial Intelligence Accidents Waiting to Happen?
abstract
Artificial Intelligence (AI) is at a crucial point in its development: stable enough to be used in production systems, and increasingly pervasive in our lives. What does that mean for its safety? In his book Normal Accidents, the sociologist Charles Perrow proposed a framework to analyze new technologies and the risks they entail. He showed that major accidents are nearly unavoidable in complex systems with tightly coupled components if they are run long enough. In this essay, we apply and extend Perrow’s framework to AI to assess its potential risks. Today’s AI systems are already highly complex, and their complexity is steadily increasing. As they become more ubiquitous, different algorithms will interact directly, leading to tightly coupled systems whose capacity to cause harm we will be unable to predict. We argue that under the current paradigm, Perrow’s normal accidents apply to AI systems and it is only a matter of time before one occurs. This article appears in the AI & Society track.
Federico Bianchi 0001, Amanda Cercas Curry, Dirk Hovy
J. Artif. Intell. Res.2
2021 ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Detection in Conversational AI
abstract
We present the first English corpus study on abusive language towards three conversational AI systems gathered 'in the wild': an opendomain social bot, a rule-based chatbot, and a task-based system.To account for the complexity of the task, we take a more 'nuanced' approach where our ConvAI dataset reflects fine-grained notions of abuse, as well as views from multiple expert annotators.We find that the distribution of abuse is vastly different compared to other commonly used datasets, with more sexually tinted aggression towards the virtual persona of these systems.Finally, we report results from bench-marking existing models against this data.Unsurprisingly, we find that there is substantial room for improvement with F1 scores below 90%.Warning: This paper contains examples of language that some people may find offensive or upsetting.
Amanda Cercas Curry, Gavin Abercrombie, Verena Rieser
EMNLP (1)1
2019 A Crowd-based Evaluation of Abuse Response Strategies in Conversational Agents
abstract
How should conversational agents respond to verbal abuse through the user?To answer this question, we conduct a large-scale crowdsourced evaluation of abuse response strategies employed by current state-of-the-art systems.Our results show that some strategies, such as "polite refusal" score highly across the board, while for other strategies demographic factors, such as age, as well as the severity of the preceding abuse influence the user's perception of which response is appropriate.In addition, we find that most data-driven models lag behind rule-based or commercial systems in terms of their perceived appropriateness.
Amanda Cercas Curry, Verena Rieser
SIGdial1
2017 Why We Need New Evaluation Metrics for NLG
abstract
The majority of NLG evaluation relies on automatic metrics, such as BLEU.In this paper, we motivate the need for novel, system-and data-independent automatic evaluation methods: We investigate a wide range of metrics, including state-of-the-art word-based and novel grammar-based ones, and demonstrate that they only weakly reflect human judgements of system outputs as generated by data-driven, end-to-end NLG.We also show that metric performance is data-and system-specific.Nevertheless, our results also suggest that automatic metrics perform reliably at system-level and can support system development by finding cases where a system performs poorly.4 https://github.com/glampouras/JLOLS_NLG 5 Note that we use lexicalised versions of SFHOTEL and SFREST and a partially lexicalised version of BAGEL, where proper names and place names are replaced by placeholders ("X"), in correspondence with the outputs generated by the MR: inform(name=X, area=X, pricerange=moderate, type=restaurant) Reference: "X is a moderately priced restaurant in X."
Jekaterina Novikova, Ondrej Dusek, Amanda Cercas Curry, Verena Rieser
EMNLP3