VLDB 2026 Research / reviewers in the wild / expert
Dirk Hovy
dblp:82/8159
· DBLP profile ↗
61ranked-venue papers
13as first author
35since 2021 · last 2026
0000-0002-4618-3127ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 58 · 12 first-author · 34 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Responsible Evaluation of AI for Mental HealthabstractHiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza-del-Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz 0001, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza del Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych |
ACL (1) | 14 |
| 2026 | ACID: On the Perception of Online ClassismabstractSocioeconomic status (SES) structures social inequality and underlies class-based discrimination that is often rationalised through stereotypes expressed in public discourse. However, despite extensive research on hate speech detection in Natural Language Processing, classism detection remains an underexplored phenomenon. We introduce ACID, a cross-cultural corpus with over 1.15 million instances, to investigate classism across YouTube and Twitter from 14 English-speaking countries. We examine (i) which stereotypes are invoked towards lower-SES, (ii) whether blame for lower-SES is attributed to individuals or structural factors, and (iii) whether these people are portrayed offensively. Across platforms, explanations are predominantly framed in terms of individual responsibility. Across countries, class stereotypes consistently revolve around moralized notions of dependency, laziness, and ignorance, revealing a shared global structure of class-based stigma. Our dataset and analysis are a foundation to advance research on class-based discrimination and its representation in online discourse. Arianna Muti, Elisa Bassignana, Amanda Cercas Curry, Federica Durante, Dirk Hovy, Debora Nozza |
LREC | 5 |
| 2026 | IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing AssistanceabstractAbstract Large language models (LLMs) are helping millions of users write texts about diverse issues, and in doing so expose users to different ideas and perspectives. This creates concerns about issue bias, where an LLM tends to present just one perspective on a given issue, which in turn may influence how users think about this issue. So far, it has not been possible to measure which issue biases LLMs manifest in real user interactions, making it difficult to address the risks from biased LLMs. Therefore, we create IssueBench: a set of 2.49m realistic English-language prompts to measure issue bias in LLM writing assistance, which we construct based on 3.9k templates (e.g., “write a blog about”) and 212 political issues (e.g., “AI regulation”) from real user interactions. Using IssueBench, we show that issue biases are common and persistent in 10 state-of-the-art LLMs. We also show that biases are very similar across models, and that all models align more with US Democrat than Republican voter opinion on a subset of issues. IssueBench can easily be adapted to include other issues, templates, or tasks. By enabling robust and realistic measurement, we hope that IssueBench can bring a new quality of evidence to ongoing discussions about LLM biases and how to address them. Paul Röttger, Musashi Hinck, Valentin Hofmann, Kobi Hackenburg, Valentina Pyatkin, Faeze Brahman, Dirk Hovy |
Trans. Assoc. Comput. Linguistics | 7 |
| 2025 | SafetyPrompts: A Systematic Review of Open Datasets for Evaluating and Improving Large Language Model SafetyabstractThe last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abundance of datasets for evaluating and improving LLM safety. However, much of this work has happened in parallel, and with very different goals in mind, ranging from the mitigation of near-term risks around bias and toxic content generation to the assessment of longer-term catastrophic risk potential. This makes it difficult for researchers and practitioners to find the most relevant datasets for their use case, and to identify gaps in dataset coverage that future work may fill. To remedy these issues, we conduct a first systematic review of open datasets for evaluating and improving LLM safety. We review 144 datasets, which we identified through an iterative and community-driven process over the course of several months. We highlight patterns and trends, such as a trend towards fully synthetic datasets, as well as gaps in dataset coverage, such as a clear lack of non-English and naturalistic datasets. We also examine how LLM safety datasets are used in practice -- in LLM release publications and popular LLM benchmarks -- finding that current evaluation practices are highly idiosyncratic and make use of only a small fraction of available datasets. Our contributions are based on SafetyPrompts.com, a living catalogue of open datasets for LLM safety, which we plan to update continuously as the field of LLM safety develops. Paul Röttger, Fabio Pernisi, Bertie Vidgen, Dirk Hovy |
AAAI | 4 |
| 2025 | The AI Gap: How Socioeconomic Status Affects Language Technology InteractionsabstractSocioeconomic status (SES) fundamentally influences how people interact with each other and more recently, with digital technologies like Large Language Models (LLMs).While previous research has highlighted the interaction between SES and language technology, it was limited by reliance on proxy metrics and synthetic data.We survey 1,000 individuals from diverse socioeconomic backgrounds about their use of language technologies and generative AI, and collect 6,482 prompts from their previous interactions with LLMs.We find systematic differences across SES groups in language technology usage (i.e., frequency, performed tasks), interaction styles, and topics.Higher SES entails a higher level of abstraction, convey requests more concisely, and topics like 'inclusivity' and 'travel'.Lower SES correlates with higher anthropomorphization of LLMs (using "hello" and "thank you") and more concrete language.Our findings suggest that while generative language technologies are becoming more accessible to everyone, socioeconomic linguistic differences still stratify their use to exacerbate the digital divide.These differences underscore the importance of considering SES in developing language technologies to accommodate varying linguistic needs rooted in socioeconomic factors and limit the AI Gap across SES groups. Elisa Bassignana, Amanda Cercas Curry, Dirk Hovy |
ACL (1) | 3 |
| 2025 | Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text PerceptionsabstractMatthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano, David Jurgens, Dirk Hovy. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Matthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano, David Jurgens, Dirk Hovy |
ACL (1) | 6 |
| 2025 | Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task PerformanceabstractExpert persona prompting-assigning roles such as expert in math to language models-is widely used for task improvement.However, prior work shows mixed results on its effectiveness, and does not consider when and why personas should improve performance.We analyze the literature on persona prompting for task improvement and distill three desiderata: 1) performance advantage of expert personas, 2) robustness to irrelevant persona attributes, and 3) fidelity to persona attributes.We then evaluate 9 state-of-the-art LLMs across 27 tasks with respect to these desiderata.We find that expert personas usually lead to positive or non-significant performance changes.Surprisingly, models are highly sensitive to irrelevant persona details, with performance drops of almost 30 percentage points.In terms of fidelity, we find that while higher education, specialization, and domain-relatedness can boost performance, their effects are often inconsistent or negligible across tasks.We propose mitigation strategies to improve robustness-but find they only work for the largest, most capable models.Our findings underscore the need for more careful persona design and for evaluation schemes that reflect the intended effects of persona usage. Pedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, Benjamin Roth 0001 |
EMNLP | 3 |
| 2025 | Biased Tales: Cultural and Topic Bias in Generating Children's StoriesabstractStories play a pivotal role in human communication, shaping beliefs and morals, particularly in children.As parents increasingly rely on large language models (LLMs) to craft bedtime stories, the presence of cultural and gender stereotypes in these narratives raises significant concerns.To address this issue, we present Biased Tales, a comprehensive dataset designed to analyze how biases influence protagonists' attributes and story elements in LLM-generated stories.Our analysis uncovers striking disparities.When the protagonist is described as a girl (as compared to a boy), appearance-related attributes increase by 55.26%.Stories featuring non-Western children disproportionately emphasize cultural heritage, tradition, and family themes far more than those for Western children.Our findings highlight the role of sociocultural bias in making creative AI use more equitable and diverse. Donya Rooein, Vilém Zouhar, Debora Nozza, Dirk Hovy |
EMNLP | 4 |
| 2024 | Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion AttributionabstractFlor Miriam Plaza-del-Arco, Amanda Cercas Curry, Alba Curry, Gavin Abercrombie, Dirk Hovy. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Flor Miriam Plaza del Arco, Amanda Cercas Curry, Alba Curry, Gavin Abercrombie, Dirk Hovy |
ACL (1) | 5 |
| 2024 | Classist Tools: Social Class Correlates with Performance in NLPabstractThe field of sociolinguistics has studied factors affecting language use for the last century.Labov (1964) and Bernstein (1960) showed that socioeconomic class strongly influences our accents, syntax and lexicon.However, despite growing concerns surrounding fairness and bias in Natural Language Processing (NLP), there is a dearth of studies delving into the effects it may have on NLP systems.We show empirically that NLP systems' performance is affected by speakers' SES, potentially disadvantaging less-privileged socioeconomic groups.We annotate a corpus of 95K utterances from movies with social class, ethnicity and geographical language variety and measure the performance of NLP systems on three tasks: language modelling, automatic speech recognition, and grammar error correction.We find significant performance disparities that can be attributed to socioeconomic status as well as ethnicity and geographical differences.1 With NLP technologies becoming ever more ubiquitous and quotidian, they must accommodate all language varieties to avoid disadvantaging already marginalised groups.We argue for the inclusion of socioeconomic class in future language technologies. Amanda Cercas Curry, Giuseppe Attanasio, Zeerak Talat, Dirk Hovy |
ACL (1) | 4 |
| 2024 | Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language ModelsabstractPaul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schuetze, Dirk Hovy. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schütze, Dirk Hovy |
ACL (1) | 7 |
| 2024 | Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future DirectionsabstractEmotions are a central aspect of communication. Consequently, emotion analysis (EA) is a rapidly growing field in natural language processing (NLP). However, there is no consensus on scope, direction, or methods. In this paper, we conduct a thorough review of 154 relevant NLP publications from the last decade. Based on this review, we address four different questions: (1) How are EA tasks defined in NLP? (2) What are the most prominent emotion frameworks and which emotions are modeled? (3) Is the subjectivity of emotions considered in terms of demographics and cultural factors? and (4) What are the primary NLP applications for EA? We take stock of trends in EA and tasks, emotion frameworks used, existing datasets, methods, and applications. We then discuss four lacunae: (1) the absence of demographic and cultural aspects does not account for the variation in how emotions are perceived, but instead assumes they are universally experienced in the same manner; (2) the poor fit of emotion categories from the two main emotion theories to the task; (3) the lack of standardized EA terminology hinders gap identification, comparison, and future goals; and (4) the absence of interdisciplinary research isolates EA from insights in other fields. Our work will enable more focused research into EA and a more holistic approach to modeling emotions in NLP. Flor Miriam Plaza del Arco, Alba Curry, Amanda Cercas Curry, Dirk Hovy |
LREC/COLING | 4 |
| 2024 | Impoverished Language Technology: The Lack of (Social) Class in NLPabstractSince Labov’s foundational 1964 work on the social stratification of language, linguistics has dedicated concerted efforts towards understanding the relationships between socio-demographic factors and language production and perception. Despite the large body of evidence identifying significant relationships between socio-demographic factors and language production, relatively few of these factors have been investigated in the context of NLP technology. While age and gender are well covered, Labov’s initial target, socio-economic class, is largely absent. We survey the existing Natural Language Processing (NLP) literature and find that only 20 papers even mention socio-economic status. However, the majority of those papers do not engage with class beyond collecting information of annotator-demographics. Given this research lacuna, we provide a definition of class that can be operationalised by NLP researchers, and argue for including socio-economic class in future language technologies. Amanda Cercas Curry, Zeerak Talat, Dirk Hovy |
LREC/COLING | 3 |
| 2024 | DADIT: A Dataset for Demographic Classification of Italian Twitter Users and a Comparison of Prediction MethodsabstractSocial scientists increasingly use demographically stratified social media data to study the attitudes, beliefs, and behavior of the general public. To facilitate such analyses, we construct, validate, and release publicly the representative DADIT dataset of 30M tweets of 20k Italian Twitter users, along with their bios and profile pictures. We enrich the user data with high-quality labels for gender, age, and location. DADIT enables us to train and compare the performance of various state-of-the-art models for the prediction of the gender and age of social media users. In particular, we investigate if tweets contain valuable information for the task, since popular classifiers like M3 don’t leverage them. Our best XLM-based classifier improves upon the commonly used competitor M3 by up to 53% F1. Especially for age prediction, classifiers profit from including tweets as features. We also confirm these findings on a German test set. Lorenzo Lupo, Paul Bose, Mahyar Habibi, Dirk Hovy, Carlo Schwarz |
LREC/COLING | 4 |
| 2024 | Explaining Speech Classification Models via Word-Level Audio Segments and Paralinguistic FeaturesabstractEliana Pastor, Alkis Koudounas, Giuseppe Attanasio, Dirk Hovy, Elena Baralis. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Eliana Pastor, Alkis Koudounas, Giuseppe Attanasio, Dirk Hovy, Elena Baralis |
EACL (1) | 4 |
| 2024 | Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance GapsabstractCurrent automatic speech recognition (ASR) models are designed to be used across many languages and tasks without substantial changes.However, this broad language coverage hides performance gaps within languages, for example, across genders.Our study systematically evaluates the performance of two widely used multilingual ASR models on three datasets, encompassing 19 languages from eight language families and two speaking conditions.Our findings reveal clear gender disparities, with the advantaged group varying across languages and models.Surprisingly, those gaps are not explained by acoustic or lexical properties.However, probing internal model states reveals a correlation with gendered performance gap.That is, the easier it is to distinguish speaker gender in a language using probes, the more the gap reduces, favoring female speakers.Our results show that gender disparities persist even in state-of-the-art models.Our findings have implications for the improvement of multilingual ASR systems, underscoring the importance of accessibility to training data and nuanced evaluation to predict and mitigate gender gaps.We release all code and artifacts at https://github.com/g8a9/multilingual -asr-gender-gap. Giuseppe Attanasio, Beatrice Savoldi, Dennis Fucci, Dirk Hovy |
EMNLP | 4 |
| 2024 | XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language ModelsabstractPaul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, Dirk Hovy. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Paul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi 0001, Dirk Hovy |
NAACL-HLT | 6 |
| 2023 | What about "em"? How Commercial Machine Translation Fails to Handle (Neo-)PronounsabstractAs 3rd-person pronoun usage shifts to include novel forms, e.g., neopronouns, we need more research on identity-inclusive NLP.Exclusion is particularly harmful in one of the most popular NLP applications, machine translation (MT).Wrong pronoun translations can discriminate against marginalized groups, e.g., non-binary individuals (Dev et al., 2021).In this "reality check", we study how three commercial MT systems translate 3rd-person pronouns.Concretely, we compare the translations of gendered vs. gender-neutral pronouns from English to five other languages (Danish, Farsi, French, German, Italian), and vice versa, from Danish to English.Our error analysis shows that the presence of a gender-neutral pronoun often leads to grammatical and semantic translation errors.Similarly, gender neutrality is often not preserved.By surveying the opinions of affected native speakers from diverse languages, we provide recommendations to address the issue in future MT research. Anne Lauscher, Debora Nozza, Ehm Miltersen, Archie Crowley, Dirk Hovy |
ACL (1) | 5 |
| 2023 | Top-Down Influence? Predicting CEO Personality and Risk Impact from Speech TranscriptsabstractHow much does a CEO’s personality impact the performanceof their company? Management theory posits a great influence, but it is difficult to show empirically—there is a lack of publicly available self-reported personality data of top managers. Instead, we propose a text-based personality regressor based on crowd-sourced Myers–Briggs Type Indicator (MBTI) assessments. The ratings have a high internal and external validity and can be predicted with moderate to strong correlations for three out of four dimensions. Providing evidence for the upper echelons theory, we demonstrate that the predicted CEO personalities have explanatory power of financial risk. Christoph Kilian Theil, Dirk Hovy, Heiner Stuckenschmidt |
ICWSM | 2 |
| 2023 | Beyond Digital "Echo Chambers": The Role of Viewpoint Diversity in Political DiscussionabstractIncreasingly taking place in online spaces, modern political conversations are typically perceived to be unproductively affirming---siloed in so called "echo chambers" of exclusively like-minded discussants. Yet, to date we lack sufficient means to measure viewpoint diversity in conversations. To this end, in this paper, we operationalize two viewpoint metrics proposed for recommender systems and adapt them to the context of social media conversations. This is the first study to apply these two metrics (Representation and Fragmentation) to real world data and to consider the implications for online conversations specifically. We apply these measures to two topics---daylight savings time (DST), which serves as a control, and the more politically polarized topic of immigration. We find that the diversity scores for both Fragmentation and Representation are lower for immigration than for DST. Further, we find that while pro-immigrant views receive consistent pushback on the platform, anti-immigrant views largely operate within echo chambers. We observe less severe yet similar patterns for DST. Taken together, Representation and Fragmentation paint a meaningful and important new picture of viewpoint diversity. Rishav Hada, Amir Ebrahimi Fard, Sarah Shugars, Federico Bianchi 0001, Patrícia G. C. Rossini, Dirk Hovy, Rebekah Tromble, Nava Tintarev |
WSDM | 6 |
| 2023 | Viewpoint: Artificial Intelligence Accidents Waiting to Happen?abstractArtificial Intelligence (AI) is at a crucial point in its development: stable enough to be used in production systems, and increasingly pervasive in our lives. What does that mean for its safety? In his book Normal Accidents, the sociologist Charles Perrow proposed a framework to analyze new technologies and the risks they entail. He showed that major accidents are nearly unavoidable in complex systems with tightly coupled components if they are run long enough. In this essay, we apply and extend Perrow’s framework to AI to assess its potential risks. Today’s AI systems are already highly complex, and their complexity is steadily increasing. As they become more ubiquitous, different algorithms will interact directly, leading to tightly coupled systems whose capacity to cause harm we will be unable to predict. We argue that under the current paradigm, Perrow’s normal accidents apply to AI systems and it is only a matter of time before one occurs. This article appears in the AI & Society track. Federico Bianchi 0001, Amanda Cercas Curry, Dirk Hovy |
J. Artif. Intell. Res. | 3 |
| 2022 | SafetyKit: First Aid for Measuring Safety in Open-domain Conversational SystemsabstractEmily Dinan, Gavin Abercrombie, A. Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, Verena Rieser. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Emily Dinan, Gavin Abercrombie, A. Stevie Bergman, Shannon L. Spruit, Dirk Hovy, Y-Lan Boureau, Verena Rieser |
ACL (1) | 5 |
| 2022 | Welcome to the Modern World of Pronouns: Identity-Inclusive Natural Language Processing beyond GenderabstractThe world of pronouns is changing – from a closed word class with few members to an open set of terms to reflect identities. However, Natural Language Processing (NLP) barely reflects this linguistic shift, resulting in the possible exclusion of non-binary users, even though recent work outlined the harms of gender-exclusive language technology. The current modeling of 3rd person pronouns is particularly problematic. It largely ignores various phenomena like neopronouns, i.e., novel pronoun sets that are not (yet) widely established. This omission contributes to the discrimination of marginalized and underrepresented groups, e.g., non-binary individuals. It thus prevents gender equality, one of the UN’s sustainable development goals (goal 5). Further, other identity-expressions beyond gender are ignored by current NLP technology. This paper provides an overview of 3rd person pronoun issues for NLP. Based on our observations and ethical considerations, we define a series of five desiderata for modeling pronouns in language technology, which we validate through a survey. We evaluate existing and novel modeling approaches w.r.t. these desiderata qualitatively and quantify the impact of a more discrimination-free approach on an established benchmark dataset. Anne Lauscher, Archie Crowley, Dirk Hovy |
COLING | 3 |
| 2022 | "It's Not Just Hate": A Multi-Dimensional Perspective on Detecting Harmful Speech OnlineabstractWell-annotated data is a prerequisite for good Natural Language Processing models.Too often, though, annotation decisions are governed by optimizing time or annotator agreement.We make a case for nuanced efforts in an interdisciplinary setting for annotating offensive online speech.Detecting offensive content is rapidly becoming one of the most important real-world NLP tasks.However, most datasets use a single binary label, e.g., for hate or incivility, even though each concept is multi-faceted.This modeling choice severely limits nuanced insights, but also performance.We show that a more fine-grained multi-label approach to predicting incivility and hateful or intolerant content addresses both conceptual and performance issues.We release a novel dataset of over 40,000 tweets about immigration from the US and UK, annotated with six labels for different aspects of incivility and intolerance.Our dataset not only allows for a more nuanced understanding of harmful speech online, models trained on it also outperform or match performance on benchmark datasets.Warning: This paper contains examples of hateful language some readers might find offensive. Federico Bianchi 0001, Stefanie Anja Hills, Patrícia G. C. Rossini, Dirk Hovy, Rebekah Tromble, Nava Tintarev |
EMNLP | 4 |
| 2022 | Bridging Fairness and Environmental Sustainability in Natural Language ProcessingabstractFairness and environmental impact are important research directions for the sustainable development of artificial intelligence.However, while each topic is an active research area in natural language processing (NLP), there is a surprising lack of research on the interplay between the two fields.This lacuna is highly problematic, since there is increasing evidence that an exclusive focus on fairness can actually hinder environmental sustainability, and vice versa.In this work, we shed light on this crucial intersection in NLP by (1) investigating the efficiency of current fairness approaches through surveying example methods for reducing unfair stereotypical bias from the literature, and(2) evaluating a common technique to reduce energy consumption (and thus environmental impact) of English NLP models, knowledge distillation (KD), for its impact on fairness.In this case study, we evaluate the effect of important KD factors, including layer and dimensionality reduction, with respect to: (a) performance on the distillation task (natural language inference and semantic similarity prediction), and (b) multiple measures and dimensions of stereotypical bias (e.g., gender bias measured via the Word Embedding Association Test).Our results lead us to clarify current assumptions regarding the effect of KD on unfair bias: contrary to other findings, we show that KD can actually decrease model fairness. Marius Hessenthaler, Emma Strubell, Dirk Hovy, Anne Lauscher |
EMNLP | 3 |
| 2022 | SocioProbe: What, When, and Where Language Models Learn about SociodemographicsabstractPre-trained language models (PLMs) have outperformed other NLP models on a wide range of tasks.Opting for a more thorough understanding of their capabilities and inner workings, researchers have established the extend to which they capture lower-level knowledge like grammaticality, and mid-level semantic knowledge like factual understanding.However, there is still little understanding of their knowledge of higher-level aspects of language.In particular, despite the importance of sociodemographic aspects in shaping our language, the questions of whether, where, and how PLMs encode these aspects, e.g., gender or age, is still unexplored.We address this research gap by probing the sociodemographic knowledge of different single-GPU PLMs on multiple English data sets via traditional classifier probing and information-theoretic minimum description length probing.Our results show that PLMs do encode these sociodemographics, and that this knowledge is sometimes spread across the layers of some of the tested PLMs.We further conduct a multilingual analysis and investigate the effect of supplementary training to further explore to what extent, where, and with what amount of pre-training data the knowledge is encoded.Our overall results indicate that sociodemographic knowledge is still a major challenge for NLP.PLMs require large amounts of pre-training data to acquire the knowledge and models that excel in general language understanding do not seem to own more knowledge about these aspects. Anne Lauscher, Federico Bianchi 0001, Samuel R. Bowman, Dirk Hovy |
EMNLP | 4 |
| 2022 | Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesabstractHate speech is a global phenomenon, but most hate speech datasets so far focus on Englishlanguage content.This hinders the development of more effective hate speech detection models in hundreds of languages spoken by billions across the world.More data is needed, but annotating hateful content is expensive, timeconsuming and potentially harmful to annotators.To mitigate these issues, we explore dataefficient strategies for expanding hate speech detection into under-resourced languages.In a series of experiments with mono-and multilingual models across five non-English languages, we find that 1) a small amount of target-language fine-tuning data is needed to achieve strong performance, 2) the benefits of using more such data decrease exponentially, and 3) initial fine-tuning on readily-available English data can partially substitute targetlanguage data and improve model generalisability.Based on these findings, we formulate actionable recommendations for hate speech detection in low-resource language settings. Paul Röttger, Debora Nozza, Federico Bianchi 0001, Dirk Hovy |
EMNLP | 4 |
| 2022 | Two Contrasting Data Annotation Paradigms for Subjective NLP TasksabstractPaul Rottger, Bertie Vidgen, Dirk Hovy, Janet Pierrehumbert. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Paul Röttger, Bertie Vidgen, Dirk Hovy, Janet B. Pierrehumbert |
NAACL-HLT | 3 |
| 2022 | Guiding the Release of Safer E2E Conversational AI through Value Sensitive DesignabstractA. Stevie Bergman, Gavin Abercrombie, Shannon Spruit, Dirk Hovy, Emily Dinan, Y-Lan Boureau, Verena Rieser. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022. A. Stevie Bergman, Gavin Abercrombie, Shannon L. Spruit, Dirk Hovy, Emily Dinan, Y-Lan Boureau, Verena Rieser |
SIGDIAL | 4 |
| 2021 | Cross-lingual Contextualized Topic Models with Zero-shot LearningabstractFederico Bianchi, Silvia Terragni, Dirk Hovy, Debora Nozza, Elisabetta Fersini. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Federico Bianchi 0001, Silvia Terragni, Dirk Hovy, Debora Nozza, Elisabetta Fersini |
EACL | 3 |
| 2021 | BERTective: Language Models and Contextual Information for Deception DetectionabstractSpotting a lie is challenging but has an enormous potential impact on security as well as private and public safety.Several NLP methods have been proposed to classify texts as truthful or deceptive.In most cases, however, the target texts' preceding context is not considered.This is a severe limitation, as any communication takes place in context, not in a vacuum, and context can help to detect deception.We study a corpus of Italian dialogues containing deceptive statements and implement deep neural models that incorporate various linguistic contexts.We establish a new state-of-theart identifying deception and find that not all context is equally useful to the task.Only the texts closest to the target, if from the same speaker (rather than questions by an interlocutor), boost performance.We also find that the semantic information in language models such as BERT contributes to the performance.However, BERT alone does not capture the implicit knowledge of deception cues: its contribution is conditional on the concurrent use of attention to learn cues from BERT's representations. Tommaso Fornaciari, Federico Bianchi 0001, Massimo Poesio, Dirk Hovy |
EACL | 4 |
| 2021 | Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task LearningabstractTommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, Massimo Poesio. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, Massimo Poesio |
NAACL-HLT | 5 |
| 2021 | The Importance of Modeling Social Factors of Language: Theory and PracticeabstractNatural language processing (NLP) applications are now more powerful and ubiquitous than ever before.With rapidly developing (neural) models and ever-more available data, current NLP models have access to more information than any human speaker during their life.Still, it would be hard to argue that NLP models have reached human-level capacity.In this position paper, we argue that the reason for the current limitations is a focus on information content while ignoring language's social factors.We show that current NLP systems systematically break down when faced with interpreting the social factors of language.This limits applications to a subset of information-related tasks and prevents NLP from reaching human-level performance.At the same time, systems that incorporate even a minimum of social factors already show remarkable improvements.We formalize a taxonomy of seven social factors based on linguistic theory and exemplify current failures and emerging successes for each of them.We suggest that the NLP community address social factors to get closer to the goal of humanlike language understanding. Dirk Hovy, Diyi Yang |
NAACL-HLT | 1 |
| 2021 | HONEST: Measuring Hurtful Sentence Completion in Language ModelsabstractLanguage models have revolutionized the field of NLP.However, language models capture and proliferate hurtful stereotypes, especially in text generation.Our results show that 4.3% of the time, language models complete a sentence with a hurtful word.These cases are not random, but follow language and genderspecific patterns.We propose a score to measure hurtful sentence completions in language models (HONEST).It uses a systematic template-and lexicon-based bias evaluation methodology for six languages.Our findings suggest that these models replicate and amplify deep-seated societal stereotypes about gender roles.Sentence completions refer to sexual promiscuity when the target is female in 9% of the time, and in 4% to homosexuality when the target is male.The results raise questions about the use of these models in production settings. Debora Nozza, Federico Bianchi 0001, Dirk Hovy |
NAACL-HLT | 3 |
| 2021 | Learning from Disagreement: A SurveyabstractMany tasks in Natural Language Processing (NLP) and Computer Vision (CV) offer evidence that humans disagree, from objective tasks such as part-of-speech tagging to more subjective tasks such as classifying an image or deciding whether a proposition follows from certain premises. While most learning in artificial intelligence (AI) still relies on the assumption that a single (gold) interpretation exists for each item, a growing body of research aims to develop learning methods that do not rely on this assumption. In this survey, we review the evidence for disagreements on NLP and CV tasks, focusing on tasks for which substantial datasets containing this information have been created. We discuss the most popular approaches to training models from datasets containing multiple judgments potentially in disagreement. We systematically compare these different approaches by training them with each of the available datasets, considering several ways to evaluate the resulting models. Finally, we discuss the results in depth, focusing on four key research questions, and assess how the type of evaluation and the characteristics of a dataset determine the answers to these questions. Our results suggest, first of all, that even if we abandon the assumption of a gold standard, it is still essential to reach a consensus on how to evaluate models. This is because the relative performance of the various training methods is critically affected by the chosen form of evaluation. Secondly, we observed a strong dataset effect. With substantial datasets, providing many judgments by high-quality coders for each item, training directly with soft labels achieved better results than training from aggregated or even gold labels. This result holds for both hard and soft evaluation. But when the above conditions do not hold, leveraging both gold and soft labels generally achieved the best results in the hard evaluation. All datasets and models employed in this paper are freely available as supplementary materials. Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio |
J. Artif. Intell. Res. | 3 |
| 2020 | "You Sound Just Like Your Father" Commercial Machine Translation Systems Include Stylistic BiasesabstractThe main goal of machine translation has been to convey the correct content. Stylistic considerations have been at best secondary. We show that as a consequence, the output of three commercial machine translation systems (Bing, DeepL, Google) make demographically diverse samples from five languages "sound" older and more male than the original. Our findings suggest that translation models reflect demographic bias in the training data. This opens up interesting new research avenues in machine translation to take stylistic considerations into account. Dirk Hovy, Federico Bianchi 0001, Tommaso Fornaciari |
ACL | 1 |
| 2020 | Predictive Biases in Natural Language Processing Models: A Conceptual Framework and OverviewabstractAn increasing number of natural language processing papers address the effect of bias on predictions, introducing mitigation techniques at different parts of the standard NLP pipeline (data and models).However, these works have been conducted individually, without a unifying framework to organize efforts within the field.This situation leads to repetitive approaches, and focuses overly on bias symptoms/effects, rather than on their origins, which could limit the development of effective countermeasures.In this paper, we propose a unifying predictive bias framework for NLP.We summarize the NLP literature and suggest general mathematical definitions of predictive bias.We differentiate two consequences of bias: outcome disparities and error disparities, as well as four potential origins of biases: label bias, selection bias, model overamplification, and semantic bias.Our framework serves as an overview of predictive bias in NLP, integrating existing work into a single structure, and providing a conceptual baseline for improved frameworks. Deven Shah, H. Andrew Schwartz, Dirk Hovy |
ACL | 3 |
| 2020 | A Case for Soft Loss FunctionsabstractRecently, Peterson et al. provided evidence of the benefits of using probabilistic soft labels generated from crowd annotations for training a computer vision model, showing that using such labels maximizes performance of the models over unseen data. In this paper, we generalize these results by showing that training with soft labels is an effective method for using crowd annotations in several other ai tasks besides the one studied by Peterson et al., and also when their performance is compared with that of state-of-the-art methods for learning from crowdsourced data. Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio |
HCOMP | 3 |
| 2018 | Improving Author Attribute Prediction by Retrofitting Linguistic Representations with HomophilyabstractMost text-classification approaches represent the input based on textual features, either feature-based or continuous.However, this ignores strong non-linguistic similarities like homophily: people within a demographic group use language more similar to each other than to non-group members.We use homophily cues to retrofit text-based author representations with non-linguistic information, and introduce a trade-off parameter.This approach increases in-class similarity between authors, and improves classification performance by making classes more linearly separable.We evaluate the effect of our method on two authorattribute prediction tasks with various trainingset sizes and parameter settings.We find that our method can significantly improve classification performance, especially when the number of labels is large and limited labeled data is available.It is potentially applicable as preprocessing step to any text-classification task. Dirk Hovy, Tommaso Fornaciari |
EMNLP | 1 |
| 2018 | Capturing Regional Variation with Distributed Place Representations and Geographic RetrofittingabstractDialects are one of the main drivers of language variation, a major challenge for natural language processing tools.In most languages, dialects exist along a continuum, and are commonly discretized by combining the extent of several preselected linguistic variables.However, the selection of these variables is theorydriven and itself insensitive to change.We use Doc2Vec on a corpus of 16.8M anonymous online posts in the German-speaking area to learn continuous document representations of cities.These representations capture continuous regional linguistic distinctions, and can serve as input to downstream NLP tasks sensitive to regional variation.By incorporating geographic information via retrofitting and agglomerative clustering with structure, we recover dialect areas at various levels of granularity.Evaluating these clusters against an existing dialect map, we achieve a match of up to 0.77 V-score (harmonic mean of cluster completeness and homogeneity).Our results show that representation learning with retrofitting offers a robust general method to automatically expose dialectal differences and regional variation at a finer granularity than was previously possible. Dirk Hovy, Christoph Purschke |
EMNLP | 1 |
| 2018 | Predicting News Headline Popularity with Syntactic and Semantic Knowledge Using Multi-Task LearningabstractNewspapers need to attract readers with headlines, anticipating their readers' preferences.These preferences rely on topical, structural, and lexical factors.We model each of these factors in a multi-task GRU network to predict headline popularity.We find that pre-trained word embeddings provide significant improvements over untrained embeddings, as do the combination of two auxiliary tasks, newssection prediction and part-of-speech tagging.However, we also find that performance is very similar to that of a simple Logistic Regression model over character n-grams.Feature analysis reveals structural patterns of headline popularity, including the use of forward-looking deictic expressions and second person pronouns. Sotiris Lamprinidis, Daniel Hardt, Dirk Hovy |
EMNLP | 3 |
| 2018 | Comparing Bayesian Models of AnnotationabstractThe analysis of crowdsourced annotations in natural language processing is concerned with identifying (1) gold standard labels, (2) annotator accuracies and biases, and (3) item difficulties and error patterns. Traditionally, majority voting was used for 1, and coefficients of agreement for 2 and 3. Lately, model-based analysis of corpus annotations have proven better at all three tasks. But there has been relatively little work comparing them on the same datasets. This paper aims to fill this gap by analyzing six models of annotation, covering different approaches to annotator ability, item difficulty, and parameter pooling (tying) across annotators and items. We evaluate these models along four aspects: comparison to gold labels, predictive accuracy for new annotations, annotator characterization, and item difficulty, using four datasets with varying degrees of noise in the form of random (spammy) annotators. We conclude with guidelines for model selection, application, and implementation. Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, Massimo Poesio |
Trans. Assoc. Comput. Linguistics | 4 |
| 2017 | Multitask Learning for Mental Health Conditions with Limited Social Media DataabstractWe introduce initial groundwork for estimating suicide risk and mental health in a deep learning framework.By modeling multiple conditions, the system learns to make predictions about suicide risk and mental health at a low false positive rate.Conditions are modeled as tasks in a multitask learning (MTL) framework, with gender prediction as an additional auxiliary task.We demonstrate the effectiveness of multi-task learning by comparison to a well-tuned single-task baseline with the same number of parameters.Our best MTL model predicts potential suicide attempt, as well as the presence of atypical mental health, with AUC > 0.8.We also find additional large improvements using multi-task learning on mental health tasks with limited training data.* Now at Google Research. 1 https://www.nami.org/Learn-More/Mental-Health-Conditions/Related-Conditions/Suicide#sthash.dMAhrKTU.dpuf 2 Communication with clinicians at the 2016 JSALT workshop (Hollingshead, 2016). Adrian Benton, Margaret Mitchell, Dirk Hovy |
EACL (1) | 3 |
| 2016 | Exploring Language Variation Across Europe - A Web-based Tool for Computational Sociolinguistics
Dirk Hovy, Anders Johannsen |
LREC | 1 |
| 2016 | Learning a POS tagger for AAVE-like languageabstractPart-of-speech (POS) taggers trained on newswire perform much worse on domains such as subtitles, lyrics, or tweets.In addition, these domains are also heterogeneous, e.g., with respect to registers and dialects.In this paper, we consider the problem of learning a POS tagger for subtitles, lyrics, and tweets associated with African-American Vernacular English (AAVE).We learn from a mixture of randomly sampled and manually annotated Twitter data and unlabeled data, which we automatically and partially label using mined tag dictionaries.Our POS tagger obtains a tagging accuracy of 89% on subtitles, 85% on lyrics, and 83% on tweets, with up to 55% error reductions over a state-of-the-art newswire POS tagger, and 15-25% error reductions over a state-of-the-art Twitter POS tagger. Anna Jørgensen, Dirk Hovy, Anders Søgaard |
HLT-NAACL | 2 |
| 2015 | Demographic Factors Improve Classification PerformanceabstractExtra-linguistic factors influence language use, and are accounted for by speakers and listeners.Most natural language processing (NLP) tasks to date, however, treat language as uniform.This assumption can harm performance.We investigate the effect of including demographic information on performance in a variety of text-classification tasks.We find that by including age or gender information, we consistently and significantly improve performance over demographic-agnostic models.These results hold across three text-classification tasks in five languages. Dirk Hovy |
ACL (1) | 1 |
| 2015 | Cross-lingual syntactic variation over age and genderabstractMost computational sociolinguistics studies have focused on phonological and lexical variation.We present the first large-scale study of syntactic variation among demographic groups (age and gender) across several languages.We harvest data from online user-review sites and parse it with universal dependencies.We show that several age and gender-specific variations hold across languages, for example that women are more likely to use VP conjunctions. Anders Johannsen, Dirk Hovy, Anders Søgaard |
CoNLL | 2 |
| 2015 | The Rating Game: Sentiment Rating Reproducibility from TextabstractSentiment analysis models often use ratings as labels, assuming that these ratings reflect the sentiment of the accompanying text.We investigate (i) whether human readers can infer ratings from review text, (ii) how human performance compares to a regression model, and (iii) whether model performance is affected by the rating "source" (i.e.original author vs. annotator).We collect IMDb movie reviews with author-provided ratings, and have them re-annotated by crowdsourced and trained annotators.Annotators reproduce the original ratings better than a model, but are still far off in more than 5% of the cases.Models trained on annotator-labels outperform those trained on author-labels, questioning the usefulness of author-rated reviews as training data for sentiment analysis. Lasse Borgholt, Peter Simonsen, Dirk Hovy |
EMNLP | 3 |
| 2015 | Mining for unambiguous instances to adapt part-of-speech taggers to new domainsabstractDirk Hovy, Barbara Plank, Héctor Martínez Alonso, Anders Søgaard. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Dirk Hovy, Barbara Plank, Héctor Martínez Alonso, Anders Søgaard |
HLT-NAACL | 1 |
| 2015 | User Review Sites as a Resource for Large-Scale Sociolinguistic StudiesabstractSociolinguistic studies investigate the relation between language and extra-linguistic variables. This requires both representative text data and the associated socio-economic meta-data of the subjects. Traditionally, sociolinguistic studies use small samples of hand-curated data and meta-data. This can lead to exaggerated or false conclusions. Using social media data offers a large-scale source of language data, but usually lacks reliable socio-economic meta-data. Our research aims to remedy both problems by exploring a large new data source, international review websites with user profiles. They provide more text data than manually collected studies, and more meta-data than most available social media text. We describe the data and present various pilot studies, illustrating the usefulness of this resource for sociolinguistic studies. Our approach can help generate new research hypotheses based on data-driven findings across several countries and languages. Dirk Hovy, Anders Johannsen, Anders Søgaard |
WWW | 1 |
| 2014 | Adapting taggers to Twitter with not-so-distant supervision
Barbara Plank, Dirk Hovy, Ryan T. McDonald, Anders Søgaard |
COLING | 2 |
| 2014 | What's in a p-value in NLP?abstractIn NLP, we need to document that our pro-posed methods perform significantly bet-ter with respect to standard metrics than previous approaches, typically by re-porting p-values obtained by rank- or randomization-based tests. We show that significance results following current re-search standards are unreliable and, in ad-dition, very sensitive to sample size, co-variates such as sentence length, as well as to the existence of multiple metrics. We estimate that under the assumption of per-fect metrics and unbiased data, we need a significance cut-off at ⇠0.0025 to reduce the risk of false positive results to <5%. Since in practice we often have consider-able selection bias and poor metrics, this, however, will not do alone. 1 Anders Søgaard, Anders Johannsen, Barbara Plank, Dirk Hovy, Héctor Martínez Alonso |
CoNLL | 4 |
| 2014 | Learning part-of-speech taggers with inter-annotator agreement lossabstractIn natural language processing (NLP) an-notation projects, we use inter-annotator agreement measures and annotation guide-lines to ensure consistent annotations. However, annotation guidelines often make linguistically debatable and even somewhat arbitrary decisions, and inter-annotator agreement is often less than perfect. While annotation projects usu-ally specify how to deal with linguisti-cally debatable phenomena, annotator dis-agreements typically still stem from these “hard ” cases. This indicates that some er-rors are more debatable than others. In this paper, we use small samples of doubly-annotated part-of-speech (POS) data for Twitter to estimate annotation reliability and show how those metrics of likely inter-annotator agreement can be implemented in the loss functions of POS taggers. We find that these cost-sensitive algorithms perform better across annotation projects and, more surprisingly, even on data an-notated according to the same guidelines. Finally, we show that POS tagging mod-els sensitive to inter-annotator agreement perform better on the downstream task of chunking. 1 Barbara Plank, Dirk Hovy, Anders Søgaard |
EACL | 2 |
| 2014 | Crowdsourcing and annotating NER for Twitter #drift
Hege Fromreide, Dirk Hovy, Anders Søgaard |
LREC | 2 |
| 2014 | When POS data sets don't add up: Combatting sample bias
Dirk Hovy, Barbara Plank, Anders Søgaard |
LREC | 1 |
| 2014 | Augmenting English Adjective Senses with Supersenses
Yulia Tsvetkov, Nathan Schneider 0001, Dirk Hovy, Archna Bhatia, Manaal Faruqui, Chris Dyer |
LREC | 3 |
| 2013 | A Walk-Based Semantically Enriched Tree Kernel Over Distributed Word RepresentationsabstractIn this paper, we propose a walk-based graph kernel that generalizes the notion of treekernels to continuous spaces.Our proposed approach subsumes a general framework for word-similarity, and in particular, provides a flexible way to incorporate distributed representations.Using vector representations, such an approach captures both distributional semantic similarities among words as well as the structural relations between them (encoded as the structure of the parse tree).We show an efficient formulation to compute this kernel using simple matrix operations.We present our results on three diverse NLP tasks, showing state-of-the-art results. Dirk Hovy, Eduard H. Hovy |
EMNLP | 2 |
| 2013 | Analysis and modeling of "focus" in contextabstractThis paper uses a crowd-sourced definition of a speech phe-nomenon we have called “focus”. Given sentences, text and speech, in isolation and in context, we asked annotators to iden-tify what we term the “focus ” word. We present their consis-tency in identifying the focused word, when presented with text or speech stimuli. We then build models to show how well we predict that focus word from lexical (and higher) level features. Also, using spectral and prosodic information, we show the dif-ferences in these focus words when spoken with and without context. Finally, we show how we can improve speech synthe-sis of these utterances given focus information. Dirk Hovy, Gopala Krishna Anumanchipalli, Alok Parlikar, Caroline Vaughn, Adam C. Lammert, Eduard H. Hovy, Alan W. Black |
INTERSPEECH | 1 |
| 2013 | Learning Whom to Trust with MACE
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, Eduard H. Hovy |
HLT-NAACL | 1 |
| 2012 | When Did that Happen? - Linking Events and Relations to Timestamps
Dirk Hovy, James Fan, Alfio Massimiliano Gliozzo, Siddharth Patwardhan, Christopher A. Welty |
EACL | 1 |
| 2011 | Unsupervised Discovery of Domain-Specific Knowledge from Text
Dirk Hovy, Chunliang Zhang, Eduard H. Hovy, Anselmo Peñas |
ACL | 1 |