Anne Lauscher

dblp:209/6857 · DBLP profile ↗
← Back
38ranked-venue papers
10as first author
32since 2021 · last 2026
0000-0001-8590-9827ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 10 first-author · 32 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Issue Detection and Category Classification in Domain-Specific Technical Logbooks
Afshin Karimi, Ingmar Hartl, Henrik Tünnermann, Anne Lauscher
LREC4
2026 Detecting Hallucinations in Authentic LLM-Human Interactions
Yujie Ren, Niklas Gruhlke, Anne Lauscher
LREC3
2026 Aligned Probing: Relating Toxic Behavior and Model Internals
abstract
Abstract Warning: This paper contains offensive text. We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (internals). Using this framework, we examine over 20 OLMo, Llama, and Mistral models, bridging behavioral and internal perspectives for toxicity for the first time. Our results show that LMs strongly encode information about the toxicity level of inputs and subsequent outputs, particularly in lower layers. Focusing on how unique LMs differ offers both correlative and causal evidence that they generate less toxic output when strongly encoding information about the input toxicity. We also highlight the heterogeneity of toxicity, as model behavior and internals vary across unique attributes such as Threat. Finally, four case studies analyzing detoxification, multi-prompt evaluations, model quantization, and pre-training dynamics underline the practical impact of aligned probing with further concrete insights. Our findings contribute to a more holistic understanding of LMs, both within and beyond the context of toxicity. alignedprobing.github.io
Andreas Waldis, Vagrant Gautam, Anne Lauscher, Dietrich Klakow, Iryna Gurevych
Trans. Assoc. Comput. Linguistics3
2025 Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
abstract
Gregor Geigle, Florian Schneider, Carolin Holtermann, Chris Biemann, Radu Timofte, Anne Lauscher, Goran Glavaš. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Gregor Geigle, Florian Schneider 0001, Carolin Holtermann, Chris Biemann, Radu Timofte, Anne Lauscher, Goran Glavas
ACL (1)6
2025 Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
abstract
Reasoning over time and space is essential for understanding our world.However, the abilities of language models in this area are largely unexplored as previous work has tested their abilities for logical reasoning in terms of time and space in isolation or only in simple or artificial environments.In this paper, we present the first evaluation of the ability of language models to jointly reason over time and space.To enable our analysis, we create GEOTEMP, a dataset of 320k prompts covering 289 cities in 217 countries and 37 time zones.Using GEOTEMP, we evaluate eight open chat models from three model families for different combinations of temporal and geographic knowledge.We find that most models perform well on reasoning tasks involving only temporal knowledge and that overall performance improves with scale.However, performance remains poor in tasks that require connecting temporal and geographical information.We do not find clear correlations of performance with specific geographic regions.Instead, we find a significant performance increase for location names with low model perplexity, suggesting their repeated occurrence during model training.We further demonstrate that model performance is heavily influenced by prompt formulation -a direct injection of geographical knowledge leads to performance gains, whereas, surprisingly, techniques like chain-of-thought prompting decrease performance on simpler tasks. 1
Carolin Holtermann, Paul Röttger, Anne Lauscher
ACL (1)3
2025 LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
abstract
Peer review is a cornerstone of quality control in scientific publishing.With the increasing workload, the unintended use of 'quick' heuristics, referred to as lazy thinking, has emerged as a recurring issue compromising review quality.Automated methods to detect such heuristics can help improve the peer-reviewing process.However, there is limited NLP research on this issue, and no real-world dataset exists to support the development of detection tools.This work introduces LAZYREVIEW, a dataset of peer-review sentences annotated with finegrained lazy thinking categories.Our analysis reveals that Large Language Models (LLMs) struggle to detect these instances in a zeroshot setting.However, instruction-based finetuning on our dataset significantly boosts performance by 10-20 performance points, highlighting the importance of high-quality training data.Furthermore, a controlled experiment demonstrates that reviews revised with lazy thinking feedback are more comprehensive and actionable than those written without such feedback.We will release our dataset and the enhanced guidelines that can be used to train junior reviewers in the community.1 Heuristics Description Example review segmentsThe results are not surprising Many findings seem obvious in retrospect, but this does not mean that the community is already aware of them and can use them as building blocks for future work.
Sukannya Purkayastha, Zhuang Li 0001, Anne Lauscher, Lizhen Qu, Iryna Gurevych
ACL (1)3
2025 Large Language Models Discriminate Against Speakers of German Dialects
abstract
Dialects represent a significant component of human culture and are found across all regions of the world.In Germany, more than 40% of the population speaks a regional dialect (Adler and Hansen, 2022).However, despite cultural importance, individuals speaking dialects often face negative societal stereotypes.We examine whether such stereotypes are mirrored by large language models (LLMs).We draw on the sociolinguistic literature on dialect perception to analyze traits commonly associated with dialect speakers.Based on these traits, we assess the dialect naming bias and dialect usage bias expressed by LLMs in two tasks: an association task and a decision task.To assess a model's dialect usage bias, we construct a novel evaluation corpus that pairs sentences from seven regional German dialects (e.g., Alemannic and Bavarian) with their standard German counterparts.We find that: (1) in the association task, all evaluated LLMs exhibit significant dialect naming and dialect usage bias against German dialect speakers, reflected in negative adjective associations; (2) all models reproduce these dialect naming and dialect usage biases in their decision making; and (3) contrary to prior work showing minimal bias with explicit demographic mentions, we find that explicitly labeling linguistic demographics-German dialect speakers-amplifies bias more than implicit cues like dialect usage.* Equal contribution. 1 The literature disagrees on an exact definition; we give more information in Appendix A.1.
Minh Duc Bui, Carolin Holtermann, Valentin Hofmann, Anne Lauscher, Katharina von der Wense
EMNLP4
2025 How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
abstract
In the age of misinformation, hallucinationthe tendency of Large Language Models (LLMs) to generate non-factual or unfaithful responses-represents the main risk for their global utility.Despite LLMs becoming increasingly multilingual, the vast majority of research on detecting and quantifying LLM hallucination are (a) English-centric and (b) focus on machine translation (MT) and summarization, tasks that are less common in realistic settings than open information seeking.In contrast, we aim to quantify the extent of LLM hallucination across languages in knowledge-intensive longform question answering (LFQA).To this end, we train a multilingual hallucination detection model and conduct a large-scale study across 30 languages and 6 open-source LLM families.We start from an English hallucination detection dataset and rely on MT to translate-train a detection model.We also manually annotate gold data for five high-resource languages; we then demonstrate, for these languages, that the estimates of hallucination rates are similar between silver (LLM-generated) and gold test sets, validating the use of silver data for estimating hallucination rates for other languages.For the final rates estimation, we build opendomain QA dataset for 30 languages with LLMgenerated prompts and Wikipedia articles as references.Our analysis shows that LLMs, in absolute terms, hallucinate more tokens in highresource languages due to longer responses, but that the actual hallucination rates (i.e., normalized for length) seems uncorrelated with the sizes of languages' digital footprints.We also find that smaller LLMs hallucinate more, and significantly, LLMs with broader language support display higher hallucination rates.
Saad Obaid ul Islam, Anne Lauscher, Goran Glavas
EMNLP2
2025 Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTE
abstract
Beatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou, Janiça Hackenbuchner, Anne Lauscher, Matteo Negri, Andrea Piergentili, Manjinder Thind, Luisa Bentivogli. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Beatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou, Janiça Hackenbuchner, Anne Lauscher, Matteo Negri, Andrea Piergentili, Manjinder Thind, Luisa Bentivogli
EMNLP6
2025 Multi³Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models
abstract
Minh Duc Bui, Katharina Von Der Wense, Anne Lauscher. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Minh Duc Bui, Katharina von der Wense, Anne Lauscher
NAACL (Long Papers)3
2025 Ethical Concern Identification in NLP: A Corpus of ACL Anthology Ethics Statements
abstract
Antonia Karamolegkou, Sandrine Schiller Hansen, Ariadni Christopoulou, Filippos Stamatiou, Anne Lauscher, Anders Søgaard. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Antonia Karamolegkou, Sandrine Schiller Hansen, Ariadni Christopoulou, Filippos Stamatiou, Anne Lauscher, Anders Søgaard
NAACL (Long Papers)5
2025 SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models
abstract
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna-Adriana Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L. Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir R. Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh D. Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat
NAACL (Long Papers)17
2024 The Echoes of Multilinguality: Tracing Cultural Value Shifts during Language Model Fine-tuning
abstract
Texts written in different languages reflect different culturally-dependent beliefs of their writers.Thus, we expect multilingual LMs (MLMs), that are jointly trained on a concatenation of text in multiple languages, to encode different cultural values for each language.Yet, as the 'multilinguality' of these LMs is driven by cross-lingual sharing, we also have reason to belief that cultural values bleed over from one language into another.This limits the use of MLMs in practice, as apart from being proficient in generating text in multiple languages, creating language technology that can serve a community also requires the output of LMs to be sensitive to their biases (Naous et al., 2023).Yet, little is known about how cultural values emerge and evolve in MLMs (Hershcovich et al., 2022a).We are the first to study how languages can exert influence on the cultural values encoded for different test languages, by studying how such values are revised during fine-tuning.Focusing on the finetuning stage allows us to study the interplay between value shifts when exposed to new linguistic experience from different data sources and languages.Lastly, we use a training data attribution method to find patterns in the finetuning examples, and the languages that they come from, that tend to instigate value shifts.
Rochelle Choenni, Anne Lauscher, Ekaterina Shutova
ACL (1)2
2024 Decoding Multilingual Moral Preferences: Unveiling LLM's Biases through the Moral Machine Experiment
abstract
Large language models (LLMs) increasingly find their way into the most diverse areas of our everyday lives. They indirectly influence people's decisions or opinions through their daily use. Therefore, understanding how and which moral judgements these LLMs make is crucial. However, morality is not universal and depends on the cultural background. This raises the question of whether these cultural preferences are also reflected in LLMs when prompted in different languages or whether moral decision-making is consistent across different languages. So far, most research has focused on investigating the inherent values of LLMs in English. While a few works conduct multilingual analyses of moral bias in LLMs in a multilingual setting, these analyses do not go beyond atomic actions. To the best of our knowledge, a multilingual analysis of moral bias in dilemmas has not yet been conducted. To address this, our paper builds on the moral machine experiment (MME) to investigate the moral preferences of five LLMs, Falcon, Gemini, Llama, GPT, and MPT, in a multilingual setting and compares them with the preferences collected from humans belonging to different cultures. To accomplish this, we generate 6500 scenarios of the MME and prompt the models in ten languages on which action to take. Our analysis reveals that all LLMs inhibit different moral biases to some degree and that they not only differ from the human preferences but also across multiple languages within the models themselves. Moreover, we find that almost all models, particularly Llama 3, divert greatly from human values and, for instance, prefer saving fewer people over saving more.
Karina Vida, Fabian Damken, Anne Lauscher
AIES (1)3
2024 Argument Quality Assessment in the Age of Instruction-Following Large Language Models
abstract
The computational treatment of arguments on controversial issues has been subject to extensive NLP research, due to its envisioned impact on opinion formation, decision making, writing education, and the like. A critical task in any such application is the assessment of an argument’s quality - but it is also particularly challenging. In this position paper, we start from a brief survey of argument quality research, where we identify the diversity of quality notions and the subjectiveness of their perception as the main hurdles towards substantial progress on argument quality assessment. We argue that the capabilities of instruction-following large language models (LLMs) to leverage knowledge across contexts enable a much more reliable assessment. Rather than just fine-tuning LLMs towards leaderboard chasing on assessment tasks, they need to be instructed systematically with argumentation theories and scenarios as well as with ways to solve argument-related problems. We discuss the real-world opportunities and ethical issues emerging thereby.
Henning Wachsmuth, Gabriella Lapesa, Elena Cabrio, Anne Lauscher, Joonsuk Park, Eva Maria Vecchi, Serena Villata, Timon Ziegenbein
LREC/COLING4
2024 Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting
abstract
Annotators' sociodemographic backgrounds (i.e., the individual compositions of their gender, age, educational background, etc.) have a strong impact on their decisions when working on subjective NLP tasks, such as toxic language detection.Often, heterogeneous backgrounds result in high disagreements.To model this variation, recent work has explored sociodemographic prompting, a technique, which steers the output of prompt-based models towards answers that humans with specific sociodemographic profiles would give.However, the available NLP literature disagrees on the efficacy of this technique -it remains unclear for which tasks and scenarios it can help, and the role of the individual factors in sociodemographic prompting is still unexplored.We address this research gap by presenting the largest and most comprehensive study of sociodemographic prompting today.We analyze its influence on model sensitivity, performance and robustness across seven datasets and six instruction-tuned model families.We show that sociodemographic information affects model predictions and can be beneficial for improving zero-shot learning in subjective NLP tasks.However, its outcomes largely vary for different model types, sizes, and datasets, and are subject to large variance with regards to prompt formulations.Most importantly, our results show that sociodemographic prompting should be used with care for sensitive applications, such as toxicity annotation or when studying LLM alignment.1
Tilman Beck, Hendrik Schuff, Anne Lauscher, Iryna Gurevych
EACL (1)3
2024 GeFMT: Gender-Fair Language in German Machine Translation
abstract
Research on gender bias in Machine Translation (MT) predominantly focuses on binary gender or few languages. In this project, we investigate the ability of commercial MT systems and neural models to translate using gender-fair language (GFL) from English into German. We enrich a community-created GFL dictionary, and sample multi-sentence test instances from encyclopedic text and parliamentary speeches. We translate our resources with different MT systems and open-weights models. We also plan to post-edit biased outputs with professionals and share them publicly. The outcome will constitute a new resource for automatic evaluation and modeling gender-fair EN-DE MT.
Manuel Lardelli, Anne Lauscher, Giuseppe Attanasio
EAMT (2)2
2024 Local Contrastive Editing of Gender Stereotypes
abstract
Stereotypical bias encoded in language models (LMs) poses a threat to safe language technology, yet our understanding of how bias manifests in the parameters of LMs remains incomplete.We introduce local contrastive editing that enables the localization and editing of a subset of weights in a target model in relation to a reference model.We deploy this approach to identify and modify subsets of weights that are associated with gender stereotypes in LMs.Through a series of experiments, we demonstrate that local contrastive editing can precisely localize and control a small subset (<0.5%) of weights that encode gender bias.Our work (i) advances our understanding of how stereotypical biases can manifest in the parameter space of LMs and (ii) opens up new avenues for developing parameter-efficient strategies for controlling model properties in a contrastive manner.
Marlene Lutz, Rochelle Choenni, Markus Strohmaier, Anne Lauscher
EMNLP4
2024 The Lou Dataset - Exploring the Impact of Gender-Fair Language in German Text Classification
abstract
Gender-fair language, an evolving linguistic variation in German, fosters inclusion by addressing all genders or using neutral forms. However, there is a notable lack of resources to assess the impact of this language shift on language models (LMs) might not been trained on examples of this variation. Addressing this gap, we present Lou, the first dataset providing high-quality reformulations for German text classification covering seven tasks, like stance detection and toxicity classification. We evaluate 16 mono- and multi-lingual LMs and find substantial label flips, reduced prediction certainty, and significantly altered attention patterns. However, existing evaluations remain valid, as LM rankings are consistent across original and reformulated instances. Our study provides initial insights into the impact of gender-fair language on classification for German. However, these findings are likely transferable to other languages, as we found consistent patterns in multi-lingual and English LMs.
Andreas Waldis, Joel Birrer, Anne Lauscher, Iryna Gurevych
EMNLP3
2024 AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
abstract
Generating bitmap graphics from text has gained considerable attention, yet for scientific figures, vector graphics are often preferred. Given that vector graphics are typically encoded using low-level graphics primitives, generating them directly is difficult. To address this, we propose the use of TikZ, a well-known abstract graphics language that can be compiled to vector graphics, as an intermediate representation of scientific figures. TikZ offers human-oriented, high-level commands, thereby facilitating conditional language modeling with any large language model. To this end, we introduce DaTikZ the first large-scale TikZ dataset, consisting of 120k TikZ drawings aligned with captions. We fine-tune LLaMA on DaTikZ, as well as our new model CLiMA, which augments LLaMA with multimodal CLIP embeddings. In both human and automatic evaluation, CLiMA and LLaMA outperform commercial GPT-4 and Claude 2 in terms of similarity to human-created figures, with CLiMA additionally improving text-image alignment. Our detailed analysis shows that all models generalize well and are not susceptible to memorization. GPT-4 and Claude 2, however, tend to generate more simplistic figures compared to both humans and our models. We make our framework, AutomaTikZ, along with model weights and datasets, publicly available.
Jonas Belouadi, Anne Lauscher, Steffen Eger
ICLR2
2024 Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
abstract
Abstract Robust, faithful, and harm-free pronoun use for individuals is an important goal for language model development as their use increases, but prior work tends to study only one or two of these characteristics at a time. To measure progress towards the combined goal, we introduce the task of pronoun fidelity: Given a context introducing a co-referring entity and pronoun, the task is to reuse the correct pronoun later. We present RUFF, a carefully designed dataset of over 5 million instances to measure robust pronoun fidelity in English, and we evaluate 37 model variants from nine popular families, across architectures (encoder-only, decoder-only, and encoder-decoder) and scales (11M-70B parameters). When an individual is introduced with a pronoun, models can mostly faithfully reuse this pronoun in the next sentence, but they are significantly worse with she/her/her, singular they, and neopronouns. Moreover, models are easily distracted by non-adversarial sentences discussing other people; even one sentence with a distractor pronoun causes accuracy to drop on average by 34 percentage points. Our results show that pronoun fidelity is not robust, in a simple, naturalistic setting where humans achieve nearly 100% accuracy. We encourage researchers to bridge the gaps we find and to carefully evaluate reasoning in settings where superficial repetition might inflate perceptions of model performance.
Vagrant Gautam, Eileen Bingert, Anne Lauscher, Dietrich Klakow
Trans. Assoc. Comput. Linguistics4
2023 What about "em"? How Commercial Machine Translation Fails to Handle (Neo-)Pronouns
abstract
As 3rd-person pronoun usage shifts to include novel forms, e.g., neopronouns, we need more research on identity-inclusive NLP.Exclusion is particularly harmful in one of the most popular NLP applications, machine translation (MT).Wrong pronoun translations can discriminate against marginalized groups, e.g., non-binary individuals (Dev et al., 2021).In this "reality check", we study how three commercial MT systems translate 3rd-person pronouns.Concretely, we compare the translations of gendered vs. gender-neutral pronouns from English to five other languages (Danish, Farsi, French, German, Italian), and vice versa, from Danish to English.Our error analysis shows that the presence of a gender-neutral pronoun often leads to grammatical and semantic translation errors.Similarly, gender neutrality is often not preserved.By surveying the opinions of affected native speakers from diverse languages, we provide recommendations to address the issue in future MT research.
Anne Lauscher, Debora Nozza, Ehm Miltersen, Archie Crowley, Dirk Hovy
ACL (1)1
2023 A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine Translation
abstract
Recent instruction fine-tuned models can solve multiple NLP tasks when prompted to do so, with machine translation (MT) being a prominent use case.However, current research often focuses on standard performance benchmarks, leaving compelling fairness and ethical considerations behind.In MT, this might lead to misgendered translations, resulting, among other harms, in the perpetuation of stereotypes and prejudices.In this work, we address this gap by investigating whether and to what extent such models exhibit gender bias in machine translation and how we can mitigate it.Concretely, we compute established gender bias metrics on the WinoMT corpus from English to German and Spanish.We discover that IFT models default to male-inflected translations, even disregarding female occupational stereotypes.Next, using interpretability methods, we unveil that models systematically overlook the pronoun indicating the gender of a target occupation in misgendered translations.Finally, based on this finding, we propose an easy-to-implement and effective bias mitigation solution based on fewshot learning that leads to significantly fairer translations.1
Giuseppe Attanasio, Flor Miriam Plaza del Arco, Debora Nozza, Anne Lauscher
EMNLP4
2023 Exploring Jiu-Jitsu Argumentation for Writing Peer Review Rebuttals
abstract
In many domains of argumentation, people's arguments are driven by so-called attitude roots, i.e., underlying beliefs and world views, and their corresponding attitude themes.Given the strength of these latent drivers of arguments, recent work in psychology suggests that instead of directly countering surface-level reasoning (e.g., falsifying given premises), one should follow an argumentation style inspired by the Jiu-Jitsu "soft" combat system (Hornsey and Fielding, 2017): first, identify an arguer's attitude roots and themes, and then choose a prototypical rebuttal that is aligned with those drivers instead of invalidating those.In this work, we are the first to explore Jiu-Jitsu argumentation for peer review by proposing the novel task of attitude and theme-guided rebuttal generation.To this end, we enrich an existing dataset for discourse structure in peer reviews with attitude roots, attitude themes, and canonical rebuttals.To facilitate this process, we recast established annotation concepts from the domain of peer reviews (e.g., aspects a review sentence is relating to) and train domain-specific models.We then propose strong rebuttal generation strategies, which we benchmark on our novel dataset for the task of end-to-end attitude and themeguided rebuttal generation and two subtasks.1
Sukannya Purkayastha, Anne Lauscher, Iryna Gurevych
EMNLP2
2022 Fair and Argumentative Language Modeling for Computational Argumentation
abstract
Although much work in NLP has focused on measuring and mitigating stereotypical bias in semantic spaces, research addressing bias in computational argumentation is still in its infancy.In this paper, we address this research gap and conduct a thorough investigation of bias in argumentative language models.To this end, we introduce AB BA , a novel resource for bias measurement specifically tailored to argumentation.We employ our resource to assess the effect of argumentative fine-tuning and debiasing on the intrinsic bias found in transformer-based language models using a lightweight adapter-based approach that is more sustainable and parameterefficient than full fine-tuning.Finally, we analyze the potential impact of language model debiasing on the performance in argument quality prediction, a downstream task of computational argumentation.Our results show that we are able to successfully and sustainably remove bias in general and argumentative language models while preserving (and sometimes improving) model performance in downstream tasks.We make all experimental code and data available at https://github.com/ umanlp/FairArgumentativeLM.
Carolin Holtermann, Anne Lauscher, Simone Paolo Ponzetto
ACL (1)2
2022 Welcome to the Modern World of Pronouns: Identity-Inclusive Natural Language Processing beyond Gender
abstract
The world of pronouns is changing – from a closed word class with few members to an open set of terms to reflect identities. However, Natural Language Processing (NLP) barely reflects this linguistic shift, resulting in the possible exclusion of non-binary users, even though recent work outlined the harms of gender-exclusive language technology. The current modeling of 3rd person pronouns is particularly problematic. It largely ignores various phenomena like neopronouns, i.e., novel pronoun sets that are not (yet) widely established. This omission contributes to the discrimination of marginalized and underrepresented groups, e.g., non-binary individuals. It thus prevents gender equality, one of the UN’s sustainable development goals (goal 5). Further, other identity-expressions beyond gender are ignored by current NLP technology. This paper provides an overview of 3rd person pronoun issues for NLP. Based on our observations and ethical considerations, we define a series of five desiderata for modeling pronouns in language technology, which we validate through a survey. We evaluate existing and novel modeling approaches w.r.t. these desiderata qualitatively and quantify the impact of a more discrimination-free approach on an established benchmark dataset.
Anne Lauscher, Archie Crowley, Dirk Hovy
COLING1
2022 Bridging Fairness and Environmental Sustainability in Natural Language Processing
abstract
Fairness and environmental impact are important research directions for the sustainable development of artificial intelligence.However, while each topic is an active research area in natural language processing (NLP), there is a surprising lack of research on the interplay between the two fields.This lacuna is highly problematic, since there is increasing evidence that an exclusive focus on fairness can actually hinder environmental sustainability, and vice versa.In this work, we shed light on this crucial intersection in NLP by (1) investigating the efficiency of current fairness approaches through surveying example methods for reducing unfair stereotypical bias from the literature, and(2) evaluating a common technique to reduce energy consumption (and thus environmental impact) of English NLP models, knowledge distillation (KD), for its impact on fairness.In this case study, we evaluate the effect of important KD factors, including layer and dimensionality reduction, with respect to: (a) performance on the distillation task (natural language inference and semantic similarity prediction), and (b) multiple measures and dimensions of stereotypical bias (e.g., gender bias measured via the Word Embedding Association Test).Our results lead us to clarify current assumptions regarding the effect of KD on unfair bias: contrary to other findings, we show that KD can actually decrease model fairness.
Marius Hessenthaler, Emma Strubell, Dirk Hovy, Anne Lauscher
EMNLP4
2022 SocioProbe: What, When, and Where Language Models Learn about Sociodemographics
abstract
Pre-trained language models (PLMs) have outperformed other NLP models on a wide range of tasks.Opting for a more thorough understanding of their capabilities and inner workings, researchers have established the extend to which they capture lower-level knowledge like grammaticality, and mid-level semantic knowledge like factual understanding.However, there is still little understanding of their knowledge of higher-level aspects of language.In particular, despite the importance of sociodemographic aspects in shaping our language, the questions of whether, where, and how PLMs encode these aspects, e.g., gender or age, is still unexplored.We address this research gap by probing the sociodemographic knowledge of different single-GPU PLMs on multiple English data sets via traditional classifier probing and information-theoretic minimum description length probing.Our results show that PLMs do encode these sociodemographics, and that this knowledge is sometimes spread across the layers of some of the tested PLMs.We further conduct a multilingual analysis and investigate the effect of supplementary training to further explore to what extent, where, and with what amount of pre-training data the knowledge is encoded.Our overall results indicate that sociodemographic knowledge is still a major challenge for NLP.PLMs require large amounts of pre-training data to acquire the knowledge and models that excel in general language understanding do not seem to own more knowledge about these aspects.
Anne Lauscher, Federico Bianchi 0001, Samuel R. Bowman, Dirk Hovy
EMNLP1
2022 Multi2WOZ: A Robust Multilingual Dataset and Conversational Pretraining for Task-Oriented Dialog
abstract
Chia-Chien Hung, Anne Lauscher, Ivan Vulić, Simone Ponzetto, Goran Glavaš. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Chia-Chien Hung, Anne Lauscher, Ivan Vulic, Simone Paolo Ponzetto, Goran Glavas
NAACL-HLT2
2022 MultiCite: Modeling realistic citations requires moving beyond the single-sentence single-label setting
abstract
Anne Lauscher, Brandon Ko, Bailey Kuehl, Sophie Johnson, Arman Cohan, David Jurgens, Kyle Lo. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Anne Lauscher, Brandon Ko, Bailey Kuehl, Sophie Johnson, Arman Cohan, David Jurgens, Kyle Lo
NAACL-HLT1
2022 Scientia Potentia Est - On the Role of Knowledge in Computational Argumentation
abstract
Abstract Despite extensive research efforts in recent years, computational argumentation (CA) remains one of the most challenging areas of natural language processing. The reason for this is the inherent complexity of the cognitive processes behind human argumentation, which integrate a plethora of different types of knowledge, ranging from topic-specific facts and common sense to rhetorical knowledge. The integration of knowledge from such a wide range in CA requires modeling capabilities far beyond many other natural language understanding tasks. Existing research on mining, assessing, reasoning over, and generating arguments largely acknowledges that much more knowledge is needed to accurately model argumentation computationally. However, a systematic overview of the types of knowledge introduced in existing CA models is missing, hindering targeted progress in the field. Adopting the operational definition of knowledge as any task-relevant normative information not provided as input, the survey paper at hand fills this gap by (1) proposing a taxonomy of types of knowledge required in CA tasks, (2) systematizing the large body of CA work according to the reliance on and exploitation of these knowledge types for the four main research areas in CA, and (3) outlining and discussing directions for future research efforts in CA.
Anne Lauscher, Henning Wachsmuth, Iryna Gurevych, Goran Glavas
Trans. Assoc. Comput. Linguistics1
2021 RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models
abstract
Soumya Barikeri, Anne Lauscher, Ivan Vulić, Goran Glavaš. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Soumya Barikeri, Anne Lauscher, Ivan Vulic, Goran Glavas
ACL/IJCNLP (1)2
2020 A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector Spaces
abstract
Distributional word vectors have recently been shown to encode many of the human biases, most notably gender and racial biases, and models for attenuating such biases have consequently been proposed. However, existing models and studies (1) operate on under-specified and mutually differing bias definitions, (2) are tailored for a particular bias (e.g., gender bias) and (3) have been evaluated inconsistently and non-rigorously. In this work, we introduce a general framework for debiasing word embeddings. We operationalize the definition of a bias by discerning two types of bias specification: explicit and implicit. We then propose three debiasing models that operate on explicit or implicit bias specifications and that can be composed towards more robust debiasing. Finally, we devise a full-fledged evaluation framework in which we couple existing bias metrics with newly proposed ones. Experimental findings across three embedding methods suggest that the proposed debiasing models are robust and widely applicable: they often completely remove the bias both implicitly and explicitly without degradation of semantic information encoded in any of the input distributional spaces. Moreover, we successfully transfer debiasing models, by means of cross-lingual embedding spaces, and remove or attenuate biases in distributional word vector spaces of languages that lack readily available bias specifications.
Anne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan Vulic
AAAI1
2020 Rhetoric, Logic, and Dialectic: Advancing Theory-based Argument Quality Assessment in Natural Language Processing
abstract
Though preceding work in computational argument quality (AQ) mostly focuses on assessing overall AQ, researchers agree that writers would benefit from feedback targeting individual dimensions of argumentation theory.However, a large-scale theory-based corpus and corresponding computational models are missing.We fill this gap by conducting an extensive analysis covering three diverse domains of online argumentative writing and presenting GAQCorpus: the first largescale English multi-domain (community Q&A forums, debate forums, review forums) corpus annotated with theory-based AQ scores.We then propose the first computational approaches to theory-based assessment, which can serve as strong baselines for future work.We demonstrate the feasibility of large-scale AQ annotation, show that exploiting relations between dimensions yields performance improvements, and explore the synergies between theory-based prediction and practical AQ assessment.
Anne Lauscher, Lily Ng, Courtney Napoles, Joel R. Tetreault
COLING1
2020 Specializing Unsupervised Pretraining Models for Word-Level Semantic Similarity
abstract
Unsupervised pretraining models have been shown to facilitate a wide range of downstream NLP applications. These models, however, retain some of the limitations of traditional static word embeddings. In particular, they encode only the distributional knowledge available in raw text corpora, incorporated through language modeling objectives. In this work, we complement such distributional knowledge with external lexical knowledge, that is, we integrate the discrete knowledge on word-level semantic similarity into pretraining. To this end, we generalize the standard BERT model to a multi-task learning setting where we couple BERT’s masked language modeling and next sentence prediction objectives with an auxiliary task of binary word relation classification. Our experiments suggest that our "Lexically Informed” BERT (LIBERT), specialized for the word-level semantic similarity, yields better performance than the lexically blind “vanilla” BERT on several language understanding tasks. Concretely, LIBERT outperforms BERT in 9 out of 10 tasks of the GLUE benchmark and is on a par with BERT in the remaining one. Moreover, we show consistent gains on 3 benchmarks for lexical simplification, a task where knowledge about word-level semantic similarity is paramount, as well as large gains on lexical reasoning probes.
Anne Lauscher, Ivan Vulic, Edoardo Maria Ponti, Anna Korhonen, Goran Glavas
COLING1
2020 From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual Transformers
abstract
bak probing, til beskrivelser av transformerarkitekturen og debatten rundt tolkbarhet av oppmerksomhet. I hvert introduksjonskapittel prver vi derfor kontekstualisere vrt eget arbeid.
Anne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran Glavas
EMNLP (1)1
2020 The OpenCitations Data Model
abstract
A variety of schemas and ontologies are currently used for the machine-readable description of bibliographic entities and citations. This diversity, and the reuse of the same ontology terms with different nuances, generates inconsistencies in data. Adoption of a single data model would facilitate data integration tasks regardless of the data supplier or context application. In this paper we present the OpenCitations Data Model (OCDM), a generic data model for describing bibliographic entities and citations, developed using Semantic Web technologies. We also evaluate the effective reusability of OCDM according to ontology evaluation practices, mention existing users of OCDM, and discuss the use and impact of OCDM in the wider open science community.
Marilena Daquino, Silvio Peroni, David M. Shotton, Giovanni Colavizza, Behnam Ghavimi, Anne Lauscher, Philipp Mayr 0001, Matteo Romanello, Philipp Zumstein
ISWC (2)6
2018 Investigating the Role of Argumentation in the Rhetorical Analysis of Scientific Publications with Neural Multi-Task Learning Models
abstract
Exponential growth in the number of scientific publications yields the need for effective automatic analysis of rhetorical aspects of scientific writing. Acknowledging the argumentative nature of scientific text, in this work we investigate the link between the argumentative structure of scientific publications and rhetorical aspects such as discourse categories or citation contexts. To this end, we (1) augment a corpus of scientific publications annotated with four layers of rhetoric annotations with argumentation annotations and (2) investigate neural multi-task learning architectures combining argument extraction with a set of rhetorical classification tasks. By coupling rhetorical classifiers with the extraction of argumentative components in a joint multi-task learning setting, we obtain significant performance gains for different rhetorical analysis tasks.
Anne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Kai Eckert 0001
EMNLP1