Gabriella Lapesa

dblp:86/8156 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
16since 2021 · last 2025
0000-0002-4418-3609ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 2 first-author · 16 since 2021
YearPublicationVenuePosition
2025 Mining the uncertainty patterns of humans and models in the annotation of moral foundations and human values
abstract
The NLP community has converged on considering disagreement in annotation (or human label variation, HLV) as a constitutive feature of subjective tasks. This paper makes a further step by investigating the relationship between HLV and model uncertainty, and the impact of linguistic features of the items on both. We focus on the identification of moral foundations (e.g., care, fairness, loyalty) and human values (e.g., be polite, be honest) in text. We select three standard datasets and proceed into two steps. First, we focus on HLV and analyze the linguistic features (complexity, polarity, pragmatic phenomena, lexical choices) that correlate with HLV. Next, we proceed to uncertainty and its relationship to HLV. We experiment with RoBERTa and Flan-T5 in a number of training setups and evaluation metrics that test the calibration of uncertainty to HLV and its relationship to performance beyond majority vote; next, we analyze the impact of linguistic features on uncertainty. We find that RoBERTa with soft loss is better calibrated to HLV, and we find alignment between calibrated models and humans in the features (textual complexity and polarity) triggering variation.
Neele Falk, Gabriella Lapesa
ACL (1)2
2025 AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contexts
abstract
Distinguishing LLM-generated text from human-written is a key challenge for safe and ethical NLP, particularly in high-stake settings such as persuasive online discourse.While recent work focuses on detection, real-world use cases also demand interpretable tools to help humans understand and distinguish LLMgenerated texts.To this end, we present an analysis framework comparing human-and LLM-authored arguments using two easilyinterpretable feature sets: general-purpose linguistic features (e.g., lexical richness, syntactic complexity) and domain-specific features related to argument quality (e.g., logical soundness, engagement strategies).Applied to /r/ChangeMyView arguments by humans and three LLMs, our method reveals clear patterns: LLM-generated counter-arguments show lower type-token and lemma-token ratios but higher emotional intensity -particularly in anticipation and trust.They more closely resemble textbook-quality arguments -cogent, justified, explicitly respectful toward others, and positive in tone.Moreover, counter-arguments generated by LLMs converge more closely with the original post's style and quality than those written by humans.Finally, we demonstrate that these differences enable a lightweight, interpretable, and highly effective classifier for detecting LLM-generated comments in CMV.Milad Alshomary and Henning Wachsmuth.2023. Conclusion-based counter-argument generation. In
Esra Dönmez, Maximilian Maurer, Gabriella Lapesa, Agnieszka Falenska
EMNLP3
2025 PerspectiveMod: A Perspectivist Resource for Deliberative Moderation
abstract
Human moderators in online discussions face a heterogeneous range of tasks, which go beyond content moderation, or policing.They also support and improve discussion quality, which is challenging to model (and evaluate) in NLP due to its inherent subjectivity and the scarcity of annotated resources.We address this gap by introducing PerspectiveMod, a dataset of online comments annotated for the question: "Does this comment require moderation, and why?" Annotations were collected from both expert moderators and trained nonexperts.PerspectiveMod is unique in its intentional variation across (a) the level of moderation experience embedded in the source data (professional vs. non-professional moderation environments), (b) the annotator profiles (experts vs. trained crowdworkers), and (c) the richness of each moderation judgment, both in terms on fine-grained comment properties (drawn from argumentation and deliberative theory) and in the representation of the individuality of the annotator (socio-demographics and attitudes towards the task).We advance understanding of the task's complexity by providing interpretation layers that account for its subjectivity.Our statistical analysis highlights the value of collecting annotator perspectives, including their experiences, attitudes, and views on AI, as a foundation for developing more context-aware and interpretively robust moderation tools.
Eva Maria Vecchi, Neele Falk, Carlotta Quensel, Iman Jundi, Gabriella Lapesa
EMNLP5
2025 It Is Not Only the Negative that Deserves Attention! Understanding, Generation & Evaluation of (Positive) Moderation
abstract
Iman Jundi, Eva Maria Vecchi, Carlotta Quensel, Neele Falk, Gabriella Lapesa. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Iman Jundi, Eva Maria Vecchi, Carlotta Quensel, Neele Falk, Gabriella Lapesa
NAACL (Long Papers)5
2025 Towards a Perspectivist Turn in Argument Quality Assessment
abstract
Julia Romberg, Maximilian Maurer, Henning Wachsmuth, Gabriella Lapesa. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Julia Romberg, Maximilian Maurer, Henning Wachsmuth, Gabriella Lapesa
NAACL (Long Papers)4
2024 Self-reported Demographics and Discourse Dynamics in a Persuasive Online Forum
abstract
Research on language as interactive discourse underscores the deliberate use of demographic parameters such as gender, ethnicity, and class to shape social identities. For example, by explicitly disclosing one’s information and enforcing one’s social identity to an online community, the reception by and interaction with the said community is impacted, e.g., strengthening one’s opinions by depicting the speaker as credible through their experience in the subject. Here, we present a first thorough study of the role and effects of self-disclosures on online discourse dynamics, focusing on a pervasive type of self-disclosure: author gender. Concretely, we investigate the contexts and properties of gender self-disclosures and their impact on interaction dynamics in an online persuasive forum, ChangeMyView. Our contribution is twofold. At the level of the target phenomenon, we fill a research gap in the understanding of the impact of these self-disclosures on the discourse by bringing together features related to forum activity (votes, number of comments), linguistic/stylistic features from the literature, and discourse topics. At the level of the contributed resource, we enrich and release a comprehensive dataset that will provide a further impulse for research on the interplay between gender disclosures, community interaction, and persuasion in online discourse.
Agnieszka Falenska, Eva Maria Vecchi, Gabriella Lapesa
LREC/COLING3
2024 Stories and Personal Experiences in the COVID-19 Discourse
abstract
Storytelling, i.e., the use of of anecdotes and personal experiences, plays a crucial role in everyday argumentation. This is particularly true for the highly controversial debates that spark in times of crisis - where the focus of the discussion is on heterogeneous aspects of everyday life. For individuals, stories can have a strong persuasive power; for a larger collective, stories can help decision-makers to develop strategies for addressing the challenges people are facing, especially in times of crisis. In this paper, we analyse the use of storytelling in the COVID-19 discourse. We carry out our analysis on three publicly available Reddit datasets, for a total of 367K comments. We automatically annotate the Reddit datasets by detecting spans containing storytelling and classifying them into: a) personal vs. general – is the story experienced by the speaker? b) argumentative function (Does the story clarify a problem, potentially consisting in harm to a specific group? Does it exemplify a solution to a problem, or does it establish the credibility of the speaker?), and c) topic. We then carry out an analysis which establishes the relevance of storytelling in the COVID discourse and further uncovers interactions between topics and types of stories associated to them.
Neele Falk, Gabriella Lapesa
LREC/COLING2
2024 Argument Quality Assessment in the Age of Instruction-Following Large Language Models
abstract
The computational treatment of arguments on controversial issues has been subject to extensive NLP research, due to its envisioned impact on opinion formation, decision making, writing education, and the like. A critical task in any such application is the assessment of an argument’s quality - but it is also particularly challenging. In this position paper, we start from a brief survey of argument quality research, where we identify the diversity of quality notions and the subjectiveness of their perception as the main hurdles towards substantial progress on argument quality assessment. We argue that the capabilities of instruction-following large language models (LLMs) to leverage knowledge across contexts enable a much more reliable assessment. Rather than just fine-tuning LLMs towards leaderboard chasing on assessment tasks, they need to be instructed systematically with argumentation theories and scenarios as well as with ways to solve argument-related problems. We discuss the real-world opportunities and ethical issues emerging thereby.
Henning Wachsmuth, Gabriella Lapesa, Elena Cabrio, Anne Lauscher, Joonsuk Park, Eva Maria Vecchi, Serena Villata, Timon Ziegenbein
LREC/COLING2
2024 Moderation in the Wild: Investigating User-Driven Moderation in Online Discussions
abstract
Effective content moderation is imperative for fostering healthy and productive discussions in online domains.Despite the substantial efforts of moderators, the overwhelming nature of discussion flow can limit their effectiveness.However, it is not only trained moderators who intervene in online discussions to improve their quality."Ordinary" users also act as moderators, actively intervening to correct information of other users' posts, enhance arguments, and steer discussions back on course.This paper introduces the phenomenon of user moderation, documenting and releasing UMOD, the first dataset of comments in which users act as moderators.UMOD contains 1000 comment-reply pairs from the subreddit r/changemyview with crowdsourced annotations from a large annotator pool and with a fine-grained annotation schema targeting the functions of moderation, stylistic properties (aggressiveness, subjectivity, sentiment), constructiveness, as well as the individual perspectives of the annotators on the task.The release of UMOD is complemented by two analyses which focus on the constitutive features of constructiveness in user moderation and on the sources of annotator disagreements, given the high subjectivity of the task.
Neele Falk, Eva Maria Vecchi, Iman Jundi, Gabriella Lapesa
EACL (1)4
2023 StoryARG: a corpus of narratives and personal experiences in argumentative texts
abstract
Humans are storytellers, even in communication scenarios which are assumed to be more rationality-oriented, such as argumentation.Indeed, supporting arguments with narratives or personal experiences (henceforth, stories) is a very natural thing to do -and yet, this phenomenon is largely unexplored in computational argumentation.Which role do stories play in an argument?Do they make the argument more effective?What are their narrative properties?To address these questions, we collected and annotated StoryARG, a dataset sampled from well-established corpora in computational argumentation (ChangeMyView and RegulationRoom), and the Social Sciences (Europolis), as well as comments to New York Times articles.StoryARG contains 2451 textual spans annotated at two levels.At the argumentative level, we annotate the function of the story (e.g., clarification, disclosure of harm, search for a solution, establishing speaker's authority), as well as its impact on the effectiveness of the argument and its emotional load.At the level of narrative properties, we annotate whether the story has a plot-like development, is factual or hypothetical, and who the protagonist is.What makes a story effective in an argument?Our analysis of the annotations in StoryARG uncover a positive impact on effectiveness for stories which illustrate a solution to a problem, and in general, annotator-specific preferences that we investigate with regression analysis.
Neele Falk, Gabriella Lapesa
ACL (1)2
2023 Node Placement in Argument Maps: Modeling Unidirectional Relations in High & Low-Resource Scenarios
abstract
Argument maps structure discourse into nodes in a tree with each node being an argument that supports or opposes its parent argument.This format is more comprehensible and less redundant compared to an unstructured one.Exploring those maps and maintaining their structure by placing new arguments under suitable parents is more challenging for users with huge maps that are typical in online discussions.To support those users, we introduce the task of node placement: suggesting candidate nodes as parents for a new contribution.We establish an upper-bound of human performance, and conduct experiments with models of various sizes and training strategies.We experiment with a selection of maps from Kialo, drawn from a heterogeneous set of domains.Based on an annotation study, we highlight the ambiguity of the task that makes it challenging for both humans and models.We examine the unidirectional relation between tree nodes and show that encoding a node into different embeddings for each of the parent and child cases improves performance.We further show the few-shot effectiveness of our approach.
Iman Jundi, Neele Falk, Eva Maria Vecchi, Gabriella Lapesa
ACL (1)4
2022 Reports of personal experiences and stories in argumentation: datasets and analysis
abstract
Reports of personal experiences or stories can play a crucial role in argumentation, as they represent an immediate and (often) relatable way to back up one's position with respect to a given topic.They are easy to understand and increase empathy: this makes them powerful in argumentation.The impact of personal reports and stories in argumentation has been studied in the Social Sciences, but it is still largely underexplored in NLP.Our work is the first step towards filling this gap: our goal is to develop robust classifiers to identify documents containing personal experiences and reports.The main challenge is the scarcity of annotated data: our solution is to leverage existing annotations to be able to scale-up the analysis.Our contribution is two-fold.First, we conduct a set of in-domain and cross-domain experiments involving three datasets (two from Argument Mining, one from the Social Sciences), modeling architectures, training setups and fine-tuning options tailored to the involved domains.We show that despite the differences among datasets and annotations, robust crossdomain classification is possible.Second, we employ linear regression for performance mining, identifying performance trends both for overall classification performance and individual classifier predictions.
Neele Falk, Gabriella Lapesa
ACL (1)2
2022 Investigating Independence vs. Control: Agenda-Setting in Russian News Coverage on Social Media
abstract
Agenda-setting is a widely explored phenomenon in political science: powerful stakeholders (governments or their financial supporters) have control over the media and set their agenda: political and economical powers determine which news should be salient. This is a clear case of targeted manipulation to divert the public attention from serious issues affecting internal politics (such as economic downturns and scandals) by flooding the media with potentially distracting information. We investigate agenda-setting in the Russian social media landscape, exploring the relation between economic indicators and mentions of foreign geopolitical entities, as well as of Russia itself. Our contributions are at three levels: at the level of the domain of the investigation, our study is the first to substructure the Russian media landscape in state-controlled vs. independent outlets in the context of strategic distraction from negative economic trends; at the level of the scope of the investigation, we involve a large set of geopolitical entities (while previous work has focused on the U.S.); at the qualitative level, our analysis of posts on Ukraine, whose relationship with Russia is of high geopolitical relevance, provides further insights into the contrast between state-controlled and independent outlets.
Annerose Eichel, Gabriella Lapesa, Sabine Schulte im Walde
LREC2
2022 Scaling up Discourse Quality Annotation for Political Science
abstract
The empirical quantification of the quality of a contribution to a political discussion is at the heart of deliberative theory, the subdiscipline of political science which investigates decision-making in deliberative democracy. Existing annotation on deliberative quality is time-consuming and carried out by experts, typically resulting in small datasets which also suffer from strong class imbalance. Scaling up such annotations with automatic tools is desirable, but very challenging. We take up this challenge and explore different strategies to improve the prediction of deliberative quality dimensions (justification, common good, interactivity, respect) in a standard dataset. Our results show that simple data augmentation techniques successfully alleviate data imbalance. Classifiers based on linguistic features (textual complexity and sentiment/polarity) and classifiers integrating argument quality annotations (from the argument mining community in NLP) were consistently outperformed by transformer-based models, with or without data augmentation.
Neele Falk, Gabriella Lapesa
LREC2
2021 Towards Argument Mining for Social Good: A Survey
abstract
Eva Maria Vecchi, Neele Falk, Iman Jundi, Gabriella Lapesa. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Eva Maria Vecchi, Neele Falk, Iman Jundi, Gabriella Lapesa
ACL/IJCNLP (1)4
2021 FAST: A carefully sampled and cognitively motivated dataset for distributional semantic evaluation
abstract
What is the first word that comes to your mind when you hear giraffe, or damsel, or freedom?Such free associations contain a huge amount of information on the mental representations of the corresponding concepts, and are thus an extremely valuable testbed for the evaluation of semantic representations extracted from corpora.In this paper, we present FAST (Free ASsociation Tasks), a free association dataset for English rigorously sampled from two standard free association norms collections (the Edinburgh Associative Thesaurus and the University of South Florida Free Association Norms), discuss two evaluation tasks, and provide baseline results.In parallel, we discuss methodological considerations concerning the desiderata for a proper evaluation of semantic representations.
Stefan Evert, Gabriella Lapesa
CoNLL2
2020 DEbateNet-mig15: Tracing the 2015 Immigration Debate in Germany Over Time
abstract
DEbateNet-migr15 is a manually annotated dataset for German which covers the public debate on immigration in 2015. The building block of our annotation is the political science notion of a claim, i.e., a statement made by a political actor (a politician, a party, or a group of citizens) that a specific action should be taken (e.g., vacant flats should be assigned to refugees). We identify claims in newspaper articles, assign them to actors and fine-grained categories and annotate their polarity and date. The aim of this paper is two-fold: first, we release the full DEbateNet-mig15 corpus and document it by means of a quantitative and qualitative analysis; second, we demonstrate its application in a discourse network analysis framework, which enables us to capture the temporal dynamics of the political debate
Gabriella Lapesa, André Blessing, Nico Blokker, Erenay Dayanik, Sebastian Haunss, Jonas Kuhn, Sebastian Padó
LREC1
2014 A Large Scale Evaluation of Distributional Semantic Models: Parameters, Interactions and Model Selection
abstract
This paper presents the results of a large-scale evaluation study of window-based Distributional Semantic Models on a wide variety of tasks. Our study combines a broad coverage of model parameters with a model selection methodology that is robust to overfitting and able to capture parameter interactions. We show that our strategy allows us to identify parameter configurations that achieve good performance across different datasets and tasks.
Gabriella Lapesa, Stefan Evert
Trans. Assoc. Comput. Linguistics1
2012 LexIt: A Computational Resource on Italian Argument Structure
Alessandro Lenci, Gabriella Lapesa, Giulia Bonansinga
LREC2
2010 Building an Italian FrameNet through Semi-automatic Corpus Analysis
Alessandro Lenci, Martina Johnson, Gabriella Lapesa
LREC3