EDBT 2026 Demo / reviewers in the wild / expert
Neele Falk
dblp:266/0722
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Are LLMs Doing to Scientific Communication? Measuring Changes in Writing Practices and Reading ExperienceabstractHas the style of scientific communication changed due to the growing use of large language models in the writing process? We address this question in the domain of Natural Language Processing by leveraging two data resources we create: a naturalistic corpus of over 37,000 papers from the ACL Anthology (2020-2024); and a synthetic dataset of 3,000 human-written passages and their LLM-generated improvements. We first implement a series of diachronic lexical analyses, showing that both word frequency and usage contexts have changed significantly over time, indicating semantic specialization in some cases and generalization in others. Broadening our perspective, we then model a range of more complex stylistic features and find that LLM-modified texts more frequently contain certain syntactic constructions, more complex and longer words and a lower lexical diversity. Finally, we connect these changes in writing practices to subjective reading experience through a pilot annotation study with 20 domain experts. They overall rate LLM-improved texts as more understandable and exciting, but also express negative qualitative attitudes towards LLMs, highlighting the strongly subjective effect of AI-assisted writing on reading experience. Filip Miletic 0002, Neele Falk |
LREC | 2 |
| 2025 | Mining the uncertainty patterns of humans and models in the annotation of moral foundations and human valuesabstractThe NLP community has converged on considering disagreement in annotation (or human label variation, HLV) as a constitutive feature of subjective tasks. This paper makes a further step by investigating the relationship between HLV and model uncertainty, and the impact of linguistic features of the items on both. We focus on the identification of moral foundations (e.g., care, fairness, loyalty) and human values (e.g., be polite, be honest) in text. We select three standard datasets and proceed into two steps. First, we focus on HLV and analyze the linguistic features (complexity, polarity, pragmatic phenomena, lexical choices) that correlate with HLV. Next, we proceed to uncertainty and its relationship to HLV. We experiment with RoBERTa and Flan-T5 in a number of training setups and evaluation metrics that test the calibration of uncertainty to HLV and its relationship to performance beyond majority vote; next, we analyze the impact of linguistic features on uncertainty. We find that RoBERTa with soft loss is better calibrated to HLV, and we find alignment between calibrated models and humans in the features (textual complexity and polarity) triggering variation. Neele Falk, Gabriella Lapesa |
ACL (1) | 1 |
| 2025 | "Feels Feminine to Me": Understanding Perceived Gendered Style through Human AnnotationsabstractIn NLP, language-gender associations are commonly grounded in the author's gender identity, inferred from their language use.However, this identity-based framing risks reinforcing stereotypes and marginalizing individuals who do not conform to normative language-gender associations.To address this, we operationalize the language-gender association as a perceived gender expression of language, focusing on how such expression is externally interpreted by humans, independent of the author's gender identity.We present the first dataset of its kind: 5,100 human annotations of perceived gendered style-human-written texts rated on a five-point scale from very feminine to very masculine.While perception is inherently subjective, our analysis identifies textual features associated with higher agreement among annotators: formal expressions and lower emotional intensity.Moreover, annotator demographics influence their perception: women annotators are more likely to label texts as feminine, and men and non-binary annotators as masculine.Finally, feature analysis reveals that text's perceived gendered style is shaped by both affective and function words, partially overlapping with known patterns of language variation across gender identities.Our findings lay the groundwork for operationalizing gendered style through human annotation, while also highlighting annotators' subjective judgments as meaningful signals to understand perceptionbased concepts.1 n_dependency_npmod has_missing_values n_dependency_root has_missing_values n_tokens high collinearity n_types high collinearity n_characters high collinearity maas_index high collinearity n_hapax_legomena high collinearity n_global_token_hapax_legomena high collinearity n_hapax_dislegomena high collinearity n_global_lemma_hapax_dislegomena high collinearity n_global_token_hapax_dislegomena high collinearity n_syllables high collinearity flesch_reading_ease high collinearity flesch_kincaid_grade high collinearity Feature Reason ari high collinearity cli high collinearity gunning_fog high collinearity lix high collinearity rix high collinearity n_dependency_advmod high collinearity n_dependency_prep high collinearity n_dependency_punct high collinearity raw_sequence_length high collinearity lemma_token_ratio high collinearity n_lemmas high collinearity cttr high collinearity ttr high collinearity herdan_c high collinearity rttr high collinearity mattr high collinearity yule_k high collinearity n_cconj high collinearity n_det high collinearity Neele Falk, Agnieszka Falenska |
EMNLP | 2 |
| 2025 | PerspectiveMod: A Perspectivist Resource for Deliberative ModerationabstractHuman moderators in online discussions face a heterogeneous range of tasks, which go beyond content moderation, or policing.They also support and improve discussion quality, which is challenging to model (and evaluate) in NLP due to its inherent subjectivity and the scarcity of annotated resources.We address this gap by introducing PerspectiveMod, a dataset of online comments annotated for the question: "Does this comment require moderation, and why?" Annotations were collected from both expert moderators and trained nonexperts.PerspectiveMod is unique in its intentional variation across (a) the level of moderation experience embedded in the source data (professional vs. non-professional moderation environments), (b) the annotator profiles (experts vs. trained crowdworkers), and (c) the richness of each moderation judgment, both in terms on fine-grained comment properties (drawn from argumentation and deliberative theory) and in the representation of the individuality of the annotator (socio-demographics and attitudes towards the task).We advance understanding of the task's complexity by providing interpretation layers that account for its subjectivity.Our statistical analysis highlights the value of collecting annotator perspectives, including their experiences, attitudes, and views on AI, as a foundation for developing more context-aware and interpretively robust moderation tools. Eva Maria Vecchi, Neele Falk, Carlotta Quensel, Iman Jundi, Gabriella Lapesa |
EMNLP | 2 |
| 2025 | It Is Not Only the Negative that Deserves Attention! Understanding, Generation & Evaluation of (Positive) ModerationabstractIman Jundi, Eva Maria Vecchi, Carlotta Quensel, Neele Falk, Gabriella Lapesa. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Iman Jundi, Eva Maria Vecchi, Carlotta Quensel, Neele Falk, Gabriella Lapesa |
NAACL (Long Papers) | 4 |
| 2024 | Stories and Personal Experiences in the COVID-19 DiscourseabstractStorytelling, i.e., the use of of anecdotes and personal experiences, plays a crucial role in everyday argumentation. This is particularly true for the highly controversial debates that spark in times of crisis - where the focus of the discussion is on heterogeneous aspects of everyday life. For individuals, stories can have a strong persuasive power; for a larger collective, stories can help decision-makers to develop strategies for addressing the challenges people are facing, especially in times of crisis. In this paper, we analyse the use of storytelling in the COVID-19 discourse. We carry out our analysis on three publicly available Reddit datasets, for a total of 367K comments. We automatically annotate the Reddit datasets by detecting spans containing storytelling and classifying them into: a) personal vs. general – is the story experienced by the speaker? b) argumentative function (Does the story clarify a problem, potentially consisting in harm to a specific group? Does it exemplify a solution to a problem, or does it establish the credibility of the speaker?), and c) topic. We then carry out an analysis which establishes the relevance of storytelling in the COVID discourse and further uncovers interactions between topics and types of stories associated to them. Neele Falk, Gabriella Lapesa |
LREC/COLING | 1 |
| 2024 | Moderation in the Wild: Investigating User-Driven Moderation in Online DiscussionsabstractEffective content moderation is imperative for fostering healthy and productive discussions in online domains.Despite the substantial efforts of moderators, the overwhelming nature of discussion flow can limit their effectiveness.However, it is not only trained moderators who intervene in online discussions to improve their quality."Ordinary" users also act as moderators, actively intervening to correct information of other users' posts, enhance arguments, and steer discussions back on course.This paper introduces the phenomenon of user moderation, documenting and releasing UMOD, the first dataset of comments in which users act as moderators.UMOD contains 1000 comment-reply pairs from the subreddit r/changemyview with crowdsourced annotations from a large annotator pool and with a fine-grained annotation schema targeting the functions of moderation, stylistic properties (aggressiveness, subjectivity, sentiment), constructiveness, as well as the individual perspectives of the annotators on the task.The release of UMOD is complemented by two analyses which focus on the constitutive features of constructiveness in user moderation and on the sources of annotator disagreements, given the high subjectivity of the task. Neele Falk, Eva Maria Vecchi, Iman Jundi, Gabriella Lapesa |
EACL (1) | 1 |
| 2024 | Annotator-Centric Active Learning for Subjective NLP TasksabstractActive Learning (AL) addresses the high costs of collecting human annotations by strategically annotating the most informative samples.However, for subjective NLP tasks, incorporating a wide range of perspectives in the annotation process is crucial to capture the variability in human judgments.We introduce Annotator-Centric Active Learning (ACAL), which incorporates an annotator selection strategy following data sampling.Our objective is two-fold:(1) to efficiently approximate the full diversity of human judgments, and (2) to assess model performance using annotator-centric metrics, which value minority and majority perspectives equally.We experiment with multiple annotator selection strategies across seven subjective NLP tasks, employing both traditional and novel, human-centered evaluation metrics.Our findings indicate that ACAL improves data efficiency and excels in annotator-centric performance evaluations.However, its success depends on the availability of a sufficiently large and diverse pool of annotators to sample from. Michiel van der Meer, Neele Falk, Pradeep K. Murukannaiah, Enrico Liscio |
EMNLP | 2 |
| 2024 | Beyond Prompt Brittleness: Evaluating the Reliability and Consistency of Political Worldviews in LLMsabstractAbstract Due to the widespread use of large language models (LLMs), we need to understand whether they embed a specific “worldview” and what these views reflect. Recent studies report that, prompted with political questionnaires, LLMs show left-liberal leanings (Feng et al., 2023; Motoki et al., 2024). However, it is as yet unclear whether these leanings are reliable (robust to prompt variations) and whether the leaning is consistent across policies and political leaning. We propose a series of tests which assess the reliability and consistency of LLMs’ stances on political statements based on a dataset of voting-advice questionnaires collected from seven EU countries and annotated for policy issues. We study LLMs ranging in size from 7B to 70B parameters and find that their reliability increases with parameter count. Larger models show overall stronger alignment with left-leaning parties but differ among policy programs: They show a (left-wing) positive stance towards environment protection, social welfare state, and liberal society but also (right-wing) law and order, with no consistent preferences in the areas of foreign policy and migration. Tanise Ceron, Neele Falk, Ana Baric, Dmitry Nikolaev 0002, Sebastian Padó |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | StoryARG: a corpus of narratives and personal experiences in argumentative textsabstractHumans are storytellers, even in communication scenarios which are assumed to be more rationality-oriented, such as argumentation.Indeed, supporting arguments with narratives or personal experiences (henceforth, stories) is a very natural thing to do -and yet, this phenomenon is largely unexplored in computational argumentation.Which role do stories play in an argument?Do they make the argument more effective?What are their narrative properties?To address these questions, we collected and annotated StoryARG, a dataset sampled from well-established corpora in computational argumentation (ChangeMyView and RegulationRoom), and the Social Sciences (Europolis), as well as comments to New York Times articles.StoryARG contains 2451 textual spans annotated at two levels.At the argumentative level, we annotate the function of the story (e.g., clarification, disclosure of harm, search for a solution, establishing speaker's authority), as well as its impact on the effectiveness of the argument and its emotional load.At the level of narrative properties, we annotate whether the story has a plot-like development, is factual or hypothetical, and who the protagonist is.What makes a story effective in an argument?Our analysis of the annotations in StoryARG uncover a positive impact on effectiveness for stories which illustrate a solution to a problem, and in general, annotator-specific preferences that we investigate with regression analysis. Neele Falk, Gabriella Lapesa |
ACL (1) | 1 |
| 2023 | Node Placement in Argument Maps: Modeling Unidirectional Relations in High & Low-Resource ScenariosabstractArgument maps structure discourse into nodes in a tree with each node being an argument that supports or opposes its parent argument.This format is more comprehensible and less redundant compared to an unstructured one.Exploring those maps and maintaining their structure by placing new arguments under suitable parents is more challenging for users with huge maps that are typical in online discussions.To support those users, we introduce the task of node placement: suggesting candidate nodes as parents for a new contribution.We establish an upper-bound of human performance, and conduct experiments with models of various sizes and training strategies.We experiment with a selection of maps from Kialo, drawn from a heterogeneous set of domains.Based on an annotation study, we highlight the ambiguity of the task that makes it challenging for both humans and models.We examine the unidirectional relation between tree nodes and show that encoding a node into different embeddings for each of the parent and child cases improves performance.We further show the few-shot effectiveness of our approach. Iman Jundi, Neele Falk, Eva Maria Vecchi, Gabriella Lapesa |
ACL (1) | 2 |
| 2022 | Reports of personal experiences and stories in argumentation: datasets and analysisabstractReports of personal experiences or stories can play a crucial role in argumentation, as they represent an immediate and (often) relatable way to back up one's position with respect to a given topic.They are easy to understand and increase empathy: this makes them powerful in argumentation.The impact of personal reports and stories in argumentation has been studied in the Social Sciences, but it is still largely underexplored in NLP.Our work is the first step towards filling this gap: our goal is to develop robust classifiers to identify documents containing personal experiences and reports.The main challenge is the scarcity of annotated data: our solution is to leverage existing annotations to be able to scale-up the analysis.Our contribution is two-fold.First, we conduct a set of in-domain and cross-domain experiments involving three datasets (two from Argument Mining, one from the Social Sciences), modeling architectures, training setups and fine-tuning options tailored to the involved domains.We show that despite the differences among datasets and annotations, robust crossdomain classification is possible.Second, we employ linear regression for performance mining, identifying performance trends both for overall classification performance and individual classifier predictions. Neele Falk, Gabriella Lapesa |
ACL (1) | 1 |
| 2022 | Scaling up Discourse Quality Annotation for Political ScienceabstractThe empirical quantification of the quality of a contribution to a political discussion is at the heart of deliberative theory, the subdiscipline of political science which investigates decision-making in deliberative democracy. Existing annotation on deliberative quality is time-consuming and carried out by experts, typically resulting in small datasets which also suffer from strong class imbalance. Scaling up such annotations with automatic tools is desirable, but very challenging. We take up this challenge and explore different strategies to improve the prediction of deliberative quality dimensions (justification, common good, interactivity, respect) in a standard dataset. Our results show that simple data augmentation techniques successfully alleviate data imbalance. Classifiers based on linguistic features (textual complexity and sentiment/polarity) and classifiers integrating argument quality annotations (from the argument mining community in NLP) were consistently outperformed by transformer-based models, with or without data augmentation. Neele Falk, Gabriella Lapesa |
LREC | 1 |
| 2021 | Towards Argument Mining for Social Good: A SurveyabstractEva Maria Vecchi, Neele Falk, Iman Jundi, Gabriella Lapesa. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Eva Maria Vecchi, Neele Falk, Iman Jundi, Gabriella Lapesa |
ACL/IJCNLP (1) | 2 |
| 2020 | All That Glitters is Not Gold: A Gold Standard of Adjective-Noun Collocations for GermanabstractIn this paper we present the GerCo dataset of adjective-noun collocations for German, such as alter Freund ‘old friend’ and tiefe Liebe ‘deep love’. The annotation has been performed by experts based on the annotation scheme introduced in this paper. The resulting dataset contains 4,732 positive and negative instances of collocations and covers all the 16 semantic classes of adjectives as defined in the German wordnet GermaNet. The dataset can serve as a reliable empirical basis for comparing different theoretical frameworks concerned with collocations or as material for data-driven approaches to the studies of collocations including different machine learning experiments. This paper addresses the latter issue by using the GerCo dataset for evaluating different models on the task of automatic collocation identification. We compare lexical association measures with static and contextualized word embeddings. The experiments show that word embeddings outperform methods based on statistical association measures by a wide margin. Yana Strakatova, Neele Falk, Isabel Fuhrmann, Erhard W. Hinrichs, Daniela Rossmann |
LREC | 2 |