VLDB 2026 Research / reviewers in the wild / expert
Barbara Plank
dblp:46/521
· DBLP profile ↗
103ranked-venue papers
14as first author
62since 2021 · last 2026
0000-0002-4394-1965ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 100 · 14 first-author · 62 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language ModelsabstractMany in-silico simulations of human survey responses with large language models (LLMs) focus on generating closed-ended survey responses, whereas LLMs are typically trained to generate open-ended text instead.Previous research has used a diverse range of methods for generating closed-ended survey responses with LLMs, and a standard practice remains to be identified.In this paper, we systematically investigate the impact that various Survey Response Generation Methods have on predicted survey responses.We present the results of 32 mio.simulated survey responses across 8 Survey Response Generation Methods, 4 political attitude surveys, and 10 openweight language models.We find significant differences between the Survey Response Generation Methods in both individual-level and subpopulation-level alignment.Our results show that Restricted Generation Methods perform best overall, and that reasoning output does not consistently improve alignment.Our work underlines the significant impact that Survey Response Generation Methods have on simulated survey responses, and we develop practical recommendations on the application of Survey Response Generation Methods. Georg Ahnert, Anna-Carolina Haensch, Barbara Plank, Markus Strohmaier |
ACL (1) | 3 |
| 2026 | Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German DialectsabstractResearch on cross-dialectal transfer from a standard to a non-standard dialect variety has typically focused on text data.However, dialects are primarily spoken, and non-standard spellings cause issues in text processing.We compare standard-to-dialect transfer in three settings: text models, speech models, and cascaded systems where speech first gets automatically transcribed and then further processed by a text model.We focus on German dialects in the context of written and spoken intent classification -releasing the first dialectal audio intent classification dataset -with supporting experiments on topic classification.The speech-only setup provides the best results on the dialect data while the text-only setup works best on the standard data.While the cascaded systems lag behind the text-only models for German, they perform relatively well on the dialectal data if the transcription system generates normalized, standard-like output. Verena Blaschke, Miriam Winkler, Barbara Plank |
ACL (1) | 3 |
| 2026 | Resource-Lean Lexicon Induction for German Dialects
Robert Litschko, Barbara Plank, Diego Frassinelli |
LREC | 2 |
| 2026 | Variation Is the Norm: Embracing Sociolinguistics in NLP
Anne-Marie Lutgen, Alistair Plum, Verena Blaschke, Barbara Plank, Christoph Purschke |
LREC | 4 |
| 2026 | Information Asymmetry across Language Varieties: A Case Study on Cantonese-Mandarin and Bavarian-German QA
Renhao Pei, Siyao Peng, Verena Blaschke, Robert Litschko, Barbara Plank |
LREC | 5 |
| 2026 | Indirect Question Answering in English, German and Bavarian: A Challenging Task for High- and Low-Resource Languages Alike
Miriam Winkler, Verena Blaschke, Barbara Plank |
LREC | 3 |
| 2026 | I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes
Shijia Zhou, Saif M. Mohammad, Barbara Plank, Diego Frassinelli |
LREC | 3 |
| 2025 | Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies Between Model Predictions and Human Responses in VQAabstractLarge vision-language models struggle to accurately predict responses provided by multiple human annotators, particularly when those responses exhibit high uncertainty. In this study, we focus on a Visual Question Answering (VQA) task and comprehensively evaluate how well the output of the state-of-the-art vision-language model correlates with the distribution of human responses. To do so, we categorize our samples based on their levels (low, medium, high) of human uncertainty in disagreement (HUD) and employ, not only accuracy, but also three new human-correlated metrics for the first time in VQA, to investigate the impact of HUD. We also verify the effect of common calibration and human calibration (Baan et al. 2022) on the alignment of models and humans. Our results show that even BEiT3, currently the best model for this task, struggles to capture the multi-label distribution inherent in diverse human responses. Additionally, we observe that the commonly used accuracy-oriented calibration technique adversely affects BEiT3’s ability to capture HUD, further widening the gap between model predictions and human distributions. In contrast, we show the benefits of calibrating models towards human distributions for VQA, to better align model confidence with human uncertainty. Our findings highlight that for VQA, the alignment between human responses and model predictions is understudied and is an important target for future studies. Diego Frassinelli, Barbara Plank |
AAAI | 3 |
| 2025 | Probing LLMs for Multilingual Discourse Generalization Through a Unified Label SetabstractDiscourse understanding is essential for many NLP tasks, yet most existing work remains constrained by framework-dependent discourse representations.This work investigates whether large language models (LLMs) capture discourse knowledge that generalizes across languages and frameworks.We address this question along two dimensions: (1) developing a unified discourse relation label set to facilitate cross-lingual and cross-framework discourse analysis, and (2) probing LLMs to assess whether they encode generalizable discourse abstractions.Using multilingual discourse relation classification as a testbed, we examine a comprehensive set of 23 LLMs of varying sizes and multilingual capabilities.Our results show that LLMs, especially those with multilingual training corpora, can generalize discourse information across languages and frameworks.Further layer-wise analyses reveal that language generalization at the discourse level is most salient in the intermediate layers.Lastly, our error analysis provides an account of challenging relation classes. Florian Eichin, Yang Janet Liu, Barbara Plank, Michael A. Hedderich |
ACL (1) | 3 |
| 2025 | What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token PatternsabstractPrompt engineering for large language models is challenging, as even small prompt perturbations or model changes can significantly impact the generated output texts.Existing evaluation methods of LLM outputs, either automated metrics or human evaluation, have limitations, such as providing limited insights or being labor-intensive.We propose Spotlight, a new approach that combines both automation and human analysis.Based on data mining techniques, we automatically distinguish between random (decoding) variations and systematic differences in language model outputs.This process provides token patterns that describe the systematic differences and guide the user in manually analyzing the effects of their prompts and changes in models efficiently.We create three benchmarks to quantitatively test the reliability of token pattern extraction methods and demonstrate that our approach provides new insights into established prompt data.From a human-centric perspective, through demonstration studies and a user study, we show that our token pattern approach helps users understand the systematic differences of language model outputs.We are further able to discover relevant differences caused by prompt and model changes (e.g.related to gender or culture), thus supporting the prompt engineering process and human-centric model behavior research. Michael A. Hedderich, Anyi Wang, Raoyuan Zhao, Florian Eichin, Jonas Fischer, Barbara Plank |
ACL (1) | 6 |
| 2025 | Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and ChallengesabstractUnderstanding pragmatics—the use of language in context—is crucial for developing NLP systems capable of interpreting nuanced language use. Despite recent advances in language technologies, including large language models, evaluating their ability to handle pragmatic phenomena such as implicatures and references remains challenging. To advance pragmatic abilities in models, it is essential to understand current evaluation trends and identify existing limitations. In this survey, we provide a comprehensive review of resources designed for evaluating pragmatic capabilities in NLP, categorizing datasets by the pragmatic phenomena they address. We analyze task designs, data collection methods, evaluation approaches, and their relevance to real-world applications. By examining these resources in the context of modern language models, we highlight emerging trends, challenges, and gaps in existing benchmarks. Our survey aims to clarify the landscape of pragmatic evaluation and guide the development of more comprehensive and targeted benchmarks, ultimately contributing to more nuanced and context-aware NLP models. Bolei Ma, Wei Zhou 0067, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank |
ACL (1) | 10 |
| 2025 | Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case StudyabstractBolei Ma, Berk Yoztyurk, Anna-Carolina Haensch, Xinpeng Wang, Markus Herklotz, Frauke Kreuter, Barbara Plank, Matthias Aßenmacher. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Bolei Ma, Berk Yoztyurk, Anna-Carolina Haensch, Xinpeng Wang 0003, Markus Herklotz, Frauke Kreuter, Barbara Plank, Matthias Aßenmacher |
ACL (1) | 7 |
| 2025 | Circuit Compositions: Exploring Modular Structures in Transformer-Based Language ModelsabstractA fundamental question in interpretability research is to what extent neural networks, particularly language models, implement reusable functions through subnetworks that can be composed to perform more complex tasks.Recent advances in mechanistic interpretability have made progress in identifying circuits, which represent the minimal computational subgraphs responsible for a model's behavior on specific tasks.However, most studies focus on identifying circuits for individual tasks without investigating how functionally similar circuits relate to each other.To address this gap, we study the modularity of neural networks by analyzing circuits for highly compositional subtasks within a transformer-based language model.Specifically, given a probabilistic context-free grammar, we identify and compare circuits responsible for ten modular string-edit operations.Our results indicate that functionally similar circuits exhibit both notable node overlap and crosstask faithfulness.Moreover, we demonstrate that the circuits identified can be reused and combined through set operations to represent more complex functional model capabilities. Philipp Mondorf, Sondre Wold, Barbara Plank |
ACL (1) | 3 |
| 2025 | Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance LanguagesabstractA large amount of local and culture-specific knowledge (e.g., people, traditions, food) can only be found in documents written in dialects. While there has been extensive research conducted on cross-lingual information retrieval (CLIR), the field of cross-dialect retrieval (CDIR) has received limited attention. Dialect retrieval poses unique challenges due to the limited availability of resources to train retrieval models and the high variability in non-standardized languages. We study these challenges on the example of German dialects and introduce the first German dialect retrieval dataset, dubbed WikiDIR, which consists of seven German dialects extracted from Wikipedia. Using WikiDIR, we demonstrate the weakness of lexical methods in dealing with high lexical variation in dialects. We further show that commonly used CLIR methods such as query translation or zero-shot cross-lingual transfer with multilingual encoders do not transfer well to extremely low-resource setups, motivating the need for resource-lean and dialect-specific retrieval models. Robert Litschko, Oliver Kraus, Verena Blaschke, Barbara Plank |
COLING | 4 |
| 2025 | Evaluating Pixel Language Models on Non-Standardized LanguagesabstractWe explore the potential of pixel-based models for transfer learning from standard languages to dialects. These models convert text into images that are divided into patches, enabling a continuous vocabulary representation that proves especially useful for out-of-vocabulary words common in dialectal data. Using German as a case study, we compare the performance of pixel-based models to token-based models across various syntactic and semantic tasks. Our results show that pixel-based models outperform token-based models in part-of-speech tagging, dependency parsing and intent detection for zero-shot dialect evaluation by up to 26 percentage points in some scenarios, though not in Standard German. However, pixel-based models fall short in topic classification. These findings emphasize the potential of pixel-based models for handling dialectal data, though further research should be conducted to assess their effectiveness in various linguistic contexts. Alberto Muñoz-Ortiz, Verena Blaschke, Barbara Plank |
COLING | 3 |
| 2025 | Disentangling Subjectivity and Uncertainty for Hate Speech Annotation and Modeling using GazeabstractÖzge Alacam, Sanne Hoeken, Andreas Säuberli, Hannes Gröner, Diego Frassinelli, Sina Zarrieß, Barbara Plank. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Özge Alaçam, Sanne Hoeken, Andreas Säuberli, Hannes Gröner, Diego Frassinelli, Sina Zarrieß, Barbara Plank |
EMNLP | 7 |
| 2025 | The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate ItabstractThe ability of large language models (LLMs) to validate their output and identify potential errors is crucial for ensuring robustness and reliability. However, current research indicates that LLMs struggle with self-correction, encountering significant challenges in detecting errors. While studies have explored methods to enhance self-correction in LLMs, relatively little attention has been given to understanding the models’ internal mechanisms underlying error detection. In this paper, we present a mechanistic analysis of error detection in LLMs, focusing on simple arithmetic problems. Through circuit analysis, we identify the computational subgraphs responsible for detecting arithmetic errors across four smaller-sized LLMs. Our findings reveal that all models heavily rely on \textit{consistency heads}\textemdash{}attention heads that assess surface-level alignment of numerical values in arithmetic solutions. Moreover, we observe that the models’ internal arithmetic computation primarily occurs in higher layers, whereas validation takes place in middle layers, before the final arithmetic results are fully encoded. This structural dissociation between arithmetic computation and validation seems to explain why smaller-sized LLMs struggle to detect even simple arithmetic errors. Leonardo Bertolazzi, Philipp Mondorf, Barbara Plank, Raffaella Bernardi |
EMNLP | 3 |
| 2025 | Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label VariationabstractThe recent rise of reasoning-tuned Large Language Models (LLMs)-which generate chains of thought (CoTs) before giving the final answer-has attracted significant attention and offers new opportunities for gaining insights into human label variation, which refers to plausible differences in how multiple annotators label the same data instance.Prior work has shown that LLM-generated explanations can help align model predictions with human label distributions, but typically adopt a reverse paradigm: producing explanations based on given answers.In contrast, CoTs provide a forward reasoning path that may implicitly embed rationales for each answer option, before generating the answers.We thus propose a novel LLM-based pipeline enriched with linguistically-grounded discourse segmenters to extract supporting and opposing statements for each answer option from CoTs with improved accuracy.We also propose a rank-based HLV evaluation framework that prioritizes the ranking of answers over exact scores, which instead favor direct comparison of label distributions.Our method outperforms a direct generation method as well as baselines on three datasets, and shows better alignment of ranking methods with humans, highlighting the effectiveness of our approach. Beiduo Chen, Yang Janet Liu, Anna Korhonen, Barbara Plank |
EMNLP | 4 |
| 2025 | Reason to Rote: Rethinking Memorization in ReasoningabstractLarge language models readily memorize arbitrary training instances, such as label noise, yet they perform strikingly well on reasoning tasks.In this work, we investigate how language models memorize label noise, and why such memorization in many cases does not heavily affect generalizable reasoning capabilities.Using two controllable synthetic reasoning datasets with noisy labels, four-digit addition (FDA) and two-hop relational reasoning (THR), we discover a reliance of memorization on generalizable reasoning mechanisms: models continue to compute intermediate reasoning outputs even when retrieving memorized noisy labels, and intervening reasoning adversely affects memorization.We further show that memorization operates through distributed encoding, i.e., aggregating various inputs and intermediate results, rather than building a look-up mechanism from inputs to noisy labels.Moreover, our FDA case study reveals memorization occurs via outlier heuristics, where existing neuron activation patterns are slightly shifted to fit noisy labels.Together, our findings suggest that memorization of label noise in language models builds on, rather than overrides, the underlying reasoning mechanisms, shedding lights on the intriguing phenomenon of benign memorization.1 Yupei Du, Philipp Mondorf, Silvia Casola, Yuekun Yao, Robert Litschko, Barbara Plank |
EMNLP | 6 |
| 2025 | LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language InferenceabstractThere is increasing evidence of Human Label Variation (HLV) in Natural Language Inference (NLI), where annotators assign different labels to the same premise-hypothesis pair.However, within-label variation-cases where annotators agree on the same label but provide divergent reasoning-poses an additional and mostly overlooked challenge.Several NLI datasets contain highlighted words in the NLI item as explanations, but the same spans on the NLI item can be highlighted for different reasons, as evidenced by free-text explanations, which offer a window into annotators' reasoning.To systematically understand this problem and gain insight into the rationales behind NLI labels, we introduce LITEX, a linguisticallyinformed taxonomy for categorizing free-text explanations in English.Using this taxonomy, we annotate a subset of the e-SNLI dataset, validate the taxonomy's reliability, and analyze how it aligns with NLI labels, highlights, and explanations.We further assess the taxonomy's role in explanation generation, demonstrating that conditioning generation on LITEX yields explanations that are linguistically closer to human explanations than those generated using only labels or highlights.Our approach thus not only captures within-label variation but also shows how taxonomy-guided generation for reasoning can bridge the gap between human and model explanations more effectively than existing strategies. Pingjun Hong, Beiduo Chen, Siyao Peng, Marie-Catherine de Marneffe, Barbara Plank |
EMNLP | 5 |
| 2025 | RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMsabstractAmbiguity resolution is key to effective communication.While humans effortlessly address ambiguity through conversational grounding strategies, the extent to which current language models can emulate these strategies remains unclear.In this work, we examine referential ambiguity in image-based question answering by introducing RACQUET, a carefully curated dataset targeting distinct aspects of ambiguity.Through a series of evaluations, we reveal significant limitations and problems of overconfidence of state-of-the-art large multimodal language models in addressing ambiguity in their responses.The overconfidence issue becomes particularly relevant for RACQUET-BIAS, a subset designed to analyze a critical yet underexplored problem: failing to address ambiguity leads to stereotypical, socially biased responses.Our results underscore the urgency of equipping models with robust strategies to deal with uncertainty without resorting to undesirable stereotypes. Alberto Testoni, Barbara Plank, Raquel Fernández |
EMNLP | 2 |
| 2025 | M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment AnalysisabstractChengYan Wu, Bolei Ma, Yihong Liu, Zheyu Zhang, Ningyuan Deng, Yanshu Li, Baolan Chen, Yi Zhang, Yun Xue, Barbara Plank. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. ChengYan Wu, Bolei Ma, Yihong Liu 0001, Zheyu Zhang 0007, Ningyuan Deng, Yanshu Li, Baolan Chen, Barbara Plank |
EMNLP | 10 |
| 2025 | Human-centered LLMs for Inclusive Language Technology: The Need to Embrace Variation Holistically in NLPabstractLarge Language Models (LLMs) have advanced rapidly but often still cater primarily to a narrow set of users.This position paper advocates for a human-centered approach to NLP technology-one that embraces linguistic variation, improves reasoning and safety, and better serves diverse communities.We outline key challenges with current LLMs, highlight opportunities in modeling variation in both language and human annotation, and outline a path toward more inclusive and trustworthy language technologies. Barbara Plank |
FedCSIS | 1 |
| 2025 | Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector AblationabstractTraining a language model to be both helpful and harmless requires careful calibration of refusal behaviours: Models should refuse to follow malicious instructions or give harmful advice (e.g."how do I kill someone?"), but they should not refuse safe requests, even if they superficially resemble unsafe ones (e.g. "how do I kill a Python process?"). Avoiding such false refusal, as prior work has shown, is challenging even for highly-capable language models. In this paper, we propose a simple and surgical method for mitigating false refusal in language models via single vector ablation. For a given model, we extract a false refusal vector and show that ablating this vector reduces false refusal rate while preserving the model's safety and general capabilities. We also show that our approach can be used for fine-grained calibration of model safety. Our approach is training-free and model-agnostic, making it useful for mitigating the problem of false refusal in current and future language models. Xinpeng Wang 0003, Chengzhi Hu, Paul Röttger, Barbara Plank |
ICLR | 4 |
| 2025 | References Matter: Investigating the Impact of Reference Set Variation on Summarization EvaluationabstractHuman language production exhibits remarkable richness and variation, reflecting diverse communication styles and intents. However, this variation is often overlooked in summarization evaluation. While having multiple reference summaries is known to improve correlation with human judgments, the impact of the reference set on reference-based metrics has not been systematically investigated. This work examines the sensitivity of widely used reference-based metrics in relation to the choice of reference sets, analyzing three diverse multi-reference summarization datasets: SummEval, GUMSum, and DUC2004. We demonstrate that many popular metrics exhibit significant instability. This instability is particularly concerning for n-gram-based metrics like ROUGE, where model rankings vary depending on the reference sets, undermining the reliability of model comparisons. We also collect human judgments on LLM outputs for genre-diverse data and examine their correlation with metrics to supplement existing findings beyond newswire summaries, finding weak-to-no correlation. Taken together, we recommend incorporating reference set variation into summarization evaluation to enhance consistency alongside correlation with human judgments, especially when evaluating LLMs. Silvia Casola, Yang Janet Liu, Siyao Peng, Oliver Kraus, Albert Gatt, Barbara Plank |
INLG | 6 |
| 2025 | A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech TranslationabstractAlthough Germany has a diverse landscape of dialects, they are underrepresented in current automatic speech recognition (ASR) research. To enable studies of how robust models are towards dialectal variation, we present Betthupferl, an evaluation dataset containing four hours of read speech in three dialect groups spoken in Southeast Germany (Franconian, Bavarian, Alemannic), and half an hour of Standard German speech. We provide both dialectal and Standard German transcriptions, and analyze the linguistic differences between them. We benchmark several multilingual state-of-the-art ASR models on speech translation into Standard German, and find differences between how much the output resembles the dialectal vs. standardized transcriptions. Qualitative error analyses of the best ASR model reveal that it sometimes normalizes grammatical differences, but often stays closer to the dialectal constructions. Verena Blaschke, Miriam Winkler, Constantin Förster, Gabriele Wenger-Glemser, Barbara Plank |
INTERSPEECH | 5 |
| 2025 | Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language ModelsabstractLovish Madaan, David Esiobu, Pontus Stenetorp, Barbara Plank, Dieuwke Hupkes. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Lovish Madaan, David Esiobu, Pontus Stenetorp, Barbara Plank, Dieuwke Hupkes |
NAACL (Long Papers) | 4 |
| 2025 | Refusal Direction is Universal Across Safety-Aligned LanguagesabstractRefusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-centric context, appropriate refusal behavior is important for any language, but poorly understood. In this paper, we investigate the refusal behavior in LLMs across 14 languages using \textit{PolyRefuse}, a multilingual safety dataset created by translating malicious and benign English prompts into these languages. We uncover the surprising cross-lingual universality of the refusal direction: a vector extracted from English can bypass refusals in other languages with near-perfect effectiveness, without any additional fine-tuning. Even more remarkably, refusal directions derived from any safety-aligned language transfer seamlessly to others. We attribute this transferability to the parallelism of refusal vectors across languages in the embedding space and identify the underlying mechanism behind cross-lingual jailbreaks. These findings provide actionable insights for building more robust multilingual safety defenses and pave the way for a deeper mechanistic understanding of cross-lingual vulnerabilities in LLMs. Xinpeng Wang 0003, Mingyang Wang 0003, Yihong Liu 0001, Hinrich Schütze, Barbara Plank |
NeurIPS | 5 |
| 2024 | Comparing Inferential Strategies of Humans and Large Language Models in Deductive ReasoningabstractDeductive reasoning plays a pivotal role in the formulation of sound and cohesive arguments.It allows individuals to draw conclusions that logically follow, given the truth value of the information provided.Recent progress in the domain of large language models (LLMs) has showcased their capability in executing deductive reasoning tasks.Nonetheless, a significant portion of research primarily assesses the accuracy of LLMs in solving such tasks, often overlooking a deeper analysis of their reasoning behavior.In this study, we draw upon principles from cognitive psychology to examine inferential strategies employed by LLMs, through a detailed evaluation of their responses to propositional logic problems.Our findings indicate that LLMs display reasoning patterns akin to those observed in humans, including strategies like supposition following or chain construction.Moreover, our research demonstrates that the architecture and scale of the model significantly affect its preferred method of reasoning, with more advanced models tending to adopt strategies more frequently than less sophisticated ones.Importantly, we assert that a model's accuracy, that is the correctness of its final conclusion, does not necessarily reflect the validity of its reasoning process.This distinction underscores the necessity for more nuanced evaluation procedures in the field. Problem 1:Statements: 1.Either there is a white marble in the box or else there is a black marble in the box, but not both.2. Either there is a black marble in the box or else there is a pink marble in the box, but not both.3.There is a pink marble in the box if and only if there is a gray marble in the box.Conclusion: If there is a white marble in the box then there is a gray marble in the box.Response: Sure, I'd be happy to help!Here's my reasoning process for this problem:1. First, I'll start by analyzing the given statements: From Statement 1, we know that the box contains either a white marble or a black marble, but not both.From Statement 2, we know that the box contains either a black marble or a pink marble, but not both.From Statement 3, we know that if there is a pink marble in the box, then there must be a gray marble in the box.2. Now, let's consider the conclusion: If there is a white marble in the box then there is a gray marble in the box.Can we logically deduce this conclusion from the given statements?Well, if there is a white marble in the box, then by Statement 1, there cannot be a black marble in the box.If there is no black marble in the box, then by Statement 2, there must be a pink marble in the box.But if there is a pink marble in the box, then by Statement 3, there must be a gray marble in the box.So, yes, we can logically deduce the conclusion that if there is a white marble in the box, then there is a gray marble in the box. Philipp Mondorf, Barbara Plank |
ACL (1) | 2 |
| 2024 | VariErr NLI: Separating Annotation Error from Human Label VariationabstractHuman label variation arises when annotators assign different labels to the same item for valid reasons, while annotation errors occur when labels are assigned for invalid reasons.These two issues are prevalent in NLP benchmarks, yet existing research has studied them in isolation.To the best of our knowledge, there exists no prior work that focuses on teasing apart error from signal, especially in cases where signal is beyond black-and-white.To fill this gap, we introduce a systematic methodology and a new dataset, VARIERR (variation versus error), focusing on the NLI task in English.We propose a 2-round annotation procedure with annotators explaining each label and subsequently judging the validity of label-explanation pairs.VARIERR contains 7,732 validity judgments on 1,933 explanations for 500 re-annotated MNLI items.We assess the effectiveness of various automatic error detection (AED) methods and GPTs in uncovering errors versus human label variation.We find that state-of-the-art AED methods significantly underperform GPTs and humans.While GPT-4 is the best system, it still falls short of human performance.Our methodology is applicable beyond NLI, offering fertile ground for future research on error versus plausible variation, which in turn can yield better and more trustworthy NLP systems. Leon Weber-Genzel, Siyao Peng, Marie-Catherine de Marneffe, Barbara Plank |
ACL (1) | 4 |
| 2024 | Through the Lens of Split Vote: Exploring Disagreement, Difficulty and Calibration in Legal Case Outcome ClassificationabstractIn legal decisions, split votes (SV) occur when judges cannot reach a unanimous decision, posing a difficulty for lawyers who must navigate diverse legal arguments and opinions.In high-stakes domains, understanding the alignment of perceived difficulty between humans and AI systems is crucial to build trust.However, existing NLP calibration methods focus on a classifier's awareness of predictive performance, measured against the human majority class, overlooking inherent human label variation (HLV).This paper explores split votes as naturally observable human disagreement and value pluralism.We collect judges' vote distributions from the European Court of Human Rights (ECHR), and present SV-ECHR 1 a case outcome classification (COC) dataset with SV information.We build a taxonomy of disagreement with SV-specific subcategories.We further assess the alignment of perceived difficulty between models and humans, as well as confidence-and human-calibration of COC models.We observe limited alignment with the judge vote distribution.To our knowledge, this is the first systematic exploration of calibration to human judgements in legal NLP.Our study underscores the necessity for further research on measuring and enhancing model calibration considering HLV in legal decision tasks.* Following Chalkidis et al. 2022a; Santosh et al. 2022 We use only the 10 most prominent ECHR articles.* https://hudoc.echr.coe.int* App B offers details on the quality assessment process.* See App D for more details of our metadata correction.* For a comprehensive understanding of each taxonomy category, we direct the reader to Xu et al. 2023b T. Y. S. S. Santosh, Oana Ichim, Barbara Plank, Matthias Grabmair |
ACL (1) | 4 |
| 2024 | How to Encode Domain Information in Relation ClassificationabstractCurrent language models require a lot of training data to obtain high performance. For Relation Classification (RC), many datasets are domain-specific, so combining datasets to obtain better performance is non-trivial. We explore a multi-domain training setup for RC, and attempt to improve performance by encoding domain information. Our proposed models improve > 2 Macro-F1 against the baseline setup, and our analysis reveals that not all the labels benefit the same: The classes which occupy a similar space across domains (i.e., their interpretation is close across them, for example “physical”) benefit the least, while domain-dependent relations (e.g., “part-of”) improve the most when encoding domain information. Elisa Bassignana, Viggo Unmack Gascou, Frida Nøhr Laustsen, Gustav Kristensen, Marie Haahr Petersen, Rob van der Goot, Barbara Plank |
LREC/COLING | 7 |
| 2024 | MaiBaam: A Multi-Dialectal Bavarian Universal Dependency TreebankabstractDespite the success of the Universal Dependencies (UD) project exemplified by its impressive language breadth, there is still a lack in ‘within-language breadth’: most treebanks focus on standard languages. Even for German, the language with the most annotations in UD, so far no treebank exists for one of its language varieties spoken by over 10M people: Bavarian. To contribute to closing this gap, we present the first multi-dialect Bavarian treebank (MaiBaam) manually annotated with part-of-speech and syntactic dependency information in UD, covering multiple text genres (wiki, fiction, grammar examples, social, non-fiction). We highlight the morphosyntactic differences between the closely-related Bavarian and German and showcase the rich variability of speakers’ orthographies. Our corpus includes 15k tokens, covering dialects from all Bavarian-speaking areas spanning three countries. We provide baseline parsing and POS tagging results, which are lower than results obtained on German and vary substantially between different graph-based parsers. To support further research on Bavarian syntax, we make our dataset, language-specific guidelines and code publicly available. Verena Blaschke, Barbara Kovacic, Siyao Peng, Hinrich Schütze, Barbara Plank |
LREC/COLING | 5 |
| 2024 | IndirectQA: Understanding Indirect Answers to Implicit Polar Questions in French and SpanishabstractPolar questions are common in dialogue and expect exactly one of two answers (yes/no). It is however not uncommon for speakers to bypass these expected choices and answer, for example, “Islands are generally by the sea” to the question: “An island? By the sea?”. While such answers are natural in spoken dialogues, conversational systems still struggle to interpret them. Seminal work to interpret indirect answers were made in recent years—but only for English and with strict question formulations. In this work, we present a new corpus for French and Spanish—IndirectQA —where we mine subtitle data for indirect answers to study the labeling task with six different labels, while broadening polar questions to include also implicit polar questions (statements that trigger a yes/no-answer which are not necessarily formulated as a question). We opted for subtitles since they are a readily available source of conversation in various languages, but also come with peculiarities and challenges which we will discuss. Overall, we provide the first results on French and Spanish. They show that the task is challenging: the baseline accuracy scores drop from 61.43 on English to 44.06 for French and Spanish. Christin Müller, Barbara Plank |
LREC/COLING | 2 |
| 2024 | Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal DataabstractNamed Entity Recognition (NER) is a fundamental task to extract key information from texts, but annotated resources are scarce for dialects. This paper introduces the first dialectal NER dataset for German, BarNER, with 161K tokens annotated on Bavarian Wikipedia articles (bar-wiki) and tweets (bar-tweet), using a schema adapted from German CoNLL 2006 and GermEval. The Bavarian dialect differs from standard German in lexical distribution, syntactic construction, and entity information. We conduct in-domain, cross-domain, sequential, and joint experiments on two Bavarian and three German corpora and present the first comprehensive NER results on Bavarian. Incorporating knowledge from the larger German NER (sub-)datasets notably improves on bar-wiki and moderately on bar-tweet. Inversely, training first on Bavarian contributes slightly to the seminal German CoNLL 2006 corpus. Moreover, with gold dialect labels on Bavarian tweets, we assess multi-task learning between five NER and two Bavarian-German dialect identification tasks and achieve NER SOTA on bar-wiki. We substantiate the necessity of our low-resource BarNER corpus and the importance of diversity in dialects, genres, and topics in enhancing model performance. Siyao Peng, Zihang Sun, Huangyan Shan, Marie Kolm, Verena Blaschke, Ekaterina Artemova, Barbara Plank |
LREC/COLING | 7 |
| 2024 | Slot and Intent Detection Resources for Bavarian and Lithuanian: Assessing Translations vs Natural Queries to Digital AssistantsabstractDigital assistants perform well in high-resource languages like English, where tasks like slot and intent detection (SID) are well-supported. Many recent SID datasets start including multiple language varieties. However, it is unclear how realistic these translated datasets are. Therefore, we extend one such dataset, namely xSID-0.4, to include two underrepresented languages: Bavarian, a German dialect, and Lithuanian, a Baltic language. Both language variants have limited speaker populations and are often not included in multilingual projects. In addition to translations we provide “natural” queries to digital assistants generated by native speakers. We further include utterances from another dataset for Bavarian to build the richest SID dataset available today for a low-resource dialect without standard orthography. We then set out to evaluate models trained on English in a zero-shot scenario on our target language variants. Our evaluation reveals that translated data can produce overly optimistic scores. However, the error patterns in translated and natural datasets are highly similar. Cross-dataset experiments demonstrate that data collection methods influence performance, with scores lower than those achieved with single-dataset translations. This work contributes to enhancing SID datasets for underrepresented languages, yielding NaLiBaSID, a new evaluation dataset for Bavarian and Lithuanian. Miriam Winkler, Virginija Juozapaityte, Rob van der Goot, Barbara Plank |
LREC/COLING | 4 |
| 2024 | Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German VarietiesabstractMainstream cross-lingual task-oriented dialogue (ToD) systems leverage the transfer learning paradigm by training a joint model for intent recognition and slot-filling in English and applying it, zero-shot, to other languages.We address a gap in prior research, which often overlooked the transfer to lower-resource colloquial varieties due to limited test data.Inspired by prior work on English varieties, we craft and manually evaluate perturbation rules that transform German sentences into colloquial forms and use them to synthesize test sets in four ToD datasets.Our perturbation rules cover 18 distinct language phenomena, enabling us to explore the impact of each perturbation on slot and intent performance.Using these new datasets, we conduct an experimental evaluation across six different transformers.Here, we demonstrate that when applied to colloquial varieties, ToD systems maintain their intent recognition performance, losing 6% (4.62 percentage points) in accuracy on average.However, they exhibit a significant drop in slot detection, with a decrease of 31% (21 percentage points) in slot F 1 score.Our findings are further supported by a transfer experiment from Standard American English to synthetic Urban African American Vernacular English. Ekaterina Artemova, Verena Blaschke, Barbara Plank |
EACL (1) | 3 |
| 2024 | NNOSE: Nearest Neighbor Occupational Skill ExtractionabstractThe labor market is changing rapidly, prompting increased interest in the automatic extraction of occupational skills from text. With the advent of English benchmark job description datasets, there is a need for systems that handle their diversity well. We tackle the complexity in occupational skill datasets tasks—combining and leveraging multiple datasets for skill extraction, to identify rarely observed skills within a dataset, and overcoming the scarcity of skills across datasets. In particular, we investigate the retrieval-augmentation of language models, employing an external datastore for retrieving similar skills in a dataset-unifying manner. Our proposed method, Nearest Neighbor Occupational Skill Extraction (NNOSE) effectively leverages multiple datasets by retrieving neighboring skills from other datasets in the datastore. This improves skill extraction without additional fine-tuning. Crucially, we observe a performance gain in predicting infrequent patterns, with substantial gains of up to 30% span-F1 in cross-dataset settings. Mike Zhang, Rob van der Goot, Min-Yen Kan, Barbara Plank |
EACL (1) | 4 |
| 2024 | Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language ModelsabstractKnights and knaves problems represent a classic genre of logical puzzles where characters either tell the truth or lie.The objective is to logically deduce each character's identity based on their statements.The challenge arises from the truth-telling or lying behavior, which influences the logical implications of each statement.Solving these puzzles requires not only direct deductions from individual statements, but the ability to assess the truthfulness of statements by reasoning through various hypothetical scenarios.As such, knights and knaves puzzles serve as compelling examples of suppositional reasoning.In this paper, we introduce TruthQuest, a benchmark for suppositional reasoning based on the principles of knights and knaves puzzles.Our benchmark presents problems of varying complexity, considering both the number of characters and the types of logical statements involved.Evaluations on TruthQuest show that large language models like Llama 3 and Mixtral-8x7B exhibit significant difficulties solving these tasks.A detailed error analysis of the models' output reveals that lower-performing models exhibit a diverse range of reasoning errors, frequently failing to grasp the concept of truth and lies.In comparison, more proficient models primarily struggle with accurately inferring the logical implications of potentially false statements. Philipp Mondorf, Barbara Plank |
EMNLP | 2 |
| 2024 | Position: Insights from Survey Methodology can Improve Training DataabstractWhether future AI models are fair, trustworthy, and aligned with the public's interests rests in part on our ability to collect accurate data about what we want the models to do. However, collecting high-quality data is difficult, and few AI/ML researchers are trained in data collection methods. Recent research in data-centric AI has show that higher quality training data leads to better performing models, making this the right moment to introduce AI/ML researchers to the field of survey methodology, the science of data collection. We summarize insights from the survey methodology literature and discuss how they can improve the quality of training and feedback data. We also suggest collaborative research ideas into how biases in data collection can be mitigated, making models more accurate and human-centric. Stephanie Eckman, Barbara Plank, Frauke Kreuter |
ICML | 2 |
| 2024 | Universal NER: A Gold-Standard Multilingual Named Entity Recognition BenchmarkabstractStephen Mayhew, Terra Blevins, Shuheng Liu, Marek Šuppa, Hila Gonen, Joseph Marvin Imperial, Börje F. Karlsson, Peiqin Lin, Nikola Ljubešić, LJ Miranda, Barbara Plank, Arij Riabi, Yuval Pinter. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Stephen Mayhew 0002, Terra Blevins, Shuheng Liu 0002, Marek Suppa, Hila Gonen, Joseph Marvin Imperial, Börje Karlsson 0001, Peiqin Lin, Nikola Ljubesic, Lester James V. Miranda, Barbara Plank, Arij Riabi, Yuval Pinter |
NAACL-HLT | 11 |
| 2023 | ESCOXLM-R: Multilingual Taxonomy-driven Pre-training for the Job Market DomainabstractThe increasing number of benchmarks for Natural Language Processing (NLP) tasks in the computational job market domain highlights the demand for methods that can handle job-related tasks such as skill extraction, skill classification, job title classification, and de-identification.While some approaches have been developed that are specific to the job market domain, there is a lack of generalized, multilingual models and benchmarks for these tasks.In this study, we introduce a language model called ESCOXLM-R, based on XLM-R large , which uses domain-adaptive pre-training on the European Skills, Competences, Qualifications and Occupations (ESCO) taxonomy, covering 27 languages.The pre-training objectives for ESCOXLM-R include dynamic masked language modeling and a novel additional objective for inducing multilingual taxonomical ESCO relations.We comprehensively evaluate the performance of ESCOXLM-R on 6 sequence labeling and 3 classification tasks in 4 languages and find that it achieves state-of-the-art results on 6 out of 9 datasets.Our analysis reveals that ESCOXLM-R performs better on short spans and outperforms XLM-R large on entity-level and surface-level span-F1, likely due to ESCO containing short skill and occupation titles, and encoding information on the entity-level. Mike Zhang, Rob van der Goot, Barbara Plank |
ACL (1) | 3 |
| 2023 | What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production VariabilityabstractIn Natural Language Generation (NLG) tasks, for any input, multiple communicative goals are plausible, and any goal can be put into words, or produced, in multiple ways.We characterise the extent to which human production varies lexically, syntactically, and semantically across four NLG tasks, connecting human production variability to aleatoric or data uncertainty.We then inspect the space of output strings shaped by a generation system's predicted probability distribution and decoding algorithm to probe its uncertainty.For each test input, we measure the generator's calibration to human production variability.Following this instance-level approach, we analyse NLG models and decoding strategies, demonstrating that probing a generator with multiple samples and, when possible, multiple references, provides the level of detail necessary to gain understanding of a model's representation of uncertainty. 1 * Equal contribution. 1 https://github.com/dmg-illc/nlg-uncertainty-probes Mario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández, Barbara Plank |
EMNLP | 5 |
| 2023 | Establishing Trustworthiness: Rethinking Tasks and Model EvaluationabstractLanguage understanding is a multi-faceted cognitive capability, which the Natural Language Processing (NLP) community has striven to model computationally for decades.Traditionally, facets of linguistic intelligence have been compartmentalized into tasks with specialized model architectures and corresponding evaluation protocols.With the advent of large language models (LLMs) the community has witnessed a dramatic shift towards general purpose, task-agnostic approaches powered by generative models.As a consequence, the traditional compartmentalized notion of language tasks is breaking down, followed by an increasing challenge for evaluation and analysis.At the same time, LLMs are being deployed in more real-world scenarios, including previously unforeseen zero-shot setups, increasing the need for trustworthy and reliable systems.Therefore, we argue that it is time to rethink what constitutes tasks and model evaluation in NLP, and pursue a more holistic view on language, placing trustworthiness at the center.Towards this goal, we review existing compartmentalized approaches for understanding the origins of a model's functional capacity, and provide recommendations for more multifaceted evaluation protocols."Trust arises from knowledge of origin as well as from knowledge of functional capacity." Robert Litschko, Max Müller-Eberstein, Rob van der Goot, Leon Weber-Genzel, Barbara Plank |
EMNLP | 5 |
| 2023 | ACTOR: Active Learning with Annotator-specific Classification Heads to Embrace Human Label VariationabstractLabel aggregation such as majority voting is commonly used to resolve annotator disagreement in dataset creation.However, this may disregard minority values and opinions.Recent studies indicate that learning from individual annotations outperforms learning from aggregated labels, though they require a considerable amount of annotation.Active learning, as an annotation cost-saving strategy, has not been fully explored in the context of learning from disagreement.We show that in the active learning setting, a multi-head model performs significantly better than a single-head model in terms of uncertainty estimation.By designing and evaluating acquisition functions with annotator-specific heads on two datasets, we show that group-level entropy works generally well on both datasets.Importantly, it achieves performance in terms of both prediction and uncertainty estimation comparable to full-scale training from disagreement, while saving 70% of the annotation budget. Key FindingsWe made several key observations:• The multi-head model works significantly better than the single-head model on uncertainty estimation. Xinpeng Wang 0003, Barbara Plank |
EMNLP | 2 |
| 2023 | From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome ClassificationabstractIn legal NLP, Case Outcome Classification (COC) must not only be accurate but also trustworthy and explainable.Existing work in explainable COC has been limited to annotations by a single expert.However, it is well-known that lawyers may disagree in their assessment of case facts.We hence collect a novel dataset RAVE: Rationale Variation in ECHR 1 , which is obtained from two experts in the domain of international human rights law, for whom we observe weak agreement.We study their disagreements and build a two-level task-independent taxonomy, supplemented with COC-specific subcategories.We quantitatively assess different taxonomy categories and find that disagreements mainly stem from underspecification of the legal context, which poses challenges given the typically limited granularity and noise in COC metadata.To our knowledge, this is the first work in the legal NLP that focuses on building a taxonomy over human label variation.We further assess the explainablility of state-of-the-art COC models on RAVE and observe limited agreement between models and experts.Overall, our case study reveals hitherto underappreciated complexities in creating benchmark datasets in legal NLP that revolve around identifying aspects of a case's facts supposedly relevant to its outcome. T. Y. S. S. Santosh, Oana Ichim, Isabella Risini, Barbara Plank, Matthias Grabmair |
EMNLP | 5 |
| 2022 | Probing for Labeled Dependency TreesabstractProbing has become an important tool for analyzing representations in Natural Language Processing (NLP). For graphical NLP tasks such as dependency parsing, linear probes are currently limited to extracting undirected or unlabeled parse trees which do not capture the full task. This work introduces DepProbe, a linear probe which can extract labeled and directed dependency parse trees from embeddings while using fewer parameters and compute than prior methods. Leveraging its full task coverage and lightweight parametrization, we investigate its predictive power for selecting the best transfer language for training a full biaffine attention parser. Across 13 languages, our proposed method identifies the best source treebank 94% of the time, outperforming competitive baselines and prior work. Finally, we analyze the informativeness of task-specific subspaces in contextual embeddings as well as which benefits a full parser’s non-linear parametrization provides. Max Müller-Eberstein, Rob van der Goot, Barbara Plank |
ACL (1) | 3 |
| 2022 | On Language Spaces, Scales and Cross-Lingual Transfer of UD ParsersabstractTanja Samardžić, Ximena Gutierrez-Vasques, Rob van der Goot, Max Müller-Eberstein, Olga Pelloni, Barbara Plank. Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL). 2022. Tanja Samardzic, Ximena Gutierrez-Vasques, Rob van der Goot, Max Müller-Eberstein, Olga Pelloni, Barbara Plank |
CoNLL | 6 |
| 2022 | Stop Measuring Calibration When Humans DisagreeabstractCalibration is a popular framework to evaluate whether a classifier knows when it does not know-i.e., its predictive probabilities are a good indication of how likely a prediction is to be correct.Correctness is commonly estimated against the human majority class.Recently, calibration to human majority has been measured on tasks where humans inherently disagree about which class applies.We show that measuring calibration to human majority given inherent disagreements is theoretically problematic, demonstrate this empirically on the ChaosNLI dataset, and derive several instancelevel measures of calibration that capture key statistical properties of human judgementsclass frequency, ranking and entropy. 1 Joris Baan, Wilker Aziz, Barbara Plank, Raquel Fernández |
EMNLP | 3 |
| 2022 | Evidence \textgreater Intuition: Transferability Estimation for Encoder SelectionabstractWith the increase in availability of large pre-trained language models (LMs) in Natural Language Processing (NLP), it becomes critical to assess their fit for a specific target task a priori-as fine-tuning the entire space of available LMs is computationally prohibitive and unsustainable.However, encoder transferability estimation has received little to no attention in NLP.In this paper, we propose to generate quantitative evidence to predict which LM, out of a pool of models, will perform best on a target task without having to fine-tune all candidates.We provide a comprehensive study on LM ranking for 10 NLP tasks spanning the two fundamental problem types of classification and structured prediction.We adopt the state-of-the-art Logarithm of Maximum Evidence (LogME) measure from Computer Vision (CV) and find that it positively correlates with final LM performance in 94% of the setups.In the first study of its kind, we further compare transferability measures with the de facto standard of human practitioner ranking, finding that evidence from quantitative metrics is more robust than pure intuition and can help identify unexpected LM candidates. Elisa Bassignana, Max Müller-Eberstein, Mike Zhang, Barbara Plank |
EMNLP | 4 |
| 2022 | Spectral ProbingabstractLinguistic information is encoded at varying timescales (subwords, phrases, etc.) and communicative levels, such as syntax and semantics.Contextualized embeddings have analogously been found to capture these phenomena at distinctive layers and frequencies.Leveraging these findings, we develop a fully learnable frequency filter to identify spectral profiles for any given task.It enables vastly more granular analyses than prior handcrafted filters, and improves on efficiency.After demonstrating the informativeness of spectral probing over manual filters in a monolingual setting, we investigate its multilingual characteristics across seven diverse NLP tasks in six languages.Our analyses identify distinctive spectral profiles which quantify cross-task similarity in a linguistically intuitive manner, while remaining consistent across languages-highlighting their potential as robust, lightweight task descriptors. Max Müller-Eberstein, Rob van der Goot, Barbara Plank |
EMNLP | 3 |
| 2022 | The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and EvaluationabstractHuman variation in labeling is often considered noise.Annotation projects for machine learning (ML) aim at minimizing human label variation, with the assumption to maximize data quality and in turn optimize and maximize machine learning metrics.However, this conventional practice assumes that there exists a ground truth, and neglects that there exists genuine human variation in labeling due to disagreement, subjectivity in annotation or multiple plausible answers.In this position paper, we argue that this big open problem of human label variation persists and critically needs more attention to move our field forward.This is because human label variation impacts all stages of the ML pipeline: data, modeling and evaluation.However, few works consider all of these dimensions jointly; and existing research is fragmented.We reconcile different previously proposed notions of human label variation, provide a repository of publicly-available datasets with un-aggregated labels, depict approaches proposed so far, identify gaps and suggest ways forward.As datasets are becoming increasingly available, we hope that this synthesized view on the "problem" will lead to an open discussion on possible strategies to devise fundamentally new directions. Barbara Plank |
EMNLP | 1 |
| 2022 | Frustratingly Easy Performance Improvements for Low-resource Setups: A Tale on BERT and Segment EmbeddingsabstractAs input representation for each sub-word, the original BERT architecture proposes the sum of the sub-word embedding, position embedding and a segment embedding. Sub-word and position embeddings are well-known and studied, and encode lexical information and word position, respectively. In contrast, segment embeddings are less known and have so far received no attention, despite being ubiquitous in large pre-trained language models. The key idea of segment embeddings is to encode to which of the two sentences (segments) a word belongs to — the intuition is to inform the model about the separation of sentences for the next sentence prediction pre-training task. However, little is known on whether the choice of segment impacts performance. In this work, we try to fill this gap and empirically study the impact of the segment embedding during inference time for a variety of pre-trained embeddings and target tasks. We hypothesize that for single-sentence prediction tasks performance is not affected — neither in mono- nor multilingual setups — while it matters when swapping segment IDs in paired-sentence tasks. To our surprise, this is not the case. Although for classification tasks and monolingual BERT models no large differences are observed, particularly word-level multilingual prediction tasks are heavily impacted. For low-resource syntactic tasks, we observe impacts of segment embedding and multilingual BERT choice. We find that the default setting for the most used multilingual BERT model underperforms heavily, and a simple swap of the segment embeddings yields an average improvement of 2.5 points absolute LAS score for dependency parsing over 9 different treebanks. Rob van der Goot, Max Müller-Eberstein, Barbara Plank |
LREC | 3 |
| 2022 | Fine-tuning vs From Scratch: Do Vision & Language Models Have Similar Capabilities on Out-of-Distribution Visual Question Answering?abstractFine-tuning general-purpose pre-trained models has become a de-facto standard, also for Vision and Language tasks such as Visual Question Answering (VQA). In this paper, we take a step back and ask whether a fine-tuned model has superior linguistic and reasoning capabilities than a prior state-of-the-art architecture trained from scratch on the training data alone. We perform a fine-grained evaluation on out-of-distribution data, including an analysis on robustness due to linguistic variation (rephrasings). Our empirical results confirm the benefit of pre-training on overall performance and rephrasing in particular. But our results also uncover surprising limitations, particularly for answering questions involving boolean operations. To complement the empirical evaluation, this paper also surveys relevant earlier work on 1) available VQA data sets, 2) models developed for VQA, 3) pre-trained Vision+Language models, and 4) earlier fine-grained evaluation of pre-trained Vision+Language models. Kristian Nørgaard Jensen, Barbara Plank |
LREC | 2 |
| 2022 | Kompetencer: Fine-grained Skill Classification in Danish Job Postings via Distant Supervision and Transfer LearningabstractSkill Classification (SC) is the task of classifying job competences from job postings. This work is the first in SC applied to Danish job vacancy data. We release the first Danish job posting dataset: Kompetencer (en: competences), annotated for nested spans of competences. To improve upon coarse-grained annotations, we make use of The European Skills, Competences, Qualifications and Occupations (ESCO; le Vrang et al., (2014)) taxonomy API to obtain fine-grained labels via distant supervision. We study two setups: The zero-shot and few-shot classification setting. We fine-tune English-based models and RemBERT (Chung et al., 2020) and compare them to in-language Danish models. Our results show RemBERT significantly outperforms all other models in both the zero-shot and the few-shot setting. Mike Zhang, Kristian Nørgaard Jensen, Barbara Plank |
LREC | 3 |
| 2022 | Sort by Structure: Language Model Ranking as Dependency ProbingabstractMaking an informed choice of pre-trained language model (LM) is critical for performance, yet environmentally costly, and as such widely underexplored. The field of Computer Vision has begun to tackle encoder ranking, with promising forays into Natural Language Processing, however they lack coverage of linguistic tasks such as structured prediction. We propose probing to rank LMs, specifically for parsing dependencies in a given language, by measuring the degree to which labeled trees are recoverable from an LM’s contextualized embeddings. Across 46 typologically and architecturally diverse LM-language pairs, our probing approach predicts the best LM choice 79% of the time using orders of magnitude less compute than training a full parser. Within this study, we identify and analyze one recently proposed decoupled LM—RemBERT—and find it strikingly contains less inherent dependency information, but often yields the best parser after full fine-tuning. Without this outlier our approach identifies the best LM in 89% of cases. Max Müller-Eberstein, Rob van der Goot, Barbara Plank |
NAACL-HLT | 3 |
| 2022 | SkillSpan: Hard and Soft Skill Extraction from English Job PostingsabstractMike Zhang, Kristian Jensen, Sif Sonniks, Barbara Plank. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Mike Zhang, Kristian Nørgaard Jensen, Sif Dam Sonniks, Barbara Plank |
NAACL-HLT | 4 |
| 2022 | Neural Natural Language Generation: A Survey on Multilinguality, Multimodality, Controllability and LearningabstractDeveloping artificial learning systems that can understand and generate natural language has been one of the long-standing goals of artificial intelligence. Recent decades have witnessed an impressive progress on both of these problems, giving rise to a new family of approaches. Especially, the advances in deep learning over the past couple of years have led to neural approaches to natural language generation (NLG). These methods combine generative language learning techniques with neural-networks based frameworks. With a wide range of applications in natural language processing, neural NLG (NNLG) is a new and fast growing field of research. In this state-of-the-art report, we investigate the recent developments and applications of NNLG in its full extent from a multidimensional view, covering critical perspectives such as multimodality, multilinguality, controllability and learning strategies. We summarize the fundamental building blocks of NNLG approaches from these aspects and provide detailed reviews of commonly used preprocessing steps and basic neural architectures. This report also focuses on the seminal applications of these NNLG models such as machine translation, description generation, automatic speech recognition, abstractive summarization, text simplification, question answering and generation, and dialogue generation. Finally, we conclude with a thorough discussion of the described frameworks by pointing out some open research directions. Erkut Erdem, Menekse Kuyu, Semih Yagcioglu, Anette Frank, Letitia Parcalabescu, Barbara Plank, Andrii Babii, Oleksii Turuta, Aykut Erdem, Iacer Calixto, Elena Lloret, Elena Apostol, Ciprian-Octavian Truica, Branislava Sandrih, Sanda Martincic-Ipsic, Gábor Berend, Albert Gatt, Grazina Korvel |
J. Artif. Intell. Res. | 6 |
| 2021 | Genre as Weak Supervision for Cross-lingual Dependency ParsingabstractRecent work has shown that monolingual masked language models learn to represent data-driven notions of language variation which can be used for domain-targeted training data selection.Dataset genre labels are already frequently available, yet remain largely unexplored in cross-lingual setups.We harness this genre metadata as a weak supervision signal for targeted data selection in zeroshot dependency parsing.Specifically, we project treebank-level genre information to the finer-grained sentence level, with the goal to amplify information implicitly stored in unsupervised contextualized representations.We demonstrate that genre is recoverable from multilingual contextual embeddings and that it provides an effective signal for training data selection in cross-lingual, zero-shot scenarios.For 12 low-resource language treebanks, six of which are test-only, our genre-specific methods significantly outperform competitive baselines as well as recent embedding-based methods for data selection.Moreover, genre-based data selection provides new state-of-the-art results for three of these target languages. Max Müller-Eberstein, Rob van der Goot, Barbara Plank |
EMNLP (1) | 3 |
| 2021 | Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task LearningabstractTommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, Massimo Poesio. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, Massimo Poesio |
NAACL-HLT | 4 |
| 2021 | From Masked Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zero-shot Spoken Language UnderstandingabstractRob van der Goot, Ibrahim Sharaf, Aizhan Imankulova, Ahmet Üstün, Marija Stepanović, Alan Ramponi, Siti Oryza Khairunnisa, Mamoru Komachi, Barbara Plank. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Rob van der Goot, Ibrahim Sharaf, Aizhan Imankulova, Ahmet Üstün, Marija Stepanovic, Alan Ramponi, Siti Oryza Khairunnisa, Mamoru Komachi, Barbara Plank |
NAACL-HLT | 9 |
| 2021 | Learning from Disagreement: A SurveyabstractMany tasks in Natural Language Processing (NLP) and Computer Vision (CV) offer evidence that humans disagree, from objective tasks such as part-of-speech tagging to more subjective tasks such as classifying an image or deciding whether a proposition follows from certain premises. While most learning in artificial intelligence (AI) still relies on the assumption that a single (gold) interpretation exists for each item, a growing body of research aims to develop learning methods that do not rely on this assumption. In this survey, we review the evidence for disagreements on NLP and CV tasks, focusing on tasks for which substantial datasets containing this information have been created. We discuss the most popular approaches to training models from datasets containing multiple judgments potentially in disagreement. We systematically compare these different approaches by training them with each of the available datasets, considering several ways to evaluate the resulting models. Finally, we discuss the results in depth, focusing on four key research questions, and assess how the type of evaluation and the characteristics of a dataset determine the answers to these questions. Our results suggest, first of all, that even if we abandon the assumption of a gold standard, it is still essential to reach a consensus on how to evaluate models. This is because the relative performance of the various training methods is critically affected by the chosen form of evaluation. Secondly, we observed a strong dataset effect. With substantial datasets, providing many judgments by high-quality coders for each item, training directly with soft labels achieved better results than training from aggregated or even gold labels. This result holds for both hard and soft evaluation. But when the above conditions do not hold, leveraging both gold and soft labels generally achieved the best results in the hard evaluation. All datasets and models employed in this paper are freely available as supplementary materials. Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio |
J. Artif. Intell. Res. | 5 |
| 2020 | DaN+: Danish Nested Named Entities and Lexical NormalizationabstractThis paper introduces DAN+, a new multi-domain corpus and annotation guidelines for Danish nested named entities (NEs) and lexical normalization to support research on cross-lingual cross-domain learning for a less-resourced language.We empirically assess three strategies to model the two-layer Named Entity Recognition (NER) task.We compare transfer capabilities from German versus in-language annotation from scratch.We examine language-specific versus multilingual BERT, and study the effect of lexical normalization on NER.Our results show that 1) the most robust strategy is multi-task learning which is rivaled by multi-label decoding, 2) BERT-based NER models are sensitive to domain shifts, and 3) in-language BERT and lexical normalization are the most beneficial on the least canonical data.Our results also show that an out-of-domain setup remains challenging, while performance on news plateaus quickly.This highlights the importance of cross-domain evaluation of cross-lingual transfer. Barbara Plank, Kristian Nørgaard Jensen, Rob van der Goot |
COLING | 1 |
| 2020 | Neural Unsupervised Domain Adaptation in NLP - A SurveyabstractDeep neural networks excel at learning from labeled data and achieve state-of-the-art results on a wide array of Natural Language Processing tasks.In contrast, learning from unlabeled data, especially under domain shift, remains a challenge.Motivated by the latest advances, in this survey we review neural unsupervised domain adaptation techniques which do not require labeled target domain data.This is a more challenging yet a more widely applicable setup.We outline methods, from early traditional non-neural methods to pre-trained model transfer.We also revisit the notion of domain, and we uncover a bias in the type of Natural Language Processing tasks which received most attention.Lastly, we outline future directions, particularly the broader need for out-of-distribution generalization of future NLP. 1 Alan Ramponi, Barbara Plank |
COLING | 2 |
| 2020 | Biomedical Event Extraction as Sequence LabelingabstractWe introduce Biomedical Event Extraction as Sequence Labeling (BEESL), a joint endto-end neural information extraction model.BEESL recasts the task as sequence labeling, taking advantage of a multi-label aware encoding strategy and jointly modeling the intermediate tasks via multi-task learning.BEESL is fast, accurate, end-to-end, and unlike current methods does not require any external knowledge base or preprocessing tools.BEESL outperforms the current best system (Li et al., 2019) on the Genia 2011 benchmark by 1.57% absolute F1 score reaching 60.22% F1, establishing a new state of the art for the task.Importantly, we also provide first results on biomedical event extraction without gold entity information.Empirical results show that BEESL's speed and accuracy makes it a viable approach for large-scale real-world scenarios.1 Alan Ramponi, Rob van der Goot, Rosario Lombardo, Barbara Plank |
EMNLP (1) | 4 |
| 2020 | A Case for Soft Loss FunctionsabstractRecently, Peterson et al. provided evidence of the benefits of using probabilistic soft labels generated from crowd annotations for training a computer vision model, showing that using such labels maximizes performance of the models over unseen data. In this paper, we generalize these results by showing that training with soft labels is an effective method for using crowd annotations in several other ai tasks besides the one studied by Peterson et al., and also when their performance is compared with that of state-of-the-art methods for learning from crowdsourced data. Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio |
HCOMP | 5 |
| 2020 | FT Speech: Danish Parliament Speech CorpusabstractThis paper introduces FT Speech, a new speech corpus created from the recorded meetings of the Danish Parliament, otherwise known as the Folketing (FT). The corpus contains over 1,800 hours of transcribed speech by a total of 434 speakers. It is significantly larger in duration, vocabulary, and amount of spontaneous speech than the existing public speech corpora for Danish, which are largely limited to read-aloud and dictation data. We outline design considerations, including the preprocessing methods and the alignment procedure. To evaluate the quality of the corpus, we train automatic speech recognition systems on the new resource and compare them to the systems trained on the Danish part of Spr\r{a}kbanken, the largest public ASR corpus for Danish to date. Our baseline results show that we achieve a 14.01 WER on the new corpus. A combination of FT Speech with in-domain language data provides comparable results to models trained specifically on Spr\r{a}kbanken, showing that FT Speech transfers well to this data set. Interestingly, our results demonstrate that the opposite is not the case. This shows that FT Speech provides a valuable resource for promoting research on Danish ASR with more spontaneous speech. Andreas Kirkedal, Marija Stepanovic, Barbara Plank |
INTERSPEECH | 3 |
| 2020 | Cross-Domain Evaluation of Edge Detection for Biomedical Event ExtractionabstractBiomedical event extraction is a crucial task in order to automatically extract information from the increasingly growing body of biomedical literature. Despite advances in the methods in recent years, most event extraction systems are still evaluated in-domain and on complete event structures only. This makes it hard to determine the performance of intermediate stages of the task, such as edge detection, across different corpora. Motivated by these limitations, we present the first cross-domain study of edge detection for biomedical event extraction. We analyze differences between five existing gold standard corpora, create a standardized benchmark corpus, and provide a strong baseline model for edge detection. Experiments show a large drop in performance when the baseline is applied on out-of-domain data, confirming the need for domain adaptation methods for the task. To encourage research efforts in this direction, we make both the data and the baseline available to the research community: https://www.cosbi.eu/cfx/9985. Alan Ramponi, Barbara Plank, Rosario Lombardo |
LREC | 2 |
| 2019 | Psycholinguistics Meets Continual Learning: Measuring Catastrophic Forgetting in Visual Question AnsweringabstractWe study the issue of catastrophic forgetting in the context of neural multimodal approaches to Visual Question Answering (VQA).Motivated by evidence from psycholinguistics, we devise a set of linguistically-informed VQA tasks, which differ by the types of questions involved (Wh-questions and polar questions).We test what impact task difficulty has on continual learning, and whether the order in which a child acquires question types facilitates computational models.Our results show that dramatic forgetting is at play and that task difficulty and order matter.Two well-known current continual learning methods mitigate the problem only to a limiting degree. Claudio Greco 0002, Barbara Plank, Raquel Fernández, Raffaella Bernardi |
ACL (1) | 2 |
| 2018 | Strong Baselines for Neural Semi-Supervised Learning under Domain ShiftabstractNovel neural models have been proposed in recent years for learning under domain shift.Most models, however, only evaluate on a single task, on proprietary datasets, or compare to weak baselines, which makes comparison of models difficult.In this paper, we re-evaluate classic general-purpose bootstrapping approaches in the context of neural networks under domain shifts vs. recent neural approaches and propose a novel multi-task tri-training method that reduces the time and space complexity of classic tri-training.Extensive experiments on two benchmarks are negative: while our novel method establishes a new state-of-the-art for sentiment analysis, it does not fare consistently the best.More importantly, we arrive at the somewhat surprising conclusion that classic tri-training, with some additions, outperforms the state of the art.We conclude that classic approaches constitute an important and strong baseline. Sebastian Ruder, Barbara Plank |
ACL (1) | 2 |
| 2018 | Distant Supervision from Disparate Sources for Low-Resource Part-of-Speech TaggingabstractWe introduce DSDS: a cross-lingual neural part-of-speech tagger that learns from disparate sources of distant supervision, and realistically scales to hundreds of low-resource languages.The model exploits annotation projection, instance selection, tag dictionaries, morphological lexicons, and distributed representations, all in a uniform framework.The approach is simple, yet surprisingly effective, resulting in a new state of the art without access to any gold annotated data. Barbara Plank, Zeljko Agic |
EMNLP | 1 |
| 2017 | When is multitask learning effective? Semantic sequence prediction under varying data conditionsabstractMultitask learning has been applied successfully to a range of tasks, mostly morphosyntactic.However, little is known on when MTL works and whether there are data characteristics that help to determine its success.In this paper we evaluate a range of semantic sequence labeling tasks in a MTL setup.We examine different auxiliary tasks, amongst which a novel setup, and correlate their impact to datadependent conditions.Our results show that MTL is not always effective, significant improvements are obtained only for 1 out of 5 tasks.When successful, auxiliary tasks with compact and more uniform label distributions are preferable. Héctor Martínez Alonso, Barbara Plank |
EACL (1) | 2 |
| 2017 | Parsing Universal Dependencies without trainingabstractHéctor Martínez Alonso, Željko Agić, Barbara Plank, Anders Søgaard. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Héctor Martínez Alonso, Zeljko Agic, Barbara Plank, Anders Søgaard |
EACL (1) | 3 |
| 2017 | Learning to select data for transfer learning with Bayesian OptimizationabstractDomain similarity measures can be used to gauge adaptability and select suitable data for transfer learning, but existing approaches define ad hoc measures that are deemed suitable for respective tasks.Inspired by work on curriculum learning, we propose to learn data selection measures using Bayesian Optimization and evaluate them across models, domains and tasks.Our learned measures outperform existing domain similarity measures significantly on three tasks: sentiment analysis, partof-speech tagging, and parsing.We show the importance of complementing similarity with diversity, and that learned measures are-to some degree-transferable across models, domains, and even tasks. Sebastian Ruder, Barbara Plank |
EMNLP | 2 |
| 2017 | Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures (Extended Abstract)abstractAutomatic image description generation is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the known approaches based on how they conceptualise this problem and provide a review of existing models, highlighting their advantages and disadvantages. Moreover, we give an overview of the benchmark image-text datasets and the evaluation measures that have been developed to assess the quality of machine-generated descriptions. Finally we explore future directions in the area of automatic image description. Raffaella Bernardi, Ruken Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, Barbara Plank |
IJCAI | 9 |
| 2017 | Sharing Is Caring: The Future of Shared TasksabstractShared tasks are indisputably drivers of progress and interest for problems in NLP. This is reflected by their increasing popularity, as well as by the fact that new shared tasks regularly emerge for under-researched and under-resourced topics, especially at workshops and smaller conferences.The general procedures and conventions for organizing a shared task have arisen organically over time (Paroubek, Chaudiron, and Hirschman, 2007, Section 7). There is no consistent framework that describes how shared tasks should be organized. This is not a harmful thing per se, but we believe that shared tasks, and by extension the field in general, would benefit from some reflection on the existing conventions. This, in turn, could lead to the future harmonization of shared task procedures.Shared tasks revolve around two aspects: research advancement and competition. We see research advancement as the driving force and main goal behind organizing them. Competition is an instrument to encourage and promote participation. However, just because these two forces are intrinsic to shared tasks does not mean that they always act in the same direction: Ensuring that the competition is fair is not a necessary requirement for advancing the field, and might even slow down progress.Our position in this respect is clear: We do believe that (i) advancing the field should be given priority over ensuring fair competition, also because (ii) inequality is partly unsolvable and intrinsic to life. In other words: Equality between competitors is desirable if it does not hinder research advancement.In the recently established workshop on ethics in NLP,1Parra Escartín et al. (2017) raise a set of considerations involving shared tasks, mainly focusing on areas where general ethical concerns regarding good scientific practice intersect with certain aspects of shared tasks. We find that they raise valid concerns, and in this contribution, we address some of them. However, we take a different perspective. Instead of focusing on ethical issues and potential negative effects of the competition aspect, we rather concentrate on how to bolster scientific progress.We make a simple proposal for the improvement of shared tasks, and discuss how it can help to mitigate the problems raised by Parra Escartín et al. (2017), while not necessarily tackling them directly. We start with assessing the concrete impact and significance of such concerns first.Recently, Parra Escartín et al. (2017) drew attention to a list of potential negative effects and ethical issues concerning shared tasks in NLP. In this section, we take this list as a starting point and examine each problem with respect to the main goal of shared tasks—to advance research in the field. Some issues were regarded as potential concerns rather than definite problems, because it is unclear how large their actual impact is. We believe that some of these concerns needed to be quantified in order to be properly assessed.To this end, we reviewed about 100 recent shared tasks from various campaigns (SemEval, EVALITA, CoNLL, WMT, and CLEF) between 2014 and 2016. We focused on several aspects, such as participation of companies, participation of organizers, closed versus open tracks, and the submission of papers by participants. Note that this is not an exhaustive overview of all shared tasks in NLP, but rather an arbitrary sample to investigate general trends in recent times. We use information drawn from this annotation exercise for assessing some of the problems we report in the following sections. The figures that are relevant for the discussion are reported in Table 1. The annotated spreadsheets used to collect this information are publicly available, together with some basic statistics and additional explanations.2Some potential concerns, although being ethically relevant, are not necessarily a problem in terms of research advancement, and fixing them directly should not be a priority. Here, we assess issues raised by Parra Escartín et al. (2017) that we believe fall into this category.Potential Conflicts of Interest. Parra Escartín et al. (2017) state that participation of organizers or annotators in their own shared task raises questions about inequality among participants, as organizers have earlier access to the data than the regular participants. In our survey, we found that in 5.8% of shared tasks, organizers did indeed participate. However, we also observe that this happens in connection with few participants (average 3.5 compared with 12.1, see Table 1), thus typically smaller shared tasks. This indicates that organizers' participation is more common in small, specialized tasks. The low number of participants can also explain why organizers perform better on average compared with non-organizers (see average normalized rank in Table 1).Unequal Playing Field. An unequal playing field mainly reflects the starting point that the different teams have. Parra Escartín et al. (2017) report on the issue of differences in processing power. An extreme example of this issue is the submission of Durrani et al. (2013) at WMT13, in which they reached the highest scores because they were able to boost the BLEU score by approximately 0.8% by making use of 1TB RAM, which was probably unavailable to the other teams at that time. Computational resources are not the only reason for an unequal playing field, though. There are many other causes that could lead to an unequal playing field—for example, some teams might have access to more proprietary data, proprietary software, or research equipment.The competitive nature of shared tasks can be fun, and stimulating for a variety of reasons (visibility, grant applications, beating state of the art, etc.). Such reasons might not necessarily be positively correlated with advancing the field, though. Here we discuss issues also raised by Parra Escartín et al. (2017) that we believe fall into this category.Secretiveness. As a result of the competitive nature of shared tasks, it can be desirable for participating teams to keep their “secret sauce” private, as this could mean an advantage for a re-run of the same task, or a shared task on a related problem. As a possible effect of secretiveness, Parra Escartín et al. (2017) also mention “Unconscious overlooking of ethical concerns,” actually referring to an unacceptable level of vagueness in papers. In other words, participants may unintentionally describe their systems in an abstract and vague way due to a previously established practice in systems' descriptions.Lack of Description of Negative Results. Given that negative results are informative, their under-representation in shared tasks is a concern. A lack of knowledge about negative results might lead to a research redundancy, which is clearly undesirable. The issue of under-represented negative results is a global concern for the entire field, and for science in general. However, shared tasks provide an excellent opportunity for publishing negative results, as the acceptance of papers for publication does not particularly favor positive results.Redundancy and Replicability in the Field. Parra Escartín et al. (2017) raise issues concerning two types of redundancy, (a) when optimal parameter settings of a previous shared task do not carry over to the new version of the task, therefore it is not clear what is learned; and (b) when algorithms are reimplemented for replicability purposes.Regarding (a), we think that differences in used parameter settings are not actually a bad thing; we learn from this that we overfit on the previous task, or that we need to adapt our systems to another data set or domain.Regarding (b), this is a real problem because starting from scratch to reimplement existing systems is unnecessarily time-consuming. In addition, it would always be desirable to be able to directly reproduce the same results of the same model for the same task (Pedersen, 2008; Fokkens et al., 2013).Withdrawal from Competition. Participants may withdraw from a shared task if their ranking in the competition can negatively affect their reputation and/or future funding. For example, Parra Escartín et al. (2017) suggest that companies might prefer to withdraw from the competition if they are not highly ranked, to avoid blemishing their reputation. This is something that we could not quantify in our survey, as in case of withdrawal there would be no evidence of participation in reports. There are two aspects, though, that we can quantify. The first aspect is the number of teams that do not publish their system's description, which amounts to approximately 9%, and could indeed be related to withdrawals. However, exactly because the paper is missing, information on why a team withdrew is not available. The second aspect is the total number of industry participants, which in our sample amounts to 20% (“Some company” and “Only company” in Table 1). Thus, although there is not much that can be done about withdrawal—and this might not be a problem anyway—we believe that, considering the substantial presence and interest of industries so far, their participation should be accommodated.Potentially Gaming the System. Shared tasks are usually bound to data sets and evaluation metrics. This could lead to competition-oriented participants focusing more on tuning their systems on a given data set and metrics rather than finding a scientifically sound and scalable method for solving the problem. This can, in turn, result in an “unfair” ranking or a misleading relation between a methodology and its value with respect to the research problem. A potential negative outcome of the latter is a scenario where “optimal” methods of a shared task do not carry over to related shared tasks. These problems become more severe when system gaming is combined with a secretive attitude. While tackling this issue, we should take into account that participants might be less eager to write about ad hoc solutions, for example tuning pre-processing components or tailoring a system too closely to specifics of the annotation.Our proposal for future shared tasks is not revolutionary. It simply revolves around the key aspects of sharing, not only resources but also experiences, including negative ones. Specifically, with research progress in mind, we believe sharing should be encouraged and even partially enforced. We therefore suggest an explicit setting for shared tasks in NLP, and reflect on the issue of what organizers could do in order to maximize sharing of information regarding participating systems. We also show how such a simple strategy can help to overcome the problems raised that can hinder research advancement.One of the challenges faced by shared tasks is to ensure a level playing field permitting a transparent comparison of the merits of different methods. The problem is that system A might come out on top of system B not because its method is superior, but because, for example, it was trained on more data. This would favor teams with access to more resources, like companies with large quantities of proprietary in-house data.Traditionally, this problem has been mitigated by establishing “closed tracks.” In closed tracks, participating teams are not allowed to use any training data other than that provided by the shared task organizers. The rationale behind this is that if all systems use exactly the same data, the playing field is equal, and the competition results will show the strengths of the different methods. In order to study the effect of additional training data, many shared tasks have a separate competition, the so-called “open track.”However, this open–closed division is increasingly impractical and ineffective. The main problem is that it is only concerned with training data, whereas the performance of systems can crucially depend on other resources. Examples are external components with pre-trained models, such as part-of-speech taggers and dependency parsers, auxiliary data-derived resources like word embeddings, and other influential factors like the availability of computational resources. Because such models are almost always derived from external language data, it is unclear where to draw the line between closed and open. Should such data be disallowed or not? If not, teams still do not really participate on an equal footing.From a research perspective, banning external models is completely impractical and nonsensical, as most state-of-the-art systems now depend on them. Likewise, trying to force all teams to use the same set of external models, and no other, would place a heavy burden on both organizers and participants. Moreover, restricting the resources participants can use is questionable, because, for research to progress quickly, teams should use the best resources available, or the resources best fitting their system.It is therefore unsurprising that the use of closed tracks has declined in shared tasks in general in the last few years, as we have observed during our review of shared tasks. However, the original problem of unequal playing field, and thus a bias in favor of teams with ample resources, remains.We propose, then, to rethink the problem, not in terms of equal training data, but in terms of equal opportunities. This is closely connected to the wider issue of reproducibility and replicability: Like all published research, shared task results should ideally be fully reproducible by anyone (Pedersen, 2008; Fokkens et al., 2013).3 Moreover, it should be easy to build on others' work to try out new variations of a method, without having to reimplement things from scratch. To ensure this, it is desirable that everything needed to reproduce experimental results is publicly and freely available, including code, data, pre-trained models, and so on. Interestingly, at the CoNLL-2013 shared task a similar step was taken, but only in terms of pre-condition: “While all teams in the shared task use the NUCLE corpus, they are also allowed to use additional external resources (both corpora and tools) so long as they are publicly available and not proprietary” (Ng et al., 2013). We would like to take this a step further, by enforcing the sharing of whatever resource teams might choose to use, so as to favor the injection of new resources in the field.Applying this principle to shared tasks in practice, we propose making the primary competition a “public track,” where participants can use any code, data, and pre-trained models they want, as long as others can then freely obtain them. In other words: All resources used to participate in the shared task should be subsequently shared with the community. Although this does not ensure equal access to resources for the current edition, it will still ensure a progressively more equal footing for the future. We believe this is the crucial step to move the field forward, as everyone will have access to the resources used in state-of-the-art systems. To keep participation possible for teams who cannot or will not make all resources available, a secondary, “proprietary track” can be established.Ranking of systems forms a large part of the appeal of shared tasks. However, rankings should not be overemphasized and are far from being the final goal of shared tasks. Research is supposed to teach us about the merits and characteristics of methods, including insights of what does not work, rather than about which team built the system that performed best on the test data.Negative results are very informative for future developments. Although publishing negative results is difficult, shared tasks do provide the ideal context for disclosing and explaining low performance methods and choices. We believe that shared task organizers should explicitly and strongly solicit the inclusion of what did not work in the reports written by participating teams. This could be even solicited via an online form that participants submit after the evaluation phase, where they comment on what worked well (as commonly done), but also provides a separate section to explain what did not work. This information could in turn be valuable data for organizers when compiling the overview report. Moreover, a clear explanation of what did not work, in connection with availability of code, would help to better understand whether something does not work as an idea or because of a specific implementation.More generally, organizers should encourage—and to some point ensure through the reviewing process—that all participants provide exhaustive reports, potentially including ablation/addition tests, so as to have a picture as comprehensive as possible. Because participating in shared tasks directly implies getting a paper accepted for publication, not everyone describes their system to the satisfaction of external reviewers. This should change, and acceptance should be conditional on clarity and exhaustiveness.We stated that the goal of advancing research should be prioritized over competition. This is especially the case when focusing on the competition aspect would encourage undesired practices like secretiveness and gaming the system. We suggest a simple solution based on the principle of maximizing resource- and information-sharing. As a byproduct, some problematic competition-related issues will be overcome, too. Some outstanding ethical issues cannot be solved, as they are intrinsic in human nature and cannot be controlled for by means of specific guidelines.Introducing proprietary and public tracks will stimulate participants to release their systems and resources. This will directly reduce Secretiveness and the issue of Redundancy and Replicability in the Field. It will also partially address the Unequal Playing Field problem, at least in the long run: Even if at the same competition different teams will have access to different resources, all resources will be available to everyone for the next round. Moreover, being able to access and run systems on different data sets will uncover limitations that might have been due to tailoring systems to the specifics of a given shared task (Potential Gaming the System). The presence of a proprietary track still allows for industrial participation (see Withdrawal from Competition in Section 2.2), where distribution of resources might not be as easy as for other teams.Encouraging participants to write comprehensive reports that include negative results will be a valid instrument towards advancing research, at the same time solving some outstanding problems. The reviewers should probably spend extra time in assessing the single reports and accept them conditionally on clarity requirements, but we believe this is worth the effort. Indeed, enforcing that systems are described properly will ensure and the of there are some issues We believe these are issues that cannot or need not be Withdrawal of teams cannot be controlled if of negative results is encouraged and common practice, it is possible that teams will choose to the competition. on progress and sharing rather than will also The of interest is not relevant in our We do believe that organizers should be allowed to and this is especially for shared tasks that might a number of participating teams due to the nature of the As long as these are explicitly reported in both the overview paper and the system the of results and ranking is to the closed tracks also an equal playing will it make it to different methods over the same We do not think Equality will be increasingly by resource sharing, to the that it is as inequality is part of the As we in Section comparison of methods has not been transparent in closed tracks the use of resources is not clarity in reports and release of systems will make it possible for the to assess which methods work and which do not, and to progressively on the state of the that our and will discussion on shared tasks, to make them more for driving progress in and each task will to have their own settings that the organizers will most However, we do believe that participants to release their in terms of resources, and of and should be a common to from long we together and on various with many It is to everyone for their However, we to mention a few who have to into better The discussion we at of the of the Computational at the of was the actual for this are to everyone and in to and for their and valuable has and to the on what does not work with the open versus closed track setting as it We and Parra Escartín for on earlier of this We are also to for Malvina Nissim, Lasha Abzianidze, Kilian Evang, Rob van der Goot, Hessel Haagsma, Barbara Plank, Martijn Wieling 0001 |
Comput. Linguistics | 6 |
| 2016 | Semantic Tagging with Deep Residual NetworksabstractWe propose a novel semantic tagging task, semtagging, tailored for the purpose of multilingual semantic parsing, and present the first tagger using deep residual networks (ResNets). Our tagger uses both word and character representations, and includes a novel residual bypass architecture. We evaluate the tagset both intrinsically on the new task of semantic tagging, as well as on Part-of-Speech (POS) tagging. Our system, consisting of a ResNet and an auxiliary loss function predicting our semantic tags, significantly outperforms prior results on English Universal Dependencies POS tagging (95.71% accuracy on UD v1.2 and 95.67% accuracy on UD v1.3). Johannes Bjerva, Barbara Plank, Johan Bos |
COLING | 2 |
| 2016 | Multi-view and multi-task training of RST discourse parsersabstractWe experiment with different ways of training LSTM networks to predict RST discourse trees. The main challenge for RST discourse parsing is the limited amounts of training data. We combat this by regularizing our models using task supervision from related tasks as well as alternative views on discourse structures. We show that a simple LSTM sequential discourse parser takes advantage of this multi-view and multi-task framework with 12-15% error reductions over our baseline (depending on the metric) and results that rival more complex state-of-the-art parsers. Chloé Braud, Barbara Plank, Anders Søgaard |
COLING | 2 |
| 2016 | Keystroke dynamics as signal for shallow syntactic parsingabstractKeystroke dynamics have been extensively used in psycholinguistic and writing research to gain insights into cognitive processing. But do keystroke logs contain actual signal that can be used to learn better natural language processing models? We postulate that keystroke dynamics contain information about syntactic structure that can inform shallow syntactic parsing. To test this hypothesis, we explore labels derived from keystroke logs as auxiliary task in a multi-task bidirectional Long Short-Term Memory (bi-LSTM). Our results show promising results on two shallow syntactic parsing tasks, chunking and CCG supertagging. Our model is simple, has the advantage that data can come from distinct sources, and produces models that are significantly better than models trained on the text annotations alone. Barbara Plank |
COLING | 1 |
| 2016 | TwiSty: A Multilingual Twitter Stylometry Corpus for Gender and Personality Profiling
Ben Verhoeven, Walter Daelemans, Barbara Plank |
LREC | 3 |
| 2016 | Multi-lingual opinion mining on YouTube
Aliaksei Severyn, Alessandro Moschitti, Olga Uryupina, Barbara Plank, Katja Filippova |
Inf. Process. Manag. | 4 |
| 2016 | Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation MeasuresabstractAutomatic description generation from natural images is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the existing approaches based on how they conceptualize this problem, viz., models that cast description as either generation problem or as a retrieval problem over a visual or multimodal representational space. We provide a detailed review of existing models, highlighting their advantages and disadvantages. Moreover, we give an overview of the benchmark image datasets and the evaluation measures that have been developed to assess the quality of machine-generated image descriptions. Finally we extrapolate future directions in the area of automatic image description generation. Raffaella Bernardi, Ruken Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, Barbara Plank |
J. Artif. Intell. Res. | 9 |
| 2016 | Multilingual Projection for Parsing Truly Low-Resource LanguagesabstractWe propose a novel approach to cross-lingual part-of-speech tagging and dependency parsing for truly low-resource languages. Our annotation projection-based approach yields tagging and parsing models for over 100 languages. All that is needed are freely available parallel texts, and taggers and parsers for resource-rich languages. The empirical evaluation across 30 test languages shows that our method consistently provides top-level accuracies, close to established upper bounds, and outperforms several competitive baselines. Zeljko Agic, Anders Johannsen, Barbara Plank, Héctor Martínez Alonso, Natalie Schluter, Anders Søgaard |
Trans. Assoc. Comput. Linguistics | 3 |
| 2015 | Using Frame Semantics for Knowledge Extraction from TwitterabstractKnowledge bases have the potential to advance artificial intelligence, but often suffer from recall problems, i.e., lack of knowledge of new entities and relations. On the contrary, social media such as Twitter provide abundance of data, in a timely manner: information spreads at an incredible pace and is posted long before it makes it into more commonly used resources for knowledge extraction. In this paper we address the question whether we can exploit social media to extract new facts, which may at first seem like finding needles in haystacks. We collect tweets about 60 entities in Freebase and compare four methods to extract binary relation candidates, based on syntactic and semantic parsing and simple mechanism for factuality scoring. The extracted facts are manually evaluated in terms of their correctness and relevance for search. We show that moving from bottom-up syntactic or semantic dependency parsing formalisms to top-down frame-semantic processing improves the robustness of knowledge extraction, producing more intelligible fact candidates of better quality. In order to evaluate the quality of frame semantic parsing on Twitter intrinsically, we make a multiply frame-annotated dataset of tweets publicly available. Anders Søgaard, Barbara Plank, Héctor Martínez Alonso |
AAAI | 2 |
| 2015 | Semantic Representations for Domain Adaptation: A Case Study on the Tree Kernel-based Method for Relation ExtractionabstractThien Huu Nguyen, Barbara Plank, Ralph Grishman. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Thien Huu Nguyen, Barbara Plank, Ralph Grishman |
ACL (1) | 2 |
| 2015 | Inverted indexing for cross-lingual NLPabstractAnders Søgaard, Željko Agić, Héctor Martínez Alonso, Barbara Plank, Bernd Bohnet, Anders Johannsen. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Anders Søgaard, Zeljko Agic, Héctor Martínez Alonso, Barbara Plank, Bernd Bohnet, Anders Johannsen |
ACL (1) | 4 |
| 2015 | Do dependency parsing metrics correlate with human judgments?abstractUsing automatic measures such as labeled and unlabeled attachment scores is common practice in dependency parser evaluation.In this paper, we examine whether these measures correlate with human judgments of overall parse quality.We ask linguists with experience in dependency annotation to judge system outputs.We measure the correlation between their judgments and a range of parse evaluation metrics across five languages.The humanmetric correlation is lower for dependency parsing than for other NLP tasks.Also, inter-annotator agreement is sometimes higher than the agreement between judgments and metrics, indicating that the standard metrics fail to capture certain aspects of parse quality, such as the relevance of root attachment or the relative importance of the different parts of speech. Barbara Plank, Héctor Martínez Alonso, Zeljko Agic, Danijela Merkler, Anders Søgaard |
CoNLL | 1 |
| 2015 | Using Knowledge Components for Collaborative Filtering in Adaptive Tutoring Systems
Peter Halkier Nicolajsen, Barbara Plank |
EDM | 2 |
| 2015 | Learning to parse with IAA-weighted lossabstractHéctor Martínez Alonso, Barbara Plank, Arne Skjærholt, Anders Søgaard. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Héctor Martínez Alonso, Barbara Plank, Arne Skjærholt, Anders Søgaard |
HLT-NAACL | 2 |
| 2015 | Mining for unambiguous instances to adapt part-of-speech taggers to new domainsabstractDirk Hovy, Barbara Plank, Héctor Martínez Alonso, Anders Søgaard. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Dirk Hovy, Barbara Plank, Héctor Martínez Alonso, Anders Søgaard |
HLT-NAACL | 2 |
| 2014 | Opinion Mining on YouTubeabstractAliaksei Severyn, Alessandro Moschitti, Olga Uryupina, Barbara Plank, Katja Filippova. Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2014. Aliaksei Severyn, Alessandro Moschitti, Olga Uryupina, Barbara Plank, Katja Filippova |
ACL (1) | 4 |
| 2014 | Adapting taggers to Twitter with not-so-distant supervision
Barbara Plank, Dirk Hovy, Ryan T. McDonald, Anders Søgaard |
COLING | 1 |
| 2014 | What's in a p-value in NLP?abstractIn NLP, we need to document that our pro-posed methods perform significantly bet-ter with respect to standard metrics than previous approaches, typically by re-porting p-values obtained by rank- or randomization-based tests. We show that significance results following current re-search standards are unreliable and, in ad-dition, very sensitive to sample size, co-variates such as sentence length, as well as to the existence of multiple metrics. We estimate that under the assumption of per-fect metrics and unbiased data, we need a significance cut-off at ⇠0.0025 to reduce the risk of false positive results to <5%. Since in practice we often have consider-able selection bias and poor metrics, this, however, will not do alone. 1 Anders Søgaard, Anders Johannsen, Barbara Plank, Dirk Hovy, Héctor Martínez Alonso |
CoNLL | 3 |
| 2014 | Learning part-of-speech taggers with inter-annotator agreement lossabstractIn natural language processing (NLP) an-notation projects, we use inter-annotator agreement measures and annotation guide-lines to ensure consistent annotations. However, annotation guidelines often make linguistically debatable and even somewhat arbitrary decisions, and inter-annotator agreement is often less than perfect. While annotation projects usu-ally specify how to deal with linguisti-cally debatable phenomena, annotator dis-agreements typically still stem from these “hard ” cases. This indicates that some er-rors are more debatable than others. In this paper, we use small samples of doubly-annotated part-of-speech (POS) data for Twitter to estimate annotation reliability and show how those metrics of likely inter-annotator agreement can be implemented in the loss functions of POS taggers. We find that these cost-sensitive algorithms perform better across annotation projects and, more surprisingly, even on data an-notated according to the same guidelines. Finally, we show that POS tagging mod-els sensitive to inter-annotator agreement perform better on the downstream task of chunking. 1 Barbara Plank, Dirk Hovy, Anders Søgaard |
EACL | 1 |
| 2014 | Importance weighting and unsupervised domain adaptation of POS taggers: a negative resultabstractImportance weighting is a generalization of various statistical bias correction techniques.While our labeled data in NLP is heavily biased, importance weighting has seen only few applications in NLP, most of them relying on a small amount of labeled target data.The publication bias toward reporting positive results makes it hard to say whether researchers have tried.This paper presents a negative result on unsupervised domain adaptation for POS tagging.In this setup, we only have unlabeled data and thus only indirect access to the bias in emission and transition probabilities.Moreover, most errors in POS tagging are due to unseen words, and there, importance weighting cannot help.We present experiments with a wide variety of weight functions, quantilizations, as well as with randomly generated weights, to support these claims. Barbara Plank, Anders Johannsen, Anders Søgaard |
EMNLP | 1 |
| 2014 | When POS data sets don't add up: Combatting sample bias
Dirk Hovy, Barbara Plank, Anders Søgaard |
LREC | 2 |
| 2014 | SenTube: A Corpus for Sentiment Analysis on YouTube Social Media
Olga Uryupina, Barbara Plank, Aliaksei Severyn, Agata Rotondi, Alessandro Moschitti |
LREC | 2 |
| 2013 | Embedding Semantic Similarity in Tree Kernels for Domain Adaptation of Relation Extraction
Barbara Plank, Alessandro Moschitti |
ACL (1) | 1 |
| 2011 | Effective Measures of Domain Similarity for Parsing
Barbara Plank, Gertjan van Noord |
ACL | 1 |
| 2010 | Improved Statistical Measures to Assess Natural Language Parser Performance across Domains
Barbara Plank |
LREC | 1 |
| 2008 | Parsing with subdomain instance weighting from raw corpora
Barbara Plank, Khalil Sima'an |
INTERSPEECH | 1 |
| 2008 | Subdomain Sensitive Statistical Parsing using Raw Corpora
Barbara Plank, Khalil Sima'an |
LREC | 1 |
| 2006 | Multilingual Search in Libraries. The case-study of the Free University of Bozen-Bolzano
Raffaella Bernardi, Diego Calvanese, Luca Dini, Vittorio Di Tomaso, Elisabeth Frasnelli, Ulrike Kugler, Barbara Plank |
LREC | 7 |