Rob van der Goot

dblp:184/8526 · DBLP profile ↗
← Back
33ranked-venue papers
9as first author
24since 2021 · last 2026
0009-0003-1999-4156ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 9 first-author · 24 since 2021
YearPublicationVenuePosition
2026 CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
abstract
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa, Nadia Ghezaiel Hammouda, Verrah Akinyi Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Chalamalasetti Kranti, Carol Muchemi, Khang Nguyen, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Jayne Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob Van Der Goot, Lanwenn ar C’horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Juan Pablo Martínez, Amit Agarwal, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin L Rice, Azril Hafizi Amirudin, Jesujoba Oluwadara Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, Akshata, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Rufaro Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Celikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah K. K. Luger. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa 0001, Nadia Ghezaiel Hammouda, Verrah Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Kranti Chalamalasetti, Carol Muchemi, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Jayne Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada A. Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob van der Goot, Lanwenn Ar C'horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin Rice, Azril Hafizi Amirudin, Jesujoba O. Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, Akshata A, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Çelikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah K. K. Luger
ACL (1)39
2026 Dynaword: From One-shot to Continuously Developed Datasets
abstract
Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguously licensed sources restricting use, sharing, and derivative works; (2) static dataset releases that prevent community contributions and diminish longevity; and (3) quality assurance processes restricted to publishing teams rather than leveraging community expertise. To address these limitations, we introduce two contributions: the Dynaword approach and Danish Dynaword. The Dynaword approach is a framework for creating large-scale, open datasets that can be continuously updated through community collaboration. Danish Dynaword is a concrete implementation that validates this approach and demonstrates its potential. Danish Dynaword contains over four times as many tokens as comparable releases, is exclusively openly licensed, and has received multiple contributions across industry and research. The repository includes light-weight tests to ensure data formatting, quality, and documentation, establishing a sustainable framework for ongoing community contributions and dataset evolution.
Kenneth C. Enevoldsen, Kristian Nørgaard Jensen, Jan Kostkan, Balázs Szabó, Márton Kardos, Kirsten Vad, Johan Heinsen, Andrea Blasi Núñez, Gianluca Barmina, Jacob Nielsen, Rasmus Larsen 0001, Rob van der Goot, Peter Bjerregaard Vahlstrup, Per Møldrup-Dalum, Desmond Elliott, Lukas Galke Poech, Peter Schneider-Kamp, Kristoffer L. Nielbo
LREC12
2026 TækTåk: Syntactic Analysis of Language Use on Danish TikTok
Thea Kristensen, Rob van der Goot
LREC2
2026 MultiLexNorm++: A Unified Benchmark and a Generative Model for Lexical Normalization for Asian Languages
abstract
Social media data has been of interest to Natural Language Processing (NLP) practitioners for over a decade, because of its richness in information, but also challenges for automatic processing. Since language use is more informal, spontaneous, and adheres to many different sociolects, the performance of NLP models often deteriorates. One solution to this problem is to transform data to a standard variant before processing it, which is also called lexical normalization. There has been a wide variety of benchmarks and models proposed for this task. The MultiLexNorm benchmark proposed to unify these efforts, but it consists almost solely of languages from the Indo-European language family in the Latin script. Hence, we propose an extension to MultiLexNorm, which covers five Asian languages from different language families in four different scripts. We show that the previous state-of-the-art model performs worse on the new languages and propose a new architecture based on Large Language Models (LLMs), which shows more robust performance. Finally, we analyze remaining errors, revealing future directions for this task.
Weerayut Buaphet, Thanh-Nhi Nguyen, Risa Kondo, Tomoyuki Kajiwara, Yumin Kim, Jimin Lee 0001, Hwanhee Lee, Holy Lovenia, Peerat Limkonchotiwat, Sarana Nutanong, Rob van der Goot
ACM Trans. Asian Low Resour. Lang. Inf. Process.11
2025 Identifying Open Challenges in Language Identification
abstract
Automatic language identification is a core problem of many Natural Language Processing (NLP) pipelines.A wide variety of architectures and benchmarks have been proposed with often near-perfect performance.Although previous studies have focused on certain challenging setups (i.e.cross-domain, short inputs), a systematic comparison is missing.We propose a benchmark that allows us to test for the effect of input size, training data size, domain, number of languages, scripts, and language families on performance.We evaluate five popular models on this benchmark and identify which open challenges remain for this task as well as which architectures achieve robust performance.We find that cross-domain setups are the most challenging (although arguably most relevant), and that number of languages, variety in scripts, and variety in language families have only a small impact on performance.We also contribute practical takeaways: training with 1,000 instances per language and a maximum input length of 100 characters is enough for robust language identification.Based on our findings, we train an accurate (94.41%) multidomain language identification model on 2,034 languages, for which we also provide an analysis of the remaining errors.
Rob van der Goot
ACL (1)1
2025 data2lang2vec: Data Driven Typological Features Completion
abstract
Language typology databases enhance multi-lingual Natural Language Processing (NLP) by improving model adaptability to diverse linguistic structures. The widely-used lang2vec toolkit integrates several such databases, but its coverage remains limited at 28.9%. Previous work on automatically increasing coverage predicts missing values based on features from other languages or focuses on single features, we propose to use textual data for better-informed feature prediction. To this end, we introduce a multi-lingual Part-of-Speech (POS) tagger, achieving over 70% accuracy across 1,749 languages, and experiment with external statistical features and a variety of machine learning algorithms. We also introduce a more realistic evaluation setup, focusing on likely to be missing typology features, and show that our approach outperforms previous work in both setups.
Hamidreza Amirzadeh, Sadegh Jafari, Anika Harju, Rob van der Goot
COLING4
2025 Iterative Structured Knowledge Distillation: Optimizing Language Models Through Layer-by-Layer Distillation
abstract
Traditional language model compression techniques, like knowledge distillation, require a fixed architecture, limiting flexibility, while structured pruning methods often fail to preserve performance. This paper introduces Iterative Structured Knowledge Distillation (ISKD), which integrates knowledge distillation and structured pruning by progressively replacing transformer blocks with smaller, efficient versions during training. This study validates ISKD on two transformer-based language models: GPT-2 and Phi-1. ISKD outperforms L1 pruning and achieves similar performance to knowledge distillation while offering greater flexibility. ISKD reduces model parameters - 30.68% for GPT-2 and 30.16% for Phi-1 - while maintaining at least four-fifths of performance on both language modeling and commonsense reasoning tasks. These findings suggest that this method offers a promising balance between model efficiency and accuracy.
Malthe Have Musaeus, Rob van der Goot
COLING2
2024 Can Humans Identify Domains?
abstract
Textual domain is a crucial property within the Natural Language Processing (NLP) community due to its effects on downstream model performance. The concept itself is, however, loosely defined and, in practice, refers to any non-typological property, such as genre, topic, medium or style of a document. We investigate the core notion of domains via human proficiency in identifying related intrinsic textual properties, specifically the concepts of genre (communicative purpose) and topic (subject matter). We publish our annotations in TGeGUM: A collection of 9.1k sentences from the GUM dataset (Zeldes, 2017) with single sentence and larger context (i.e., prose) annotations for one of 11 genres (source type), and its topic/subtopic as per the Dewey Decimal library classification system (Dewey, 1979), consisting of 10/100 hierarchical topics of increased granularity. Each instance is annotated by three annotators, for a total of 32.7k annotations, allowing us to examine the level of human disagreement and the relative difficulty of each annotation task. With a Fleiss’ kappa of at most 0.53 on the sentence level and 0.66 at the prose level, it is evident that despite the ubiquity of domains in NLP, there is little human consensus on how to define them. By training classifiers to perform the same task, we find that this uncertainty also extends to NLP models.
Maria Barrett, Max Müller-Eberstein, Elisa Bassignana, Amalie Brogaard Pauli, Mike Zhang, Rob van der Goot
LREC/COLING6
2024 How to Encode Domain Information in Relation Classification
abstract
Current language models require a lot of training data to obtain high performance. For Relation Classification (RC), many datasets are domain-specific, so combining datasets to obtain better performance is non-trivial. We explore a multi-domain training setup for RC, and attempt to improve performance by encoding domain information. Our proposed models improve > 2 Macro-F1 against the baseline setup, and our analysis reveals that not all the labels benefit the same: The classes which occupy a similar space across domains (i.e., their interpretation is close across them, for example “physical”) benefit the least, while domain-dependent relations (e.g., “part-of”) improve the most when encoding domain information.
Elisa Bassignana, Viggo Unmack Gascou, Frida Nøhr Laustsen, Gustav Kristensen, Marie Haahr Petersen, Rob van der Goot, Barbara Plank
LREC/COLING6
2024 Enough Is Enough! a Case Study on the Effect of Data Size for Evaluation Using Universal Dependencies
abstract
When creating a new dataset for evaluation, one of the first considerations is the size of the dataset. If our evaluation data is too small, we risk making unsupported claims based on the results on such data. If, on the other hand, the data is too large, we waste valuable annotation time and costs that could have been used to widen the scope of our evaluation (i.e. annotate for more domains/languages). Hence, we investigate the effect of the size and a variety of sampling strategies of evaluation data to optimize annotation efforts, using dependency parsing as a test case. We show that for in-language in-domain datasets, 5,000 tokens is enough to obtain a reliable ranking of different parsers; especially if the data is distant enough from the training split (otherwise, we recommend 10,000). In cross-domain setups, the same amounts are required, but in cross-lingual setups much less (2,000 tokens) is enough.
Rob van der Goot, Zoey Liu, Max Müller-Eberstein
LREC/COLING1
2024 Slot and Intent Detection Resources for Bavarian and Lithuanian: Assessing Translations vs Natural Queries to Digital Assistants
abstract
Digital assistants perform well in high-resource languages like English, where tasks like slot and intent detection (SID) are well-supported. Many recent SID datasets start including multiple language varieties. However, it is unclear how realistic these translated datasets are. Therefore, we extend one such dataset, namely xSID-0.4, to include two underrepresented languages: Bavarian, a German dialect, and Lithuanian, a Baltic language. Both language variants have limited speaker populations and are often not included in multilingual projects. In addition to translations we provide “natural” queries to digital assistants generated by native speakers. We further include utterances from another dataset for Bavarian to build the richest SID dataset available today for a low-resource dialect without standard orthography. We then set out to evaluate models trained on English in a zero-shot scenario on our target language variants. Our evaluation reveals that translated data can produce overly optimistic scores. However, the error patterns in translated and natural datasets are highly similar. Cross-dataset experiments demonstrate that data collection methods influence performance, with scores lower than those achieved with single-dataset translations. This work contributes to enhancing SID datasets for underrepresented languages, yielding NaLiBaSID, a new evaluation dataset for Bavarian and Lithuanian.
Miriam Winkler, Virginija Juozapaityte, Rob van der Goot, Barbara Plank
LREC/COLING3
2024 NNOSE: Nearest Neighbor Occupational Skill Extraction
abstract
The labor market is changing rapidly, prompting increased interest in the automatic extraction of occupational skills from text. With the advent of English benchmark job description datasets, there is a need for systems that handle their diversity well. We tackle the complexity in occupational skill datasets tasks—combining and leveraging multiple datasets for skill extraction, to identify rarely observed skills within a dataset, and overcoming the scarcity of skills across datasets. In particular, we investigate the retrieval-augmentation of language models, employing an external datastore for retrieving similar skills in a dataset-unifying manner. Our proposed method, Nearest Neighbor Occupational Skill Extraction (NNOSE) effectively leverages multiple datasets by retrieving neighboring skills from other datasets in the datastore. This improves skill extraction without additional fine-tuning. Crucially, we observe a performance gain in predicting infrequent patterns, with substantial gains of up to 30% span-F1 in cross-dataset settings.
Mike Zhang, Rob van der Goot, Min-Yen Kan, Barbara Plank
EACL (1)2
2023 ESCOXLM-R: Multilingual Taxonomy-driven Pre-training for the Job Market Domain
abstract
The increasing number of benchmarks for Natural Language Processing (NLP) tasks in the computational job market domain highlights the demand for methods that can handle job-related tasks such as skill extraction, skill classification, job title classification, and de-identification.While some approaches have been developed that are specific to the job market domain, there is a lack of generalized, multilingual models and benchmarks for these tasks.In this study, we introduce a language model called ESCOXLM-R, based on XLM-R large , which uses domain-adaptive pre-training on the European Skills, Competences, Qualifications and Occupations (ESCO) taxonomy, covering 27 languages.The pre-training objectives for ESCOXLM-R include dynamic masked language modeling and a novel additional objective for inducing multilingual taxonomical ESCO relations.We comprehensively evaluate the performance of ESCOXLM-R on 6 sequence labeling and 3 classification tasks in 4 languages and find that it achieves state-of-the-art results on 6 out of 9 datasets.Our analysis reveals that ESCOXLM-R performs better on short spans and outperforms XLM-R large on entity-level and surface-level span-F1, likely due to ESCO containing short skill and occupation titles, and encoding information on the entity-level.
Mike Zhang, Rob van der Goot, Barbara Plank
ACL (1)2
2023 Establishing Trustworthiness: Rethinking Tasks and Model Evaluation
abstract
Language understanding is a multi-faceted cognitive capability, which the Natural Language Processing (NLP) community has striven to model computationally for decades.Traditionally, facets of linguistic intelligence have been compartmentalized into tasks with specialized model architectures and corresponding evaluation protocols.With the advent of large language models (LLMs) the community has witnessed a dramatic shift towards general purpose, task-agnostic approaches powered by generative models.As a consequence, the traditional compartmentalized notion of language tasks is breaking down, followed by an increasing challenge for evaluation and analysis.At the same time, LLMs are being deployed in more real-world scenarios, including previously unforeseen zero-shot setups, increasing the need for trustworthy and reliable systems.Therefore, we argue that it is time to rethink what constitutes tasks and model evaluation in NLP, and pursue a more holistic view on language, placing trustworthiness at the center.Towards this goal, we review existing compartmentalized approaches for understanding the origins of a model's functional capacity, and provide recommendations for more multifaceted evaluation protocols."Trust arises from knowledge of origin as well as from knowledge of functional capacity."
Robert Litschko, Max Müller-Eberstein, Rob van der Goot, Leon Weber-Genzel, Barbara Plank
EMNLP3
2022 Probing for Labeled Dependency Trees
abstract
Probing has become an important tool for analyzing representations in Natural Language Processing (NLP). For graphical NLP tasks such as dependency parsing, linear probes are currently limited to extracting undirected or unlabeled parse trees which do not capture the full task. This work introduces DepProbe, a linear probe which can extract labeled and directed dependency parse trees from embeddings while using fewer parameters and compute than prior methods. Leveraging its full task coverage and lightweight parametrization, we investigate its predictive power for selecting the best transfer language for training a full biaffine attention parser. Across 13 languages, our proposed method identifies the best source treebank 94% of the time, outperforming competitive baselines and prior work. Finally, we analyze the informativeness of task-specific subspaces in contextual embeddings as well as which benefits a full parser’s non-linear parametrization provides.
Max Müller-Eberstein, Rob van der Goot, Barbara Plank
ACL (1)2
2022 Tafsir Dataset: A Novel Multi-Task Benchmark for Named Entity Recognition and Topic Modeling in Classical Arabic Literature
abstract
Various historical languages, which used to be lingua franca of science and arts, deserve the attention of current NLP research. In this work, we take the first data-driven steps towards this research line for Classical Arabic (CA) by addressing named entity recognition (NER) and topic modeling (TM) on the example of CA literature. We manually annotate the encyclopedic work of Tafsir Al-Tabari with span-based NEs, sentence-based topics, and span-based subtopics, thus creating the Tafsir Dataset with over 51,000 sentences, the first large-scale multi-task benchmark for CA. Next, we analyze our newly generated dataset, which we make open-source available, with current language models (lightweight BiLSTM, transformer-based MaChAmP) along a novel script compression method, thereby achieving state-of-the-art performance for our target task CA-NER. We also show that CA-TM from the perspective of historical topic models, which are central to Arabic studies, is very challenging. With this interdisciplinary work, we lay the foundations for future research on automatic analysis of CA literature.
Sajawel Ahmed, Rob van der Goot, Misbahur Rehman, Carl Kruse, Ömer Özsoy, Alexander Mehler, Gemma Roig
COLING2
2022 On Language Spaces, Scales and Cross-Lingual Transfer of UD Parsers
abstract
Tanja Samardžić, Ximena Gutierrez-Vasques, Rob van der Goot, Max Müller-Eberstein, Olga Pelloni, Barbara Plank. Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL). 2022.
Tanja Samardzic, Ximena Gutierrez-Vasques, Rob van der Goot, Max Müller-Eberstein, Olga Pelloni, Barbara Plank
CoNLL3
2022 Spectral Probing
abstract
Linguistic information is encoded at varying timescales (subwords, phrases, etc.) and communicative levels, such as syntax and semantics.Contextualized embeddings have analogously been found to capture these phenomena at distinctive layers and frequencies.Leveraging these findings, we develop a fully learnable frequency filter to identify spectral profiles for any given task.It enables vastly more granular analyses than prior handcrafted filters, and improves on efficiency.After demonstrating the informativeness of spectral probing over manual filters in a monolingual setting, we investigate its multilingual characteristics across seven diverse NLP tasks in six languages.Our analyses identify distinctive spectral profiles which quantify cross-task similarity in a linguistically intuitive manner, while remaining consistent across languages-highlighting their potential as robust, lightweight task descriptors.
Max Müller-Eberstein, Rob van der Goot, Barbara Plank
EMNLP2
2022 Frustratingly Easy Performance Improvements for Low-resource Setups: A Tale on BERT and Segment Embeddings
abstract
As input representation for each sub-word, the original BERT architecture proposes the sum of the sub-word embedding, position embedding and a segment embedding. Sub-word and position embeddings are well-known and studied, and encode lexical information and word position, respectively. In contrast, segment embeddings are less known and have so far received no attention, despite being ubiquitous in large pre-trained language models. The key idea of segment embeddings is to encode to which of the two sentences (segments) a word belongs to — the intuition is to inform the model about the separation of sentences for the next sentence prediction pre-training task. However, little is known on whether the choice of segment impacts performance. In this work, we try to fill this gap and empirically study the impact of the segment embedding during inference time for a variety of pre-trained embeddings and target tasks. We hypothesize that for single-sentence prediction tasks performance is not affected — neither in mono- nor multilingual setups — while it matters when swapping segment IDs in paired-sentence tasks. To our surprise, this is not the case. Although for classification tasks and monolingual BERT models no large differences are observed, particularly word-level multilingual prediction tasks are heavily impacted. For low-resource syntactic tasks, we observe impacts of segment embedding and multilingual BERT choice. We find that the default setting for the most used multilingual BERT model underperforms heavily, and a simple swap of the segment embeddings yields an average improvement of 2.5 points absolute LAS score for dependency parsing over 9 different treebanks.
Rob van der Goot, Max Müller-Eberstein, Barbara Plank
LREC1
2022 Sort by Structure: Language Model Ranking as Dependency Probing
abstract
Making an informed choice of pre-trained language model (LM) is critical for performance, yet environmentally costly, and as such widely underexplored. The field of Computer Vision has begun to tackle encoder ranking, with promising forays into Natural Language Processing, however they lack coverage of linguistic tasks such as structured prediction. We propose probing to rank LMs, specifically for parsing dependencies in a given language, by measuring the degree to which labeled trees are recoverable from an LM’s contextualized embeddings. Across 46 typologically and architecturally diverse LM-language pairs, our probing approach predicts the best LM choice 79% of the time using orders of magnitude less compute than training a full parser. Within this study, we identify and analyze one recently proposed decoupled LM—RemBERT—and find it strikingly contains less inherent dependency information, but often yields the best parser after full fine-tuning. Without this outlier our approach identifies the best LM in 89% of cases.
Max Müller-Eberstein, Rob van der Goot, Barbara Plank
NAACL-HLT2
2021 Lexical Normalization for Code-switched Data and its Effect on POS Tagging
abstract
Lexical normalization, the translation of noncanonical data to standard language, has shown to improve the performance of many natural language processing tasks on social media.Yet, using multiple languages in one utterance, also called code-switching (CS), is frequently overlooked by these normalization systems, despite its common use in social media.In this paper, we propose three normalization models specifically designed to handle codeswitched data which we evaluate for two language pairs: Indonesian-English (Id-En) and Turkish-German (Tr-De).For the latter, we introduce novel normalization layers and their corresponding language ID and POS tags for the dataset, and evaluate the downstream effect of normalization on POS tagging.Results show that our CS-tailored normalization models outperform Id-En state of the art and Tr-De monolingual models, and lead to 5.4% relative performance increase for POS tagging as compared to unnormalized input. 1
Rob van der Goot, Özlem Çetinoglu
EACL1
2021 We Need to Talk About train-dev-test Splits
abstract
Standard train-dev-test splits used to benchmark multiple models against each other are ubiquitously used in Natural Language Processing (NLP). In this setup, the train data is used for training the model, the development set for evaluating different versions of the proposed model(s) during development, and the test set to confirm the answers to the main research question(s). However, the introduction of neural networks in NLP has led to a different use of these standard splits; the development set is now often used for model selection during the training procedure. Because of this, comparing multiple versions of the same model during development leads to overestimation on the development data. As an effect, people have started to compare an increasing amount of models on the test data, leading to faster overfitting and ``expiration'' of our test sets. We propose to use a tune-set when developing neural network methods, which can be used for model picking so that comparing the different versions of a new model can safely be done on the development data.
Rob van der Goot
EMNLP (1)1
2021 Genre as Weak Supervision for Cross-lingual Dependency Parsing
abstract
Recent work has shown that monolingual masked language models learn to represent data-driven notions of language variation which can be used for domain-targeted training data selection.Dataset genre labels are already frequently available, yet remain largely unexplored in cross-lingual setups.We harness this genre metadata as a weak supervision signal for targeted data selection in zeroshot dependency parsing.Specifically, we project treebank-level genre information to the finer-grained sentence level, with the goal to amplify information implicitly stored in unsupervised contextualized representations.We demonstrate that genre is recoverable from multilingual contextual embeddings and that it provides an effective signal for training data selection in cross-lingual, zero-shot scenarios.For 12 low-resource language treebanks, six of which are test-only, our genre-specific methods significantly outperform competitive baselines as well as recent embedding-based methods for data selection.Moreover, genre-based data selection provides new state-of-the-art results for three of these target languages.
Max Müller-Eberstein, Rob van der Goot, Barbara Plank
EMNLP (1)2
2021 From Masked Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zero-shot Spoken Language Understanding
abstract
Rob van der Goot, Ibrahim Sharaf, Aizhan Imankulova, Ahmet Üstün, Marija Stepanović, Alan Ramponi, Siti Oryza Khairunnisa, Mamoru Komachi, Barbara Plank. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Rob van der Goot, Ibrahim Sharaf, Aizhan Imankulova, Ahmet Üstün, Marija Stepanovic, Alan Ramponi, Siti Oryza Khairunnisa, Mamoru Komachi, Barbara Plank
NAACL-HLT1
2020 DaN+: Danish Nested Named Entities and Lexical Normalization
abstract
This paper introduces DAN+, a new multi-domain corpus and annotation guidelines for Danish nested named entities (NEs) and lexical normalization to support research on cross-lingual cross-domain learning for a less-resourced language.We empirically assess three strategies to model the two-layer Named Entity Recognition (NER) task.We compare transfer capabilities from German versus in-language annotation from scratch.We examine language-specific versus multilingual BERT, and study the effect of lexical normalization on NER.Our results show that 1) the most robust strategy is multi-task learning which is rivaled by multi-label decoding, 2) BERT-based NER models are sensitive to domain shifts, and 3) in-language BERT and lexical normalization are the most beneficial on the least canonical data.Our results also show that an out-of-domain setup remains challenging, while performance on news plateaus quickly.This highlights the importance of cross-domain evaluation of cross-lingual transfer.
Barbara Plank, Kristian Nørgaard Jensen, Rob van der Goot
COLING3
2020 Biomedical Event Extraction as Sequence Labeling
abstract
We introduce Biomedical Event Extraction as Sequence Labeling (BEESL), a joint endto-end neural information extraction model.BEESL recasts the task as sequence labeling, taking advantage of a multi-label aware encoding strategy and jointly modeling the intermediate tasks via multi-task learning.BEESL is fast, accurate, end-to-end, and unlike current methods does not require any external knowledge base or preprocessing tools.BEESL outperforms the current best system (Li et al., 2019) on the Genia 2011 benchmark by 1.57% absolute F1 score reaching 60.22% F1, establishing a new state of the art for the task.Importantly, we also provide first results on biomedical event extraction without gold entity information.Empirical results show that BEESL's speed and accuracy makes it a viable approach for large-scale real-world scenarios.1
Alan Ramponi, Rob van der Goot, Rosario Lombardo, Barbara Plank
EMNLP (1)2
2020 Synthetic Data for English Lexical Normalization: How Close Can We Get to Manually Annotated Data?
abstract
Social media is a valuable data resource for various natural language processing (NLP) tasks. However, standard NLP tools were often designed with standard texts in mind, and their performance decreases heavily when applied to social media data. One solution to this problem is to adapt the input text to a more standard form, a task also referred to as normalization. Automatic approaches to normalization have shown that they can be used to improve performance on a variety of NLP tasks. However, all of these systems are supervised, thereby being heavily dependent on the availability of training data for the correct language and domain. In this work, we attempt to overcome this dependence by automatically generating training data for lexical normalization. Starting with raw tweets, we attempt two directions, to insert non-standardness (noise) and to automatically normalize in an unsupervised setting. Our best results are achieved by automatically inserting noise. We evaluate our approaches by using an existing lexical normalization system; our best scores are achieved by custom error generation system, which makes use of some manually created datasets. With this system, we score 94.29 accuracy on the test data, compared to 95.22 when it is trained on human-annotated data. Our best system which does not depend on any type of annotation is based on word embeddings and scores 92.04 accuracy. Finally, we perform an experiment in which we asked humans to predict whether a sentence was written by a human or generated by our best model. This experiment showed that in most cases it is hard for a human to detect automatically generated sentences.
Kelly Dekker, Rob van der Goot
LREC2
2020 Norm It! Lexical Normalization for Italian and Its Downstream Effects for Dependency Parsing
abstract
Lexical normalization is the task of translating non-standard social media data to a standard form. Previous work has shown that this is beneficial for many downstream tasks in multiple languages. However, for Italian, there is no benchmark available for lexical normalization, despite the presence of many benchmarks for other tasks involving social media data. In this paper, we discuss the creation of a lexical normalization dataset for Italian. After two rounds of annotation, a Cohen’s kappa score of 78.64 is obtained. During this process, we also analyze the inter-annotator agreement for this task, which is only rarely done on datasets for lexical normalization,and when it is reported, the analysis usually remains shallow. Furthermore, we utilize this dataset to train a lexical normalization model and show that it can be used to improve dependency parsing of social media data. All annotated data and the code to reproduce the results are available at: http://bitbucket.org/robvanderg/normit.
Rob van der Goot, Alan Ramponi, Tommaso Caselli, Michele Cafagna, Lorenzo De Mattei
LREC1
2020 Fair Is Better than Sensational: Man Is to Doctor as Woman Is to Doctor
abstract
Analogies such as man is to king as woman is to X are often used to illustrate the amazing power of word embeddings. Concurrently, they have also been used to expose how strongly human biases are encoded in vector spaces trained on natural language, with examples like man is to computer programmer as woman is to homemaker. Recent work has shown that analogies are in fact not an accurate diagnostic for bias, but this does not mean that they are not used anymore, or that their legacy is fading. Instead of focusing on the intrinsic problems of the analogy task as a bias detection tool, we discuss a series of issues involving implementation as well as subjective choices that might have yielded a distorted picture of bias in word embeddings. We stand by the truth that human biases are present in word embeddings, and, of course, the need to address them. But analogies are not an accurate tool to do so, and the way they have been most often used has exacerbated some possibly non-existing biases and perhaps hidden others. Because they are still widely popular, and some of them have become classics within and outside the NLP community, we deem it important to provide a series of clarifications that should put well-known, and potentially new analogies, into the right perspective.
Malvina Nissim, Rik van Noord, Rob van der Goot
Comput. Linguistics3
2018 Modeling Input Uncertainty in Neural Network Dependency Parsing
abstract
Recently introduced neural network parsers allow for new approaches to circumvent data sparsity issues by modeling character level information and by exploiting raw data in a semi-supervised setting.Data sparsity is especially prevailing when transferring to nonstandard domains.In this setting, lexical normalization has often been used in the past to circumvent data sparsity.In this paper, we investigate whether these new neural approaches provide similar functionality as lexical normalization, or whether they are complementary.We provide experimental results which show that a separate normalization component improves performance of a neural network parser even if it has access to character level information as well as external word embeddings.Further improvements are obtained by a straightforward but novel approach in which the top-N best candidates provided by the normalization component are available to the parser.
Rob van der Goot, Gertjan van Noord
EMNLP1
2018 A Taxonomy for In-depth Evaluation of Normalization for User Generated Content
Rob van der Goot, Rik van Noord, Gertjan van Noord
LREC1
2017 Sharing Is Caring: The Future of Shared Tasks
abstract
Shared tasks are indisputably drivers of progress and interest for problems in NLP. This is reflected by their increasing popularity, as well as by the fact that new shared tasks regularly emerge for under-researched and under-resourced topics, especially at workshops and smaller conferences.The general procedures and conventions for organizing a shared task have arisen organically over time (Paroubek, Chaudiron, and Hirschman, 2007, Section 7). There is no consistent framework that describes how shared tasks should be organized. This is not a harmful thing per se, but we believe that shared tasks, and by extension the field in general, would benefit from some reflection on the existing conventions. This, in turn, could lead to the future harmonization of shared task procedures.Shared tasks revolve around two aspects: research advancement and competition. We see research advancement as the driving force and main goal behind organizing them. Competition is an instrument to encourage and promote participation. However, just because these two forces are intrinsic to shared tasks does not mean that they always act in the same direction: Ensuring that the competition is fair is not a necessary requirement for advancing the field, and might even slow down progress.Our position in this respect is clear: We do believe that (i) advancing the field should be given priority over ensuring fair competition, also because (ii) inequality is partly unsolvable and intrinsic to life. In other words: Equality between competitors is desirable if it does not hinder research advancement.In the recently established workshop on ethics in NLP,1Parra Escartín et al. (2017) raise a set of considerations involving shared tasks, mainly focusing on areas where general ethical concerns regarding good scientific practice intersect with certain aspects of shared tasks. We find that they raise valid concerns, and in this contribution, we address some of them. However, we take a different perspective. Instead of focusing on ethical issues and potential negative effects of the competition aspect, we rather concentrate on how to bolster scientific progress.We make a simple proposal for the improvement of shared tasks, and discuss how it can help to mitigate the problems raised by Parra Escartín et al. (2017), while not necessarily tackling them directly. We start with assessing the concrete impact and significance of such concerns first.Recently, Parra Escartín et al. (2017) drew attention to a list of potential negative effects and ethical issues concerning shared tasks in NLP. In this section, we take this list as a starting point and examine each problem with respect to the main goal of shared tasks—to advance research in the field. Some issues were regarded as potential concerns rather than definite problems, because it is unclear how large their actual impact is. We believe that some of these concerns needed to be quantified in order to be properly assessed.To this end, we reviewed about 100 recent shared tasks from various campaigns (SemEval, EVALITA, CoNLL, WMT, and CLEF) between 2014 and 2016. We focused on several aspects, such as participation of companies, participation of organizers, closed versus open tracks, and the submission of papers by participants. Note that this is not an exhaustive overview of all shared tasks in NLP, but rather an arbitrary sample to investigate general trends in recent times. We use information drawn from this annotation exercise for assessing some of the problems we report in the following sections. The figures that are relevant for the discussion are reported in Table 1. The annotated spreadsheets used to collect this information are publicly available, together with some basic statistics and additional explanations.2Some potential concerns, although being ethically relevant, are not necessarily a problem in terms of research advancement, and fixing them directly should not be a priority. Here, we assess issues raised by Parra Escartín et al. (2017) that we believe fall into this category.Potential Conflicts of Interest. Parra Escartín et al. (2017) state that participation of organizers or annotators in their own shared task raises questions about inequality among participants, as organizers have earlier access to the data than the regular participants. In our survey, we found that in 5.8% of shared tasks, organizers did indeed participate. However, we also observe that this happens in connection with few participants (average 3.5 compared with 12.1, see Table 1), thus typically smaller shared tasks. This indicates that organizers' participation is more common in small, specialized tasks. The low number of participants can also explain why organizers perform better on average compared with non-organizers (see average normalized rank in Table 1).Unequal Playing Field. An unequal playing field mainly reflects the starting point that the different teams have. Parra Escartín et al. (2017) report on the issue of differences in processing power. An extreme example of this issue is the submission of Durrani et al. (2013) at WMT13, in which they reached the highest scores because they were able to boost the BLEU score by approximately 0.8% by making use of 1TB RAM, which was probably unavailable to the other teams at that time. Computational resources are not the only reason for an unequal playing field, though. There are many other causes that could lead to an unequal playing field—for example, some teams might have access to more proprietary data, proprietary software, or research equipment.The competitive nature of shared tasks can be fun, and stimulating for a variety of reasons (visibility, grant applications, beating state of the art, etc.). Such reasons might not necessarily be positively correlated with advancing the field, though. Here we discuss issues also raised by Parra Escartín et al. (2017) that we believe fall into this category.Secretiveness. As a result of the competitive nature of shared tasks, it can be desirable for participating teams to keep their “secret sauce” private, as this could mean an advantage for a re-run of the same task, or a shared task on a related problem. As a possible effect of secretiveness, Parra Escartín et al. (2017) also mention “Unconscious overlooking of ethical concerns,” actually referring to an unacceptable level of vagueness in papers. In other words, participants may unintentionally describe their systems in an abstract and vague way due to a previously established practice in systems' descriptions.Lack of Description of Negative Results. Given that negative results are informative, their under-representation in shared tasks is a concern. A lack of knowledge about negative results might lead to a research redundancy, which is clearly undesirable. The issue of under-represented negative results is a global concern for the entire field, and for science in general. However, shared tasks provide an excellent opportunity for publishing negative results, as the acceptance of papers for publication does not particularly favor positive results.Redundancy and Replicability in the Field. Parra Escartín et al. (2017) raise issues concerning two types of redundancy, (a) when optimal parameter settings of a previous shared task do not carry over to the new version of the task, therefore it is not clear what is learned; and (b) when algorithms are reimplemented for replicability purposes.Regarding (a), we think that differences in used parameter settings are not actually a bad thing; we learn from this that we overfit on the previous task, or that we need to adapt our systems to another data set or domain.Regarding (b), this is a real problem because starting from scratch to reimplement existing systems is unnecessarily time-consuming. In addition, it would always be desirable to be able to directly reproduce the same results of the same model for the same task (Pedersen, 2008; Fokkens et al., 2013).Withdrawal from Competition. Participants may withdraw from a shared task if their ranking in the competition can negatively affect their reputation and/or future funding. For example, Parra Escartín et al. (2017) suggest that companies might prefer to withdraw from the competition if they are not highly ranked, to avoid blemishing their reputation. This is something that we could not quantify in our survey, as in case of withdrawal there would be no evidence of participation in reports. There are two aspects, though, that we can quantify. The first aspect is the number of teams that do not publish their system's description, which amounts to approximately 9%, and could indeed be related to withdrawals. However, exactly because the paper is missing, information on why a team withdrew is not available. The second aspect is the total number of industry participants, which in our sample amounts to 20% (“Some company” and “Only company” in Table 1). Thus, although there is not much that can be done about withdrawal—and this might not be a problem anyway—we believe that, considering the substantial presence and interest of industries so far, their participation should be accommodated.Potentially Gaming the System. Shared tasks are usually bound to data sets and evaluation metrics. This could lead to competition-oriented participants focusing more on tuning their systems on a given data set and metrics rather than finding a scientifically sound and scalable method for solving the problem. This can, in turn, result in an “unfair” ranking or a misleading relation between a methodology and its value with respect to the research problem. A potential negative outcome of the latter is a scenario where “optimal” methods of a shared task do not carry over to related shared tasks. These problems become more severe when system gaming is combined with a secretive attitude. While tackling this issue, we should take into account that participants might be less eager to write about ad hoc solutions, for example tuning pre-processing components or tailoring a system too closely to specifics of the annotation.Our proposal for future shared tasks is not revolutionary. It simply revolves around the key aspects of sharing, not only resources but also experiences, including negative ones. Specifically, with research progress in mind, we believe sharing should be encouraged and even partially enforced. We therefore suggest an explicit setting for shared tasks in NLP, and reflect on the issue of what organizers could do in order to maximize sharing of information regarding participating systems. We also show how such a simple strategy can help to overcome the problems raised that can hinder research advancement.One of the challenges faced by shared tasks is to ensure a level playing field permitting a transparent comparison of the merits of different methods. The problem is that system A might come out on top of system B not because its method is superior, but because, for example, it was trained on more data. This would favor teams with access to more resources, like companies with large quantities of proprietary in-house data.Traditionally, this problem has been mitigated by establishing “closed tracks.” In closed tracks, participating teams are not allowed to use any training data other than that provided by the shared task organizers. The rationale behind this is that if all systems use exactly the same data, the playing field is equal, and the competition results will show the strengths of the different methods. In order to study the effect of additional training data, many shared tasks have a separate competition, the so-called “open track.”However, this open–closed division is increasingly impractical and ineffective. The main problem is that it is only concerned with training data, whereas the performance of systems can crucially depend on other resources. Examples are external components with pre-trained models, such as part-of-speech taggers and dependency parsers, auxiliary data-derived resources like word embeddings, and other influential factors like the availability of computational resources. Because such models are almost always derived from external language data, it is unclear where to draw the line between closed and open. Should such data be disallowed or not? If not, teams still do not really participate on an equal footing.From a research perspective, banning external models is completely impractical and nonsensical, as most state-of-the-art systems now depend on them. Likewise, trying to force all teams to use the same set of external models, and no other, would place a heavy burden on both organizers and participants. Moreover, restricting the resources participants can use is questionable, because, for research to progress quickly, teams should use the best resources available, or the resources best fitting their system.It is therefore unsurprising that the use of closed tracks has declined in shared tasks in general in the last few years, as we have observed during our review of shared tasks. However, the original problem of unequal playing field, and thus a bias in favor of teams with ample resources, remains.We propose, then, to rethink the problem, not in terms of equal training data, but in terms of equal opportunities. This is closely connected to the wider issue of reproducibility and replicability: Like all published research, shared task results should ideally be fully reproducible by anyone (Pedersen, 2008; Fokkens et al., 2013).3 Moreover, it should be easy to build on others' work to try out new variations of a method, without having to reimplement things from scratch. To ensure this, it is desirable that everything needed to reproduce experimental results is publicly and freely available, including code, data, pre-trained models, and so on. Interestingly, at the CoNLL-2013 shared task a similar step was taken, but only in terms of pre-condition: “While all teams in the shared task use the NUCLE corpus, they are also allowed to use additional external resources (both corpora and tools) so long as they are publicly available and not proprietary” (Ng et al., 2013). We would like to take this a step further, by enforcing the sharing of whatever resource teams might choose to use, so as to favor the injection of new resources in the field.Applying this principle to shared tasks in practice, we propose making the primary competition a “public track,” where participants can use any code, data, and pre-trained models they want, as long as others can then freely obtain them. In other words: All resources used to participate in the shared task should be subsequently shared with the community. Although this does not ensure equal access to resources for the current edition, it will still ensure a progressively more equal footing for the future. We believe this is the crucial step to move the field forward, as everyone will have access to the resources used in state-of-the-art systems. To keep participation possible for teams who cannot or will not make all resources available, a secondary, “proprietary track” can be established.Ranking of systems forms a large part of the appeal of shared tasks. However, rankings should not be overemphasized and are far from being the final goal of shared tasks. Research is supposed to teach us about the merits and characteristics of methods, including insights of what does not work, rather than about which team built the system that performed best on the test data.Negative results are very informative for future developments. Although publishing negative results is difficult, shared tasks do provide the ideal context for disclosing and explaining low performance methods and choices. We believe that shared task organizers should explicitly and strongly solicit the inclusion of what did not work in the reports written by participating teams. This could be even solicited via an online form that participants submit after the evaluation phase, where they comment on what worked well (as commonly done), but also provides a separate section to explain what did not work. This information could in turn be valuable data for organizers when compiling the overview report. Moreover, a clear explanation of what did not work, in connection with availability of code, would help to better understand whether something does not work as an idea or because of a specific implementation.More generally, organizers should encourage—and to some point ensure through the reviewing process—that all participants provide exhaustive reports, potentially including ablation/addition tests, so as to have a picture as comprehensive as possible. Because participating in shared tasks directly implies getting a paper accepted for publication, not everyone describes their system to the satisfaction of external reviewers. This should change, and acceptance should be conditional on clarity and exhaustiveness.We stated that the goal of advancing research should be prioritized over competition. This is especially the case when focusing on the competition aspect would encourage undesired practices like secretiveness and gaming the system. We suggest a simple solution based on the principle of maximizing resource- and information-sharing. As a byproduct, some problematic competition-related issues will be overcome, too. Some outstanding ethical issues cannot be solved, as they are intrinsic in human nature and cannot be controlled for by means of specific guidelines.Introducing proprietary and public tracks will stimulate participants to release their systems and resources. This will directly reduce Secretiveness and the issue of Redundancy and Replicability in the Field. It will also partially address the Unequal Playing Field problem, at least in the long run: Even if at the same competition different teams will have access to different resources, all resources will be available to everyone for the next round. Moreover, being able to access and run systems on different data sets will uncover limitations that might have been due to tailoring systems to the specifics of a given shared task (Potential Gaming the System). The presence of a proprietary track still allows for industrial participation (see Withdrawal from Competition in Section 2.2), where distribution of resources might not be as easy as for other teams.Encouraging participants to write comprehensive reports that include negative results will be a valid instrument towards advancing research, at the same time solving some outstanding problems. The reviewers should probably spend extra time in assessing the single reports and accept them conditionally on clarity requirements, but we believe this is worth the effort. Indeed, enforcing that systems are described properly will ensure and the of there are some issues We believe these are issues that cannot or need not be Withdrawal of teams cannot be controlled if of negative results is encouraged and common practice, it is possible that teams will choose to the competition. on progress and sharing rather than will also The of interest is not relevant in our We do believe that organizers should be allowed to and this is especially for shared tasks that might a number of participating teams due to the nature of the As long as these are explicitly reported in both the overview paper and the system the of results and ranking is to the closed tracks also an equal playing will it make it to different methods over the same We do not think Equality will be increasingly by resource sharing, to the that it is as inequality is part of the As we in Section comparison of methods has not been transparent in closed tracks the use of resources is not clarity in reports and release of systems will make it possible for the to assess which methods work and which do not, and to progressively on the state of the that our and will discussion on shared tasks, to make them more for driving progress in and each task will to have their own settings that the organizers will most However, we do believe that participants to release their in terms of resources, and of and should be a common to from long we together and on various with many It is to everyone for their However, we to mention a few who have to into better The discussion we at of the of the Computational at the of was the actual for this are to everyone and in to and for their and valuable has and to the on what does not work with the open versus closed track setting as it We and Parra Escartín for on earlier of this We are also to for
Malvina Nissim, Lasha Abzianidze, Kilian Evang, Rob van der Goot, Hessel Haagsma, Barbara Plank, Martijn Wieling 0001
Comput. Linguistics4
2016 The Denoised Web Treebank: Evaluating Dependency Parsing under Noisy Input Conditions
Joachim Daiber, Rob van der Goot
LREC2