EDBT 2026 Demo / reviewers in the wild / expert
Anders Søgaard
dblp:30/2756
· DBLP profile ↗
134ranked-venue papers
20as first author
50since 2021 · last 2026
0000-0001-5250-4276ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 129 · 20 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review GenerationabstractThe rapid growth of scientific publications has made it increasingly difficult to keep literature reviews comprehensive and up-to-date. Though prior work has focused on automating retrieval and screening, the writing phase of systematic reviews remains largely under-explored, especially with regard to readability and factual accuracy. To address this, we present LiRA (Literature Review Agents), a multi-agent collaborative workflow which emulates the human literature review process. LiRA utilizes specialized agents for content outlining, subsection writing, editing, and reviewing, producing cohesive and comprehensive review articles. Evaluated on SciReviewGen and a proprietary ScienceDirect dataset, LiRA outperforms current baselines such as AutoSurvey and MASS-Survey in writing and citation quality, while maintaining competitive similarity to human-written reviews. We further evaluate LiRA in real-world scenarios using document retrieval and assess its robustness to reviewer model variation. Our findings highlight the potential of agentic LLM workflows, even without domain-specific tuning, to improve the reliability and usability of automated scientific writing. Gregory Hok Tjoan Go, Khang Ly, Anders Søgaard, Seyed Amin Tabatabaei, Maarten de Rijke, Xinyi Chen 0005 |
AAAI | 3 |
| 2026 | Realist and Pluralist Conceptions of Intelligence and Their Implications on AI ResearchabstractIn this paper, we argue that current AI research operates on a spectrum between two different underlying conceptions of intelligence: Intelligence Realism, which holds that intelligence represents a single, universal capacity measurable across all systems, and Intelligence Pluralism, which views intelligence as diverse, context-dependent capacities that cannot be reduced to a single universal measure. Through an analysis of current debates in AI research, we demonstrate how the conceptions remain largely implicit yet fundamentally shape how empirical evidence gets interpreted across a wide range of areas. These underlying views generate fundamentally different research approaches across three areas. Methodologically, they produce different approaches to model selection, benchmark design, and experimental validation. Interpretively, they lead to contradictory readings of the same empirical phenomena, from capability emergence to system limitations. Regarding AI risk, they generate categorically different assessments: realists view superintelligence as the primary risk and search for unified alignment solutions, while pluralists see diverse threats across different domains requiring context-specific solutions. We argue that making explicit these underlying assumptions can contribute to a clearer understanding of disagreements in AI research. Ninell Oldenburg, Ruchira Dhar, Anders Søgaard |
AAAI | 3 |
| 2026 | LLM Beliefs Are in Their HeadsabstractWe investigate belief-like representations in decoder-only autoregressive LLMs using linear controlled probes on residual stream activations and single attention heads. Following Herrmann and Levinstein’s (2025) criteria (Accuracy, Use, Coherence, and Uniformity) we find that large models exhibit strong truth sensitivity (Accuracy), and steering activations along probe directions reliably changes downstream behavior (Use). Coherence, measured via calibrated probes and cross-dataset probing, is moderate across models, while training on diverse data yields domain-consistent truth directions (Uniformity). The results are particularly encouraging at the head level and align with some standard philosophical accounts of belief, e.g., minimal functionalism, supporting the view that LLMs can maintain propositional attitudes under such theoretical frameworks. Alessandro Corona Mendozza, Anders Søgaard |
ACL (1) | 2 |
| 2025 | Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired UsersabstractAntonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Antonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard |
ACL (1) | 8 |
| 2025 | Do Language Models Have Semantics? On the Five Standard PositionsabstractWe identify five positions on whether large language models (LLMs) and chatbots can be said to exhibit semantic understanding. These positions differ in whether they attribute semantics to LLMs and/or chatbots trained on feedback, what kind of semantics they attribute (inferential or referential), and in virtue of what they attribute referential semantics (internal or external causes). This allows for 24 = 16 logically possible positions, only five of which have been argued for. Based on a pairwise comparison of these five positions, we conclude that the better theory of semantics in large language models is, in fact, a sixth combination: Both large language models and chatbots have inferential and referential semantics, grounded in both internal and external causes. Anders Søgaard |
ACL (1) | 1 |
| 2025 | Understanding Subword Compositionality of Large Language ModelsabstractLarge language models (LLMs) take sequences of subwords as input, requiring them to effective compose subword representations into meaningful word-level representations.In this paper, we present a comprehensive set of experiments to probe how LLMs compose subword information, focusing on three key aspects: structural similarity, semantic decomposability, and form retention.Our analysis of the experiments suggests that five LLM families can be classified into three distinct groups, likely reflecting difference in their underlying composition strategies.Specifically, we observe (i) three distinct patterns in the evolution of structural similarity between subword compositions and whole-word representations across layers; (ii) great performance when probing layer by layer their sensitivity to semantic decompositionality; and (iii) three distinct patterns when probing sensitivity to formal features, e.g., character sequence length.These findings provide valuable insights into the compositional dynamics of LLMs and highlight different compositional pattens in how LLMs encode and integrate subword information. Qiwei Peng 0003, Yekun Chai, Anders Søgaard |
EMNLP | 3 |
| 2025 | Debiasing Multilingual LLMs in Cross-lingual Latent SpaceabstractDebiasing techniques such as SentDebias aim to reduce bias in large language models (LLMs).Previous studies have evaluated their cross-lingual transferability by directly applying these methods to LLM representations, revealing their limited effectiveness across languages.In this work, we therefore propose to perform debiasing in a joint latent space rather than directly on LLM representations.We construct a well-aligned cross-lingual latent space using an autoencoder trained on parallel TED talk scripts.Our experiments with Aya-expanse and two debiasing techniques across four languages (English, French, German, Dutch) demonstrate that a) autoencoders effectively construct a well-aligned cross-lingual latent space, and b) applying debiasing techniques in the learned cross-lingual latent space significantly improves both the overall debiasing performance and cross-lingual transferability. Qiwei Peng 0003, Guimin Hu, Yekun Chai, Anders Søgaard |
EMNLP | 4 |
| 2025 | A Study on The Impact of Foundation Models on Automatic Depression Detection from Speech SignalsabstractAn automatic depression detection (ADD) system using spoken language offers the opportunity to develop practical, low-cost tools to detect symptoms early. However, limited data availability, privacy concerns, and transcription efforts pose significant challenges. Recent advancements in foundational models, capable of understanding and processing multimodal inputs, present opportunities for enhancing ADD systems. This study explores various speech foundation models to investigate their impact on ADD. We leverage Whisper and MMS for automatic transcription and integrate speech and text embeddings into a language model optimized with low-rank adaptation (LoRA). In addition, we examine the effects of fine-tuning strategies and prompt formats on model performance. We used English and Bengali datasets to demonstrate the potential of our method in ADD, even with moderate-quality transcriptions. The best speech and language foundation models outperform baseline models on both datasets. Bubai Maji, Monorama Swain, Shazia Nasreen, Debabrata Majumdar, Rajlakshmi Guha, Aurobinda Routray, Anders Søgaard |
INTERSPEECH | 7 |
| 2025 | Ethical Concern Identification in NLP: A Corpus of ACL Anthology Ethics StatementsabstractAntonia Karamolegkou, Sandrine Schiller Hansen, Ariadni Christopoulou, Filippos Stamatiou, Anne Lauscher, Anders Søgaard. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Antonia Karamolegkou, Sandrine Schiller Hansen, Ariadni Christopoulou, Filippos Stamatiou, Anne Lauscher, Anders Søgaard |
NAACL (Long Papers) | 6 |
| 2024 | Evaluating Webcam-based Gaze Data as an Alternative for Human Rationale AnnotationsabstractRationales in the form of manually annotated input spans usually serve as ground truth when evaluating explainability methods in NLP. They are, however, time-consuming and often biased by the annotation process. In this paper, we debate whether human gaze, in the form of webcam-based eye-tracking recordings, poses a valid alternative when evaluating importance scores. We evaluate the additional information provided by gaze data, such as total reading times, gaze entropy, and decoding accuracy with respect to human rationale annotations. We compare WebQAmGaze, a multilingual dataset for information-seeking QA, with attention and explainability-based importance scores for 4 different multilingual Transformer-based language models (mBERT, distil-mBERT, XLMR, and XLMR-L) and 3 languages (English, Spanish, and German). Our pipeline can easily be applied to other tasks and languages. Our findings suggest that gaze data offers valuable linguistic insights that could be leveraged to infer task difficulty and further show a comparable ranking of explainability methods to that of human rationales. Stephanie Brandl, Oliver Eberle, Anders Søgaard, Nora Hollenstein |
LREC/COLING | 4 |
| 2024 | Unlocking Markets: A Multilingual Benchmark to Cross-Market Question AnsweringabstractUsers post numerous product-related questions on e-commerce platforms, affecting their purchase decisions.Product-related question answering (PQA) entails utilizing product-related resources to provide precise responses to users.We propose a novel task of Multilingual Crossmarket Product-based Question Answering (MCPQA) and define the task as providing answers to product-related questions in a main marketplace by utilizing information from another resource-rich auxiliary marketplace in a multilingual context.We introduce a largescale dataset comprising over 7 million questions from 17 marketplaces across 11 languages.We then perform automatic translation on the Electronics category of our dataset, naming it as McMarket.We focus on two subtasks: review-based answer generation and productrelated question ranking.For each subtask, we label a subset of McMarket using an LLM and further evaluate the quality of the annotations via human assessment.We then conduct experiments to benchmark our dataset, using models ranging from traditional lexical models to LLMs in both single-market and cross-market scenarios across McMarket and the corresponding LLM subset.Results show that incorporating cross-market information significantly enhances performance in both tasks. Yifei Yuan 0002, Yang Deng 0002, Anders Søgaard, Mohammad Aliannejadi |
EMNLP | 3 |
| 2024 | Concept Space Alignment in Multilingual LLMsabstractMultilingual large language models (LLMs) seem to generalize somewhat across languages.We hypothesize this is a result of implicit vector space alignment.Evaluating such alignment, we see that larger models exhibit very highquality linear alignments between corresponding concepts in different languages.Our experiments show that multilingual LLMs suffer from two familiar weaknesses: generalization works best for languages with similar typology, and for abstract concepts.For some models, e.g., the Llama-2 family of models, prompt-based embeddings align better than word embeddings, but the projections are less linear -an observation that holds across almost all model families, indicating that some of the implicitly learned alignments are broken somewhat by promptbased methods. Qiwei Peng 0003, Anders Søgaard |
EMNLP | 2 |
| 2024 | Defining Knowledge: Bridging Epistemology and Large Language ModelsabstractKnowledge claims are abundant in the literature on large language models (LLMs); but can we say that GPT-4 truly "knows" the Earth is round?To address this question, we review standard definitions of knowledge in epistemology and we formalize interpretations applicable to LLMs.In doing so, we identify inconsistencies and gaps in how current NLP research conceptualizes knowledge with respect to epistemological frameworks.Additionally, we conduct a survey of 100 professional philosophers and computer scientists to compare their preferences in knowledge definitions and their views on whether LLMs can really be said to know.Finally, we suggest evaluation protocols for testing knowledge in accordance to the most relevant definitions. Constanza Fierro, Ruchira Dhar, Filippos Stamatiou, Nicolas Garneau, Anders Søgaard |
EMNLP | 5 |
| 2024 | FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food CultureabstractWenyan Li, Crystina Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu, Yifei Yuan, Anders Søgaard, Daniel Hershcovich, Desmond Elliott. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Wenyan Li 0001, Xinyu Zhang 0018, Jiaang Li 0002, Qiwei Peng 0003, Raphael Tang, Li Zhou 0010, Weijia Zhang 0004, Guimin Hu, Yifei Yuan 0002, Anders Søgaard, Daniel Hershcovich, Desmond Elliott |
EMNLP | 10 |
| 2024 | On Mitigating Performance Disparities in Multilingual Speech RecognitionabstractHow far have we come in mitigating performance disparities across genders in multilingual speech recognition?We compare the impact on gender disparity of different finetuning algorithms for automated speech recognition across model sizes, languages and gender.We look at both performance-focused and fairness-promoting algorithms.Across languages, we see slightly better performance for female speakers for larger models regardless of the fine-tuning algorithm.The best tradeoff between performance and parity is found using adapter fusion.Fairness-promoting finetuning algorithms (Group-DRO and Spectral Decoupling) hurt performance compared to adapter fusion with only slightly better performance parity.LoRA increases disparities slightly.Fairness-mitigating fine-tuning techniques led to slightly higher variance in performance across languages, with the exception of adapter fusion. Monorama Swain, Anna Zee, Anders Søgaard |
EMNLP | 3 |
| 2024 | CreoleVal: Multilingual Multitask Benchmarks for CreolesabstractAbstract Creoles represent an under-explored and marginalized group of languages, with few available resources for NLP research. While the genealogical ties between Creoles and a number of highly resourced languages imply a significant potential for transfer learning, this potential is hampered due to this lack of annotated data. In this work we present CreoleVal, a collection of benchmark datasets spanning 8 different NLP tasks, covering up to 28 Creole languages; it is an aggregate of novel development datasets for reading comprehension relation classification, and machine translation for Creoles, in addition to a practical gateway to a handful of preexisting benchmarks. For each benchmark, we conduct baseline experiments in a zero-shot setting in order to further ascertain the capabilities and limitations of transfer learning for Creoles. Ultimately, we see CreoleVal as an opportunity to empower research on Creoles in NLP and computational linguistics, and in general, a step towards more equitable language technology around the globe. Heather C. Lent, Kushal Tatariya, Raj Dabre, Yiyi Chen 0002, Marcell Fekete, Esther Ploeger, Li Zhou 0010, Ruth-Ann Armstrong, Abee Eijansantos, Catriona Malau, Hans Erik Heje, Ernests Lavrinovics, Diptesh Kanojia, Paul Belony, Marcel Bollmann, Loïc Grobol, Miryam de Lhoneux, Daniel Hershcovich, Michel DeGraff, Anders Søgaard, Johannes Bjerva |
Trans. Assoc. Comput. Linguistics | 20 |
| 2024 | Do Vision and Language Models Share Concepts? A Vector Space Alignment StudyabstractAbstract Large-scale pretrained language models (LMs) are said to “lack the ability to connect utterances to the world” (Bender and Koller, 2020), because they do not have “mental models of the world” (Mitchell and Krakauer, 2023). If so, one would expect LM representations to be unrelated to representations induced by vision models. We present an empirical evaluation across four families of LMs (BERT, GPT-2, OPT, and LLaMA-2) and three vision model architectures (ResNet, SegFormer, and MAE). Our experiments show that LMs partially converge towards representations isomorphic to those of vision models, subject to dispersion, polysemy, and frequency. This has important implications for both multi-modal processing and the LM understanding debate (Mitchell and Krakauer, 2023).1 Jiaang Li 0002, Yova Kementchedjhieva, Constanza Fierro, Anders Søgaard |
Trans. Assoc. Comput. Linguistics | 4 |
| 2023 | LeXFiles and LegalLAMA: Facilitating English Multinational Legal Language Model DevelopmentabstractIn this work, we conduct a detailed analysis on the performance of legal-oriented pretrained language models (PLMs).We examine the interplay between their original objective, acquired knowledge, and legal language understanding capacities which we define as the upstream, probing, and downstream performance, respectively.We consider not only the models' size but also the pre-training corpora used as important dimensions in our study.To this end, we release a multinational English legal corpus (LeXFiles) and a legal knowledge probing benchmark (LegalLAMA) to facilitate training and detailed analysis of legal-oriented PLMs.We release two new legal PLMs trained on LeXFiles and evaluate them alongside others on LegalLAMA and LexGLUE.We find that probing performance strongly correlates with upstream performance in related legal topics.On the other hand, downstream performance is mainly driven by the model's size and prior legal knowledge which can be estimated by upstream and probing performance.Based on these findings, we can conclude that both dimensions are important for those seeking the development of domain-specific PLMs. Ilias Chalkidis, Nicolas Garneau, Catalina Goanta, Daniel Martin Katz, Anders Søgaard |
ACL (1) | 5 |
| 2023 | What does the Failure to Reason with "Respectively" in Zero/Few-Shot Settings Tell Us about Language Models?abstractHumans can effortlessly understand the coordinate structure of sentences such as “Niels Bohr and Kurt Cobain were born in Copenhagen and Seattle, respectively”. In the context of natural language inference (NLI), we examine how language models (LMs) reason with respective readings (Gawron and Kehler, 2004) from two perspectives: syntactic-semantic and commonsense-world knowledge. We propose a controlled synthetic dataset WikiResNLI and a naturally occurring dataset NatResNLI to encompass various explicit and implicit realizations of “respectively”. We show that fine-tuned NLI models struggle with understanding such readings without explicit supervision. While few-shot learning is easy in the presence of explicit cues, longer training is required when the reading is evoked implicitly, leaving models to rely on common sense inferences. Furthermore, our fine-grained analysis indicates models fail to generalize across different constructions. To conclude, we demonstrate that LMs still lag behind humans in generalizing to the long tail of linguistic constructions. Ruixiang Cui, Seolhwa Lee, Daniel Hershcovich, Anders Søgaard |
ACL (1) | 4 |
| 2023 | Being Right for Whose Right Reasons?abstractExplainability methods are used to benchmark the extent to which model predictions align with human rationales i.e., are 'right for the right reasons'.Previous work has failed to acknowledge, however, that what counts as a rationale is sometimes subjective.This paper presents what we think is a first of its kind, a collection of human rationale annotations augmented with the annotators demographic information.We cover three datasets spanning sentiment analysis and common-sense reasoning, and six demographic groups (balanced across age and ethnicity).Such data enables us to ask both what demographics our predictions align with and whose reasoning patterns our models' rationales align with.We find systematic inter-group annotator disagreement and show how 16 Transformer-based models align better with rationales provided by certain demographic groups: We find that models are biased towards aligning best with older and/or white annotators.We zoom in on the effects of model size and model distillation, finding -contrary to our expectations -negative correlations between model size and rationale agreement as well as no evidence that either model size or model distillation improves fairness. Terne Sasha Thorn Jakobsen, Laura Cabello Piqueras, Anders Søgaard |
ACL (1) | 3 |
| 2023 | A Two-Sided Discussion of Preregistration of NLP ResearchabstractVan Miltenburg et al. (2021) suggest NLP research should adopt preregistration to prevent fishing expeditions and to promote publication of negative results.At face value, this is a very reasonable suggestion, seemingly solving many methodological problems with NLP research.We discuss pros and cons-some old, some new: a) Preregistration is challenged by the practice of retrieving hypotheses after the results are known; b) preregistration may bias NLP toward confirmatory research; c) preregistration must allow for reclassification of research as exploratory; d) preregistration may increase publication bias; e) preregistration may increase flag-planting; f) preregistration may increase p-hacking; and finally, g) preregistration may make us less risk tolerant.We cast our discussion as a dialogue, presenting both sides of the debate. Anders Søgaard, Daniel Hershcovich, Miryam de Lhoneux |
EACL | 1 |
| 2023 | Copyright Violations and Large Language ModelsabstractLanguage models may memorize more than just facts, including entire chunks of texts seen during training.Fair use exemptions to copyright laws typically allow for limited use of copyrighted material without permission from the copyright holder, but typically for extraction of information from copyrighted materials, rather than verbatim reproduction.This work explores the issue of copyright violations and large language models through the lens of verbatim memorization, focusing on possible redistribution of copyrighted text.We present experiments with a range of language models over a collection of popular books and coding problems, providing a conservative characterization of the extent to which language models can redistribute these materials.Overall, this research highlights the need for further examination and the potential impact on future developments in natural language processing to ensure adherence to copyright regulations. Antonia Karamolegkou, Jiaang Li 0002, Li Zhou 0010, Anders Søgaard |
EMNLP | 4 |
| 2023 | Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language ModelsabstractLanguage models such as mBERT, XLM-R, and BLOOM aim to achieve multilingual generalization or compression to facilitate transfer to a large number of (potentially unseen) languages. However, these models should ideally also be private, linguistically fair, and transparent, by relating their predictions to training data. Can these requirements be simultaneously satisfied? We show that multilingual compression and linguistic fairness are compatible with differential privacy, but that differential privacy is at odds with training data influence sparsity, an objective for transparency. We further present a series of experiments on two common NLP tasks and evaluate multilingual compression and training data influence sparsity under different privacy guarantees, exploring these trade-offs in more detail. Our results suggest that we need to develop ways to jointly optimize for these objectives in order to find practical trade-offs. Phillip Rust, Anders Søgaard |
ICML | 2 |
| 2023 | Re-Framing Case Law Citation Prediction from a Paragraph PerspectiveabstractCase law citation prediction, i.e., predicting what historical cases are relevant for your current case, can assist legal discovery and decision-making, but legal documents are long, and often only parts of them are relevant for a particular use case. We therefore reframe case law citation prediction as a paragraph-to-paragraph citation task, introduce a new dataset, and train and evaluate new models. We also evaluate our models qualitatively. Our resources provide a first step toward discovering citation patterns and modeling legal rules in EU law from precedent documents. Henrik Palmer Olsen, Nicolas Garneau, Yannis Panagis, Johan Lindholm, Anders Søgaard |
JURIX | 5 |
| 2023 | Private Meeting Summarization Without Performance LossabstractMeeting summarization has an enormous business potential, but in addition to being a hard problem, roll-out is challenged by privacy concerns. We explore the problem of meeting summarization under differential privacy constraints and find, to our surprise, that while differential privacy leads to slightly lower performance on in-sample data, differential privacy improves performance when evaluated on unseen meeting types. Since meeting summarization systems will encounter a great variety of meeting types in practical employment scenarios, this observation makes safe meeting summarization seem much more feasible. We perform extensive error analysis and identify potential risks in meeting summarization under differential privacy, including a faithfulness analysis. Seolhwa Lee, Anders Søgaard |
SIGIR | 2 |
| 2022 | Word Order Does Matter and Shuffled Language Models Know ItabstractRecent studies have shown that language models pretrained and/or fine-tuned on randomly permuted sentences exhibit competitive performance on GLUE, putting into question the importance of word order information. Somewhat counter-intuitively, some of these studies also report that position embeddings appear to be crucial for models' good performance with shuffled text. We probe these language models for word order information and investigate what position embeddings learned from shuffled text encode, showing that these models retain information pertaining to the original, naturalistic word order. We show this is in part due to a subtlety in how shuffling is implemented in previous work -before rather than after subword segmentation. Surprisingly, we find even Language models trained on text shuffled after subword segmentation retain some semblance of information about word order because of the statistical dependencies between sentence length and unigram probabilities. Finally, we show that beyond GLUE, a variety of language understanding tasks do require word order information, often to an extent that cannot be learned through fine-tuning. * Equal contribution. Order was decided by a coin toss. Mostafa Abdou, Vinit Ravishankar, Artur Kulmizev, Anders Søgaard |
ACL (1) | 4 |
| 2022 | FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text ProcessingabstractIlias Chalkidis, Tommaso Pasini, Sheng Zhang, Letizia Tomada, Sebastian Schwemer, Anders Søgaard. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Ilias Chalkidis, Tommaso Pasini, Sheng Zhang 0022, Letizia Tomada, Sebastian Felix Schwemer, Anders Søgaard |
ACL (1) | 6 |
| 2022 | Do Transformer Models Show Similar Attention Patterns to Task-Specific Human Gaze?abstractLearned self-attention functions in state-of-theart NLP models often correlate with human attention.We investigate whether self-attention in large-scale pre-trained language models is as predictive of human eye fixation patterns during task-reading as classical cognitive models of human attention.We compare attention functions across two task-specific reading datasets for sentiment analysis and relation extraction.We find the predictiveness of large-scale pretrained self-attention for human attention depends on 'what is in the tail', e.g., the syntactic nature of rare contexts.Further, we observe that task-specific fine-tuning does not increase the correlation with human task-specific reading.Through an input reduction experiment we give complementary insights on the sparsity and fidelity trade-off, showing that lowerentropy attention vectors are more faithful. Oliver Eberle, Stephanie Brandl, Jonas Pilot, Anders Søgaard |
ACL (1) | 4 |
| 2022 | Challenges and Strategies in Cross-Cultural NLPabstractDaniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Daniel Hershcovich, Stella Frank, Heather C. Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Aikaterini Margatina, Phillip Rust, Anders Søgaard |
ACL (1) | 14 |
| 2022 | Are Pretrained Multilingual Models Equally Fair across Languages?abstractPretrained multilingual language models can help bridge the digital language divide, enabling high-quality NLP models for lower-resourced languages. Studies of multilingual models have so far focused on performance, consistency, and cross-lingual generalisation. However, with their wide-spread application in the wild and downstream societal impact, it is important to put multilingual models under the same scrutiny as monolingual models. This work investigates the group fairness of multilingual models, asking whether these models are equally fair across languages. To this end, we create a new four-way multilingual dataset of parallel cloze test examples (MozArt), equipped with demographic information (balanced with regard to gender and native tongue) about the test participants. We evaluate three multilingual models on MozArt –mBERT, XLM-R, and mT5– and show that across the four target languages, the three models exhibit different levels of group disparity, e.g., exhibiting near-equal risk for Spanish, but high levels of disparity for German. Laura Cabello Piqueras, Anders Søgaard |
COLING | 2 |
| 2022 | Should We Ban English NLP for a Year?abstractAround two thirds of NLP research at top venues is devoted exclusively to developing technology for speakers of English, most speech data comes from young urban speakers, and most texts used to train language models come from male writers.These biases feed into consumer technologies to widen existing inequality gaps, not only within, but also across, societies.Many have argued that it is almost impossible to mitigate inequality amplification.I argue that, on the contrary, it is quite simple to do so, and that counter-measures would have little-to-no negative impact, except for, perhaps, in the very short term. Anders Søgaard |
EMNLP | 1 |
| 2022 | Date Recognition in Historical Parish Records
Laura Cabello Piqueras, Constanza Fierro, Jonas F. Lotz, Phillip Rust, Joen Rommedahl, Jeppe Klok Due, Christian Igel, Desmond Elliott, Carsten B. Pedersen, Israfel Salazar, Anders Søgaard |
ICFHR | 11 |
| 2022 | Evaluating Deep Taylor Decomposition for Reliability Assessment in the Wild
Stephanie Brandl, Daniel Hershcovich, Anders Søgaard |
ICWSM | 3 |
| 2022 | What a Creole Wants, What a Creole NeedsabstractIn recent years, the natural language processing (NLP) community has given increased attention to the disparity of efforts directed towards high-resource languages over low-resource ones. Efforts to remedy this delta often begin with translations of existing English datasets into other languages. However, this approach ignores that different language communities have different needs. We consider a group of low-resource languages, creole languages. Creoles are both largely absent from the NLP literature, and also often ignored by society at large due to stigma, despite these languages having sizable and vibrant communities. We demonstrate, through conversations with creole experts and surveys of creole-speaking communities, how the things needed from language technology can change dramatically from one language to another, even when the languages are considered to be very similar to each other, as with creoles. We discuss the prominent themes arising from these conversations, and ultimately demonstrate that useful language technology cannot be built without involving the relevant community. Heather C. Lent, Kelechi Ogueji, Miryam de Lhoneux, Orevaoghene Ahia, Anders Søgaard |
LREC | 5 |
| 2022 | How Conservative are Language Models? Adapting to the Introduction of Gender-Neutral PronounsabstractGender-neutral pronouns have recently been introduced in many languages to a) include non-binary people and b) as a generic singular.Recent results from psycholinguistics suggest that gender-neutral pronouns (in Swedish) are not associated with human processing difficulties.This, we show, is in sharp contrast with automated processing.We show that gender-neutral pronouns in Danish, English, and Swedish are associated with higher perplexity, more dispersed attention patterns, and worse downstream performance.We argue that such conservativity in language models may limit widespread adoption of gender-neutral pronouns and must therefore be resolved. Stephanie Brandl, Ruixiang Cui, Anders Søgaard |
NAACL-HLT | 3 |
| 2022 | Generalized Quantifiers as a Source of Error in Multilingual NLU BenchmarksabstractLogical approaches to representing language have developed and evaluated computational models of quantifier words since the 19th century, but today's NLU models still struggle to capture their semantics.We rely on Generalized Quantifier Theory for languageindependent representations of the semantics of quantifier words, to quantify their contribution to the errors of NLU models.We find that quantifiers are pervasive in NLU benchmarks, and their occurrence at test time is associated with performance drops.Multilingual models also exhibit unsatisfying quantifier reasoning abilities, but not necessarily worse for non-English languages.To facilitate directlytargeted probing, we present an adversarial generalized quantifier NLI task (GQNLI) and show that pre-trained language models have a clear lack of robustness in generalized quantifier reasoning. Ruixiang Cui, Daniel Hershcovich, Anders Søgaard |
NAACL-HLT | 3 |
| 2021 | Joint Semantic Analysis with Document-Level Cross-Task Coherence RewardsabstractCoreference resolution and semantic role labeling are NLP tasks that capture different aspects of semantics, indicating respectively, which expressions refer to the same entity, and what semantic roles expressions serve in the sentence. However, they are often closely interdependent, and both generally necessitate natural language understanding. Do they form a coherent abstract representation of documents? We present a neural network architecture for joint coreference resolution and semantic role labeling for English, and train graph neural networks to model the 'coherence' of the combined shallow semantic graph. Using the resulting coherence score as a reward for our joint semantic analyzer, we use reinforcement learning to encourage global coherence over the document and between semantic annotations. This leads to improvements on both tasks in multiple datasets from different domains, and across a range of encoders of different expressivity, calling, we believe, for a more holistic approach for semantics in NLP. Rahul Aralikatte, Mostafa Abdou, Heather C. Lent, Daniel Hershcovich, Anders Søgaard |
AAAI | 5 |
| 2021 | Analogy Training Multilingual EncodersabstractLanguage encoders encode words and phrases in ways that capture their local semantic relatedness, but are known to be globally inconsistent. Global inconsistency can seemingly be corrected for, in part, by leveraging signals from knowledge bases, but previous results are partial and limited to monolingual English encoders. We extract a large-scale multilingual, multi-word analogy dataset from Wikidata for diagnosing and correcting for global inconsistencies, and then implement a four-way Siamese BERT architecture for grounding multilingual BERT (mBERT) in Wikidata through analogy training. We show that analogy training not only improves the global consistency of mBERT, as well as the isomorphism of language-specific subspaces, but also leads to consistent gains on downstream tasks such as bilingual dictionary induction and sentence retrieval. Nicolas Garneau, Mareike Hartmann, Anders Sandholm 0001, Sebastian Ruder, Ivan Vulic, Anders Søgaard |
AAAI | 6 |
| 2021 | Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in ColorabstractPretrained language models have been shown to encode relational information, such as the relations between entities or concepts in knowledge-bases -(Paris, Capital, France).However, simple relations of this type can often be recovered heuristically and the extent to which models implicitly reflect topological structure that is grounded in world, such as perceptual structure, is unknown.To explore this question, we conduct a thorough case study on color.Namely, we employ a dataset of monolexemic color terms and color chips represented in CIELAB, a color space with a perceptually meaningful distance metric.Using two methods of evaluating the structural alignment of colors in this space with textderived color term representations, we find significant correspondence.Analyzing the differences in alignment across the color spectrum, we find that warmer colors are, on average, better aligned to the perceptual color space than cooler ones, suggesting an intriguing connection to findings from recent work on efficient communication in color naming.Further analysis suggests that differences in alignment are, in part, mediated by collocationality and differences in syntactic usage, posing questions as to the relationship between color perception and usage and context. Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, Anders Søgaard |
CoNLL | 6 |
| 2021 | A Multilingual Benchmark for Probing Negation-Awareness with Minimal PairsabstractMareike Hartmann, Miryam de Lhoneux, Daniel Hershcovich, Yova Kementchedjhieva, Lukas Nielsen, Chen Qiu, Anders Søgaard. Proceedings of the 25th Conference on Computational Natural Language Learning. 2021. Mareike Hartmann, Miryam de Lhoneux, Daniel Hershcovich, Yova Kementchedjhieva, Lukas Nielsen, Chen Qiu 0005, Anders Søgaard |
CoNLL | 7 |
| 2021 | On Language Models for CreolesabstractCreole languages such as Nigerian Pidgin English and Haitian Creole are under-resourced and largely ignored in the NLP literature.Creoles typically result from the fusion of a foreign language with multiple local languages, and what grammatical and lexical features are transferred to the creole is a complex process (Sessarego, 2020).While creoles are generally stable, the prominence of some features may be much stronger with certain demographics or in some linguistic situations (Winford, 1999;Patrick, 1999).This paper makes several contributions: We collect existing corpora and release models for Haitian Creole, Nigerian Pidgin English, and Singaporean Colloquial English.We evaluate these models on intrinsic and extrinsic tasks.Motivated by the above literature, we compare standard language models with distributionally robust ones and find that, somewhat surprisingly, the standard language models are superior to the distributionally robust ones.We investigate whether this is an effect of overparameterization or relative distributional stability, and find that the difference persists in the absence of over-parameterization, and that drift is limited, confirming the relative stability of creole languages. Heather C. Lent, Emanuele Bugliarello, Miryam de Lhoneux, Chen Qiu 0005, Anders Søgaard |
CoNLL | 5 |
| 2021 | Ellipsis Resolution as Question Answering: An EvaluationabstractMost, if not all forms of ellipsis (e.g., 'so does Mary') are similar to reading comprehension questions ('what does Mary do'), in that in order to resolve them, we need to identify an appropriate text span in the preceding discourse.Following this observation, we present an alternative approach for English ellipsis resolution relying on architectures developed for question answering (QA).We present both single-task models, and joint models trained on auxiliary QA and coreference resolution datasets, clearly outperforming the current state of the art for Sluice Ellipsis (from 70.00 to 86.01 F 1 ) and Verb Phrase Ellipsis (from 72.89 to 78.66 F 1 ). Rahul Aralikatte, Matthew Lamm, Daniel Hardt, Anders Søgaard |
EACL | 4 |
| 2021 | Error Analysis and the Role of MorphologyabstractWe evaluate two common conjectures in error analysis of NLP models: (i) Morphology is predictive of errors; and (ii) the importance of morphology increases with the morphological complexity of a language. We show across four different tasks and up to 57 languages that of these conjectures, somewhat surprisingly, only (i) is true. Using morphological features does improve error prediction across tasks; however, this effect is less pronounced with morphologically complex languages. We speculate this is because morphology is more discriminative in morphologically simple languages. Across all four tasks, case and gender are the morphological features most predictive of error. Marcel Bollmann, Anders Søgaard |
EACL | 2 |
| 2021 | Attention Can Reflect Syntactic Structure (If You Let It)abstractVinit Ravishankar, Artur Kulmizev, Mostafa Abdou, Anders Søgaard, Joakim Nivre. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Vinit Ravishankar, Artur Kulmizev, Mostafa Abdou, Anders Søgaard, Joakim Nivre |
EACL | 4 |
| 2021 | We Need To Talk About Random SplitsabstractGorman and Bedrick (2019) argued for using random splits rather than standard splits in NLP experiments.We argue that random splits, like standard splits, lead to overly optimistic performance estimates.We can also split data in biased or adversarial ways, e.g., training on short sentences and evaluating on long ones.Biased sampling has been used in domain adaptation to simulate real-world drift; this is known as the covariate shift assumption.In NLP, however, even worst-case splits, maximizing bias, often under-estimate the error observed on new samples of in-domain data, i.e., the data that models should minimally generalize to at test time.This invalidates the covariate shift assumption.Instead of using multiple random splits, future benchmarks should ideally include multiple, independent test sets instead; if infeasible, we argue that multiple biased splits leads to more realistic performance estimates than multiple random splits. Anders Søgaard, Sebastian Ebert, Jasmijn Bastings, Katja Filippova |
EACL | 1 |
| 2021 | Sociolectal Analysis of Pretrained Language ModelsabstractUsing data from English cloze tests, in which subjects also self-reported their gender, age, education, and race, we examine performance differences of pretrained language models across demographic groups, defined by these (protected) attributes.We demonstrate wide performance gaps across demographic groups and show that pretrained language models systematically disfavor young non-white male speakers; i.e., not only do pretrained language models learn social biases (stereotypical associations) -pretrained language models also learn sociolectal biases, learning to speak more like some than like others.We show, however, that, with the exception of BERT models, larger pretrained language models reduce some the performance gaps between majority and minority groups. Sheng Zhang 0022, Xin Zhang 0018, Anders Søgaard |
EMNLP (1) | 4 |
| 2021 | The Effect of Round-Trip Translation on Fairness in Sentiment AnalysisabstractSentiment analysis systems have been shown to exhibit sensitivity to protected attributes.Round-trip translation, on the other hand, has been shown to normalize text.We explore the impact of round-trip translation on the demographic parity of sentiment classifiers and show how round-trip translation consistently improves classification fairness at test time (reducing up to 47% of between-group gaps).We also explore the idea of retraining sentiment classifiers on round-trip-translated data. Jonathan Gabel Christiansen, Mathias Gammelgaard, Anders Søgaard |
EMNLP (1) | 3 |
| 2021 | Dynamic Forecasting of Conversation DerailmentabstractOnline conversations can sometimes take a turn for the worse, either due to systematic cultural differences, accidental misunderstandings, or mere malice.Automatically forecasting derailment in public online conversations provides an opportunity to take early action to moderate it.Previous work in this space is limited, and we extend it in several ways.We apply a pretrained language encoder to the task, which outperforms earlier approaches.We further experiment with shifting the training paradigm for the task from a static to a dynamic one to increase the forecast horizon.This approach shows mixed results: in a highquality data setting, a longer average forecast horizon can be achieved at the cost of a small drop in F1; in a low-quality data setting, however, dynamic training propagates the noise and is highly detrimental to performance. Yova Kementchedjhieva, Anders Søgaard |
EMNLP (1) | 2 |
| 2021 | The Impact of Positional Encodings on Multilingual CompressionabstractIn order to preserve word-order information in a non-autoregressive setting, transformer architectures tend to include positional knowledge, by (for instance) adding positional encodings to token embeddings.Several modifications have been proposed over the sinusoidal positional encodings used in the original transformer architecture; these include, for instance, separating position encodings and token embeddings, or directly modifying attention weights based on the distance between word pairs.We first show that surprisingly, while these modifications tend to improve monolingual language models, none of them result in better multilingual language models.We then answer why that is: Sinusoidal encodings were explicitly designed to facilitate compositionality by allowing linear projections over arbitrary time steps.Higher variances in multilingual training distributions requires higher compression, in which case, compositionality becomes indispensable.Learned absolute positional encodings (e.g., in mBERT) tend to approximate sinusoidal embeddings in multilingual settings, but more complex positional encoding architectures lack the inductive bias to effectively learn compositionality and cross-lingual alignment.In other words, while sinusoidal positional encodings were originally designed for monolingual applications, they are particularly useful in multilingual language models. Vinit Ravishankar, Anders Søgaard |
EMNLP (1) | 2 |
| 2021 | Locke's Holiday: Belief Bias in Machine ReadingabstractI highlight a simple failure mode of state-ofthe-art machine reading systems: when contexts do not align with commonly shared beliefs.For example, machine reading systems fail to answer What did Elizabeth want?correctly in the context of 'My kingdom for a cough drop, cried Queen Elizabeth.'Biased by co-occurrence statistics in the training data of pretrained language models, systems predict my kingdom, rather than a cough drop.I argue such biases are analogous to human belief biases and present a carefully designed challenge dataset for English machine reading, called AUTO-LOCKE, to quantify such effects.Evaluations of machine reading systems on AUTO-LOCKE show the pervasiveness of belief bias in machine reading. *Each author wishes the others had contributed more. ContextIndonesia is the Germany of the Asean.So then, Malaysia is the France. QuestionWhat country is Indonesia similar to?Answer Germany Prediction Malaysia Anders Søgaard |
EMNLP (1) | 1 |
| 2020 | What Do You Mean 'Why?': Resolving Sluices in ConversationsabstractIn conversation, we often ask one-word questions such as ‘Why?’ or ‘Who?’. Such questions are typically easy for humans to answer, but can be hard for computers, because their resolution requires retrieving both the right semantic frames and the right arguments from context. This paper introduces the novel ellipsis resolution task of resolving such one-word questions, referred to as sluices in linguistics. We present a crowd-sourced dataset containing annotations of sluices from over 4,000 dialogues collected from conversational QA datasets, as well as a series of strong baseline architectures. Victor Petrén Bach Hansen, Anders Søgaard |
AAAI | 2 |
| 2020 | Weakly Supervised POS Taggers Perform Poorly on Truly Low-Resource LanguagesabstractPart-of-speech (POS) taggers for low-resource languages which are exclusively based on various forms of weak supervision – e.g., cross-lingual transfer, type-level supervision, or a combination thereof – have been reported to perform almost as well as supervised ones. However, weakly supervised POS taggers are commonly only evaluated on languages that are very different from truly low-resource languages, and the taggers use sources of information, like high-coverage and almost error-free dictionaries, which are likely not available for resource-poor languages. We train and evaluate state-of-the-art weakly supervised POS taggers for a typologically diverse set of 15 truly low-resource languages. On these languages, given a realistic amount of resources, even our best model gets only less than half of the words right. Our results highlight the need for new and different approaches to POS tagging for truly low-resource languages. Katharina Kann, Ophélie Lacroix, Anders Søgaard |
AAAI | 3 |
| 2020 | Parsing as PretrainingabstractRecent analyses suggest that encoders pretrained for language modeling capture certain morpho-syntactic structure. However, probing frameworks for word vectors still do not report results on standard setups such as constituent and dependency parsing. This paper addresses this problem and does full parsing (on English) relying only on pretraining architectures – and no decoding. We first cast constituent and dependency parsing as sequence tagging. We then use a single feed-forward layer to directly map word vectors to labels that encode a linearized tree. This is used to: (i) see how far we can reach on syntax modelling with just pretrained encoders, and (ii) shed some light about the syntax-sensitivity of different word vectors (by freezing the weights of the pretraining network during training). For evaluation, we use bracketing F1-score and las, and analyze in-depth differences across representations for span lengths and dependency displacements. The overall results surpass existing sequence tagging parsers on the ptb (93.5%) and end-to-end en-ewt ud (78.8%). David Vilares 0001, Michalina Strzyz, Anders Søgaard, Carlos Gómez-Rodríguez |
AAAI | 3 |
| 2020 | The Sensitivity of Language Models and Humans to Winograd Schema PerturbationsabstractLarge-scale pretrained language models are the major driving force behind recent improvements in performance on the Winograd Schema Challenge, a widely employed test of commonsense reasoning ability.We show, however, with a new diagnostic dataset, that these models are sensitive to linguistic perturbations of the Winograd examples that minimally affect human understanding.Our results highlight interesting differences between humans and language models: language models are more sensitive to number or gender alternations and synonym replacements than humans, and humans are more stable and consistent in their predictions, maintain a much higher absolute performance, and perform better on non-associative instances than associative ones.Overall, humans are correct more often than out-of-the-box models, and the models are sometimes right for the wrong reasons.Finally, we show that fine-tuning on a large, task-specific dataset can offer a solution to these issues. Mostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov, Desmond Elliott, Anders Søgaard |
ACL | 6 |
| 2020 | Grammatical Error Correction in Low Error Density Domains: A New Benchmark and AnalysesabstractEvaluation of grammatical error correction (GEC) systems has primarily focused on essays written by non-native learners of English, which however is only part of the full spectrum of GEC applications.We aim to broaden the target domain of GEC and release CWEB, a new benchmark for GEC consisting of website text generated by English speakers of varying levels of proficiency.Website data is a common and important domain that contains far fewer grammatical errors than learner essays, which we show presents a challenge to stateof-the-art GEC systems.We demonstrate that a factor behind this is the inability of systems to rely on a strong internal language model in low error density domains.We hope this work shall facilitate the development of opendomain GEC models that generalize to different topics and genres. Simon Flachs, Ophélie Lacroix, Helen Yannakoudakis, Marek Rei, Anders Søgaard |
EMNLP (1) | 5 |
| 2020 | Type B Reflexivization as an Unambiguous Testbed for Multilingual Multi-Task Gender BiasabstractThe one-sided focus on English in previous studies of gender bias in NLP misses out on opportunities in other languages: English challenge datasets such as GAP and Wino-Gender highlight model preferences that are "hallucinatory", e.g., disambiguating genderambiguous occurrences of 'doctor' as male doctors.We show that for languages with type B reflexivization, e.g., Swedish and Russian, we can construct multi-task challenge datasets for detecting gender bias that lead to unambiguously wrong model predictions: In these languages, the direct translation of 'the doctor removed his mask' is not ambiguous between a coreferential reading and a disjoint reading.Instead, the coreferential reading requires a non-gendered pronoun, and the gendered, possessive pronouns are anti-reflexive.We present a multilingual, multi-task challenge dataset, which spans four languages and four NLP tasks and focuses only on this phenomenon.We find evidence for gender bias across all task-language combinations and correlate model bias with national labor market statistics. Ana Valeria González-Garduño, Maria Barrett, Rasmus Hvingelby, Kellie Webster, Anders Søgaard |
EMNLP (1) | 5 |
| 2020 | Some Languages Seem Easier to Parse Because Their Treebanks LeakabstractCross-language differences in dependency parsing performance are mostly attributed to treebank size, average sentence length, average dependency length, morphological complexity, and domain differences.In this paper I point to a factor not previously discussed: If we abstract away from words and dependency labels, how many graphs in the test data were seen in the training data?I discuss how to compute graph isomorphisms, and show that, treebank size aside, overlap between training and test graphs explains more of the observed variation than standard explanations such as the above. Anders Søgaard |
EMNLP (1) | 1 |
| 2020 | Are All Good Word Vector Spaces Isomorphic?abstractExisting algorithms for aligning cross-lingual word vector spaces assume that vector spaces are approximately isomorphic.As a result, they perform poorly or fail completely on nonisomorphic spaces.Such non-isomorphism has been hypothesised to result from typological differences between languages.In this work, we ask whether non-isomorphism is also crucially a sign of degenerate word vector spaces.We present a series of experiments across diverse languages which show that variance in performance across language pairs is not only due to typological differences, but can mostly be attributed to the size of the monolingual resources available, and to the properties and duration of monolingual training (e.g."under-training"). Ivan Vulic, Sebastian Ruder, Anders Søgaard |
EMNLP (1) | 3 |
| 2020 | Do End-to-End Speech Recognition Models Care About Context?abstractThe two most common paradigms for end-to-end speech recognition are connectionist temporal classification (CTC) and attention-based encoder-decoder (AED) models. It has been argued that the latter is better suited for learning an implicit language model. We test this hypothesis by measuring temporal context sensitivity and evaluate how the models perform when we constrain the amount of contextual information in the audio input. We find that the AED model is indeed more context sensitive, but that the gap can be closed by adding self-attention to the CTC model. Furthermore, the two models perform similarly when contextual information is constrained. Finally, in contrast to previous research, our results show that the CTC model is highly competitive on WSJ and LibriSpeech without the help of an external language model. Lasse Borgholt, Jakob D. Havtorn, Zeljko Agic, Anders Søgaard, Lars Maaløe, Christian Igel |
INTERSPEECH | 4 |
| 2020 | Model-based Annotation of CoreferenceabstractHumans do not make inferences over texts, but over models of what texts are about. When annotators are asked to annotate coreferent spans of text, it is therefore a somewhat unnatural task. This paper presents an alternative in which we preprocess documents, linking entities to a knowledge base, and turn the coreference annotation task – in our case limited to pronouns – into an annotation task where annotators are asked to assign pronouns to entities. Model-based annotation is shown to lead to faster annotation and higher inter-annotator agreement, and we argue that it also opens up an alternative approach to coreference resolution. We present two new coreference benchmark datasets, for English Wikipedia and English teacher-student dialogues, and evaluate state-of-the-art coreference resolvers on them. Rahul Aralikatte, Anders Søgaard |
LREC | 2 |
| 2020 | DaNE: A Named Entity Resource for DanishabstractWe present a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme: DaNE. It is the largest publicly available, Danish named entity gold annotation. We evaluate the quality of our annotations intrinsically by double annotating the entire treebank and extrinsically by comparing our annotations to a recently released named entity annotation of the validation and test sections of the Danish Universal Dependencies treebank. We benchmark the new resource by training and evaluating competitive architectures for supervised named entity recognition (NER), including FLAIR, monolingual (Danish) BERT and multilingual BERT. We explore cross-lingual transfer in multilingual BERT from five related languages in zero-shot and direct transfer setups, and we show that even with our modestly-sized training set, we improve Danish NER over a recent cross-lingual approach, as well as over zero-shot transfer from five related languages. Using multilingual BERT, we achieve higher performance by fine-tuning on both DaNE and a larger Bokmål (Norwegian) training set compared to only using DaNE. However, the highest performance isachieved by using a Danish BERT fine-tuned on DaNE. Our dataset enables improvements and applicability for Danish NER beyond cross-lingual methods. We employ a thorough error analysis of the predictions of the best models for seen and unseen entities, as well as their robustness on un-capitalized text. The annotated dataset and all the trained models are made publicly available. Rasmus Hvingelby, Amalie Brogaard Pauli, Maria Barrett, Christina Rosted, Lasse Malm Lidegaard, Anders Søgaard |
LREC | 6 |
| 2020 | WikiBank: Using Wikidata to Improve Multilingual Frame-Semantic ParsingabstractFrame-semantic annotations exist for a tiny fraction of the world’s languages, Wikidata, however, links knowledge base triples to texts in many languages, providing a common, distant supervision signal for semantic parsers. We present WikiBank, a multilingual resource of partial semantic structures that can be used to extend pre-existing resources rather than creating new man-made resources from scratch. We also integrate this form of supervision into an off-the-shelf frame-semantic parser and allow cross-lingual transfer. Using Google’s Sling architecture, we show significant improvements on the English and Spanish CoNLL 2009 datasets, whether training on the full available datasets or small subsamples thereof. Cezar Sas, Meriem Beloucif, Anders Søgaard |
LREC | 3 |
| 2019 | Predicting Concrete and Abstract Entities in Modern Poetry
Fiammetta Caccavale, Anders Søgaard |
AAAI | 2 |
| 2019 | Jointly Learning to Label Sentences and TokensabstractLearning to construct text representations in end-to-end systems can be difficult, as natural languages are highly compositional and task-specific annotated datasets are often limited in size. Methods for directly supervising language composition can allow us to guide the models based on existing knowledge, regularizing them towards more robust and interpretable representations. In this paper, we investigate how objectives at different granularities can be used to learn better language representations and we propose an architecture for jointly learning to label sentences and tokens. The predictions at each level are combined together using an attention mechanism, with token-level labels also acting as explicit supervision for composing sentence-level representations. Our experiments show that by learning to perform these tasks jointly on multiple levels, the model achieves substantial improvements for both sentence classification and sequence labeling. Marek Rei, Anders Søgaard |
AAAI | 2 |
| 2019 | Latent Multi-Task Architecture LearningabstractMulti-task learning (MTL) allows deep neural networks to learn from related tasks by sharing parameters with other networks. In practice, however, MTL involves searching an enormous space of possible parameter sharing architectures to find (a) the layers or subspaces that benefit from sharing, (b) the appropriate amount of sharing, and (c) the appropriate relative weights of the different task losses. Recent work has addressed each of the above problems in isolation. In this work we present an approach that learns a latent multi-task architecture that jointly addresses (a)–(c). We present experiments on synthetic data and data from OntoNotes 5.0, including four different tasks and seven different domains. Our extension consistently outperforms previous approaches to learning latent architectures for multi-task problems and achieves up to 15% average error reductions over common approaches to MTL. Sebastian Ruder, Joachim Bingel, Isabelle Augenstein, Anders Søgaard |
AAAI | 4 |
| 2019 | Historical Text Normalization with Delayed RewardsabstractTraining neural sequence-to-sequence models with simple token-level log-likelihood is now a standard approach to historical text normalization, albeit often outperformed by phrasebased models.Policy gradient training enables direct optimization for exact matches, and while the small datasets in historical text normalization are prohibitive of from-scratch reinforcement learning, we show that policy gradient fine-tuning leads to significant improvements across the board.Policy gradient training, in particular, leads to more accurate normalizations for long or unseen words. Simon Flachs, Marcel Bollmann, Anders Søgaard |
ACL (1) | 3 |
| 2019 | Multi-Task Semantic Dependency Parsing with Policy Gradient for Learning Easy-First StrategiesabstractIn Semantic Dependency Parsing (SDP), semantic relations form directed acyclic graphs, rather than trees.We propose a new iterative predicate selection (IPS) algorithm for SDP.Our IPS algorithm combines the graph-based and transition-based parsing approaches in order to handle multiple semantic head words.We train the IPS model using a combination of multi-task learning and task-specific policy gradient training.Trained this way, IPS achieves a new state of the art on the SemEval 2015 Task 18 datasets.Furthermore, we observe that policy gradient training learns an easy-first strategy. Shuhei Kurita, Anders Søgaard |
ACL (1) | 2 |
| 2019 | Higher-order Comparisons of Sentence Encoder RepresentationsabstractMostafa Abdou, Artur Kulmizev, Felix Hill, Daniel M. Low, Anders Søgaard. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mostafa Abdou, Artur Kulmizev, Felix Hill, Daniel M. Low, Anders Søgaard |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Rewarding Coreference Resolvers for Being Consistent with World KnowledgeabstractRahul Aralikatte, Heather Lent, Ana Valeria Gonzalez, Daniel Herschcovich, Chen Qiu, Anders Sandholm, Michael Ringaard, Anders Søgaard. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Rahul Aralikatte, Heather C. Lent, Ana Valeria González-Garduño, Daniel Hershcovich, Chen Qiu 0005, Anders Sandholm 0001, Michael Ringaard, Anders Søgaard |
EMNLP/IJCNLP (1) | 8 |
| 2019 | Adversarial Removal of Demographic Attributes RevisitedabstractMaria Barrett, Yova Kementchedjhieva, Yanai Elazar, Desmond Elliott, Anders Søgaard. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Maria Barrett, Yova Kementchedjhieva, Yanai Elazar, Desmond Elliott, Anders Søgaard |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Lost in Evaluation: Misleading Benchmarks for Bilingual Dictionary InductionabstractYova Kementchedjhieva, Mareike Hartmann, Anders Søgaard. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yova Kementchedjhieva, Mareike Hartmann, Anders Søgaard |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A systematic comparison of methods for low-resource dependency parsing on genuinely low-resource languagesabstractClara Vania, Yova Kementchedjhieva, Anders Søgaard, Adam Lopez. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Clara Vania, Yova Kementchedjhieva, Anders Søgaard, Adam Lopez |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Comparing Unsupervised Word Translation Methods Step by StepabstractCross-lingual word vector space alignment is the task of mapping the vocabularies of two languages into a shared semantic space, which can be used for dictionary induction, unsupervised machine translation, and transfer learning. In the unsupervised regime, an initial seed dictionary is learned in the absence of any known correspondences between words, through {\bf distribution matching}, and the seed dictionary is then used to supervise the induction of the final alignment in what is typically referred to as a (possibly iterative) {\bf refinement} step. We focus on the first step and compare distribution matching techniques in the context of language pairs for which mixed training stability and evaluation scores have been reported. We show that, surprisingly, when looking at this initial step in isolation, vanilla GANs are superior to more recent methods, both in terms of precision and robustness. The improvements reported by more recent methods thus stem from the refinement techniques, and we show that we can obtain state-of-the-art performance combining vanilla GANs with such refinement techniques. Mareike Hartmann, Yova Kementchedjhieva, Anders Søgaard |
NeurIPS | 3 |
| 2019 | A Survey of Cross-lingual Word Embedding ModelsabstractCross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we provide a comprehensive typology of cross-lingual word embedding models. We compare their data requirements and objective functions. The recurring theme of the survey is that many of the models presented in the literature optimize for the same objectives, and that seemingly different models are often equivalent, modulo optimization strategies, hyper-parameters, and such. We also discuss the different ways cross-lingual word embeddings are evaluated, as well as future challenges and research horizons. Sebastian Ruder, Ivan Vulic, Anders Søgaard |
J. Artif. Intell. Res. | 3 |
| 2018 | Learning to Predict Readability Using Eye-Movement Data From Natives and LearnersabstractReadability assessment can improve the quality of assisting technologies aimed at language learners. Eye-tracking data has been used for both inducing and evaluating general-purpose NLP/AI models, and below we show that unsurprisingly, gaze data from language learners can also improve multi-task readability assessment models. This is unsurprising, since the gaze data records the reading difficulties ofthe learners. Unfortunately, eye-tracking data from language learners is often much harder to obtain than eye-tracking data from native speakers. We therefore compare the performance of deep learning readability models that use nativespeaker eye movement data to models using data from language learners. Somewhat surprisingly, we observe no significant drop in performance when replacing learners with natives, making approaches that rely on native speaker gaze information, more scalable. In other words, our finding is that language learner difficulties can be efficiently estimated from native speakers, which suggests that, more generally, readily available gaze data can be used to improve educational NLP/AI models targeted towards language learners. Ana Valeria González-Garduño, Anders Søgaard |
AAAI | 2 |
| 2018 | On the Limitations of Unsupervised Bilingual Dictionary InductionabstractUnsupervised machine translation-i.e., not assuming any cross-lingual supervision signal, whether a dictionary, translations, or comparable corpora-seems impossible, but nevertheless, Lample et al. (2018a) recently proposed a fully unsupervised machine translation (MT) model.The model relies heavily on an adversarial, unsupervised alignment of word embedding spaces for bilingual dictionary induction (Conneau et al., 2018), which we examine here.Our results identify the limitations of current unsupervised MT: unsupervised bilingual dictionary induction performs much worse on morphologically rich languages that are not dependent marking, when monolingual corpora from different domains or different embedding algorithms are used.We show that a simple trick, exploiting a weak supervision signal from identical words, enables more robust induction, and establish a near-perfect correlation between unsupervised bilingual dictionary induction performance and a previously unexplored graph similarity metric. Anders Søgaard, Sebastian Ruder, Ivan Vulic |
ACL (1) | 1 |
| 2018 | Lexi: A tool for adaptive, personalized text simplificationabstractMost previous research in text simplification has aimed to develop generic solutions, assuming very homogeneous target audiences with consistent intra-group simplification needs. We argue that this assumption does not hold, and that instead we need to develop simplification systems that adapt to the individual needs of specific users. As a first step towards personalized simplification, we propose a framework for adaptive lexical simplification and introduce Lexi, a free open-source and easily extensible tool for adaptive, personalized text simplification. Lexi is easily installed as a browser extension, enabling easy access to the service for its users. Joachim Bingel, Gustavo Paetzold, Anders Søgaard |
COLING | 3 |
| 2018 | Sequence Classification with Human AttentionabstractLearning attention functions requires large volumes of data, but many NLP tasks simulate human behavior, and in this paper, we show that human attention really does provide a good inductive bias on many attention functions in NLP.Specifically, we use estimated human attention derived from eyetracking corpora to regularize attention functions in recurrent neural networks.We show substantial improvements across a range of tasks, including sentiment analysis, grammatical error detection, and detection of abusive language. Maria Barrett, Joachim Bingel, Nora Hollenstein, Marek Rei, Anders Søgaard |
CoNLL | 5 |
| 2018 | Generalizing Procrustes Analysis for Better Bilingual Dictionary InductionabstractMost recent approaches to bilingual dictionary induction find a linear alignment between the word vector spaces of two languages. We show that projecting the two languages onto a third, latent space, rather than directly onto each other, while equivalent in terms of expressivity, makes it easier to learn approximate alignments. Our modified approach also allows for supporting languages to be included in the alignment process, to obtain an even better performance in low resource settings. Yova Kementchedjhieva, Sebastian Ruder, Ryan Cotterell, Anders Søgaard |
CoNLL | 4 |
| 2018 | A strong baseline for question relevancy rankingabstractThe best systems at the SemEval-16 and SemEval-17 community question answering shared tasks -a task that amounts to question relevancy ranking -involve complex pipelines and manual feature engineering.Despite this, many of these still fail at beating the IR baseline, i.e., the rankings provided by Google's search engine.We present a strong baseline for question relevancy ranking by training a simple multi-task feed forward network on a bag of 14 distance measures for the input question pair.This baseline model, which is fast to train and uses only language-independent features, outperforms the best shared task systems on the task of retrieving relevant previously asked questions. Ana Valeria González-Garduño, Isabelle Augenstein, Anders Søgaard |
EMNLP | 3 |
| 2018 | Why is unsupervised alignment of English embeddings from different algorithms so hard?abstractThis paper presents a challenge to the community: Generative adversarial networks (GANs) can perfectly align independent English word embeddings induced using the same algorithm, based on distributional information alone; but fails to do so, for two different embeddings algorithms.Why is that?We believe understanding why, is key to understand both modern word embedding algorithms and the limitations and instability dynamics of GANs.This paper shows that (a) in all these cases, where alignment fails, there exists a linear transform between the two embeddings (so algorithm biases do not lead to non-linear differences), and (b) similar effects can not easily be obtained by varying hyper-parameters.One plausible suggestion based on our initial experiments is that the differences in the inductive biases of the embedding algorithms lead to an optimization landscape that is riddled with local optima, leading to a very small basin of convergence, but we present this more as a challenge paper than a technical contribution. Mareike Hartmann, Yova Kementchedjhieva, Anders Søgaard |
EMNLP | 3 |
| 2018 | Parameter sharing between dependency parsers for related languagesabstractPrevious work has suggested that parameter sharing between transition-based neural dependency parsers for related languages can lead to better performance, but there is no consensus on what parameters to share.We present an evaluation of 27 different parameter sharing strategies across 10 languages, representing five pairs of related languages, each pair from a different language family.We find that sharing transition classifier parameters always helps, whereas the usefulness of sharing word and/or character LSTM parameters varies.Based on this result, we propose an architecture where the transition classifier is shared, and the sharing of word and character parameters is controlled by a parameter that can be tuned on validation data.This model is linguistically motivated and obtains significant improvements over a mono-lingually trained baseline.We also find that sharing transition classifier parameters helps when training a parser on unrelated language pairs, but we find that, in the case of unrelated languages, sharing too many parameters does not help. Miryam de Lhoneux, Johannes Bjerva, Isabelle Augenstein, Anders Søgaard |
EMNLP | 4 |
| 2018 | A Discriminative Latent-Variable Model for Bilingual Lexicon InductionabstractWe introduce a novel discriminative latentvariable model for the task of bilingual lexicon induction.Our model combines the bipartite matching dictionary prior of Haghighi et al. (2008) with a state-of-the-art embeddingbased approach.To train the model, we derive an efficient Viterbi EM algorithm.We provide empirical improvements on six language pairs under two metrics and show that the prior theoretically and empirically helps to mitigate the hubness problem.We also demonstrate how previous work may be viewed as a similarly fashioned latent-variable model, albeit with a different prior. 1 * The first two authors contributed equally. Sebastian Ruder, Ryan Cotterell, Yova Kementchedjhieva, Anders Søgaard |
EMNLP | 4 |
| 2018 | A Danish FrameNet Lexicon and an Annotated Corpus Used for Training and Evaluating a Semantic Frame Classifier
Bolette S. Pedersen, Sanni Nimb, Anders Søgaard, Mareike Hartmann, Sussi Olsen |
LREC | 3 |
| 2018 | Multi-Task Learning of Pairwise Sequence Classification Tasks over Disparate Label SpacesabstractIsabelle Augenstein, Sebastian Ruder, Anders Søgaard. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Isabelle Augenstein, Sebastian Ruder, Anders Søgaard |
NAACL-HLT | 3 |
| 2018 | Unsupervised Induction of Linguistic Categories with Records of Reading, Speaking, and WritingabstractMaria Barrett, Ana Valeria González-Garduño, Lea Frermann, Anders Søgaard. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Maria Barrett, Ana Valeria González-Garduño, Lea Frermann, Anders Søgaard |
NAACL-HLT | 4 |
| 2018 | Zero-Shot Sequence Labeling: Transferring Knowledge from Sentences to TokensabstractCan attention-or gradient-based visualization techniques be used to infer token-level labels for binary sequence tagging problems, using networks trained only on sentence-level labels?We construct a neural network architecture based on soft attention, train it as a binary sentence classifier and evaluate against tokenlevel annotation on four different datasets.Inferring token labels from a network provides a method for quantitatively evaluating what the model is learning, along with generating useful feedback in assistance systems.Our results indicate that attention-based methods are able to predict token-level labels more accurately, compared to gradient-based methods, sometimes even rivaling the supervised oracle network. Marek Rei, Anders Søgaard |
NAACL-HLT | 2 |
| 2017 | Learning attention for historical text normalization by learning to pronounceabstractAutomated processing of historical texts often relies on pre-normalization to modern word forms.Training encoder-decoder architectures to solve such problems typically requires a lot of training data, which is not available for the named task.We address this problem by using several novel encoder-decoder architectures, including a multi-task learning (MTL) architecture using a grapheme-to-phoneme dictionary as auxiliary data, pushing the state-of-theart by an absolute 2% increase in performance.We analyze the induced models across 44 different texts from Early New High German.Interestingly, we observe that, as previously conjectured, multi-task learning can learn to focus attention during decoding, in ways remarkably similar to recently proposed attention mechanisms.This, we believe, is an important step toward understanding how MTL works. Marcel Bollmann, Joachim Bingel, Anders Søgaard |
ACL (1) | 3 |
| 2017 | Cross-lingual RST Discourse ParsingabstractDiscourse parsing is an integral part of understanding information flow and argumentative structure in documents.Most previous research has focused on inducing and evaluating models from the English RST Discourse Treebank.However, discourse treebanks for other languages exist, including Spanish, German, Basque, Dutch and Brazilian Portuguese.The treebanks share the same underlying linguistic theory, but differ slightly in the way documents are annotated.In this paper, we present (a) a new discourse parser which is simpler, yet competitive (significantly better on 2/3 metrics) to state of the art for English, (b) a harmonization of discourse treebanks across languages, enabling us to present (c) what to the best of our knowledge are the first experiments on crosslingual discourse parsing. Chloé Braud, Maximin Coavoux, Anders Søgaard |
EACL (1) | 3 |
| 2017 | A Strong Baseline for Learning Cross-Lingual Word Embeddings from Sentence AlignmentsabstractWhile cross-lingual word embeddings have been studied extensively in recent years, the qualitative differences between the different algorithms remain vague.We observe that whether or not an algorithm uses a particular feature set (sentence IDs) accounts for a significant performance gap among these algorithms.This feature set is also used by traditional alignment algorithms, such as IBM Model-1, which demonstrate similar performance to stateof-the-art embedding algorithms on a variety of benchmarks.Overall, we observe that different algorithmic approaches for utilizing the sentence ID feature space result in similar performance.This paper draws both empirical and theoretical parallels between the embedding and alignment literature, and suggests that adding additional sources of information, which go beyond the traditional signal of bilingual sentence-aligned corpora, may substantially improve cross-lingual word embeddings, and that future baselines should at least take such features into account. Omer Levy, Anders Søgaard, Yoav Goldberg |
EACL (1) | 2 |
| 2017 | Parsing Universal Dependencies without trainingabstractHéctor Martínez Alonso, Željko Agić, Barbara Plank, Anders Søgaard. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Héctor Martínez Alonso, Zeljko Agic, Barbara Plank, Anders Søgaard |
EACL (1) | 4 |
| 2017 | Cross-Lingual Dependency Parsing with Late Decoding for Truly Low-Resource LanguagesabstractIn cross-lingual dependency annotation projection, information is often lost during transfer because of early decoding.We present an end-to-end graph-based neural network dependency parser that can be trained to reproduce matrices of edge scores, which can be directly projected across word alignments.We show that our approach to cross-lingual dependency parsing is not only simpler, but also achieves an absolute improvement of 2.25% averaged across 10 languages compared to the previous state of the art. Michael Sejr Schlichtkrull, Anders Søgaard |
EACL (1) | 2 |
| 2017 | Does syntax help discourse segmentation? Not so muchabstractDiscourse segmentation is the first step in building discourse parsers.Most work on discourse segmentation does not scale to real-world discourse parsing across languages, for two reasons: (i) models rely on constituent trees, and (ii) experiments have relied on gold standard identification of sentence and token boundaries.We therefore investigate to what extent constituents can be replaced with universal dependencies, or left out completely, as well as how state-of-the-art segmenters fare in the absence of sentence boundaries.Our results show that dependency information is less useful than expected, but we provide a fully scalable, robust model that only relies on part-of-speech information, and show that it performs well across languages in the absence of any gold-standard annotation. Chloé Braud, Ophélie Lacroix, Anders Søgaard |
EMNLP | 3 |
| 2017 | Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasmabstractNLP tasks are often limited by scarcity of manually annotated data. In social media sentiment analysis and related tasks, researchers have therefore used binarized emoticons and specific hashtags as forms of distant supervision. Our paper shows that by extending the distant supervision to a more diverse set of noisy labels, the models can learn richer representations. Through emoji prediction on a dataset of 1246 million tweets containing one of 64 common emojis we obtain state-of-the-art performance on 8 benchmark datasets within sentiment, emotion and sarcasm detection using a single pretrained model. Our analyses confirm that the diversity of our emotional labels yield a performance improvement over previous distant supervision approaches. Bjarke Felbo, Alan Mislove, Anders Søgaard, Iyad Rahwan, Sune Lehmann |
EMNLP | 3 |
| 2017 | Spikes as regularizers
Anders Søgaard |
ESANN | 1 |
| 2016 | Extracting token-level signals of syntactic processing from fMRI - with an application to PoS inductionabstractNeuro-imaging studies on reading different parts of speech (PoS) report somewhat mixed results, yet some of them indicate different activations with different PoS.This paper addresses the difficulty of using fMRI to discriminate between linguistic tokens in reading of running text because of low temporal resolution.We show that once we solve this problem, fMRI data contains a signal of PoS distinctions to the extent that it improves PoS induction with error reductions of more than 4%. Joachim Bingel, Maria Barrett, Anders Søgaard |
ACL (1) | 3 |
| 2016 | Cross-lingual Transfer of Correlations between Parts of Speech and Gaze FeaturesabstractSeveral recent studies have shown that eye movements during reading provide information about grammatical and syntactic processing, which can assist the induction of NLP models. All these studies have been limited to English, however. This study shows that gaze and part of speech (PoS) correlations largely transfer across English and French. This means that we can replicate previous studies on gaze-based PoS tagging for French, but also that we can use English gaze data to assist the induction of French NLP models. Maria Barrett, Frank Keller, Anders Søgaard |
COLING | 3 |
| 2016 | Improving historical spelling normalization with bi-directional LSTMs and multi-task learningabstractNatural-language processing of historical documents is complicated by the abundance of variant spellings and lack of annotated data. A common approach is to normalize the spelling of historical words to modern forms. We explore the suitability of a deep neural network architecture for this task, particularly a deep bi-LSTM network applied on a character level. Our model compares well to previously established normalization algorithms when evaluated on a diverse set of texts from Early New High German. We show that multi-task learning with additional normalization data can improve our model’s performance further. Marcel Bollmann, Anders Søgaard |
COLING | 2 |
| 2016 | Multi-view and multi-task training of RST discourse parsersabstractWe experiment with different ways of training LSTM networks to predict RST discourse trees. The main challenge for RST discourse parsing is the limited amounts of training data. We combat this by regularizing our models using task supervision from related tasks as well as alternative views on discourse structures. We show that a simple LSTM sequential discourse parser takes advantage of this multi-view and multi-task framework with 12-15% error reductions over our baseline (depending on the metric) and results that rival more complex state-of-the-art parsers. Chloé Braud, Barbara Plank, Anders Søgaard |
COLING | 3 |
| 2016 | The SemDaX Corpus ― Sense Annotations with Scalable Sense Inventories
Bolette S. Pedersen, Anna Braasch, Anders Johannsen, Héctor Martínez Alonso, Sanni Nimb, Sussi Olsen, Anders Søgaard, Nicolai Hartvig Sørensen |
LREC | 7 |
| 2016 | Learning a POS tagger for AAVE-like languageabstractPart-of-speech (POS) taggers trained on newswire perform much worse on domains such as subtitles, lyrics, or tweets.In addition, these domains are also heterogeneous, e.g., with respect to registers and dialects.In this paper, we consider the problem of learning a POS tagger for subtitles, lyrics, and tweets associated with African-American Vernacular English (AAVE).We learn from a mixture of randomly sampled and manually annotated Twitter data and unlabeled data, which we automatically and partially label using mined tag dictionaries.Our POS tagger obtains a tagging accuracy of 89% on subtitles, 85% on lyrics, and 83% on tweets, with up to 55% error reductions over a state-of-the-art newswire POS tagger, and 15-25% error reductions over a state-of-the-art Twitter POS tagger. Anna Jørgensen, Dirk Hovy, Anders Søgaard |
HLT-NAACL | 3 |
| 2016 | Improving sentence compression by learning to predict gazeabstractWe show how eye-tracking corpora can be used to improve sentence compression models, presenting a novel multi-task learning algorithm based on multi-layer LSTMs.We obtain performance competitive with or better than state-of-the-art approaches. Sigrid Klerke, Yoav Goldberg, Anders Søgaard |
HLT-NAACL | 3 |
| 2016 | Multilingual Projection for Parsing Truly Low-Resource LanguagesabstractWe propose a novel approach to cross-lingual part-of-speech tagging and dependency parsing for truly low-resource languages. Our annotation projection-based approach yields tagging and parsing models for over 100 languages. All that is needed are freely available parallel texts, and taggers and parsers for resource-rich languages. The empirical evaluation across 30 test languages shows that our method consistently provides top-level accuracies, close to established upper bounds, and outperforms several competitive baselines. Zeljko Agic, Anders Johannsen, Barbara Plank, Héctor Martínez Alonso, Natalie Schluter, Anders Søgaard |
Trans. Assoc. Comput. Linguistics | 6 |
| 2015 | Modeling Eye Movements when Reading MicroblogsabstractThis PhD project aims at a quantitative description of reading patterns from eye movements when reading tweets and the development of an eye movement relevance model. Maria Barrett, Anders Søgaard |
AAAI | 2 |
| 2015 | Using Frame Semantics for Knowledge Extraction from TwitterabstractKnowledge bases have the potential to advance artificial intelligence, but often suffer from recall problems, i.e., lack of knowledge of new entities and relations. On the contrary, social media such as Twitter provide abundance of data, in a timely manner: information spreads at an incredible pace and is posted long before it makes it into more commonly used resources for knowledge extraction. In this paper we address the question whether we can exploit social media to extract new facts, which may at first seem like finding needles in haystacks. We collect tweets about 60 entities in Freebase and compare four methods to extract binary relation candidates, based on syntactic and semantic parsing and simple mechanism for factuality scoring. The extracted facts are manually evaluated in terms of their correctness and relevance for search. We show that moving from bottom-up syntactic or semantic dependency parsing formalisms to top-down frame-semantic processing improves the robustness of knowledge extraction, producing more intelligible fact candidates of better quality. In order to evaluate the quality of frame semantic parsing on Twitter intrinsically, we make a multiply frame-annotated dataset of tweets publicly available. Anders Søgaard, Barbara Plank, Héctor Martínez Alonso |
AAAI | 1 |
| 2015 | Inverted indexing for cross-lingual NLPabstractAnders Søgaard, Željko Agić, Héctor Martínez Alonso, Barbara Plank, Bernd Bohnet, Anders Johannsen. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Anders Søgaard, Zeljko Agic, Héctor Martínez Alonso, Barbara Plank, Bernd Bohnet, Anders Johannsen |
ACL (1) | 1 |
| 2015 | Reading behavior predicts syntactic categoriesabstractIt is well-known that readers are less likely to fixate their gaze on closed class syntactic categories such as prepositions and pronouns.This paper investigates to what extent the syntactic category of a word in context can be predicted from gaze features obtained using eye-tracking equipment.If syntax can be reliably predicted from eye movements of readers, it can speed up linguistic annotation substantially, since reading is considerably faster than doing linguistic annotation by hand.Our results show that gaze features do discriminate between most pairs of syntactic categories, and we show how we can use this to annotate words with part of speech across domains, when tag dictionaries enable us to narrow down the set of potential categories. Maria Barrett, Anders Søgaard |
CoNLL | 2 |
| 2015 | Cross-lingual syntactic variation over age and genderabstractMost computational sociolinguistics studies have focused on phonological and lexical variation.We present the first large-scale study of syntactic variation among demographic groups (age and gender) across several languages.We harvest data from online user-review sites and parse it with universal dependencies.We show that several age and gender-specific variations hold across languages, for example that women are more likely to use VP conjunctions. Anders Johannsen, Dirk Hovy, Anders Søgaard |
CoNLL | 3 |
| 2015 | Do dependency parsing metrics correlate with human judgments?abstractUsing automatic measures such as labeled and unlabeled attachment scores is common practice in dependency parser evaluation.In this paper, we examine whether these measures correlate with human judgments of overall parse quality.We ask linguists with experience in dependency annotation to judge system outputs.We measure the correlation between their judgments and a range of parse evaluation metrics across five languages.The humanmetric correlation is lower for dependency parsing than for other NLP tasks.Also, inter-annotator agreement is sometimes higher than the agreement between judgments and metrics, indicating that the standard metrics fail to capture certain aspects of parse quality, such as the relevance of root attachment or the relative importance of the different parts of speech. Barbara Plank, Héctor Martínez Alonso, Zeljko Agic, Danijela Merkler, Anders Søgaard |
CoNLL | 5 |
| 2015 | Any-language frame-semantic parsingabstractWe present a multilingual corpus of Wikipedia and Twitter texts annotated with FRAMENET 1.5 semantic frames in nine different languages, as well as a novel technique for weakly supervised cross-lingual frame-semantic parsing. Our approach only assumes the existence of linked, comparable source and target lan-guage corpora (e.g., Wikipedia) and a bilingual dictionary (e.g., Wiktionary or BABELNET). Our approach uses a truly interlingual representation, enabling us to use the same model across all nine lan-guages. We present average error reduc-tions over running a state-of-the-art parser on word-to-word translations of 46 % for target identification, 37 % for frame identi-fication, and 14 % for argument identifica-tion. 1 Anders Johannsen, Héctor Martínez Alonso, Anders Søgaard |
EMNLP | 3 |
| 2015 | Learning to parse with IAA-weighted lossabstractHéctor Martínez Alonso, Barbara Plank, Arne Skjærholt, Anders Søgaard. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Héctor Martínez Alonso, Barbara Plank, Arne Skjærholt, Anders Søgaard |
HLT-NAACL | 4 |
| 2015 | Simple task-specific bilingual word embeddingsabstractWe introduce a simple wrapper method that uses off-the-shelf word embedding algorithms to learn task-specific bilingual word embeddings.We use a small dictionary of easily-obtainable task-specific word equivalence classes to produce mixed context-target pairs that we use to train off-the-shelf embedding models.Our model has the advantage that it (a) is independent of the choice of embedding algorithm, (b) does not require parallel data, and (c) can be adapted to specific tasks by re-defining the equivalence classes.We show how our method outperforms off-the-shelf bilingual embeddings on the task of unsupervised cross-language partof-speech (POS) tagging, as well as on the task of semi-supervised cross-language super sense (SuS) tagging. Stephan Gouws, Anders Søgaard |
HLT-NAACL | 2 |
| 2015 | Mining for unambiguous instances to adapt part-of-speech taggers to new domainsabstractDirk Hovy, Barbara Plank, Héctor Martínez Alonso, Anders Søgaard. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Dirk Hovy, Barbara Plank, Héctor Martínez Alonso, Anders Søgaard |
HLT-NAACL | 4 |
| 2015 | User Review Sites as a Resource for Large-Scale Sociolinguistic StudiesabstractSociolinguistic studies investigate the relation between language and extra-linguistic variables. This requires both representative text data and the associated socio-economic meta-data of the subjects. Traditionally, sociolinguistic studies use small samples of hand-curated data and meta-data. This can lead to exaggerated or false conclusions. Using social media data offers a large-scale source of language data, but usually lacks reliable socio-economic meta-data. Our research aims to remedy both problems by exploring a large new data source, international review websites with user profiles. They provide more text data than manually collected studies, and more meta-data than most available social media text. We describe the data and present various pilot studies, illustrating the usefulness of this resource for sociolinguistic studies. Our approach can help generate new research hypotheses based on data-driven findings across several countries and languages. Dirk Hovy, Anders Johannsen, Anders Søgaard |
WWW | 3 |
| 2014 | Adapting taggers to Twitter with not-so-distant supervision
Barbara Plank, Dirk Hovy, Ryan T. McDonald, Anders Søgaard |
COLING | 4 |
| 2014 | What's in a p-value in NLP?abstractIn NLP, we need to document that our pro-posed methods perform significantly bet-ter with respect to standard metrics than previous approaches, typically by re-porting p-values obtained by rank- or randomization-based tests. We show that significance results following current re-search standards are unreliable and, in ad-dition, very sensitive to sample size, co-variates such as sentence length, as well as to the existence of multiple metrics. We estimate that under the assumption of per-fect metrics and unbiased data, we need a significance cut-off at ⇠0.0025 to reduce the risk of false positive results to <5%. Since in practice we often have consider-able selection bias and poor metrics, this, however, will not do alone. 1 Anders Søgaard, Anders Johannsen, Barbara Plank, Dirk Hovy, Héctor Martínez Alonso |
CoNLL | 1 |
| 2014 | Learning part-of-speech taggers with inter-annotator agreement lossabstractIn natural language processing (NLP) an-notation projects, we use inter-annotator agreement measures and annotation guide-lines to ensure consistent annotations. However, annotation guidelines often make linguistically debatable and even somewhat arbitrary decisions, and inter-annotator agreement is often less than perfect. While annotation projects usu-ally specify how to deal with linguisti-cally debatable phenomena, annotator dis-agreements typically still stem from these “hard ” cases. This indicates that some er-rors are more debatable than others. In this paper, we use small samples of doubly-annotated part-of-speech (POS) data for Twitter to estimate annotation reliability and show how those metrics of likely inter-annotator agreement can be implemented in the loss functions of POS taggers. We find that these cost-sensitive algorithms perform better across annotation projects and, more surprisingly, even on data an-notated according to the same guidelines. Finally, we show that POS tagging mod-els sensitive to inter-annotator agreement perform better on the downstream task of chunking. 1 Barbara Plank, Dirk Hovy, Anders Søgaard |
EACL | 3 |
| 2014 | Importance weighting and unsupervised domain adaptation of POS taggers: a negative resultabstractImportance weighting is a generalization of various statistical bias correction techniques.While our labeled data in NLP is heavily biased, importance weighting has seen only few applications in NLP, most of them relying on a small amount of labeled target data.The publication bias toward reporting positive results makes it hard to say whether researchers have tried.This paper presents a negative result on unsupervised domain adaptation for POS tagging.In this setup, we only have unlabeled data and thus only indirect access to the bias in emission and transition probabilities.Moreover, most errors in POS tagging are due to unseen words, and there, importance weighting cannot help.We present experiments with a wide variety of weight functions, quantilizations, as well as with randomly generated weights, to support these claims. Barbara Plank, Anders Johannsen, Anders Søgaard |
EMNLP | 3 |
| 2014 | Crowdsourcing and annotating NER for Twitter #drift
Hege Fromreide, Dirk Hovy, Anders Søgaard |
LREC | 3 |
| 2014 | When POS data sets don't add up: Combatting sample bias
Dirk Hovy, Barbara Plank, Anders Søgaard |
LREC | 3 |
| 2013 | With Blinkers on: Robust Prediction of Eye Movements across ReadersabstractNilsson and Nivre (2009) introduced a treebased model of persons' eye movements in reading.The individual variation between readers reportedly made application across readers impossible.While a tree-based model seems plausible for eye movements, we show that competitive results can be obtained with a linear CRF model.Increasing the inductive bias also makes learning across readers possible.In fact we observe next-to-no performance drop when evaluating models trained on gaze records of multiple readers on new readers. Franz Matthies, Anders Søgaard |
EMNLP | 2 |
| 2013 | Using Crowdsourcing to get Representations based on Regular ExpressionsabstractOften the bottleneck in document classification is finding good representations that zoom in on the most important aspects of the documents.Most research uses n-gram representations, but relevant features often occur discontinuously, e.g., not. . .good in sentiment analysis.In this paper we present experiments getting experts to provide regular expressions, as well as crowdsourced annotation tasks from which regular expressions can be derived.Somewhat surprisingly, it turns out that these crowdsourced feature combinations outperform automatic feature combination methods, as well as expert features, by a very large margin and reduce error by 24-41% over n-gram representations. Anders Søgaard, Héctor Martínez Alonso, Jakob Elming, Anders Johannsen |
EMNLP | 1 |
| 2013 | Cross-Domain Answer Ranking using Importance Sampling
Anders Johannsen, Anders Søgaard |
IJCNLP | 2 |
| 2013 | Disambiguating Explicit Discourse Connectives without Oracles
Anders Johannsen, Anders Søgaard |
IJCNLP | 2 |
| 2013 | Down-stream effects of tree-to-dependency conversions
Jakob Elming, Anders Johannsen, Sigrid Klerke, Emanuele Lapponi, Héctor Martínez Alonso, Anders Søgaard |
HLT-NAACL | 6 |
| 2013 | Estimating effect size across datasets
Anders Søgaard |
HLT-NAACL | 1 |
| 2013 | Zipfian corruptions for robust POS tagging
Anders Søgaard |
HLT-NAACL | 1 |
| 2012 | DSim, a Danish Parallel Corpus for Text Simplification
Sigrid Klerke, Anders Søgaard |
LREC | 2 |
| 2012 | Unsupervised dependency parsing without trainingabstractAbstract Usually unsupervised dependency parsers try to optimize the probability of a corpus by revising the dependency model that is assumed to have generated the corpus. In this paper we explore a different view in which a dependency structure is, among other things, a partial order on the nodes in terms of centrality or saliency. Under this assumption we directly model centrality and derive dependency trees from the ordering of words. The result is an approach to unsupervised dependency parsing that is very different from standard ones in that it requires no training data. The input words are ordered by centrality, and a parse is derived from the ranking using a simple deterministic parsing algorithm, relying on the universal dependency rules defined by Naseem et al. (Naseem, T., Chen, H., Barzilay, R., Johnson, M. 2010. Using universal linguistic knowledge to guide grammar induction. In Proceedings of Empirical Methods in Natural Language Processing, Boston, MA, USA, pp. 1234–44.). Our approach is evaluated on data from twelve different languages and is remarkably competitive. Anders Søgaard |
Nat. Lang. Eng. | 1 |
| 2011 | A O(|G|n6) time extension of inversion transduction grammars
Anders Søgaard |
Mach. Transl. | 1 |
| 2010 | Semi-supervised dependency parsing using generalized tri-training
Anders Søgaard, Christian Rishøj |
COLING | 1 |
| 2010 | Can inversion transduction grammars generate hand alignments
Anders Søgaard |
EAMT | 1 |
| 2008 | Learning context-sensitive synchronous rules
Anders Søgaard |
EAMT | 1 |
| 2006 | Logical investigations on the adequacy of certain feature-based theories of natural language
Anders Søgaard |
HLT-NAACL | 1 |