EDBT 2026 Demo / reviewers in the wild / expert
Johannes Bjerva
dblp:148/4464
· DBLP profile ↗
29ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-9512-0739ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization FrameworkabstractLarge Language Models (LLMs) have been reported to "leak" Personally Identifiable Information (PII), with successful PII reconstruction often interpreted as evidence of memorization.We propose a principled revision of memorization evaluation for LLMs, arguing that PII leakage should be evaluated under low lexical cue conditions, where target PII cannot be reconstructed through prompt-induced generalization or pattern completion.We formalize Cue-Resistant Memorization (CRM) as a cue-controlled evaluation framework and a necessary condition for valid memorization evaluation, explicitly conditioning on prompt-target overlap cues.Using CRM, we conduct a largescale multilingual re-evaluation of PII leakage across 32 languages and multiple memorization paradigms.Revisiting reconstruction-based settings, including verbatim prefix-suffix completion and associative reconstruction, we find that their apparent effectiveness is driven primarily by direct surface-form cues rather than by true memorization.When such cues are controlled for, reconstruction success diminishes substantially.We further examine cue-free generation and membership inference, both of which exhibit extremely low true positive rates.Overall, our results suggest that previously reported PII leakage is better explained by cue-driven behavior than by genuine memorization, highlighting the importance of cue-controlled evaluation for reliably quantifying privacy-relevant memorization in LLMs 1 . Yiyi Chen 0002, Qiongxiu Li, Johannes Bjerva |
ACL (1) | 4 |
| 2026 | How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLPabstractKushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather Lent, Miryam de Lhoneux. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather C. Lent, Miryam de Lhoneux |
ACL (1) | 6 |
| 2026 | HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings FilingsabstractAccurate tagging of earnings reports can yield significant short-term returns for stakeholders. The machine-readable inline eXtensible Business Reporting Language (iXBRL) is mandated for public financial filings. Yet, its complex, fine-grained taxonomy limits the cross-company transferability of tagged Key Performance Indicators (KPIs). To address this, we introduce the Hierarchical Financial Key Performance Indicator (HiFi-KPI) dataset, a large-scale corpus of 1.65M paragraphs and 198k unique, hierarchically organized labels linked to iXBRL taxonomies. HiFi-KPI supports multiple tasks and we evaluate three: KPI classification, KPI extraction, and structured KPI extraction. For rapid evaluation, we also release HiFi-KPI-Lite, a manually curated 2.5K-instance subset. Baselines on HiFi-KPI-Lite show that encoder-based models achieve over 0.906 macro-F1 on classification, while Large Language Models (LLMs) reach 0.440 F1 on structured extraction. Finally, a qualitative analysis reveals that extraction errors primarily relate to dates. We open-source all code and data. Rasmus Aavang, Giovanni Rizzi, Rasmus Bøggild, Alexandre Iolov, Mike Zhang, Johannes Bjerva |
LREC | 6 |
| 2026 | SEFL: A Framework for Generating Synthetic Educational Assignment Feedback with LLM AgentsabstractProviding high-quality feedback on student assignments is crucial for student success, but it is heavily limited by time and budgetary constraints. In this work, we introduce Synthetic Educational Feedback Loops (SEFL), a synthetic data framework designed to generate data that resembles immediate, on-demand feedback at scale without relying on extensive, real-world student assignments and teacher feedback. To obtain this type of data, two large language models (LLMs) operate in a teacher-student role to simulate assignment completion and formative feedback, generating 19.8K synthetic pairs of student work and corresponding critiques and actionable improvements from a teacher. With this data, we fine-tune smaller, more computationally efficient LLMs on these synthetic pairs, enabling them to replicate key features of high-quality, goal-oriented feedback. Through comprehensive evaluations with three LLM judges and three human experts, across a subset of 900 outputs, we demonstrate that SEFL-tuned models outperform both their untuned counterparts and an existing baseline in terms of feedback quality. The potential for societal impact is reinforced by extensive qualitative comments and ratings from human stakeholders -- both students and higher education instructors. SEFL has the potential to transform feedback processes for higher education and beyond. Mike Zhang, Amalie Pernille Dilling, Léon Gondelman, Niels Erik Ruan Lyngdorf, Euan D. Lindsay, Johannes Bjerva |
LREC | 6 |
| 2026 | A Hierarchical Evaluation Framework for LLM-driven Threat Modelling ToolsabstractAI adoption has accelerated with the rise of LLMs, and people within the field of security are increasingly exploring their practical value. Threat modelling is central to secure system development, yet it remains largely manual and the value of LLM-driven tools is unclear. Even when LLMs prove useful, selecting the right one can be more challenging than using it. This paper introduces a systematic evaluation framework for LLM-driven threat modelling tools to support tool selection, observing the general LLM-integration, governance risks, and allowing for comparison of tool output. Using Goal-Question-Metric, we derive evaluation metrics and show the value of the framework on a set of state-of-the-art LLM-driven threat modelling tools. The results show our metrics distinguish both performance and governance risks, providing a basis for organisations to ensure automation strengthens rather than burdens their threat modelling process. Josephine Marie Bakka, Andreas Kjeldgaard Brandhøj, Tobias Worm Bøgedal, Johannes Bjerva |
SECRYPT (1) | 4 |
| 2025 | Against All Odds: Overcoming Typology, Script, and Language Confusion in Multilingual Embedding Inversion AttacksabstractLarge Language Models (LLMs) are susceptible to malicious influence by cyber attackers through intrusions such as adversarial, backdoor, and embedding inversion attacks. In response, the burgeoning field of LLM Security aims to study and defend against such threats. Thus far, the majority of works in this area have focused on monolingual English models; however, emerging research suggests that multilingual LLMs may be more vulnerable to various attacks than their monolingual counterparts. While previous work has investigated embedding inversion over a small subset of European languages, it is challenging to extrapolate these findings to languages from different linguistic families and with differing scripts. To this end, we explore the security of multilingual LLMs in the context of embedding inversion attacks and investigate cross-lingual and cross-script inversion across 20 languages, spanning over 8 language families and 12 scripts. Our findings indicate that languages written in Arabic and Cyrillic scripts are particularly vulnerable to embedding inversion, as are languages within the Indo-Aryan language family. We further observe that inversion models tend to suffer from language confusion, sometimes significantly reducing the efficacy of an attack. Accordingly, we systematically explore this bottleneck for inversion models, uncovering predictable patterns attackers could leverage. Ultimately, this study aims to further the field's understanding of the outstanding security vulnerabilities facing multilingual LLMs and raise awareness for the languages most at risk of negative impact from these attacks. Yiyi Chen 0002, Russa Biswas, Heather C. Lent, Johannes Bjerva |
AAAI | 4 |
| 2025 | ALGEN: Few-shot Inversion Attacks on Textual Embeddings via Cross-Model Alignment and GenerationabstractWith the growing popularity of Large Language Models (LLMs) and vector databases, private textual data is increasingly processed and stored as numerical embeddings. However, recent studies have proven that such embeddings are vulnerable to inversion attacks, where original text is reconstructed to reveal sensitive information. Previous research has largely assumed access to millions of sentences to train attack models, e.g., through data leakage or nearly unrestricted API access. With our method, a single data point is sufficient for a partially successful inversion attack. With as little as 1k data samples, performance reaches an optimum across a range of black-box encoders, without training on leaked data. We present a Few-shot Textual Embedding Inversion Attack using Cross-Model ALignment and GENeration (ALGEN), by aligning victim embeddings to the attack space and using a generative model to reconstruct text. We find that ALGEN attacks can be effectively transferred across domains and languages, revealing key information. We further examine a variety of defense mechanisms against ALGEN, and find that none are effective, highlighting the vulnerabilities posed by inversion attacks. By significantly lowering the cost of inversion and proving that embedding spaces can be aligned through one-step optimization, we establish a new textual embedding inversion paradigm with broader applications for embedding alignment in NLP. Yiyi Chen 0002, Qiongkai Xu, Johannes Bjerva |
ACL (1) | 3 |
| 2025 | How Do Hackathons Foster Creativity? Towards Automated Evaluation of Creativity at ScaleabstractHackathons have become popular collaborative events for accelerating the development of creative ideas and prototypes. There are several case studies showcasing creative outcomes across domains such as industry, education, and research. However, there are no large-scale studies on creativity in hackathons which can advance theory on how hackathon formats lead to creative outcomes. We conducted a computational analysis of 193,353 hackathon projects. By operationalizing creativity through usefulness and novelty, we refined our dataset to 10,363 projects, allowing us to analyze how participant characteristics, collaboration patterns, and hackathon setups influence the development of creative projects. The contribution of our paper is twofold: We identified means for organizers to foster creativity in hackathons. We also explore the use of large language models (LLMs) to augment the evaluation of creative outcomes and discuss challenges and opportunities of doing this, which has implications for creativity research at large. Jeanette Falk, Yiyi Chen 0002, Janet Rafner, Mike Zhang, Johannes Bjerva, Alexander Nolte |
CHI | 5 |
| 2025 | The Responsible Development of Automated Student Feedback with Generative AIabstractProviding rich, constructive feedback to students is essential for supporting and enhancing their learning. Recent advancements in Generative Artificial Intelligence (AI), particularly with large language models (LLMs), present new opportunities to deliver scalable, repeatable, and instant feedback, effectively making abundant a resource that has historically been scarce and costly. From a technical perspective, this approach is now feasible due to breakthroughs in AI and Natural Language Processing (NLP). While the potential educational benefits are compelling, implementing these technologies also introduces a host of ethical considerations that must be thoughtfully addressed. One of the core advantages of AI systems is their ability to automate routine and mundane tasks, potentially freeing up human educators for more nuanced work. However, the ease of automation risks a “tyranny of the majority”, where the diverse needs of minority or unique learners are overlooked, as they may be harder to systematize and less straightforward to accommodate. Ensuring inclusivity and equity in AI-generated feedback, therefore, becomes a critical aspect of responsible AI implementation in education. The process of developing machine learning models that produce valuable, personalized, and authentic feedback also requires significant input from human domain experts. Decisions around whose expertise is incorporated, how it is captured, and when it is applied have profound implications for the relevance and quality of the resulting feedback. Additionally, the maintenance and continuous refinement of these models are necessary to adapt feedback to evolving contextual, theoretical, and student-related factors. Without ongoing adaptation, feedback risks becoming obsolete or mismatched with the current needs of diverse student populations. Addressing these challenges is essential not only for ethical integrity but also for building the operational trust needed to integrate AI-driven systems as valuable tools in contemporary education. Thoughtful planning and deliberate choices are needed to ensure that these solutions truly benefit all students, allowing AI to support an inclusive and dynamic learning environment. Euan D. Lindsay, Mike Zhang, Aditya Johri, Johannes Bjerva |
EDUCON | 4 |
| 2025 | Shared Path: Unraveling Memorization in Multilingual LLMs through Language SimilaritiesabstractWe present the first comprehensive study of Memorization in Multilingual Large Language Models (MLLMs), analyzing 95 languages using models across diverse model scales, architectures, and memorization definitions.As MLLMs are increasingly deployed, understanding their memorization behavior has become critical.Yet prior work has focused primarily on monolingual models, leaving multilingual memorization underexplored, despite the inherently long-tailed nature of training corpora.We find that the prevailing assumption, that memorization is highly correlated with training data availability, fails to fully explain memorization patterns in MLLMs.We hypothesize that the conventional focus on monolingual settings, effectively treating languages in isolation, may obscure the true patterns of memorization.To address this, we propose a novel graph-based correlation metric that incorporates language similarity to analyze cross-lingual memorization.Our analysis reveals that among similar languages, those with fewer training tokens tend to exhibit higher memorization, a trend that only emerges when cross-lingual relationships are explicitly modeled.These findings underscore the importance of a languageaware perspective in evaluating and mitigating memorization vulnerabilities in MLLMs.This also constitutes empirical evidence that language similarity both explains Memorization in MLLMs and underpins Cross-lingual Transferability, with broad implications for multilingual NLP 1 . Yiyi Chen 0002, Johannes Bjerva, Qiongxiu Li |
EMNLP | 3 |
| 2025 | NLP Security and Ethics, in the WildabstractAbstract As NLP models are used by a growing number of end-users, an area of increasing importance is NLP Security (NLPSec): assessing the vulnerability of models to malicious attacks and developing comprehensive countermeasures against them. While work at the intersection of NLP and cybersecurity has the potential to create safer NLP for all, accidental oversights can result in tangible harm (e.g., breaches of privacy or proliferation of malicious models). In this emerging field, however, the research ethics of NLP have not yet faced many of the long-standing conundrums pertinent to cybersecurity, until now. We thus examine contemporary works across NLPSec, and explore their engagement with cybersecurity’s ethical norms. We identify trends across the literature, ultimately finding alarming gaps on topics like harm minimization and responsible disclosure. To alleviate these concerns, we provide concrete recommendations to help NLP researchers navigate this space more ethically, bridging the gap between traditional cybersecurity and NLP ethics, which we frame as “white hat NLP”. The goal of this work is to help cultivate an intentional culture of ethical research for those working in NLP Security. Heather C. Lent, Erick Galinkin, Yiyi Chen 0002, Jens Myrup Pedersen, Leon Derczynski, Johannes Bjerva |
Trans. Assoc. Comput. Linguistics | 6 |
| 2025 | Knowledge Graphs, Large Language Models, and Hallucinations: An NLP PerspectiveabstractLarge Language Models (LLMs) have revolutionized Natural Language Processing (NLP) based applications including automated text generation, question answering, chatbots, and others. However, they face a significant challenge: hallucinations, where models produce plausible-sounding but factually incorrect responses. This undermines trust and limits the applicability of LLMs in different domains. Knowledge Graphs (KGs), on the other hand, provide a structured collection of interconnected facts represented as entities (nodes) and their relationships (edges). In recent research, KGs have been leveraged to provide context that can fill gaps in an LLM’s understanding of certain topics offering a promising approach to mitigate hallucinations in LLMs, enhancing their reliability and accuracy while benefiting from their wide applicability. Nonetheless, it is still a very active area of research with various unresolved open problems. In this paper, we discuss these open challenges covering state-of-the-art datasets and benchmarks as well as methods for knowledge integration and evaluating hallucinations. In our discussion, we consider the current use of KGs in LLM systems and identify future directions within each of these challenges. Ernests Lavrinovics, Russa Biswas, Johannes Bjerva, Katja Hose |
J. Web Semant. | 3 |
| 2024 | Text Embedding Inversion Security for Multilingual Language ModelsabstractTextual data is often represented as realnumbered embeddings in NLP, particularly with the popularity of large language models (LLMs) and Embeddings as a Service (EaaS).However, storing sensitive information as embeddings can be susceptible to security breaches, as research shows that text can be reconstructed from embeddings, even without knowledge of the underlying model.While defence mechanisms have been explored, these are exclusively focused on English, leaving other languages potentially exposed to attacks.This work explores LLM security through multilingual embedding inversion.We define the problem of black-box multilingual and crosslingual inversion attacks, and explore their potential implications.Our findings suggest that multilingual LLMs may be more vulnerable to inversion attacks, in part because English-based defences may be ineffective.To alleviate this, we propose a simple masking defense effective for both monolingual and multilingual models.This study is the first to investigate multilingual inversion attacks, shedding light on the differences in attacks and defenses across monolingual and multilingual settings. Yiyi Chen 0002, Heather C. Lent, Johannes Bjerva |
ACL (1) | 3 |
| 2024 | What is "Typological Diversity" in NLP?abstractThe NLP research community has devoted increased attention to languages beyond English, resulting in considerable improvements for multilingual NLP.However, these improvements only apply to a small subset of the world's languages.An increasing number of papers aspires to enhance generalizable multilingual performance across languages.To this end, linguistic typology is commonly used to motivate language selection, on the basis that a broad typological sample ought to imply generalization across a broad range of languages.These selections are often described as being 'typologically diverse'.In this meta-analysis, we systematically investigate NLP research that includes claims regarding typological diversity.We find there are no set definitions or criteria for such claims.We introduce metrics to approximate the diversity of resulting language samples along several axes and find that the results vary considerably across papers.Crucially, we show that skewed language selection can lead to overestimated multilingual performance.We recommend future work to include an operationalization of typological diversity that empirically justifies the diversity of language samples.To help facilitate this, we release the code for our diversity measures.1 * Equal contribution. 1 Our code and data are publicly available: https://github.com/WPoelman/typ-div Esther Ploeger, Wessel Poelman, Miryam de Lhoneux, Johannes Bjerva |
EMNLP | 4 |
| 2024 | The Role of Typological Feature Prediction in NLP and LinguisticsabstractAbstract Computational typology has gained traction in the field of Natural Language Processing (NLP) in recent years, as evidenced by the increasing number of papers on the topic and the establishment of a Special Interest Group on the topic (SIGTYP), including the organization of successful workshops and shared tasks. A considerable amount of work in this sub-field is concerned with prediction of typological features, for example, for databases such as the World Atlas of Language Structures (WALS) or Grambank. Prediction is argued to be useful either because (1) it allows for obtaining feature values for relatively undocumented languages, alleviating the sparseness in WALS, in turn argued to be useful for both NLP and linguistics; and (2) it allows us to probe models to see whether or not these typological features are encapsulated in, for example, language representations. In this article, we present a critical stance concerning prediction of typological features, investigating to what extent this line of research is aligned with purported needs—both from the perspective of NLP practitioners, and perhaps more importantly, from the perspective of linguists specialized in typology and language documentation. We provide evidence that this line of research in its current state suffers from a lack of interdisciplinary alignment. Based on an extensive survey of the linguistic typology community, we present concrete recommendations for future research in order to improve this alignment between linguists and NLP researchers, beyond the scope of typological feature prediction. Johannes Bjerva |
Comput. Linguistics | 1 |
| 2024 | Recommending tasks based on search queries and missionsabstractAbstract Web search is an experience that naturally lends itself to recommendations, including query suggestions and related entities. In this article, we propose to recommend specific tasks to users, based on their search queries, such as planning a holiday trip or organizing a party. Specifically, we introduce the problem of query-based task recommendation and develop methods that combine well-established term-based ranking techniques with continuous semantic representations, including sentence representations from several transformer-based models. Using a purpose-built test collection, we find that our method is able to significantly outperform a strong text-based baseline. Further, we extend our approach to using a set of queries that all share the same underlying task, referred to as search mission, as input. The study is rounded off with a detailed feature and query analysis. Darío Garigliotti, Krisztian Balog, Katja Hose, Johannes Bjerva |
Nat. Lang. Eng. | 4 |
| 2024 | CreoleVal: Multilingual Multitask Benchmarks for CreolesabstractAbstract Creoles represent an under-explored and marginalized group of languages, with few available resources for NLP research. While the genealogical ties between Creoles and a number of highly resourced languages imply a significant potential for transfer learning, this potential is hampered due to this lack of annotated data. In this work we present CreoleVal, a collection of benchmark datasets spanning 8 different NLP tasks, covering up to 28 Creole languages; it is an aggregate of novel development datasets for reading comprehension relation classification, and machine translation for Creoles, in addition to a practical gateway to a handful of preexisting benchmarks. For each benchmark, we conduct baseline experiments in a zero-shot setting in order to further ascertain the capabilities and limitations of transfer learning for Creoles. Ultimately, we see CreoleVal as an opportunity to empower research on Creoles in NLP and computational linguistics, and in general, a step towards more equitable language technology around the globe. Heather C. Lent, Kushal Tatariya, Raj Dabre, Yiyi Chen 0002, Marcell Fekete, Esther Ploeger, Li Zhou 0010, Ruth-Ann Armstrong, Abee Eijansantos, Catriona Malau, Hans Erik Heje, Ernests Lavrinovics, Diptesh Kanojia, Paul Belony, Marcel Bollmann, Loïc Grobol, Miryam de Lhoneux, Daniel Hershcovich, Michel DeGraff, Anders Søgaard, Johannes Bjerva |
Trans. Assoc. Comput. Linguistics | 21 |
| 2022 | Quantifying Synthesis and Fusion and their Impact on Machine TranslationabstractArturo Oncevay, Duygu Ataman, Niels Van Berkel, Barry Haddow, Alexandra Birch, Johannes Bjerva. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Arturo Oncevay, Duygu Ataman, Niels van Berkel, Barry Haddow, Alexandra Birch, Johannes Bjerva |
NAACL-HLT | 6 |
| 2021 | Does Typological Blinding Impede Cross-Lingual Sharing?abstractBridging the performance gap between highand low-resource languages has been the focus of much previous work.Typological features from databases such as the World Atlas of Language Structures (WALS) are a prime candidate for this, as such data exists even for very low-resource languages.However, previous work has only found minor benefits from using typological information.Our hypothesis is that a model trained in a cross-lingual setting will pick up on typological cues from the input data, thus overshadowing the utility of explicitly using such features.We verify this hypothesis by blinding a model to typological information, and investigate how cross-lingual sharing and performance is impacted.Our model is based on a cross-lingual architecture in which the latent weights governing the sharing between languages is learnt during training.We show that (i) preventing this model from exploiting typology severely reduces performance, while a control experiment reaffirms that (ii) encouraging sharing according to typology somewhat improves performance. Johannes Bjerva, Isabelle Augenstein |
EACL | 1 |
| 2020 | Back to the Future - Temporal Adaptation of Text RepresentationsabstractLanguage evolves over time in many ways relevant to natural language processing tasks. For example, recent occurrences of tokens 'BERT' and 'ELMO' in publications refer to neural network architectures rather than persons. This type of temporal signal is typically overlooked, but is important if one aims to deploy a machine learning model over an extended period of time. In particular, language evolution causes data drift between time-steps in sequential decision-making tasks. Examples of such tasks include prediction of paper acceptance for yearly conferences (regular intervals) or author stance prediction for rumours on Twitter (irregular intervals). Inspired by successes in computer vision, we tackle data drift by sequentially aligning learned representations. We evaluate on three challenging tasks varying in terms of time-scales, linguistic units, and domains. These tasks show our method outperforming several strong baselines, including using all available data. We argue that, due to its low computational expense, sequential alignment is a practical solution to dealing with language evolution. Johannes Bjerva, Wouter M. Kouw, Isabelle Augenstein |
AAAI | 1 |
| 2020 | SubjQA: A Dataset for Subjectivity and Review ComprehensionabstractSubjectivity is the expression of internal opinions or beliefs which cannot be objectively observed or verified, and has been shown to be important for sentiment analysis and wordsense disambiguation.Furthermore, subjectivity is an important aspect of user-generated data.In spite of this, subjectivity has not been investigated in contexts where such data is widespread, such as in question answering (QA).We develop a new dataset which allows us to investigate this relationship.We find that subjectivity is an important feature in the case of QA, albeit with more intricate interactions between subjectivity and QA performance than found in previous work on sentiment analysis.For instance, a subjective question may or may not be associated with a subjective answer.We release an English QA dataset (SUBJQA) based on customer reviews, containing subjectivity annotations for questions and answer spans across 6 domains. Johannes Bjerva, Nikita Bhutani, Behzad Golshan, Wang Chiew Tan, Isabelle Augenstein |
EMNLP (1) | 1 |
| 2020 | Zero-Shot Cross-Lingual Transfer with Meta LearningabstractLearning what to share between tasks has become a topic of great importance, as strategic sharing of knowledge has been shown to improve downstream task performance.This is particularly important for multilingual applications, as most languages in the world are under-resourced.Here, we consider the setting of training models on multiple different languages at the same time, when little or no data is available for languages other than English.We show that this challenging setup can be approached using meta-learning: in addition to training a source language model, another model learns to select which training instances are the most beneficial to the first.We experiment using standard supervised, zero-shot cross-lingual, as well as fewshot cross-lingual settings for different natural language understanding tasks (natural language inference, question answering).Our extensive experimental setup demonstrates the consistent effectiveness of meta-learning for a total of 15 languages.We improve upon the state-of-the-art for zero-shot and few-shot NLI (on MultiNLI and XNLI) and QA (on the MLQA dataset).A comprehensive error analysis indicates that the correlation of typological features between languages can partly explain when parameter sharing learned via meta-learning is beneficial. Farhad Nooralahzadeh, Giannis Bekoulis, Johannes Bjerva, Isabelle Augenstein |
EMNLP (1) | 3 |
| 2019 | Uncovering Probabilistic Implications in Typological Knowledge BasesabstractThe study of linguistic typology is rooted in the implications we find between linguistic features, such as the fact that languages with object-verb word ordering tend to have postpositions. Uncovering such implications typically amounts to time-consuming manual processing by trained and experienced linguists, which potentially leaves key linguistic universals unexplored. In this paper, we present a computational model which successfully identifies known universals, including Greenberg universals, but also uncovers new ones, worthy of further linguistic investigation. Our approach outperforms baselines previously used for this problem, as well as a strong baseline from knowledge base population. Johannes Bjerva, Yova Kementchedjhieva, Ryan Cotterell, Isabelle Augenstein |
ACL (1) | 1 |
| 2019 | What Do Language Representations Really Represent?abstractA neural language model trained on a text corpus can be used to induce distributed representations of words, such that similar words end up with similar representations. If the corpus is multilingual, the same model can be used to learn distributed representations of languages, such that similar languages end up with similar representations. We show that this holds even when the multilingual corpus has been translated into English, by picking up the faint signal left by the source languages. However, just as it is a thorny problem to separate semantic from syntactic similarity in word representations, it is not obvious what type of similarity is captured by language representations. We investigate correlations and causal relationships between language representations learned from translations on one hand, and genetic, geographical, and several levels of structural similarity between languages on the other. Of these, structural similarity is found to correlate most strongly with language representation similarity, whereas genetic relationships—a convenient benchmark used for evaluation in previous work—appears to be a confounding factor. Apart from implications about translation effects, we see this more generally as a case where NLP and linguistic typology can interact and benefit one another. Johannes Bjerva, Robert Östling, Maria Han Veiga, Jörg Tiedemann, Isabelle Augenstein |
Comput. Linguistics | 1 |
| 2018 | Parameter sharing between dependency parsers for related languagesabstractPrevious work has suggested that parameter sharing between transition-based neural dependency parsers for related languages can lead to better performance, but there is no consensus on what parameters to share.We present an evaluation of 27 different parameter sharing strategies across 10 languages, representing five pairs of related languages, each pair from a different language family.We find that sharing transition classifier parameters always helps, whereas the usefulness of sharing word and/or character LSTM parameters varies.Based on this result, we propose an architecture where the transition classifier is shared, and the sharing of word and character parameters is controlled by a parameter that can be tuned on validation data.This model is linguistically motivated and obtains significant improvements over a mono-lingually trained baseline.We also find that sharing transition classifier parameters helps when training a parser on unrelated language pairs, but we find that, in the case of unrelated languages, sharing too many parameters does not help. Miryam de Lhoneux, Johannes Bjerva, Isabelle Augenstein, Anders Søgaard |
EMNLP | 2 |
| 2018 | From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language EmbeddingsabstractJohannes Bjerva, Isabelle Augenstein. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Johannes Bjerva, Isabelle Augenstein |
NAACL-HLT | 1 |
| 2017 | Articulation Rate in Swedish Child-Directed Speech Increases as a Function of the Age of the Child Even When Surprisal is Controlled forabstractIn earlier work, we have shown that articulation rate in Swedish child-directed speech (CDS) increases as a function of the age of the child, even when utterance length and differences in articulation rate between subjects are controlled for. In this paper we show on utterance level in spontaneous Swedish speech that i) for the youngest children, articulation rate in CDS is lower than in adult-directed speech (ADS), ii) there is a significant negative correlation between articulation rate and surprisal (the negative log probability) in ADS, and iii) the increase in articulation rate in Swedish CDS as a function of the age of the child holds, even when surprisal along with utterance length and differences in articulation rate between speakers are controlled for. These results indicate that adults adjust their articulation rate to make it fit the linguistic capacity of the child. Johan Sjons, Thomas Hörberg, Robert Östling, Johannes Bjerva |
INTERSPEECH | 4 |
| 2016 | Semantic Tagging with Deep Residual NetworksabstractWe propose a novel semantic tagging task, semtagging, tailored for the purpose of multilingual semantic parsing, and present the first tagger using deep residual networks (ResNets). Our tagger uses both word and character representations, and includes a novel residual bypass architecture. We evaluate the tagset both intrinsically on the new task of semantic tagging, as well as on Part-of-Speech (POS) tagging. Our system, consisting of a ResNet and an auxiliary loss function predicting our semantic tags, significantly outperforms prior results on English Universal Dependencies POS tagging (95.71% accuracy on UD v1.2 and 95.67% accuracy on UD v1.3). Johannes Bjerva, Barbara Plank, Johan Bos |
COLING | 1 |
| 2014 | Multi-class Animacy Classification with Semantic FeaturesabstractAnimacy is the semantic property of nouns denoting whether an entity can act, or is perceived as acting, of its own will.This property is marked grammatically in various languages, albeit rarely in English.It has recently been highlighted as a relevant property for NLP applications such as parsing and anaphora resolution.In order for animacy to be used in conjunction with other semantic features for such applications, appropriate data is necessary.However, the few corpora which do contain animacy annotation, rarely contain much other semantic information.The addition of such an annotation layer to a corpus already containing deep semantic annotation should therefore be of particular interest.The work presented in this paper contains three main contributions.Firstly, we improve upon the state of the art in multiclass animacy classification.Secondly, we use this classifier to contribute to the annotation of an openly available corpus containing deep semantic annotation.Finally, we provide source code, as well as trained models and scripts needed to reproduce the results presented in this paper, or aid in annotation of other texts. 1 Johannes Bjerva |
EACL | 1 |