VLDB 2026 Research / reviewers in the wild / expert
Alexandra Birch
dblp:24/6740
· DBLP profile ↗
55ranked-venue papers
7as first author
32since 2021 · last 2026
0000-0002-9022-3405ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 7 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Prosody of EmojisabstractProsodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure.In text-based settings, where these cues are absent, emojis act as visual surrogates that add affective and pragmatic nuance.This study examines how emojis influence prosodic realisation in speech and how listeners interpret prosodic cues to recover emoji meanings.Unlike previous work, we directly link prosody and emojis by analysing human speech data collected through a controlled elicited production task 1 .Using Bayesian multilevel modelling, we show that speakers systematically adapt their prosody based on emoji cues, and that listeners can recover intended meanings significantly above chance.Furthermore, our results reveal a clear hierarchy in prosodic shifts: greater semantic differences between emojis correspond to increased prosodic divergence.These findings suggest that emojis are meaningful carriers of prosodic intent that bridge the gap between digital text and spoken production. Giulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry Haddow |
ACL (1) | 3 |
| 2025 | Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual InterventionabstractLarge Language Models (LLMs) have shown remarkable capabilities in natural language processing but exhibit significant performance gaps among different languages.Most existing approaches to address these disparities rely on pretraining or fine-tuning, which are resourceintensive.To overcome these limitations without incurring significant costs, we propose Inference-Time Cross-Lingual Intervention (INCLINE), a novel framework that enhances LLM performance on low-performing (source) languages by aligning their internal representations with those of high-performing (target) languages during inference.INCLINE initially learns alignment matrices using parallel sentences from source and target languages through a Least-Squares optimization, and then applies these matrices during inference to transform the low-performing language representations toward the high-performing language space.Extensive experiments on nine benchmarks with five LLMs demonstrate that IN-CLINE significantly improves performance across diverse tasks and languages, compared to recent strong baselines.Our analysis demonstrates that INCLINE is highly cost-effective and applicable to a wide range of applications.In addition, we release the code to foster research along this line. Minghao Wu, Barry Haddow, Alexandra Birch |
ACL (1) | 4 |
| 2025 | Path encoding and manner salience in motion event descriptions: the case of Bulgarian and English
Radina Dobreva, Annie Holtz, Alexandra Birch, Frank Keller |
CogSci | 3 |
| 2025 | Generics are puzzling. Can language models find the missing piece?abstractGeneric sentences express generalisations about the world without explicit quantification. Although generics are central to everyday communication, building a precise semantic framework has proven difficult, in part because speakers use generics to generalise properties with widely different statistical prevalence. In this work, we study the implicit quantification and context-sensitivity of generics by leveraging language models as models of language. We create ConGen, a dataset of 2873 naturally occurring generic and quantified sentences in context, and define p-acceptability, a metric based on surprisal that is sensitive to quantification. Our experiments show generics are more context-sensitive than determiner quantifiers and about 20% of naturally occurring generics we analyze express weak generalisations. We also explore how human biases in stereotypes can be observed in language models. Gustavo Cilleruelo Calderón, Emily Allaway, Barry Haddow, Alexandra Birch |
COLING | 4 |
| 2025 | No Train but Gain: Language Arithmetic for training-free Language Adapters enhancementabstractModular deep learning is the state-of-the-art solution for lifting the curse of multilinguality, preventing the impact of negative interference and enabling cross-lingual performance in Multilingual Pre-trained Language Models. However, a trade-off of this approach is the reduction in positive transfer learning from closely related languages. In response, we introduce a novel method called language arithmetic, which enables training-free post-processing to address this limitation. Extending the task arithmetic framework, we apply learning via addition to the language adapters, transitioning the framework from a multi-task to a multilingual setup. The effectiveness of the proposed solution is demonstrated on three downstream tasks in a MAD-X-based set of cross-lingual schemes, acting as a post-processing procedure. Language arithmetic consistently improves the baselines with significant gains, especially in the most challenging case of zero-shot application. Our code and models are available at https://github.com/mklimasz/language-arithmetic. Mateusz Klimaszewski, Piotr Andruszkiewicz, Alexandra Birch |
COLING | 3 |
| 2025 | The Only Way is Ethics: A Guide to Ethical Research with Large Language ModelsabstractThere is a significant body of work looking at the ethical considerations of large language models (LLMs): critiquing tools to measure performance and harms; proposing toolkits to aid in ideation; discussing the risks to workers; considering legislation around privacy and security etc. As yet there is no work that integrates these resources into a single practical guide that focuses on LLMs; we attempt this ambitious goal. We introduce LLM Ethics Whitepaper, which we provide as an open and living resource for NLP practitioners, and those tasked with evaluating the ethical implications of others’ work. Our goal is to translate ethics literature into concrete recommendations for computer scientists. LLM Ethics Whitepaper distils a thorough literature review into clear Do’s and Don’ts, which we present also in this paper. We likewise identify useful toolkits to support ethical work. We refer the interested reader to the full LLM Ethics Whitepaper, which provides a succinct discussion of ethical considerations at each stage in a project lifecycle, as well as citations for the hundreds of papers from which we drew our recommendations. The present paper can be thought of as a pocket guide to conducting ethical research with LLMs. Eddie L. Ungless, Nikolas Vitsakis, Zeerak Talat, James Garforth, Björn Ross, Arno Onken, Atoosa Kasirzadeh, Alexandra Birch |
COLING | 8 |
| 2025 | Improving Multilingual Retrieval-Augmented Language Models through Dialectic Reasoning ArgumentationsabstractRetrieval-augmented generation (RAG) is key to improving large language models (LLMs) in systematically accessing richer factual knowledge.Yet, using RAG mechanisms brings intrinsic challenges, as LLMs must deal with conflicting knowledge, especially in multilingual retrieval, where the heterogeneity of knowledge retrieved may deliver different outlooks.To make RAG more analytical, critical and grounded, we introduce Dialectic-RAG (D-RAG), a modular approach guided by Argumentative Explanations, i.e., structured reasoning process that systematically evaluates retrieved information by comparing, contrasting, and resolving conflicting perspectives.Given a query and a set of multilingual related documents, D-RAG selects and exemplifies relevant knowledge for delivering dialectic explanations that, by critically weighing opposing arguments and filtering extraneous content, clearly determine the final response.We show the impact of our framework both as an incontext learning strategy and for constructing demonstrations to instruct smaller models.Our experiments demonstrate that D-RAG significantly improves RAG approaches, requiring low-impact computational effort and providing robustness to knowledge perturbations.英格兰有多少位女王[1] [英国曾有 13 位君主... Retrieval Augmented LLM Baseline Dialectic RAG -[3] states that eight queens ruled the Kingdom of England, naming Matilda, Lady Jane Grey, Mary I, Elizabeth I, Mary II, Anne, Victoria, and Elizabeth II.-[4] focuses specifically on Queen Anne, who ruled from 1702 and became the first monarch of the UK in 1707 following the union of England and Scotland.-[1] outlines the formation of the UK and mention 13 monarchs.-[5] mentions Elizabeth II and her rule over the UK but does not address the total number of queens in England.-Doc [3]: Claims there were eight queens in England's history.(Relevant): -Doc [4]: Discusses Queen Anne's reign and her role as the monarch when England and Scotland unified in 1707.(Partially Relevant) -Doc [1]: Explains the timeline of the United Kingdom's formation but does not reference the specific number of queens.(Partially Relevant) -Doc [5]: Mentions Elizabeth II's reign over the United Kingdom but focuses on her role in the Commonwealth.(Irrelevant) [3] In less than 200 years, eight queens ruled the Kingdom of England. Leonardo Ranaldi, Federico Ranaldi, Fabio Massimo Zanzotto, Barry Haddow, Alexandra Birch |
EMNLP | 5 |
| 2025 | EuroGEST: Investigating gender stereotypes in multilingual language modelsabstractLarge language models increasingly support multiple languages, yet most benchmarks for gender bias remain English-centric.We introduce EuroGEST, a dataset designed to measure gender-stereotypical reasoning in LLMs across English and 29 European languages.Eu-roGEST builds on an existing expert-informed benchmark covering 16 gender stereotypes, expanded in this work using translation tools, quality estimation metrics, and morphological heuristics.Human evaluations confirm that our data generation method results in high accuracy of both translations and gender labels across languages.We use EuroGEST to evaluate 24 multilingual language models from six model families, demonstrating that the strongest stereotypes in all models across all languages are that women are beautiful, empathetic and neat and men are leaders, strong, tough and professional.We also show that larger models encode gendered stereotypes more strongly and that instruction finetuned models continue to exhibit gendered stereotypes.Our work highlights the need for more multilingual studies of fairness in LLMs and offers scalable methods and resources to audit gender bias across languages. Jacqueline Rowe, Mateusz Klimaszewski, Liane Guillou, Shannon Vallor, Alexandra Birch |
EMNLP | 5 |
| 2025 | Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine TranslationabstractThe faithful transfer of contextually-embedded meaning continues to challenge contemporary machine translation (MT), particularly in the rendering of culture-bound terms-expressions or concepts rooted in specific languages or cultures, resisting direct linguistic transfer.Existing computational approaches to explicitating these terms have focused exclusively on in-text solutions, overlooking paratextual apparatus in the footnotes and endnotes employed by professional translators.In this paper, we formalize Genette's (1987) theory of paratexts from literary and translation studies to introduce the task of paratextual explicitation for MT.We construct a dataset of 560 expert-aligned paratexts from four English translations of the classical Chinese short story collection Liaozhai and evaluate LLMs with and without reasoning traces on choice and content of explicitation.Experiments across intrinsic prompting and agentic retrieval methods establish the difficulty of this task, with human evaluation showing that LLMgenerated paratexts improve audience comprehension, though remain considerably less effective than translator-authored ones.Beyond model performance, statistical analysis reveals that even professional translators vary widely in their use of paratexts, suggesting that cultural mediation is inherently open-ended rather than prescriptive.Our findings demonstrate the potential of paratextual explicitation in advancing MT beyond linguistic equivalence, with promising extensions to monolingual explanation and personalized adaptation. Sherrie Shen, Alexandra Birch |
EMNLP | 3 |
| 2025 | Machine Translation Meta Evaluation through Translation Accuracy Challenge SetsabstractAbstract Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgment. However, these results are often obtained by averaging predictions across large test sets without any insights into the strengths and weaknesses of these metrics across different error types. Challenge sets are used to probe specific dimensions of metric behavior but there are very few such datasets and they either focus on a limited number of phenomena or a limited number of language pairs. We introduce ACES, a contrastive challenge set spanning 146 language pairs, aimed at discovering whether metrics can identify 68 translation accuracy errors. These phenomena range from basic alterations at the word/character level to more intricate errors based on discourse and real-world knowledge. We conducted a large-scale study by benchmarking ACES on 47 metrics submitted to the WMT 2022 and WMT 2023 metrics shared tasks. We also measure their sensitivity to a range of linguistic phenomena. We further investigate claims that large language models (LLMs) are effective as MT evaluators, addressing the limitations of previous studies by using a dataset that covers a range of linguistic phenomena and language pairs and includes both low- and medium-resource languages. Our results demonstrate that different metric families struggle with different phenomena and that LLM-based methods are unreliable. We expose a number of major flaws with existing methods: Most metrics ignore the source sentence; metrics tend to prefer surface level overlap; and over-reliance on language-agnostic representations leads to confusion when the target language is similar to the source language. To further encourage detailed evaluation beyond singular scores, we expand ACES to include error span annotations, denoted as SPAN-ACES, and we use this dataset to evaluate span-based error metrics, showing that these metrics also need considerable improvement. Based on our observations, we provide a set of recommendations for building better MT metrics, including focusing on error labels instead of scores, ensembling, designing metrics to explicitly focus on the source sentence, focusing on semantic content rather than relying on the lexical overlap, and choosing the right pre-trained model for obtaining representations. Nikita Moghe, Arnisa Fazla, Chantal Amrhein, Tom Kocmi, Mark Steedman, Alexandra Birch, Rico Sennrich, Liane Guillou |
Comput. Linguistics | 6 |
| 2024 | Document-Level Machine Translation with Large-Scale Public Parallel CorporaabstractDespite the fact that document-level machine translation has inherent advantages over sentence-level machine translation due to additional information available to a model from document context, most translation systems continue to operate at a sentence level.This is primarily due to the severe lack of publicly available large-scale parallel corpora at the document level.We release a large-scale open parallel corpus with document context extracted from ParaCrawl in five language pairs, along with code to compile document-level datasets for any language pair supported by ParaCrawl.We train context-aware models on these datasets and find improvements in terms of overall translation quality and targeted document-level phenomena.We also analyse how much long-range information is useful to model some of these discourse phenomena and find models are able to utilise context from several preceding sentences. Proyag Pal, Alexandra Birch, Kenneth Heafield |
ACL (1) | 2 |
| 2024 | Retrieval-Augmented Multilingual Knowledge EditingabstractKnowledge represented in Large Language Models (LLMs) is quite often incorrect and can also become obsolete over time.Updating knowledge via fine-tuning is computationally resource-hungry and not reliable, and so knowledge editing (KE) has developed as an effective and economical alternative to inject new knowledge or to fix factual errors in LLMs.Although there has been considerable interest in this area, current KE research exclusively focuses on monolingual settings, typically in English.However, what happens if the new knowledge is supplied in one language, but we would like to query an LLM in a different language?To address the problem of multilingual knowledge editing, we propose Retrieval-Augmented Multilingual Knowledge Editor (ReMaKE) to update knowledge in LLMs.Re-MaKE can be used to perform model-agnostic knowledge editing in a multilingual setting.ReMaKE concatenates the new knowledge retrieved from a multilingual knowledge base with users' prompts before querying an LLM.Our experimental results show that ReMaKE outperforms baseline knowledge editing methods by a significant margin and is scalable to real-word application scenarios.Our multilingual knowledge editing dataset (MzsRE) in 12 languages, the code, and additional project information are available at https://github. com/weixuan-wang123/ReMaKE. Barry Haddow, Alexandra Birch |
ACL (1) | 3 |
| 2024 | Effects of Context on the Use of Descriptive Verbs
Radina Dobreva, Frank Keller, Alexandra Birch |
CogSci | 3 |
| 2024 | Is Modularity Transferable? A Case Study through the Lens of Knowledge DistillationabstractThe rise of Modular Deep Learning showcases its potential in various Natural Language Processing applications. Parameter-efficient fine-tuning (PEFT) modularity has been shown to work for various use cases, from domain adaptation to multilingual setups. However, all this work covers the case where the modular components are trained and deployed within one single Pre-trained Language Model (PLM). This model-specific setup is a substantial limitation on the very modularity that modular architectures are trying to achieve. We ask whether current modular approaches are transferable between models and whether we can transfer the modules from more robust and larger PLMs to smaller ones. In this work, we aim to fill this gap via a lens of Knowledge Distillation, commonly used for model compression, and present an extremely straightforward approach to transferring pre-trained, task-specific PEFT modules between same-family PLMs. Moreover, we propose a method that allows the transfer of modules between incompatible PLMs without any change in the inference complexity. The experiments on Named Entity Recognition, Natural Language Inference, and Paraphrase Identification tasks over multiple languages and PEFT methods showcase the initial potential of transferable modularity. Mateusz Klimaszewski, Piotr Andruszkiewicz, Alexandra Birch |
LREC/COLING | 3 |
| 2024 | Code-Switched Language Identification is Harder Than You ThinkabstractLaurie Burchell, Alexandra Birch, Robert Thompson, Kenneth Heafield. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Laurie Burchell, Alexandra Birch, Robert P. Thompson, Kenneth Heafield |
EACL (1) | 2 |
| 2024 | Contrastive Decoding Reduces Hallucinations in Large Multilingual Machine Translation ModelsabstractIn Neural Machine Translation (NMT), models will sometimes generate repetitive or fluent output that is not grounded in the source sentence.This phenomenon is known as hallucination and is a problem even in large-scale multilingual translation models.We propose to use Contrastive Decoding, an algorithm developed to improve generation from unconditional language models, to mitigate hallucinations in NMT.Specifically, we maximise the log-likelihood difference between a model and the same model with reduced contribution from the encoder outputs.Additionally, we propose an alternative implementation of Contrastive Decoding that dynamically weights the difference based on the maximum probability in the output distribution to reduce the effect of CD when the model is confident of its prediction.We evaluate our methods using the Small (418M) and Medium (1.2B) M2M models across 21 low and medium-resource language pairs.Our results show a 14.6 ± 0.5 and 11.0 ± 0.6 maximal increase in the mean COMET scores for the Small and Medium models (respectively) on those sentences for which the M2M models initially generate a hallucination. Jonas Waldendorf, Barry Haddow, Alexandra Birch |
EACL (1) | 3 |
| 2024 | Empowering Multi-step Reasoning across Languages via Program-Aided Language ModelsabstractIn-context learning methods are commonly employed as inference strategies, where Large Language Models (LLMs) are elicited to solve a task by leveraging provided demonstrations without requiring parameter updates.Among these approaches are the reasoning methods, exemplified by Chain-of-Thought (CoT) and Program-Aided Language Models (PAL), which encourage LLMs to generate reasoning steps, leading to improved accuracy.Despite their success, the ability to deliver multi-step reasoning remains limited to a single language, making it challenging to generalize to other languages and hindering global development.In this work, we propose Cross-lingual Program-Aided Language Models (Cross-PAL), a method for aligning reasoning programs across languages.Our method delivers programs as intermediate reasoning steps in different languages through a double-step cross-lingual prompting mechanism inspired by the Program-Aided approach.Moreover, we introduce Self-consistent Cross-PAL (SCross-PAL) to ensemble different reasoning paths across languages.Our experimental evaluations show that Cross-PAL outperforms existing methods, reducing the number of interactions and achieving state-of-the-art performance. Leonardo Ranaldi, Giulia Pucci, Barry Haddow, Alexandra Birch |
EMNLP | 4 |
| 2024 | When Does Monolingual Data Help Multilingual Translation: The Role of Domain and Model ScaleabstractChristos Baziotis, Biao Zhang, Alexandra Birch, Barry Haddow. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Christos Baziotis, Biao Zhang 0006, Alexandra Birch, Barry Haddow |
NAACL-HLT | 3 |
| 2024 | Assessing Factual Reliability of Large Language Model KnowledgeabstractWeixuan Wang, Barry Haddow, Alexandra Birch, Wei Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Barry Haddow, Alexandra Birch, Wei Peng 0011 |
NAACL-HLT | 3 |
| 2024 | Can GPT-3.5 generate and code discharge summaries?abstractOBJECTIVES: The aim of this study was to investigate GPT-3.5 in generating and coding medical documents with International Classification of Diseases (ICD)-10 codes for data augmentation on low-resource labels. MATERIALS AND METHODS: Employing GPT-3.5 we generated and coded 9606 discharge summaries based on lists of ICD-10 code descriptions of patients with infrequent (or generation) codes within the MIMIC-IV dataset. Combined with the baseline training set, this formed an augmented training set. Neural coding models were trained on baseline and augmented data and evaluated on an MIMIC-IV test set. We report micro- and macro-F1 scores on the full codeset, generation codes, and their families. Weak Hierarchical Confusion Matrices determined within-family and outside-of-family coding errors in the latter codesets. The coding performance of GPT-3.5 was evaluated on prompt-guided self-generated data and real MIMIC-IV data. Clinicians evaluated the clinical acceptability of the generated documents. RESULTS: Data augmentation results in slightly lower overall model performance but improves performance for the generation candidate codes and their families, including 1 absent from the baseline training data. Augmented models display lower out-of-family error rates. GPT-3.5 identifies ICD-10 codes by their prompted descriptions but underperforms on real data. Evaluators highlight the correctness of generated concepts while suffering in variety, supporting information, and narrative. DISCUSSION AND CONCLUSION: While GPT-3.5 alone given our prompt setting is unsuitable for ICD-10 coding, it supports data augmentation for training neural models. Augmentation positively affects generation code families but mainly benefits codes with existing examples. Augmentation reduces out-of-family errors. Documents generated by GPT-3.5 state prompted concepts correctly but lack variety, and authenticity in narratives. Matús Falis, Aryo Pradipta Gema, Hang Dong 0002, Luke Daines, Siddharth Basetti, Michael Holder, Rose S. Penfold, Alexandra Birch, Beatrice Alex |
J. Am. Medical Informatics Assoc. | 8 |
| 2023 | Extrinsic Evaluation of Machine Translation MetricsabstractAutomatic machine translation (MT) metrics are widely used to distinguish the quality of machine translation systems across large test sets (i.e., system-level evaluation).However, it is unclear if automatic metrics can reliably distinguish good translations from bad at the sentence level (i.e., segment-level evaluation).We investigate how useful MT metrics are at detecting segment-level quality by correlating metrics with the translation utility for downstream tasks.We evaluate the segment-level performance of widespread MT metrics (chrF, COMET, BERTScore, etc.) on three downstream cross-lingual tasks (dialogue state tracking, question answering, and semantic parsing).For each task, we have access to a monolingual task-specific model and a translation model.We calculate the correlation between the metric's ability to predict a good/bad translation with the success/failure on the final task for machine-translated test sentences.Our experiments demonstrate that all metrics exhibit negligible correlation with the extrinsic evaluation of downstream outcomes.We also find that the scores provided by neural metrics are not interpretable, in large part due to having undefined ranges.We synthesise our analysis into recommendations for future MT metrics to produce labels rather than scores for more informative interaction between machine translation and multilingual language understanding. Nikita Moghe, Tom Sherborne, Mark Steedman, Alexandra Birch |
ACL (1) | 4 |
| 2023 | Prompting Large Language Model for Machine Translation: A Case StudyabstractResearch on prompting has shown excellent performance with little or even no supervised training across many tasks. However, prompting for machine translation is still under-explored in the literature. We fill this gap by offering a systematic study on prompting strategies for translation, examining various factors for prompt template and demonstration example selection. We further explore the use of monolingual data and the feasibility of cross-lingual, cross-domain, and sentence-to-document transfer learning in prompting. Extensive experiments with GLM-130B (Zeng et al., 2022) as the testbed show that 1) the number and the quality of prompt examples matter, where using suboptimal examples degenerates translation; 2) several features of prompt examples, such as semantic similarity, show significant Spearman correlation with their prompting performance; yet, none of the correlations are strong enough; 3) using pseudo parallel prompt examples constructed from monolingual data via zero-shot prompting could improve translation; and 4) improved performance is achievable by transferring knowledge from prompt examples selected in other settings. We finally provide an analysis on the model outputs and discuss several problems that prompting still suffers from. Biao Zhang 0006, Barry Haddow, Alexandra Birch |
ICML | 3 |
| 2023 | Hallucinations in Large Multilingual Translation ModelsabstractAbstract Hallucinated translations can severely undermine and raise safety issues when machine translation systems are deployed in the wild. Previous research on the topic focused on small bilingual models trained on high-resource languages, leaving a gap in our understanding of hallucinations in multilingual models across diverse translation scenarios. In this work, we fill this gap by conducting a comprehensive analysis—over 100 language pairs across various resource levels and going beyond English-centric directions—on both the M2M neural machine translation (NMT) models and GPT large language models (LLMs). Among several insights, we highlight that models struggle with hallucinations primarily in low-resource directions and when translating out of English, where, critically, they may reveal toxic patterns that can be traced back to the training data. We also find that LLMs produce qualitatively different hallucinations to those of NMT models. Finally, we show that hallucinations are hard to reverse by merely scaling models trained with the same data. However, employing more diverse models, trained on different data or with different procedures, as fallback systems can improve translation quality and virtually eliminate certain pathologies. Nuno Miguel Guerreiro, Duarte M. Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 5 |
| 2022 | GoURMET - Machine Translation for Low-Resourced LanguagesabstractThe GoURMET project, funded by the European Commission’s H2020 program (under grant agreement 825299), develops models for machine translation, in particular for low-resourced languages. Data, models and software releases as well as the GoURMET Translate Tool are made available as open source. Peggy van der Kreeft, Alexandra Birch, Sevi Sariisik, Felipe Sánchez-Martínez, Wilker Aziz |
EAMT | 2 |
| 2022 | Non-Autoregressive Machine Translation: It's Not as Fast as it SeemsabstractEfficient machine translation models are commercially important as they can increase inference speeds, and reduce costs and carbon emissions.Recently, there has been much interest in non-autoregressive (NAR) models, which promise faster translation.In parallel to the research on NAR models, there have been successful attempts to create optimized autoregressive models as part of the WMT shared task on efficient translation.In this paper, we point out flaws in the evaluation methodology present in the literature on NAR models and we provide a fair comparison between a stateof-the-art NAR model and the autoregressive submissions to the shared task.We make the case for consistent evaluation of NAR models, and also for the importance of comparing NAR models with other widely used methods for improving efficiency.We run experiments with a connectionist-temporal-classification-based (CTC) NAR model implemented in C++ and compare it with AR models using wall clock times.Our results show that, although NAR models are faster on GPUs, with small batch sizes, they are almost always slower under more realistic usage conditions.We call for more realistic and extensive evaluation of NAR models in future work. Jindrich Helcl, Barry Haddow, Alexandra Birch |
NAACL-HLT | 3 |
| 2022 | Quantifying Synthesis and Fusion and their Impact on Machine TranslationabstractArturo Oncevay, Duygu Ataman, Niels Van Berkel, Barry Haddow, Alexandra Birch, Johannes Bjerva. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Arturo Oncevay, Duygu Ataman, Niels van Berkel, Barry Haddow, Alexandra Birch, Johannes Bjerva |
NAACL-HLT | 5 |
| 2022 | Survey of Low-Resource Machine TranslationabstractAbstract We present a survey covering the state of the art in low-resource machine translation (MT) research. There are currently around 7,000 languages spoken in the world and almost all language pairs lack significant resources for training machine translation models. There has been increasing interest in research addressing the challenge of producing useful translation models when very little translated training data is available. We present a summary of this topical research field and provide a description of the techniques evaluated by researchers in several recent shared tasks in low-resource MT. Barry Haddow, Rachel Bawden, Antonio Valerio Miceli Barone, Jindrich Helcl, Alexandra Birch |
Comput. Linguistics | 5 |
| 2021 | Few-shot learning through contextual data augmentationabstractMachine translation (MT) models used in industries with constantly changing topics, such as translation or news agencies, need to adapt to new data to maintain their performance over time.Our aim is to teach a pre-trained MT model to translate previously unseen words accurately, based on very few examples.We propose (i) an experimental setup allowing us to simulate novel vocabulary appearing in human-submitted translations, and (ii) corresponding evaluation metrics to compare our approaches.We extend a data augmentation approach using a pre-trained language model to create training examples with similar contexts for novel words.We compare different fine-tuning and data augmentation approaches and show that adaptation on the scale of one to five examples is possible.Combining data augmentation with randomly selected training sentences leads to the highest BLEU score and accuracy improvements.Impressively, with only 1 to 5 examples, our model reports better accuracy scores than a reference system trained with on average 313 parallel examples. Farid Arthaud, Rachel Bawden, Alexandra Birch |
EACL | 3 |
| 2021 | CoPHE: A Count-Preserving Hierarchical Evaluation Metric in Large-Scale Multi-Label Text ClassificationabstractLarge-Scale Multi-Label Text Classification (LMTC) includes tasks with hierarchical label spaces, such as automatic assignment of ICD-9 codes to discharge summaries.Performance of models in prior art is evaluated with standard precision, recall, and F 1 measures without regard for the rich hierarchical structure.In this work we argue for hierarchical evaluation of the predictions of neural LMTC models.With the example of the ICD-9 ontology we describe a structural issue in the representation of the structured label space in prior art, and propose an alternative representation based on the depth of the ontology.We propose a set of metrics for hierarchical evaluation using the depthbased representation.We compare the evaluation scores from the proposed metrics with previously used metrics on prior art LMTC models for ICD-9 coding in MIMIC-III.We also propose further avenues of research involving the proposed ontological representation. Matús Falis, Hang Dong 0002, Alexandra Birch, Beatrice Alex |
EMNLP (1) | 3 |
| 2021 | Cross-lingual Intermediate Fine-tuning improves Dialogue State TrackingabstractRecent progress in task-oriented neural dialogue systems is largely focused on a handful of languages, as annotation of training data is tedious and expensive.Machine translation has been used to make systems multilingual, but this can introduce a pipeline of errors.Another promising solution is using cross-lingual transfer learning through pretrained multilingual models.Existing methods train multilingual models with additional codemixed task data or refine the cross-lingual representations through parallel ontologies.In this work, we enhance the transfer learning process by intermediate fine-tuning of pretrained multilingual models, where the multilingual models are fine-tuned with different but related data and/or tasks.Specifically, we use parallel and conversational movie subtitles datasets to design cross-lingual intermediate tasks suitable for downstream dialogue tasks.We use only 200K lines of parallel data for intermediate fine-tuning which is already available for 1782 language pairs.We test our approach on the cross-lingual dialogue state tracking task for the parallel Mul-tiWoZ (English→Chinese, Chinese→English) and Multilingual WoZ (English→German, English→Italian) datasets.We achieve impressive improvements (> 20% on joint goal accuracy) on the parallel MultiWoZ dataset and the Multilingual WoZ dataset over the vanilla baseline with only 10% of the target language task data and zero-shot setup respectively. Nikita Moghe, Mark Steedman, Alexandra Birch |
EMNLP (1) | 3 |
| 2021 | Surprise Language Challenge: Developing a Neural Machine Translation System between Pashto and English in Two MonthsabstractIn the media industry and the focus of global reporting can shift overnight. There is a compelling need to be able to develop new machine translation systems in a short period of time and in order to more efficiently cover quickly developing stories. As part of the EU project GoURMET and which focusses on low-resource machine translation and our media partners selected a surprise language for which a machine translation system had to be built and evaluated in two months(February and March 2021). The language selected was Pashto and an Indo-Iranian language spoken in Afghanistan and Pakistan and India. In this period we completed the full pipeline of development of a neural machine translation system: data crawling and cleaning and aligning and creating test sets and developing and testing models and and delivering them to the user partners. In this paperwe describe rapid data creation and experiments with transfer learning and pretraining for this low-resource language pair. We find that starting from an existing large model pre-trained on 50languages leads to far better BLEU scores than pretraining on one high-resource language pair with a smaller model. We also present human evaluation of our systems and which indicates that the resulting systems perform better than a freely available commercial system when translating from English into Pashto direction and and similarly when translating from Pashto into English. Alexandra Birch, Barry Haddow, Antonio Valerio Miceli Barone, Jindrich Helcl, Jonas Waldendorf, Felipe Sánchez-Martínez, Mikel L. Forcada, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Miquel Esplà-Gomis, Wilker Aziz, Lina Murady, Sevi Sariisik, Peggy van der Kreeft, Kay Macquarrie |
MTSummit (1) | 1 |
| 2021 | Neural Machine Translation 2020, by Philipp Koehn, Cambridge, Cambridge University Press, ISBN 978-1-108-49732-9, pages 393
Alexandra Birch |
Nat. Lang. Eng. | 1 |
| 2020 | Language Model Prior for Low-Resource Neural Machine TranslationabstractThe scarcity of large parallel corpora is an important obstacle for neural machine translation.A common solution is to exploit the knowledge of language models (LM) trained on abundant monolingual data.In this work, we propose a novel approach to incorporate a LM as prior in a neural translation model (TM).Specifically, we add a regularization term, which pushes the output distributions of the TM to be probable under the LM prior, while avoiding wrong predictions when the TM "disagrees" with the LM.This objective relates to knowledge distillation, where the LM can be viewed as teaching the TM about the target language.The proposed approach does not compromise decoding speed, because the LM is used only at training time, unlike previous work that requires it during inference.We present an analysis on the effects that different methods have on the distributions of the TM.Results on two low-resource machine translation datasets show clear improvements even with limited monolingual data. Christos Baziotis, Barry Haddow, Alexandra Birch |
EMNLP (1) | 3 |
| 2020 | Bridging Linguistic Typology and Multilingual Machine Translation with Multi-View Language RepresentationsabstractSparse language vectors from linguistic typology databases and learned embeddings from tasks like multilingual machine translation have been investigated in isolation, without analysing how they could benefit from each other's language characterisation.We propose to fuse both views using singular vector canonical correlation analysis and study what kind of information is induced from each source.By inferring typological features and language phylogenies, we observe that our representations embed typology and strengthen correlations with language relationships.We then take advantage of our multi-view language vector space for multilingual machine translation, where we achieve competitive overall translation accuracy in tasks that require information about language similarities, such as language clustering and ranking candidates for multilingual transfer.With our method, which is also released as a tool, we can easily project and assess new languages without expensive retraining of massive multilingual or ranking models, which are major disadvantages of related approaches. Arturo Oncevay, Barry Haddow, Alexandra Birch |
EMNLP (1) | 3 |
| 2020 | A Latent Morphology Model for Open-Vocabulary Neural Machine Translation
Duygu Ataman, Wilker Aziz, Alexandra Birch |
ICLR | 3 |
| 2020 | Towards Making the Most of Context in Neural Machine TranslationabstractDocument-level machine translation manages to outperform sentence level models by a small margin, but have failed to be widely adopted. We argue that previous research did not make a clear use of the global context, and propose a new document-level NMT framework that deliberately models the local context of each sentence with the awareness of the global context of the document in both source and target languages. We specifically design the model to be able to deal with documents containing any number of sentences, including single sentences. This unified approach allows our model to be trained elegantly on standard datasets without needing to train on sentence and document level data separately. Experimental results demonstrate that our model outperforms Transformer baselines and previous document-level NMT models with substantial margins of up to 2.1 BLEU on state-of-the-art baselines. We also provide analyses which show the benefit of context far beyond the neighboring two or three sentences, which previous studies have typically incorporated. Zaixiang Zheng, Xiang Yue, Shujian Huang, Jiajun Chen 0001, Alexandra Birch |
IJCAI | 5 |
| 2020 | Multiword Expression aware Neural Machine TranslationabstractMultiword Expressions (MWEs) are a frequently occurring phenomenon found in all natural languages that is of great importance to linguistic theory, natural language processing applications, and machine translation systems. Neural Machine Translation (NMT) architectures do not handle these expressions well and previous studies have rarely addressed MWEs in this framework. In this work, we show that annotation and data augmentation, using external linguistic resources, can improve both translation of MWEs that occur in the source, and the generation of MWEs on the target, and increase performance by up to 5.09 BLEU points on MWE test sets. We also devise a MWE score to specifically assess the quality of MWE translation which agrees with human evaluation. We make available the MWE score implementation – along with MWE-annotated training sets and corpus-based lists of MWEs – for reproduction and extension. Andrea Zaninello, Alexandra Birch |
LREC | 2 |
| 2019 | Global Under-Resourced Media Translation (GoURMET)
Alexandra Birch, Barry Haddow, Ivan Titov 0001, Antonio Valerio Miceli Barone, Rachel Bawden, Felipe Sánchez-Martínez, Mikel L. Forcada, Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Wilker Aziz, Andrew Secker, Peggy van der Kreeft |
MTSummit (2) | 1 |
| 2018 | The SUMMA Platform: Scalable Understanding of Multilingual MediaabstractWe present the latest version of the SUMMA platform, an open-source software platform for monitoring and interpreting multi-lingual media, from written news published on the internet to live media broadcasts via satellite or internet streaming. Ulrich Germann, Peggy van der Kreeft, Guntis Barzdins, Alexandra Birch |
EAMT | 4 |
| 2018 | Evaluating Discourse Phenomena in Neural Machine TranslationabstractRachel Bawden, Rico Sennrich, Alexandra Birch, Barry Haddow. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Rachel Bawden, Rico Sennrich, Alexandra Birch, Barry Haddow |
NAACL-HLT | 3 |
| 2016 | Improving Neural Machine Translation Models with Monolingual DataabstractNeural Machine Translation (NMT) has obtained state-of-the art performance for several language pairs, while only using parallel data for training.Targetside monolingual data plays an important role in boosting fluency for phrasebased statistical machine translation, and we investigate the use of monolingual data for NMT.In contrast to previous work, which combines NMT models with separately trained language models, we note that encoder-decoder NMT architectures already have the capacity to learn the same information as a language model, and we explore strategies to train with monolingual data without changing the neural network architecture.By pairing monolingual training data with an automatic backtranslation, we can treat it as additional parallel training data, and we obtain substantial improvements on the WMT 15 task English↔German (+2.8-3.7 BLEU), and for the low-resourced IWSLT 14 task Turkish→English (+2.1-3.4BLEU), obtaining new state-of-the-art results.We also show that fine-tuning on in-domain monolingual and parallel data gives substantial improvements for the IWSLT 15 task English→German. Rico Sennrich, Barry Haddow, Alexandra Birch |
ACL (1) | 3 |
| 2016 | Neural Machine Translation of Rare Words with Subword UnitsabstractNeural machine translation (NMT) models typically operate with a fixed vocabulary, but translation is an open-vocabulary problem.Previous work addresses the translation of out-of-vocabulary words by backing off to a dictionary.In this paper, we introduce a simpler and more effective approach, making the NMT model capable of open-vocabulary translation by encoding rare and unknown words as sequences of subword units.This is based on the intuition that various word classes are translatable via smaller units than words, for instance names (via character copying or transliteration), compounds (via compositional translation), and cognates and loanwords (via phonological and morphological transformations).We discuss the suitability of different word segmentation techniques, including simple character ngram models and a segmentation based on the byte pair encoding compression algorithm, and empirically show that subword models improve over a back-off dictionary baseline for the WMT 15 translation tasks English→German and English→Russian by up to 1.1 and 1.3 BLEU, respectively. Rico Sennrich, Barry Haddow, Alexandra Birch |
ACL (1) | 3 |
| 2016 | HUME: Human UCCA-Based Evaluation of Machine TranslationabstractHuman evaluation of machine translation normally uses sentence-level measures such as relative ranking or adequacy scales. However, these provide no insight into possible errors, and do not scale well with sentence length. We argue for a semantics-based evaluation, which captures what meaning components are retained in the MT output, thus providing a more fine-grained analysis of translation quality, and enabling the construction and tuning of semantics-based MT. We present a novel human semantic evaluation measure, Human UCCA-based MT Evaluation (HUME), building on the UCCA semantic representation scheme. HUME covers a wider range of semantic phenomena than previous methods and does not rely on semantic annotation of the potentially garbled MT output. We experiment with four language pairs, demonstrating HUME’s broad applicability, and report good inter-annotator agreement rates and correlation with human adequacy scores. Alexandra Birch, Omri Abend, Ondrej Bojar, Barry Haddow |
EMNLP | 1 |
| 2016 | Controlling Politeness in Neural Machine Translation via Side ConstraintsabstractMany languages use honorifics to express politeness, social distance, or the relative social status between the speaker and their addressee(s). In machine translation from a language without honorifics such as English, it is difficult to predict the appropriate honorific, but users may want to control the level of politeness in the output. In this paper, we perform a pilot study to control honorifics in neural machine translation (NMT) via side constraints, focusing on English!German. We show that by marking up the (English) source side of the training data with a feature that encodes the use of honorifics on the (German) target side, we can control the honorifics produced at test time. Experiments show that the choice of honorifics has a big impact on translation quality as measured by BLEU, and oracle experiments show that substantial improvements are possible by constraining the translation to the desired level of politeness. Rico Sennrich, Barry Haddow, Alexandra Birch |
HLT-NAACL | 3 |
| 2015 | A system for automatic broadcast news summarisation, geolocation and translation
Peter Bell 0001, Catherine Lai, Clare Llewellyn, Alexandra Birch, Mark Sinclair |
INTERSPEECH | 4 |
| 2015 | Mixed domain vs. multi-domain statistical machine translation
Matthias Huck, Alexandra Birch, Barry Haddow |
MTSummit | 2 |
| 2014 | Generalizing a Strongly Lexicalized Parser using Unlabeled DataabstractStatistical parsers trained on labeled data suffer from sparsity, both grammatical and lexical.For parsers based on strongly lexicalized grammar formalisms (such as CCG, which has complex lexical categories but simple combinatory rules), the problem of sparsity can be isolated to the lexicon.In this paper, we show that semi-supervised Viterbi-EM can be used to extend the lexicon of a generative CCG parser.By learning complex lexical entries for low-frequency and unseen words from unlabeled data, we obtain improvements over our supervised model for both indomain (WSJ) and out-of-domain (questions and Wikipedia) data.Our learnt lexicons when used with a discriminative parser such as C&C also significantly improve its performance on unseen words. Tejaswini Deoskar, Christos Christodoulopoulos 0001, Alexandra Birch, Mark Steedman |
EACL | 3 |
| 2014 | Automated production of true-cased punctuated subtitles for weather and news broadcasts
Joris Driesen, Alexandra Birch, Simon Grimsey, Saeid Safarfashandi, Juliet Gauthier, Matt Simpson, Steve Renals |
INTERSPEECH | 2 |
| 2014 | A semi-Markov model for speech segmentation with an utterance-break priorabstractSpeech segmentation is the problem of finding the end points of a speech utterance for passing to an automatic speech recogni-tion (ASR) system. The quality of this segmentation can have a large impact on the accuracy of the ASR system; in this pa-per we demonstrate that it can have an even larger impact on downstream natural language processing tasks – in this case, machine translation. We develop a novel semi-Markov model which allows the segmentation of audio streams into speech ut-terances which are optimised for the desired distribution of sen-tence lengths for the target domain. We compare this with exist-ing state-of-the-art methods and show that it is able to achieve not only improved ASR performance, but also to yield signifi-cant benefits to a speech translation task. Index Terms: speech activity detection, speech segmentation, machine translation, speech recognition Mark Sinclair, Peter Bell 0001, Alexandra Birch, Fergus R. McInnes |
INTERSPEECH | 3 |
| 2011 | Reordering Metrics for MT
Alexandra Birch, Miles Osborne |
ACL | 1 |
| 2011 | Soft Dependency Constraints for Reordering in Hierarchical Phrase-Based Translation
Yang Gao 0005, Philipp Koehn, Alexandra Birch |
EMNLP | 3 |
| 2010 | Metrics for MT evaluation: evaluating reordering
Alexandra Birch, Miles Osborne, Phil Blunsom |
Mach. Transl. | 1 |
| 2009 | 462 Machine Translation Systems for Europe
Philipp Koehn, Alexandra Birch, Ralf Steinberger |
MTSummit | 2 |
| 2008 | Predicting Success in Machine Translation
Alexandra Birch, Miles Osborne, Philipp Koehn |
EMNLP | 1 |
| 2007 | Moses: Open Source Toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, Evan Herbst |
ACL | 3 |