Barry Haddow

dblp:12/5915 · DBLP profile ↗
← Back
49ranked-venue papers
4as first author
28since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 4 first-author · 28 since 2021
YearPublicationVenuePosition
2026 The Prosody of Emojis
abstract
Prosodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure.In text-based settings, where these cues are absent, emojis act as visual surrogates that add affective and pragmatic nuance.This study examines how emojis influence prosodic realisation in speech and how listeners interpret prosodic cues to recover emoji meanings.Unlike previous work, we directly link prosody and emojis by analysing human speech data collected through a controlled elicited production task 1 .Using Bayesian multilevel modelling, we show that speakers systematically adapt their prosody based on emoji cues, and that listeners can recover intended meanings significantly above chance.Furthermore, our results reveal a clear hierarchy in prosodic shifts: greater semantic differences between emojis correspond to increased prosodic divergence.These findings suggest that emojis are meaningful carriers of prosodic intent that bridge the gap between digital text and spoken production.
Giulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry Haddow
ACL (1)4
2026 HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
abstract
We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder–decoder models, as well as about 30 “smallish” monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation.
Stephan Oepen, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Maja Buljan, Laurie Burchell, Lucas Georges Gabriel Charpentier, Pinzhen Chen, Mariia Fedorova, Ona de Gibert Bonet, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Andrey Kutuzov, Veronika Laippala, Bhavitvya Malik, Vladislav Mikhailov, Amanda Myntti, Dayyán O'Brien, Lucie Poláková, Gema Ramírez-Sánchez, Janine Siewert, Pavel Stepachev, Jörg Tiedemann, Teemu Vahtola, Dusan Varis, Fedor Vitiugin, Jaume Zaragoza
LREC11
2025 An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
abstract
Laurie Burchell, Ona De Gibert Bonet, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajič, Jindřich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, Amanda Myntti, Dayyán O’Brien, Stephan Oepen, Proyag Pal, Jousia Piha, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dušan Variš, Tereza Vojtěchová, Jaume Zaragoza-Bernabeu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Laurie Burchell, Ona de Gibert Bonet, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, Amanda Myntti, Dayyán O'Brien, Stephan Oepen, Proyag Pal, Jousia Piha, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dusan Varis, Tereza Vojtechová, Jaume Zaragoza-Bernabeu
ACL (1)9
2025 Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual Intervention
abstract
Large Language Models (LLMs) have shown remarkable capabilities in natural language processing but exhibit significant performance gaps among different languages.Most existing approaches to address these disparities rely on pretraining or fine-tuning, which are resourceintensive.To overcome these limitations without incurring significant costs, we propose Inference-Time Cross-Lingual Intervention (INCLINE), a novel framework that enhances LLM performance on low-performing (source) languages by aligning their internal representations with those of high-performing (target) languages during inference.INCLINE initially learns alignment matrices using parallel sentences from source and target languages through a Least-Squares optimization, and then applies these matrices during inference to transform the low-performing language representations toward the high-performing language space.Extensive experiments on nine benchmarks with five LLMs demonstrate that IN-CLINE significantly improves performance across diverse tasks and languages, compared to recent strong baselines.Our analysis demonstrates that INCLINE is highly cost-effective and applicable to a wide range of applications.In addition, we release the code to foster research along this line.
Minghao Wu, Barry Haddow, Alexandra Birch
ACL (1)3
2025 Generics are puzzling. Can language models find the missing piece?
abstract
Generic sentences express generalisations about the world without explicit quantification. Although generics are central to everyday communication, building a precise semantic framework has proven difficult, in part because speakers use generics to generalise properties with widely different statistical prevalence. In this work, we study the implicit quantification and context-sensitivity of generics by leveraging language models as models of language. We create ConGen, a dataset of 2873 naturally occurring generic and quantified sentences in context, and define p-acceptability, a metric based on surprisal that is sensitive to quantification. Our experiments show generics are more context-sensitive than determiner quantifiers and about 20% of naturally occurring generics we analyze express weak generalisations. We also explore how human biases in stereotypes can be observed in language models.
Gustavo Cilleruelo Calderón, Emily Allaway, Barry Haddow, Alexandra Birch
COLING3
2025 Improving Multilingual Retrieval-Augmented Language Models through Dialectic Reasoning Argumentations
abstract
Retrieval-augmented generation (RAG) is key to improving large language models (LLMs) in systematically accessing richer factual knowledge.Yet, using RAG mechanisms brings intrinsic challenges, as LLMs must deal with conflicting knowledge, especially in multilingual retrieval, where the heterogeneity of knowledge retrieved may deliver different outlooks.To make RAG more analytical, critical and grounded, we introduce Dialectic-RAG (D-RAG), a modular approach guided by Argumentative Explanations, i.e., structured reasoning process that systematically evaluates retrieved information by comparing, contrasting, and resolving conflicting perspectives.Given a query and a set of multilingual related documents, D-RAG selects and exemplifies relevant knowledge for delivering dialectic explanations that, by critically weighing opposing arguments and filtering extraneous content, clearly determine the final response.We show the impact of our framework both as an incontext learning strategy and for constructing demonstrations to instruct smaller models.Our experiments demonstrate that D-RAG significantly improves RAG approaches, requiring low-impact computational effort and providing robustness to knowledge perturbations.英格兰有多少位女王[1] [英国曾有 13 位君主... Retrieval Augmented LLM Baseline Dialectic RAG -[3] states that eight queens ruled the Kingdom of England, naming Matilda, Lady Jane Grey, Mary I, Elizabeth I, Mary II, Anne, Victoria, and Elizabeth II.-[4] focuses specifically on Queen Anne, who ruled from 1702 and became the first monarch of the UK in 1707 following the union of England and Scotland.-[1] outlines the formation of the UK and mention 13 monarchs.-[5] mentions Elizabeth II and her rule over the UK but does not address the total number of queens in England.-Doc [3]: Claims there were eight queens in England's history.(Relevant): -Doc [4]: Discusses Queen Anne's reign and her role as the monarch when England and Scotland unified in 1707.(Partially Relevant) -Doc [1]: Explains the timeline of the United Kingdom's formation but does not reference the specific number of queens.(Partially Relevant) -Doc [5]: Mentions Elizabeth II's reign over the United Kingdom but focuses on her role in the Commonwealth.(Irrelevant) [3] In less than 200 years, eight queens ruled the Kingdom of England.
Leonardo Ranaldi, Federico Ranaldi, Fabio Massimo Zanzotto, Barry Haddow, Alexandra Birch
EMNLP4
2025 HPLT's Second Data Release
abstract
We describe the progress of the High Performance Language Technologies (HPLT) project, a 3-year EU-funded project that started in September 2022. We focus on the up-to-date results on the release of free text datasets derived from web crawls, one of the central objectives of the project. The second release used a revised processing pipeline, and an enlarged set of input crawls. From 4.5 petabytes of web crawls we extracted 7.6T tokens of monolingual text in 193 languages, plus 380 million parallel sentences in 51 language pairs. We also release MultiHPLT, a cross-combination of the parallel data, which produces 1,275 pairs, as well as releasing the containing documents for all parallel sentences in order to enable research in document-level MT. We report changes in the pipeline, analysis and evaluation results for the second parallel data release based on machine translation systems. All datasets are released under a permissive CC0 licence.
Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Laurie Burchell, Pinzhen Chen, Mariia Fedorova, Ona de Gibert Bonet, Liane Guillou, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Erik Henriksson, Andrey Kutuzov, Veronika Laippala, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Amanda Myntti, Dayyán O'Brien, Stephan Oepen, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dusan Varis, Jaume Zaragoza-Bernabeu
MTSummit (2)9
2025 Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
abstract
Tsz Kin Lam, Marco Gaido, Sara Papi, Luisa Bentivogli, Barry Haddow. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Tsz Kin Lam, Marco Gaido, Sara Papi, Luisa Bentivogli, Barry Haddow
NAACL (Long Papers)5
2024 Retrieval-Augmented Multilingual Knowledge Editing
abstract
Knowledge represented in Large Language Models (LLMs) is quite often incorrect and can also become obsolete over time.Updating knowledge via fine-tuning is computationally resource-hungry and not reliable, and so knowledge editing (KE) has developed as an effective and economical alternative to inject new knowledge or to fix factual errors in LLMs.Although there has been considerable interest in this area, current KE research exclusively focuses on monolingual settings, typically in English.However, what happens if the new knowledge is supplied in one language, but we would like to query an LLM in a different language?To address the problem of multilingual knowledge editing, we propose Retrieval-Augmented Multilingual Knowledge Editor (ReMaKE) to update knowledge in LLMs.Re-MaKE can be used to perform model-agnostic knowledge editing in a multilingual setting.ReMaKE concatenates the new knowledge retrieved from a multilingual knowledge base with users' prompts before querying an LLM.Our experimental results show that ReMaKE outperforms baseline knowledge editing methods by a significant margin and is scalable to real-word application scenarios.Our multilingual knowledge editing dataset (MzsRE) in 12 languages, the code, and additional project information are available at https://github. com/weixuan-wang123/ReMaKE.
Barry Haddow, Alexandra Birch
ACL (1)2
2024 Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
abstract
Human evaluation is a critical component in machine translation system development and has received much attention in text translation research. However, little prior work exists on the topic of human evaluation for speech translation, which adds additional challenges such as noisy data and segmentation mismatches. We take the first steps to fill this gap by conducting a comprehensive human evaluation of the results of several shared tasks from the last International Workshop on Spoken Language Translation (IWSLT 2023). We propose an effective evaluation strategy based on automatic resegmentation and direct assessment with segment context. Our analysis revealed that: 1) the proposed evaluation strategy is robust and scores well-correlated with other types of human judgements; 2) automatic metrics are usually, but not always, well-correlated with direct assessment scores; and 3) COMET as a slightly stronger automatic metric than chrF, despite the segmentation noise introduced by the resegmentation step systems. We release the collected human-annotated data in order to encourage further investigation.
Matthias Sperber, Ondrej Bojar, Barry Haddow, Dávid Javorský, Xutai Ma, Matteo Negri, Jan Niehues, Peter Polak, Elizabeth Salesky, Katsuhito Sudoh, Marco Turchi
LREC/COLING3
2024 Contrastive Decoding Reduces Hallucinations in Large Multilingual Machine Translation Models
abstract
In Neural Machine Translation (NMT), models will sometimes generate repetitive or fluent output that is not grounded in the source sentence.This phenomenon is known as hallucination and is a problem even in large-scale multilingual translation models.We propose to use Contrastive Decoding, an algorithm developed to improve generation from unconditional language models, to mitigate hallucinations in NMT.Specifically, we maximise the log-likelihood difference between a model and the same model with reduced contribution from the encoder outputs.Additionally, we propose an alternative implementation of Contrastive Decoding that dynamically weights the difference based on the maximum probability in the output distribution to reduce the effect of CD when the model is confident of its prediction.We evaluate our methods using the Small (418M) and Medium (1.2B) M2M models across 21 low and medium-resource language pairs.Our results show a 14.6 ± 0.5 and 11.0 ± 0.6 maximal increase in the mean COMET scores for the Small and Medium models (respectively) on those sentences for which the M2M models initially generate a hallucination.
Jonas Waldendorf, Barry Haddow, Alexandra Birch
EACL (1)2
2024 HPLT's First Release of Data and Models
abstract
The High Performance Language Technologies (HPLT) project is a 3-year EU-funded project that started in September 2022. It aims to deliver free, sustainable, and reusable datasets, models, and workflows at scale using high-performance computing. We describe the first results of the project. The data release includes monolingual data in 75 languages at 5.6T tokens and parallel data in 18 language pairs at 96M pairs, derived from 1.8 petabytes of web crawls. Building upon automated and transparent pipelines, the first machine translation (MT) models as well as large language models (LLMs) have been trained and released. Multiple data processing tools and pipelines have also been made public.
Nikolay Arefyev, Mikko Aulamo, Pinzhen Chen, Ona de Gibert Bonet, Barry Haddow, Jindrich Helcl, Bhavitvya Malik, Gema Ramírez-Sánchez, Pavel Stepachev, Jörg Tiedemann, Dusan Varis, Jaume Zaragoza-Bernabeu
EAMT (2)5
2024 Iterative Translation Refinement with Large Language Models
abstract
We propose iteratively prompting a large language model to self-correct a translation, with inspiration from their strong language capability as well as a human-like translation approach. Interestingly, multi-turn querying reduces the output’s string-based metric scores, but neural metrics suggest comparable or improved quality after two or more iterations. Human evaluations indicate better fluency and naturalness compared to initial translations and even human references, all while maintaining quality. Ablation studies underscore the importance of anchoring the refinement to the source and a reasonable seed translation for quality considerations. We also discuss the challenges in evaluation and relation to human performance and translationese.
Pinzhen Chen, Zhicheng Guo, Barry Haddow, Kenneth Heafield
EAMT (1)3
2024 Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
abstract
Multilingual large language models are designed, claimed, and expected to cater to speakers of varied languages.We hypothesise that the current practices of fine-tuning and evaluating these models may not perfectly align with this objective owing to a heavy reliance on translation, which cannot cover languagespecific knowledge but can introduce translation defects.It remains unknown whether the nature of the instruction data has an impact on the model output; conversely, it is questionable whether translated test sets can capture such nuances.Due to the often coupled practices of using translated data in both stages, such imperfections could have been overlooked.This work investigates these issues using controlled native or translated data during the instruction tuning and evaluation stages.We show that native or generation benchmarks reveal a notable difference between native and translated instruction data especially when model performance is high, whereas other types of test sets cannot.The comparison between round-trip and single-pass translations reflects the importance of knowledge from language-native resources.Finally, we demonstrate that regularization is beneficial to bridging this gap on structured but not generative tasks. 1
Pinzhen Chen, Simon Yu, Zhicheng Guo, Barry Haddow
EMNLP4
2024 Empowering Multi-step Reasoning across Languages via Program-Aided Language Models
abstract
In-context learning methods are commonly employed as inference strategies, where Large Language Models (LLMs) are elicited to solve a task by leveraging provided demonstrations without requiring parameter updates.Among these approaches are the reasoning methods, exemplified by Chain-of-Thought (CoT) and Program-Aided Language Models (PAL), which encourage LLMs to generate reasoning steps, leading to improved accuracy.Despite their success, the ability to deliver multi-step reasoning remains limited to a single language, making it challenging to generalize to other languages and hindering global development.In this work, we propose Cross-lingual Program-Aided Language Models (Cross-PAL), a method for aligning reasoning programs across languages.Our method delivers programs as intermediate reasoning steps in different languages through a double-step cross-lingual prompting mechanism inspired by the Program-Aided approach.Moreover, we introduce Self-consistent Cross-PAL (SCross-PAL) to ensemble different reasoning paths across languages.Our experimental evaluations show that Cross-PAL outperforms existing methods, reducing the number of interactions and achieving state-of-the-art performance.
Leonardo Ranaldi, Giulia Pucci, Barry Haddow, Alexandra Birch
EMNLP3
2024 Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?
abstract
Traditionally, success in multilingual machine translation can be attributed to three key factors in training data: large volume, diverse translation directions, and high quality.In the current practice of fine-tuning large language models (LLMs) for translation, we revisit the importance of these factors.We find that LLMs display strong translation capability after being fine-tuned on as few as 32 parallel sentences and that fine-tuning on a single translation direction enables translation in multiple directions.However, the choice of direction is critical: fine-tuning LLMs with only English on the target side can lead to task misinterpretation, which hinders translation into non-English languages.Problems also arise when noisy synthetic data is placed on the target side, especially when the target language is wellrepresented in LLM pre-training.Yet interestingly, synthesized data in an under-represented language has a less pronounced effect.Our findings suggest that when adapting LLMs to translation, the requirement on data quantity can be eased but careful considerations are still crucial to prevent an LLM from exploiting unintended data biases.
Pinzhen Chen, Miaoran Zhang, Barry Haddow, Xiaoyu Shen 0001, Dietrich Klakow
EMNLP4
2024 When Does Monolingual Data Help Multilingual Translation: The Role of Domain and Model Scale
abstract
Christos Baziotis, Biao Zhang, Alexandra Birch, Barry Haddow. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Christos Baziotis, Biao Zhang 0006, Alexandra Birch, Barry Haddow
NAACL-HLT4
2024 Assessing Factual Reliability of Large Language Model Knowledge
abstract
Weixuan Wang, Barry Haddow, Alexandra Birch, Wei Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Barry Haddow, Alexandra Birch, Wei Peng 0011
NAACL-HLT2
2023 Self-training Reduces Flicker in Retranslation-based Simultaneous Translation
abstract
In simultaneous translation, the retranslation approach has the advantage of requiring no modifications to the inference engine.However, in order to reduce the undesirable flicker in the output, previous work has resorted to increasing the latency through masking, and introducing specialised inference, thus losing the simplicity of the approach.In this work, we show that self-training improves the flickerlatency tradeoff, while maintaining similar translation quality to the original.Our analysis indicates that self-training reduces flicker by controlling monotonicity.Furthermore, selftraining can be combined with biased beam search to further improve the flicker-latency tradeoff.
Sukanta Sen, Rico Sennrich, Biao Zhang 0006, Barry Haddow
EACL4
2023 Efficient CTC Regularization via Coarse Labels for End-to-End Speech Translation
abstract
For end-to-end speech translation, regularizing the encoder with the Connectionist Temporal Classification (CTC) objective using the source transcript or target translation as labels can greatly improve quality metrics.However, CTC demands an extra prediction layer over the vocabulary space, bringing in nonnegligible model parameters and computational overheads, although this layer is typically not used for inference.In this paper, we re-examine the need for genuine vocabulary labels for CTC for regularization and explore strategies to reduce the CTC label space, targeting improved efficiency without quality degradation.We propose coarse labeling for CTC (CoLaCTC), which merges vocabulary labels via simple heuristic rules, such as using truncation, division or modulo (MOD) operations.Despite its simplicity, our experiments on 4 source and 8 target languages show that CoLaCTC with MOD particularly can compress the label space aggressively to 256 and even further, gaining training efficiency (1.18× ∼ 1.77× speedup depending on the original vocabulary size) yet still delivering comparable or better performance than the CTC baseline.We also show that CoLaCTC successfully generalizes to CTC regularization regardless of using transcript or translation for labeling.
Biao Zhang 0006, Barry Haddow, Rico Sennrich
EACL2
2023 Prompting Large Language Model for Machine Translation: A Case Study
abstract
Research on prompting has shown excellent performance with little or even no supervised training across many tasks. However, prompting for machine translation is still under-explored in the literature. We fill this gap by offering a systematic study on prompting strategies for translation, examining various factors for prompt template and demonstration example selection. We further explore the use of monolingual data and the feasibility of cross-lingual, cross-domain, and sentence-to-document transfer learning in prompting. Extensive experiments with GLM-130B (Zeng et al., 2022) as the testbed show that 1) the number and the quality of prompt examples matter, where using suboptimal examples degenerates translation; 2) several features of prompt examples, such as semantic similarity, show significant Spearman correlation with their prompting performance; yet, none of the correlations are strong enough; 3) using pseudo parallel prompt examples constructed from monolingual data via zero-shot prompting could improve translation; and 4) improved performance is achievable by transferring knowledge from prompt examples selected in other settings. We finally provide an analysis on the model outputs and discuss several problems that prompting still suffers from.
Biao Zhang 0006, Barry Haddow, Alexandra Birch
ICML2
2023 Hallucinations in Large Multilingual Translation Models
abstract
Abstract Hallucinated translations can severely undermine and raise safety issues when machine translation systems are deployed in the wild. Previous research on the topic focused on small bilingual models trained on high-resource languages, leaving a gap in our understanding of hallucinations in multilingual models across diverse translation scenarios. In this work, we fill this gap by conducting a comprehensive analysis—over 100 language pairs across various resource levels and going beyond English-centric directions—on both the M2M neural machine translation (NMT) models and GPT large language models (LLMs). Among several insights, we highlight that models struggle with hallucinations primarily in low-resource directions and when translating out of English, where, critically, they may reveal toxic patterns that can be traced back to the training data. We also find that LLMs produce qualitatively different hallucinations to those of NMT models. Finally, we show that hallucinations are hard to reverse by merely scaling models trained with the same data. However, employing more diverse models, trained on different data or with different procedures, as fallback systems can improve translation quality and virtually eliminate certain pathologies.
Nuno Miguel Guerreiro, Duarte M. Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, André F. T. Martins
Trans. Assoc. Comput. Linguistics4
2022 Revisiting End-to-End Speech-to-Text Translation From Scratch
abstract
End-to-end (E2E) speech-to-text translation (ST) often depends on pretraining its encoder and/or decoder using source transcripts via speech recognition or text translation tasks, without which translation performance drops substantially. However, transcripts are not always available, and how significant such pretraining is for E2E ST has rarely been studied in the literature. In this paper, we revisit this question and explore the extent to which the quality of E2E ST trained on speech-translation pairs alone can be improved. We reexamine several techniques proven beneficial to ST previously, and offer a set of best practices that biases a Transformer-based E2E ST system toward training from scratch. Besides, we propose parameterized distance penalty to facilitate the modeling of locality in the self-attention model for speech. On four benchmarks covering 23 languages, our experiments show that, without using any transcripts or pretraining, the proposed system reaches and even outperforms previous studies adopting pretraining, although the gap remains in (extremely) low-resource settings. Finally, we discuss neural acoustic feature modeling, where a neural model is designed to extract acoustic features from raw speech signals directly, with the goal to simplify inductive biases and add freedom to the model in describing speech. For the first time, we demonstrate its feasibility and show encouraging results on ST tasks.
Biao Zhang 0006, Barry Haddow, Rico Sennrich
ICML2
2022 Non-Autoregressive Machine Translation: It's Not as Fast as it Seems
abstract
Efficient machine translation models are commercially important as they can increase inference speeds, and reduce costs and carbon emissions.Recently, there has been much interest in non-autoregressive (NAR) models, which promise faster translation.In parallel to the research on NAR models, there have been successful attempts to create optimized autoregressive models as part of the WMT shared task on efficient translation.In this paper, we point out flaws in the evaluation methodology present in the literature on NAR models and we provide a fair comparison between a stateof-the-art NAR model and the autoregressive submissions to the shared task.We make the case for consistent evaluation of NAR models, and also for the importance of comparing NAR models with other widely used methods for improving efficiency.We run experiments with a connectionist-temporal-classification-based (CTC) NAR model implemented in C++ and compare it with AR models using wall clock times.Our results show that, although NAR models are faster on GPUs, with small batch sizes, they are almost always slower under more realistic usage conditions.We call for more realistic and extensive evaluation of NAR models in future work.
Jindrich Helcl, Barry Haddow, Alexandra Birch
NAACL-HLT2
2022 Quantifying Synthesis and Fusion and their Impact on Machine Translation
abstract
Arturo Oncevay, Duygu Ataman, Niels Van Berkel, Barry Haddow, Alexandra Birch, Johannes Bjerva. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Arturo Oncevay, Duygu Ataman, Niels van Berkel, Barry Haddow, Alexandra Birch, Johannes Bjerva
NAACL-HLT4
2022 Survey of Low-Resource Machine Translation
abstract
Abstract We present a survey covering the state of the art in low-resource machine translation (MT) research. There are currently around 7,000 languages spoken in the world and almost all language pairs lack significant resources for training machine translation models. There has been increasing interest in research addressing the challenge of producing useful translation models when very little translated training data is available. We present a summary of this topical research field and provide a description of the techniques evaluated by researchers in several recent shared tasks in low-resource MT.
Barry Haddow, Rachel Bawden, Antonio Valerio Miceli Barone, Jindrich Helcl, Alexandra Birch
Comput. Linguistics1
2021 Beyond Sentence-Level End-to-End Speech Translation: Context Helps
abstract
Biao Zhang, Ivan Titov, Barry Haddow, Rico Sennrich. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Biao Zhang 0006, Ivan Titov 0001, Barry Haddow, Rico Sennrich
ACL/IJCNLP (1)3
2021 Surprise Language Challenge: Developing a Neural Machine Translation System between Pashto and English in Two Months
abstract
In the media industry and the focus of global reporting can shift overnight. There is a compelling need to be able to develop new machine translation systems in a short period of time and in order to more efficiently cover quickly developing stories. As part of the EU project GoURMET and which focusses on low-resource machine translation and our media partners selected a surprise language for which a machine translation system had to be built and evaluated in two months(February and March 2021). The language selected was Pashto and an Indo-Iranian language spoken in Afghanistan and Pakistan and India. In this period we completed the full pipeline of development of a neural machine translation system: data crawling and cleaning and aligning and creating test sets and developing and testing models and and delivering them to the user partners. In this paperwe describe rapid data creation and experiments with transfer learning and pretraining for this low-resource language pair. We find that starting from an existing large model pre-trained on 50languages leads to far better BLEU scores than pretraining on one high-resource language pair with a smaller model. We also present human evaluation of our systems and which indicates that the resulting systems perform better than a freely available commercial system when translating from English into Pashto direction and and similarly when translating from Pashto into English.
Alexandra Birch, Barry Haddow, Antonio Valerio Miceli Barone, Jindrich Helcl, Jonas Waldendorf, Felipe Sánchez-Martínez, Mikel L. Forcada, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Miquel Esplà-Gomis, Wilker Aziz, Lina Murady, Sevi Sariisik, Peggy van der Kreeft, Kay Macquarrie
MTSummit (1)2
2020 ParaCrawl: Web-Scale Acquisition of Parallel Corpora
abstract
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, Jaume Zaragoza. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz-Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strong, Brian Thompson 0001, William Waites, Dion Wiggins, Jaume Zaragoza
ACL3
2020 ELITR: European Live Translator
abstract
ELITR (European Live Translator) project aims to create a speech translation system for simultaneous subtitling of conferences and online meetings targetting up to 43 languages. The technology is tested by the Supreme Audit Office of the Czech Republic and by alfaview®, a German online conferencing system. Other project goals are to advance document-level and multilingual machine translation, automatic speech recognition, and automatic minuting.
Ondrej Bojar, Dominik Machácek, Sangeet Sagar, Otakar Smrz, Jonás Kratochvíl, Ebrahim Ansari, Dario Franceschini, Chiara Canton, Ivan Simonini, Thai Son Nguyen, Sebastian Stüker, Alex Waibel, Barry Haddow, Rico Sennrich, Philip Williams
EAMT14
2020 Language Model Prior for Low-Resource Neural Machine Translation
abstract
The scarcity of large parallel corpora is an important obstacle for neural machine translation.A common solution is to exploit the knowledge of language models (LM) trained on abundant monolingual data.In this work, we propose a novel approach to incorporate a LM as prior in a neural translation model (TM).Specifically, we add a regularization term, which pushes the output distributions of the TM to be probable under the LM prior, while avoiding wrong predictions when the TM "disagrees" with the LM.This objective relates to knowledge distillation, where the LM can be viewed as teaching the TM about the target language.The proposed approach does not compromise decoding speed, because the LM is used only at training time, unlike previous work that requires it during inference.We present an analysis on the effects that different methods have on the distributions of the TM.Results on two low-resource machine translation datasets show clear improvements even with limited monolingual data.
Christos Baziotis, Barry Haddow, Alexandra Birch
EMNLP (1)2
2020 Statistical Power and Translationese in Machine Translation Evaluation
abstract
The term translationese has been used to describe features of translated text, and in this paper, we provide detailed analysis of potential adverse effects of translationese on machine translation evaluation.Our analysis shows differences in conclusions drawn from evaluations that include translationese in test data compared to experiments that tested only with text originally composed in that language.For this reason we recommend that reverse-created test data be omitted from future machine translation test sets.In addition, we provide a reevaluation of a past machine translation evaluation claiming human-parity of MT.One important issue not previously considered is statistical power of significance tests applied to comparison of human and machine translation.Since the very aim of past evaluations was the investigation of ties between human and MT systems, power analysis is of particular importance, to avoid, for example, claims of human parity simply corresponding to Type II error resulting from the application of a low powered test.We provide detailed analysis of tests used in such evaluations to provide an indication of a suitable minimum sample size for future studies.
Yvette Graham, Barry Haddow, Philipp Koehn
EMNLP (1)2
2020 Bridging Linguistic Typology and Multilingual Machine Translation with Multi-View Language Representations
abstract
Sparse language vectors from linguistic typology databases and learned embeddings from tasks like multilingual machine translation have been investigated in isolation, without analysing how they could benefit from each other's language characterisation.We propose to fuse both views using singular vector canonical correlation analysis and study what kind of information is induced from each source.By inferring typological features and language phylogenies, we observe that our representations embed typology and strengthen correlations with language relationships.We then take advantage of our multi-view language vector space for multilingual machine translation, where we achieve competitive overall translation accuracy in tasks that require information about language similarities, such as language clustering and ranking candidates for multilingual transfer.With our method, which is also released as a tool, we can easily project and assess new languages without expensive retraining of massive multilingual or ranking models, which are major disadvantages of related approaches.
Arturo Oncevay, Barry Haddow, Alexandra Birch
EMNLP (1)2
2019 Global Under-Resourced Media Translation (GoURMET)
Alexandra Birch, Barry Haddow, Ivan Titov 0001, Antonio Valerio Miceli Barone, Rachel Bawden, Felipe Sánchez-Martínez, Mikel L. Forcada, Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Wilker Aziz, Andrew Secker, Peggy van der Kreeft
MTSummit (2)2
2018 Evaluating Discourse Phenomena in Neural Machine Translation
abstract
Rachel Bawden, Rico Sennrich, Alexandra Birch, Barry Haddow. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Rachel Bawden, Rico Sennrich, Alexandra Birch, Barry Haddow
NAACL-HLT4
2017 Regularization techniques for fine-tuning in neural machine translation
abstract
We investigate techniques for supervised domain adaptation for neural machine translation where an existing model trained on a large out-of-domain dataset is adapted to a small in-domain dataset.In this scenario, overfitting is a major challenge.We investigate a number of techniques to reduce overfitting and improve transfer learning, including regularization techniques such as dropout and L2regularization towards an out-of-domain prior.In addition, we introduce tuneout, a novel regularization technique inspired by dropout.We apply these techniques, alone and in combination, to neural machine translation, obtaining improvements on IWSLT datasets for English→German and English→Russian.We also investigate the amounts of in-domain training data needed for domain adaptation in NMT, and find a logarithmic relationship between the amount of training data and gain in BLEU score.
Antonio Valerio Miceli Barone, Barry Haddow, Ulrich Germann, Rico Sennrich
EMNLP2
2016 Improving Neural Machine Translation Models with Monolingual Data
abstract
Neural Machine Translation (NMT) has obtained state-of-the art performance for several language pairs, while only using parallel data for training.Targetside monolingual data plays an important role in boosting fluency for phrasebased statistical machine translation, and we investigate the use of monolingual data for NMT.In contrast to previous work, which combines NMT models with separately trained language models, we note that encoder-decoder NMT architectures already have the capacity to learn the same information as a language model, and we explore strategies to train with monolingual data without changing the neural network architecture.By pairing monolingual training data with an automatic backtranslation, we can treat it as additional parallel training data, and we obtain substantial improvements on the WMT 15 task English↔German (+2.8-3.7 BLEU), and for the low-resourced IWSLT 14 task Turkish→English (+2.1-3.4BLEU), obtaining new state-of-the-art results.We also show that fine-tuning on in-domain monolingual and parallel data gives substantial improvements for the IWSLT 15 task English→German.
Rico Sennrich, Barry Haddow, Alexandra Birch
ACL (1)2
2016 Neural Machine Translation of Rare Words with Subword Units
abstract
Neural machine translation (NMT) models typically operate with a fixed vocabulary, but translation is an open-vocabulary problem.Previous work addresses the translation of out-of-vocabulary words by backing off to a dictionary.In this paper, we introduce a simpler and more effective approach, making the NMT model capable of open-vocabulary translation by encoding rare and unknown words as sequences of subword units.This is based on the intuition that various word classes are translatable via smaller units than words, for instance names (via character copying or transliteration), compounds (via compositional translation), and cognates and loanwords (via phonological and morphological transformations).We discuss the suitability of different word segmentation techniques, including simple character ngram models and a segmentation based on the byte pair encoding compression algorithm, and empirically show that subword models improve over a back-off dictionary baseline for the WMT 15 translation tasks English→German and English→Russian by up to 1.1 and 1.3 BLEU, respectively.
Rico Sennrich, Barry Haddow, Alexandra Birch
ACL (1)2
2016 HUME: Human UCCA-Based Evaluation of Machine Translation
abstract
Human evaluation of machine translation normally uses sentence-level measures such as relative ranking or adequacy scales. However, these provide no insight into possible errors, and do not scale well with sentence length. We argue for a semantics-based evaluation, which captures what meaning components are retained in the MT output, thus providing a more fine-grained analysis of translation quality, and enabling the construction and tuning of semantics-based MT. We present a novel human semantic evaluation measure, Human UCCA-based MT Evaluation (HUME), building on the UCCA semantic representation scheme. HUME covers a wider range of semantic phenomena than previous methods and does not rely on semantic annotation of the potentially garbled MT output. We experiment with four language pairs, demonstrating HUME’s broad applicability, and report good inter-annotator agreement rates and correlation with human adequacy scores.
Alexandra Birch, Omri Abend, Ondrej Bojar, Barry Haddow
EMNLP4
2016 Controlling Politeness in Neural Machine Translation via Side Constraints
abstract
Many languages use honorifics to express politeness, social distance, or the relative social status between the speaker and their addressee(s). In machine translation from a language without honorifics such as English, it is difficult to predict the appropriate honorific, but users may want to control the level of politeness in the output. In this paper, we perform a pilot study to control honorifics in neural machine translation (NMT) via side constraints, focusing on English!German. We show that by marking up the (English) source side of the training data with a feature that encodes the use of honorifics on the (German) target side, we can control the honorifics produced at test time. Experiments show that the choice of honorifics has a big impact on translation quality as measured by BLEU, and oracle experiments show that substantial improvements are possible by constraining the translation to the desired level of politeness.
Rico Sennrich, Barry Haddow, Alexandra Birch
HLT-NAACL2
2015 HimL (Health in my Language)
Barry Haddow
EAMT1
2015 A Joint Dependency Model of Morphological and Syntactic Structure for Statistical Machine Translation
abstract
When translating between two languages that differ in their degree of morphological synthesis, syntactic structures in one language may be realized as morphological structures in the other, and SMT models need a mechanism to learn such translations.Prior work has used morpheme splitting with flat representations that do not encode the hierarchical structure between morphemes, but this structure is relevant for learning morphosyntactic constraints and selectional preferences.We propose to model syntactic and morphological structure jointly in a dependency translation model, allowing the system to generalize to the level of morphemes.We present a dependency representation of German compounds and particle verbs that results in improvements in translation quality of 1.4-1.8BLEU in the WMT English-German translation task.
Rico Sennrich, Barry Haddow
EMNLP2
2015 Mixed domain vs. multi-domain statistical machine translation
Matthias Huck, Alexandra Birch, Barry Haddow
MTSummit3
2014 Dynamic Topic Adaptation for Phrase-based MT
abstract
Translating text from diverse sources poses a challenge to current machine translation systems which are rarely adapted to structure beyond corpus level. We explore topic adaptation on a diverse data set and present a new bilingual vari-ant of Latent Dirichlet Allocation to com-pute topic-adapted, probabilistic phrase translation features. We dynamically in-fer document-specific translation proba-bilities for test sets of unknown origin, thereby capturing the effects of document context on phrase translations. We show gains of up to 1.26 BLEU over the base-line and 1.04 over a domain adaptation benchmark. We further provide an anal-ysis of the domain-specific data and show additive gains of our model in combination with other types of topic-adapted features. 1
Eva Hasler, Phil Blunsom, Philipp Koehn, Barry Haddow
EACL4
2013 Applying Pairwise Ranked Optimisation to Improve the Interpolation of Translation Models
Barry Haddow
HLT-NAACL1
2010 Monte Carlo techniques for phrase-based translation
Abhishek Arun, Barry Haddow, Philipp Koehn, Adam Lopez, Chris Dyer, Phil Blunsom
Mach. Transl.2
2009 Monte Carlo inference and maximization for phrase-based translation
Abhishek Arun, Chris Dyer, Barry Haddow, Phil Blunsom, Adam Lopez, Philipp Koehn
CoNLL3
2009 Interactive Assistance to Human Translators using Statistical Machine Translation Methods
Philipp Koehn, Barry Haddow
MTSummit2
2008 Exploiting Multiply Annotated Corpora in Biomedical Information Extraction Tasks
Barry Haddow, Beatrice Alex
LREC1