VLDB 2026 Research / reviewers in the wild / expert
Rico Sennrich
dblp:00/8341
· DBLP profile ↗
76ranked-venue papers
12as first author
34since 2021 · last 2026
0000-0002-1438-4741ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 12 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in TokenizationabstractNegar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus, Sina Ahmadi, Antoine Bosselut, Rico Sennrich. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Negar Foroutan Eghlidi, Clara Meister, Debjit Paul, Joel Niklaus, Sina Ahmadi, Antoine Bosselut, Rico Sennrich |
ACL (1) | 7 |
| 2026 | SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related DocumentsabstractRecognizing semantic differences across documents is crucial for text generation evaluation and content alignment, especially in crosslingual settings.However, as a standalone task, it has received little attention.We address this by introducing SwissGov-RSD, the first naturalistic, document-level, cross-lingual dataset for semantic difference recognition.It encompasses a total of 224 multi-parallel documents in English-German, English-French, and English-Italian with token-level difference annotations by human annotators.We evaluate a variety of open-source and closed-source large language models as well as encoder models across different fine-tuning settings on this new benchmark.Our results show that current automatic approaches perform poorly compared to their performance on monolingual, sentence-level, and synthetic benchmarks, revealing a considerable gap for both LLMs and encoder models.We make our code and dataset publicly available 1 . Michelle Wastl, Jannis Vamvas, Rico Sennrich |
ACL (1) | 3 |
| 2026 | CommonMorph: Participatory Morphological Documentation Platform
Aso Mahmudi, Sina Ahmadi, Kemal Kurniawan, Rico Sennrich, Eduard H. Hovy, Ekaterina Vylomova |
LREC | 4 |
| 2026 | Video and Language Alignment in 2D Systems for 3D Multi-object Scenes with Multi-Information Derivative-Free ControlabstractCross-modal systems trained on 2D visual inputs are presented with a dimensional shift when processing 3D scenes. An in-scene camera bridges the dimensionality gap but requires learning a control module. We introduce a new method that improves multivariate mutual information estimates by regret minimisation with derivative-free optimisation. Our algorithm enables off-the-shelf cross-modal systems trained on 2D visual inputs to adapt online to object occlusions and differentiate features. The pairing of expressive measures and value-based optimisation assists control of an in-scene camera to learn directly from the noisy outputs of vision-language models. The resulting pipeline improves performance in cross-modal tasks on multi-object 3D scenes without resorting to pretraining or finetuning. Jason Armitage, Rico Sennrich |
WACV | 2 |
| 2025 | Leveraging In-Context Learning for Political Bias Testing of LLMsabstractA growing body of work has been querying LLMs with political questions to evaluate their potential biases.However, this probing method has limited stability, making comparisons between models unreliable.In this paper, we argue that LLMs need more context.We propose a new probing task, Questionnaire Modeling (QM), that uses human survey data as incontext examples.We show that QM improves the stability of question-based bias evaluation, and demonstrate that it may be used to compare instruction-tuned models to their base versions.Experiments with LLMs of various sizes indicate that instruction tuning can indeed change the direction of bias.Furthermore, we observe a trend that larger models are able to leverage in-context examples more effectively, and generally exhibit smaller bias scores in QM.Data and code are publicly available.1 Patrick Haller 0001, Jannis Vamvas, Rico Sennrich, Lena A. Jäger |
ACL (1) | 3 |
| 2025 | ConLoan: A Contrastive Multilingual Dataset for Evaluating LoanwordsabstractSina Ahmadi, Micha David Hess, Elena Álvarez-Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang, Andrianos Michail, Peshmerge Morad, Joel Niklaus, Maria Christina Panagiotopoulou, Stefano Perrella, Juri Opitz, Anastassia Shaitarova, Rico Sennrich. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Sina Ahmadi, Micha David Hess, Elena Álvarez Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang, Andrianos Michail, Peshmerge Morad, Joel Niklaus, Maria Christina Panagiotopoulou, Stefano Perrella, Juri Opitz, Anastassia Shaitarova, Rico Sennrich |
ACL (1) | 16 |
| 2025 | PARME: Parallel Corpora for Low-Resourced Middle Eastern LanguagesabstractSina Ahmadi, Rico Sennrich, Erfan Karami, Ako Marani, Parviz Fekrazad, Gholamreza Akbarzadeh Baghban, Hanah Hadi, Semko Heidari, Mahîr Dogan, Pedram Asadi, Dashne Bashir, Mohammad Amin Ghodrati, Kourosh Amini, Zeynab Ashourinezhad, Mana Baladi, Farshid Ezzati, Alireza Ghasemifar, Daryoush Hosseinpour, Behrooz Abbaszadeh, Amin Hassanpour, Bahaddin Jalal Hamaamin, Saya Kamal Hama, Ardeshir Mousavi, Sarko Nazir Hussein, Isar Nejadgholi, Mehmet Ölmez, Horam Osmanpour, Rashid Roshan Ramezani, Aryan Sediq Aziz, Ali Salehi, Mohammadreza Yadegari, Kewyar Yadegari, Sedighe Zamani Roodsari. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Sina Ahmadi, Rico Sennrich, Erfan Karami, Ako Marani, Parviz Fekrazad, Gholamreza Akbarzadeh Baghban, Hanah Hadi, Semko Heidari, Mahîr Dogan, Pedram Asadi, Dashne Bashir, Mohammad Amin Ghodrati, Kourosh Amini, Zeynab Ashourinezhad, Mana Baladi, Farshid Ezzati, Alireza Ghasemifar, Daryoush Hosseinpour, Behrooz Abbaszadeh, Amin Hassanpour, Bahaddin Jalal Hamaamin, Saya Kamal Hama, Ardeshir Mousavi, Sarko Nazir Hussein, Isar Nejadgholi, Mehmet Ölmez, Horam Osmanpour, Rashid Roshan Ramezani, Aryan Sediq Aziz, Ali Salehi, Mohammadreza Yadegari, Kewyar Yadegari, Sedighe Zamani Roodsari |
ACL (1) | 2 |
| 2025 | Measuring the Effect of Disfluency in Multilingual Knowledge Probing BenchmarksabstractFor multilingual factual knowledge assessment of LLMs, benchmarks such as MLAMA use template translations that do not take into account the grammatical and semantic information of the named entities inserted in the sentence.This leads to numerous instances of ungrammaticality or wrong wording of the final prompts, which complicates the interpretation of scores, especially for languages that have a rich morphological inventory.In this work, we sample 4 Slavic languages from the MLAMA dataset and compare the knowledge retrieval scores between the initial (templated) MLAMA dataset and its sentence-level translations made by Google Translate and ChatGPT.We observe a significant increase in knowledge retrieval scores, and provide a qualitative analysis for possible reasons behind it.We also make an additional analysis of 5 more languages from different families and see similar patterns.Therefore, we encourage the community to control the grammaticality of highly multilingual datasets for higher and more interpretable results, which is well approximated by whole sentence translation with neural MT or LLM systems. 1 Kirill Semenov, Rico Sennrich |
EMNLP | 2 |
| 2025 | Robust Native Language Identification through Agentic DecompositionabstractLarge language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual clues such as names, locations, and cultural stereotypes, rather than the underlying linguistic patterns indicative of native language (L1) influence.To improve robustness, previous work has instructed LLMs to disregard such clues.In this work, we demonstrate that such a strategy is unreliable and model predictions can be easily altered by misleading hints.To address this problem, we introduce an agentic NLI pipeline inspired by forensic linguistics, where specialized agents accumulate and categorize diverse linguistic evidence before an independent final overall assessment.In this final assessment, a goal-aware coordinating agent synthesizes all evidence to make the NLI prediction.On two benchmark datasets, our approach significantly enhances NLI robustness against misleading contextual clues and performance consistency compared to standard prompting methods. 1 Ahmet Yavuz Uluslu, Tannon Kew, Tilia Ellendorff, Gerold Schneider, Rico Sennrich |
EMNLP | 5 |
| 2025 | Automatic Speech Recognition for Low-Resourced Middle Eastern Languages
Razhan Hameed, Sina Ahmadi, Hanah Hadi, Rico Sennrich |
INTERSPEECH | 4 |
| 2025 | Machine Translation Meta Evaluation through Translation Accuracy Challenge SetsabstractAbstract Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgment. However, these results are often obtained by averaging predictions across large test sets without any insights into the strengths and weaknesses of these metrics across different error types. Challenge sets are used to probe specific dimensions of metric behavior but there are very few such datasets and they either focus on a limited number of phenomena or a limited number of language pairs. We introduce ACES, a contrastive challenge set spanning 146 language pairs, aimed at discovering whether metrics can identify 68 translation accuracy errors. These phenomena range from basic alterations at the word/character level to more intricate errors based on discourse and real-world knowledge. We conducted a large-scale study by benchmarking ACES on 47 metrics submitted to the WMT 2022 and WMT 2023 metrics shared tasks. We also measure their sensitivity to a range of linguistic phenomena. We further investigate claims that large language models (LLMs) are effective as MT evaluators, addressing the limitations of previous studies by using a dataset that covers a range of linguistic phenomena and language pairs and includes both low- and medium-resource languages. Our results demonstrate that different metric families struggle with different phenomena and that LLM-based methods are unreliable. We expose a number of major flaws with existing methods: Most metrics ignore the source sentence; metrics tend to prefer surface level overlap; and over-reliance on language-agnostic representations leads to confusion when the target language is similar to the source language. To further encourage detailed evaluation beyond singular scores, we expand ACES to include error span annotations, denoted as SPAN-ACES, and we use this dataset to evaluate span-based error metrics, showing that these metrics also need considerable improvement. Based on our observations, we provide a set of recommendations for building better MT metrics, including focusing on error labels instead of scores, ensembling, designing metrics to explicitly focus on the source sentence, focusing on semantic content rather than relying on the lexical overlap, and choosing the right pre-trained model for obtaining representations. Nikita Moghe, Arnisa Fazla, Chantal Amrhein, Tom Kocmi, Mark Steedman, Alexandra Birch, Rico Sennrich, Liane Guillou |
Comput. Linguistics | 7 |
| 2024 | SwissSLi: The Multi-parallel Sign Language Corpus for SwitzerlandabstractIn this work, we introduce SwissSLi, the first sign language corpus that contains parallel data of all three Swiss sign languages, namely Swiss German Sign Language (DSGS), French Sign Language of Switzerland (LSF-CH), and Italian Sign Language of Switzerland (LIS-CH). The data underlying this corpus originates from television programs in three spoken languages: German, French, and Italian. The programs have for the most part been translated into sign language by deaf translators, resulting in a unique, up to six-way multi-parallel dataset between spoken and sign languages. We describe and release the sign language videos and spoken language subtitles as well as the overall statistics and some derivatives of the raw material. These derived components include cropped videos, pose estimation, phrase/sign-segmented videos, and sentence-segmented subtitles, all of which facilitate downstream tasks such as sign language transcription (glossing) and machine translation. The corpus is publicly available on the SWISSUbase data platform for research purposes only under a CC BY-NC-SA 4.0 license. Zifan Jiang, Anne Göhring, Amit Moryossef, Rico Sennrich, Sarah Ebling |
LREC/COLING | 4 |
| 2024 | SignCLIP: Connecting Text and Sign Language by Contrastive LearningabstractWe present SignCLIP, which re-purposes CLIP (Contrastive Language-Image Pretraining) to project spoken language text and sign language videos, two classes of natural languages of distinct modalities, into the same space.SignCLIP is an efficient method of learning useful visual representations for sign language processing from large-scale, multilingual video-text pairs, without directly optimizing for a specific task or sign language which is often of limited size.We pretrain SignCLIP on Spreadthesign, a prominent sign language dictionary consisting of ∼500 thousand video clips in up to 44 sign languages, and evaluate it with various downstream datasets.SignCLIP discerns in-domain signing with notable text-to-video/video-to-text retrieval accuracy.It also performs competitively for out-of-domain downstream tasks such as isolated sign language recognition upon essential few-shot prompting or fine-tuning.We analyze the latent space formed by the spoken language text and sign language poses, which provides additional linguistic insights.Our code and models are openly available 1 . Zifan Jiang, Gerard Sant, Amit Moryossef, Mathias Müller 0002, Rico Sennrich, Sarah Ebling |
EMNLP | 5 |
| 2023 | Improving the Cross-Lingual Generalisation in Visual Question AnsweringabstractWhile several benefits were realized for multilingual vision-language pretrained models, recent benchmarks across various tasks and languages showed poor cross-lingual generalisation when multilingually pre-trained vision-language models are applied to non-English data, with a large gap between (supervised) English performance and (zero-shot) cross-lingual transfer. In this work, we explore the poor performance of these models on a zero-shot cross-lingual visual question answering (VQA) task, where models are fine-tuned on English visual-question data and evaluated on 7 typologically diverse languages. We improve cross-lingual transfer with three strategies: (1) we introduce a linguistic prior objective to augment the cross-entropy loss with a similarity-based loss to guide the model during training, (2) we learn a task-specific subnetwork that improves cross-lingual generalisation and reduces variance without model modification, (3) we augment training examples using synthetic code-mixing to promote alignment of embeddings between source and target languages. Our experiments on xGQA using the pretrained multilingual multimodal transformers UC2 and M3P demonstrates the consistent effectiveness of the proposed fine-tuning strategy for 7 languages, outperforming existing transfer methods with sparse models. Farhad Nooralahzadeh, Rico Sennrich |
AAAI | 2 |
| 2023 | Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting ModelabstractNatural language generation models reproduce and often amplify the biases present in their training data.Previous research explored using sequence-to-sequence rewriting models to transform biased model outputs (or original texts) into more gender-fair language by creating pseudo training data through linguistic rules.However, this approach is not practical for languages with more complex morphology than English.We hypothesise that creating training data in the reverse direction, i.e. starting from gender-fair text, is easier for morphologically complex languages and show that it matches the performance of state-of-the-art rewriting models for English.To eliminate the rule-based nature of data creation, we instead propose using machine translation models to create gender-biased text from real gender-fair text via round-trip translation.Our approach allows us to train a rewriting model for German without the need for elaborate handcrafted rules.The outputs of this model increased genderfairness as shown in a human evaluation study. 1 Chantal Amrhein, Florian Schottmann, Rico Sennrich, Samuel Läubli |
ACL (1) | 3 |
| 2023 | What's the Meaning of Superhuman Performance in Today's NLU?abstractSimone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajič, Daniel Hershcovich, Eduard Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajic 0001, Daniel Hershcovich, Eduard H. Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli |
ACL (1) | 10 |
| 2023 | Self-training Reduces Flicker in Retranslation-based Simultaneous TranslationabstractIn simultaneous translation, the retranslation approach has the advantage of requiring no modifications to the inference engine.However, in order to reduce the undesirable flicker in the output, previous work has resorted to increasing the latency through masking, and introducing specialised inference, thus losing the simplicity of the approach.In this work, we show that self-training improves the flickerlatency tradeoff, while maintaining similar translation quality to the original.Our analysis indicates that self-training reduces flicker by controlling monotonicity.Furthermore, selftraining can be combined with biased beam search to further improve the flicker-latency tradeoff. Sukanta Sen, Rico Sennrich, Biao Zhang 0006, Barry Haddow |
EACL | 2 |
| 2023 | Efficient CTC Regularization via Coarse Labels for End-to-End Speech TranslationabstractFor end-to-end speech translation, regularizing the encoder with the Connectionist Temporal Classification (CTC) objective using the source transcript or target translation as labels can greatly improve quality metrics.However, CTC demands an extra prediction layer over the vocabulary space, bringing in nonnegligible model parameters and computational overheads, although this layer is typically not used for inference.In this paper, we re-examine the need for genuine vocabulary labels for CTC for regularization and explore strategies to reduce the CTC label space, targeting improved efficiency without quality degradation.We propose coarse labeling for CTC (CoLaCTC), which merges vocabulary labels via simple heuristic rules, such as using truncation, division or modulo (MOD) operations.Despite its simplicity, our experiments on 4 source and 8 target languages show that CoLaCTC with MOD particularly can compress the label space aggressively to 256 and even further, gaining training efficiency (1.18× ∼ 1.77× speedup depending on the original vocabulary size) yet still delivering comparable or better performance than the CTC baseline.We also show that CoLaCTC successfully generalizes to CTC regularization regardless of using transcript or translation for labeling. Biao Zhang 0006, Barry Haddow, Rico Sennrich |
EACL | 3 |
| 2023 | Towards Unsupervised Recognition of Token-level Semantic Differences in Related DocumentsabstractAutomatically highlighting words that cause semantic differences between two documents could be useful for a wide range of applications.We formulate recognizing semantic differences (RSD) as a token-level regression task and study three unsupervised approaches that rely on a masked language model.To assess the approaches, we begin with basic English sentences and gradually move to more complex, cross-lingual document pairs.Our results show that an approach based on word alignment and sentence-level contrastive learning has a robust correlation to gold labels.However, all unsupervised approaches still leave a large margin of improvement.Code to reproduce our experiments is available.1 Jannis Vamvas, Rico Sennrich |
EMNLP | 2 |
| 2023 | SLTUNET: A Simple Unified Model for Sign Language Translation
Biao Zhang 0006, Mathias Müller 0002, Rico Sennrich |
ICLR | 3 |
| 2023 | A Priority Map for Vision-and-Language Navigation with Trajectory Plans and Feature-Location CuesabstractIn a busy city street, a pedestrian surrounded by distractions can pick out a single sign if it is relevant to their route. Artificial agents in outdoor Vision-and-Language Navigation (VLN) are also confronted with detecting supervisory signal on environment features and location in inputs. To boost the prominence of relevant features in transformer-based systems without costly preprocessing and pretraining, we take inspiration from priority maps - a mechanism described in neuropsychological studies. We implement a novel priority map module and pretrain on auxiliary tasks using low-sample datasets with high-level representations of routes and environment-related references to urban features. A hierarchical process of trajectory planning -with subsequent parameterised visual boost filtering on visual inputs and prediction of corresponding textual spans - addresses the core challenge of cross-modal alignment and feature-level localisation. The priority map module is integrated into a feature-location framework that doubles the task completion rates of standalone transformers and attains state-of-the-art performance for transformer-based systems on the Touchdown benchmark for VLN. We release code (https://github.com/JasonArmitage-res/PM-VLN) and data (https://zenodo.org/record/6891965.YtwoS3ZBxD8). Jason Armitage, Leonardo Impett, Rico Sennrich |
WACV | 3 |
| 2022 | Revisiting End-to-End Speech-to-Text Translation From ScratchabstractEnd-to-end (E2E) speech-to-text translation (ST) often depends on pretraining its encoder and/or decoder using source transcripts via speech recognition or text translation tasks, without which translation performance drops substantially. However, transcripts are not always available, and how significant such pretraining is for E2E ST has rarely been studied in the literature. In this paper, we revisit this question and explore the extent to which the quality of E2E ST trained on speech-translation pairs alone can be improved. We reexamine several techniques proven beneficial to ST previously, and offer a set of best practices that biases a Transformer-based E2E ST system toward training from scratch. Besides, we propose parameterized distance penalty to facilitate the modeling of locality in the self-attention model for speech. On four benchmarks covering 23 languages, our experiments show that, without using any transcripts or pretraining, the proposed system reaches and even outperforms previous studies adopting pretraining, although the gap remains in (extremely) low-resource settings. Finally, we discuss neural acoustic feature modeling, where a neural model is designed to extract acoustic features from raw speech signals directly, with the goal to simplify inductive biases and add freedom to the model in describing speech. For the first time, we demonstrate its feasibility and show encouraging results on ST tasks. Biao Zhang 0006, Barry Haddow, Rico Sennrich |
ICML | 3 |
| 2022 | BlonDe: An Automatic Evaluation Metric for Document-level Machine TranslationabstractYuchen Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, Ming Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Tianyu Liu 0004, Shuming Ma, Dongdong Zhang 0001, Jian Yang 0030, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, Ming Zhou 0001 |
NAACL-HLT | 7 |
| 2021 | Understanding the Properties of Minimum Bayes Risk Decoding in Neural Machine TranslationabstractMathias Müller, Rico Sennrich. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mathias Müller 0002, Rico Sennrich |
ACL/IJCNLP (1) | 2 |
| 2021 | Analyzing the Source and Target Contributions to Predictions in Neural Machine TranslationabstractElena Voita, Rico Sennrich, Ivan Titov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Elena Voita, Rico Sennrich, Ivan Titov 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Beyond Sentence-Level End-to-End Speech Translation: Context HelpsabstractBiao Zhang, Ivan Titov, Barry Haddow, Rico Sennrich. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Biao Zhang 0006, Ivan Titov 0001, Barry Haddow, Rico Sennrich |
ACL/IJCNLP (1) | 4 |
| 2021 | Wino-X: Multilingual Winograd Schemas for Commonsense Reasoning and Coreference ResolutionabstractWinograd schemas are a well-established tool for evaluating coreference resolution (CoR) and commonsense reasoning (CSR) capabilities of computational models.So far, schemas remained largely confined to English, limiting their utility in multilingual settings.This work presents Wino-X, a parallel dataset of German, French, and Russian schemas, aligned with their English counterparts.We use this resource to investigate whether neural machine translation (NMT) models can perform CoR that requires commonsense knowledge and whether multilingual language models (MLLMs) are capable of CSR across multiple languages.Our findings show Wino-X to be exceptionally challenging for NMT systems that are prone to undesireable biases and unable to detect disambiguating information.We quantify biases using established statistical methods and define ways to address both of these issues.We furthermore present evidence of active cross-lingual knowledge transfer in MLLMs, whereby fine-tuning models on English schemas yields CSR improvements in other languages.1 Denis Emelin, Rico Sennrich |
EMNLP (1) | 2 |
| 2021 | Vision Matters When It Should: Sanity Checking Multimodal Machine Translation ModelsabstractMultimodal machine translation (MMT) systems have been shown to outperform their textonly neural machine translation (NMT) counterparts when visual context is available.However, recent studies have also shown that the performance of MMT models is only marginally impacted when the associated image is replaced with an unrelated image or noise, which suggests that the visual context might not be exploited by the model at all.We hypothesize that this might be caused by the nature of the commonly used evaluation benchmark, also known as Multi30K, where the translations of image captions were prepared without actually showing the images to human translators.In this paper, we present a qualitative study that examines the role of datasets in stimulating the leverage of visual modality and we propose methods to highlight the importance of visual signals in the datasets which demonstrate improvements in reliance of models on the source images.Our findings suggest the research on effective MMT architectures is currently impaired by the lack of suitable datasets and careful consideration must be taken in creation of future MMT datasets, for which we also provide useful insights.1 Jiaoda Li, Duygu Ataman, Rico Sennrich |
EMNLP (1) | 3 |
| 2021 | Contrastive Conditioning for Assessing Disambiguation in MT: A Case Study of Distilled BiasabstractLexical disambiguation is a major challenge for machine translation systems, especially if some senses of a word are trained less often than others.Identifying patterns of overgeneralization requires evaluation methods that are both reliable and scalable.We propose contrastive conditioning as a reference-free blackbox method for detecting disambiguation errors.Specifically, we score the quality of a translation by conditioning on variants of the source that provide contrastive disambiguation cues.After validating our method, we apply it in a case study to perform a targeted evaluation of sequence-level knowledge distillation.By probing word sense disambiguation and translation of gendered occupation names, we show that distillation-trained models tend to overgeneralize more than other models with a comparable BLEU score.Contrastive conditioning thus highlights a side effect of distillation that is not fully captured by standard evaluation metrics.Code and data to reproduce our findings are publicly available.1 Jannis Vamvas, Rico Sennrich |
EMNLP (1) | 2 |
| 2021 | Language Modeling, Lexical Translation, Reordering: The Training Process of NMT through the Lens of Classical SMTabstractDifferently from the traditional statistical MT that decomposes the translation task into distinct separately learned components, neural machine translation uses a single neural network to model the entire translation process.Despite neural machine translation being defacto standard, it is still not clear how NMT models acquire different competences over the course of training, and how this mirrors the different models in traditional SMT.In this work, we look at the competences related to three core SMT components and find that during training, NMT first focuses on learning targetside language modeling, then improves translation quality approaching word-by-word translation, and finally learns more complicated reordering patterns.We show that this behavior holds for several models and language pairs.Additionally, we explain how such an understanding of the training process can be useful in practice and, as an example, show how it can be used to improve vanilla nonautoregressive neural machine translation by guiding teacher model selection. Elena Voita, Rico Sennrich, Ivan Titov 0001 |
EMNLP (1) | 2 |
| 2021 | Sparse Attention with Linear UnitsabstractRecently, it has been argued that encoderdecoder models can be made more interpretable by replacing the softmax function in the attention with its sparse variants.In this work, we introduce a novel, simple method for achieving sparsity in attention: we replace the softmax activation with a ReLU, and show that sparsity naturally emerges from such a formulation.Training stability is achieved with layer normalization with either a specialized initialization or an additional gating function.Our model, which we call Rectified Linear Attention (ReLA), is easy to implement and more efficient than previously proposed sparse attention mechanisms.We apply ReLA to the Transformer and conduct experiments on five machine translation tasks.ReLA achieves translation performance comparable to several strong baselines, with training and decoding speed similar to that of the vanilla attention.Our analysis shows that ReLA delivers high sparsity rate and head diversity, and the induced cross attention achieves better accuracy with respect to source-target word alignment than recent sparsified softmax-based models.Intriguingly, ReLA heads also learn to attend to nothing (i.e.'switch off') for some queries, which is not possible with sparsified softmax alternatives.1 Biao Zhang 0006, Ivan Titov 0001, Rico Sennrich |
EMNLP (1) | 3 |
| 2021 | Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual Translation
Biao Zhang 0006, Ankur Bapna, Rico Sennrich, Orhan Firat |
ICLR | 3 |
| 2021 | On Biasing Transformer Attention Towards MonotonicityabstractAnnette Rios, Chantal Amrhein, Noëmi Aepli, Rico Sennrich. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Annette Rios, Chantal Amrhein, Noëmi Aepli, Rico Sennrich |
NAACL-HLT | 4 |
| 2021 | Revisiting Negation in Neural Machine TranslationabstractIn this paper, we evaluate the translation of negation both automatically and manually, in English–German (EN–DE) and English– Chinese (EN–ZH). We show that the ability of neural machine translation (NMT) models to translate negation has improved with deeper and more advanced networks, although the performance varies between language pairs and translation directions. The accuracy of manual evaluation in EN→DE, DE→EN, EN→ZH, and ZH→EN is 95.7%, 94.8%, 93.4%, and 91.7%, respectively. In addition, we show that under-translation is the most significant error type in NMT, which contrasts with the more diverse error profile previously observed for statistical machine translation. To better understand the root of the under-translation of negation, we study the model’s information flow and training data. While our information flow analysis does not reveal any deficiencies that could be used to detect or fix the under-translation of negation, we find that negation is often rephrased during training, which could make it more difficult for the model to learn a reliable link between source and target negation. We finally conduct intrinsic analysis and extrinsic probing tasks on negation, showing that NMT models can distinguish negation and non-negation tokens very well and encode a lot of information about negation in hidden states but nevertheless leave room for improvement. Gongbo Tang, Philipp Rönchen, Rico Sennrich, Joakim Nivre |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | In Neural Machine Translation, What Does Transfer Learning Transfer?abstractTransfer learning improves quality for lowresource machine translation, but it is unclear what exactly it transfers.We perform several ablation studies that limit information transfer, then measure the quality impact across three language pairs to gain a black-box understanding of transfer learning.Word embeddings play an important role in transfer learning, particularly if they are properly aligned.Although transfer learning can be performed without embeddings, results are sub-optimal.In contrast, transferring only the embeddings but nothing else yields catastrophic results.We then investigate diagonal alignments with auto-encoders over real languages and randomly generated sequences, finding even randomly generated sequences as parents yield noticeable but smaller gains.Finally, transfer learning can eliminate the need for a warmup phase when training transformer models in high resource language pairs. Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, Rico Sennrich |
ACL | 4 |
| 2020 | On Exposure Bias, Hallucination and Domain Shift in Neural Machine TranslationabstractThe standard training algorithm in neural machine translation (NMT) suffers from exposure bias, and alternative algorithms have been proposed to mitigate this.However, the practical impact of exposure bias is under debate.In this paper, we link exposure bias to another well-known problem in NMT, namely the tendency to generate hallucinations under domain shift.In experiments on three datasets with multiple test domains, we show that exposure bias is partially to blame for hallucinations, and that training with Minimum Risk Training, which avoids exposure bias, can mitigate this.Our analysis explains why exposure bias is more problematic under domain shift, and also links exposure bias to the beam search problem, i.e. performance deterioration with increasing beam size.Our results provide a new justification for methods that reduce exposure bias: even if they do not increase performance on in-domain test sets, they can increase model robustness to domain shift. Chaojun Wang, Rico Sennrich |
ACL | 2 |
| 2020 | Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationabstractMassively multilingual models for neural machine translation (NMT) are theoretically attractive, but often underperform bilingual models and deliver poor zero-shot translations.In this paper, we explore ways to improve them.We argue that multilingual NMT requires stronger modeling capacity to support language pairs with varying typological characteristics, and overcome this bottleneck via language-specific components and deepening NMT architectures.We identify the off-target translation issue (i.e.translating into a wrong target language) as the major source of the inferior zero-shot performance, and propose random online backtranslation to enforce the translation of unseen training language pairs.Experiments on OPUS-100 (a novel multilingual dataset with 100 languages) show that our approach substantially narrows the performance gap with bilingual models in both oneto-many and many-to-many settings, and improves zero-shot performance by ∼10 BLEU, approaching conventional pivot-based methods. 1 Biao Zhang 0006, Philip Williams, Ivan Titov 0001, Rico Sennrich |
ACL | 4 |
| 2020 | Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into EnglishabstractRecent work has shown that deeper character-based neural machine translation (NMT) models can outperform subword-based models.However, it is still unclear what makes deeper character-based models successful.In this paper, we conduct an investigation into pure character-based models in the case of translating Finnish into English, including exploring the ability to learn word senses and morphological inflections and the attention mechanism.We demonstrate that word-level information is distributed over the entire character sequence rather than over a single character, and characters at different positions play different roles in learning linguistic knowledge.In addition, character-based models need more layers to encode word senses which explains why only deeper models outperform subword-based models.The attention distribution pattern shows that separators attract a lot of attention and we explore a sparse word-level attention to enforce character hidden states to capture the full word-level information.Experimental results show that the word-level attention with a single head results in 1.2 BLEU points drop. Gongbo Tang, Rico Sennrich, Joakim Nivre |
COLING | 2 |
| 2020 | ELITR: European Live TranslatorabstractELITR (European Live Translator) project aims to create a speech translation system for simultaneous subtitling of conferences and online meetings targetting up to 43 languages. The technology is tested by the Supreme Audit Office of the Czech Republic and by alfaview®, a German online conferencing system. Other project goals are to advance document-level and multilingual machine translation, automatic speech recognition, and automatic minuting. Ondrej Bojar, Dominik Machácek, Sangeet Sagar, Otakar Smrz, Jonás Kratochvíl, Ebrahim Ansari, Dario Franceschini, Chiara Canton, Ivan Simonini, Thai Son Nguyen, Sebastian Stüker, Alex Waibel, Barry Haddow, Rico Sennrich, Philip Williams |
EAMT | 15 |
| 2020 | Detecting Word Sense Disambiguation Biases in Machine Translation for Model-Agnostic Adversarial AttacksabstractWord sense disambiguation is a well-known source of translation errors in NMT.We posit that some of the incorrect disambiguation choices are due to models' over-reliance on dataset artifacts found in training data, specifically superficial word co-occurrences, rather than a deeper understanding of the source text.We introduce a method for the prediction of disambiguation errors based on statistical data properties, demonstrating its effectiveness across several domains and model types.Moreover, we develop a simple adversarial attack strategy that minimally perturbs sentences in order to elicit disambiguation errors to further probe the robustness of translation models.Our findings indicate that disambiguation robustness varies substantially between domains and that different models trained on the same data are vulnerable to different attacks. 1 Denis Emelin, Ivan Titov 0001, Rico Sennrich |
EMNLP (1) | 3 |
| 2020 | Zero-Shot Crosslingual Sentence SimplificationabstractSentence simplification aims to make sentences easier to read and understand.Recent approaches have shown promising results with encoder-decoder models trained on large amounts of parallel data which often only exists in English.We propose a zero-shot modeling framework which transfers simplification knowledge from English to another language (for which no parallel simplification corpus exists) while generalizing across languages and tasks.A shared transformer encoder constructs language-agnostic representations, with a combination of task-specific encoder layers added on top (e.g., for translation and simplification).Empirical results using both human and automatic metrics show that our approach produces better simplifications than unsupervised and pivot-based methods. Jonathan Mallinson, Rico Sennrich, Mirella Lapata |
EMNLP (1) | 2 |
| 2020 | A Set of Recommendations for Assessing Human-Machine Parity in Language TranslationabstractThe quality of machine translation has increased remarkably over the past years, to the degree that it was found to be indistinguishable from professional human translation in a number of empirical investigations. We reassess Hassan et al.'s 2018 investigation into Chinese to English news translation, showing that the finding of human–machine parity was owed to weaknesses in the evaluation design—which is currently considered best practice in the field. We show that the professional human translations contained significantly fewer errors, and that perceived quality in human evaluation depends on the choice of raters, the availability of linguistic context, and the creation of reference translations. Our results call for revisiting current best practices to assess strong machine translation systems in general and human–machine parity in particular, for which we offer a set of recommendations based on our empirical findings. Samuel Läubli, Sheila Castilho, Graham Neubig, Rico Sennrich, Qinlan Shen, Antonio Toral |
J. Artif. Intell. Res. | 4 |
| 2019 | Revisiting Low-Resource Neural Machine Translation: A Case StudyabstractIt has been shown that the performance of neural machine translation (NMT) drops starkly in low-resource conditions, underperforming phrase-based statistical machine translation (PBSMT) and requiring large amounts of auxiliary data to achieve competitive results.In this paper, we re-assess the validity of these results, arguing that they are the result of lack of system adaptation to low-resource settings.We discuss some pitfalls to be aware of when training low-resource NMT systems, and recent techniques that have shown to be especially helpful in low-resource settings, resulting in a set of best practices for low-resource NMT.In our experiments on German-English with different amounts of IWSLT14 training data, we show that, without the use of any auxiliary monolingual or multilingual data, an optimized NMT system can outperform PBSMT with far less data than previously claimed.We also apply these techniques to a low-resource Korean-English dataset, surpassing previously reported results by 4 BLEU. Rico Sennrich, Biao Zhang 0006 |
ACL (1) | 1 |
| 2019 | When a Good Translation is Wrong in Context: Context-Aware Machine Translation Improves on Deixis, Ellipsis, and Lexical CohesionabstractThough machine translation errors caused by the lack of context beyond one sentence have long been acknowledged, the development of context-aware NMT systems is hampered by several problems.Firstly, standard metrics are not sensitive to improvements in consistency in document-level translations.Secondly, previous work on context-aware NMT assumed that the sentence-aligned parallel data consisted of complete documents while in most practical scenarios such document-level data constitutes only a fraction of the available parallel data.To address the first issue, we perform a human study on an English-Russian subtitles dataset and identify deixis, ellipsis and lexical cohesion as three main sources of inconsistency.We then create test sets targeting these phenomena.To address the second shortcoming, we consider a set-up in which a much larger amount of sentence-level data is available compared to that aligned at the document level.We introduce a model that is suitable for this scenario and demonstrate major gains over a context-agnostic baseline on our new benchmarks without sacrificing performance as measured with BLEU. 1 Elena Voita, Rico Sennrich, Ivan Titov 0001 |
ACL (1) | 2 |
| 2019 | Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be PrunedabstractMulti-head self-attention is a key component of the Transformer, a state-of-the-art architecture for neural machine translation.In this work we evaluate the contribution made by individual attention heads in the encoder to the overall performance of the model and analyze the roles played by them.We find that the most important and confident heads play consistent and often linguistically-interpretable roles.When pruning heads using a method based on stochastic gates and a differentiable relaxation of the L 0 penalty, we observe that specialized heads are last to be pruned.Our novel pruning method removes the vast majority of heads without seriously affecting performance.For example, on the English-Russian WMT dataset, pruning 38 out of 48 encoder heads results in a drop of only 0.15 BLEU. 1 Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, Ivan Titov 0001 |
ACL (1) | 4 |
| 2019 | A Lightweight Recurrent Network for Sequence ModelingabstractRecurrent networks have achieved great success on various sequential tasks with the assistance of complex recurrent units, but suffer from severe computational inefficiency due to weak parallelization.One direction to alleviate this issue is to shift heavy computations outside the recurrence.In this paper, we propose a lightweight recurrent network, or LRN.LRN uses input and forget gates to handle long-range dependencies as well as gradient vanishing and explosion, with all parameterrelated calculations factored outside the recurrence.The recurrence in LRN only manipulates the weight assigned to each token, tightly connecting LRN with self-attention networks.We apply LRN as a drop-in replacement of existing recurrent units in several neural sequential models.Extensive experiments on six NLP tasks show that LRN yields the best running efficiency with little or no loss in model performance.1 Biao Zhang 0006, Rico Sennrich |
ACL (1) | 2 |
| 2019 | Encoders Help You Disambiguate Word Senses in Neural Machine TranslationabstractGongbo Tang, Rico Sennrich, Joakim Nivre. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Gongbo Tang, Rico Sennrich, Joakim Nivre |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Context-Aware Monolingual Repair for Neural Machine TranslationabstractElena Voita, Rico Sennrich, Ivan Titov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Elena Voita, Rico Sennrich, Ivan Titov 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling ObjectivesabstractElena Voita, Rico Sennrich, Ivan Titov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Elena Voita, Rico Sennrich, Ivan Titov 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Improving Deep Transformer with Depth-Scaled Initialization and Merged AttentionabstractBiao Zhang, Ivan Titov, Rico Sennrich. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Biao Zhang 0006, Ivan Titov 0001, Rico Sennrich |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Root Mean Square Layer NormalizationabstractLayer normalization (LayerNorm) has been successfully applied to various deep neural networks to help stabilize training and boost model convergence because of its capability in handling re-centering and re-scaling of both inputs and weight matrix. However, the computational overhead introduced by LayerNorm makes these improvements expensive and significantly slows the underlying network, e.g. RNN in particular. In this paper, we hypothesize that re-centering invariance in LayerNorm is dispensable and propose root mean square layer normalization, or RMSNorm. RMSNorm regularizes the summed inputs to a neuron in one layer according to root mean square (RMS), giving the model re-scaling invariance property and implicit learning rate adaptation ability. RMSNorm is computationally simpler and thus more efficient than LayerNorm. We also present partial RMSNorm, or pRMSNorm where the RMS is estimated from p% of the summed inputs without breaking the above properties. Extensive experiments on several tasks using diverse network architectures show that RMSNorm achieves comparable performance against LayerNorm but reduces the running time by 7%~64% on different models. Source code is available at https://github.com/bzhangGo/rmsnorm. Biao Zhang 0006, Rico Sennrich |
NeurIPS | 2 |
| 2018 | Context-Aware Neural Machine Translation Learns Anaphora ResolutionabstractStandard machine translation systems process sentences in isolation and hence ignore extra-sentential information, even though extended context can both prevent mistakes in ambiguous cases and improve translation coherence.We introduce a context-aware neural machine translation model designed in such way that the flow of information from the extended context to the translation model can be controlled and analyzed.We experiment with an English-Russian subtitles dataset, and observe that much of what is captured by our model deals with improving pronoun translation.We measure correspondences between induced attention distributions and coreference relations and observe that the model implicitly captures anaphora.It is consistent with gains for sentences where pronouns need to be gendered in translation.Beside improvements in anaphoric cases, the model also improves in overall BLEU, both over its context-agnostic version (+0.7) and over simple concatenation of the context and source sentences (+0.6). Elena Voita, Pavel Serdyukov, Rico Sennrich, Ivan Titov 0001 |
ACL (1) | 3 |
| 2018 | Has Machine Translation Achieved Human Parity? A Case for Document-level EvaluationabstractRecent research suggests that neural machine translation achieves parity with professional human translation on the WMT Chinese-English news translation task.We empirically test this claim with alternative evaluation protocols, contrasting the evaluation of single sentences and entire documents.In a pairwise ranking experiment, human raters assessing adequacy and fluency show a stronger preference for human over machine translation when evaluating documents as compared to isolated sentences.Our findings emphasise the need to shift towards document-level evaluation as machine translation improves to the degree that errors which are hard or impossible to spot at the sentence-level become decisive in discriminating quality of different translation outputs. Samuel Läubli, Rico Sennrich, Martin Volk 0001 |
EMNLP | 2 |
| 2018 | Sentence Compression for Arbitrary Languages via Multilingual PivotingabstractIn this paper we advocate the use of bilingual corpora which are abundantly available for training sentence compression models.Our approach borrows much of its machinery from neural machine translation and leverages bilingual pivoting: compressions are obtained by translating a source string into a foreign language and then back-translating it into the source while controlling the translation length.Our model can be trained for any language as long as a bilingual corpus is available and performs arbitrary rewrites without access to compression specific data.We release 1 MOSS, a new parallel Multilingual Compression dataset for English, German, and French which can be used to evaluate compression models across languages and genres. Jonathan Mallinson, Rico Sennrich, Mirella Lapata |
EMNLP | 2 |
| 2018 | Why Self-Attention? A Targeted Evaluation of Neural Machine Translation ArchitecturesabstractRecently, non-recurrent architectures (convolutional, self-attentional) have outperformed RNNs in neural machine translation.CNNs and self-attentional networks can connect distant words via shorter network paths than RNNs, and it has been speculated that this improves their ability to model long-range dependencies.However, this theoretical argument has not been tested empirically, nor have alternative explanations for their strong performance been explored in-depth.We hypothesize that the strong performance of CNNs and self-attentional networks could also be due to their ability to extract semantic features from the source text, and we evaluate RNNs, CNNs and self-attention networks on two tasks: subject-verb agreement (where capturing long-range dependencies is required) and word sense disambiguation (where semantic feature extraction is required).Our experimental results show that: 1) self-attentional networks and CNNs do not outperform RNNs in modeling subject-verb agreement over long distances; 2) self-attentional networks perform distinctly better than RNNs and CNNs on word sense disambiguation. Gongbo Tang, Mathias Müller 0002, Annette Rios, Rico Sennrich |
EMNLP | 4 |
| 2018 | Improving Machine Translation of Educational Content via Crowdsourcing
Maximiliana Behnke, Antonio Valerio Miceli Barone, Rico Sennrich, Vilelmini Sosoni, Thanasis Naskos, Eirini Takoulidou, Maria Stasimioti, Menno van Zaanen, Sheila Castilho, Federico Gaspari, Panayota Georgakopoulou, Valia Kordoni, Markus Egg, Katia Kermanidis |
LREC | 3 |
| 2018 | Evaluating Machine Translation Performance on Chinese Idioms with a Blacklist Method
Yutong Shao, Rico Sennrich, Bonnie L. Webber, Federico Fancellu |
LREC | 2 |
| 2018 | Evaluating Discourse Phenomena in Neural Machine TranslationabstractRachel Bawden, Rico Sennrich, Alexandra Birch, Barry Haddow. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Rachel Bawden, Rico Sennrich, Alexandra Birch, Barry Haddow |
NAACL-HLT | 2 |
| 2018 | Evaluating MT for massive open online courses - A multifaceted comparison between PBSMT and NMT systems
Sheila Castilho, Joss Moorkens, Federico Gaspari, Rico Sennrich, Andy Way, Panayota Georgakopoulou |
Mach. Transl. | 4 |
| 2017 | Paraphrasing Revisited with Neural Machine TranslationabstractRecognizing and generating paraphrases is an important component in many natural language processing applications.A wellestablished technique for automatically extracting paraphrases leverages bilingual corpora to find meaning-equivalent phrases in a single language by "pivoting" over a shared translation in another language.In this paper we revisit bilingual pivoting in the context of neural machine translation and present a paraphrasing model based purely on neural networks.Our model represents paraphrases in a continuous space, estimates the degree of semantic relatedness between text segments of arbitrary length, or generates candidate paraphrases for any source input.Experimental results across tasks and datasets show that neural paraphrases outperform those obtained with conventional phrase-based pivoting approaches. Jonathan Mallinson, Rico Sennrich, Mirella Lapata |
EACL (1) | 2 |
| 2017 | Regularization techniques for fine-tuning in neural machine translationabstractWe investigate techniques for supervised domain adaptation for neural machine translation where an existing model trained on a large out-of-domain dataset is adapted to a small in-domain dataset.In this scenario, overfitting is a major challenge.We investigate a number of techniques to reduce overfitting and improve transfer learning, including regularization techniques such as dropout and L2regularization towards an out-of-domain prior.In addition, we introduce tuneout, a novel regularization technique inspired by dropout.We apply these techniques, alone and in combination, to neural machine translation, obtaining improvements on IWSLT datasets for English→German and English→Russian.We also investigate the amounts of in-domain training data needed for domain adaptation in NMT, and find a logarithmic relationship between the amount of training data and gain in BLEU score. Antonio Valerio Miceli Barone, Barry Haddow, Ulrich Germann, Rico Sennrich |
EMNLP | 4 |
| 2017 | Image Pivoting for Learning Multilingual Multimodal RepresentationsabstractIn this paper we propose a model to learn multimodal multilingual representations for matching images and sentences in different languages, with the aim of advancing multilingual versions of image search and image understanding.Our model learns a common representation for images and their descriptions in two different languages (which need not be parallel) by considering the image as a pivot between two languages.We introduce a new pairwise ranking loss function which can handle both symmetric and asymmetric similarity between the two modalities.We evaluate our models on image-description ranking for German and English, and on semantic textual similarity of image descriptions in English.In both cases we achieve state-of-the-art performance. Spandana Gella, Rico Sennrich, Frank Keller, Mirella Lapata |
EMNLP | 2 |
| 2017 | A Comparative Quality Evaluation of PBSMT and NMT using Professional Translators
Sheila Castilho, Joss Moorkens, Federico Gaspari, Rico Sennrich, Vilelmini Sosoni, Panayota Georgakopoulou, Pintu Lohar, Andy Way, Antonio Valerio Miceli Barone, Maria Gialama |
MTSummit (1) | 4 |
| 2017 | Syntax-Based Statistical Machine Translation Philip Williams, Rico Sennrich, Matt Post, Philipp Koehn (University of Edinburgh, University of Edinburgh, Johns Hopkins University, Johns Hopkins University), edited by Graeme Hirst, volume 33), 2016, xvii+190 pp; paperback, ISBN 978-1-62705-900-8; ebook, ISBN 978-1-62705-502-4; doi: 10.2200/S00716ED1V04Y201604HLT033, $70abstractIn its early development, machine translation adopted rule-based approaches, which can include the use of language syntax. The late 1980s and early 1990s saw the inception of the statistical machine translation (SMT) approach, where translation models can be learned automatically from a parallel corpus rather than created manually by humans. Initial SMT models were word-based and phrase-based, without the use of syntactic knowledge. In phrase-based SMT, a source sentence is first segmented into phrases and then translated phrase-by-phrase with some reordering of the translated phrases in the target sentence. This has posed challenges when translating between two syntactically different languages. Syntax-based SMT approaches take advantage of syntactic knowledge within the framework of SMT. This book provides an introduction to syntax-based SMT approaches. It is a valuable resource for those who are interested in syntax-based SMT.The book consists of seven chapters. There is not an introduction chapter in this book, aside from the preface, which can be considered as a brief introduction. Readers are referred to Koehn (2010) for background knowledge. I think an introduction chapter categorized into sections would have been useful, before proceeding to describe the various models. The first two chapters provide principles applicable across various syntax-based SMT approaches. The next three chapters describe syntax-based SMT decoding in detail; this constitutes half of the book. Selected extended topics are provided in the next chapter, which is followed by a concluding chapter.Chapter 1 describes the models and formalisms applicable to syntax-based SMT. The first section describes the phrasal translation units in phrase-based SMT, its limitations, and how tree structures address the limitations of the phrase-based approach. This explanation is useful as translation units are the key difference between the phrase-based and syntax-based SMT approaches. The next two sections describe the grammar formalisms and the statistical models that define syntax-based SMT. The section that covers the grammar formalisms (i.e., synchronous context-free grammar [SCFG] and synchronous tree-substitution grammar [STSG]), would have been clearer if their differences were presented in a side-by-side illustrating example. The remainder of the chapter discusses different categories of syntax-based SMT approaches and the history of these approaches, which include string-to-string, string-to-tree, tree-to-string, and tree-to-tree SMT approaches. Although the syntax-based translation model in Galley et al. (2006) falls under the string-to-tree category, I wonder why hierarchical phrase-based SMT, or Hiero (Chiang, 2007), is not explicitly put under the string-to-string category, since Hiero also uses “unlabeled hierarchical phrases where there is no representation of linguistic categories.”Chapter 2 focuses on how the statistical framework of a syntax-based SMT approach learns its model from a word-aligned and parsed parallel text. The first section explains how phrase pairs are extracted as translation rules from a word-aligned sentence pair in phrase-based SMT (Koehn, Och, and Marcu, 2003), highlighting the definition of a phrase as a sequence of words and the alignment-consistency property of a phrase pair as defined in Och and Ney (2004). The remainder of the chapter introduces three predominant instantiations of syntax-based models: hierarchical phrase-based SMT (Hiero) (Chiang, 2007), which is a non-labeled syntax-based SMT approach arising from the phrase-based approach; syntax-augmented machine translation (SAMT), which introduces the notion of soft labels while keeping the nonlinguistic phrase notion; and GHKM (Galley et al., 2004), which only extracts translation rules consistent with constituency parse subtrees. This chapter is nicely organized and it is easy to follow the gradual evolution from phrase-based SMT to GHKM.Chapter 3 introduces the decoding formalism in the form of a directed hypergraph, defined as a set of vertices and a set of directed hyperedges. The first section introduces the notion of a weighted parse forest represented in a weighted hypergraph, representing alternative parse trees of a sentence. I found it important to pay careful attention to this section, in order to understand the next section and the following chapters. The next section presents various algorithms on a hypergraph to translate a sentence in a hypergraph representation of possible tree derivations. Overall, I found this chapter to contain many technical details. The last section of this chapter provides historical notes on the sources of these concepts. This chapter needs to be read before the next chapter, which assumes understanding of the concepts introduced in Chapter 3.Chapter 4 describes tree decoding—that is, decoding with the constituency parse tree of a source sentence as its input, focusing on the tree-to-string approach. The first two sections highlight decoding with local and non-local features, where non-local features accommodate n-gram language models and are more complex than local features. The next section is devoted to an in-depth description of a beam search algorithm on the parse tree of a source sentence. The description could have been improved if the running example showed the decoding steps. The next two sections present extensions to the concepts introduced in the earlier part of this chapter, by providing references to more efficient hypergraph operations. The content of this section requires readers who are interested in implementing an efficient tree-based algorithm to go through the cited references. Brief historical notes conclude this chapter nicely, by pointing to relevant materials for further reading.Chapter 5 describes string decoding with a source sentence string as its input. The first two sections describe beam search decoding algorithms in a binary SCFG, namely, a maximum of two non-terminal symbols on the right-hand side of each rule, adopted in Hiero and SAMT. The algorithms covered are a basic algorithm and an optimized algorithm. The complexity comparison between the two is nicely presented here, emphasizing the complexity reduction achieved by algorithm optimization. The handling of non-binary rules is described in the following section, illustrated by GHKM rule extraction. A mid-chapter summary section divides this chapter into two parts: beam search decoding and parsing. The second part describes parsing algorithms in the context of shared-category SCFG, assuming the same set of non-terminal symbols for the left-hand and right-hand sides of a rule, followed by a section extending the algorithm to STSG and distinct-category SCFG. The organization of this chapter is excellent. However, I feel that the inclusion of distinct-category SCFG decoding does not fit well into this chapter, as string decoding in string-to-tree SMT requires no knowledge of the source syntax. The historical notes also do not provide any references of prior work on string decoding using distinct-category SCFG.Chapter 6 contains various selected topics on syntax-based SMT. The first section discusses tree transformations, which make translation rule learning more effective. The description of non-context-free models serves as a prelude to the next section on dependency-based SMT, which covers dependency treelet (equivalent to the tree-to-string approach) and string-to-dependency (equivalent to the string-to-tree approach). The next section focuses on the ability of syntax-based SMT to have a more grammatical output compared with phrase-based SMT, although there is still room for improvement, including the use of unification grammars and semantic properties. Finally, the last section of this chapter explains how MT evaluation benefits from syntax-based SMT principles. Overall, this chapter enriches readers' knowledge beyond basic syntax-based SMT in the earlier chapters. I would also suggest the inclusion of phrase-based decoding approaches that use syntax-based features (Cherry, 2008; Chang et al., 2009).Chapter 7 nicely concludes this book by discussing the comparison between phrase-based and syntax-based SMT approaches and proposing possible future developments of syntax-based SMT. The chapter also highlights that syntax-driven MT predates statistical MT, as I mentioned at the beginning of this review.Overall, I found this book to be a useful reference book for those interested in syntax-based SMT. The book is well organized, which makes it easy for readers to refer to specific aspects of syntax-based SMT. An improvement can be made to the presentation of ideas in this book. Throughout the book, there are many technical keywords, resulting from the complexity of syntax-based SMT. It would be useful to highlight these keywords in a side bar to remind readers that they are important keywords. In addition, although examples are given throughout the book, it would be even more useful to use these examples to illustrate how the algorithms work, so that readers can gain a better understanding of the algorithms. Philip Williams, Rico Sennrich, Matt Post, Philipp Koehn, Graeme Hirst, Christian Hadiwinoto |
Comput. Linguistics | 2 |
| 2016 | Improving Neural Machine Translation Models with Monolingual DataabstractNeural Machine Translation (NMT) has obtained state-of-the art performance for several language pairs, while only using parallel data for training.Targetside monolingual data plays an important role in boosting fluency for phrasebased statistical machine translation, and we investigate the use of monolingual data for NMT.In contrast to previous work, which combines NMT models with separately trained language models, we note that encoder-decoder NMT architectures already have the capacity to learn the same information as a language model, and we explore strategies to train with monolingual data without changing the neural network architecture.By pairing monolingual training data with an automatic backtranslation, we can treat it as additional parallel training data, and we obtain substantial improvements on the WMT 15 task English↔German (+2.8-3.7 BLEU), and for the low-resourced IWSLT 14 task Turkish→English (+2.1-3.4BLEU), obtaining new state-of-the-art results.We also show that fine-tuning on in-domain monolingual and parallel data gives substantial improvements for the IWSLT 15 task English→German. Rico Sennrich, Barry Haddow, Alexandra Birch |
ACL (1) | 1 |
| 2016 | Neural Machine Translation of Rare Words with Subword UnitsabstractNeural machine translation (NMT) models typically operate with a fixed vocabulary, but translation is an open-vocabulary problem.Previous work addresses the translation of out-of-vocabulary words by backing off to a dictionary.In this paper, we introduce a simpler and more effective approach, making the NMT model capable of open-vocabulary translation by encoding rare and unknown words as sequences of subword units.This is based on the intuition that various word classes are translatable via smaller units than words, for instance names (via character copying or transliteration), compounds (via compositional translation), and cognates and loanwords (via phonological and morphological transformations).We discuss the suitability of different word segmentation techniques, including simple character ngram models and a segmentation based on the byte pair encoding compression algorithm, and empirically show that subword models improve over a back-off dictionary baseline for the WMT 15 translation tasks English→German and English→Russian by up to 1.1 and 1.3 BLEU, respectively. Rico Sennrich, Barry Haddow, Alexandra Birch |
ACL (1) | 1 |
| 2016 | Controlling Politeness in Neural Machine Translation via Side ConstraintsabstractMany languages use honorifics to express politeness, social distance, or the relative social status between the speaker and their addressee(s). In machine translation from a language without honorifics such as English, it is difficult to predict the appropriate honorific, but users may want to control the level of politeness in the output. In this paper, we perform a pilot study to control honorifics in neural machine translation (NMT) via side constraints, focusing on English!German. We show that by marking up the (English) source side of the training data with a feature that encodes the use of honorifics on the (German) target side, we can control the honorifics produced at test time. Experiments show that the choice of honorifics has a big impact on translation quality as measured by BLEU, and oracle experiments show that substantial improvements are possible by constraining the translation to the desired level of politeness. Rico Sennrich, Barry Haddow, Alexandra Birch |
HLT-NAACL | 1 |
| 2015 | A Joint Dependency Model of Morphological and Syntactic Structure for Statistical Machine TranslationabstractWhen translating between two languages that differ in their degree of morphological synthesis, syntactic structures in one language may be realized as morphological structures in the other, and SMT models need a mechanism to learn such translations.Prior work has used morpheme splitting with flat representations that do not encode the hierarchical structure between morphemes, but this structure is relevant for learning morphosyntactic constraints and selectional preferences.We propose to model syntactic and morphological structure jointly in a dependency translation model, allowing the system to generalize to the level of morphemes.We present a dependency representation of German compounds and particle verbs that results in improvements in translation quality of 1.4-1.8BLEU in the WMT English-German translation task. Rico Sennrich, Barry Haddow |
EMNLP | 1 |
| 2015 | A tree does not make a well-formed sentence: Improving syntactic string-to-tree statistical machine translation with more linguistic knowledgeabstractSynchronous context-free grammars (SCFGs) can be learned from parallel texts that are annotated with target-side syntax, and can produce translations by building target-side syntactic trees from source strings. Ideally, producing syntactic trees would entail that the translation is grammatically well-formed, but in reality, this is often not the case. Focusing on translation into German, we discuss various ways in which string-to-tree translation models over- or undergeneralise. We show how these problems can be addressed by choosing a suitable parser and modifying its output, by introducing linguistic constraints that enforce morphological agreement and constrain subcategorisation, and by modelling the productive generation of German compounds. Rico Sennrich, Philip Williams, Matthias Huck |
Comput. Speech Lang. | 1 |
| 2015 | Modelling and Optimizing on Syntactic N-Grams for Statistical Machine TranslationabstractThe role of language models in SMT is to promote fluent translation output, but traditional n-gram language models are unable to capture fluency phenomena between distant words, such as some morphological agreement phenomena, subcategorisation, and syntactic collocations with string-level gaps. Syntactic language models have the potential to fill this modelling gap. We propose a language model for dependency structures that is relational rather than configurational and thus particularly suited for languages with a (relatively) free word order. It is trainable with Neural Networks, and not only improves over standard n-gram language models, but also outperforms related syntactic language models. We empirically demonstrate its effectiveness in terms of perplexity and as a feature function in string-to-tree SMT from English to German and Russian. We also show that using a syntactic evaluation metric to tune the log-linear parameters of an SMT system further increases translation quality when coupled with a syntactic language model. Rico Sennrich |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | Handling technical OOVs in SMT
Mark Fishel, Rico Sennrich |
EAMT | 2 |
| 2014 | Zmorge: A German Morphological Lexicon Extracted from Wiktionary
Rico Sennrich, Beat Kunz |
LREC | 1 |
| 2013 | A Multi-Domain Translation Model Framework for Statistical Machine Translation
Rico Sennrich, Holger Schwenk, Walid Aransa |
ACL (1) | 1 |
| 2012 | Perplexity Minimization for Translation Model Domain Adaptation in Statistical Machine Translation
Rico Sennrich |
EACL | 1 |
| 2012 | Mixture-Modeling with Unsupervised Clusters for Domain Adaptation in Statistical Machine Translation
Rico Sennrich |
EAMT | 1 |
| 2011 | Combining Multi-Engine Machine Translation and Online Learning through Dynamic Phrase Tables
Rico Sennrich |
EAMT | 1 |