EDBT 2026 Demo / reviewers in the wild / expert
Jörg Tiedemann
dblp:15/670
· DBLP profile ↗
62ranked-venue papers
24as first author
17since 2021 · last 2026
0000-0003-3065-7989ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 24 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Challenge of Finding Robust and Efficient Strategies for Training Machine Translation Models with Noisy DataabstractMost machine translation datasets come with a certain level of noise, and strategies for handling such data need to be robust and efficient. Data selection and filtering are challenging and may depend on expensive language-specific tools that are not necessarily available, especially for low-resource languages. This paper looks at training strategies that combine cheap heuristic filters with curriculum learning to implement iterative procedures that robustly operate on raw noisy data without expensive prior preprocessing and data selection. The intuition is that we can cluster data into buckets with varying noise levels and use different sets of buckets at different stages of MT model training. We test various strategies and compare them to pre-filtering approaches for a diverse set of low-resource languages and conclude that curriculum learning can improve robustness but does not necessarily lead to improved translation performance. Overall, the experiments demonstrate the importance of proper experimental workflows, which cannot easily generalize from one language pair and scenario to another. Mikko Aulamo, Sami Virpioja, Yves Scherrer, Jörg Tiedemann |
EAMT (1) | 4 |
| 2026 | HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained ModelsabstractWe present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder–decoder models, as well as about 30 “smallish” monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation. Stephan Oepen, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Maja Buljan, Laurie Burchell, Lucas Georges Gabriel Charpentier, Pinzhen Chen, Mariia Fedorova, Ona de Gibert Bonet, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Andrey Kutuzov, Veronika Laippala, Bhavitvya Malik, Vladislav Mikhailov, Amanda Myntti, Dayyán O'Brien, Lucie Poláková, Gema Ramírez-Sánchez, Janine Siewert, Pavel Stepachev, Jörg Tiedemann, Teemu Vahtola, Dusan Varis, Fedor Vitiugin, Jaume Zaragoza |
LREC | 25 |
| 2026 | OpenSubtitles2024: A Massively Parallel Dataset of Movie Subtitles for MT Development and Evaluation
Jörg Tiedemann, Hengyu Luo |
LREC | 1 |
| 2025 | An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)abstractLaurie Burchell, Ona De Gibert Bonet, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajič, Jindřich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, Amanda Myntti, Dayyán O’Brien, Stephan Oepen, Proyag Pal, Jousia Piha, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dušan Variš, Tereza Vojtěchová, Jaume Zaragoza-Bernabeu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Laurie Burchell, Ona de Gibert Bonet, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, Amanda Myntti, Dayyán O'Brien, Stephan Oepen, Proyag Pal, Jousia Piha, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dusan Varis, Tereza Vojtechová, Jaume Zaragoza-Bernabeu |
ACL (1) | 32 |
| 2025 | Scaling Low-Resource MT via Synthetic Data Generation with LLMsabstractOna de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo, Zihao Li, Raúl Vázquez, Tiancheng Hu, Jörg Tiedemann. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ona de Gibert Bonet, Joseph Attieh, Teemu Vahtola, Mikko Aulamo, Raúl Vázquez, Tiancheng Hu, Jörg Tiedemann |
EMNLP | 8 |
| 2025 | HPLT's Second Data ReleaseabstractWe describe the progress of the High Performance Language Technologies (HPLT) project, a 3-year EU-funded project that started in September 2022. We focus on the up-to-date results on the release of free text datasets derived from web crawls, one of the central objectives of the project. The second release used a revised processing pipeline, and an enlarged set of input crawls. From 4.5 petabytes of web crawls we extracted 7.6T tokens of monolingual text in 193 languages, plus 380 million parallel sentences in 51 language pairs. We also release MultiHPLT, a cross-combination of the parallel data, which produces 1,275 pairs, as well as releasing the containing documents for all parallel sentences in order to enable research in document-level MT. We report changes in the pipeline, analysis and evaluation results for the second parallel data release based on machine translation systems. All datasets are released under a permissive CC0 licence. Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Laurie Burchell, Pinzhen Chen, Mariia Fedorova, Ona de Gibert Bonet, Liane Guillou, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Erik Henriksson, Andrey Kutuzov, Veronika Laippala, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Amanda Myntti, Dayyán O'Brien, Stephan Oepen, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dusan Varis, Jaume Zaragoza-Bernabeu |
MTSummit (2) | 25 |
| 2024 | A New Massive Multilingual Dataset for High-Performance Language TechnologiesabstractWe present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from CommonCrawl and previously unused web crawls from the Internet Archive. We describe our methods for data acquisition, management and processing of large corpora, which rely on open-source software tools and high-performance computing. Our monolingual collection focuses on low- to medium-resourced languages and covers 75 languages and a total of ≈ 5.6 trillion word tokens de-duplicated on the document level. Our English-centric parallel corpus is derived from its monolingual counterpart and covers 18 language pairs and more than 96 million aligned sentence pairs with roughly 1.4 billion English tokens. The HPLT language resources are one of the largest open text corpora ever released, providing a great resource for language modeling and machine translation training. We publicly release the corpora, the software, and the tools used in this work. Ona de Gibert Bonet, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, Jörg Tiedemann |
LREC/COLING | 13 |
| 2024 | Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?abstractMultilingual pretraining and fine-tuning have remarkably succeeded in various natural language processing tasks. Transferring representations from one language to another is especially crucial for cross-lingual learning. One can expect machine translation objectives to be well suited to fostering such capabilities, as they involve the explicit alignment of semantically equivalent sentences from different languages. This paper investigates the potential benefits of employing machine translation as a continued training objective to enhance language representation learning, bridging multilingual pretraining and cross-lingual applications. We study this question through two lenses: a quantitative evaluation of the performance of existing models and an analysis of their latent representations. Our results show that, contrary to expectations, machine translation as the continued training fails to enhance cross-lingual representation learning in multiple cross-lingual natural language understanding tasks. We conclude that explicit sentence-level alignment in the cross-lingual scenario is detrimental to cross-lingual transfer pretraining, which has important implications for future cross-lingual transfer studies. We furthermore provide evidence through similarity measures and investigation of parameters that this lack of positive influence is due to output separability—which we argue is of use for machine translation but detrimental elsewhere. Shaoxiong Ji, Timothee Mickus, Vincent Segonne, Jörg Tiedemann |
LREC/COLING | 4 |
| 2024 | HPLT's First Release of Data and ModelsabstractThe High Performance Language Technologies (HPLT) project is a 3-year EU-funded project that started in September 2022. It aims to deliver free, sustainable, and reusable datasets, models, and workflows at scale using high-performance computing. We describe the first results of the project. The data release includes monolingual data in 75 languages at 5.6T tokens and parallel data in 18 language pairs at 96M pairs, derived from 1.8 petabytes of web crawls. Building upon automated and transparent pipelines, the first machine translation (MT) models as well as large language models (LLMs) have been trained and released. Multiple data processing tools and pipelines have also been made public. Nikolay Arefyev, Mikko Aulamo, Pinzhen Chen, Ona de Gibert Bonet, Barry Haddow, Jindrich Helcl, Bhavitvya Malik, Gema Ramírez-Sánchez, Pavel Stepachev, Jörg Tiedemann, Dusan Varis, Jaume Zaragoza-Bernabeu |
EAMT (2) | 10 |
| 2024 | A Comparison of Language Modeling and Translation as Multilingual Pretraining ObjectivesabstractPretrained language models (PLMs) display impressive performances and have captured the attention of the NLP community.Establishing best practices in pretraining has, therefore, become a major focus of NLP research, especially since insights gained from monolingual English models may not necessarily apply to more complex multilingual models.One significant caveat of the current state of the art is that different works are rarely comparable: they often discuss different parameter counts, training data, and evaluation methodology.This paper proposes a comparison of multilingual pretraining objectives in a controlled methodological environment.We ensure that training data and model architectures are comparable, and discuss the downstream performances across 6 languages that we observe in probing and fine-tuning scenarios.We make two key observations: (1) the architecture dictates which pretraining objective is optimal;(2) multilingual translation is a very effective pretraining objective under the right conditions.We make our code, data, and model weights available at https://github. com/Helsinki-NLP/lm-vs-mt. Shaoxiong Ji, Timothee Mickus, Vincent Segonne, Jörg Tiedemann |
EMNLP | 5 |
| 2023 | HPLT: High Performance Language TechnologiesabstractWe describe the High Performance Language Technologies project (HPLT), a 3-year EU-funded project started in September 2022. HPLT will build a space combining petabytes of natural language data with large-scale model training. It will derive monolingual and bilingual datasets from the Internet Archive and CommonCrawl and build efficient and solid machine translation (MT) as well as large language models (LLMs). HPLT aims at providing free, sustainable and reusable datasets, models and workflows at scale using high-performance computing (HPC). Mikko Aulamo, Nikolay Bogoychev, Shaoxiong Ji, Graeme Nail, Gema Ramírez-Sánchez, Jörg Tiedemann, Jelmer van der Linde, Jaume Zaragoza |
EAMT | 6 |
| 2023 | Unsupervised Feature Selection for Effective Parallel Corpus FilteringabstractThis work presents an unsupervised method of selecting filters and threshold values for the OpusFilter parallel corpus cleaning toolbox. The method clusters sentence pairs into noisy and clean categories and uses the features of the noisy cluster center as filtering parameters. Our approach utilizes feature importance analysis to disregard filters that do not differentiate between clean and noisy data. A randomly sampled subset of a given corpus is used for filter selection and ineffective filters are not run for the full corpus. We use a set of automatic evaluation metrics to assess the quality of translation models trained with data filtered by our method and data filtered with OpusFilter’s default parameters. The trained models cover English-German and English-Ukrainian in both directions. The proposed method outperforms the default parameters in all translation directions for almost all evaluation metrics. Mikko Aulamo, Ona de Gibert Bonet, Sami Virpioja, Jörg Tiedemann |
EAMT | 4 |
| 2022 | When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and Its IntensityabstractPrerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatically detecting humor in the Friends TV show using multimodal data. Our model is capable of recognizing whether an utterance is humorous or not and assess the intensity of it. We use the prerecorded laughter in the show as annotation as it marks humor and the length of the audience’s laughter tells us how funny a given joke is. We evaluate the model on episodes the model has not been exposed to during the training phase. Our results show that the model is capable of correctly detecting whether an utterance is humorous 78% of the time and how long the audience’s laughter reaction should last with a mean absolute error of 600 milliseconds. Khalid Al-Najjar, Mika Hämäläinen, Jörg Tiedemann, Jorma Laaksonen, Mikko Kurimo |
COLING | 3 |
| 2022 | A Closer Look at Parameter Contributions When Training Neural Language and Translation ModelsabstractWe analyze the learning dynamics of neural language and translation models using Loss Change Allocation (LCA), an indicator that enables a fine-grained analysis of parameter updates when optimizing for the loss function. In other words, we can observe the contributions of different network components at training time. In this article, we systematically study masked language modeling, causal language modeling, and machine translation. We show that the choice of training objective leads to distinctive optimization procedures, even when performed on comparable Transformer architectures. We demonstrate how the various Transformer parameters are used during training, supporting that the feed-forward components of each layer are the main contributors to the optimization procedure. Finally, we find that the learning dynamics are not affected by data size and distribution but rather determined by the learning objective. Raúl Vázquez, Hande Çelikkanat, Vinit Ravishankar, Mathias Creutz, Jörg Tiedemann |
COLING | 5 |
| 2022 | Latest Development in the FoTran Project - Scaling Up Language Coverage in Neural Machine Translation Using Distributed Training with Language-Specific ComponentsabstractWe describe the enhancement of a multilingual NMT toolkit developed as part of the FoTran project. We devise our modular attention-bridge model, which connects language-specific components through a shared network layer. The system now supports distributed training over many nodes and GPUs in order to substantially scale up the number of languages that can be included in a modern neural translation architecture. The model enables the study of emerging language-agnostic representations and also provides a modular toolkit for efficient machine translation. Raúl Vázquez, Michele Boggia, Alessandro Raganato, Niki Andreas Lopi, Stig-Arne Grönroos, Jörg Tiedemann |
EAMT | 6 |
| 2022 | Modeling Noise in Paraphrase DetectionabstractNoisy labels in training data present a challenging issue in classification tasks, misleading a model towards incorrect decisions during training. In this paper, we propose the use of a linear noise model to augment pre-trained language models to account for label noise in fine-tuning. We test our approach in a paraphrase detection task with various levels of noise and five different languages. Our experiments demonstrate the effectiveness of the additional noise model in making the training procedures more robust and stable. Furthermore, we show that this model can be applied without further knowledge about annotation confidence and reliability of individual training examples and we analyse our results in light of data selection and sampling strategies. Teemu Vahtola, Eetu Sjöblom, Jörg Tiedemann, Mathias Creutz |
LREC | 3 |
| 2021 | An Empirical Investigation of Word Alignment Supervision for Zero-Shot Multilingual Neural Machine TranslationabstractZero-shot translations is a fascinating feature of Multilingual Neural Machine Translation (MNMT) systems. These MNMT models are usually trained on English-centric data, i.e. English either as the source or target language, and with a language label prepended to the input indicating the target language. However, recent work has highlighted several flaws of these models in zero-shot scenarios where language labels are ignored and the wrong language is generated or different runs show highly unstable results. In this paper, we investigate the benefits of an explicit alignment to language labels in Transformer-based MNMT models in the zero-shot context, by jointly training one cross attention head with word alignment supervision to stress the focus on the target language label. We compare and evaluate several MNMT systems on three multilingual MT benchmarks of different sizes, showing that simply supervising one cross attention head to focus both on word alignments and language labels reduces the bias towards translating into the wrong language, improving the zero-shot performance overall. Moreover, as an additional advantage, we find that our alignment supervision leads to more stable results across different training runs. Alessandro Raganato, Raúl Vázquez, Mathias Creutz, Jörg Tiedemann |
EMNLP (1) | 4 |
| 2020 | XED: A Multilingual Dataset for Sentiment Analysis and Emotion DetectionabstractWe introduce XED, a multilingual fine-grained emotion dataset.The dataset consists of humanannotated Finnish (25k) and English sentences (30k), as well as projected annotations for 30 additional languages, providing new resources for many low-resource languages.We use Plutchik's core emotions to annotate the dataset with the addition of neutral to create a multilabel multiclass dataset.The dataset is carefully evaluated using language-specific BERT models and SVMs to show that XED performs on par with other similar datasets and is therefore a useful tool for sentiment analysis and emotion detection. Emily Öhman, Marc Pàmies, Kaisla Kajava, Jörg Tiedemann |
COLING | 4 |
| 2020 | MT for subtitling: User evaluation of post-editing productivityabstractThis paper presents a user evaluation of machine translation and post-editing for TV subtitles. Based on a process study where 12 professional subtitlers translated and post-edited subtitles, we compare effort in terms of task time and number of keystrokes. We also discuss examples of specific subtitling features like condensation, and how these features may have affected the post-editing results. In addition to overall MT quality, segmentation and timing of the subtitles are found to be important issues to be addressed in future work. Maarit Koponen, Umut Sulubacak, Kaisa Vitikainen, Jörg Tiedemann |
EAMT | 4 |
| 2020 | OPUS-MT - Building open translation services for the WorldabstractThis paper presents OPUS-MT a project that focuses on the development of free resources and tools for machine translation. The current status is a repository of over 1,000 pre-trained neural machine translation models that are ready to be launched in on-line translation services. For this we also provide open source implementations of web applications that can run efficiently on average desktop hardware with a straightforward setup and installation. Jörg Tiedemann, Santhosh Thottingal |
EAMT | 1 |
| 2020 | OpusTools and Parallel Corpus DiagnosticsabstractThis paper introduces OpusTools, a package for downloading and processing parallel corpora included in the OPUS corpus collection. The package implements tools for accessing compressed data in their archived release format and make it possible to easily convert between common formats. OpusTools also includes tools for language identification and data filtering as well as tools for importing data from various sources into the OPUS format. We show the use of these tools in parallel corpus creation and data diagnostics. The latter is especially useful for the identification of potential problems and errors in the extensive data set. Using these tools, we can now monitor the validity of data sets and improve the overall quality and consitency of the data collection. Mikko Aulamo, Umut Sulubacak, Sami Virpioja, Jörg Tiedemann |
LREC | 4 |
| 2020 | An Evaluation Benchmark for Testing the Word Sense Disambiguation Capabilities of Machine Translation SystemsabstractLexical ambiguity is one of the many challenging linguistic phenomena involved in translation, i.e., translating an ambiguous word with its correct sense. In this respect, previous work has shown that the translation quality of neural machine translation systems can be improved by explicitly modeling the senses of ambiguous words. Recently, several evaluation test sets have been proposed to measure the word sense disambiguation (WSD) capability of machine translation systems. However, to date, these evaluation test sets do not include any training data that would provide a fair setup measuring the sense distributions present within the training data itself. In this paper, we present an evaluation benchmark on WSD for machine translation for 10 language pairs, comprising training data with known sense distributions. Our approach for the construction of the benchmark builds upon the wide-coverage multilingual sense inventory of BabelNet, the multilingual neural parsing pipeline TurkuNLP, and the OPUS collection of translated texts from the web. The test suite is available at http://github.com/Helsinki-NLP/MuCoW. Alessandro Raganato, Yves Scherrer, Jörg Tiedemann |
LREC | 3 |
| 2020 | The FISKMÖ Project: Resources and Tools for Finnish-Swedish Machine Translation and Cross-Linguistic ResearchabstractThis paper presents FISKMÖ, a project that focuses on the development of resources and tools for cross-linguistic research and machine translation between Finnish and Swedish. The goal of the project is the compilation of a massive parallel corpus out of translated material collected from web sources, public and private organisations and language service providers in Finland with its two official languages. The project also aims at the development of open and freely accessible translation services for those two languages for the general purpose and for domain-specific use. We have released new data sets with over 3 million translation units, a benchmark test set for MT development, pre-trained neural MT models with high coverage and competitive performance and a self-contained MT plugin for a popular CAT tool. The latter enables offline translation without dependencies on external services making it possible to work with highly sensitive data without compromising security concerns. Jörg Tiedemann, Tommi Nieminen, Mikko Aulamo, Jenna Kanerva, Akseli Leino, Filip Ginter, Niko Papula |
LREC | 1 |
| 2020 | A Systematic Study of Inner-Attention-Based Sentence Representations in Multilingual Neural Machine TranslationabstractNeural machine translation has considerably improved the quality of automatic translations by learning good representations of input sentences. In this article, we explore a multilingual translation model capable of producing fixed-size sentence representations by incorporating an intermediate crosslingual shared layer, which we refer to as attention bridge. This layer exploits the semantics from each language and develops into a language-agnostic meaning representation that can be efficiently used for transfer learning. We systematically study the impact of the size of the attention bridge and the effect of including additional languages in the model. In contrast to related previous work, we demonstrate that there is no conflict between translation performance and the use of sentence representations in downstream tasks. In particular, we show that larger intermediate layers not only improve translation quality, especially for long sentences, but also push the accuracy of trainable classification tasks. Nevertheless, shorter representations lead to increased compression that is beneficial in non-trainable similarity tasks. Similarly, we show that trainable downstream tasks benefit from multilingual models, whereas additional language signals do not improve performance in non-trainable benchmarks. This is an important insight that helps to properly design models for specific applications. Finally, we also include an in-depth analysis of the proposed attention bridge and its ability to encode linguistic properties. We carefully analyze the information that is captured by individual attention heads and identify interesting patterns that explain the performance of specific settings in linguistic probing tasks. Raúl Vázquez, Alessandro Raganato, Mathias Creutz, Jörg Tiedemann |
Comput. Linguistics | 4 |
| 2020 | Multimodal machine translation through visuals and speechabstractAbstract Multimodal machine translation involves drawing information from more than one modality, based on the assumption that the additional modalities will contain useful alternative views of the input data. The most prominent tasks in this area are spoken language translation, image-guided translation, and video-guided translation, which exploit audio and visual modalities, respectively. These tasks are distinguished from their monolingual counterparts of speech recognition, image captioning, and video captioning by the requirement of models to generate outputs in a different language. This survey reviews the major data resources for these tasks, the evaluation campaigns concentrated around them, the state of the art in end-to-end and pipeline approaches, and also the challenges in performance evaluation. The paper concludes with a discussion of directions for future research in these areas: the need for more expansive and challenging datasets, for targeted evaluations of model performance, and for multimodality in both the input and output space. Umut Sulubacak, Ozan Caglayan, Stig-Arne Grönroos, Aku Rouhe, Desmond Elliott, Lucia Specia, Jörg Tiedemann |
Mach. Transl. | 7 |
| 2019 | What Do Language Representations Really Represent?abstractA neural language model trained on a text corpus can be used to induce distributed representations of words, such that similar words end up with similar representations. If the corpus is multilingual, the same model can be used to learn distributed representations of languages, such that similar languages end up with similar representations. We show that this holds even when the multilingual corpus has been translated into English, by picking up the faint signal left by the source languages. However, just as it is a thorny problem to separate semantic from syntactic similarity in word representations, it is not obvious what type of similarity is captured by language representations. We investigate correlations and causal relationships between language representations learned from translations on one hand, and genetic, geographical, and several levels of structural similarity between languages on the other. Of these, structural similarity is found to correlate most strongly with language representation similarity, whereas genetic relationships—a convenient benchmark used for evaluation in previous work—appears to be a confounding factor. Apart from implications about translation effects, we see this more generally as a case where NLP and linguistic typology can interact and benefit one another. Johannes Bjerva, Robert Östling, Maria Han Veiga, Jörg Tiedemann, Isabelle Augenstein |
Comput. Linguistics | 4 |
| 2019 | Sentence embeddings in NLI with iterative refinement encodersabstractAbstract Sentence-level representations are necessary for various natural language processing tasks. Recurrent neural networks have proven to be very effective in learning distributed representations and can be trained efficiently on natural language inference tasks. We build on top of one such model and propose a hierarchy of bidirectional LSTM and max pooling layers that implements an iterative refinement strategy and yields state of the art results on the SciTail dataset as well as strong results for Stanford Natural Language Inference and Multi-Genre Natural Language Inference. We can show that the sentence embeddings learned in this way can be utilized in a wide variety of transfer learning tasks, outperforming InferSent on 7 out of 10 and SkipThought on 8 out of 9 SentEval sentence embedding evaluation tasks. Furthermore, our model beats the InferSent model in 8 out of 10 recently published SentEval probing tasks designed to evaluate sentence embeddings’ ability to capture some of the important linguistic properties of sentences. Aarne Talman, Anssi Yli-Jyrä, Jörg Tiedemann |
Nat. Lang. Eng. | 3 |
| 2018 | OpenSubtitles2018: Statistical Rescoring of Sentence Alignments in Large, Noisy Parallel Corpora
Pierre Lison, Jörg Tiedemann, Milen Kouylekov |
LREC | 2 |
| 2017 | Character-based Joint Segmentation and POS Tagging for Chinese using Bidirectional RNN-CRFabstractWe present a character-based model for joint segmentation and POS tagging for Chinese. The bidirectional RNN-CRF architecture for general sequence tagging is adapted and applied with novel vector representations of Chinese characters that capture rich contextual information and lower-than-character level features. The proposed model is extensively evaluated and compared with a state-of-the-art tagger respectively on CTB5, CTB9 and UD Chinese. The experimental results indicate that our model is accurate and robust across datasets in different sizes, genres and annotation schemes. We obtain state-of-the-art performance on CTB5, achieving 94.38 F1-score for joint segmentation and POS tagging. Christian Hardmeier, Jörg Tiedemann, Joakim Nivre |
IJCNLP(1) | 3 |
| 2016 | Climbing Mont BLEU: The Strange World of Reachable High-BLEU Translations
Aaron Smith, Christian Hardmeier, Jörg Tiedemann |
EAMT | 3 |
| 2016 | OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles
Pierre Lison, Jörg Tiedemann |
LREC | 2 |
| 2016 | Finding Alternative Translations in a Large Corpus of Movie Subtitle
Jörg Tiedemann |
LREC | 1 |
| 2016 | Synthetic Treebanking for Cross-Lingual Dependency ParsingabstractHow do we parse the languages for which no treebanks are available? This contribution addresses the cross-lingual viewpoint on statistical dependency parsing, in which we attempt to make use of resource-rich source language treebanks to build and adapt models for the under-resourced target languages. We outline the benefits, and indicate the drawbacks of the current major approaches. We emphasize synthetic treebanking: the automatic creation of target language treebanks by means of annotation projection and machine translation. We present competitive results in cross-lingual dependency parsing using a combination of various techniques that contribute to the overall success of the method. We further include a detailed discussion about the impact of part-of-speech label accuracy on parsing results that provide guidance in practical applications of cross-lingual methods for truly under-resourced languages. Jörg Tiedemann, Zeljko Agic |
J. Artif. Intell. Res. | 1 |
| 2014 | Improved Text Extraction from PDF Documents for Large-Scale Natural Language Processing
Jörg Tiedemann |
CICLing (1) | 1 |
| 2014 | Rediscovering Annotation Projection for Cross-Lingual Parser Induction
Jörg Tiedemann |
COLING | 1 |
| 2014 | Treebank Translation for Cross-Lingual Parser InductionabstractCross-lingual learning has become a popular approach to facilitate the development of resources and tools for low density languages. Its underlying idea is to make use of existing tools and annotations in resource-rich languages to create similar tools and resources for resource-poor languages. Typically, this is achieved by either projecting annotations across parallel corpora, or by transferring models from one or more source languages to a target language. In this paper, we explore a third strategy by using machine translation to create synthetic training data from the original source-side annotations. Specifically, we apply this technique to dependency parsing, using a cross-lingually unified treebank for adequate evaluation. Our approach draws on annotation projection but avoids the use of noisy source-side annotation of an unrelated parallel corpus and instead relies on manual treebank annotation in combination with statistical machine translation, which makes it possible to train fully lexicalized parsers. We show that this approach significantly outperforms delexicalized transfer parsing.% despite the error-prone translation step. Jörg Tiedemann, Zeljko Agic, Joakim Nivre |
CoNLL | 1 |
| 2014 | ParCor 1.0: A Parallel Pronoun-Coreference Corpus to Support Statistical MT
Liane Guillou, Christian Hardmeier, Aaron Smith, Jörg Tiedemann, Bonnie L. Webber |
LREC | 4 |
| 2014 | Billions of Parallel Words for Free: Building and Using the EU Bookshop Corpus
Raivis Skadins, Jörg Tiedemann, Roberts Rozis, Daiga Deksne |
LREC | 2 |
| 2013 | Latent Anaphora Resolution for Cross-Lingual Pronoun PredictionabstractThis paper addresses the task of predicting the correct French translations of third-person subject pronouns in English discourse, a problem that is relevant as a prerequisite for machine translation and that requires anaphora resolution.We present an approach based on neural networks that models anaphoric links as latent variables and show that its performance is competitive with that of a system with separate anaphora resolution while not requiring any coreference-annotated training data.This demonstrates that the information contained in parallel bitexts can successfully be used to acquire knowledge about pronominal anaphora in an unsupervised way. Christian Hardmeier, Jörg Tiedemann, Joakim Nivre |
EMNLP | 2 |
| 2013 | Markus Dickinson, Chris Brew and Detmar Meurers: Language and Computers - Wiley-Blackwell, 2013, ISBN: 978-1 4051 8305 5, xviii $$+$$ + 232 pp
Jörg Tiedemann |
Mach. Transl. | 1 |
| 2012 | Efficient Discrimination Between Closely Related Languages
Jörg Tiedemann, Nikola Ljubesic |
COLING | 1 |
| 2012 | Character-Based Pivot Translation for Under-Resourced Languages and Domains
Jörg Tiedemann |
EACL | 1 |
| 2012 | Document-Wide Decoding for Phrase-Based Statistical Machine Translation
Christian Hardmeier, Joakim Nivre, Jörg Tiedemann |
EMNLP-CoNLL | 3 |
| 2012 | Large aligned treebanks for syntax-based machine translation
Gideon Kotzé, Vincent Vandeghinste, Scott Martens, Jörg Tiedemann |
LREC | 4 |
| 2012 | Parallel Data, Tools and Interfaces in OPUS
Jörg Tiedemann |
LREC | 1 |
| 2012 | A Distributed Resource Repository for Cloud-Based Machine Translation
Jörg Tiedemann, Dorte Haltrup Hansen, Lene Offersgaard, Sussi Olsen, Matthias Zumpe |
LREC | 1 |
| 2010 | English to Bangla Phrase-Based Machine Translation
Zahurul Islam, Jörg Tiedemann, Andreas Eisele 0001 |
EAMT | 2 |
| 2010 | Lingua-Align: An Experimental Toolbox for Automatic Tree-to-Tree Alignment
Jörg Tiedemann |
LREC | 1 |
| 2009 | Character-Based PSMT for Closely Related Languages
Jörg Tiedemann |
EAMT | 1 |
| 2009 | Translating Questions for Cross-Lingual QA
Jörg Tiedemann |
EAMT | 1 |
| 2008 | Synchronizing Translated Movie Subtitles
Jörg Tiedemann |
LREC | 1 |
| 2006 | Finding Synonyms Using Automatic Word Alignment and Measures of Distributional Similarity
Lonneke van der Plas, Jörg Tiedemann |
ACL | 2 |
| 2006 | ISA & ICA - Two Web Interfaces for Interactive Alignment of Bitexts alignment of parallel texts
Jörg Tiedemann |
LREC | 1 |
| 2005 | Optimization of word alignment cluesabstractStatistical, linguistic, and heuristic clues can be used for the alignment of words and multi-word units in parallel texts. This article describes the clue alignment approach and the optimization of its parameters using a genetic algorithm. Word alignment clues can come from various sources such as statistical alignment models, co-occurrence tests, string similarity scores and static dictionaries. A genetic algorithm implementing an evolutionary procedure can be used to optimize the parameters necessary for combining available clues. Experiments on English/Swedish bitext show a significant improvement of about 6% in F-scores compared to the baseline produced by statistical word alignment.Most of the work described in this paper was carried out at the Department of Linguistics and Philology at Uppsala University. I would like to acknowledge technical and scientific support by people at the department in Uppsala. Jörg Tiedemann |
Nat. Lang. Eng. | 1 |
| 2004 | Word to word alignment strategies
Jörg Tiedemann |
COLING | 1 |
| 2004 | The OPUS Corpus - Parallel and Free: http: //logos.uio.no/opus
Jörg Tiedemann, Lars Nygaard |
LREC | 1 |
| 2004 | MT Goes Farming: Comparing Two Machine Translation Approaches on a New Domain
Per Weijnitz, Eva Forsbom, Ebba Gustavii, Eva Pettersson, Jörg Tiedemann |
LREC | 5 |
| 2003 | Combining Clues for Word Alignment
Jörg Tiedemann |
EACL | 1 |
| 2002 | Scaling Up an MT Prototype for Industrial Use - Databases and Data Flow
Anna Sågvall Hein, Eva Forsbom, Jörg Tiedemann, Per Weijnitz, Ingrid Almqvist, Leif-Jöran Olsson, Sten Thaning |
LREC | 3 |
| 2002 | MatsLex - a Multilingual Lexical Database for Machine Translation
Jörg Tiedemann |
LREC | 1 |
| 2000 | Evaluation of Word Alignment Systems
Lars Ahrenberg, Magnus Merkel, Anna Sågvall Hein, Jörg Tiedemann |
LREC | 4 |
| 1999 | Automatic Construction of Weighted String Similarity Measures
Jörg Tiedemann |
EMNLP | 1 |