EDBT 2026 Demo / reviewers in the wild / expert
Mikel Artetxe
dblp:168/0354
· DBLP profile ↗
42ranked-venue papers
19as first author
24since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 19 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Emergent Abilities of Large Language Models under Continued Pre-training for Language AdaptationabstractContinued pretraining (CPT) is a popular approach to adapt existing large language models (LLMs) to new languages.When doing so, it is common practice to include a portion of English data in the mixture, but its role has not been carefully studied to date.In this work, we show that including English does not impact validation perplexity, yet it is critical for the emergence of downstream capabilities in the target language.We introduce a language-agnostic benchmark for in-context learning (ICL), which reveals catastrophic forgetting early on CPT when English is not included.This in turn damages the ability of the model to generalize to downstream prompts in the target language as measured by perplexity, even if it does not manifest in terms of accuracy until later in training, and can be tied to a big shift in the model parameters.Based on these insights, we introduce curriculum learning and exponential moving average (EMA) of weights as effective alternatives to mitigate the need for English.All in all, our work sheds light into the dynamics by which emergent abilities arise when doing CPT for language adaptation, and can serve as a foundation to design more effective methods in the future. Ahmed Elhady, Eneko Agirre, Mikel Artetxe |
ACL (1) | 3 |
| 2025 | BOUQuET : dataset, Benchmark and Open initiative for Universal Quality Evaluation in TranslationabstractPierre Andrews, Mikel Artetxe, Mariano Coria Meglioli, Marta R. Costa-jussà, Joe Chuang, David Dale, Mark Duppenthaler, Nathanial Paul Ekberg, Cynthia Gao, Daniel Edward Licht, Jean Maillard, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Eduardo Sánchez, Ioannis Tsiamas, Arina Turkatenko, Albert Ventayol-Boada, Shireen Yates. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Pierre Andrews, Mikel Artetxe, Mariano Coria Meglioli, Marta R. Costa-jussà, Joe Chuang, David Dale, Mark Duppenthaler, Nathanial Paul Ekberg, Cynthia Gao, Daniel Edward Licht, Jean Maillard, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Eduardo Sánchez, Ioannis Tsiamas, Arina Turkatenko, Albert Ventayol-Boada, Shireen Yates |
EMNLP | 2 |
| 2025 | Instructing Large Language Models for Low-Resource Languages: A Systematic Study for BasqueabstractOscar Sainz, Naiara Perez, Julen Etxaniz, Joseba Fernandez de Landa, Itziar Aldabe, Iker García-Ferrero, Aimar Zabala, Ekhi Azurmendi, German Rigau, Eneko Agirre, Mikel Artetxe, Aitor Soroa. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Oscar Sainz, Naiara Pérez, Julen Etxaniz, Joseba Fernandez de Landa, Itziar Aldabe, Iker García-Ferrero, Aimar Zabala, Ekhi Azurmendi, German Rigau, Eneko Agirre, Mikel Artetxe, Aitor Soroa |
EMNLP | 11 |
| 2025 | Linguini: A benchmark for language-agnostic linguistic reasoningabstractWe propose a new benchmark to measure a language model's linguistic reasoning skills without relying on pre-existing language-specific knowledge. The test covers 894 questions grouped in 160 problems across 75 (mostly) extremely low-resource languages, extracted from the International Linguistic Olympiad corpus. To attain high accuracy on this benchmark, models don't need previous knowledge of the tested language, as all the information needed to solve the linguistic puzzle is presented in the context. We find that, while all analyzed models rank below 25% accuracy, there is a significant gap between open and closed models, with the best-performing proprietary model scoring 24.05% and the best-performing open model 8.84%. Eduardo Sánchez, Belen Alastruey, Christophe Ropers, Arina Turkatenko, Pontus Stenetorp, Mikel Artetxe, Marta R. Costa-jussà |
NeurIPS | 6 |
| 2024 | The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsabstractLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal 0001, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa |
ACL (1) | 4 |
| 2024 | Latxa: An Open Language Model and Evaluation Suite for BasqueabstractJulen Etxaniz, Oscar Sainz, Naiara Perez, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Julen Etxaniz, Oscar Sainz, Naiara Miguel, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa |
ACL (1) | 8 |
| 2024 | BertaQA: How Much Do Language Models Know About Local Culture?abstractLarge Language Models (LLMs) exhibit extensive knowledge about the world, but most evaluations have been limited to global or anglocentric subjects. This raises the question of how well these models perform on topics relevant to other cultures, whose presence on the web is not that prominent. To address this gap, we introduce BertaQA, a multiple-choice trivia dataset that is parallel in English and Basque. The dataset consists of a local subset with questions pertinent to the Basque culture, and a global subset with questions of broader interest. We find that state-of-the-art LLMs struggle with local cultural knowledge, even as they excel on global topics. However, we show that continued pre-training in Basque significantly improves the models' performance on Basque culture, even when queried in English. To our knowledge, this is the first solid evidence of knowledge transfer from a low-resource to a high-resource language. Our analysis sheds light on the complex interplay between language and knowledge, and reveals that some prior findings do not fully hold when reassessed on local topics. Our dataset and evaluation code are available under open licenses at https://github.com/juletx/BertaQA. Julen Etxaniz, Gorka Azkune, Aitor Soroa, Oier Lopez de Lacalle, Mikel Artetxe |
NeurIPS | 5 |
| 2023 | Training Trajectories of Language Models Across ScalesabstractMengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen 0001, Luke Zettlemoyer, Veselin Stoyanov |
ACL (1) | 2 |
| 2023 | Revisiting Machine Translation for Cross-lingual ClassificationabstractMachine Translation (MT) has been widely used for cross-lingual classification, either by translating the test set into English and running inference with a monolingual model (translatetest), or translating the training set into the target languages and finetuning a multilingual model (translate-train).However, most research in the area focuses on the multilingual models rather than the MT component.We show that, by using a stronger MT system and mitigating the mismatch between training on original text and running inference on machine translated text, translate-test can do substantially better than previously assumed.The optimal approach, however, is highly task dependent, as we identify various sources of cross-lingual transfer gap that affect different tasks and approaches differently.Our work calls into question the dominance of multilingual models for cross-lingual classification, and prompts to pay more attention to MTbased baselines. Mikel Artetxe, Vedanuj Goswami, Shruti Bhosale, Angela Fan, Luke Zettlemoyer |
EMNLP | 1 |
| 2023 | CombLM: Adapting Black-Box Language Models through Small Fine-Tuned ModelsabstractMethods for adapting language models (LMs) to new tasks and domains have traditionally assumed white-box access to the model, and work by modifying its parameters.However, this is incompatible with a recent trend in the field, where the highest quality models are only available as black-boxes through inference APIs.Even when the model weights are available, the computational cost of fine-tuning large LMs can be prohibitive for most practitioners.In this work, we present a lightweight method for adapting large LMs to new domains and tasks, assuming no access to their weights or intermediate activations.Our approach fine-tunes a small white-box LM and combines it with the large black-box LM at the probability level through a small network, learned on a small validation set.We validate our approach by adapting a large LM (OPT-30B) to several domains and a downstream task (machine translation), observing improved performance in all cases, of up to 9%, while using a domain expert 23x smaller. Aitor Ormazabal, Mikel Artetxe, Eneko Agirre |
EMNLP | 2 |
| 2023 | Improving Language Plasticity via Pretraining with Active ForgettingabstractPretrained language models (PLMs) are today the primary model for natural language processing. Despite their impressive downstream performance, it can be difficult to apply PLMs to new languages, a barrier to making their capabilities universally accessible. While prior work has shown it possible to address this issue by learning a new embedding layer for the new language, doing so is both data and compute inefficient. We propose to use an active forgetting mechanism during pretraining, as a simple way of creating PLMs that can quickly adapt to new languages. Concretely, by resetting the embedding layer every K updates during pretraining, we encourage the PLM to improve its ability of learning new embeddings within limited number of updates, similar to a meta-learning effect. Experiments with RoBERTa show that models pretrained with our forgetting mechanism not only demonstrate faster convergence during language adaptation, but also outperform standard ones in a low-data regime, particularly for languages that are distant from English. Code will be available at https://github.com/facebookresearch/language-model-plasticity. Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani, Pontus Stenetorp, Sebastian Riedel 0001, Mikel Artetxe |
NeurIPS | 7 |
| 2022 | Principled Paraphrase Generation with Parallel CorporaabstractRound-trip Machine Translation (MT) is a popular choice for paraphrase generation, which leverages readily available parallel corpora for supervision.In this paper, we formalize the implicit similarity function induced by this approach, and show that it is susceptible to nonparaphrase pairs sharing a single ambiguous translation.Based on these insights, we design an alternative similarity metric that mitigates this issue by requiring the entire translation distribution to match, and implement a relaxation of it through the Information Bottleneck method.Our approach incorporates an adversarial term into MT training in order to learn representations that encode as much information about the reference translation as possible, while keeping as little information about the input as possible.Paraphrases can be generated by decoding back to the source from this representation, without having to generate pivot translations.In addition to being more principled and efficient than round-trip MT, our approach offers an adjustable parameter to control the fidelity-diversity trade-off, and obtains better results in our experiments. Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre |
ACL (1) | 2 |
| 2022 | Does Corpus Quality Really Matter for Low-Resource Languages?abstractThe vast majority of non-English corpora are derived from automatically filtered versions of CommonCrawl.While prior work has identified major issues on the quality of these datasets (Kreutzer et al., 2021), it is not clear how this impacts downstream performance.Taking representation learning in Basque as a case study, we explore tailored crawling-manually identifying and scraping websites with high-quality content-as an alternative to filtering Common-Crawl.Our new corpus, called EusCrawl, is similar in size to the Basque portion of popular multilingual corpora like CC100 and mC4, yet it has a much higher quality according to native annotators.For instance, 66% of documents are rated as high-quality for EusCrawl, in contrast with < 33% for both mC4 and CC100.Nevertheless, we obtain similar results on downstream NLU tasks regardless of the corpus used for pre-training.Our work suggests that NLU performance in low-resource languages is not primarily constrained by the quality of the data, and other factors like corpus size and domain coverage can play a more important role. Mikel Artetxe, Itziar Aldabe, Rodrigo Agerri, Olatz Perez-de-Viñaspre, Aitor Soroa |
EMNLP | 1 |
| 2022 | Efficient Large Scale Language Modeling with Mixtures of ExpertsabstractMikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giridharan Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O’Horo, Jeffrey Wang, Luke Zettlemoyer, Mona Diab, Zornitsa Kozareva, Veselin Stoyanov. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Mikel Artetxe, Shruti Bhosale, Naman Goyal 0001, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer 0001, Ramakanth Pasunuru, Giri Anantharaman, Xian Li 0003, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Punit Singh Koura, Brian O'Horo, Jeffrey Wang, Luke Zettlemoyer, Mona T. Diab, Zornitsa Kozareva, Veselin Stoyanov |
EMNLP | 1 |
| 2022 | Multilingual Machine Translation with Hyper-AdaptersabstractMultilingual machine translation suffers from negative interference across languages.A common solution is to relax parameter sharing with language-specific modules like adapters.However, adapters of related languages are unable to transfer information, and their total number of parameters becomes prohibitively expensive as the number of languages grows.In this work, we overcome these drawbacks using hyper-adapters-hyper-networks that generate adapters from language and layer embeddings.While past work had poor results when scaling hyper-networks, we propose a rescaling fix that significantly improves convergence and enables training larger hyper-networks.We find that hyper-adapters are more parameter efficient than regular adapters, reaching the same performance with up to 12 times less parameters.When using the same number of parameters and FLOPS, our approach consistently outperforms regular adapters.Also, hyper-adapters converge faster than alternative approaches and scale better than regular dense networks.Our analysis shows that hyperadapters learn to encode language relatedness, enabling positive transfer across languages. Christos Baziotis, Mikel Artetxe, James Cross 0003, Shruti Bhosale |
EMNLP | 2 |
| 2022 | Don't Prompt, Search! Mining-based Zero-Shot Learning with Language ModelsabstractMasked language models like BERT can perform text classification in a zero-shot fashion by reformulating downstream tasks as text infilling.However, this approach is highly sensitive to the template used to prompt the model, yet practitioners are blind when designing them in strict zero-shot settings.In this paper, we propose an alternative mining-based approach for zero-shot learning.Instead of prompting language models, we use regular expressions to mine labeled examples 1 from unlabeled corpora, which can optionally be filtered through prompting, and used to finetune a pretrained model.Our method is more flexible and interpretable than prompting, and outperforms it on a wide range of tasks when using comparable templates.Our results suggest that the success of prompting can partly be explained by the model being exposed to similar examples during pretraining, which can be directly retrieved through regular expressions. Task Lbl VerbalizersSent.Pos.good good good Mozes van de Kar, Mengzhou Xia, Danqi Chen 0001, Mikel Artetxe |
EMNLP | 4 |
| 2022 | Few-shot Learning with Multilingual Generative Language ModelsabstractXi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, Xian Li. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal 0001, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, Veselin Stoyanov, Xian Li 0003 |
EMNLP | 3 |
| 2022 | Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?abstractLarge language models (LMs) are able to incontext learn-perform a new task via inference alone by conditioning on a few input-label pairs (demonstrations) and making predictions for new inputs.However, there has been little understanding of how the model learns and which aspects of the demonstrations contribute to end task performance.In this paper, we show that ground truth demonstrations are in fact not required-randomly replacing labels in the demonstrations barely hurts performance on a range of classification and multi-choce tasks, consistently over 12 different models including GPT-3.Instead, we find that other aspects of the demonstrations are the key drivers of end task performance, including the fact that they provide a few examples of (1) the label space, (2) the distribution of the input text, and (3) the overall format of the sequence.Together, our analysis provides a new way of understanding how and why in-context learning works, while opening up new questions about how much can be learned from large language models through inference alone. Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP | 4 |
| 2022 | Prompting ELECTRA: Few-Shot Learning with Discriminative Pre-Trained ModelsabstractPre-trained masked language models successfully perform few-shot learning by formulating downstream tasks as text infilling.However, as a strong alternative in full-shot settings, discriminative pre-trained models like ELECTRA do not fit into the paradigm.In this work, we adapt prompt-based few-shot learning to ELECTRA and show that it outperforms masked language models in a wide range of tasks.ELECTRA is pre-trained to distinguish if a token is generated or original.We naturally extend that to prompt-based few-shot learning by training to score the originality of the target options without introducing new parameters.Our method can be easily adapted to tasks involving multi-token predictions without extra computation overhead.Analysis shows that ELECTRA learns distributions that align better with downstream tasks. 1 Mengzhou Xia, Mikel Artetxe, Jingfei Du, Danqi Chen 0001, Veselin Stoyanov |
EMNLP | 2 |
| 2022 | Lifting the Curse of Multilinguality by Pre-training Modular TransformersabstractJonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, Mikel Artetxe. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jonas Pfeiffer, Naman Goyal 0001, Xi Victoria Lin, Xian Li 0003, James Cross 0003, Sebastian Riedel 0001, Mikel Artetxe |
NAACL-HLT | 7 |
| 2022 | PARADISE: Exploiting Parallel Data for Multilingual Sequence-to-Sequence PretrainingabstractDespite the success of multilingual sequenceto-sequence pretraining, most existing approaches rely on monolingual corpora, and do not make use of the strong cross-lingual signal contained in parallel data.In this paper, we present PARADISE (PARAllel & Denoising Integration in SEquence-to-sequence models), which extends the conventional denoising objective used to train these models by (i) replacing words in the noised sequence according to a multilingual dictionary, and (ii) predicting the reference translation according to a parallel corpus instead of recovering the original sequence.Our experiments on machine translation and cross-lingual natural language inference show an average improvement of 2.0 BLEU points and 6.7 accuracy points from integrating parallel data into pretraining, respectively, obtaining results that are competitive with several popular models at a fraction of their computational cost. 1 Machel Reid, Mikel Artetxe |
NAACL-HLT | 2 |
| 2022 | Multilingual Autoregressive Entity LinkingabstractAbstract We present mGENRE, a sequence-to- sequence system for the Multilingual Entity Linking (MEL) problem—the task of resolving language-specific mentions to a multilingual Knowledge Base (KB). For a mention in a given language, mGENRE predicts the name of the target entity left-to-right, token-by-token in an autoregressive fashion. The autoregressive formulation allows us to effectively cross-encode mention string and entity names to capture more interactions than the standard dot product between mention and entity vectors. It also enables fast search within a large KB even for mentions that do not appear in mention tables and with no need for large-scale vector indices. While prior MEL works use a single representation for each entity, we match against entity names of as many languages as possible, which allows exploiting language connections between source input and target name. Moreover, in a zero-shot setting on languages with no training data at all, mGENRE treats the target language as a latent variable that is marginalized at prediction time. This leads to over 50% improvements in average accuracy. We show the efficacy of our approach through extensive evaluation including experiments on three popular MEL benchmarks where we establish new state-of-the-art results. Source code available at https://github.com/facebookresearch/GENRE. Nicola De Cao, Ledell Wu, Kashyap Popat, Mikel Artetxe, Naman Goyal 0001, Mikhail Plekhanov, Luke Zettlemoyer, Nicola Cancedda, Sebastian Riedel 0001, Fabio Petroni |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context AnchoringabstractRecent research on cross-lingual word embeddings has been dominated by unsupervised mapping approaches that align monolingual embeddings. Such methods critically rely on those embeddings having a similar structure, but it was recently shown that the separate training in different languages causes departures from this assumption. In this paper, we propose an alternative approach that does not have this limitation, while requiring a weak seed dictionary (e.g., a list of identical words) as the only form of supervision. Rather than aligning two fixed embedding spaces, our method works by fixing the target language embeddings, and learning a new set of embeddings for the source language that are aligned with them. To that end, we use an extension of skip-gram that leverages translated context words as anchor points, and incorporates self-learning and iterative restarts to reduce the dependency on the initial dictionary. Our approach outperforms conventional mapping methods on bilingual lexicon induction, and obtains competitive results in the downstream XNLI task. Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre |
ACL/IJCNLP (1) | 2 |
| 2021 | Multilingual Machine Translation: Closing the Gap between Shared and Language-specific Encoder-DecodersabstractCarlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa, Mikel Artetxe. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa, Mikel Artetxe |
EACL | 4 |
| 2020 | On the Cross-lingual Transferability of Monolingual RepresentationsabstractState-of-the-art unsupervised multilingual models (e.g., multilingual BERT) have been shown to generalize in a zero-shot cross-lingual setting. This generalization ability has been attributed to the use of a shared subword vocabulary and joint training across multiple languages giving rise to deep multilingual abstractions. We evaluate this hypothesis by designing an alternative approach that transfers a monolingual model to new languages at the lexical level. More concretely, we first train a transformer-based masked language model on one language, and transfer it to a new language by learning a new embedding matrix with the same masked language modeling objective, freezing parameters of all other layers. This approach does not rely on a shared vocabulary or joint training. However, we show that it is competitive with multilingual BERT on standard cross-lingual classification benchmarks and on a new Cross-lingual Question Answering Dataset (XQuAD). Our results contradict common beliefs of the basis of the generalization ability of multilingual models and suggest that deep monolingual models learn some abstractions that generalize across languages. We also release XQuAD as a more comprehensive cross-lingual benchmark, which comprises 240 paragraphs and 1190 question-answer pairs from SQuAD v1.1 translated into ten languages by professional translators. Mikel Artetxe, Sebastian Ruder, Dani Yogatama |
ACL | 1 |
| 2020 | A Call for More Rigor in Unsupervised Cross-lingual LearningabstractWe review motivations, definition, approaches, and methodology for unsupervised crosslingual learning and call for a more rigorous position in each of them.An existing rationale for such research is based on the lack of parallel data for many of the world's languages.However, we argue that a scenario without any parallel data and abundant monolingual data is unrealistic in practice.We also discuss different training signals that have been used in previous work, which depart from the pure unsupervised setting.We then describe common methodological issues in tuning and evaluation of unsupervised cross-lingual models and present best practices.Finally, we provide a unified outlook for different types of research in this area (i.e., cross-lingual word embeddings, deep multilingual pretraining, and unsupervised machine translation) and argue for comparable evaluation of these models. Mikel Artetxe, Sebastian Ruder, Dani Yogatama, Gorka Labaka, Eneko Agirre |
ACL | 1 |
| 2020 | Translation Artifacts in Cross-lingual Transfer LearningabstractBoth human and machine translation play a central role in cross-lingual transfer learning: many multilingual datasets have been created through professional translation services, and using machine translation to translate either the test set or the training set is a widely used transfer technique.In this paper, we show that such translation process can introduce subtle artifacts that have a notable impact in existing cross-lingual models.For instance, in natural language inference, translating the premise and the hypothesis independently can reduce the lexical overlap between them, which current models are highly sensitive to.We show that some previous findings in cross-lingual transfer learning need to be reconsidered in the light of this phenomenon.Based on the gained insights, we also improve the state-of-the-art in XNLI for the translate-test and zero-shot approaches by 4.3 and 2.8 points, respectively.3 We use FastAlign (Dyer et al., 2013) for word alignment, and discard the few questions for which the mapping method fails (when none of the tokens in the answer span are aligned).4 We use the same procedure as for the training set except that (i) given the small size of the test set, we combine it with WikiMatrix (Schwenk et al., 2019) to aid word alignment, (ii) we use Jieba for Chinese segmentation instead of the Moses tokenizer, and (iii) for the few unaligned spans, we return the English answer. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
EMNLP (1) | 1 |
| 2020 | Do all roads lead to Rome? Understanding the role of initialization in iterative back-translation
Mikel Artetxe, Gorka Labaka, Noe Casas, Eneko Agirre |
Knowl. Based Syst. | 1 |
| 2019 | An Effective Approach to Unsupervised Machine TranslationabstractWhile machine translation has traditionally relied on large amounts of parallel corpora, a recent research line has managed to train both Neural Machine Translation (NMT) and Statistical Machine Translation (SMT) systems using monolingual corpora only. In this paper, we identify and address several deficiencies of existing unsupervised SMT approaches by exploiting subword information, developing a theoretically well founded unsupervised tuning method, and incorporating a joint refinement procedure. Moreover, we use our improved SMT system to initialize a dual NMT model, which is further fine-tuned through on-the-fly back-translation. Together, we obtain large improvements over the previous state-of-the-art in unsupervised machine translation. For instance, we get 22.5 BLEU points in English-to-German WMT 2014, 5.5 points more than the previous best unsupervised system, and 0.5 points more than the (supervised) shared task winner back in 2014. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
ACL (1) | 1 |
| 2019 | Bilingual Lexicon Induction through Unsupervised Machine TranslationabstractA recent research line has obtained strong results on bilingual lexicon induction by aligning independently trained word embeddings in two languages and using the resulting crosslingual embeddings to induce word translation pairs through nearest neighbor or related retrieval methods.In this paper, we propose an alternative approach to this problem that builds on the recent work on unsupervised machine translation.This way, instead of directly inducing a bilingual lexicon from cross-lingual embeddings, we use them to build a phrasetable, combine it with a language model, and use the resulting machine translation system to generate a synthetic parallel corpus, from which we extract the bilingual lexicon using statistical word alignment techniques.As such, our method can work with any word embedding and cross-lingual mapping technique, and it does not require any additional resource besides the monolingual corpus used to train the embeddings.When evaluated on the exact same cross-lingual embeddings, our proposed method obtains an average improvement of 6 accuracy points over nearest neighbor and 4 points over CSLS retrieval, establishing a new state-of-the-art in the standard MUSE dataset. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
ACL (1) | 1 |
| 2019 | Margin-based Parallel Corpus Mining with Multilingual Sentence EmbeddingsabstractMachine translation is highly sensitive to the size and quality of the training data, which has led to an increasing interest in collecting and filtering large parallel corpora. In this paper, we propose a new method for this task based on multilingual sentence embeddings. In contrast to previous approaches, which rely on nearest neighbor retrieval with a hard threshold over cosine similarity, our proposed method accounts for the scale inconsistencies of this measure, considering the margin between a given sentence pair and its closest candidates instead. Our experiments show large improvements over existing methods. We outperform the best published results on the BUCC mining task and the UN reconstruction task by more than 10 F1 and 30 precision points, respectively. Filtering the English-German ParaCrawl corpus with our approach, we obtain 31.2 BLEU points on newstest2014, an improvement of more than one point over the best official filtered version. Mikel Artetxe, Holger Schwenk |
ACL (1) | 1 |
| 2019 | Analyzing the Limitations of Cross-lingual Word Embedding MappingsabstractRecent research in cross-lingual word embeddings has almost exclusively focused on offline methods, which independently train word embeddings in different languages and map them to a shared space through linear transformations.While several authors have questioned the underlying isomorphism assumption, which states that word embeddings in different languages have approximately the same structure, it is not clear whether this is an inherent limitation of mapping approaches or a more general issue when learning crosslingual embeddings.So as to answer this question, we experiment with parallel corpora, which allows us to compare offline mapping to an extension of skip-gram that jointly learns both embedding spaces.We observe that, under these ideal conditions, joint learning yields to more isomorphic embeddings, is less sensitive to hubness, and obtains stronger results in bilingual lexicon induction.We thus conclude that current mapping methods do have strong limitations, calling for further research to jointly learn cross-lingual embeddings with a weaker cross-lingual signal. Aitor Ormazabal, Mikel Artetxe, Gorka Labaka, Aitor Soroa, Eneko Agirre |
ACL (1) | 2 |
| 2019 | Contextualized Translations of Phrasal Verbs with Distributional Compositional Semantics and Monolingual CorporaabstractThis article describes a compositional distributional method to generate contextualized senses of words and identify their appropriate translations in the target language using monolingual corpora. Word translation is modeled in the same way as contextualization of word meaning, but in a bilingual vector space. The contextualization of meaning is carried out by means of distributional composition within a structured vector space with syntactic dependencies, and the bilingual space is created by means of transfer rules and a bilingual dictionary. A phrase in the source language, consisting of a head and a dependent, is translated into the target language by selecting both the nearest neighbor of the head given the dependent, and the nearest neighbor of the dependent given the head. This process is expanded to larger phrases by means of incremental composition. Experiments were performed on English and Spanish monolingual corpora in order to translate phrasal verbs in context. A new bilingual data set to evaluate strategies aimed at translating phrasal verbs in restricted syntactic domains has been created and released. Pablo Gamallo 0001, Susana Sotelo, José Ramom Pichel Campos, Mikel Artetxe |
Comput. Linguistics | 4 |
| 2019 | Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and BeyondabstractWe introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts. Our system uses a single BiLSTM encoder with a shared BPE vocabulary for all languages, which is coupled with an auxiliary decoder and trained on publicly available parallel corpora. This enables us to learn a classifier on top of the resulting embeddings using English annotated data only, and transfer it to any of the 93 languages without any modification. Our experiments in cross-lingual natural language inference (XNLI dataset), cross-lingual document classification (MLDoc dataset) and parallel corpus mining (BUCC dataset) show the effectiveness of our approach. We also introduce a new test set of aligned sentences in 112 languages, and show that our sentence embeddings obtain strong results in multilingual similarity search even for low-resource languages. Our implementation, the pre-trained encoder and the multilingual test set are available at https://github.com/facebookresearch/LASER Mikel Artetxe, Holger Schwenk |
Trans. Assoc. Comput. Linguistics | 1 |
| 2018 | Generalizing and Improving Bilingual Word Embedding Mappings with a Multi-Step Framework of Linear TransformationsabstractUsing a dictionary to map independently trained word embeddings to a shared space has shown to be an effective approach to learn bilingual word embeddings. In this work, we propose a multi-step framework of linear transformations that generalizes a substantial body of previous work. The core step of the framework is an orthogonal transformation, and existing methods can be explained in terms of the additional normalization, whitening, re-weighting, de-whitening and dimensionality reduction steps. This allows us to gain new insights into the behavior of existing methods, including the effectiveness of inverse regression, and design a novel variant that obtains the best published results in zero-shot bilingual lexicon extraction. The corresponding software is released as an open source project. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
AAAI | 1 |
| 2018 | A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddingsabstractRecent work has managed to learn crosslingual word embeddings without parallel data by mapping monolingual embeddings to a shared space through adversarial training.However, their evaluation has focused on favorable conditions, using comparable corpora or closely-related languages, and we show that they often fail in more realistic scenarios.This work proposes an alternative approach based on a fully unsupervised initialization that explicitly exploits the structural similarity of the embeddings, and a robust self-learning algorithm that iteratively improves this solution.Our method succeeds in all tested scenarios and obtains the best published results in standard datasets, even surpassing previous supervised systems.Our implementation is released as an open source project at https://github. com/artetxem/vecmap. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
ACL (1) | 1 |
| 2018 | Uncovering Divergent Linguistic Information in Word Embeddings with Lessons for Intrinsic and Extrinsic EvaluationabstractFollowing the recent success of word embeddings, it has been argued that there is no such thing as an ideal representation for words, as different models tend to capture divergent and often mutually incompatible aspects like semantics/syntax and similarity/relatedness.In this paper, we show that each embedding model captures more information than directly apparent.A linear transformation that adjusts the similarity order of the model without any external resource can tailor it to achieve better results in those aspects, providing a new perspective on how embeddings encode divergent linguistic information.In addition, we explore the relation between intrinsic and extrinsic evaluation, as the effect of our transformations in downstream tasks is higher for unsupervised systems than for supervised ones. Mikel Artetxe, Gorka Labaka, Iñigo Lopez-Gazpio, Eneko Agirre |
CoNLL | 1 |
| 2018 | Unsupervised Statistical Machine TranslationabstractWhile modern machine translation has relied on large parallel corpora, a recent line of work has managed to train Neural Machine Translation (NMT) systems from monolingual corpora only (Artetxe et al., 2018c;Lample et al., 2018).Despite the potential of this approach for low-resource settings, existing systems are far behind their supervised counterparts, limiting their practical interest.In this paper, we propose an alternative approach based on phrase-based Statistical Machine Translation (SMT) that significantly closes the gap with supervised systems.Our method profits from the modular architecture of SMT: we first induce a phrase table from monolingual corpora through cross-lingual embedding mappings, combine it with an n-gram language model, and fine-tune hyperparameters through an unsupervised MERT variant.In addition, iterative backtranslation improves results further, yielding, for instance, 14.08 and 26.22 BLEU points in WMT 2014 English-German and English-French, respectively, an improvement of more than 7-10 BLEU points over previous unsupervised systems, and closing the gap with supervised SMT (Moses trained on Europarl) down to 2-5 BLEU points.Our implementation is available at https:// github.com/artetxem/monoses. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
EMNLP | 1 |
| 2018 | Unsupervised Neural Machine Translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, Kyunghyun Cho |
ICLR (Poster) | 1 |
| 2017 | Learning bilingual word embeddings with (almost) no bilingual dataabstractMost methods to learn bilingual word embeddings rely on large parallel corpora, which is difficult to obtain for most language pairs.This has motivated an active research line to relax this requirement, with methods that use document-aligned corpora or bilingual dictionaries of a few thousand words instead.In this work, we further reduce the need of bilingual resources using a very simple self-learning approach that can be combined with any dictionary-based mapping technique.Our method exploits the structural similarity of embedding spaces, and works with as little bilingual evidence as a 25 word dictionary or even an automatically generated list of numerals, obtaining results comparable to those of systems that use richer resources. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
ACL (1) | 1 |
| 2016 | Learning principled bilingual mappings of word embeddings while preserving monolingual invarianceabstractMapping word embeddings of different languages into a single space has multiple applications.In order to map from a source space into a target space, a common approach is to learn a linear mapping that minimizes the distances between equivalences listed in a bilingual dictionary.In this paper, we propose a framework that generalizes previous work, provides an efficient exact method to learn the optimal linear transformation and yields the best bilingual results in translation induction while preserving monolingual performance in an analogy task. Mikel Artetxe, Gorka Labaka, Eneko Agirre |
EMNLP | 1 |
| 2015 | Building hybrid machine translation systems by using an EBMT preprocessor to create partial translations
Mikel Artetxe, Gorka Labaka, Kepa Sarasola |
EAMT | 1 |