Gorka Labaka

dblp:63/5247 · DBLP profile ↗
← Back
43ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-4611-2502ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Conditioning LLMs to Generate Code-Switched Text
Maite Heredia, Gorka Labaka, Jeremy Barnes 0001, Aitor Soroa
LREC2
2024 SignON - a Co-creative Machine Translation for Sign and Spoken Languages (end-of-project results, contributions and lessons learned)
abstract
SignON, a 3-year Horizon 20202 project addressing the lack of technology and services for MT between sign languages (SLs) and spoken languages (SpLs) ended in December 2023. SignON was unprecedented. Not only it addressed the wider complexity of the aforementioned problem – from research and development of recognition, translation and synthesis, through development of easy-to-use mobile applications and a cloud-based framework to do the “heavy lifting” as well as to establishing ethical, privacy and inclusivenesspolicies and operation guidelines – but also engaged with the deaf and hard of hearing communities in an effective co-creation approach where these main stakeholders drove the development in the right direction and had the final say.Currently we are witnessing advances in natural language processing for SLs, including MT. SignON was one of the largest projects that contributed to this surge with 17 partners and more than 60 consortium members, working in parallel with other international and European initiatives, such as project EASIER and others.
Dimitar Sht. Shterionov, Vincent Vandeghinste, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Andy Way, Josep Blat, Frankie Picron, Davy Van Landuyt, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Caro Brosens, Jorn Rijckaert, Víctor Ubieto Nogales, Bram Vanroy, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion
EAMT (2)22
2023 What Works When in Context-aware Neural Machine Translation?
abstract
Document-level Machine Translation has emerged as a promising means to enhance automated translation quality, but it is currently unclear how effectively context-aware models use the available context during translation. This paper aims to provide insight into the current state of models based on input concatenation, with an in-depth evaluation on English–German and English–French standard datasets. We notably evaluate the impact of data bias, antecedent part-of-speech, context complexity, and the syntactic function of the elements involved in discursive phenomena. Our experimental results indicate that the selected models do improve the overall translation in context, with varying sensitivity to the different factors we examined. We notably show that the selected context-aware models operate markedly better on regular syntactic configurations involving subject antecedents and pronouns, with degraded performance as the configurations become more dissimilar.
Harritxu Gete, Thierry Etchegoyhen, Gorka Labaka
EAMT3
2023 SignON: Sign Language Translation. Progress and challenges
abstract
SignON (https://signon-project.eu/) is a Horizon 2020 project, running from 2021 until the end of 2023, which addresses the lack of technology and services for the automatic translation between sign languages (SLs) and spoken languages, through an inclusive, human-centric solution, hence contributing to the repertoire of communication media for deaf, hard of hearing (DHH) and hearing individuals. In this paper, we present an update of the status of the project, describing the approaches developed to address the challenges and peculiarities of SL machine translation (SLMT).
Vincent Vandeghinste, Dimitar Sht. Shterionov, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Josep Blat, Frankie Picron, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Jorn Rijckaert, Bram Vanroy, Víctor Ubieto Nogales, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion
EAMT19
2023 Targeted Data Augmentation Improves Context-aware Neural Machine Translation
abstract
Progress in document-level Machine Translation is hindered by the lack of parallel training data that include context information. In this work, we evaluate the potential of data augmentation techniques to circumvent these limitations, showing that significant gains can be achieved via upsampling, similar context sampling and back-translations, targeted on context-relevant data. We apply these methods on standard document-level datasets in English-German and English-French and demonstrate their relevance to improve the translation of contextual phenomena. In particular, we show that relatively small volumes of targeted data augmentation lead to significant improvements over a strong context-concatenation baseline and standard back-translation of document-level data. We also compare the accuracy of the selected methods depending on data volumes or distance to relevant context information, and explore their use in combination.
Harritxu Gete, Thierry Etchegoyhen, Gorka Labaka
MTSummit (1)3
2023 A survey on Sign Language machine translation
abstract
Sign Languages (SLs) are employed by deaf and hard-of-hearing (DHH) people to communicate on a daily basis. However, the communication with hearing people still faces some barriers, mainly because of the scarce knowledge about SLs among hearing people. Hence, tools to allow the communication between users of either sign or spoken languages must be encouraged. A stepping stone in this direction is the research of the sign language translation (SLT) task, which aims to produce a spoken language translation of a sign language video or vice versa. By implementing these types of translators in portable devices, we will make considerable progress towards a barrier-free communication between DHH and hearing people. That is why, in this work, we focus on reviewing the literature on SLT and provide the necessary background about SLs. Besides, we summarise the available datasets and the results found in the literature for one of the most used datasets, the RWTH-PHOENIX-2014T. Moreover, the survey lists the challenges that need to be tackled within the SLT research and also for the adoption of SLT technologies, and proposes future research lines.
Adrián Núñez-Marcos, Olatz Perez-de-Viñaspre, Gorka Labaka
Expert Syst. Appl.3
2022 Principled Paraphrase Generation with Parallel Corpora
abstract
Round-trip Machine Translation (MT) is a popular choice for paraphrase generation, which leverages readily available parallel corpora for supervision.In this paper, we formalize the implicit similarity function induced by this approach, and show that it is susceptible to nonparaphrase pairs sharing a single ambiguous translation.Based on these insights, we design an alternative similarity metric that mitigates this issue by requiring the entire translation distribution to match, and implement a relaxation of it through the Information Bottleneck method.Our approach incorporates an adversarial term into MT training in order to learn representations that encode as much information about the reference translation as possible, while keeping as little information about the input as possible.Paraphrases can be generated by decoding back to the source from this representation, without having to generate pivot translations.In addition to being more principled and efficient than round-trip MT, our approach offers an adjustable parameter to control the fidelity-diversity trade-off, and obtains better results in our experiments.
Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre
ACL (1)4
2022 Comparing and combining tagging with different decoding algorithms for back-translation in NMT: learnings from a low resource scenario
abstract
Back-translation is a well established approach to improve the performance of Neural Machine Translation (NMT) systems when large monolingual corpora of the target language and domain are available. Recently, diverse approaches have been proposed to get better automatic evaluation results of NMT models using back-translation, including the use of sampling instead of beam search as decoding algorithm for creating the synthetic corpus. Alternatively, it has been proposed to append a tag to the back-translated corpus for helping the NMT system to distinguish the synthetic bilingual corpus from the authentic one. However, not all the combinations of the previous approaches have been tested, and thus it is not clear which is the best approach for developing a given NMT system. In this work, we empirically compare and combine existing techniques for back-translation in a real low resource setting: the translation of clinical notes from Basque into Spanish. Apart from automatically evaluating the MT systems, we ask bilingual healthcare workers to perform a human evaluation, and analyze the different synthetic corpora by measuring their lexical diversity (LD). For reproducibility and generalizability, we repeat our experiments for German to English translation using public data. The results suggest that in lower resource scenarios tagging only helps when using sampling for decoding, in contradiction with the previous literature using bigger corpora from the news domain. When fine-tuning with a few thousand bilingual in-domain sentences, one of our proposed method (tagged restricted sampling) obtains the best results both in terms of automatic and human evaluation. We will publish the code upon acceptance.
Xabier Soto, Olatz Perez-de-Viñaspre, Gorka Labaka, Maite Oronoz
EAMT3
2022 Unsupervised Machine Translation in Real-World Scenarios
abstract
In this work, we present the work that has been carried on in the MT4All CEF project and the resources that it has generated by leveraging recent research carried out in the field of unsupervised learning. In the course of the project 18 monolingual corpora for specific domains and languages have been collected, and 12 bilingual dictionaries and translation models have been generated. As part of the research, the unsupervised MT methodology based only on monolingual corpora (Artetxe et al., 2017) has been tested on a variety of languages and domains. Results show that in specialised domains, when there is enough monolingual in-domain data, unsupervised results are comparable to those of general domain supervised translation, and that, at any rate, unsupervised techniques can be used to boost results whenever very little data is available.
Ona de Gibert Bonet, Iakes Goenaga, Jordi Armengol-Estapé, Olatz Perez-de-Viñaspre, Carla Parra Escartín, Marina Sanchez, Marcis Pinnis, Gorka Labaka, Maite Melero
LREC8
2022 TANDO: A Corpus for Document-level Machine Translation
abstract
Document-level Neural Machine Translation aims to increase the quality of neural translation models by taking into account contextual information. Properly modelling information beyond the sentence level can result in improved machine translation output in terms of coherence, cohesion and consistency. Suitable corpora for context-level modelling are necessary to both train and evaluate context-aware systems, but are still relatively scarce. In this work we describe TANDO, a document-level corpus for the under-resourced Basque-Spanish language pair, which we share with the scientific community. The corpus is composed of parallel data from three different domains and has been prepared with context-level information. Additionally, the corpus includes contrastive test sets for fine-grained evaluations of gender and register contextual phenomena on both source and target language sides. To establish the usefulness of the corpus, we trained and evaluated baseline Transformer models and context-aware variants based on context concatenation. Our results indicate that the corpus is suitable for fine-grained evaluation of document-level machine translation systems.
Harritxu Gete, Thierry Etchegoyhen, David Ponce, Gorka Labaka, Nora Aranberri, Ander Corral, Xabier Saralegi, Igor Ellakuria, Maite Martín
LREC4
2021 Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context Anchoring
abstract
Recent research on cross-lingual word embeddings has been dominated by unsupervised mapping approaches that align monolingual embeddings. Such methods critically rely on those embeddings having a similar structure, but it was recently shown that the separate training in different languages causes departures from this assumption. In this paper, we propose an alternative approach that does not have this limitation, while requiring a weak seed dictionary (e.g., a list of identical words) as the only form of supervision. Rather than aligning two fixed embedding spaces, our method works by fixing the target language embeddings, and learning a new set of embeddings for the source language that are aligned with them. To that end, we use an extension of skip-gram that leverages translated context words as anchor points, and incorporates self-learning and iterative restarts to reduce the dependency on the initial dictionary. Our approach outperforms conventional mapping methods on bilingual lexicon induction, and obtains competitive results in the downstream XNLI task.
Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre
ACL/IJCNLP (1)4
2021 Label Verbalization and Entailment for Effective Zero and Few-Shot Relation Extraction
abstract
Relation extraction systems require large amounts of labeled examples which are costly to annotate.In this work we reformulate relation extraction as an entailment task, with simple, hand-made, verbalizations of relations produced in less than 15 minutes per relation.The system relies on a pretrained textual entailment engine which is run as-is (no training examples, zero-shot) or further fine-tuned on labeled examples (few-shot or fully trained).In our experiments on TACRED we attain 63% F1 zero-shot, 69% with 16 examples per relation (17% points better than the best supervised system on the same conditions), and only 4 points short of the state-of-the-art (which uses 20 times more training data).We also show that the performance can be improved significantly with larger entailment models, up to 12 points in zero-shot, giving the best results to date on TACRED when fully trained.The analysis shows that our few-shot systems are especially effective when discriminating between relations, and that the performance difference in low data regimes comes mainly from identifying no-relation cases.
Oscar Sainz, Oier Lopez de Lacalle, Gorka Labaka, Ander Barrena, Eneko Agirre
EMNLP (1)3
2020 A Call for More Rigor in Unsupervised Cross-lingual Learning
abstract
We review motivations, definition, approaches, and methodology for unsupervised crosslingual learning and call for a more rigorous position in each of them.An existing rationale for such research is based on the lack of parallel data for many of the world's languages.However, we argue that a scenario without any parallel data and abundant monolingual data is unrealistic in practice.We also discuss different training signals that have been used in previous work, which depart from the pure unsupervised setting.We then describe common methodological issues in tuning and evaluation of unsupervised cross-lingual models and present best practices.Finally, we provide a unified outlook for different types of research in this area (i.e., cross-lingual word embeddings, deep multilingual pretraining, and unsupervised machine translation) and argue for comparable evaluation of these models.
Mikel Artetxe, Sebastian Ruder, Dani Yogatama, Gorka Labaka, Eneko Agirre
ACL4
2020 Translation Artifacts in Cross-lingual Transfer Learning
abstract
Both human and machine translation play a central role in cross-lingual transfer learning: many multilingual datasets have been created through professional translation services, and using machine translation to translate either the test set or the training set is a widely used transfer technique.In this paper, we show that such translation process can introduce subtle artifacts that have a notable impact in existing cross-lingual models.For instance, in natural language inference, translating the premise and the hypothesis independently can reduce the lexical overlap between them, which current models are highly sensitive to.We show that some previous findings in cross-lingual transfer learning need to be reconsidered in the light of this phenomenon.Based on the gained insights, we also improve the state-of-the-art in XNLI for the translate-test and zero-shot approaches by 4.3 and 2.8 points, respectively.3 We use FastAlign (Dyer et al., 2013) for word alignment, and discard the few questions for which the mapping method fails (when none of the tokens in the answer span are aligned).4 We use the same procedure as for the training set except that (i) given the small size of the test set, we combine it with WikiMatrix (Schwenk et al., 2019) to aid word alignment, (ii) we use Jieba for Chinese segmentation instead of the Moses tokenizer, and (iii) for the few unaligned spans, we return the English answer.
Mikel Artetxe, Gorka Labaka, Eneko Agirre
EMNLP (1)2
2020 Do all roads lead to Rome? Understanding the role of initialization in iterative back-translation
Mikel Artetxe, Gorka Labaka, Noe Casas, Eneko Agirre
Knowl. Based Syst.2
2019 An Effective Approach to Unsupervised Machine Translation
abstract
While machine translation has traditionally relied on large amounts of parallel corpora, a recent research line has managed to train both Neural Machine Translation (NMT) and Statistical Machine Translation (SMT) systems using monolingual corpora only. In this paper, we identify and address several deficiencies of existing unsupervised SMT approaches by exploiting subword information, developing a theoretically well founded unsupervised tuning method, and incorporating a joint refinement procedure. Moreover, we use our improved SMT system to initialize a dual NMT model, which is further fine-tuned through on-the-fly back-translation. Together, we obtain large improvements over the previous state-of-the-art in unsupervised machine translation. For instance, we get 22.5 BLEU points in English-to-German WMT 2014, 5.5 points more than the previous best unsupervised system, and 0.5 points more than the (supervised) shared task winner back in 2014.
Mikel Artetxe, Gorka Labaka, Eneko Agirre
ACL (1)2
2019 Bilingual Lexicon Induction through Unsupervised Machine Translation
abstract
A recent research line has obtained strong results on bilingual lexicon induction by aligning independently trained word embeddings in two languages and using the resulting crosslingual embeddings to induce word translation pairs through nearest neighbor or related retrieval methods.In this paper, we propose an alternative approach to this problem that builds on the recent work on unsupervised machine translation.This way, instead of directly inducing a bilingual lexicon from cross-lingual embeddings, we use them to build a phrasetable, combine it with a language model, and use the resulting machine translation system to generate a synthetic parallel corpus, from which we extract the bilingual lexicon using statistical word alignment techniques.As such, our method can work with any word embedding and cross-lingual mapping technique, and it does not require any additional resource besides the monolingual corpus used to train the embeddings.When evaluated on the exact same cross-lingual embeddings, our proposed method obtains an average improvement of 6 accuracy points over nearest neighbor and 4 points over CSLS retrieval, establishing a new state-of-the-art in the standard MUSE dataset.
Mikel Artetxe, Gorka Labaka, Eneko Agirre
ACL (1)2
2019 Analyzing the Limitations of Cross-lingual Word Embedding Mappings
abstract
Recent research in cross-lingual word embeddings has almost exclusively focused on offline methods, which independently train word embeddings in different languages and map them to a shared space through linear transformations.While several authors have questioned the underlying isomorphism assumption, which states that word embeddings in different languages have approximately the same structure, it is not clear whether this is an inherent limitation of mapping approaches or a more general issue when learning crosslingual embeddings.So as to answer this question, we experiment with parallel corpora, which allows us to compare offline mapping to an extension of skip-gram that jointly learns both embedding spaces.We observe that, under these ideal conditions, joint learning yields to more isomorphic embeddings, is less sensitive to hubness, and obtains stronger results in bilingual lexicon induction.We thus conclude that current mapping methods do have strong limitations, calling for further research to jointly learn cross-lingual embeddings with a weaker cross-lingual signal.
Aitor Ormazabal, Mikel Artetxe, Gorka Labaka, Aitor Soroa, Eneko Agirre
ACL (1)3
2019 Neural machine translation of clinical texts between long distance languages
abstract
OBJECTIVE: To analyze techniques for machine translation of electronic health records (EHRs) between long distance languages, using Basque and Spanish as a reference. We studied distinct configurations of neural machine translation systems and used different methods to overcome the lack of a bilingual corpus of clinical texts or health records in Basque and Spanish. MATERIALS AND METHODS: We trained recurrent neural networks on an out-of-domain corpus with different hyperparameter values. Subsequently, we used the optimal configuration to evaluate machine translation of EHR templates between Basque and Spanish, using manual translations of the Basque templates into Spanish as a standard. We successively added to the training corpus clinical resources, including a Spanish-Basque dictionary derived from resources built for the machine translation of the Spanish edition of SNOMED CT into Basque, artificial sentences in Spanish and Basque derived from frequently occurring relationships in SNOMED CT, and Spanish monolingual EHRs. Apart from calculating bilingual evaluation understudy (BLEU) values, we tested the performance in the clinical domain by human evaluation. RESULTS: We achieved slight improvements from our reference system by tuning some hyperparameters using an out-of-domain bilingual corpus, obtaining 10.67 BLEU points for Basque-to-Spanish clinical domain translation. The inclusion of clinical terminology in Spanish and Basque and the application of the back-translation technique on monolingual EHRs significantly improved the performance, obtaining 21.59 BLEU points. This was confirmed by the human evaluation performed by 2 clinicians, ranking our machine translations close to the human translations. DISCUSSION: We showed that, even after optimizing the hyperparameters out-of-domain, the inclusion of available resources from the clinical domain and applied methods were beneficial for the described objective, managing to obtain adequate translations of EHR templates. CONCLUSION: We have developed a system which is able to properly translate health record templates from Basque to Spanish without making use of any bilingual corpus of clinical texts or health records.
Xabier Soto, Olatz Perez-de-Viñaspre, Gorka Labaka, Maite Oronoz
J. Am. Medical Informatics Assoc.3
2018 Generalizing and Improving Bilingual Word Embedding Mappings with a Multi-Step Framework of Linear Transformations
abstract
Using a dictionary to map independently trained word embeddings to a shared space has shown to be an effective approach to learn bilingual word embeddings. In this work, we propose a multi-step framework of linear transformations that generalizes a substantial body of previous work. The core step of the framework is an orthogonal transformation, and existing methods can be explained in terms of the additional normalization, whitening, re-weighting, de-whitening and dimensionality reduction steps. This allows us to gain new insights into the behavior of existing methods, including the effectiveness of inverse regression, and design a novel variant that obtains the best published results in zero-shot bilingual lexicon extraction. The corresponding software is released as an open source project.
Mikel Artetxe, Gorka Labaka, Eneko Agirre
AAAI2
2018 A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
abstract
Recent work has managed to learn crosslingual word embeddings without parallel data by mapping monolingual embeddings to a shared space through adversarial training.However, their evaluation has focused on favorable conditions, using comparable corpora or closely-related languages, and we show that they often fail in more realistic scenarios.This work proposes an alternative approach based on a fully unsupervised initialization that explicitly exploits the structural similarity of the embeddings, and a robust self-learning algorithm that iteratively improves this solution.Our method succeeds in all tested scenarios and obtains the best published results in standard datasets, even surpassing previous supervised systems.Our implementation is released as an open source project at https://github. com/artetxem/vecmap.
Mikel Artetxe, Gorka Labaka, Eneko Agirre
ACL (1)2
2018 Uncovering Divergent Linguistic Information in Word Embeddings with Lessons for Intrinsic and Extrinsic Evaluation
abstract
Following the recent success of word embeddings, it has been argued that there is no such thing as an ideal representation for words, as different models tend to capture divergent and often mutually incompatible aspects like semantics/syntax and similarity/relatedness.In this paper, we show that each embedding model captures more information than directly apparent.A linear transformation that adjusts the similarity order of the model without any external resource can tailor it to achieve better results in those aspects, providing a new perspective on how embeddings encode divergent linguistic information.In addition, we explore the relation between intrinsic and extrinsic evaluation, as the effect of our transformations in downstream tasks is higher for unsupervised systems than for supervised ones.
Mikel Artetxe, Gorka Labaka, Iñigo Lopez-Gazpio, Eneko Agirre
CoNLL2
2018 Neural Machine Translation of Basque
abstract
We describe the first experimental results in neural machine translation for Basque. As a synthetic language featuring agglutinative morphology, an extended case system, complex verbal morphology and relatively free word order, Basque presents a large number of challenging characteristics for machine translation in general, and for data-driven approaches such as attentionbased encoder-decoder models in particular. We present our results on a large range of experiments in Basque-Spanish translation, comparing several neural machine translation system variants with both rule-based and statistical machine translation systems. We demonstrate that significant gains can be obtained with a neural network approach for this challenging language pair, and describe optimal configurations in terms of word segmentation and decoding parameters, measured against test sets that feature multiple references to account for word order variability.
Thierry Etchegoyhen, Eva Martínez Garcia, Andoni Azpeitia, Gorka Labaka, Iñaki Alegria, Itziar Cortés Etxabe, Amaia Jauregi, Igor Ellakuria, Maite Martín, Eusebi Calonge
EAMT4
2018 Unsupervised Statistical Machine Translation
abstract
While modern machine translation has relied on large parallel corpora, a recent line of work has managed to train Neural Machine Translation (NMT) systems from monolingual corpora only (Artetxe et al., 2018c;Lample et al., 2018).Despite the potential of this approach for low-resource settings, existing systems are far behind their supervised counterparts, limiting their practical interest.In this paper, we propose an alternative approach based on phrase-based Statistical Machine Translation (SMT) that significantly closes the gap with supervised systems.Our method profits from the modular architecture of SMT: we first induce a phrase table from monolingual corpora through cross-lingual embedding mappings, combine it with an n-gram language model, and fine-tune hyperparameters through an unsupervised MERT variant.In addition, iterative backtranslation improves results further, yielding, for instance, 14.08 and 26.22 BLEU points in WMT 2014 English-German and English-French, respectively, an improvement of more than 7-10 BLEU points over previous unsupervised systems, and closing the gap with supervised SMT (Moses trained on Europarl) down to 2-5 BLEU points.Our implementation is available at https:// github.com/artetxem/monoses.
Mikel Artetxe, Gorka Labaka, Eneko Agirre
EMNLP2
2018 Unsupervised Neural Machine Translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, Kyunghyun Cho
ICLR (Poster)2
2018 Building Named Entity Recognition Taggers via Parallel Corpora
Rodrigo Agerri, Yiling Chung, Itziar Aldabe, Nora Aranberri, Gorka Labaka, German Rigau
LREC5
2018 Konbitzul: an MWE-specific database for Spanish-Basque
Uxoa Iñurrieta Urmeneta, Itziar Aduriz, Arantza Díaz de Ilarraza, Gorka Labaka, Kepa Sarasola
LREC4
2017 Learning bilingual word embeddings with (almost) no bilingual data
abstract
Most methods to learn bilingual word embeddings rely on large parallel corpora, which is difficult to obtain for most language pairs.This has motivated an active research line to relax this requirement, with methods that use document-aligned corpora or bilingual dictionaries of a few thousand words instead.In this work, we further reduce the need of bilingual resources using a very simple self-learning approach that can be combined with any dictionary-based mapping technique.Our method exploits the structural similarity of embedding spaces, and works with as little bilingual evidence as a 25 word dictionary or even an automatically generated list of numerals, obtaining results comparable to those of systems that use richer resources.
Mikel Artetxe, Gorka Labaka, Eneko Agirre
ACL (1)2
2016 Using Linguistic Data for English and Spanish Verb-Noun Combination Identification
abstract
We present a linguistic analysis of a set of English and Spanish verb+noun combinations (VNCs), and a method to use this information to improve VNC identification. Firstly, a sample of frequent VNCs are analysed in-depth and tagged along lexico-semantic and morphosyntactic dimensions, obtaining satisfactory inter-annotator agreement scores. Then, a VNC identification experiment is undertaken, where the analysed linguistic data is combined with chunking information and syntactic dependencies. A comparison between the results of the experiment and the results obtained by a basic detection method shows that VNC identification can be greatly improved by using linguistic information, as a large number of additional occurrences are detected with high precision.
Uxoa Iñurrieta Urmeneta, Arantza Díaz de Ilarraza, Gorka Labaka, Kepa Sarasola, Itziar Aduriz, John Carroll 0001
COLING3
2016 Learning principled bilingual mappings of word embeddings while preserving monolingual invariance
abstract
Mapping word embeddings of different languages into a single space has multiple applications.In order to map from a source space into a target space, a common approach is to learn a linear mapping that minimizes the distances between equivalences listed in a bilingual dictionary.In this paper, we propose a framework that generalizes previous work, provides an efficient exact method to learn the optimal linear transformation and yields the best bilingual results in translation induction while preserving monolingual performance in an analogy task.
Mikel Artetxe, Gorka Labaka, Eneko Agirre
EMNLP2
2016 Domain Adaptation in MT Using Titles in Wikipedia as a Parallel Corpus: Resources and Evaluation
Gorka Labaka, Iñaki Alegria, Kepa Sarasola
LREC1
2015 Exploiting portability to build an RBMT prototype for a new source language
Nora Aranberri, Gorka Labaka, Arantza Díaz de Ilarraza, Kepa Sarasola
EAMT2
2015 Building hybrid machine translation systems by using an EBMT preprocessor to create partial translations
Mikel Artetxe, Gorka Labaka, Kepa Sarasola
EAMT2
2014 A hybrid machine translation architecture guided by syntax
Gorka Labaka, Cristina España-Bonet, Lluís Màrquez, Kepa Sarasola
Mach. Transl.1
2011 Hybrid Machine Translation Guided by a Rule-Based System
Cristina España-Bonet, Gorka Labaka, Arantza Díaz de Ilarraza, Lluís Màrquez
MTSummit2
2011 Matxin, an open-source rule-based machine translation system for Basque
Aingeru Mayor, Iñaki Alegria, Arantza Díaz de Ilarraza, Gorka Labaka, Mikel Lersundi, Kepa Sarasola
Mach. Transl.4
2010 Plagiarism Detection across Distant Language Pairs
Alberto Barrón-Cedeño, Paolo Rosso, Eneko Agirre, Gorka Labaka
COLING4
2010 A Morphological Processor Based on Foma for Biscayan (a Basque dialect)
Iñaki Alegria, Garbiñe Aranbarri, Klara Ceberio, Gorka Labaka, Bittor Laskurain, Ruben Urizar
LREC4
2009 Use of Rich Linguistic Information to Translate Prepositions and Grammar Cases to Basque
Eneko Agirre, Aitziber Atutxa, Gorka Labaka, Mikel Lersundi, Aingeru Mayor, Kepa Sarasola
EAMT3
2009 Relevance of Different Segmentation Options on Spanish-Basque SMT
Arantza Díaz de Ilarraza, Gorka Labaka, Kepa Sarasola
EAMT2
2008 Strategies for sustainable MT for Basque: incremental design, reusability, standardization and open-source
Iñaki Alegria, Xabier Arregi, Xabier Artola, Arantza Díaz de Ilarraza, Gorka Labaka, Mikel Lersundi, Aingeru Mayor, Kepa Sarasola
IJCNLP5
2007 Transfer-Based MT from Spanish into Basque: Reusability, Standardization and Open Source
Iñaki Alegria, Arantza Díaz de Ilarraza, Gorka Labaka, Mikel Lersundi, Aingeru Mayor, Kepa Sarasola
CICLing3
2007 Comparing rule-based and data-driven approaches to Spanish-to-Basque machine translation
Gorka Labaka, Nicolas Stroppa, Andy Way, Kepa Sarasola
MTSummit1