Cristina España-Bonet

dblp:59/7935 · DBLP profile ↗
← Back
37ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0001-5414-4710ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 5 first-author · 21 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KinyCOMET: Automatic Evaluation of Machine Translation Systems for Kinyarwanda-English
Prince Chris Mazimpaka, Jan Nehring, Samuel Rutunda, Cristina España-Bonet
LREC4
2026 PETra: A Multilingual Corpus of Pragmatic Explicitation in Translation
abstract
Translators often enrich texts with background details that make implicit cultural meanings explicit for new audiences. This phenomenon, known as pragmatic explicitation, has been widely discussed in translation theory but rarely modeled computationally. We introduce PragExTra, the first multilingual corpus and detection framework for pragmatic explicitation. The corpus covers eight language pairs from TED-Multi and Europarl and includes additions such as entity descriptions, measurement conversions, and translator remarks. We identify candidate explicitation cases through null alignments and refined using active learning with human annotation. Our results show that entity and system-level explicitations are most frequent, and that active learning improves classifier accuracy by 7-8 percentage points, achieving up to 0.88 accuracy and 0.82 F1 across languages. PragExTra establishes pragmatic explicitation as a measurable, cross-linguistic phenomenon and takes a step towards building culturally aware machine translation. Keywords: translation, multilingualism, explicitation
Doreen Osmelak, Koel Dutta Chowdhury, Uliana Sentsova, Cristina España-Bonet, Josef van Genabith
LREC4
2026 A Critical Study of Automatic Evaluation in Sign Language Translation
Shakib Yazdani, Yasser Hamidullah, Cristina España-Bonet, Eleftherios Avramidis, Josef van Genabith
LREC3
2026 LaFresCat: A studio-quality Catalan multi-accent speech dataset for text-to-speech synthesis
abstract
Current text-to-speech (TTS) systems are capable of learning the phonetics of a language accurately given that the speech data used to train such models covers all phonetic phenomena. For languages with different varieties, this includes all their richness and accents. This is the case of Catalan, a mid-resourced language with several dialects or accents. Although there are various publicly available corpora, there is a lack of high-quality open-access data for speech technologies covering its variety of accents. Common Voice includes recordings of Catalan speakers from different regions; however, accent labeling has been shown to be inaccurate, and artificially enhanced samples may be unsuitable for TTS. To address these limitations, we present LaFresCat, the first studio-quality Catalan multi-accent dataset. LaFresCat comprises 3.5 h of professionally recording speech covering four of the most prominent Catalan accents: Balearic, Central, North-Western, and Valencian. In this work, we provide a detailed description of the dataset design: utterances were selected to be phonetically balanced, detailed speaker instructions were provided, native speakers from the regions corresponding to the Catalan accents were hired, and the recordings were formatted and post-processed. The resulting dataset, LaFresCat, is publicly available. To preliminarily evaluate the dataset, we trained and assessed a lightweight flow-based TTS system, which is also provided as a by-product. We also analyzed LaFresCat samples and the corresponding TTS-generated samples at the phonetic level, employing expert annotations and Pillai scores to quantify acoustic vowel overlap. Preliminary results suggest a significant improvement in predicted mean opinion score (UTMOS), with an increase of 0.42 points when the TTS system is fine-tuned on LaFresCat rather than trained from scratch, starting from a pre-trained version based on Central Catalan data from Common Voice. Subsequent human expert annotations achieved nearly 90% accuracy in accent classification for LaFresCat recordings. However, although the TTS tends to homogenize pronunciation, it still learns distinct dialectal patterns. This assessment offers key insights for establishing a baseline to guide future evaluations of Catalan multi-accent TTS systems and further studies of LaFresCat.
Alex Peiró Lilja, Carme Armentano-Oller, José Giraldo, Wendy Elvira-García, Ignasi Esquerra, Rodolfo Zevallos, Cristina España-Bonet, Martí Llopart-Font, Baybars Külebi, Mireia Farrús
Comput. Speech Lang.7
2025 AFRIDOC-MT: Document-level MT Corpus for African Languages
abstract
Jesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, Dawei Zhu, David Ifeoluwa Adelani, Clement Oyeleke Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jesujoba O. Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, David Ifeoluwa Adelani, Clement Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow
EMNLP4
2025 Evaluating Speech Enhancement Performance Across Demographics and Language
José Giraldo, Alex Peiró Lilja, Carme Armentano-Oller, Rodolfo Zevallos, Cristina España-Bonet
INTERSPEECH5
2025 Towards Domain-Specific Spoken Language Understanding for a Catalan Voice-Controlled Video Game
Alex Peiró Lilja, Rodolfo Zevallos, Carme Armentano-Oller, José Giraldo, Cristina España-Bonet, Mireia Farrús
INTERSPEECH5
2025 Continual Learning in Multilingual Sign Language Translation
abstract
Shakib Yazdani, Josef Van Genabith, Cristina España-Bonet. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Shakib Yazdani, Josef van Genabith, Cristina España-Bonet
NAACL (Long Papers)3
2024 When Your Cousin Has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages
abstract
Most existing approaches for unsupervised bilingual lexicon induction (BLI) depend on good quality static or contextual embeddings requiring large monolingual corpora for both languages. However, unsupervised BLI is most likely to be useful for low-resource languages (LRLs), where large datasets are not available. Often we are interested in building bilingual resources for LRLs against related high-resource languages (HRLs), resulting in severely imbalanced data settings for BLI. We first show that state-of-the-art BLI methods in the literature exhibit near-zero performance for severely data-imbalanced language pairs, indicating that these settings require more robust techniques. We then present a new method for unsupervised BLI between a related LRL and HRL that only requires inference on a masked language model of the HRL, and demonstrate its effectiveness on truly low-resource languages Bhojpuri and Magahi (with <5M monolingual tokens each), against Hindi. We further present experiments on (mid-resource) Marathi and Nepali to compare approach performances by resource range, and release our resulting lexicons for five low-resource Indic languages: Bhojpuri, Magahi, Awadhi, Braj, and Maithili, against Hindi.
Niyati Bafna, Cristina España-Bonet, Josef van Genabith, Benoît Sagot, Rachel Bawden
LREC/COLING2
2024 DGS-Fabeln-1: A Multi-Angle Parallel Corpus of Fairy Tales between German Sign Language and German Text
abstract
We present the acquisition process and the data of DGS-Fabeln-1, a parallel corpus of German text and videos containing German fairy tales interpreted into the German Sign Language (DGS) by a native DGS signer. The corpus contains 573 segments of videos with a total duration of 1 hour and 32 minutes, corresponding with 1428 written sentences. It is the first corpus of semi-naturally expressed DGS that has been filmed from 7 angles, and one of the few sign language (SL) corpora globally which have been filmed from more than 3 angles and where the listener has been simultaneously filmed. The corpus aims at aiding research at SL linguistics, SL machine translation and affective computing, and is freely available for research purposes at the following address: https://doi.org/10.5281/zenodo.10822097.
Fabrizio Nunnari, Eleftherios Avramidis, Cristina España-Bonet, Marco González, Anna Hennes, Patrick Gebhard
LREC/COLING3
2024 Mitigating Translationese with GPT-4: Strategies and Performance
abstract
Translations differ in systematic ways from texts originally authored in the same language.These differences, collectively known as translationese, can pose challenges in cross-lingual natural language processing: models trained or tested on translated input might struggle when presented with non-translated language. Translationese mitigation can alleviate this problem. This study investigates the generative capacities of GPT-4 to reduce translationese in human-translated texts. The task is framed as a rewriting process aimed at modified translations indistinguishable from the original text in the target language. Our focus is on prompt engineering that tests the utility of linguistic knowledge as part of the instruction for GPT-4. Through a series of prompt design experiments, we show that GPT4-generated revisions are more similar to originals in the target language when the prompts incorporate specific linguistic instructions instead of relying solely on the model’s internal knowledge. Furthermore, we release the segment-aligned bidirectional German-English data built from the Europarl corpus that underpins this study.
Maria Kunilovskaya, Koel Dutta Chowdhury, Heike Przybyl, Cristina España-Bonet, Josef van Genabith
EAMT (1)4
2024 Elote, Choclo and Mazorca: on the Varieties of Spanish
abstract
Cristina España-Bonet, Alberto Barrón-Cedeño. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Cristina España-Bonet, Alberto Barrón-Cedeño
NAACL-HLT1
2023 First WMT Shared Task on Sign Language Translation (WMT-SLT22)
abstract
This paper is a brief summary of the First WMT Shared Task on Sign Language Translation (WMT-SLT22), a project partly funded by EAMT. The focus of this shared task is automatic translation between signed and spoken languages. Details can be found on our website (https://www.wmt-slt.com/) or in the findings paper (Müller et al., 2022).
Mathias Müller 0002, Sarah Ebling, Eleftherios Avramidis, Alessia Battisti, Michèle Berger, Richard Bowden, Annelies Braffort, Necati Cihan Camgöz, Cristina España-Bonet, Roman Grundkiewicz, Zifan Jiang, Oscar Koller, Amit Moryossef, Regula Perrollaz, Sabine Reinhard, Annette Rios, Dimitar Sht. Shterionov, Sandra Sidler-Miserez, Katja Tissi, Davy Van Landuyt
EAMT9
2023 Translating away Translationese without Parallel Data
abstract
Translated texts exhibit systematic linguistic differences compared to original texts in the same language, and these differences are referred to as translationese.Translationese has effects on various cross-lingual natural language processing tasks, potentially leading to biased results.In this paper, we explore a novel approach to reduce translationese in translated texts: translation-based style transfer.As there are no parallel human-translated and original data in the same language, we use a selfsupervised approach that can learn from comparable (rather than parallel) mono-lingual original and translated data.However, even this self-supervised approach requires some parallel data for validation.We show how we can eliminate the need for parallel validation data by combining the self-supervised loss with an unsupervised loss.This unsupervised loss leverages the original language model loss over the style-transferred output and a semantic similarity loss between the input and style-transferred output.We evaluate our approach in terms of original vs. translationese binary classification in addition to measuring content preservation and target-style fluency.The results show that our approach is able to reduce translationese classifier accuracy to a level of a random classifier after style transfer while adequately preserving the content and fluency in the target original style.
Rricha Jalota, Koel Dutta Chowdhury, Cristina España-Bonet, Josef van Genabith
EMNLP3
2023 Tailoring and evaluating the Wikipedia for in-domain comparable corpora extraction
abstract
Abstract We propose a language-independent graph-based method to build à-la-carte article collections on user-defined domains from the Wikipedia. The core model is based on the exploration of the encyclopedia’s category graph and can produce both mono- and multilingual comparable collections. We run thorough experiments to assess the quality of the obtained corpora in 10 languages and 743 domains. According to an extensive manual evaluation, our graph model reaches an average precision of $$84\%$$ 84 % on in-domain articles, outperforming an alternative model based on information retrieval techniques. As manual evaluations are costly, we introduce the concept of domainness and design several automatic metrics to account for the quality of the collections. Our best metric for domainness shows a strong correlation with human judgments, representing a reasonable automatic alternative to assess the quality of domain-specific corpora. We release the toolkit with the implementation of the extraction methods, the evaluation measures and several utilities.
Cristina España-Bonet, Alberto Barrón-Cedeño, Lluís Màrquez
Knowl. Inf. Syst.1
2022 Combining Noisy Semantic Signals with Orthographic Cues: Cognate Induction for the Indic Dialect Continuum
abstract
We present a novel method for unsupervised cognate/borrowing identification from monolingual corpora designed for low and extremely low resource scenarios, based on combining noisy semantic signals from joint bilingual spaces with orthographic cues modelling sound change.We apply our method to the North Indian dialect continuum, containing several dozens of dialects and languages spoken by more than 100 million people.Many of these languages are zero-resource and therefore natural language processing for them is nonexistent.We first collect monolingual data for 26 Indic languages, 16 of which were previously zero-resource, and perform exploratory character, lexical and subword cross-lingual alignment experiments for the first time at this scale on this dialect continuum.We create bilingual evaluation lexicons against Hindi for 20 of the languages.We then apply our cognate identification method on the data, and show that our method outperforms both traditional orthography baselines as well as EM-style learnt edit distance matrices.To the best of our knowledge, this is the first work to combine traditional orthographic cues with noisy bilingual embeddings to tackle unsupervised cognate detection in a (truly) low-resource setup, showing that even noisy bilingual embeddings can act as good guides for this task.We release our multilingual dialect corpus, called HinDialect, as well as our scripts for evaluation data collection and cognate induction.2
Niyati Bafna, Josef van Genabith, Cristina España-Bonet, Zdenek Zabokrtský
CoNLL3
2022 The (Undesired) Attenuation of Human Biases by Multilinguality
abstract
Some human preferences are universal.The odor of vanilla is perceived as pleasant all around the world.We expect neural models trained on human texts to exhibit these kind of preferences, i.e. biases, but we show that this is not always the case.We explore 16 static and contextual embedding models in 9 languages and, when possible, compare them under similar training conditions.We introduce and release CA-WEAT, multilingual cultural aware tests to quantify biases, and compare them to previous English-centric tests.Our experiments confirm that monolingual static embeddings do exhibit human biases, but values differ across languages, being far from universal.Biases are less evident in contextual models, to the point that the original human association might be reversed.Multilinguality proves to be another variable that attenuates and even reverses the effect of the bias, specially in contextual multilingual models.In order to explain this variance among models and languages, we examine the effect of asymmetries in the training corpus, departures from isomorphism in multilingual embedding spaces and discrepancies in the testing measures between languages.
Cristina España-Bonet, Alberto Barrón-Cedeño
EMNLP1
2022 Towards Debiasing Translation Artifacts
abstract
Koel Dutta Chowdhury, Rricha Jalota, Cristina España-Bonet, Josef Genabith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Koel Dutta Chowdhury, Rricha Jalota, Cristina España-Bonet, Josef van Genabith
NAACL-HLT3
2021 Comparing Feature-Engineering and Feature-Learning Approaches for Multilingual Translationese Classification
abstract
Daria Pylypenko, Kwabena Amponsah-Kaakyire, Koel Dutta Chowdhury, Josef van Genabith, Cristina España-Bonet. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Daria Pylypenko, Kwabena Amponsah-Kaakyire, Koel Dutta Chowdhury, Josef van Genabith, Cristina España-Bonet
EMNLP (1)5
2021 A Data Augmentation Approach for Sign-Language-To-Text Translation In-The-Wild
abstract
In this paper, we describe the current main approaches to sign language translation which use deep neural networks with videos as input and text as output. We highlight that, under our point of view, their main weakness is the lack of generalization in daily life contexts. Our goal is to build a state-of-the-art system for the automatic interpretation of sign language in unpredictable video framing conditions. Our main contribution is the shift from image features to landmark positions in order to diminish the size of the input data and facilitate the combination of data augmentation techniques for landmarks. We describe the set of hypotheses to build such a system and the list of experiments that will lead us to their verification.
Fabrizio Nunnari, Cristina España-Bonet, Eleftherios Avramidis
LDK2
2021 The Effect of Domain and Diacritics in Yoruba-English Neural Machine Translation
abstract
Massively multilingual machine translation (MT) has shown impressive capabilities and including zero and few-shot translation between low-resource language pairs. However and these models are often evaluated on high-resource languages with the assumption that they generalize to low-resource ones. The difficulty of evaluating MT models on low-resource pairs is often due to lack of standardized evaluation datasets. In this paper and we present MENYO-20k and the first multi-domain parallel corpus with a especially curated orthography for Yoruba–English with standardized train-test splits for benchmarking. We provide several neural MT benchmarks and compare them to the performance of popular pre-trained (massively multilingual) MT models both for the heterogeneous test set and its subdomains. Since these pre-trained models use huge amounts of data with uncertain quality and we also analyze the effect of diacritics and a major characteristic of Yoruba and in the training data. We investigate how and when this training condition affects the final quality of a translation and its understandability.Our models outperform massively multilingual models such as Google (+8.7 BLEU) and Facebook M2M (+9.1) when translating to Yoruba and setting a high quality benchmark for future research.
David Ifeoluwa Adelani, Dana Ruiter, Jesujoba O. Alabi, Damilola Adebonojo, Adesina Ayeni, Mofe Adeyemi, Ayodele Awokoya, Cristina España-Bonet
MTSummit (1)8
2021 Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource Languages
abstract
For most language combinations and parallel data is either scarce or simply unavailable. To address this and unsupervised machine translation (UMT) exploits large amounts of monolingual data by using synthetic data generation techniques such as back-translation and noising and while self-supervised NMT (SSNMT) identifies parallel sentences in smaller comparable data and trains on them. To this date and the inclusion of UMT data generation techniques in SSNMT has not been investigated. We show that including UMT techniques into SSNMT significantly outperforms SSNMT (up to +4.3 BLEU and af2en) as well as statistical (+50.8 BLEU) and hybrid UMT (+51.5 BLEU) baselines on related and distantly-related and unrelated language pairs.
Dana Ruiter, Dietrich Klakow, Josef van Genabith, Cristina España-Bonet
MTSummit (1)4
2020 Understanding Translationese in Multi-view Embedding Spaces
abstract
Recent studies use a combination of lexical and syntactic features to show that footprints of the source language remain visible in translations, to the extent that it is possible to predict the original source language from the translation.In this paper, we focus on embedding-based semantic spaces, exploiting departures from isomorphism between spaces built from original target language and translations into this target language to predict relations between languages in an unsupervised way.We use different views of the data -words, parts of speech, semantic tags and synsets -to track translationese.Our analysis shows that (i) semantic distances between original target language and translations into this target language can be detected using the notion of isomorphism, (ii) language family ties with characteristics similar to linguistically motivated phylogenetic trees can be inferred from the distances and (iii) with delexicalised embeddings exhibiting source-language interference most significantly, other levels of abstraction display the same tendency, indicating the lexicalised results to be not "just" due to possible topic differences between original and translated texts.To the best of our knowledge, this is the first time departures from isomorphism between embedding spaces are used to track translationese.
Koel Dutta Chowdhury, Cristina España-Bonet, Josef van Genabith
COLING2
2020 Self-Induced Curriculum Learning in Self-Supervised Neural Machine Translation
abstract
Self-supervised neural machine translation (SSNMT) jointly learns to identify and select suitable training data from comparable (rather than parallel) corpora and to translate, in a way that the two tasks support each other in a virtuous circle.In this study, we provide an in-depth analysis of the sampling choices the SSNMT model makes during training.We show how, without it having been told to do so, the model self-selects samples of increasing (i) complexity and (ii) task-relevance in combination with (iii) performing a denoising curriculum.We observe that the dynamics of the mutual-supervision signals of both system internal representation types are vital for the extraction and translation performance.We show that in terms of the Gunning-Fog Readability index, SSNMT starts extracting and learning from Wikipedia data suitable for high school students and quickly moves towards content suitable for first year undergraduate students.
Dana Ruiter, Josef van Genabith, Cristina España-Bonet
EMNLP (1)3
2020 Massive vs. Curated Embeddings for Low-Resourced Languages: the Case of Yorùbá and Twi
abstract
The success of several architectures to learn semantic representations from unannotated text and the availability of these kind of texts in online multilingual resources such as Wikipedia has facilitated the massive and automatic creation of resources for multiple languages. The evaluation of such resources is usually done for the high-resourced languages, where one has a smorgasbord of tasks and test sets to evaluate on. For low-resourced languages, the evaluation is more difficult and normally ignored, with the hope that the impressive capability of deep learning architectures to learn (multilingual) representations in the high-resourced setting holds in the low-resourced setting too. In this paper we focus on two African languages, Yorùbá and Twi, and compare the word embeddings obtained in this way, with word embeddings obtained from curated corpora and a language-dependent processing. We analyse the noise in the publicly available corpora, collect high quality and noisy data for the two languages and quantify the improvements that depend not only on the amount of data but on the quality too. We also use different architectures that learn word representations both from surface forms and characters to further exploit all the available information which showed to be important for these languages. For the evaluation, we manually translate the wordsim-353 word pairs dataset from English into Yorùbá and Twi. We extend the analysis to contextual word embeddings and evaluate multilingual BERT on a named entity recognition task. For this, we annotate with named entities the Global Voices corpus for Yorùbá. As output of the work, we provide corpora, embeddings and the test suits for both languages.
Jesujoba O. Alabi, Kwabena Amponsah-Kaakyire, David Ifeoluwa Adelani, Cristina España-Bonet
LREC4
2020 GeBioToolkit: Automatic Extraction of Gender-Balanced Multilingual Corpus of Wikipedia Biographies
abstract
We introduce GeBioToolkit, a tool for extracting multilingual parallel corpora at sentence level, with document and gender information from Wikipedia biographies. Despite the gender inequalities present in Wikipedia, the toolkit has been designed to extract corpus balanced in gender. While our toolkit is customizable to any number of languages (and different domains), in this work we present a corpus of 2,000 sentences in English, Spanish and Catalan, which has been post-edited by native speakers to become a high-quality dataset for machine translation evaluation. While GeBioCorpus aims at being one of the first non-synthetic gender-balanced test datasets, GeBioToolkit aims at paving the path to standardize procedures to produce gender-balanced datasets.
Marta R. Costa-jussà, Pau Li Lin, Cristina España-Bonet
LREC3
2020 Multilingual and Interlingual Semantic Representations for Natural Language Processing: A Brief Introduction
abstract
We introduce the Computational Linguistics special issue on Multilingual and Interlingual Semantic Representations for Natural Language Processing. We situate the special issue’s five articles in the context of our fast-changing field, explaining our motivation for this project. We offer a brief summary of the work in the issue, which includes developments on lexical and sentential semantic representations, from symbolic and neural perspectives.
Marta R. Costa-jussà, Cristina España-Bonet, Pascale Fung, Noah A. Smith
Comput. Linguistics2
2019 Self-Supervised Neural Machine Translation
abstract
We present a simple new method where an emergent NMT system is used for simultaneously selecting training data and learning internal NMT representations.This is done in a self-supervised way without parallel data, in such a way that both tasks enhance each other during training.The method is language independent, introduces no additional hyper-parameters, and achieves BLEU scores of 29.21 (en2f r) and 27.36 (f r2en) on new-stest2014 using English and French Wikipedia data for training.
Dana Ruiter, Cristina España-Bonet, Josef van Genabith
ACL (1)2
2016 TweetMT: A Parallel Microblog Corpus
Iñaki San Vicente, Iñaki Alegria, Cristina España-Bonet, Pablo Gamallo 0001, Hugo Gonçalo Oliveira, Eva Martínez Garcia, Antonio Toral, Arkaitz Zubiaga, Nora Aranberri
LREC3
2015 Document-Level Machine Translation with Word Vector Models
Eva Martínez Garcia, Cristina España-Bonet, Lluís Màrquez
EAMT2
2014 A hybrid machine translation architecture guided by syntax
Gorka Labaka, Cristina España-Bonet, Lluís Màrquez, Kepa Sarasola
Mach. Transl.2
2012 A Hybrid System for Patent Translation
Ramona Enache, Cristina España-Bonet, Aarne Ranta, Lluís Màrquez
EAMT2
2012 Context-Aware Machine Translation for Software Localization
Victor Muntés-Mulero, Patricia Paladini Adell, Cristina España-Bonet, Lluís Màrquez
EAMT3
2011 Hybrid Machine Translation Guided by a Rule-Based System
Cristina España-Bonet, Gorka Labaka, Arantza Díaz de Ilarraza, Lluís Màrquez
MTSummit1
2010 Robust Estimation of Feature Weights in Statistical Machine Translation
Cristina España-Bonet, Lluís Màrquez
EAMT1
2010 Language Technology Challenges of a 'Small' Language (Catalan)
Maite Melero, Gemma Boleda, Montse Cuadros, Cristina España-Bonet, Lluís Padró 0001, Martí Quixal, Carlos Rodríguez Penagos, Roser Saurí
LREC4
2009 Discriminative Phrase-Based Models for Arabic Machine Translation
abstract
A design for an Arabic-to-English translation system is presented. The core of the system implements a standard phrase-based statistical machine translation architecture, but it is extended by incorporating a local discriminative phrase selection model to address the semantic ambiguity of Arabic. Local classifiers are trained using linguistic information and context to translate a phrase, and this significantly increases the accuracy in phrase selection with respect to the most frequent translation traditionally considered. These classifiers are integrated into the translation system so that the global task gets benefits from the discriminative learning. As a result, we obtain significant improvements in the full translation task at the lexical, syntactic, and semantic levels as measured by an heterogeneous set of automatic evaluation metrics.
Cristina España-Bonet, Jesús Giménez, Lluís Màrquez
ACM Trans. Asian Lang. Inf. Process.1