Miquel Esplà-Gomis

dblp:33/10489 · DBLP profile ↗
← Back
32ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0002-2682-066XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 11 first-author · 13 since 2021
YearPublicationVenuePosition
2026 When Translations Surprise: Human Awareness of Predictability in Translations
Cristian García-Romero, Miquel Esplà-Gomis, Felipe Sánchez-Martinez
LREC2
2026 Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model
abstract
Abstract Modern machine translation (MT) systems depend on large parallel corpora, often collected from the Internet. However, recent evidence indicates that (i) a substantial portion of these texts are machine-generated translations, and (ii) an overreliance on such synthetic content in training data can significantly degrade translation quality. As a result, filtering out non-human translations is becoming an essential pre-processing step in building high-quality MT systems. In this work, we propose a novel approach that directly exploits the internal representations of a surrogate multilingual MT model to distinguish between human and machine-translated sentences. Experimental results show that our method outperforms current state-of-the-art techniques, particularly for non-English language pairs, achieving gains of at least 5 percentage points of accuracy.
Cristian García-Romero, Miquel Esplà-Gomis, Felipe Sánchez-Martínez
Trans. Assoc. Comput. Linguistics2
2025 Quality Beyond A Glance: Revealing Large Quality Differences Between Web-Crawled Parallel Corpora
abstract
Parallel corpora play a vital role in advanced multilingual natural language processing tasks, notably in machine translation (MT). The recent emergence of numerous large parallel corpora, often extracted from multilingual documents on the Internet, has expanded the available resources. Nevertheless, the quality of these corpora remains largely unexplored, while there are large differences in how the corpora are constructed. Moreover, how the potential differences affect the performance of neural MT (NMT) systems has also received limited attention. This study addresses this gap by manually and automatically evaluating four well-known publicly available parallel corpora across eleven language pairs. Our findings are quite concerning: all corpora contain a substantial amount of noisy sentence pairs, with CCMatrix and CCAligned having well below of 50% reasonably clean pairs. MaCoCu and ParaCrawl generally have higher quality texts, though around a third of the texts still have clear issues. While corpus size impacts NMT models’ performance, our study highlights the critical role of quality: higher-quality corpora consistently yield better-performing NMT models when controlling for size.
Rik van Noord, Miquel Esplà-Gomis, Malina Chichirau, Gema Ramírez-Sánchez, Antonio Toral
COLING2
2025 DeMINT: Automated Language Debriefing for English Learners via AI Chatbot Analysis of Meeting Transcripts
abstract
The objective of the DeMINT project is to develop a conversational tutoring system aimed at enhancing non-native English speakers’ language skills through post-meeting analysis of the transcriptions of video conferences in which they have participated. This paper describes the model developed and the results obtained through a human evaluation conducted with learners of English as a second language.
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz
MTSummit (2)1
2025 FLORES+ Mayas: Generating Textual Resources to Foster the Development of Language Technologies for Mayan Languages
abstract
A significant percentage of the population of Guatemala and Mexico belongs to various Mayan indigenous communities, for whom language barriers lead to social, economic, and digital exclusion. The Mayan languages spoken by these communities remain severely underrepresented in terms of digital resources, which prevents them from leveraging the latest advances in artificial intelligence. This project addresses that problem by means of: 1) the digitisation and release of multiple printed linguistic resources; 2) the development of a high-quality parallel machine translation (MT) evaluation corpus for six Mayan languages. In doing so, we are paving the way for the development of MT systems that will facilitate the access for Mayan speakers to essential services such as healthcare or legal aid. The resources are produced with the essential participation of indigenous communities, whereby native speakers provide the necessary translation services, QA, and linguistic expertise. The project is funded by the Google Academic Research Awards and carried out in collaboration with the Proyecto Lingüístico Francisco Marroquín Foundation in Guatemala.
Andrés Lou, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena
MTSummit (2)4
2024 Do Language Models Care about Text Quality? Evaluating Web-Crawled Corpora across 11 Languages
abstract
Large, curated, web-crawled corpora play a vital role in training language models (LMs). They form the lion’s share of the training data in virtually all recent LMs, such as the well-known GPT, LLaMA and XLM-RoBERTa models. However, despite this importance, relatively little attention has been given to the quality of these corpora. In this paper, we compare four of the currently most relevant large, web-crawled corpora (CC100, MaCoCu, mC4 and OSCAR) across eleven lower-resourced European languages. Our approach is two-fold: first, we perform an intrinsic evaluation by performing a human evaluation of the quality of samples taken from different corpora; then, we assess the practical impact of the qualitative differences by training specific LMs on each of the corpora and evaluating their performance on downstream tasks. We find that there are clear differences in quality of the corpora, with MaCoCu and OSCAR obtaining the best results. However, during the extrinsic evaluation, we actually find that the CC100 corpus achieves the highest scores. We conclude that, in our experiments, the quality of the web-crawled corpora does not seem to play a significant role when training LMs.
Rik van Noord, Taja Kuzman, Peter Rupnik, Nikola Ljubesic, Miquel Esplà-Gomis, Gema Ramírez-Sánchez, Antonio Toral
LREC/COLING5
2024 Lightweight neural translation technologies for low-resource languages
abstract
The LiLowLa (“Lightweight neural translation technologies for low-resource languages”) project aims to enhance machine translation (MT) and translation memory (TM) technologies, particularly for low-resource language pairs, where adequate linguistic resources are scarce. The project started in September 2022 and will run till August 2025.
Felipe Sánchez-Martínez, Juan Antonio Pérez-Ortiz, Víctor M. Sánchez-Cartagena, Andrés Lou, Cristian García-Romero, Aarón Galiano Jiménez, Miquel Esplà-Gomis
EAMT (2)7
2024 Non-Fluent Synthetic Target-Language Data Improve Neural Machine Translation
abstract
When the amount of parallel sentences available to train a neural machine translation is scarce, a common practice is to generate new synthetic training samples from them. A number of approaches have been proposed to produce synthetic parallel sentences that are similar to those in the parallel data available. These approaches work under the assumption that non-fluent target-side synthetic training samples can be harmful and may deteriorate translation performance. Even so, in this paper we demonstrate that synthetic training samples with non-fluent target sentences can improve translation performance if they are used in a multilingual machine translation framework as if they were sentences in another language. We conducted experiments on ten low-resource and four high-resource translation tasks and found out that this simple approach consistently improves translation performance as compared to state-of-the-art methods for generating synthetic training samples similar to those found in corpora. Furthermore, this improvement is independent of the size of the original training corpus, the resulting systems are much more robust against domain shift and produce less hallucinations.
Víctor M. Sánchez-Cartagena, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages
abstract
We present the most relevant results of the project MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages in its second year. To date, parallel and monolingual corpora have been produced for seven low-resourced European languages by crawling large amounts of textual data from selected top-level domains of the Internet; both human and automatic evaluation show its usefulness. In addition, several large language models pretrained on MaCoCu data have been published, as well as the code used to collect and curate the data.
Marta Bañón, Malina Chichirau, Miquel Esplà-Gomis, Mikel L. Forcada, Aarón Galiano Jiménez, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Peter Rupnik, Vit Suchomel, Antonio Toral, Jaume Zaragoza-Bernabeu
EAMT3
2022 MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages
abstract
We introduce the project “MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages”, funded by the Connecting Europe Facility, which is aimed at building monolingual and parallel corpora for under-resourced European languages. The approach followed consists of crawling large amounts of textual data from carefully selected top-level domains of the Internet, and then applying a curation and enrichment pipeline. In addition to corpora, the project will release successive versions of the free/open-source web crawling and curation software used.
Marta Bañón, Miquel Esplà-Gomis, Mikel L. Forcada, Cristian García-Romero, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Peter Rupnik, Vit Suchomel, Antonio Toral, Tobias van der Werff, Jaume Zaragoza
EAMT2
2022 Cross-lingual neural fuzzy matching for exploiting target-language monolingual corpora in computer-aided translation
abstract
Computer-aided translation (CAT) tools based on translation memories (MT) play a prominent role in the translation workflow of professional translators.However, the reduced availability of in-domain TMs, as compared to indomain monolingual corpora, limits its adoption for a number of translation tasks.In this paper, we introduce a novel neural approach aimed at overcoming this limitation by exploiting not only TMs, but also in-domain targetlanguage (TL) monolingual corpora, and still enabling a similar functionality to that offered by conventional TM-based CAT tools.Our approach relies on cross-lingual sentence embeddings to retrieve translation proposals from TL monolingual corpora, and on a neural model to estimate their post-editing effort.The paper presents an automatic evaluation of these techniques on four language pairs that shows that our approach can successfully exploit monolingual texts in a TM-based CAT environment, increasing the amount of useful translation proposals, and that our neural model for estimating the post-editing effort enables the combination of translation proposals obtained from monolingual corpora and from TMs in the usual way.A human evaluation performed on a single language pair confirms the results of the automatic evaluation and seems to indicate that the translation proposals retrieved with our approach are more useful than what the automatic evaluation shows.
Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez
EMNLP1
2021 Rethinking Data Augmentation for Low-Resource Neural Machine Translation: A Multi-Task Learning Approach
abstract
In the context of neural machine translation, data augmentation (DA) techniques may be used for generating additional training samples when the available parallel data are scarce.Many DA approaches aim at expanding the support of the empirical data distribution by generating new sentence pairs that contain infrequent words, thus making it closer to the true data distribution of parallel sentences.In this paper, we propose to follow a completely different approach and present a multi-task DA approach in which we generate new sentence pairs with transformations, such as reversing the order of the target sentence, which produce unfluent target sentences.During training, these augmented sentences are used as auxiliary tasks in a multi-task framework with the aim of providing new contexts where the target prefix is not informative enough to predict the next word.This strengthens the encoder and forces the decoder to pay more attention to the source representations of the encoder.Experiments carried out on six lowresource translation tasks show consistent improvements over the baseline and over DA methods aiming at extending the support of the empirical data distribution.The systems trained with our approach rely more on the source tokens, are more robust against domain shift and suffer less hallucinations.
Víctor M. Sánchez-Cartagena, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez
EMNLP (1)2
2021 Surprise Language Challenge: Developing a Neural Machine Translation System between Pashto and English in Two Months
abstract
In the media industry and the focus of global reporting can shift overnight. There is a compelling need to be able to develop new machine translation systems in a short period of time and in order to more efficiently cover quickly developing stories. As part of the EU project GoURMET and which focusses on low-resource machine translation and our media partners selected a surprise language for which a machine translation system had to be built and evaluated in two months(February and March 2021). The language selected was Pashto and an Indo-Iranian language spoken in Afghanistan and Pakistan and India. In this period we completed the full pipeline of development of a neural machine translation system: data crawling and cleaning and aligning and creating test sets and developing and testing models and and delivering them to the user partners. In this paperwe describe rapid data creation and experiments with transfer learning and pretraining for this low-resource language pair. We find that starting from an existing large model pre-trained on 50languages leads to far better BLEU scores than pretraining on one high-resource language pair with a smaller model. We also present human evaluation of our systems and which indicates that the resulting systems perform better than a freely available commercial system when translating from English into Pashto direction and and similarly when translating from Pashto into English.
Alexandra Birch, Barry Haddow, Antonio Valerio Miceli Barone, Jindrich Helcl, Jonas Waldendorf, Felipe Sánchez-Martínez, Mikel L. Forcada, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Miquel Esplà-Gomis, Wilker Aziz, Lina Murady, Sevi Sariisik, Peggy van der Kreeft, Kay Macquarrie
MTSummit (1)10
2020 ParaCrawl: Web-Scale Acquisition of Parallel Corpora
abstract
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, Jaume Zaragoza. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz-Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strong, Brian Thompson 0001, William Waites, Dion Wiggins, Jaume Zaragoza
ACL6
2020 An English-Swahili parallel corpus and its use for neural machine translation in the news domain
abstract
This paper describes our approach to create a neural machine translation system to translate between English and Swahili (both directions) in the news domain, as well as the process we followed to crawl the necessary parallel corpora from the Internet. We report the results of a pilot human evaluation performed by the news media organisations participating in the H2020 EU-funded project GoURMET.
Felipe Sánchez-Martínez, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Mikel L. Forcada, Miquel Esplà-Gomis, Andrew Secker, Susie Coleman, Julie Wall
EAMT5
2019 Global Under-Resourced Media Translation (GoURMET)
Alexandra Birch, Barry Haddow, Ivan Titov 0001, Antonio Valerio Miceli Barone, Rachel Bawden, Felipe Sánchez-Martínez, Mikel L. Forcada, Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Wilker Aziz, Andrew Secker, Peggy van der Kreeft
MTSummit (2)8
2019 ParaCrawl: Web-scale parallel corpora for the languages of the EU
Miquel Esplà-Gomis, Mikel L. Forcada, Gema Ramírez-Sánchez, Hieu Hoang
MTSummit (2)1
2018 Proceedings of the 21st Annual Conference of the European Association for Machine Translation
Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Miquel Esplà-Gomis, Maja Popovic, Celia Rico, Joachim Van den Bogaert, Mikel L. Forcada
EAMT3
2017 One-parameter models for sentence-level post-editing effort estimation
Mikel L. Forcada, Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Lucia Specia
MTSummit (1)2
2016 Stand-off Annotation of Web Content as a Legally Safer Alternative to Bitext Crawling for Distribution
Mikel L. Forcada, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz
EAMT2
2016 Producing Monolingual and Parallel Web Corpora at the Same Time - SpiderLing and Bitextor's Love Affair
Nikola Ljubesic, Miquel Esplà-Gomis, Antonio Toral, Sergio Ortiz-Rojas, Filip Klubicka
LREC2
2016 New directions in empirical translation process research exploring the CRITT TPR-DB - Springer International Publishing, Switzerland, 2016, ISBN: 978-3-319-20357-7, v + 315 pp
Miquel Esplà-Gomis
Mach. Transl.1
2015 Using on-line available sources of bilingual information for word-level machine translation quality estimation
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada
EAMT1
2015 Abu-MaTran: Automatic building of Machine Translation
Antonio Toral, Flammie A. Pirinen, Andy Way, Gema Ramírez-Sánchez, Sergio Ortiz-Rojas, Raphaël Rubino, Miquel Esplà-Gomis, Mikel L. Forcada, Vassilis Papavassiliou, Prokopis Prokopidis, Nikola Ljubesic
EAMT7
2015 Using Machine Translation to Provide Target-Language Edit Hints in Computer Aided Translation Based on Translation Memories
abstract
This paper explores the use of general-purpose machine translation (MT) in assisting the users of computer-aided translation (CAT) systems based on translation memory (TM) to identify the target words in the translation proposals that need to be changed (either replaced or removed) or kept unedited, a task we term as "word-keeping recommendation". MT is used as a black box to align source and target sub-segments on the fly in the translation units (TUs) suggested to the user. Source-language (SL) and target-language (TL) segments in the matching TUs are segmented into overlapping sub-segments of variable length and machine-translated into the TL and the SL, respectively. The bilingual sub-segments obtained and the matching between the SL segment in the TU and the segment to be translated are employed to build the features that are then used by a binary classifier to determine the target words to be changed and those to be kept unedited. In this approach, MT results are never presented to the translator. Two approaches are presented in this work: one using a word-keeping recommendation system which can be trained on the TM used with the CAT system, and a more basic approach which does not require any training. Experiments are conducted by simulating the translation of texts in several language pairs with corpora belonging to different domains and using three different MT systems. We compare the performance obtained to that of previous works that have used statistical word alignment for word-keeping recommendation, and show that the MT-based approaches presented in this paper are more accurate in most scenarios. In particular, our results confirm that the MT-based approaches are better than the alignment-based approach when using models trained on out-of-domain TMs. Additional experiments were performed to check how dependent the MT-based recommender is on the language pair and MT system used for training. These experiments confirm a high degree of reusability of the recommendation models across various MT systems, but a low level of reusability across language pairs.
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada
J. Artif. Intell. Res.1
2014 An efficient method to assist non-expert users in extending dictionaries by assigning stems and inflectional paradigms to unknknown words
Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Felipe Sánchez-Martínez, Rafael C. Carrasco, Mikel L. Forcada, Juan Antonio Pérez-Ortiz
EAMT1
2014 Extrinsic evaluation of web-crawlers in machine translation: a study on Croatian-English for the tourism domain
Antonio Toral, Raphaël Rubino, Miquel Esplà-Gomis, Flammie A. Pirinen, Andy Way, Gema Ramírez-Sánchez
EAMT3
2014 Comparing two acquisition systems for automatically building an English-Croatian parallel corpus from multilingual websites
Miquel Esplà-Gomis, Filip Klubicka, Nikola Ljubesic, Sergio Ortiz-Rojas, Vassilis Papavassiliou, Prokopis Prokopidis
LREC1
2012 Source-Language Dictionaries Help Non-Expert Users to Enlarge Target-Language Dictionaries for Machine Translation
Víctor M. Sánchez-Cartagena, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz
LREC2
2011 Using word alignments to assist computer-aided translation users by marking which target-side words to change or keep unedited
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada
EAMT1
2011 Using machine translation in computer-aided translation to suggest the target-side words to change
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada
MTSummit1
2011 Multimodal Building of Monolingual Dictionaries for Machine Translation by Non-Expert Users
Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz
MTSummit1