Felipe Sánchez-Martínez

dblp:56/6718 · DBLP profile ↗
← Back
45ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-2295-2630ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 10 first-author · 13 since 2021Databases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author
YearPublicationVenuePosition
2026 Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model
abstract
Abstract Modern machine translation (MT) systems depend on large parallel corpora, often collected from the Internet. However, recent evidence indicates that (i) a substantial portion of these texts are machine-generated translations, and (ii) an overreliance on such synthetic content in training data can significantly degrade translation quality. As a result, filtering out non-human translations is becoming an essential pre-processing step in building high-quality MT systems. In this work, we propose a novel approach that directly exploits the internal representations of a surrogate multilingual MT model to distinguish between human and machine-translated sentences. Experimental results show that our method outperforms current state-of-the-art techniques, particularly for non-English language pairs, achieving gains of at least 5 percentage points of accuracy.
Cristian García-Romero, Miquel Esplà-Gomis, Felipe Sánchez-Martínez
Trans. Assoc. Comput. Linguistics3
2025 DeMINT: Automated Language Debriefing for English Learners via AI Chatbot Analysis of Meeting Transcripts
abstract
The objective of the DeMINT project is to develop a conversational tutoring system aimed at enhancing non-native English speakers’ language skills through post-meeting analysis of the transcriptions of video conferences in which they have participated. This paper describes the model developed and the results obtained through a human evaluation conducted with learners of English as a second language.
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz
MTSummit (2)2
2025 FLORES+ Mayas: Generating Textual Resources to Foster the Development of Language Technologies for Mayan Languages
abstract
A significant percentage of the population of Guatemala and Mexico belongs to various Mayan indigenous communities, for whom language barriers lead to social, economic, and digital exclusion. The Mayan languages spoken by these communities remain severely underrepresented in terms of digital resources, which prevents them from leveraging the latest advances in artificial intelligence. This project addresses that problem by means of: 1) the digitisation and release of multiple printed linguistic resources; 2) the development of a high-quality parallel machine translation (MT) evaluation corpus for six Mayan languages. In doing so, we are paving the way for the development of MT systems that will facilitate the access for Mayan speakers to essential services such as healthcare or legal aid. The resources are produced with the essential participation of indigenous communities, whereby native speakers provide the necessary translation services, QA, and linguistic expertise. The project is funded by the Google Academic Research Awards and carried out in collaboration with the Proyecto Lingüístico Francisco Marroquín Foundation in Guatemala.
Andrés Lou, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena
MTSummit (2)3
2024 Lightweight neural translation technologies for low-resource languages
abstract
The LiLowLa (“Lightweight neural translation technologies for low-resource languages”) project aims to enhance machine translation (MT) and translation memory (TM) technologies, particularly for low-resource language pairs, where adequate linguistic resources are scarce. The project started in September 2022 and will run till August 2025.
Felipe Sánchez-Martínez, Juan Antonio Pérez-Ortiz, Víctor M. Sánchez-Cartagena, Andrés Lou, Cristian García-Romero, Aarón Galiano Jiménez, Miquel Esplà-Gomis
EAMT (2)1
2024 Curated Datasets and Neural Models for Machine Translation of Informal Registers between Mayan and Spanish Vernaculars
abstract
Andrés Lou, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Víctor Sánchez-Cartagena. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Andrés Lou, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Víctor M. Sánchez-Cartagena
NAACL-HLT3
2024 Non-Fluent Synthetic Target-Language Data Improve Neural Machine Translation
abstract
When the amount of parallel sentences available to train a neural machine translation is scarce, a common practice is to generate new synthetic training samples from them. A number of approaches have been proposed to produce synthetic parallel sentences that are similar to those in the parallel data available. These approaches work under the assumption that non-fluent target-side synthetic training samples can be harmful and may deteriorate translation performance. Even so, in this paper we demonstrate that synthetic training samples with non-fluent target sentences can improve translation performance if they are used in a multilingual machine translation framework as if they were sentences in another language. We conducted experiments on ten low-resource and four high-resource translation tasks and found out that this simple approach consistently improves translation performance as compared to state-of-the-art methods for generating synthetic training samples similar to those found in corpora. Furthermore, this improvement is independent of the size of the original training corpus, the resulting systems are much more robust against domain shift and produce less hallucinations.
Víctor M. Sánchez-Cartagena, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Exploiting large pre-trained models for low-resource neural machine translation
abstract
Pre-trained models have drastically changed the field of natural language processing by providing a way to leverage large-scale language representations to various tasks. Some pre-trained models offer general-purpose representations, while others are specialized in particular tasks, like neural machine translation (NMT). Multilingual NMT-targeted systems are often fine-tuned for specific language pairs, but there is a lack of evidence-based best-practice recommendations to guide this process. Moreover, the trend towards even larger pre-trained models has made it challenging to deploy them in the computationally restrictive environments typically found in developing regions where low-resource languages are usually spoken. We propose a pipeline to tune the mBART50 pre-trained model to 8 diverse low-resource language pairs, and then distil the resulting system to obtain lightweight and more sustainable models. Our pipeline conveniently exploits back-translation, synthetic corpus filtering, and knowledge distillation to deliver efficient, yet powerful bilingual translation models 13 times smaller than the original pre-trained ones, but with close performance in terms of BLEU.
Aarón Galiano Jiménez, Felipe Sánchez-Martínez, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz
EAMT2
2022 MultitraiNMT Erasmus+ project: Machine Translation Training for multilingual citizens (multitrainmt.eu)
abstract
The MultitraiNMT Erasmus+ project has developed an open innovative syl-labus in machine translation, focusing on neural machine translation (NMT) and targeting both language learners and translators. The training materials include an open access coursebook with more than 250 activities and a pedagogical NMT interface called MutNMT that allows users to learn how neural machine translation works. These materials will allow students to develop the technical and ethical skills and competences required to become informed, critical users of machine translation in their own language learn-ing and translation practice. The pro-ject started in July 2019 and it will end in July 2022.
Mikel L. Forcada, Pilar Sánchez-Gijón, Dorothy Kenny, Felipe Sánchez-Martínez, Juan Antonio Pérez-Ortiz, Riccardo Superbo, Gema Ramírez-Sánchez, Olga Torres-Hostench, Caroline Rossi
EAMT4
2022 GoURMET - Machine Translation for Low-Resourced Languages
abstract
The GoURMET project, funded by the European Commission’s H2020 program (under grant agreement 825299), develops models for machine translation, in particular for low-resourced languages. Data, models and software releases as well as the GoURMET Translate Tool are made available as open source.
Peggy van der Kreeft, Alexandra Birch, Sevi Sariisik, Felipe Sánchez-Martínez, Wilker Aziz
EAMT4
2022 Cross-lingual neural fuzzy matching for exploiting target-language monolingual corpora in computer-aided translation
abstract
Computer-aided translation (CAT) tools based on translation memories (MT) play a prominent role in the translation workflow of professional translators.However, the reduced availability of in-domain TMs, as compared to indomain monolingual corpora, limits its adoption for a number of translation tasks.In this paper, we introduce a novel neural approach aimed at overcoming this limitation by exploiting not only TMs, but also in-domain targetlanguage (TL) monolingual corpora, and still enabling a similar functionality to that offered by conventional TM-based CAT tools.Our approach relies on cross-lingual sentence embeddings to retrieve translation proposals from TL monolingual corpora, and on a neural model to estimate their post-editing effort.The paper presents an automatic evaluation of these techniques on four language pairs that shows that our approach can successfully exploit monolingual texts in a TM-based CAT environment, increasing the amount of useful translation proposals, and that our neural model for estimating the post-editing effort enables the combination of translation proposals obtained from monolingual corpora and from TMs in the usual way.A human evaluation performed on a single language pair confirms the results of the automatic evaluation and seems to indicate that the translation proposals retrieved with our approach are more useful than what the automatic evaluation shows.
Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez
EMNLP4
2022 Fuzzy-Match Repair Guided by Quality Estimation
abstract
Computer-aided translation tools based on translation memories are widely used to assist professional translators. A translation memory (TM) consists of a set of translation units (TU) made up of source- and target-language segment pairs. For the translation of a new source segment$s^{\prime }$, these tools search the TM and retrieve the TUs$(s,t)$whose source segments are more similar to$s^{\prime }$. The translator then chooses a TU and edit the target segment$t$to turn it into an adequate translation of$s^{\prime }$.Fuzzy-match repair(FMR) techniques can be used to automatically modify the parts of$t$that need to be edited. We describe a language-independent FMR method that first uses machine translation to generate, given$s^{\prime }$and$(s,t)$, a set of candidate fuzzy-match repaired segments, and then chooses the best one by estimating their quality. An evaluation on three different language pairs shows that the selected candidate is a good approximation to the best (oracle) candidate produced and is closer to reference translations than machine-translated segments and unrepaired fuzzy matches ($t$). In addition, a single quality estimation model trained on a mix of data from all the languages performs well on any of the languages used.
John E. Ortega, Mikel L. Forcada, Felipe Sánchez-Martínez
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Rethinking Data Augmentation for Low-Resource Neural Machine Translation: A Multi-Task Learning Approach
abstract
In the context of neural machine translation, data augmentation (DA) techniques may be used for generating additional training samples when the available parallel data are scarce.Many DA approaches aim at expanding the support of the empirical data distribution by generating new sentence pairs that contain infrequent words, thus making it closer to the true data distribution of parallel sentences.In this paper, we propose to follow a completely different approach and present a multi-task DA approach in which we generate new sentence pairs with transformations, such as reversing the order of the target sentence, which produce unfluent target sentences.During training, these augmented sentences are used as auxiliary tasks in a multi-task framework with the aim of providing new contexts where the target prefix is not informative enough to predict the next word.This strengthens the encoder and forces the decoder to pay more attention to the source representations of the encoder.Experiments carried out on six lowresource translation tasks show consistent improvements over the baseline and over DA methods aiming at extending the support of the empirical data distribution.The systems trained with our approach rely more on the source tokens, are more robust against domain shift and suffer less hallucinations.
Víctor M. Sánchez-Cartagena, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez
EMNLP (1)4
2021 Surprise Language Challenge: Developing a Neural Machine Translation System between Pashto and English in Two Months
abstract
In the media industry and the focus of global reporting can shift overnight. There is a compelling need to be able to develop new machine translation systems in a short period of time and in order to more efficiently cover quickly developing stories. As part of the EU project GoURMET and which focusses on low-resource machine translation and our media partners selected a surprise language for which a machine translation system had to be built and evaluated in two months(February and March 2021). The language selected was Pashto and an Indo-Iranian language spoken in Afghanistan and Pakistan and India. In this period we completed the full pipeline of development of a neural machine translation system: data crawling and cleaning and aligning and creating test sets and developing and testing models and and delivering them to the user partners. In this paperwe describe rapid data creation and experiments with transfer learning and pretraining for this low-resource language pair. We find that starting from an existing large model pre-trained on 50languages leads to far better BLEU scores than pretraining on one high-resource language pair with a smaller model. We also present human evaluation of our systems and which indicates that the resulting systems perform better than a freely available commercial system when translating from English into Pashto direction and and similarly when translating from Pashto into English.
Alexandra Birch, Barry Haddow, Antonio Valerio Miceli Barone, Jindrich Helcl, Jonas Waldendorf, Felipe Sánchez-Martínez, Mikel L. Forcada, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Miquel Esplà-Gomis, Wilker Aziz, Lina Murady, Sevi Sariisik, Peggy van der Kreeft, Kay Macquarrie
MTSummit (1)6
2020 Understanding the effects of word-level linguistic annotations in under-resourced neural machine translation
abstract
This paper studies the effects of word-level linguistic annotations in under-resourced neural machine translation, for which there is incomplete evidence in the literature.The study covers eight language pairs, different training corpus sizes, two architectures and three types of annotation: dummy tags (with no linguistic information at all), part-of-speech tags, and morpho-syntactic description tags, which consist of part of speech and morphological features.These linguistic annotations are interleaved in the input or output streams as a single tag placed before each word.In order to measure the performance under each scenario, we use automatic evaluation metrics and perform automatic error classification.Our experiments show that, in general, source-language annotations are helpful and morpho-syntactic descriptions outperform part of speech for some language pairs.On the contrary, when words are annotated in the target language, part-of-speech tags systematically outperform morpho-syntactic description tags in terms of automatic evaluation metrics, even though the use of morpho-syntactic description tags improves the grammaticality of the output.We provide a detailed analysis of the reasons behind this result.
Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez
COLING3
2020 A multi-source approach for Breton-French hybrid machine translation
abstract
Corpus-based approaches to machine translation (MT) have difficulties when the amount of parallel corpora to use for training is scarce, especially if the languages involved in the translation are highly inflected. This problem can be addressed from different perspectives, including data augmentation, transfer learning, and the use of additional resources, such as those used in rule-based MT. This paper focuses on the hybridisation of rule-based MT and neural MT for the Breton–French under-resourced language pair in an attempt to study to what extent the rule-based MT resources help improve the translation quality of the neural MT system for this particular under-resourced language pair. We combine both translation approaches in a multi-source neural MT architecture and find out that, even though the rule-based system has a low performance according to automatic evaluation metrics, using it leads to improved translation quality.
Víctor M. Sánchez-Cartagena, Mikel L. Forcada, Felipe Sánchez-Martínez
EAMT3
2020 An English-Swahili parallel corpus and its use for neural machine translation in the news domain
abstract
This paper describes our approach to create a neural machine translation system to translate between English and Swahili (both directions) in the news domain, as well as the process we followed to crawl the necessary parallel corpora from the Internet. We report the results of a pilot human evaluation performed by the news media organisations participating in the H2020 EU-funded project GoURMET.
Felipe Sánchez-Martínez, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Mikel L. Forcada, Miquel Esplà-Gomis, Andrew Secker, Susie Coleman, Julie Wall
EAMT1
2019 Global Under-Resourced Media Translation (GoURMET)
Alexandra Birch, Barry Haddow, Ivan Titov 0001, Antonio Valerio Miceli Barone, Rachel Bawden, Felipe Sánchez-Martínez, Mikel L. Forcada, Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Wilker Aziz, Andrew Secker, Peggy van der Kreeft
MTSummit (2)6
2019 Improving Translations by Combining Fuzzy-Match Repair with Automatic Post-Editing
John E. Ortega, Felipe Sánchez-Martínez, Marco Turchi, Matteo Negri
MTSummit (1)2
2018 Proceedings of the 21st Annual Conference of the European Association for Machine Translation
Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Miquel Esplà-Gomis, Maja Popovic, Celia Rico, Joachim Van den Bogaert, Mikel L. Forcada
EAMT2
2017 One-parameter models for sentence-level post-editing effort estimation
Mikel L. Forcada, Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Lucia Specia
MTSummit (1)3
2016 Integrating Rules and Dictionaries from Shallow-Transfer Machine Translation into Phrase-Based Statistical Machine Translation
abstract
We describe a hybridisation strategy whose objective is to integrate linguistic resources from shallow-transfer rule-based machine translation (RBMT) into phrase-based statistical machine translation (PBSMT). It basically consists of enriching the phrase table of a PBSMT system with bilingual phrase pairs matching transfer rules and dictionary entries from a shallow-transfer RBMT system. This new strategy takes advantage of how the linguistic resources are used by the RBMT system to segment the source-language sentences to be translated, and overcomes the limitations of existing hybrid approaches that treat the RBMT systems as a black box. Experimental results confirm that our approach delivers translations of higher quality than existing ones, and that it is specially useful when the parallel corpus available for training the SMT system is small or when translating out-of-domain texts that are well covered by the RBMT dictionaries. A combination of this approach with a recently proposed unsupervised shallow-transfer rule inference algorithm results in a significantly greater translation quality than that of a baseline PBSMT; in this case, the only hand-crafted resource used are the dictionaries commonly used in RBMT. Moreover, the translation quality achieved by the hybrid system built with automatically inferred rules is similar to that obtained by those built with hand-crafted rules.
Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez
J. Artif. Intell. Res.3
2015 Using on-line available sources of bilingual information for word-level machine translation quality estimation
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada
EAMT2
2015 A general framework for minimizing translation effort: towards a principled combination of translation technologies in computer-aided translation
Mikel L. Forcada, Felipe Sánchez-Martínez
EAMT2
2015 Unsupervised training of maximum-entropy models for lexical selection in rule-based machine translation
Francis M. Tyers, Felipe Sánchez-Martínez, Mikel L. Forcada
EAMT2
2015 Linguistically-Enhanced Search over an Open Diachronic Corpus
Rafael C. Carrasco, Isabel Martínez-Sempere, Enrique Mollá-Gandía, Felipe Sánchez-Martínez, Gustavo Candela, Maria Pilar Escobar Esteban
ECIR4
2015 A generalised alignment template formalism and its application to the inference of shallow-transfer machine translation rules from scarce bilingual corpora
Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez
Comput. Speech Lang.3
2015 Using Machine Translation to Provide Target-Language Edit Hints in Computer Aided Translation Based on Translation Memories
abstract
This paper explores the use of general-purpose machine translation (MT) in assisting the users of computer-aided translation (CAT) systems based on translation memory (TM) to identify the target words in the translation proposals that need to be changed (either replaced or removed) or kept unedited, a task we term as "word-keeping recommendation". MT is used as a black box to align source and target sub-segments on the fly in the translation units (TUs) suggested to the user. Source-language (SL) and target-language (TL) segments in the matching TUs are segmented into overlapping sub-segments of variable length and machine-translated into the TL and the SL, respectively. The bilingual sub-segments obtained and the matching between the SL segment in the TU and the segment to be translated are employed to build the features that are then used by a binary classifier to determine the target words to be changed and those to be kept unedited. In this approach, MT results are never presented to the translator. Two approaches are presented in this work: one using a word-keeping recommendation system which can be trained on the TM used with the CAT system, and a more basic approach which does not require any training. Experiments are conducted by simulating the translation of texts in several language pairs with corpora belonging to different domains and using three different MT systems. We compare the performance obtained to that of previous works that have used statistical word alignment for word-keeping recommendation, and show that the MT-based approaches presented in this paper are more accurate in most scenarios. In particular, our results confirm that the MT-based approaches are better than the alignment-based approach when using models trained on out-of-domain TMs. Additional experiments were performed to check how dependent the MT-based recommender is on the language pair and MT system used for training. These experiments confirm a high degree of reusability of the recommendation models across various MT systems, but a low level of reusability across language pairs.
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada
J. Artif. Intell. Res.2
2014 An efficient method to assist non-expert users in extending dictionaries by assigning stems and inflectional paradigms to unknknown words
Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Felipe Sánchez-Martínez, Rafael C. Carrasco, Mikel L. Forcada, Juan Antonio Pérez-Ortiz
EAMT3
2013 Generalized Biwords for Bitext Compression and Translation Spotting: Extended Abstract
Felipe Sánchez-Martínez, Rafael C. Carrasco, Miguel A. Martínez-Prieto, Joaquín Adiego
IJCAI1
2012 Flexible finite-state lexical selection for rule-based machine translation
Francis M. Tyers, Felipe Sánchez-Martínez, Mikel L. Forcada
EAMT2
2012 Generalized Biwords for Bitext Compression and Translation Spotting
abstract
Large bilingual parallel texts (also known as bitexts) are usually stored in a compressed form, and previous work has shown that they can be more efficiently compressed if the fact that the two texts are mutual translations is exploited. For example, a bitext can be seen as a sequence of biwords ---pairs of parallel words with a high probability of co-occurrence--- that can be used as an intermediate representation in the compression process. However, the simple biword approach described in the literature can only exploit one-to-one word alignments and cannot tackle the reordering of words. We therefore introduce a generalization of biwords which can describe multi-word expressions and reorderings. We also describe some methods for the binary compression of generalized biword sequences, and compare their performance when different schemes are applied to the extraction of the biword sequence. In addition, we show that this generalization of biwords allows for the implementation of an efficient algorithm to look on the compressed bitext for words or text segments in one of the texts and retrieve their counterpart translations in the other text ---an application usually referred to as translation spotting--- with only some minor modifications in the compression algorithm.
Felipe Sánchez-Martínez, Rafael C. Carrasco, Miguel A. Martínez-Prieto, Joaquín Adiego
J. Artif. Intell. Res.1
2011 Using word alignments to assist computer-aided translation users by marking which target-side words to change or keep unedited
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada
EAMT2
2011 Choosing the best machine translation system to translate a sentence by using only source-language information
Felipe Sánchez-Martínez
EAMT1
2011 Using machine translation in computer-aided translation to suggest the target-side words to change
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada
MTSummit2
2011 Integrating shallow-transfer rules into phrase-based statistical machine translation
Víctor M. Sánchez-Cartagena, Felipe Sánchez-Martínez, Juan Antonio Pérez-Ortiz
MTSummit2
2011 Apertium: a free/open-source platform for rule-based machine translation
Mikel L. Forcada, Mireia Ginestí-Rosell, Jacob Nordfalk, Jim O'Regan, Sergio Ortiz-Rojas, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Gema Ramírez-Sánchez, Francis M. Tyers
Mach. Transl.7
2011 Free/open-source machine translation: preface
Felipe Sánchez-Martínez, Mikel L. Forcada
Mach. Transl.1
2010 Modelling Parallel Texts for Boosting Compression
abstract
Bilingual parallel corpora, also know as bitexts, convey the same information in two different languages. This implies that when modelling bitexts one can take advantage of the fact that there exists a relation between both texts; the text alignment task allow to establish such relationship. In this paper we propose different approaches that use words and biwords (pairs made of two words, each one from a different text) as representation symbolic units. The properties of these approaches are analyzed from a statistical point of view and tested as a preprocessing step to general purpose compressors. The results obtained suggest interesting conclusions concerning the use of both words and biwords. When encoded models are used as compression boosters we achieve compression ratios improving state-of-the-art compressors up to 6.5 percentage points, being up to 40% faster.
Joaquín Adiego, Miguel A. Martínez-Prieto, Javier E. Hoyos-Torío, Felipe Sánchez-Martínez
DCC4
2010 Philipp Koehn, Statistical machine translation - Cambridge University Press, 2010, Hardcover, xii + 433 pages, ISBN 978-0-521-87415-1, Price: EUR 40, USD 60
Felipe Sánchez-Martínez, Juan Antonio Pérez-Ortiz
Mach. Transl.1
2009 On the Use of Word Alignments to Enhance Bitext Compression
abstract
This paper describes a novel approach for bilingual parallel corpora (bitexts) compression. The approach takes advantage of the fact that the two texts that form a bitext are mutual translations. First, the two texts are aligned both at the sentence and the word level. Then, word alignments are used to define biwords, that is, pairs of two words, each one from a different text, that are mutual translations. Finally, a biword-based PPM compressor is applied. The results obtained compressing the two texts of the bitext together improve the compression ratios achieved when both texts are independently compressed through a word-based PPM compressor; thus, saving storage and transmission costs.
Miguel A. Martínez-Prieto, Joaquín Adiego, Felipe Sánchez-Martínez, Pablo de la Fuente, Rafael C. Carrasco
DCC3
2009 Marker-Based Filtering of Bilingual Phrase Pairs for SMT
Felipe Sánchez-Martínez, Andy Way
EAMT1
2009 A Two-Level Structure for Compressing Aligned Bitexts
Joaquín Adiego, Nieves R. Brisaboa, Miguel A. Martínez-Prieto, Felipe Sánchez-Martínez
SPIRE4
2009 Inferring Shallow-Transfer Machine Translation Rules from Small Parallel Corpora
abstract
This paper describes a method for the automatic inference of structural transfer rules to be used in a shallow-transfer machine translation (MT) system from small parallel corpora. The structural transfer rules are based on alignment templates, like those used in statistical MT. Alignment templates are extracted from sentence-aligned parallel corpora and extended with a set of restrictions which are derived from the bilingual dictionary of the MT system and control their application as transfer rules. The experiments conducted using three different language pairs in the free/open-source MT platform Apertium show that translation quality is improved as compared to word-for-word translation (when no transfer rules are used), and that the resulting translation quality is close to that obtained using hand-coded transfer rules. The method we present is entirely unsupervised and benefits from information in the rest of modules of the MT system in which the inferred rules are applied.
Felipe Sánchez-Martínez, Mikel L. Forcada
J. Artif. Intell. Res.1
2008 Using target-language information to train part-of-speech taggers for machine translation
Felipe Sánchez-Martínez, Juan Antonio Pérez-Ortiz, Mikel L. Forcada
Mach. Transl.1
2005 An open-source shallow-transfer machine translation engine for the Romance languages of Spain
Antonio M. Corbí-Bellot, Mikel L. Forcada, Sergio Ortiz-Rojas, Juan Antonio Pérez-Ortiz, Gema Ramírez-Sánchez, Felipe Sánchez-Martínez, Iñaki Alegria, Aingeru Mayor, Kepa Sarasola
EAMT6