VLDB 2026 Research / reviewers in the wild / expert
Mikel L. Forcada
dblp:07/536
· DBLP profile ↗
52ranked-venue papers
11as first author
8since 2021 · last 2026
0000-0003-0843-6442ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 11 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Theory of computation · 2Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Machine translation · 98% Deep learning architectures and training · 1% Representation and self-supervised learning · 0% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
computer-assisted translation |
0.6 | 1 | 2022 | Fuzzy-Match Repair Guided by Quality Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Natural language and speech › Machine translation › computer-assisted translation
translation memory |
0.6 | 1 | 2022 | Fuzzy-Match Repair Guided by Quality Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation |
0.6 | 1 | 2022 | Fuzzy-Match Repair Guided by Quality Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Natural language and speech › Machine translation
parallel corpus mining |
0.4 | 1 | 2020 | ParaCrawl: Web-Scale Acquisition of Parallel Corpora · ACL 2020 |
Information retrieval
cross-language information retrieval |
0.4 | 1 | 2020 | ParaCrawl: Web-Scale Acquisition of Parallel Corpora · ACL 2020 |
Machine learning › Deep learning architectures and training
recursive neural network |
0.0 | 1 | 2001 | Simple Strategies to Encode Tree Automata in Sigmoid Recursive Neural Networks · IEEE Trans. Knowl. Data Eng. 2001 |
Automata and formal languages
tree automata |
0.0 | 1 | 2001 | Simple Strategies to Encode Tree Automata in Sigmoid Recursive Neural Networks · IEEE Trans. Knowl. Data Eng. 2001 |
Methods — techniques the papers use, named apart from their topics
web crawling · 0.9automatic alignment · 0.9quality estimation model · 0.6machine translation · 0.6threshold linear unit · 0.1sigmoid activation · 0.1one-hot encoding · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prompsit's API and CLI: planet-friendly, privacy-first, open-source translation services for everyoneabstractPrompsit Language Engineering is launching an updated API and CLI for its open-source, planet-friendly machine translation services. Operating on a freemium model, the tools offer free limited access alongside tiered pricing for advanced features like MT evaluation, quality estimation, corpus scoring, and multilingual dataset annotation. Lev Nikolaevich Berezhnoy, Gema Ramírez-Sánchez, Sergio Ortiz-Rojas, Mikel L. Forcada |
EAMT (2) | 4 |
| 2024 | SmartBiC: Smart Harvesting of Bilingual Corpora from the InternetabstractSmartBiC, an 18-month innovation project funded by the Spanish Government, aims at improving the full process of collecting, filtering and selecting in-domain parallel content to be used for machine translation and language model tuning purposes in industrial settings. Based on state-of-the-art technology in the free/open-source parallel web corpora harvester Bitextor, SmartBic develops a web-based application around it including novel components such as a language- and domain-focused crawler and a domain-specific corpora selector. SmartBic also addresses specific industrial use cases for individual components of the Bitextor pipeline, such as parallel data cleaning. Relevant improvements to the current Bitextor pipeline will be publicly released. Gema Ramírez-Sánchez, Sergio Ortiz-Rojas, Alicia Núñez Alcover, Tudor Nicolae Mateiu, Mikel L. Forcada, Pedro Orzas, Almudena Carrillo, Giuseppe Nolasco, Noelia Listón |
EAMT (2) | 5 |
| 2023 | MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languagesabstractWe present the most relevant results of the project MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages in its second year. To date, parallel and monolingual corpora have been produced for seven low-resourced European languages by crawling large amounts of textual data from selected top-level domains of the Internet; both human and automatic evaluation show its usefulness. In addition, several large language models pretrained on MaCoCu data have been published, as well as the code used to collect and curate the data. Marta Bañón, Malina Chichirau, Miquel Esplà-Gomis, Mikel L. Forcada, Aarón Galiano Jiménez, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Peter Rupnik, Vit Suchomel, Antonio Toral, Jaume Zaragoza-Bernabeu |
EAMT | 4 |
| 2022 | MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languagesabstractWe introduce the project “MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages”, funded by the Connecting Europe Facility, which is aimed at building monolingual and parallel corpora for under-resourced European languages. The approach followed consists of crawling large amounts of textual data from carefully selected top-level domains of the Internet, and then applying a curation and enrichment pipeline. In addition to corpora, the project will release successive versions of the free/open-source web crawling and curation software used. Marta Bañón, Miquel Esplà-Gomis, Mikel L. Forcada, Cristian García-Romero, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Peter Rupnik, Vit Suchomel, Antonio Toral, Tobias van der Werff, Jaume Zaragoza |
EAMT | 3 |
| 2022 | MultitraiNMT Erasmus+ project: Machine Translation Training for multilingual citizens (multitrainmt.eu)abstractThe MultitraiNMT Erasmus+ project has developed an open innovative syl-labus in machine translation, focusing on neural machine translation (NMT) and targeting both language learners and translators. The training materials include an open access coursebook with more than 250 activities and a pedagogical NMT interface called MutNMT that allows users to learn how neural machine translation works. These materials will allow students to develop the technical and ethical skills and competences required to become informed, critical users of machine translation in their own language learn-ing and translation practice. The pro-ject started in July 2019 and it will end in July 2022. Mikel L. Forcada, Pilar Sánchez-Gijón, Dorothy Kenny, Felipe Sánchez-Martínez, Juan Antonio Pérez-Ortiz, Riccardo Superbo, Gema Ramírez-Sánchez, Olga Torres-Hostench, Caroline Rossi |
EAMT | 1 |
| 2022 | Fuzzy-Match Repair Guided by Quality EstimationabstractComputer-aided translation tools based on translation memories are widely used to assist professional translators. A translation memory (TM) consists of a set of translation units (TU) made up of source- and target-language segment pairs. For the translation of a new source segment$s^{\prime }$, these tools search the TM and retrieve the TUs$(s,t)$whose source segments are more similar to$s^{\prime }$. The translator then chooses a TU and edit the target segment$t$to turn it into an adequate translation of$s^{\prime }$.Fuzzy-match repair(FMR) techniques can be used to automatically modify the parts of$t$that need to be edited. We describe a language-independent FMR method that first uses machine translation to generate, given$s^{\prime }$and$(s,t)$, a set of candidate fuzzy-match repaired segments, and then chooses the best one by estimating their quality. An evaluation on three different language pairs shows that the selected candidate is a good approximation to the best (oracle) candidate produced and is closer to reference translations than machine-translated segments and unrepaired fuzzy matches ($t$). In addition, a single quality estimation model trained on a mix of data from all the languages performs well on any of the languages used. John E. Ortega, Mikel L. Forcada, Felipe Sánchez-Martínez |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Free/Open-Source Machine Translation for the Low-Resource Languages of Spain (Invited Talk)abstractWhile machine translation has historically been rule-based, that is, based on dictionaries and rules written by experts, most present-day machine translation is corpus-based. In the last few years, statistical machine translation, the dominant corpus-based approach, has been displaced by neural machine translation in most applications, in view of the better results reported, particularly for languages with very different syntax. But both statistical and neural machine translation need to be trained on large amounts of parallel data, that is, sentences in one language carefully paired with their translations in their other language, and this is a resource that may not be available for some low-resource languages. While some of the languages of Spain may be considered to be reasonably endowed with parallel corpora connecting them to Spanish or even to English - Basque, Catalan, Galician -, and are well-served with machine translation systems, there are many other languages which cannot afford them such as Aranese Occitan, Aragonese, or Asturian/Leonese. Fortunately, languages in this last group belong to the Romance language family, as Spanish does, and this makes translation from and into Spanish under a rule-based paradigm the only feasible approach. After describing briefly the main machine translation paradigms, I will describe the Apertium free/open-source rule-based machine translation platform, which has been used to build machine translation systems for these low-resource languages of Spain, indeed, sometimes the only ones available. The free/open-source setting has made linguistic data for these languages available for anyone in their linguistic communities to build other linguistic technologies for these low-resourced languages. For example, the Apertium family of bilingual and monolingual data has been converted into RDF and they have been made accessible on the Web as linked data. Mikel L. Forcada |
LDK | 1 |
| 2021 | Surprise Language Challenge: Developing a Neural Machine Translation System between Pashto and English in Two MonthsabstractIn the media industry and the focus of global reporting can shift overnight. There is a compelling need to be able to develop new machine translation systems in a short period of time and in order to more efficiently cover quickly developing stories. As part of the EU project GoURMET and which focusses on low-resource machine translation and our media partners selected a surprise language for which a machine translation system had to be built and evaluated in two months(February and March 2021). The language selected was Pashto and an Indo-Iranian language spoken in Afghanistan and Pakistan and India. In this period we completed the full pipeline of development of a neural machine translation system: data crawling and cleaning and aligning and creating test sets and developing and testing models and and delivering them to the user partners. In this paperwe describe rapid data creation and experiments with transfer learning and pretraining for this low-resource language pair. We find that starting from an existing large model pre-trained on 50languages leads to far better BLEU scores than pretraining on one high-resource language pair with a smaller model. We also present human evaluation of our systems and which indicates that the resulting systems perform better than a freely available commercial system when translating from English into Pashto direction and and similarly when translating from Pashto into English. Alexandra Birch, Barry Haddow, Antonio Valerio Miceli Barone, Jindrich Helcl, Jonas Waldendorf, Felipe Sánchez-Martínez, Mikel L. Forcada, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Miquel Esplà-Gomis, Wilker Aziz, Lina Murady, Sevi Sariisik, Peggy van der Kreeft, Kay Macquarrie |
MTSummit (1) | 7 |
| 2020 | ParaCrawl: Web-Scale Acquisition of Parallel CorporaabstractMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, Jaume Zaragoza. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz-Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strong, Brian Thompson 0001, William Waites, Dion Wiggins, Jaume Zaragoza |
ACL | 7 |
| 2020 | A multi-source approach for Breton-French hybrid machine translationabstractCorpus-based approaches to machine translation (MT) have difficulties when the amount of parallel corpora to use for training is scarce, especially if the languages involved in the translation are highly inflected. This problem can be addressed from different perspectives, including data augmentation, transfer learning, and the use of additional resources, such as those used in rule-based MT. This paper focuses on the hybridisation of rule-based MT and neural MT for the Breton–French under-resourced language pair in an attempt to study to what extent the rule-based MT resources help improve the translation quality of the neural MT system for this particular under-resourced language pair. We combine both translation approaches in a multi-source neural MT architecture and find out that, even though the rule-based system has a low performance according to automatic evaluation metrics, using it leads to improved translation quality. Víctor M. Sánchez-Cartagena, Mikel L. Forcada, Felipe Sánchez-Martínez |
EAMT | 2 |
| 2020 | An English-Swahili parallel corpus and its use for neural machine translation in the news domainabstractThis paper describes our approach to create a neural machine translation system to translate between English and Swahili (both directions) in the news domain, as well as the process we followed to crawl the necessary parallel corpora from the Internet. We report the results of a pilot human evaluation performed by the news media organisations participating in the H2020 EU-funded project GoURMET. Felipe Sánchez-Martínez, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Mikel L. Forcada, Miquel Esplà-Gomis, Andrew Secker, Susie Coleman, Julie Wall |
EAMT | 4 |
| 2019 | Global Under-Resourced Media Translation (GoURMET)
Alexandra Birch, Barry Haddow, Ivan Titov 0001, Antonio Valerio Miceli Barone, Rachel Bawden, Felipe Sánchez-Martínez, Mikel L. Forcada, Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Wilker Aziz, Andrew Secker, Peggy van der Kreeft |
MTSummit (2) | 7 |
| 2019 | ParaCrawl: Web-scale parallel corpora for the languages of the EU
Miquel Esplà-Gomis, Mikel L. Forcada, Gema Ramírez-Sánchez, Hieu Hoang |
MTSummit (2) | 2 |
| 2018 | Proceedings of the 21st Annual Conference of the European Association for Machine Translation
Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Miquel Esplà-Gomis, Maja Popovic, Celia Rico, Joachim Van den Bogaert, Mikel L. Forcada |
EAMT | 8 |
| 2018 | Learning to use machine translation on the Translation Commons Learn portalabstractWe describe the Learn portal of Translation Commons (TC), a self-managed community of volunteer translators community aimed at sharing tools, resources and initiatives for the translation community as a whole. Members are encouraged to upload and share their free resources on the platform and to create free courses and tutorials. Specifically there are no educational material on machine translation yet and we invite experts to contribute. Jeannette Stewart, Mikel L. Forcada |
EAMT | 2 |
| 2018 | Editors' foreword to the invited issue on SMT and NMT
Andy Way, Mikel L. Forcada |
Mach. Transl. | 2 |
| 2017 | One-parameter models for sentence-level post-editing effort estimation
Mikel L. Forcada, Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Lucia Specia |
MTSummit (1) | 1 |
| 2017 | Usefulness of MT output for comprehension - an analysis from the point of view of linguistic intercomprehension
Kenneth Jordan Núñez, Mikel L. Forcada, Esteve Clua |
MTSummit (1) | 2 |
| 2016 | Stand-off Annotation of Web Content as a Legally Safer Alternative to Bitext Crawling for Distribution
Mikel L. Forcada, Miquel Esplà-Gomis, Juan Antonio Pérez-Ortiz |
EAMT | 1 |
| 2015 | Evaluating machine translation for assimilation via a gap-filling task
Ekaterina Ageeva, Mikel L. Forcada, Francis M. Tyers, Juan Antonio Pérez-Ortiz |
EAMT | 2 |
| 2015 | Using on-line available sources of bilingual information for word-level machine translation quality estimation
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada |
EAMT | 3 |
| 2015 | A general framework for minimizing translation effort: towards a principled combination of translation technologies in computer-aided translation
Mikel L. Forcada, Felipe Sánchez-Martínez |
EAMT | 1 |
| 2015 | Abu-MaTran: Automatic building of Machine Translation
Antonio Toral, Flammie A. Pirinen, Andy Way, Gema Ramírez-Sánchez, Sergio Ortiz-Rojas, Raphaël Rubino, Miquel Esplà-Gomis, Mikel L. Forcada, Vassilis Papavassiliou, Prokopis Prokopidis, Nikola Ljubesic |
EAMT | 8 |
| 2015 | Unsupervised training of maximum-entropy models for lexical selection in rule-based machine translation
Francis M. Tyers, Felipe Sánchez-Martínez, Mikel L. Forcada |
EAMT | 3 |
| 2015 | Using Machine Translation to Provide Target-Language Edit Hints in Computer Aided Translation Based on Translation MemoriesabstractThis paper explores the use of general-purpose machine translation (MT) in assisting the users of computer-aided translation (CAT) systems based on translation memory (TM) to identify the target words in the translation proposals that need to be changed (either replaced or removed) or kept unedited, a task we term as "word-keeping recommendation". MT is used as a black box to align source and target sub-segments on the fly in the translation units (TUs) suggested to the user. Source-language (SL) and target-language (TL) segments in the matching TUs are segmented into overlapping sub-segments of variable length and machine-translated into the TL and the SL, respectively. The bilingual sub-segments obtained and the matching between the SL segment in the TU and the segment to be translated are employed to build the features that are then used by a binary classifier to determine the target words to be changed and those to be kept unedited. In this approach, MT results are never presented to the translator. Two approaches are presented in this work: one using a word-keeping recommendation system which can be trained on the TM used with the CAT system, and a more basic approach which does not require any training. Experiments are conducted by simulating the translation of texts in several language pairs with corpora belonging to different domains and using three different MT systems. We compare the performance obtained to that of previous works that have used statistical word alignment for word-keeping recommendation, and show that the MT-based approaches presented in this paper are more accurate in most scenarios. In particular, our results confirm that the MT-based approaches are better than the alignment-based approach when using models trained on out-of-domain TMs. Additional experiments were performed to check how dependent the MT-based recommender is on the language pair and MT system used for training. These experiments confirm a high degree of reusability of the recommendation models across various MT systems, but a low level of reusability across language pairs. Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada |
J. Artif. Intell. Res. | 3 |
| 2014 | An efficient method to assist non-expert users in extending dictionaries by assigning stems and inflectional paradigms to unknknown words
Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Felipe Sánchez-Martínez, Rafael C. Carrasco, Mikel L. Forcada, Juan Antonio Pérez-Ortiz |
EAMT | 5 |
| 2014 | On the annotation of TMX translation memories for advanced leveraging in computer-aided translation
Mikel L. Forcada |
LREC | 1 |
| 2012 | Flexible finite-state lexical selection for rule-based machine translation
Francis M. Tyers, Felipe Sánchez-Martínez, Mikel L. Forcada |
EAMT | 3 |
| 2011 | Using Example-Based MT to Support Statistical MT when Translating Homogeneous Data in a Resource-Poor Setting
Sandipan Dandapat, Sara Morrissey, Andy Way, Mikel L. Forcada |
EAMT | 4 |
| 2011 | Using word alignments to assist computer-aided translation users by marking which target-side words to change or keep unedited
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada |
EAMT | 3 |
| 2011 | Using machine translation in computer-aided translation to suggest the target-side words to change
Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Mikel L. Forcada |
MTSummit | 3 |
| 2011 | Apertium: a free/open-source platform for rule-based machine translation
Mikel L. Forcada, Mireia Ginestí-Rosell, Jacob Nordfalk, Jim O'Regan, Sergio Ortiz-Rojas, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Gema Ramírez-Sánchez, Francis M. Tyers |
Mach. Transl. | 1 |
| 2011 | Free/open-source machine translation: preface
Felipe Sánchez-Martínez, Mikel L. Forcada |
Mach. Transl. | 2 |
| 2009 | Incremental Construction of Minimal Tree Automata
Rafael C. Carrasco, Jan Daciuk, Mikel L. Forcada |
Algorithmica | 3 |
| 2009 | Inferring Shallow-Transfer Machine Translation Rules from Small Parallel CorporaabstractThis paper describes a method for the automatic inference of structural transfer rules to be used in a shallow-transfer machine translation (MT) system from small parallel corpora. The structural transfer rules are based on alignment templates, like those used in statistical MT. Alignment templates are extracted from sentence-aligned parallel corpora and extended with a set of restrictions which are derived from the bilingual dictionary of the MT system and control their application as transfer rules. The experiments conducted using three different language pairs in the free/open-source MT platform Apertium show that translation quality is improved as compared to word-for-word translation (when no transfer rules are used), and that the resulting translation quality is close to that obtained using hand-coded transfer rules. The method we present is entirely unsupervised and benefits from information in the rest of modules of the MT system in which the inferred rules are applied. Felipe Sánchez-Martínez, Mikel L. Forcada |
J. Artif. Intell. Res. | 2 |
| 2008 | Using target-language information to train part-of-speech taggers for machine translation
Felipe Sánchez-Martínez, Juan Antonio Pérez-Ortiz, Mikel L. Forcada |
Mach. Transl. | 3 |
| 2007 | An Implementation of Deterministic Tree Automata Minimization
Rafael C. Carrasco, Jan Daciuk, Mikel L. Forcada |
CIAA | 3 |
| 2006 | Automatic induction of bilingual resources from aligned parallel corpora: application to shallow-transfer machine translation
Helena de Medeiros Caseli, Maria das Graças Volpe Nunes, Mikel L. Forcada |
Mach. Transl. | 3 |
| 2005 | An open-source shallow-transfer machine translation engine for the Romance languages of Spain
Antonio M. Corbí-Bellot, Mikel L. Forcada, Sergio Ortiz-Rojas, Juan Antonio Pérez-Ortiz, Gema Ramírez-Sánchez, Felipe Sánchez-Martínez, Iñaki Alegria, Aingeru Mayor, Kepa Sarasola |
EAMT | 2 |
| 2004 | Book Review: Brian James Baer, Geoffrey S. Koby (eds), Beyond the Ivory Tower: Rethinking Translation Pedagogy. John Benjamins Publishing Co., Amsterdam/Philadelphia, 2003, xvi + 258 pp
Mikel L. Forcada |
Mach. Transl. | 1 |
| 2002 | Incremental Construction and Maintenance of Minimal Finite-State AutomataabstractDaciuk et al. [Computational Linguistics 26(1):3–16 (2000)] describe a method for constructing incrementally minimal, deterministic, acyclic finite-state automata (dictionaries) from sets of strings. But acyclic finite-state automata have limitations: For instance, if one wants a linguistic application to accept all possible integer numbers or Internet addresses, the corresponding finite-state automaton has to be cyclic. In this article, we describe a simple and equally efficient method for modifying any minimal finite-state automaton (be it acyclic or not) so that a string is added to or removed from the language it accepts; both operations are very important when dictionary maintenance is performed and solve the dictionary construction problem addressed by Daciuk et al. as a special case. The algorithms proposed here may be straightforwardly derived from the customary textbook constructions for the intersection and the complementation of finite-state automata; the algorithms exploit the special properties of the automata resulting from the intersection operation when one of the finite-state automata accepts a single string. Rafael C. Carrasco, Mikel L. Forcada |
Comput. Linguistics | 2 |
| 2001 | Online Symbolic-Sequence Prediction with Discrete-Time Recurrent Neural Networks
Juan Antonio Pérez-Ortiz, Jorge Calera-Rubio, Mikel L. Forcada |
ICANN | 3 |
| 2001 | The Spanish<>Catalan machine translation system interNOSTRUMabstractThis paper describes interNOSTRUM, a Spanish3Catalan machine translation system currently under development that achieves great speed through the use of finite-state technologies (so that it may be integrated with internet browsing) and a reasonable accuracy using an advanced morphological transfer strategy (to produce fast translation drafts ready for light postedition). Raul Canals-Marote, Anna Esteve-Guillén, Alicia Garrido-Alenda, M. Isabel Guardiola-Savall, Amaia Iturraspe-Bellver, Sandra Montserrat-Buendia, Sergio Ortiz-Rojas, Herminia Pastor-Pina, Pedro M. Pérez-Antón, Mikel L. Forcada |
MTSummit | 10 |
| 2001 | Online Text Prediction with Recurrent Neural Networks
Juan Antonio Pérez-Ortiz, Jorge Calera-Rubio, Mikel L. Forcada |
Neural Process. Lett. | 3 |
| 2001 | Simple Strategies to Encode Tree Automata in Sigmoid Recursive Neural NetworksabstractRecently, a number of authors have explored the use of recursive neural nets (RNN) for the adaptive processing of trees or tree-like structures. One of the most important language-theoretical formalizations of the processing of tree-structured data is that of deterministic finite-state tree automata (DFSTA). DFSTA may easily be realized as RNN using discrete-state units, such as the threshold linear unit. A recent result by J. Sima (1997) shows that any threshold linear unit operating on binary inputs can be implemented in an analog unit using a continuous activation function and bounded real inputs. The constructive proof finds a scaling factor for the weights and reestimates the bias accordingly. We explore the application of this result to simulate DFSTA in sigmoid RNN (that is, analog RNN using monotonically growing activation functions) and also present an alternative scheme for one-hot encoding of the input that yields smaller weight values, and therefore works at a lower saturation level. Rafael C. Carrasco, Mikel L. Forcada |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2000 | Stable Encoding of Finite-State Machines in Discrete-Time Recurrent Neural Nets with Sigmoid UnitsabstractThere has been a lot of interest in the use of discrete-time recurrent neural nets (DTRNN) to learn finite-state tasks, with interesting results regarding the induction of simple finite-state machines from input-output strings. Parallel work has studied the computational power of DTRNN in connection with finite-state computation. This article describes a simple strategy to devise stable encodings of finite-state machines in computationally capable discrete-time recurrent neural architectures with sigmoid units and gives a detailed presentation on how this strategy may be applied to encode a general class of finite-state machines in a variety of commonly used first- and second-order recurrent neural networks. Unlike previous work that either imposed some restrictions to state values or used a detailed analysis based on fixed-point attractors, our approach applies to any positive, bounded, strictly growing, continuous activation function and uses simple bounding criteria based on a study of the conditions under which a proposed encoding scheme guarantees that the DTRNN is actually behaving as a finite-state machine. Rafael C. Carrasco, Mikel L. Forcada, M. Ángeles Valdés-Muñoz, Ramón P. Ñeco |
Neural Comput. | 2 |
| 1999 | Neural learning of approximate simple regular languages
Mikel L. Forcada, Antonio M. Corbí-Bellot, Marco Gori, Marco Maggini |
ESANN | 1 |
| 1999 | Encoding of sequential translators in discrete-time recurrent neural nets
Ramón P. Ñeco, Mikel L. Forcada, Rafael C. Carrasco, M. Ángeles Valdés-Muñoz |
ESANN | 2 |
| 1999 | Recurrent neural networks can learn simple, approximate regular languagesabstractA number of researchers have shown that discrete-time recurrent neural networks (DTRNN) are capable of inferring deterministic finite automata from sets of example and counterexample strings; however, discrete algorithmic methods are much better at this task and clearly outperform DTRNN in terms of space and time complexity. We show how DTRNN may be used to learn not the exact language that explains the whole learning set but an approximate and much simpler language that explains a great majority of the examples by using simpler rules. This is accomplished by gradually varying the error function in such a way that the DTRNN is eventually allowed to classify clearly but incorrectly those strings that it has found to be difficult to learn, which are treated as exceptions. The results show that in this way, the DTRNN usually manages to learn a simplified approximate language. Mikel L. Forcada, Antonio M. Corbí-Bellot, Marco Gori, Marco Maggini |
IJCNN | 1 |
| 1997 | A comparison between recurrent neural network architectures for digital equalizationabstractThis paper shows a comparison between three different first-order recurrent neural network (RNN) architectures (fully recurrent, partially recurrent, and Elman (1990)), trained using the real-time recurrent learning (RTRL) algorithm and the GSM training sequence ratio (26/114) for digital equalization of 2-ary PAM signals. The results show no substantial effect of the particular architecture or the number of units on the overall performance. This is due to the assumption of a suboptimal equalization scheme by the RNNs, because of the learning algorithm. The results are compared to those obtained using a classical (decision-feedback equalizer) approach. Jorge D. Ortiz-Fuentes, Mikel L. Forcada |
ICASSP | 2 |
| 1995 | Learning the Initial State of a Second-Order Recurrent Neural Network during Regular-Language InferenceabstractRecent work has shown that second-order recurrent neural networks (2ORNNs) may be used to infer regular languages. This paper presents a modified version of the real-time recurrent learning (RTRL) algorithm used to train 2ORNNs, that learns the initial state in addition to the weights. The results of this modification, which adds extra flexibility at a negligible cost in time complexity, suggest that it may be used to improve the learning of regular languages when the size of the network is small. Mikel L. Forcada, Rafael C. Carrasco |
Neural Comput. | 1 |
| 1995 | A note on the Nagendraprasad-Wang-Gupta thinning algorithm
Rafael C. Carrasco, Mikel L. Forcada |
Pattern Recognit. Lett. | 2 |