VLDB 2026 Research / reviewers in the wild / expert
Mark Fishel
dblp:18/8157 · also Mark Fisel
· DBLP profile ↗
22ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OSCAIL-OpenScience Communication through AI in EU LanguagesabstractThe Anglocentric nature of scholarly communication has many implications, such as limiting publication, discoverability and access from other language communities (even for major languages); putting minoritized languages at risk in the academic domain; and excluding many from peer review. The OSCAIL project addresses these challenges by exploring how machine translation (MT) enhanced by large language model (LLM)–based technologies can support access to scientific knowledge. Outputs will include evaluation datasets, protocols and best practices for MT in scholarly communication, and a prototype integration of MT tools into Open Journal Systems, the world’s most widely used open-source scholarly publishing platform. Sheila Castilho, Susanna Fiorini, Lynne Bowker, Petr Motlícek, Joss Moorkens, Lieve Macken, Dairazalia Sanchez-Cortes, Janne Pölönen, Sami Syrjämäki, Mikael Laakso, Mark Fishel, Anastasia Stasenko |
EAMT (2) | 11 |
| 2026 | The Two Towers for Estonian-Centric and Finno-Ugric Machine TranslationabstractWe present two open-weight translation models for Estonian and its low-resource "relatives" in the Finno-Ugric language family. The training data includes 12 languages paired with Estonian as well as 23 more Finno-Ugric languages and varieties, ranging from mid-resource examples with tens of thousands of speakers to extremely low-resource critically endangered languages with less than a hundred speakers. The translation models use Unbabel Tower+ 2B and 9B as their starting point. We compare their performance on two benchmarks to DeepL and GPT-5.2 and show that in most cases we surpass the quality of DeepL and match or nearly match the quality of GPT-5.2’s output with just a fraction of the parameters. Among other contributions we also restore the paragraph structure of a massive synthetic multiparallel corpus for Estonian translation and use it in training the models. The resulting models, training scripts and training data are released openly (to be made public upon de-anonymization). Mark Fishel, Lisa Yankovskaya |
EAMT (1) | 1 |
| 2024 | Multilinguality or Back-translation? A Case Study with EstonianabstractMachine translation quality is highly reliant on large amounts of training data, and, when a limited amount of parallel data is available, synthetic back-translated or multilingual data can be used in addition. In this work, we introduce SynEst, a synthetic corpus of translations from 11 languages into Estonian which totals over 1 billion sentence pairs. Using this corpus, we investigate whether adding synthetic or English-centric additional data yields better translation quality for translation directions that do not include English. Our results show that while both strategies are effective, synthetic data gives better results. Our final models improve the performance of the baseline No Language Left Behind model while retaining its source-side multilinguality. Elizaveta Korotkova, Taido Purason, Agnes Luhtaru, Mark Fishel |
LREC/COLING | 4 |
| 2024 | No Error Left Behind: Multilingual Grammatical Error Correction with Pre-trained Translation ModelsabstractGrammatical Error Correction (GEC) enhances language proficiency and promotes effective communication, but research has primarily centered around English.We propose a simple approach to multilingual and low-resource GEC by exploring the potential of multilingual machine translation (MT) models for error correction.We show that MT models are not only capable of error correction out-of-the-box, but that they can also be fine-tuned to even better correction quality.Results show the effectiveness of this approach, with our multilingual model outperforming similar-sized mT5-based models and even competing favourably with larger models. Agnes Luhtaru, Elizaveta Korotkova, Mark Fishel |
EACL (1) | 3 |
| 2024 | Estonian-Centric Machine Translation: Data, Models, and ChallengesabstractMachine translation (MT) research is most typically English-centric. In recent years, massively multilingual translation systems have also been increasingly popular. However, efforts purposefully focused on less-resourced languages are less widespread. In this paper, we focus on MT from and into the Estonian language. First, emphasizing the importance of data availability, we generate and publicly release a back-translation corpus of over 2 billion sentence pairs. Second, using these novel data, we create MT models covering 18 translation directions, all either from or into Estonian. We re-use the encoder of the NLLB multilingual model and train modular decoders separately for each language, surpassing the original NLLB quality. Our resulting MT models largely outperform other open-source MT systems, including previous Estonian-focused efforts, and are released as part of this submission. Elizaveta Korotkova, Mark Fishel |
EAMT (1) | 2 |
| 2024 | SMUGRI-MT - Machine Translation System for Low-Resource Finno-Ugric LanguagesabstractWe introduce SMUGRI-MT, an online neural machine translation system that covers 20 low-resource Finno-Ugric languages, along with seven high-resource languages. Taido Purason, Aleksei Ivanov, Lisa Yankovskaya, Mark Fishel |
EAMT (2) | 4 |
| 2024 | Enabling Conversational Speech Synthesis using Noisy Spontaneous DataabstractIn recent years, the quality of text-to-speech models has increased significantly, but most text-to-speech solutions are trained on datasets of read speech and do not cover conversational speaking styles due to the lack of suitable training data.This paper explores options for creating multi-style speech synthesis using speech recognition datasets that contain samples of spontaneous speech and dialogues but may also include background noise and an insufficient number of samples per speaker.We develop an Estonian multi-speaker TTS system that increases prosodic variability on conversational inputs while still being able to synthesize read speech.We show that our proposed approach can be used to train models that can be controlled to produce conversational speech with little compromise on audio quality.We also highlight a potential multilingual use case to achieve cross-lingual speaker and style transfer to lowresource languages that lack stylistically diverse speech corpora. Liisa Rätsep, Rasmus Lellep, Mark Fishel |
INTERSPEECH | 3 |
| 2022 | MTee: Open Machine Translation Platform for Estonian GovernmentabstractWe present the MTee project - a research initiative funded via an Estonian public procurement to develop machine translation technology that is open-source and free of charge. The MTee project delivered an open-source platform serving state-of-the-art machine translation systems supporting four domains for six language pairs translating from Estonian into English, German, and Russian and vice-versa. The platform also features grammatical error correction and speech translation for Estonian and allows for formatted document translation and automatic domain detection. The software, data and training workflows for machine translation engines are all made publicly available for further use and research. Toms Bergmanis, Marcis Pinnis, Roberts Rozis, Janis Slapins, Valters Sics, Berta Bernane, Guntars Puzulis, Endijs Titomers, Andre Tättar, Taido Purason, Hele-Andra Kuulmets, Agnes Luhtaru, Liisa Rätsep, Maali Tars, Annika Laumets-Tättar, Mark Fishel |
EAMT | 16 |
| 2022 | National Language Technology Platform (NLTP): overall viewabstractThe work in progress on the CEF Action National Language Technology Platform (NLTP) is presented. The Action aims at combining the most advanced Language Technology (LT) tools and solutions in a new state-of-the-art, Artificial Intelli- gence (AI) driven, National Language Technology Platform (NLTP). Arturs Vasilevskis, Janis Ziedins, Marko Tadic, Zeljka Motika, Mark Fishel, Hrafn Loftsson, Jón Guðnason, Claudia Borg, Keith Cortis, Judie Attard, Donatienne Spiteri |
EAMT | 5 |
| 2020 | Unsupervised Quality Estimation for Neural Machine TranslationabstractQuality Estimation (QE) is an important component in making Machine Translation (MT) useful in real-world applications, as it is aimed to inform the user on the quality of the MT output at test time. Existing approaches require large amounts of expert annotated data, computation, and time for training. As an alternative, we devise an unsupervised approach to QE where no training or access to additional resources besides the MT system itself is required. Different from most of the current work that treats the MT system as a black box, we explore useful information that can be extracted from the MT system as a by-product of translation. By utilizing methods for uncertainty quantification, we achieve very good correlation with human judgments of quality, rivaling state-of-the-art supervised QE models. To evaluate our approach we collect the first dataset that enables work on both black-box and glass-box approaches to QE. Marina Fomicheva, Lisa Yankovskaya, Frédéric Blain, Francisco Guzmán, Mark Fishel, Nikolaos Aletras, Vishrav Chaudhary, Lucia Specia |
Trans. Assoc. Comput. Linguistics | 6 |
| 2018 | Multi-Domain Neural Machine TranslationabstractWe present an approach to neural machine translation (NMT) that supports multiple domains in a single model and allows switching between the domains when translating. The core idea is to treat text domainsasdistinctlanguagesandusemultilingual NMT methods to create multi-domain translation systems; we show that this approach results in significant translation quality gains over fine-tuning. We also explore whether the knowledge of pre-specified text domains is necessary; turns out that it is after all, but also that when it is not known quite high translation quality can be reached, and even higher than with known domains in some cases. Sander Tars, Mark Fishel |
EAMT | 2 |
| 2017 | Confidence through Attention
Matiss Rikters, Mark Fishel |
MTSummit (1) | 2 |
| 2014 | Handling technical OOVs in SMT
Mark Fishel, Rico Sennrich |
EAMT | 1 |
| 2014 | Machine Translation for Subtitling: A Large-Scale Evaluation
Thierry Etchegoyhen, Lindsay Bywood, Mark Fishel, Panayota Georgakopoulou, Gerard van Loenhout, Arantza del Pozo, Mirjam Sepesy Maucec, Anja Turner, Martin Volk 0001 |
LREC | 3 |
| 2012 | From Subtitles to Parallel Corpora
Mark Fishel, Panayota Georgakopoulou, Sergio Penkale, Volha Petukhova, Matej Rojc, Martin Volk 0001, Andy Way |
EAMT | 1 |
| 2012 | Automatic MT Error Analysis: Hjerson Helping Addicter
Jan Berka, Ondrej Bojar, Mark Fishel, Maja Popovic, Daniel Zeman |
LREC | 3 |
| 2012 | Terra: a Collection of Translation Error-Annotated Corpora
Mark Fishel, Ondrej Bojar, Maja Popovic |
LREC | 1 |
| 2012 | SUMAT: Data Collection and Parallel Corpus Compilation for Machine Translation of Subtitles
Volha Petukhova, Rodrigo Agerri, Mark Fishel, Sergio Penkale, Arantza del Pozo, Mirjam Sepesy Maucec, Andy Way, Panayota Georgakopoulou, Martin Volk 0001 |
LREC | 3 |
| 2010 | Linguistically Motivated Unsupervised Segmentation for Machine Translation
Mark Fishel, Harri Kirik |
LREC | 1 |
| 2010 | Simpler Is Better: Re-evaluation of Default Word Alignment Models in Statistical MT
Mark Fishel |
PACLIC | 1 |
| 2008 | Mixing and Blending Syntactic and Semantic Dependencies
Yvonne Samuelsson, Oscar Täckström, Sumithra Velupillai, Johan Eklund, Mark Fishel, Markus Saers |
CoNLL | 5 |
| 2008 | Experiments on Processing Overlapping Parallel Corpora
Mark Fishel, Heiki-Jaan Kaalep |
LREC | 1 |