EDBT 2026 Demo / reviewers in the wild / expert
Sheila Castilho
dblp:126/8641
· DBLP profile ↗
31ranked-venue papers
16as first author
14since 2021 · last 2026
0000-0002-8416-6555ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 16 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Translating Under Pressure: Domain-Aware LLMs for Crisis CommunicationabstractTimely and reliable multilingual communication is critical during natural and human-induced disasters, but developing effective solutions for crisis communication is limited by the scarcity of curated parallel data. We propose a domain-adaptive pipeline that expands a small reference corpus, by retrieving and filtering data from general corpora. We use the resulting dataset to fine-tune a small language model for crisis-domain translation and then apply preference optimization to bias outputs toward CEFR A2-level English. Automatic and human evaluation shows that this approach improves readability, while maintaining strong adequacy. Our results indicate that simplified English, combined with domain adaptation, can function as a practical lingua franca for emergency communication when full multilingual coverage is not feasible. Antonio Castaldo, Maria Carmen Staiano, Johanna Monti, Sheila Castilho, Francesca Chiusaroli |
EAMT (1) | 4 |
| 2026 | OSCAIL-OpenScience Communication through AI in EU LanguagesabstractThe Anglocentric nature of scholarly communication has many implications, such as limiting publication, discoverability and access from other language communities (even for major languages); putting minoritized languages at risk in the academic domain; and excluding many from peer review. The OSCAIL project addresses these challenges by exploring how machine translation (MT) enhanced by large language model (LLM)–based technologies can support access to scientific knowledge. Outputs will include evaluation datasets, protocols and best practices for MT in scholarly communication, and a prototype integration of MT tools into Open Journal Systems, the world’s most widely used open-source scholarly publishing platform. Sheila Castilho, Susanna Fiorini, Lynne Bowker, Petr Motlícek, Joss Moorkens, Lieve Macken, Dairazalia Sanchez-Cortes, Janne Pölönen, Sami Syrjämäki, Mikael Laakso, Mark Fishel, Anastasia Stasenko |
EAMT (2) | 1 |
| 2026 | VERA: A Platform for Automatic and Human Evaluation of Machine TranslationabstractWe present VERA, an easy-to-use platform for machine translation (MT) evaluation that integrates automatic and human evaluation within a single web environment. VERA supports standard reference based metrics and multiuser annotation following the Multidimensional Quality Metrics (MQM) Core framework. The platform enables the export of annotated corpora and the generation of final PDF reports summarizing both automatic and human evaluation results, including correlations between them. Sofía García González, Inés Quintana Raña, Jorge N. Afonso Cabido, Alberto Hernández Lado, German Rigau Claramunt, Sheila Castilho |
EAMT (2) | 6 |
| 2026 | Literacy-Grounded and Industry-Oriented Translation Training with LT-LiDERabstracthe Erasmus+-funded international research consortium LT-LiDER develops a range of digital training resources which are grounded in the overarching frameworks of digital and AI literacy and oriented towards practical application contexts in the language and translation industry. These resources can be implemented on a component basis or as a complete curriculum in higher-education language and translation classrooms. Janiça Hackenbuchner, Maria Isabel Rivas Ginel, Joss Moorkens, Sheila Castilho, Nora Aranberri, Sergi Alvarez-Vidal, María Do Campo Bayón, Ralph Krüger |
EAMT (2) | 4 |
| 2025 | Extending CREAMT: Leveraging Large Language Models for Literary Translation Post-EditingabstractPost-editing machine translation (MT) for creative texts, such as literature, requires balancing efficiency with the preservation of creativity and style. While neural MT systems struggle with these challenges, large language models (LLMs) offer improved capabilities for context-aware and creative translation. This study evaluates the feasibility of post-editing literary translations generated by LLMs. Using a custom research tool, we collaborated with professional literary translators to analyze editing time, quality, and creativity. Our results indicate that post-editing (PE) LLM-generated translations significantly reduce editing time compared to human translation while maintaining a similar level of creativity. The minimal difference in creativity between PE and MT, combined with substantial productivity gains, suggests that LLMs may effectively support literary translators. Antonio Castaldo, Sheila Castilho, Joss Moorkens, Johanna Monti |
MTSummit (1) | 2 |
| 2025 | UniOr PET: An Online Platform for Translation Post-EditingabstractUniOr PET is a browser-based platform for machine translation post-editing and a modern successor to the original PET tool. It features a user-friendly interface that records detailed editing actions, including time spent, additions, and deletions. Fully compatible with PET, UniOr PET introduces two advanced timers for more precise tracking of editing time and computes widely used metrics such as hTER, BLEU, and ChrF, providing comprehensive insights into translation quality and post-editing productivity. Designed with translators and researchers in mind, UniOr PET combines the strengths of its predecessor with enhanced functionality for efficient and user-friendly post-editing projects. Antonio Castaldo, Sheila Castilho, Joss Moorkens, Johanna Monti |
MTSummit (2) | 2 |
| 2025 | Synthetic Fluency: Hallucinations, Confabulations, and the Creation of IrishWords in LLM-Generated TranslationsabstractThis study examines hallucinations in Large Language Model (LLM) translations into Irish, specifically focusing on instances where the models generate novel, non-existent words. We classify these hallucinations within verb and noun categories, identifying six distinct patterns among the latter. Additionally, we analyse whether these hallucinations adhere to Irish morphological rules and what linguistic tendencies they exhibit. Our findings show that while both GPT-4.o and GPT-4.o Mini produce similar types of hallucinations, the Mini model generates them at a significantly higher frequency. Beyond classification, the discussion raises speculative questions about the implications of these hallucinations for the Irish language. Rather than seeking definitive answers, we offer food for thought regarding the increasing use of LLMs and their potential role in shaping Irish vocabulary and linguistic evolution. We aim to prompt discussion on how such technologies might influence language over time, particularly in the context of low-resource, morphologically rich languages. Sheila Castilho, Zoe Fitzsimmons, Claire Holton, Aoife Mc Donagh |
MTSummit (1) | 1 |
| 2025 | Context-Aware Monolingual Evaluation of Machine TranslationabstractThis paper explores the potential of context-aware monolingual evaluation for assessing machine translation (MT) when no source is given for reference. To this end, we compare monolingual with bilingual evaluations (with source text), under two scenarios: the evaluation of a single MT system, and the comparative evaluation of pairwise MT systems. Four professional translators performed both monolingual and bilingual evaluations by assigning ratings and annotating errors, and providing feedback on their experience. Our findings suggest that context-aware monolingual evaluation achieves comparable outcomes to bilingual evaluations, and highlight the feasibility and potential of monolingual evaluation as an efficient approach to assessing MT. Silvio Picinini, Sheila Castilho |
MTSummit (1) | 2 |
| 2024 | Perceptions of Educators on MTQA Curriculum and InstructionabstractThis paper reports the preliminary resultsof a survey aimed at identifying and ex-ploring the attitudes and recommendationsof machine translation quality assessment(MTQA) educators. Drawing upon ele-ments from the literature on MTQA teach-ing, the survey explores themes that maypose a challenge or lead to successful im-plementation of human evaluation, as theliterature shows that there has not beenenough design and reporting. Results show educators’ awareness ofthe topic, awareness stemming from therecommendations of the literature on MTevaluation, and reports new challenges andissues. João Camargo, Sheila Castilho, Joss Moorkens |
EAMT (1) | 2 |
| 2023 | Do online Machine Translation Systems Care for Context? What About a GPT Model?abstractThis paper addresses the challenges of evaluating document-level machine translation (MT) in the context of recent advances in context-aware neural machine translation (NMT). It investigates how well online MT systems deal with six context-related issues, namely lexical ambiguity, grammatical gender, grammatical number, reference, ellipsis, and terminology, when a larger context span containing the solution for those issues is given as input. Results are compared to the translation outputs from the online ChatGPT. Our results show that, while the change of punctuation in the input yields great variability in the output translations, the context position does not seem to have a great impact. Moreover, the GPT model seems to outperform the NMT systems but performs poorly for Irish. The study aims to provide insights into the effectiveness of online MT systems in handling context and highlight the importance of considering contextual factors in evaluating MT systems. Sheila Castilho, Clodagh Quinn Mallon, Rahel Meister, Shengya Yue |
EAMT | 1 |
| 2022 | Achievements of the PRINCIPLE Project: Promoting MT for Croatian, Icelandic, Irish and NorwegianabstractThis paper provides an overview of the main achievements of the completed PRINCIPLE project, a 2-year action funded by the European Commission under the Connecting Europe Facility (CEF) programme. PRINCIPLE focused on collecting high-quality language resources for Croatian, Icelandic, Irish and Norwegian, which are severely low-resource languages, especially for building effective machine translation (MT) systems. We report the achievements of the project, primarily, in terms of the large amounts of data collected for all four low-resource languages and of promoting the uptake of neural MT (NMT) for these languages. Petra Bago, Sheila Castilho, Jane Dunne, Federico Gaspari, Andre Kåsen, Gauti Kristmannsson, Jon Arild Olsen, Natália Resende, Níels Rúnar Gíslason, Dana Davis Sheridan, Paraic Sheridan, John Tinsley, Andy Way |
EAMT | 2 |
| 2022 | DELA Project: Document-level Machine Translation EvaluationabstractThis paper presents the results of the DELA Project. We describe the testing of context span for document-level evaluation, construction of a document-level corpus, and context position, as well as the latest developments of the project when looking at human and automatic evaluation metrics for document-level evaluation. Sheila Castilho |
EAMT | 1 |
| 2022 | MT-Pese: Machine Translation and Post-EditeseabstractThis paper introduces the MT-Pese project, which aims at researching the post-editese phenomena in machine translated texts. We describe a range of experiments performed in order to gauge the effect of post-editese in dif-ferent domains, backtranslation, and quality. Sheila Castilho, Natália Resende |
EAMT | 1 |
| 2022 | How Much Context Span is Enough? Examining Context-Related Issues for Document-level MTabstractThis paper analyses how much context span is necessary to solve different context-related issues, namely, reference, ellipsis, gender, number, lexical ambiguity, and terminology when translating from English into Portuguese. We use the DELA corpus, which consists of 60 documents and six different domains (subtitles, literary, news, reviews, medical, and legislation). We find that the shortest context span to disambiguate issues can appear in different positions in the document including preceding, following, global, world knowledge. Moreover, the average length depends on the issue types as well as the domain. Moreover, we show that the standard approach of relying on only two preceding sentences as context might not be enough depending on the domain and issue types. Sheila Castilho |
LREC | 1 |
| 2020 | Document-Level Machine Translation Evaluation Project: Methodology, Effort and Inter-Annotator AgreementabstractDocument-level (doc-level) human eval-uation of machine translation (MT) has raised interest in the community after a fewattempts have disproved claims of “human parity” (Toral et al., 2018; Laubli et al.,2018). However, little is known about bestpractices regarding doc-level human evalu-ation. The goal of this project is to identifywhich methodologies better cope with i)the current state-of-the-art (SOTA) humanmetrics, ii) a possible complexity when as-signing a single score to a text consisted of‘good’ and ‘bad’ sentences, iii) a possibletiredness bias in doc-level set-ups, and iv)the difference in inter-annotator agreement(IAA) between sentence and doc-level set-ups. Sheila Castilho |
EAMT | 1 |
| 2020 | A human evaluation of English-Irish statistical and neural machine translationabstractWith official status in both Ireland and the EU, there is a need for high-quality English-Irish (EN-GA) machine translation (MT) systems which are suitable for use in a professional translation environment. While we have seen recent research on improving both statistical MT and neural MT for the EN-GA pair, the results of such systems have always been reported using automatic evaluation metrics. This paper provides the first human evaluation study of EN-GA MT using professional translators and in-domain (public administration) data for a more accurate depiction of the translation quality available via MT. Meghan Dowling, Sheila Castilho, Joss Moorkens, Teresa Lynn, Andy Way |
EAMT | 2 |
| 2020 | On Context Span Needed for Machine Translation EvaluationabstractDespite increasing efforts to improve evaluation of machine translation (MT) by going beyond the sentence level to the document level, the definition of what exactly constitutes a “document level” is still not clear. This work deals with the context span necessary for a more reliable MT evaluation. We report results from a series of surveys involving three domains and 18 target languages designed to identify the necessary context span as well as issues related to it. Our findings indicate that, despite the fact that some issues and spans are strongly dependent on domain and on the target language, a number of common patterns can be observed so that general guidelines for context-aware MT evaluation can be drawn. Sheila Castilho, Maja Popovic, Andy Way |
LREC | 1 |
| 2020 | A Set of Recommendations for Assessing Human-Machine Parity in Language TranslationabstractThe quality of machine translation has increased remarkably over the past years, to the degree that it was found to be indistinguishable from professional human translation in a number of empirical investigations. We reassess Hassan et al.'s 2018 investigation into Chinese to English news translation, showing that the finding of human–machine parity was owed to weaknesses in the evaluation design—which is currently considered best practice in the field. We show that the professional human translations contained significantly fewer errors, and that perceived quality in human evaluation depends on the choice of raters, the availability of linguistic context, and the creation of reference translations. Our results call for revisiting current best practices to assess strong machine translation systems in general and human–machine parity in particular, for which we offer a set of recommendations based on our empirical findings. Samuel Läubli, Sheila Castilho, Graham Neubig, Rico Sennrich, Qinlan Shen, Antonio Toral |
J. Artif. Intell. Res. | 2 |
| 2019 | Large-scale Machine Translation Evaluation of the iADAATPA Project
Sheila Castilho, Natália Resende, Federico Gaspari, Andy Way, Tony O'Dowd, Marek Mazur, Manuel Herranz, Alexandre Helle, Gema Ramírez-Sánchez, Víctor M. Sánchez-Cartagena, Marcis Pinnis, Valters Sics |
MTSummit (2) | 1 |
| 2019 | Chan Sin-Wai (ed): The human factor in machine translation - Routledge Studies in Translation Technology, Routledge, 1st edition (2018), xii+256pp, ISBN 781-13-855-1213 (Hardback), 978-13-151-4753-6 (e-book)
Sheila Castilho |
Mach. Transl. | 1 |
| 2019 | Editors' foreword to the special issue on human factors in neural machine translation
Sheila Castilho, Federico Gaspari, Joss Moorkens, Maja Popovic, Antonio Toral |
Mach. Transl. | 1 |
| 2018 | Reading Comprehension of Machine Translation Output: What Makes for a Better Read?abstractThis paper reports on a pilot experiment that compares two different machine translation (MT) paradigms in reading comprehension tests. To explore a suitable methodology, we set up a pilot experiment with a group of six users (with English, Spanish and Simplified Chinese languages) using an English Language Testing System (IELTS), and an eye-tracker. The users were asked to read three texts in their native language: either the original English text (for the English speakers) or the machine-translated text (for the Spanish and Simplified Chinese speakers). The original texts were machine-translated via two MT systems: neural (NMT) and statistical (SMT). The users were also asked to rank satisfaction statements on a 3-point scale after reading each text and answering the respective comprehension questions. After all tasks were completed, a post-task retrospective interview took place to gather qualitative data. The findings suggest that the users from the target languages completed more tasks in less time with a higher level of satisfaction when using translations from the NMT system. Sheila Castilho, Ana Guerberof Arenas |
EAMT | 1 |
| 2018 | Project PiPeNovel: Pilot on Post-editing NovelsabstractGiven (i) the rise of a new paradigm to machine translation based on neural networks that results in more fluent and less literal output than previous models and (ii) the maturity of machine-assisted translation via post-editing in industry, project PiPeNovel studies the feasibility of the post-editing workflow for literary text conducting experiments with professional literary translators. Antonio Toral, Martijn Wieling 0001, Sheila Castilho, Joss Moorkens, Andy Way |
EAMT | 3 |
| 2018 | Improving Machine Translation of Educational Content via Crowdsourcing
Maximiliana Behnke, Antonio Valerio Miceli Barone, Rico Sennrich, Vilelmini Sosoni, Thanasis Naskos, Eirini Takoulidou, Maria Stasimioti, Menno van Zaanen, Sheila Castilho, Federico Gaspari, Panayota Georgakopoulou, Valia Kordoni, Markus Egg, Katia Kermanidis |
LREC | 9 |
| 2018 | Translation Crowdsourcing: Creating a Multilingual Corpus of Online Educational Content
Vilelmini Sosoni, Katia Kermanidis, Maria Stasimioti, Thanasis Naskos, Eirini Takoulidou, Menno van Zaanen, Sheila Castilho, Panayota Georgakopoulou, Valia Kordoni, Markus Egg |
LREC | 7 |
| 2018 | Evaluating MT for massive open online courses - A multifaceted comparison between PBSMT and NMT systems
Sheila Castilho, Joss Moorkens, Federico Gaspari, Rico Sennrich, Andy Way, Panayota Georgakopoulou |
Mach. Transl. | 1 |
| 2017 | A Comparative Quality Evaluation of PBSMT and NMT using Professional Translators
Sheila Castilho, Joss Moorkens, Federico Gaspari, Rico Sennrich, Vilelmini Sosoni, Panayota Georgakopoulou, Pintu Lohar, Andy Way, Antonio Valerio Miceli Barone, Maria Gialama |
MTSummit (1) | 1 |
| 2017 | Translation Dictation vs. Post-editing with Cloud-based Voice Recognition: A Pilot Experiment
Julián Zapata, Sheila Castilho, Joss Moorkens |
MTSummit (2) | 2 |
| 2016 | Evaluating the Impact of Light Post-Editing on Usability
Sheila Castilho, Sharon O'Brien |
LREC | 1 |
| 2014 | Does post-editing increase usability? A study with Brazilian Portuguese as target language
Sheila Castilho, Sharon O'Brien, Fábio Alves, Morgan O'Brien |
EAMT | 1 |
| 2012 | PET: a Tool for Post-editing and Assessing Machine Translation
Wilker Aziz, Sheila Castilho, Lucia Specia |
LREC | 2 |