Ana Guerberof Arenas

dblp:152/8574 · also Ana Guerberof · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-9820-7074ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations
abstract
This article investigates the performance of automatic evaluation metrics (AEMs) and LLM-as-a-judge evaluation on literary translation across multiple languages, genres, and translation modalities. The aim is to assess how well these tools align with professionals when evaluating translation, creativity (creative shifts & errors), and see if they can substitute laborious manual annotations. A dataset of literary translations across three modalities (human translation, machine translation, and post-editing), three genres and three language pairs was created and annotated in detail for creativity by experienced professional literary translators. The results show that both AEMs and LLM-as-a-judge evaluations correlate poorly with professional evaluations on creativity, with LLM-as-a-judge showing a systematic bias in favour of machine-translated texts and penalising creative and culturally appropriate solutions. Moreover, performance is consistently worse for more literary genres such as poetry. This highlights fundamental limitations of current automatic evaluation tools for literary translation and the need to create new tools that do not frequently consider out of routine translations as errors.
Kyo Gerrits, Rik van Noord, Ana Guerberof Arenas
EAMT (1)3
2025 Optimising ChatGPT for creativity in literary translation: A case study from English into Dutch, Chinese, Catalan and Spanish
abstract
This study examines the variability of ChatGPT’s machine translation (MT) outputs across six different configurations in four languages, with a focus on creativity in a literary text. We evaluate GPT translations in different text granularity levels, temperature settings and prompting strategies with a Creativity Score formula. We found that prompting ChatGPT with a minimal instruction yields the best creative translations, with Translate the following text into [TG] creatively at the temperature of 1.0 outperforming other configurations and DeepL in Spanish, Dutch, and Chinese. Nonetheless, ChatGPT consistently underperforms compared to human translation (HT). All the code and data are available at Repository URL will be provided with camera-ready version.
Shuxiang Du, Ana Guerberof Arenas, Antonio Toral, Kyo Gerrits, Josep Marco Borillo
MTSummit (1)2
2025 To MT or not to MT: An eye-tracking study on the reception by Dutch readers of different translation and creativity levels
abstract
This article presents the results of a pilot study involving the reception of a fictional short story translated from English into Dutch under four conditions: machine translation (MT), post-editing (PE), human translation (HT) and original source text (ST). The aim is to understand how creativity and errors in different translation modalities affect readers, specifically regarding cognitive load. Eight participants filled in a questionnaire, read a story using an eye-tracker, and conducted a retrospective think-aloud (RTA) interview. The results show that units of creative potential (UCP) increase cognitive load and that this is the highest in HT and the lowest in MT; no effect of error was observed. Triangulating the data with RTAs leads us to hypothesize that the higher cognitive load in UCPs is linked to increases in reader enjoyment and immersion. The effect of translation creativity on cognitive load in different translation modalities at word-level is novel and opens up new avenues for further research.
Kyo Gerrits, Ana Guerberof Arenas
MTSummit (1)2
2025 QE4PE: Word-level Quality Estimation for Human Post-Editing
Gabriele Sarti, Vilém Zouhar, Grzegorz Chrupala, Ana Guerberof Arenas, Malvina Nissim, Arianna Bisazza
Trans. Assoc. Comput. Linguistics4
2024 INCREC: Uncovering the creative process of translated content using machine translation
abstract
The INCREC project aims to uncover professional translators’ creative stages to understand how technology can be best applied to the translation of literary and audio-visual texts, and to analyse the impact of these processes on readers and viewers. To better understand this process, INCREC triangulates data from eye-tracking, retrospective think-aloud inter-views, translated material, and questionnaires from professional translators and users.
Ana Guerberof Arenas
EAMT (2)1
2024 Literacy in Digital Environments and Resources (LT-LiDER)
abstract
LT-LiDER is an Erasmus+ cooperation project with two main aims. The first is to map the landscape of technological capabilities required to work as a language and/or translation expert in the digitalised and datafied language industry. The second is to generate training outputs that will help language and translation trainers improve their skills and adopt appropriate pedagogical approaches and strategies for integrating data-driven technology into their language or translation classrooms, with a focus on digital and AI literacy.
Joss Moorkens, Pilar Sánchez-Gijón, Esther Simon, Mireia Urpí, Nora Aranberri, Dragos Ciobanu, Ana Guerberof Arenas, Janiça Hackenbuchner, Dorothy Kenny, Ralph Krüger, Miguel Ángel Ríos-Gaona, Isabel Ginel, Caroline Rossi, Alina Secara, Antonio Toral
EAMT (2)7
2024 What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study
abstract
Gender bias in machine translation (MT) is recognized as an issue that can harm people and society.And yet, advancements in the field rarely involve people, the final MT users, or inform how they might be impacted by biased technologies.Current evaluations are often restricted to automatic methods, which offer an opaque estimate of what the downstream impact of gender disparities might be.We conduct an extensive human-centered study to examine if and to what extent bias in MT brings harms with tangible costs, such as quality of service gaps across women and men.To this aim, we collect behavioral data from ∼90 participants, who post-edited MT outputs to ensure correct gender translation.Across multiple datasets, languages, and types of users, our study shows that feminine post-editing demands significantly more technical and temporal effort, also corresponding to higher financial costs.Existing bias measurements, however, fail to reflect the found disparities.Our findings advocate for human-centered approaches that can inform the societal impact of bias.
Beatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof Arenas, Luisa Bentivogli
EMNLP4
2023 Migrant communities living in the Netherlands and their use of MT in healthcare settings
abstract
As part of a larger project on the use of MT in healthcare settings among migrant communities, this paper investigates if, when, how and with what (potential) challenges migrants use MT based on a survey of 201 non-native speakers of Dutch currently living in the Netherlands. Three main findings stand out from our analysis. First, most migrants use MT to understand health information in Dutch and communicate with health professionals. How MT is used and received varies depending on the context and the L2 language level, as well as age, but not on the educational level. Second, some users face challenges of different kinds, including a lack of trust or perceived inaccuracies. Some of these challenges are related to comprehension, which brings us to our third point. We argue that a more nuanced understanding of medical translation is needed in expert-to-non-expert health communication. This questionnaire helped us identify several topics we hope to explore in the project’s next phase.
Susana Valdez, Ana Guerberof Arenas, Kars Ligtenberg
EAMT2
2022 CREAMT: Creativity and narrative engagement of literary texts translated by translators and NMT
abstract
We present here the EU-funded project CREAMT that seeks to understand what is meant by creativity in different translation modalities, e.g. machine translation, post-editing or professional translation. Focusing on the textual elements that determine creativity in translated literary texts and the reader experience, CREAMT uses a novel, interdisciplinary approach to assess how effective MT is in literary translation considering creativity in translation and the ultimate user: the reader.
Ana Guerberof Arenas, Antonio Toral
EAMT1
2022 DivEMT: Neural Machine Translation Post-Editing Effort Across Typologically Diverse Languages
abstract
We introduce DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.Using a strictly controlled setup, 18 professional translators were instructed to translate or post-edit the same set of English documents into Arabic, Dutch, Italian, Turkish, Ukrainian, and Vietnamese.During the process, their edits, keystrokes, editing times and pauses were recorded, enabling an in-depth, cross-lingual evaluation of NMT quality and post-editing effectiveness.Using this new dataset, we assess the impact of two state-of-the-art NMT systems, Google Translate and the multilingual mBART-50 model, on translation productivity.We find that post-editing is consistently faster than translation from scratch.However, the magnitude of productivity gains varies widely across systems and languages, highlighting major disparities in post-editing effectiveness for languages at different degrees of typological relatedness to English, even when controlling for system architecture and training data size.We publicly release the complete dataset 1 including all collected behavioral data, to foster new research on the translation capabilities of NMT systems for typologically diverse languages.
Gabriele Sarti, Arianna Bisazza, Ana Guerberof Arenas, Antonio Toral
EMNLP3
2021 The impact of translation modality on user experience: an eye-tracking study of the Microsoft Word user interface
abstract
This paper presents results of the effect of different translation modalities on users when working with the Microsoft Word user interface. An experimental study was set up with 84 Japanese, German, Spanish, and English native speakers working with Microsoft Word in three modalities: the published translated version, a machine translated (MT) version (with unedited MT strings incorporated into the MS Word interface) and the published English version. An eye-tracker measured the cognitive load and usability according to the ISO/TR 16982 guidelines: i.e., effectiveness, efficiency, and satisfaction followed by retrospective think-aloud protocol. The results show that the users' effectiveness (number of tasks completed) does not significantly differ due to the translation modality. However, their efficiency (time for task completion) and self-reported satisfaction are significantly higher when working with the released product as opposed to the unedited MT version, especially when participants are less experienced. The eye-tracking results show that users experience a higher cognitive load when working with MT and with the human-translated versions as opposed to the English original. The results suggest that language and translation modality play a significant role in the usability of software products whether users complete the given tasks or not and even if they are unaware that MT was used to translate the interface.
Ana Guerberof Arenas, Joss Moorkens, Sharon O'Brien
Mach. Transl.1
2019 What is the impact of raw MT on Japanese users of Word: preliminary results of a usability study using eye-tracking
Ana Guerberof Arenas, Joss Moorkens, Sharon O'Brien
MTSummit (1)1
2018 Reading Comprehension of Machine Translation Output: What Makes for a Better Read?
abstract
This paper reports on a pilot experiment that compares two different machine translation (MT) paradigms in reading comprehension tests. To explore a suitable methodology, we set up a pilot experiment with a group of six users (with English, Spanish and Simplified Chinese languages) using an English Language Testing System (IELTS), and an eye-tracker. The users were asked to read three texts in their native language: either the original English text (for the English speakers) or the machine-translated text (for the Spanish and Simplified Chinese speakers). The original texts were machine-translated via two MT systems: neural (NMT) and statistical (SMT). The users were also asked to rank satisfaction statements on a 3-point scale after reading each text and answering the respective comprehension questions. After all tasks were completed, a post-task retrospective interview took place to gather qualitative data. The findings suggest that the users from the target languages completed more tasks in less time with a higher level of satisfaction when using translations from the NMT system.
Sheila Castilho, Ana Guerberof Arenas
EAMT2
2014 Miguel A. Jiménez-Crespo: Translation and Web Localization, Routledge, 2013, xii + 233 pp, ISBN: 978-0 41564316-0
Ana Guerberof Arenas
Mach. Transl.1
2014 Correlations between productivity and quality when post-editing in a professional context
Ana Guerberof Arenas
Mach. Transl.1