Arda Tezcan

dblp:185/5586 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-8707-6176ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 5 first-author · 10 since 2021
YearPublicationVenuePosition
2026 Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution
abstract
Machine translation (MT) systems continue to produce gender-biased translations. In a time where self-expression is paramount, mistranslations based on default behaviour and stereotyping can lead to harm for users of these systems. To better understand how these systems translate gender in the absence of clear gender cues, we need benchmarking resources that reflect gender ambiguous scenarios in a natural way. To this end, we present GAND, a gender ambiguous natural data benchmarking resource for MT consisting of English source sentences, specifically designed to analyse the influence of contextual cues on gender in translation. We leverage GAND to conduct an interpretability analysis: we translate a subset of GAND into two grammatical gender languages and extend these with manually crafted contrastive translations. A following feature attribution analysis reveals source words in context that inform the gender translation of an ambiguous referent entity in the target translation. As a newly introduced resource, GAND is designed to be benchmarked across diverse target languages and evaluated with a wide range of MT systems.
Janiça Hackenbuchner, Jasper Degraeuwe, Arda Tezcan, Joke Daems
EAMT (1)3
2026 Multilingual Communication in the Asylum Context: Evaluating LLM-Based Machine Translation with Fuzzy Match Augmentation and Adaptive NMT across Resource Conditions under Low-Data Constraints
abstract
Effective communication in asylum reception settings requires reliable machine translation (MT) across many languages, including low-resource ones. Using data from the MaTIAS project, we compare retrieval-augmented LLM translation with adaptive Neural MT across 14 target languages with varying resource levels. Working with a very small translation memory of only 358 sentences, we evaluate fuzzy match (FM) augmentation as an in-context learning strategy for open-source and commercial LLMs and benchmark these against ModernMT with and without domain adaptation. In the LLM setting, FM-based example selection consistently outperforms random selection and zero-shot prompting, with the largest gains for low-resource languages. Adaptive NMT retains an overall advantage, although Gemini Pro approaches its performance and outperforms it on 6 of 14 languages, highlighting a trade-off between translation quality and data sovereignty in privacy-sensitive contexts. These findings show that FM augmentation remains effective under severe data constraints and emphasise the importance of language-specific evaluation in multilingual MT.
Thomas Moerman 0001, Arda Tezcan, Lieve Macken
EAMT (1)2
2026 MaTIAS - Machine Translation to Inform Asylum Seekers: final results
abstract
This paper reports on the final stages of the MaTIAS project. A functional prototype of the multilingual notification tool was deployed across seven Belgian reception centres, accompanied by training and technical support. Feedback was gathered through interviews and surveys. Two rounds of machine translation evaluation revealed considerable differences in quality across languages. The translation quality of Tigrinya in particular was deemed too low to be usable.
July Wilde, Anaïs Wouters, Arda Tezcan, Simon Van den Meersschaut, Katrijn Maryns, Lieve Macken
EAMT (2)3
2025 Machine Translation to Inform Asylum Seekers: Intermediate Findings from the MaTIAS Project
abstract
We present key interim findings from the ongoing MaTIAS project, which focuses on developing a multilingual notification system for asylum reception centres in Belgium. This system integrates machine translation (MT) to enable staff to provide practical information to residents in their native language, thus fostering more effective communication. Our discussion focuses on three key aspects: the development of the multilingual messaging platform, the types of messages the system is designed to handle, and the evaluation of potential MT systems for integration.
Lieve Macken, Ella Hest, Arda Tezcan, Michaël Lumingu, Katrijn Maryns, July Wilde
MTSummit (2)3
2024 Automatic detection of (potential) factors in the source text leading to gender bias in machine translation
abstract
This research project aims to develop a comprehensive methodology to help make machine translation (MT) systems more gender-inclusive for society. The goal is the creation of a detection system, a machine learning (ML) model trained on manual annotations, that can automatically analyse source data and detect and highlight words and phrases that influence the gender bias inflection in target translations.The main research outputs will be (1) a manually annotated dataset, (2) a taxonomy, and (3) a fine-tuned model.
Janiça Hackenbuchner, Arda Tezcan, Joke Daems
EAMT (2)2
2024 MaTIAS: Machine Translation to Inform Asylum Seekers
abstract
This project aims to develop a multilingual notification system for asylum reception centres in Belgium using machine translation. The system will allow staff to communicate practical messages to residents in their own language. Ethnographically inspired fieldwork is being conducted in reception centres to understand current communication practices and ensure that the technology meets user needs. The quality and suitability of machine translation will be evaluated for three MT systems supporting all target languages. Automatic and manual evaluation methods will be used to assess translation quality, and terms of use, privacy and data protection conditions will be analysed.
Lieve Macken, Ella Hest, Arda Tezcan, Michaël Lumingu, Katrijn Maryns, July Wilde
EAMT (2)3
2023 Adapting Machine Translation Education to the Neural Era: A Case Study of MT Quality Assessment
abstract
The use of automatic evaluation metrics to assess Machine Translation (MT) quality is well established in the translation industry. Whereas it is relatively easy to cover the word- and character-based metrics in an MT course, it is less obvious to integrate the newer neural metrics. In this paper we discuss how we introduced the topic of MT quality assessment in a course for translation students. We selected three English source texts, each having a different difficulty level and style, and let the students translate the texts into their L1 and reflect upon translation difficulty. Afterwards, the students were asked to assess MT quality for the same texts using different methods and to critically reflect upon obtained results. The students had access to the MATEO web interface, which contains word- and character-based metrics as well as neural metrics. The students used two different reference translations: their own translations and professional translations of the three texts. We not only synthesise the comments of the students, but also present the results of some cross-lingual analyses on nine different language pairs.
Lieve Macken, Bram Vanroy, Arda Tezcan
EAMT3
2023 MATEO: MAchine Translation Evaluation Online
abstract
We present MAchine Translation Evaluation Online (MATEO), a project that aims to facilitate machine translation (MT) evaluation by means of an easy-to-use interface that can evaluate given machine translations with a battery of automatic metrics. It caters to both experienced and novice users who are working with MT, such as MT system builders, teachers and students of (machine) translation, and researchers.
Bram Vanroy, Arda Tezcan, Lieve Macken
EAMT2
2022 Literary translation as a three-stage process: machine translation, post-editing and revision
abstract
This study focuses on English-Dutch literary translations that were created in a professional environment using an MT-enhanced workflow consisting of a three-stage process of automatic translation followed by post-editing and (mainly) monolingual revision. We compare the three successive versions of the target texts. We used different automatic metrics to measure the (dis)similarity between the consecutive versions and analyzed the linguistic characteristics of the three translation variants. Additionally, on a subset of 200 segments, we manually annotated all errors in the machine translation output and classified the different editing actions that were carried out. The results show that more editing occurred during revision than during post-editing and that the types of editing actions were different.
Lieve Macken, Bram Vanroy, Luca Desmet, Arda Tezcan
EAMT4
2022 Dynamic Adaptation of Neural Machine-Translation Systems Through Translation Exemplars
abstract
This project aims to study the impact of adapting neural machine translation (NMT) systems through translation exemplars, determine the optimal similarity metric(s) for retrieving informative exemplars, and, verify the usefulness of this approach for domain adaptation of NMT systems.
Arda Tezcan
EAMT1
2020 Assessing the Comprehensibility of Automatic Translations (ArisToCAT)
abstract
The ArisToCAT project aims to assess the comprehensibility of ‘raw’ (unedited) MT output for readers who can only rely on the MT output. In this project description, we summarize the main results of the project and present future work.
Lieve Macken, Margot Fonteyne, Arda Tezcan, Joke Daems
EAMT3
2020 Literary Machine Translation under the Magnifying Glass: Assessing the Quality of an NMT-Translated Detective Novel on Document Level
abstract
Several studies (covering many language pairs and translation tasks) have demonstrated that translation quality has improved enormously since the emergence of neural machine translation systems. This raises the question whether such systems are able to produce high-quality translations for more creative text types such as literature and whether they are able to generate coherent translations on document level. Our study aimed to investigate these two questions by carrying out a document-level evaluation of the raw NMT output of an entire novel. We translated Agatha Christie’s novel The Mysterious Affair at Styles with Google’s NMT system from English into Dutch and annotated it in two steps: first all fluency errors, then all accuracy errors. We report on the overall quality, determine the remaining issues, compare the most frequent error types to those in general-domain MT, and investigate whether any accuracy and fluency errors co-occur regularly. Additionally, we assess the inter-annotator agreement on the first chapter of the novel.
Margot Fonteyne, Arda Tezcan, Lieve Macken
LREC2
2020 Estimating word-level quality of statistical machine translation output using monolingual information alone
abstract
Abstract Various studies show that statistical machine translation (SMT) systems suffer from fluency errors, especially in the form of grammatical errors and errors related to idiomatic word choices. In this study, we investigate the effectiveness of using monolingual information contained in the machine-translated text to estimate word-level quality of SMT output. We propose a recurrent neural network architecture which uses morpho-syntactic features and word embeddings as word representations within surface and syntactic n-grams. We test the proposed method on two language pairs and for two tasks, namely detecting fluency errors and predicting overall post-editing effort. Our results show that this method is effective for capturing all types of fluency errors at once. Moreover, on the task of predicting post-editing effort, while solely relying on monolingual information, it achieves on-par results with the state-of-the-art quality estimation systems which use both bilingual and monolingual information.
Arda Tezcan, Véronique Hoste, Lieve Macken
Nat. Lang. Eng.1
2019 Neural Fuzzy Repair: Integrating Fuzzy Matches into Neural Machine Translation
abstract
We present a simple yet powerful data augmentation method for boosting Neural Machine Translation (NMT) performance by leveraging information retrieved from a Translation Memory (TM).We propose and test two methods for augmenting NMT training data with fuzzy TM matches.Tests on the DGT-TM data set for two language pairs show consistent and substantial improvements over a range of baseline systems.The results suggest that this method is promising for any translation environment in which a sizeable TM is available and a certain amount of repetition across translations is to be expected, especially considering its ease of implementation.
Bram Bulté, Arda Tezcan
ACL (1)2
2019 Estimating post-editing time using a gold-standard set of machine translation errors
Arda Tezcan, Véronique Hoste, Lieve Macken
Comput. Speech Lang.1
2018 Smart Computer-Aided Translation Environment (SCATE): Highlights
abstract
We present the highlights of the now finished 4-year SCATE project. It was completed in February 2018 and funded by the Flemish Government IWT-SBO, project No. 130041.1
Vincent Vandeghinste, Tom Vanallemeersch, Bram Bulté, Liesbeth Augustinus, Frank Van Eynde, Joris Pelemans, Lyan Verwimp, Patrick Wambacq, Geert Heyman, Marie-Francine Moens, Iulianna Van der Lek-Ciudin, Frieda Steurs, Ayla Rigouts Terryn, Els Lefever, Arda Tezcan, Lieve Macken, Sven Coppers, Jens Brulmans, Jan Van den Bergh 0001, Kris Luyten, Karin Coninx
EAMT15
2018 A fine-grained error analysis of NMT, SMT and RBMT output for English-to-Dutch
Laura Van Brussel, Arda Tezcan, Lieve Macken
LREC2
2016 Detecting Grammatical Errors in Machine Translation Output Using Dependency Parsing and Treebank Querying
Arda Tezcan, Véronique Hoste, Lieve Macken
EAMT1
2015 Smart Computer Aided Translation Environment - SCATE
Vincent Vandeghinste, Tom Vanallemeersch, Frank Van Eynde, Geert Heyman, Marie-Francine Moens, Joris Pelemans, Patrick Wambacq, Iulianna Van der Lek-Ciudin, Arda Tezcan, Lieve Macken, Véronique Hoste, Eva Geurts, Mieke Haesen
EAMT9
2011 SMT-CAT integration in a Technical Domain: Handling XML Markup Using Pre & Post-processing Methods
Arda Tezcan, Vincent Vandeghinste
EAMT1