EDBT 2026 Demo / reviewers in the wild / expert
Lieve Macken
dblp:16/3175
· DBLP profile ↗
34ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0001-7516-7487ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 10 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OSCAIL-OpenScience Communication through AI in EU LanguagesabstractThe Anglocentric nature of scholarly communication has many implications, such as limiting publication, discoverability and access from other language communities (even for major languages); putting minoritized languages at risk in the academic domain; and excluding many from peer review. The OSCAIL project addresses these challenges by exploring how machine translation (MT) enhanced by large language model (LLM)–based technologies can support access to scientific knowledge. Outputs will include evaluation datasets, protocols and best practices for MT in scholarly communication, and a prototype integration of MT tools into Open Journal Systems, the world’s most widely used open-source scholarly publishing platform. Sheila Castilho, Susanna Fiorini, Lynne Bowker, Petr Motlícek, Joss Moorkens, Lieve Macken, Dairazalia Sanchez-Cortes, Janne Pölönen, Sami Syrjämäki, Mikael Laakso, Mark Fishel, Anastasia Stasenko |
EAMT (2) | 6 |
| 2026 | Multilingual Communication in the Asylum Context: Evaluating LLM-Based Machine Translation with Fuzzy Match Augmentation and Adaptive NMT across Resource Conditions under Low-Data ConstraintsabstractEffective communication in asylum reception settings requires reliable machine translation (MT) across many languages, including low-resource ones. Using data from the MaTIAS project, we compare retrieval-augmented LLM translation with adaptive Neural MT across 14 target languages with varying resource levels. Working with a very small translation memory of only 358 sentences, we evaluate fuzzy match (FM) augmentation as an in-context learning strategy for open-source and commercial LLMs and benchmark these against ModernMT with and without domain adaptation. In the LLM setting, FM-based example selection consistently outperforms random selection and zero-shot prompting, with the largest gains for low-resource languages. Adaptive NMT retains an overall advantage, although Gemini Pro approaches its performance and outperforms it on 6 of 14 languages, highlighting a trade-off between translation quality and data sovereignty in privacy-sensitive contexts. These findings show that FM augmentation remains effective under severe data constraints and emphasise the importance of language-specific evaluation in multilingual MT. Thomas Moerman 0001, Arda Tezcan, Lieve Macken |
EAMT (1) | 3 |
| 2026 | MaTIAS - Machine Translation to Inform Asylum Seekers: final resultsabstractThis paper reports on the final stages of the MaTIAS project. A functional prototype of the multilingual notification tool was deployed across seven Belgian reception centres, accompanied by training and technical support. Feedback was gathered through interviews and surveys. Two rounds of machine translation evaluation revealed considerable differences in quality across languages. The translation quality of Tigrinya in particular was deemed too low to be usable. July Wilde, Anaïs Wouters, Arda Tezcan, Simon Van den Meersschaut, Katrijn Maryns, Lieve Macken |
EAMT (2) | 6 |
| 2025 | LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question AnsweringabstractThe impact of Large Language Models (LLMs) has extended into literary domains.However, existing evaluation metrics for literature prioritize mechanical accuracy over artistic expression and tend to overrate machine translation as being superior to human translation from experienced professionals.In the long run, this bias could result in an irreversible decline in translation quality and cultural authenticity.In response to the urgent need for a specialized literary evaluation metric, we introduce LITRANSPROQA, a novel, referencefree, LLM-based question-answering framework designed for literary translation evaluation.LITRANSPROQA integrates humans in the loop to incorporate insights from professional literary translators and researchers, focusing on critical elements in literary quality assessment such as literary devices, cultural understanding, and authorial voice.Our extensive evaluation shows that while literaryfinetuned XCOMET-XL yields marginal gains, LITRANSPROQA substantially outperforms current metrics, achieving up to 0.07 gain in correlation and surpassing the best state-of-theart metrics by over 15 points in adequacy assessments.Incorporating professional translator insights as weights further improves performance, highlighting the value of translator inputs.Notably, LITRANSPROQA reaches an adequacy performance comparable to trained linguistic student evaluators, though it still falls behind experienced professional translators.LITRANSPROQA shows broad applicability to open-source models like LLaMA3.3-70b and Qwen2.5-32b,indicating its potential as an accessible and training-free tool for evaluating literary translations that require local processing due to copyright or ethical considerations. Ran Zhang 0013, Lieve Macken, Steffen Eger |
EMNLP | 3 |
| 2025 | Decoding Machine Translationese in English-Chinese News: LLMs vs. NMTsabstractThis study explores Machine Translationese (MTese) — the linguistic peculiarities of machine translation outputs — focusing on the under-researched English-to-Chinese language pair in news texts. We construct a large dataset consisting of 4 sub-corpora and employ a comprehensive five-layer feature set. Then, a chi-square ranking algorithm is applied for feature selection in both classification and clustering tasks. Our findings confirm the presence of MTese in both Neural Machine Translation systems (NMTs) and Large Language Models (LLMs). Original Chinese texts are nearly perfectly distinguishable from both LLM and NMT outputs. Notable linguistic patterns in MT outputs are shorter sentence lengths and increased use of adversative conjunctions. Comparing LLMs and NMTs, we achieve approximately 70% classification accuracy, with LLMs exhibiting greater lexical diversity and NMTs using more brackets. Additionally, translation-specific LLMs show lower lexical diversity but higher usage of causal conjunctions compared to generic LLMs. Lastly, we find no significant differences between LLMs developed by Chinese firms and their foreign counterparts. Delu Kong, Lieve Macken |
MTSummit (1) | 2 |
| 2025 | Machine Translation to Inform Asylum Seekers: Intermediate Findings from the MaTIAS ProjectabstractWe present key interim findings from the ongoing MaTIAS project, which focuses on developing a multilingual notification system for asylum reception centres in Belgium. This system integrates machine translation (MT) to enable staff to provide practical information to residents in their native language, thus fostering more effective communication. Our discussion focuses on three key aspects: the development of the multilingual messaging platform, the types of messages the system is designed to handle, and the evaluation of potential MT systems for integration. Lieve Macken, Ella Hest, Arda Tezcan, Michaël Lumingu, Katrijn Maryns, July Wilde |
MTSummit (2) | 1 |
| 2024 | MaTIAS: Machine Translation to Inform Asylum SeekersabstractThis project aims to develop a multilingual notification system for asylum reception centres in Belgium using machine translation. The system will allow staff to communicate practical messages to residents in their own language. Ethnographically inspired fieldwork is being conducted in reception centres to understand current communication practices and ensure that the technology meets user needs. The quality and suitability of machine translation will be evaluated for three MT systems supporting all target languages. Automatic and manual evaluation methods will be used to assess translation quality, and terms of use, privacy and data protection conditions will be analysed. Lieve Macken, Ella Hest, Arda Tezcan, Michaël Lumingu, Katrijn Maryns, July Wilde |
EAMT (2) | 1 |
| 2023 | Adapting Machine Translation Education to the Neural Era: A Case Study of MT Quality AssessmentabstractThe use of automatic evaluation metrics to assess Machine Translation (MT) quality is well established in the translation industry. Whereas it is relatively easy to cover the word- and character-based metrics in an MT course, it is less obvious to integrate the newer neural metrics. In this paper we discuss how we introduced the topic of MT quality assessment in a course for translation students. We selected three English source texts, each having a different difficulty level and style, and let the students translate the texts into their L1 and reflect upon translation difficulty. Afterwards, the students were asked to assess MT quality for the same texts using different methods and to critically reflect upon obtained results. The students had access to the MATEO web interface, which contains word- and character-based metrics as well as neural metrics. The students used two different reference translations: their own translations and professional translations of the three texts. We not only synthesise the comments of the students, but also present the results of some cross-lingual analyses on nine different language pairs. Lieve Macken, Bram Vanroy, Arda Tezcan |
EAMT | 1 |
| 2023 | Developing User-centred Approaches to Technological Innovation in Literary Translation (DUAL-T)abstractDUAL-T is an EU-funded project which aims at involving literary translators in the testing of technology-inclusive workflows. Participants will be asked to translate three short stories using, respectively, (1) a text editor combined with online resources, (2) a Computer-Aided Translation (CAT) tool, and (3) a Machine Translation Post-editing (MTPE) tool. Paola Ruffo, Joke Daems, Lieve Macken |
EAMT | 3 |
| 2023 | MATEO: MAchine Translation Evaluation OnlineabstractWe present MAchine Translation Evaluation Online (MATEO), a project that aims to facilitate machine translation (MT) evaluation by means of an easy-to-use interface that can evaluate given machine translations with a battery of automatic metrics. It caters to both experienced and novice users who are working with MT, such as MT system builders, teachers and students of (machine) translation, and researchers. Bram Vanroy, Arda Tezcan, Lieve Macken |
EAMT | 3 |
| 2022 | Writing in a second Language with Machine translation (WiLMa)abstractThe WiLMa project aims to assess the effects of using machine translation (MT) tools on the writing processes of second language (L2) learners of varying proficiency. Particular attention is given to individual variation in learners’ tool use. Margot Fonteyne, Maribel Montero Perez, Joke Daems, Lieve Macken |
EAMT | 4 |
| 2022 | Literary translation as a three-stage process: machine translation, post-editing and revisionabstractThis study focuses on English-Dutch literary translations that were created in a professional environment using an MT-enhanced workflow consisting of a three-stage process of automatic translation followed by post-editing and (mainly) monolingual revision. We compare the three successive versions of the target texts. We used different automatic metrics to measure the (dis)similarity between the consecutive versions and analyzed the linguistic characteristics of the three translation variants. Additionally, on a subset of 200 segments, we manually annotated all errors in the machine translation output and classified the different editing actions that were carried out. The results show that more editing occurred during revision than during post-editing and that the types of editing actions were different. Lieve Macken, Bram Vanroy, Luca Desmet, Arda Tezcan |
EAMT | 1 |
| 2022 | GECO-MT: The Ghent Eye-tracking Corpus of Machine TranslationabstractIn the present paper, we describe a large corpus of eye movement data, collected during natural reading of a human translation and a machine translation of a full novel. This data set, called GECO-MT (Ghent Eye tracking Corpus of Machine Translation) expands upon an earlier corpus called GECO (Ghent Eye-tracking Corpus) by Cop et al. (2017). The eye movement data in GECO-MT will be used in future research to investigate the effect of machine translation on the reading process and the effects of various error types on reading. In this article, we describe in detail the materials and data collection procedure of GECO-MT. Extensive information on the language proficiency of our participants is given, as well as a comparison with the participants of the original GECO. We investigate the distribution of a selection of important eye movement variables and explore the possibilities for future analyses of the data. GECO-MT is freely available at https://www.lt3.ugent.be/resources/geco-mt. Toon Colman, Margot Fonteyne, Joke Daems, Nicolas Dirix, Lieve Macken |
LREC | 5 |
| 2022 | LeConTra: A Learner Corpus of English-to-Dutch News TranslationabstractWe present LeConTra, a learner corpus consisting of English-to-Dutch news translations enriched with translation process data. Three students of a Master’s programme in Translation were asked to translate 50 different English journalistic texts of approximately 250 tokens each. Because we also collected translation process data in the form of keystroke logging, our dataset can be used as part of different research strands such as translation process research, learner corpus research, and corpus-based translation studies. Reference translations, without process data, are also included. The data has been manually segmented and tokenized, and manually aligned at both segment and word level, leading to a high-quality corpus with token-level process data. The data is freely accessible via the Translation Process Research DataBase, which emphasises our commitment of distributing our dataset. The tool that was built for manual sentence segmentation and tokenization, Mantis, is also available as an open-source aid for data processing. Bram Vanroy, Lieve Macken |
LREC | 2 |
| 2020 | Assessing the Comprehensibility of Automatic Translations (ArisToCAT)abstractThe ArisToCAT project aims to assess the comprehensibility of ‘raw’ (unedited) MT output for readers who can only rely on the MT output. In this project description, we summarize the main results of the project and present future work. Lieve Macken, Margot Fonteyne, Arda Tezcan, Joke Daems |
EAMT | 1 |
| 2020 | Literary Machine Translation under the Magnifying Glass: Assessing the Quality of an NMT-Translated Detective Novel on Document LevelabstractSeveral studies (covering many language pairs and translation tasks) have demonstrated that translation quality has improved enormously since the emergence of neural machine translation systems. This raises the question whether such systems are able to produce high-quality translations for more creative text types such as literature and whether they are able to generate coherent translations on document level. Our study aimed to investigate these two questions by carrying out a document-level evaluation of the raw NMT output of an entire novel. We translated Agatha Christie’s novel The Mysterious Affair at Styles with Google’s NMT system from English into Dutch and annotated it in two steps: first all fluency errors, then all accuracy errors. We report on the overall quality, determine the remaining issues, compare the most frequent error types to those in general-domain MT, and investigate whether any accuracy and fluency errors co-occur regularly. Additionally, we assess the inter-annotator agreement on the first chapter of the novel. Margot Fonteyne, Arda Tezcan, Lieve Macken |
LREC | 3 |
| 2020 | Estimating word-level quality of statistical machine translation output using monolingual information aloneabstractAbstract Various studies show that statistical machine translation (SMT) systems suffer from fluency errors, especially in the form of grammatical errors and errors related to idiomatic word choices. In this study, we investigate the effectiveness of using monolingual information contained in the machine-translated text to estimate word-level quality of SMT output. We propose a recurrent neural network architecture which uses morpho-syntactic features and word embeddings as word representations within surface and syntactic n-grams. We test the proposed method on two language pairs and for two tasks, namely detecting fluency errors and predicting overall post-editing effort. Our results show that this method is effective for capturing all types of fluency errors at once. Moreover, on the task of predicting post-editing effort, while solely relying on monolingual information, it achieves on-par results with the state-of-the-art quality estimation systems which use both bilingual and monolingual information. Arda Tezcan, Véronique Hoste, Lieve Macken |
Nat. Lang. Eng. | 3 |
| 2019 | Estimating post-editing time using a gold-standard set of machine translation errors
Arda Tezcan, Véronique Hoste, Lieve Macken |
Comput. Speech Lang. | 3 |
| 2019 | Interactive adaptive SMT versus interactive adaptive NMT: a user experience evaluation
Joke Daems, Lieve Macken |
Mach. Transl. | 2 |
| 2018 | Smart Computer-Aided Translation Environment (SCATE): HighlightsabstractWe present the highlights of the now finished 4-year SCATE project. It was completed in February 2018 and funded by the Flemish Government IWT-SBO, project No. 130041.1 Vincent Vandeghinste, Tom Vanallemeersch, Bram Bulté, Liesbeth Augustinus, Frank Van Eynde, Joris Pelemans, Lyan Verwimp, Patrick Wambacq, Geert Heyman, Marie-Francine Moens, Iulianna Van der Lek-Ciudin, Frieda Steurs, Ayla Rigouts Terryn, Els Lefever, Arda Tezcan, Lieve Macken, Sven Coppers, Jens Brulmans, Jan Van den Bergh 0001, Kris Luyten, Karin Coninx |
EAMT | 16 |
| 2018 | A fine-grained error analysis of NMT, SMT and RBMT output for English-to-Dutch
Laura Van Brussel, Arda Tezcan, Lieve Macken |
LREC | 3 |
| 2016 | Detecting Grammatical Errors in Machine Translation Output Using Dependency Parsing and Treebank Querying
Arda Tezcan, Véronique Hoste, Lieve Macken |
EAMT | 3 |
| 2016 | Multimodular Text Normalization of Dutch User-Generated ContentabstractAs social media constitutes a valuable source for data analysis for a wide range of applications, the need for handling such data arises. However, the nonstandard language used on social media poses problems for natural language processing (NLP) tools, as these are typically trained on standard language material. We propose a text normalization approach to tackle this problem. More specifically, we investigate the usefulness of a multimodular approach to account for the diversity of normalization issues encountered in user-generated content (UGC). We consider three different types of UGC written in Dutch (SNS, SMS, and tweets) and provide a detailed analysis of the performance of the different modules and the overall system. We also apply an extrinsic evaluation by evaluating the performance of a part-of-speech tagger, lemmatizer, and named-entity recognizer before and after normalization. Sarah Schulz, Guy De Pauw, Orphée De Clercq, Bart Desmet, Véronique Hoste, Walter Daelemans, Lieve Macken |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2015 | Smart Computer Aided Translation Environment - SCATE
Vincent Vandeghinste, Tom Vanallemeersch, Frank Van Eynde, Geert Heyman, Marie-Francine Moens, Joris Pelemans, Patrick Wambacq, Iulianna Van der Lek-Ciudin, Arda Tezcan, Lieve Macken, Véronique Hoste, Eva Geurts, Mieke Haesen |
EAMT | 10 |
| 2014 | On the origin of errors: A fine-grained analysis of MT and PE errors and their relationship
Joke Daems, Lieve Macken, Sonia Vandepitte |
LREC | 2 |
| 2014 | Using the crowd for readability predictionabstractAbstract While human annotation is crucial for many natural language processing tasks, it is often very expensive and time-consuming. Inspired by previous work on crowdsourcing, we investigate the viability of using non-expert labels instead of gold standard annotations from experts for a machine learning approach to automatic readability prediction. In order to do so, we evaluate two different methodologies to assess the readability of a wide variety of text material: A more traditional setup in which expert readers make readability judgments and a crowdsourcing setup for users who are not necessarily experts. To this purpose two assessment tools were implemented: a tool where expert readers can rank a batch of texts based on readability, and a lightweight crowdsourcing tool, which invites users to provide pairwise comparisons. To validate this approach, readability assessments for a corpus of written Dutch generic texts were gathered. By collecting multiple assessments per text, we explicitly wanted to level out readers' background knowledge and attitude. Our findings show that the assessments collected through both methodologies are highly consistent and that crowdsourcing is a viable alternative to expert labeling. This is a good news as crowdsourcing is more lightweight to use and can have access to a much wider audience of potential annotators. By performing a set of basic machine learning experiments using a feature set that mainly encodes basic lexical and morpho-syntactic information, we further illustrate how the collected data can be used to perform text comparisons or to assign an absolute readability score to an individual text. We do not focus on optimising the algorithms to achieve the best possible results for the learning tasks, but carry them out to illustrate the various possibilities of our data sets. The results on different data sets, however, show that our system outperforms the readability formulas and a baseline language modelling approach. We conclude that readability assessment by comparing texts is a polyvalent methodology, which can be adapted to specific domains and target audiences if required. Orphée De Clercq, Véronique Hoste, Bart Desmet, Philip van Oosten, Martine De Cock, Lieve Macken |
Nat. Lang. Eng. | 6 |
| 2012 | From keystrokes to annotated process data: Enriching the output of Inputlog with linguistic information
Lieve Macken, Véronique Hoste, Mariëlle Leijten, Luuk van Waes |
LREC | 1 |
| 2010 | A Chunk-Driven Bootstrapping Approach to Extracting Translation Patterns
Lieve Macken, Walter Daelemans |
CICLing | 1 |
| 2010 | An Annotation Scheme and Gold Standard for Dutch-English Word Alignment
Lieve Macken |
LREC | 1 |
| 2009 | Language-Independent Bilingual Terminology Extraction from a Multilingual Parallel Corpus
Els Lefever, Lieve Macken, Véronique Hoste |
EACL | 2 |
| 2008 | Linguistically-Based Sub-Sentential Alignment for Terminology Extraction from a Bilingual Automotive Corpus
Lieve Macken, Els Lefever, Véronique Hoste |
COLING | 1 |
| 2008 | Sentence Alignment in DPC: Maximizing Precision, Minimizing Human Effort
Julia S. Trushkina, Lieve Macken, Hans Paulussen |
LREC | 2 |
| 2007 | Dutch parallel corpus: MT corpus and translator's aid
Lieve Macken, Julia S. Trushkina, Lidia Rura |
MTSummit | 1 |
| 2002 | Intonation modelling for the synthesis of structured documentsabstractThis paper describes experiments concerning the prediction of a good intonation for the synthesis of structured documents. The paper extends our previous research in four important aspects: (i) models are trained and evaluated on read text material (no isolated sentences), (ii) the intonation model is evaluated while fully integrated in the entire prosody model chain, (iii) the feature selection process is completely automated, and (iv) the importance of typical text-level features such as text type, text structure and typesetting are investigated. Clearly, human readings of running texts exhibit a much richer intonation than the intonation observed in read isolated sentences. We try to capture this richness in an intonation model that can be learned automatically using data-driven techniques. Our intonation models are RNNs (Recurrent Neural Networks) which are trained from prosodically labelled databases. Objective tests have demonstrated that acceptable intonation models can be constructed in this way, and that text type and text structure are important features whereas type-setting is not. 1. Jeska Buhmann, Jean-Pierre Martens, Lieve Macken, Bert Van Coile |
INTERSPEECH | 3 |