VLDB 2026 Research / reviewers in the wild / expert
Maja Popovic
dblp:31/752
· DBLP profile ↗
43ranked-venue papers
23as first author
12since 2021 · last 2025
0000-0001-8234-8745ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 23 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Interdisciplinary Approach to Human-Centered Machine TranslationabstractMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao 0001, Alvin Grissom II, Marzena Karpinska, Elaine C. Khoong, William D. Lewis, André F. T. Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon |
EMNLP | 18 |
| 2023 | Computational analysis of different translations: by professionals, students and machinesabstractIn this work, we analyse different translated texts in terms of various text features. We compare two types of human translations, professional and students’, and machine translation outputs in terms of lexical and grammatical variety, sentence length,as well as frequencies of different POS tags and POS-trigrams. Our experimentsare carried out on parallel translations into three languages, Croatian, Finnish andRussian, all originating from the same source English texts. Our results indicatethat machine translations are closest to the source text, followed by student translations. Also, student translations are similar both to professional as well as to MT, sometimes even more to MT. Furthermore, we identify sets of features which are convenient for distinguishing machine from human translations. Maja Popovic, Ekaterina Lapshinova-Koltunski, Maarit Koponen |
EAMT | 1 |
| 2023 | Using MT for multilingual covid-19 case load prediction from social media textsabstractIn the context of an epidemiological study involving multilingual social media, this paper reports on the ability of machine translation systems to preserve content relevant for a document classification task designed to determine whether the social media text is related to covid. The results indicate that machine translation does provide a feasible basis for scaling epidemiological social media surveillance to multiple languages. Moreover, a qualitative error analysis revealed that the majority of classification errors are not caused by MT errors. Maja Popovic, Vasudevan Nedumpozhimana, Meegan Gower, Sneha Rautmare, Nishtha Jain, John D. Kelleher |
EAMT | 1 |
| 2023 | Leveraging machine translation for cross-lingual fine-grained cyberbullying classification amongst pre-adolescentsabstractAbstract Cyberbullying is the wilful and repeated infliction of harm on an individual using the Internet and digital technologies. Similar to face-to-face bullying, cyberbullying can be captured formally using the Routine Activities Model (RAM) whereby the potential victim and bully are brought into proximity of one another via the interaction on online social networking (OSN) platforms. Although the impact of the COVID-19 (SARS-CoV-2) restrictions on the online presence of minors has yet to be fully grasped, studies have reported that 44% of pre-adolescents have encountered more cyberbullying incidents during the COVID-19 lockdown. Transparency reports shared by OSN companies indicate an increased take-downs of cyberbullying-related comments, posts or content by artificially intelligen moderation tools. However, in order to efficiently and effectively detect or identify whether a social media post or comment qualifies as cyberbullying, there are a number factors based on the RAM, which must be taken into account, which includes the identification of cyberbullying roles and forms. This demands the acquisition of large amounts of fine-grained annotated data which is costly and ethically challenging to produce. In addition where fine-grained datasets do exist they may be unavailable in the target language. Manual translation is costly and expensive, however, state-of-the-art neural machine translation offers a workaround. This study presents a first of its kind experiment in leveraging machine translation to automatically translate a unique pre-adolescent cyberbullying gold standard dataset in Italian with fine-grained annotations into English for training and testing a native binary classifier for pre-adolescent cyberbullying. In addition to contributing high-quality English reference translation of the source gold standard, our experiments indicate that the performance of our target binary classifier when trained on machine-translated English output is on par with the source (Italian) classifier. Kanishk Verma, Maja Popovic, Alexandros Poulis, Yelena Cherkasova, Cathal Ó Hóbáin, Angela Mazzone, Tijana Milosevic, Brian Davis 0001 |
Nat. Lang. Eng. | 2 |
| 2022 | Quantified Reproducibility Assessment of NLP ResultsabstractThis paper describes and tests a method for carrying out quantified reproducibility assessment (QRA) that is based on concepts and definitions from metrology.QRA produces a single score estimating the degree of reproducibility of a given system and evaluation measure, on the basis of the scores from, and differences between, different reproductions.We test QRA on 18 system and evaluation measure combinations (involving diverse NLP tasks and types of evaluation), for each of which we have the original results and one to seven reproduction results.The proposed QRA method produces degree-of-reproducibility scores that are comparable across multiple reproductions not only of the same, but of different original studies.We find that the proposed method facilitates insights into causes of variation between reproductions, and allows conclusions to be drawn about what changes to system and/or evaluation design might lead to improved reproducibility. Anya Belz, Maja Popovic, Simon Mille |
ACL (1) | 2 |
| 2022 | DiHuTra: a Parallel Corpus to Analyse Differences between Human TranslationsabstractThis project aimed to design a corpus of parallel human translations (HTs) of the same source texts by professionals and students. The resulting corpus consists of English news and reviews source texts, their translations into Russian and Croatian, and translations of the reviews into Finnish. The corpus will be valuable for both studying variation in translation and evaluating machine translation (MT) systems. Ekaterina Lapshinova-Koltunski, Maja Popovic, Maarit Koponen |
EAMT | 2 |
| 2022 | Leveraging Pre-trained Language Models for Gender DebiasingabstractStudying and mitigating gender and other biases in natural language have become important areas of research from both algorithmic and data perspectives. This paper explores the idea of reducing gender bias in a language generation context by generating gender variants of sentences. Previous work in this field has either been rule-based or required large amounts of gender balanced training data. These approaches are however not scalable across multiple languages, as creating data or rules for each language is costly and time-consuming. This work explores a light-weight method to generate gender variants for a given text using pre-trained language models as the resource, without any task-specific labelled data. The approach is designed to work on multiple languages with minimal changes in the form of heuristics. To showcase that, we have tested it on a high-resourced language, namely Spanish, and a low-resourced language from a different family, namely Serbian. The approach proved to work very well on Spanish, and while the results were less positive for Serbian, it showed potential even for languages where pre-trained models are less effective. Nishtha Jain, Declan Groves, Lucia Specia, Maja Popovic |
LREC | 4 |
| 2022 | DiHuTra: a Parallel Corpus to Analyse Differences between Human TranslationsabstractThis paper describes a new corpus of human translations which contains both professional and students translations. The data consists of English sources – texts from news and reviews – and their translations into Russian and Croatian, as well as of the subcorpus containing translations of the review texts into Finnish. All target languages represent mid-resourced and less or mid-investigated ones. The corpus will be valuable for studying variation in translation as it allows a direct comparison between human translations of the same source texts. The corpus will also be a valuable resource for evaluating machine translation systems. We believe that this resource will facilitate understanding and improvement of the quality issues in both human and machine translation. In the paper, we describe how the data was collected, provide information on translator groups and summarise the differences between the human translations at hand based on our preliminary results with shallow features. Ekaterina Lapshinova-Koltunski, Maja Popovic, Maarit Koponen |
LREC | 2 |
| 2021 | Agree to Disagree: Analysis of Inter-Annotator Disagreements in Human Evaluation of Machine Translation OutputabstractThis work describes an analysis of interannotator disagreements in human evaluation of machine translation output.The errors in the analysed texts were marked by multiple annotators under guidance of different quality criteria: adequacy, comprehension, and an unspecified generic mixture of adequacy and fluency.Our results show that different criteria result in different disagreements, and indicate that a clear definition of quality criterion can improve the inter-annotator agreement.Furthermore, our results show that for certain linguistic phenomena which are not limited to one or two words (such as word ambiguity or gender) but span over several words or even entire phrases (such as negation or relative clause), disagreements do not necessarily represent "errors" or "noise" but are rather inherent to the evaluation process.On the other hand, for some other phenomena (such as omission or verb forms) agreement can be easily improved by providing more precise and detailed instructions to the evaluators. Maja Popovic |
CoNLL | 1 |
| 2021 | A Reproduction Study of an Annotation-based Human Evaluation of MT OutputsabstractIn this paper we report our reproduction study of the Croatian part of an annotation-based human evaluation of machine-translated user reviews (Popović, 2020).The work was carried out as part of the ReproGen Shared Task on Reproducibility of Human Evaluation in NLG.Our aim was to repeat the original study exactly, except for using a different set of evaluators.We describe the experimental design, characterise differences between original and reproduction study, and present the results from each study, along with analysis of the similarity between them.For the six main evaluation results of Major/Minor/All Comprehension error rates and Major/Minor/All Adequacy error rates, we find that (i) 4/6 system rankings are the same in both studies, (ii) the relative differences between systems are replicated well for Major Comprehension and Adequacy (Pearson's > 0.9), but not for the corresponding Minor error rates (Pearson's 0.36 for Adequacy, 0.67 for Comprehension), and (iii) the individual system scores for both types of Minor error rates had a higher degree of reproducibility than the corresponding Major error rates.We also examine inter-annotator agreement and compare the annotations obtained in the original and reproduction studies. Maja Popovic, Anya Belz |
INLG | 1 |
| 2021 | On nature and causes of observed MT errorsabstractThis work describes analysis of nature and causes of MT errors observed by different evaluators under guidance of different quality criteria: adequacy and comprehension and and a not specified generic mixture of adequacy and fluency. We report results for three language pairs and two domains and eleven MT systems. Our findings indicate that and despite the fact that some of the identified phenomena depend on domain and/or language and the following set of phenomena can be considered as generally challenging for modern MT systems: rephrasing groups of words and translation of ambiguous source words and translating noun phrases and and mistranslations. Furthermore and we show that the quality criterion also has impact on error perception. Our findings indicate that comprehension and adequacy can be assessed simultaneously by different evaluators and so that comprehension and as an important quality criterion and can be included more often in human evaluations. Maja Popovic |
MTSummit (1) | 1 |
| 2021 | From MT to LREV: managing the transition
Maja Popovic |
Mach. Transl. | 1 |
| 2020 | Informative Manual Evaluation of Machine Translation OutputabstractThis work proposes a new method for manual evaluation of Machine Translation (MT) output based on marking actual issues in the translated text.The novelty is that the evaluators are not assigning any scores, nor classifying errors, but marking all problematic parts (words, phrases, sentences) of the translation.The main advantage of this method is that the resulting annotations do not only provide overall scores by counting words with assigned tags, but can be further used for analysis of errors and challenging linguistic phenomena, as well as inter-annotator disagreements.Detailed analysis and understanding of actual problems are not enabled by typical manual evaluations where the annotators are asked to assign overall scores or to rank two or more translations.The proposed method is very general: it can be applied on any genre/domain and language pair, and it can be guided by various types of quality criteria.Also, it is not restricted to MT output, but can be used for other types of generated text. Maja Popovic |
COLING | 1 |
| 2020 | Relations between comprehensibility and adequacy errors in machine translation outputabstractThis work presents a detailed analysis of translation errors perceived by readers as comprehensibility and/or adequacy issues.The main finding is that good comprehensibility, similarly to good fluency, can mask a number of adequacy errors.Of all major adequacy errors, 30% were fully comprehensible, thus fully misleading the reader to accept the incorrect information.Another 25% of major adequacy errors were perceived as almost comprehensible, thus being potentially misleading.Also, a vast majority of omissions (about 70%) is hidden by comprehensibility.Further analysis of misleading translations revealed that the most frequent error types are ambiguity, mistranslation, noun phrase error, word-by-word translation, untranslated word, subject-verb agreement, and spelling error in the source text.However, none of these error types appears exclusively in misleading translations, but are also frequent in fully incorrect (incomprehensible inadequate) and discarded correct (incomprehensible adequate) translations.Deeper analysis is needed to potentially detect underlying phenomena specifically related to misleading translations. Maja Popovic |
CoNLL | 1 |
| 2020 | On the differences between human translationsabstractMany studies have confirmed that translated texts exhibit different features than texts originally written in the given language. This work explores texts translated by different translators taking into account expertise and native language. A set of computational analyses was conducted on three language pairs, English-Croatian, German-French and English-Finnish, and the results show that each of the factors has certain influence on the features of the translated texts, especially on sentence length and lexical richness. The results also indicate that for translations used for machine translation evaluation, it is important to specify these factors, especially if comparing machine translation quality with human translation quality is involved. Maja Popovic |
EAMT | 1 |
| 2020 | QRev: Machine Translation of User Reviews: What Influences the Translation Quality?abstractThis project aims to identify the important aspects of translation quality of user reviews which will represent a starting point for developing better automatic MT metrics and challenge test sets, and will be also helpful for developing MT systems for this genre. We work on two types of reviews: Amazon products and IMDb movies, written in English and translated into two closely related target languages, Croatian and Serbian. Maja Popovic |
EAMT | 1 |
| 2020 | On Context Span Needed for Machine Translation EvaluationabstractDespite increasing efforts to improve evaluation of machine translation (MT) by going beyond the sentence level to the document level, the definition of what exactly constitutes a “document level” is still not clear. This work deals with the context span necessary for a more reliable MT evaluation. We report results from a series of surveys involving three domains and 18 target languages designed to identify the necessary context span as well as issues related to it. Our findings indicate that, despite the fact that some issues and spans are strongly dependent on domain and on the target language, a number of common patterns can be observed so that general guidelines for context-aware MT evaluation can be drawn. Sheila Castilho, Maja Popovic, Andy Way |
LREC | 2 |
| 2019 | On reducing translation shifts in translations intended for MT evaluation
Maja Popovic |
MTSummit (2) | 1 |
| 2019 | Automatic error classification with multiple error labels
Maja Popovic, David Vilar |
MTSummit (1) | 1 |
| 2019 | Editors' foreword to the special issue on human factors in neural machine translation
Sheila Castilho, Federico Gaspari, Joss Moorkens, Maja Popovic, Antonio Toral |
Mach. Transl. | 4 |
| 2018 | Proceedings of the 21st Annual Conference of the European Association for Machine Translation
Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Miquel Esplà-Gomis, Maja Popovic, Celia Rico, Joachim Van den Bogaert, Mikel L. Forcada |
EAMT | 4 |
| 2018 | A Multilingual Wikified Data Set of Educational Material
Iris Hendrickx, Eirini Takoulidou, Thanasis Naskos, Katia Kermanidis, Vilelmini Sosoni, Hugo De Vos, Maria Stasimioti, Menno van Zaanen, Panayota Georgakopoulou, Valia Kordoni, Maja Popovic, Markus Egg, Antal van den Bosch |
LREC | 11 |
| 2018 | Language-related issues for NMT and PBMT for English-German and English-Serbian
Maja Popovic |
Mach. Transl. | 1 |
| 2016 | Potential and Limits of Using Post-edits as Reference Translations for MT Evaluation
Maja Popovic, Mihael Arcan, Arle Lommel |
EAMT | 1 |
| 2016 | Can Text Simplification Help Machine Translation?
Sanja Stajner, Maja Popovic |
EAMT | 2 |
| 2016 | Tools and Guidelines for Principled Machine Translation Development
Nora Aranberri, Eleftherios Avramidis, Aljoscha Burchardt, Ondrej Klejch, Martin Popel, Maja Popovic |
LREC | 6 |
| 2016 | PE2rr Corpus: Manual Error Annotation of Automatically Pre-annotated MT Post-edits
Maja Popovic, Mihael Arcan |
LREC | 1 |
| 2015 | Identifying main obstacles for statistical machine translation of morphologically rich South Slavic languages
Maja Popovic, Mihael Arcan |
EAMT | 1 |
| 2015 | Poor man's lemmatisation for automatic error classification
Maja Popovic, Mihael Arcan, Eleftherios Avramidis, Aljoscha Burchardt, Arle Lommel |
EAMT | 1 |
| 2014 | Using a new analytic measure for the annotation and analysis of MT errors on real data
Arle Lommel, Aljoscha Burchardt, Maja Popovic, Kim Harris, Eleftherios Avramidis, Hans Uszkoreit |
EAMT | 3 |
| 2014 | Relations between different types of post-editing operations, cognitive effort and temporal effort
Maja Popovic, Arle Lommel, Aljoscha Burchardt, Eleftherios Avramidis, Hans Uszkoreit |
EAMT | 1 |
| 2014 | The taraXÜ corpus of human-annotated machine translations
Eleftherios Avramidis, Aljoscha Burchardt, Sabine Hunsicker, Maja Popovic, Cindy Tscherwinka, David Vilar, Hans Uszkoreit |
LREC | 4 |
| 2012 | Involving Language Professionals in the Evaluation of Machine Translation
Eleftherios Avramidis, Aljoscha Burchardt, Christian Federmann, Maja Popovic, Cindy Tscherwinka, David Vilar |
LREC | 4 |
| 2012 | Automatic MT Error Analysis: Hjerson Helping Addicter
Jan Berka, Ondrej Bojar, Mark Fishel, Maja Popovic, Daniel Zeman |
LREC | 4 |
| 2012 | Terra: a Collection of Translation Error-Annotated Corpora
Mark Fishel, Ondrej Bojar, Maja Popovic |
LREC | 3 |
| 2012 | Study and correlation analysis of linguistic, perceptual, and automatic machine translation evaluationsabstractAbstract Evaluation of machine translation output is an important task. Various human evaluation techniques as well as automatic metrics have been proposed and investigated in the last decade. However, very few evaluation methods take the linguistic aspect into account. In this article, we use an objective evaluation method for machine translation output that classifies all translation errors into one of the five following linguistic levels: orthographic, morphological, lexical, semantic, and syntactic. Linguistic guidelines for the target language are required, and human evaluators use them in to classify the output errors. The experiments are performed on English‐to‐Catalan and Spanish‐to‐Catalan translation outputs generated by four different systems: 2 rule‐based and 2 statistical. All translations are evaluated using the 3 following methods: a standard human perceptual evaluation method, several widely used automatic metrics, and the human linguistic evaluation. Pearson and Spearman correlation coefficients between the linguistic, perceptual, and automatic results are then calculated, showing that the semantic level correlates significantly with both perceptual evaluation and automatic metrics. Mireia Farrús, Marta R. Costa-jussà, Maja Popovic |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2011 | From Human to Automatic Error Classification for Machine Translation Output
Maja Popovic, Aljoscha Burchardt |
EAMT | 1 |
| 2011 | Towards Automatic Error Analysis of Machine Translation OutputabstractEvaluation and error analysis of machine translation output are important but difficult tasks. In this article, we propose a framework for automatic error analysis and classification based on the identification of actual erroneous words using the algorithms for computation of Word Error Rate (WER) and Position-independent word Error Rate (PER), which is just a very first step towards development of automatic evaluation measures that provide more specific information of certain translation problems. The proposed approach enables the use of various types of linguistic knowledge in order to classify translation errors in many different ways. This work focuses on one possible set-up, namely, on five error categories: inflectional errors, errors due to wrong word order, missing words, extra words, and incorrect lexical choices. For each of the categories, we analyze the contribution of various POS classes. We compared the results of automatic error analysis with the results of human error analysis in order to investigate two possible applications: estimating the contribution of each error type in a given translation output in order to identify the main sources of errors for a given translation system, and comparing different translation outputs using the introduced error categories in order to obtain more information about advantages and disadvantages of different systems and possibilites for improvements, as well as about advantages and disadvantages of applied methods for improvements. We used Arabic–English Newswire and Broadcast News and Chinese–English Newswire outputs created in the framework of the GALE project, several Spanish and English European Parliament outputs generated during the TC-Star project, and three German–English outputs generated in the framework of the fourth Machine Translation Workshop. We show that our results correlate very well with the results of a human error analysis, and that all our metrics except the extra words reflect well the differences between different versions of the same translation system as well as the differences between different translation systems. Maja Popovic, Hermann Ney |
Comput. Linguistics | 1 |
| 2006 | POS-based Word Reorderings for Statistical Machine Translation
Maja Popovic, Hermann Ney |
LREC | 1 |
| 2005 | Exploiting phrasal lexica and additional morpho-syntactic language resources for statistical machine translation with scarce training data
Maja Popovic, Hermann Ney |
EAMT | 1 |
| 2004 | Improving Word Alignment Quality using Morpho-syntactic Information
Hermann Ney, Maja Popovic |
COLING | 2 |
| 2004 | Error Measures and Bayes Decision Rules Revisited with Applications to POS Tagging
Hermann Ney, Maja Popovic, David Suendermann-Oeft |
EMNLP | 2 |
| 2004 | Towards the Use of Word Stems and Suffixes for Statistical Machine Translation
Maja Popovic, Hermann Ney |
LREC | 1 |