Maja Popovic

dblp:31/752 · DBLP profile ↗
← Back
43ranked-venue papers
23as first author
12since 2021 · last 2025
0000-0001-8234-8745ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 23 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 An Interdisciplinary Approach to Human-Centered Machine Translation
abstract
Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao 0001, Alvin Grissom II, Marzena Karpinska, Elaine C. Khoong, William D. Lewis, André F. T. Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon
EMNLP18
2023 Computational analysis of different translations: by professionals, students and machines
abstract
In this work, we analyse different translated texts in terms of various text features. We compare two types of human translations, professional and students’, and machine translation outputs in terms of lexical and grammatical variety, sentence length,as well as frequencies of different POS tags and POS-trigrams. Our experimentsare carried out on parallel translations into three languages, Croatian, Finnish andRussian, all originating from the same source English texts. Our results indicatethat machine translations are closest to the source text, followed by student translations. Also, student translations are similar both to professional as well as to MT, sometimes even more to MT. Furthermore, we identify sets of features which are convenient for distinguishing machine from human translations.
Maja Popovic, Ekaterina Lapshinova-Koltunski, Maarit Koponen
EAMT1
2023 Using MT for multilingual covid-19 case load prediction from social media texts
abstract
In the context of an epidemiological study involving multilingual social media, this paper reports on the ability of machine translation systems to preserve content relevant for a document classification task designed to determine whether the social media text is related to covid. The results indicate that machine translation does provide a feasible basis for scaling epidemiological social media surveillance to multiple languages. Moreover, a qualitative error analysis revealed that the majority of classification errors are not caused by MT errors.
Maja Popovic, Vasudevan Nedumpozhimana, Meegan Gower, Sneha Rautmare, Nishtha Jain, John D. Kelleher
EAMT1
2023 Leveraging machine translation for cross-lingual fine-grained cyberbullying classification amongst pre-adolescents
abstract
Abstract Cyberbullying is the wilful and repeated infliction of harm on an individual using the Internet and digital technologies. Similar to face-to-face bullying, cyberbullying can be captured formally using the Routine Activities Model (RAM) whereby the potential victim and bully are brought into proximity of one another via the interaction on online social networking (OSN) platforms. Although the impact of the COVID-19 (SARS-CoV-2) restrictions on the online presence of minors has yet to be fully grasped, studies have reported that 44% of pre-adolescents have encountered more cyberbullying incidents during the COVID-19 lockdown. Transparency reports shared by OSN companies indicate an increased take-downs of cyberbullying-related comments, posts or content by artificially intelligen moderation tools. However, in order to efficiently and effectively detect or identify whether a social media post or comment qualifies as cyberbullying, there are a number factors based on the RAM, which must be taken into account, which includes the identification of cyberbullying roles and forms. This demands the acquisition of large amounts of fine-grained annotated data which is costly and ethically challenging to produce. In addition where fine-grained datasets do exist they may be unavailable in the target language. Manual translation is costly and expensive, however, state-of-the-art neural machine translation offers a workaround. This study presents a first of its kind experiment in leveraging machine translation to automatically translate a unique pre-adolescent cyberbullying gold standard dataset in Italian with fine-grained annotations into English for training and testing a native binary classifier for pre-adolescent cyberbullying. In addition to contributing high-quality English reference translation of the source gold standard, our experiments indicate that the performance of our target binary classifier when trained on machine-translated English output is on par with the source (Italian) classifier.
Kanishk Verma, Maja Popovic, Alexandros Poulis, Yelena Cherkasova, Cathal Ó Hóbáin, Angela Mazzone, Tijana Milosevic, Brian Davis 0001
Nat. Lang. Eng.2
2022 Quantified Reproducibility Assessment of NLP Results
abstract
This paper describes and tests a method for carrying out quantified reproducibility assessment (QRA) that is based on concepts and definitions from metrology.QRA produces a single score estimating the degree of reproducibility of a given system and evaluation measure, on the basis of the scores from, and differences between, different reproductions.We test QRA on 18 system and evaluation measure combinations (involving diverse NLP tasks and types of evaluation), for each of which we have the original results and one to seven reproduction results.The proposed QRA method produces degree-of-reproducibility scores that are comparable across multiple reproductions not only of the same, but of different original studies.We find that the proposed method facilitates insights into causes of variation between reproductions, and allows conclusions to be drawn about what changes to system and/or evaluation design might lead to improved reproducibility.
Anya Belz, Maja Popovic, Simon Mille
ACL (1)2
2022 DiHuTra: a Parallel Corpus to Analyse Differences between Human Translations
abstract
This project aimed to design a corpus of parallel human translations (HTs) of the same source texts by professionals and students. The resulting corpus consists of English news and reviews source texts, their translations into Russian and Croatian, and translations of the reviews into Finnish. The corpus will be valuable for both studying variation in translation and evaluating machine translation (MT) systems.
Ekaterina Lapshinova-Koltunski, Maja Popovic, Maarit Koponen
EAMT2
2022 Leveraging Pre-trained Language Models for Gender Debiasing
abstract
Studying and mitigating gender and other biases in natural language have become important areas of research from both algorithmic and data perspectives. This paper explores the idea of reducing gender bias in a language generation context by generating gender variants of sentences. Previous work in this field has either been rule-based or required large amounts of gender balanced training data. These approaches are however not scalable across multiple languages, as creating data or rules for each language is costly and time-consuming. This work explores a light-weight method to generate gender variants for a given text using pre-trained language models as the resource, without any task-specific labelled data. The approach is designed to work on multiple languages with minimal changes in the form of heuristics. To showcase that, we have tested it on a high-resourced language, namely Spanish, and a low-resourced language from a different family, namely Serbian. The approach proved to work very well on Spanish, and while the results were less positive for Serbian, it showed potential even for languages where pre-trained models are less effective.
Nishtha Jain, Declan Groves, Lucia Specia, Maja Popovic
LREC4
2022 DiHuTra: a Parallel Corpus to Analyse Differences between Human Translations
abstract
This paper describes a new corpus of human translations which contains both professional and students translations. The data consists of English sources – texts from news and reviews – and their translations into Russian and Croatian, as well as of the subcorpus containing translations of the review texts into Finnish. All target languages represent mid-resourced and less or mid-investigated ones. The corpus will be valuable for studying variation in translation as it allows a direct comparison between human translations of the same source texts. The corpus will also be a valuable resource for evaluating machine translation systems. We believe that this resource will facilitate understanding and improvement of the quality issues in both human and machine translation. In the paper, we describe how the data was collected, provide information on translator groups and summarise the differences between the human translations at hand based on our preliminary results with shallow features.
Ekaterina Lapshinova-Koltunski, Maja Popovic, Maarit Koponen
LREC2
2021 Agree to Disagree: Analysis of Inter-Annotator Disagreements in Human Evaluation of Machine Translation Output
abstract
This work describes an analysis of interannotator disagreements in human evaluation of machine translation output.The errors in the analysed texts were marked by multiple annotators under guidance of different quality criteria: adequacy, comprehension, and an unspecified generic mixture of adequacy and fluency.Our results show that different criteria result in different disagreements, and indicate that a clear definition of quality criterion can improve the inter-annotator agreement.Furthermore, our results show that for certain linguistic phenomena which are not limited to one or two words (such as word ambiguity or gender) but span over several words or even entire phrases (such as negation or relative clause), disagreements do not necessarily represent "errors" or "noise" but are rather inherent to the evaluation process.On the other hand, for some other phenomena (such as omission or verb forms) agreement can be easily improved by providing more precise and detailed instructions to the evaluators.
Maja Popovic
CoNLL1
2021 A Reproduction Study of an Annotation-based Human Evaluation of MT Outputs
abstract
In this paper we report our reproduction study of the Croatian part of an annotation-based human evaluation of machine-translated user reviews (Popović, 2020).The work was carried out as part of the ReproGen Shared Task on Reproducibility of Human Evaluation in NLG.Our aim was to repeat the original study exactly, except for using a different set of evaluators.We describe the experimental design, characterise differences between original and reproduction study, and present the results from each study, along with analysis of the similarity between them.For the six main evaluation results of Major/Minor/All Comprehension error rates and Major/Minor/All Adequacy error rates, we find that (i) 4/6 system rankings are the same in both studies, (ii) the relative differences between systems are replicated well for Major Comprehension and Adequacy (Pearson's > 0.9), but not for the corresponding Minor error rates (Pearson's 0.36 for Adequacy, 0.67 for Comprehension), and (iii) the individual system scores for both types of Minor error rates had a higher degree of reproducibility than the corresponding Major error rates.We also examine inter-annotator agreement and compare the annotations obtained in the original and reproduction studies.
Maja Popovic, Anya Belz
INLG1
2021 On nature and causes of observed MT errors
abstract
This work describes analysis of nature and causes of MT errors observed by different evaluators under guidance of different quality criteria: adequacy and comprehension and and a not specified generic mixture of adequacy and fluency. We report results for three language pairs and two domains and eleven MT systems. Our findings indicate that and despite the fact that some of the identified phenomena depend on domain and/or language and the following set of phenomena can be considered as generally challenging for modern MT systems: rephrasing groups of words and translation of ambiguous source words and translating noun phrases and and mistranslations. Furthermore and we show that the quality criterion also has impact on error perception. Our findings indicate that comprehension and adequacy can be assessed simultaneously by different evaluators and so that comprehension and as an important quality criterion and can be included more often in human evaluations.
Maja Popovic
MTSummit (1)1
2021 From MT to LREV: managing the transition
Maja Popovic
Mach. Transl.1
2020 Informative Manual Evaluation of Machine Translation Output
abstract
This work proposes a new method for manual evaluation of Machine Translation (MT) output based on marking actual issues in the translated text.The novelty is that the evaluators are not assigning any scores, nor classifying errors, but marking all problematic parts (words, phrases, sentences) of the translation.The main advantage of this method is that the resulting annotations do not only provide overall scores by counting words with assigned tags, but can be further used for analysis of errors and challenging linguistic phenomena, as well as inter-annotator disagreements.Detailed analysis and understanding of actual problems are not enabled by typical manual evaluations where the annotators are asked to assign overall scores or to rank two or more translations.The proposed method is very general: it can be applied on any genre/domain and language pair, and it can be guided by various types of quality criteria.Also, it is not restricted to MT output, but can be used for other types of generated text.
Maja Popovic
COLING1
2020 Relations between comprehensibility and adequacy errors in machine translation output
abstract
This work presents a detailed analysis of translation errors perceived by readers as comprehensibility and/or adequacy issues.The main finding is that good comprehensibility, similarly to good fluency, can mask a number of adequacy errors.Of all major adequacy errors, 30% were fully comprehensible, thus fully misleading the reader to accept the incorrect information.Another 25% of major adequacy errors were perceived as almost comprehensible, thus being potentially misleading.Also, a vast majority of omissions (about 70%) is hidden by comprehensibility.Further analysis of misleading translations revealed that the most frequent error types are ambiguity, mistranslation, noun phrase error, word-by-word translation, untranslated word, subject-verb agreement, and spelling error in the source text.However, none of these error types appears exclusively in misleading translations, but are also frequent in fully incorrect (incomprehensible inadequate) and discarded correct (incomprehensible adequate) translations.Deeper analysis is needed to potentially detect underlying phenomena specifically related to misleading translations.
Maja Popovic
CoNLL1
2020 On the differences between human translations
abstract
Many studies have confirmed that translated texts exhibit different features than texts originally written in the given language. This work explores texts translated by different translators taking into account expertise and native language. A set of computational analyses was conducted on three language pairs, English-Croatian, German-French and English-Finnish, and the results show that each of the factors has certain influence on the features of the translated texts, especially on sentence length and lexical richness. The results also indicate that for translations used for machine translation evaluation, it is important to specify these factors, especially if comparing machine translation quality with human translation quality is involved.
Maja Popovic
EAMT1
2020 QRev: Machine Translation of User Reviews: What Influences the Translation Quality?
abstract
This project aims to identify the important aspects of translation quality of user reviews which will represent a starting point for developing better automatic MT metrics and challenge test sets, and will be also helpful for developing MT systems for this genre. We work on two types of reviews: Amazon products and IMDb movies, written in English and translated into two closely related target languages, Croatian and Serbian.
Maja Popovic
EAMT1
2020 On Context Span Needed for Machine Translation Evaluation
abstract
Despite increasing efforts to improve evaluation of machine translation (MT) by going beyond the sentence level to the document level, the definition of what exactly constitutes a “document level” is still not clear. This work deals with the context span necessary for a more reliable MT evaluation. We report results from a series of surveys involving three domains and 18 target languages designed to identify the necessary context span as well as issues related to it. Our findings indicate that, despite the fact that some issues and spans are strongly dependent on domain and on the target language, a number of common patterns can be observed so that general guidelines for context-aware MT evaluation can be drawn.
Sheila Castilho, Maja Popovic, Andy Way
LREC2
2019 On reducing translation shifts in translations intended for MT evaluation
Maja Popovic
MTSummit (2)1
2019 Automatic error classification with multiple error labels
Maja Popovic, David Vilar
MTSummit (1)1
2019 Editors' foreword to the special issue on human factors in neural machine translation
Sheila Castilho, Federico Gaspari, Joss Moorkens, Maja Popovic, Antonio Toral
Mach. Transl.4
2018 Proceedings of the 21st Annual Conference of the European Association for Machine Translation
Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Miquel Esplà-Gomis, Maja Popovic, Celia Rico, Joachim Van den Bogaert, Mikel L. Forcada
EAMT4
2018 A Multilingual Wikified Data Set of Educational Material
Iris Hendrickx, Eirini Takoulidou, Thanasis Naskos, Katia Kermanidis, Vilelmini Sosoni, Hugo De Vos, Maria Stasimioti, Menno van Zaanen, Panayota Georgakopoulou, Valia Kordoni, Maja Popovic, Markus Egg, Antal van den Bosch
LREC11
2018 Language-related issues for NMT and PBMT for English-German and English-Serbian
Maja Popovic
Mach. Transl.1
2016 Potential and Limits of Using Post-edits as Reference Translations for MT Evaluation
Maja Popovic, Mihael Arcan, Arle Lommel
EAMT1
2016 Can Text Simplification Help Machine Translation?
Sanja Stajner, Maja Popovic
EAMT2
2016 Tools and Guidelines for Principled Machine Translation Development
Nora Aranberri, Eleftherios Avramidis, Aljoscha Burchardt, Ondrej Klejch, Martin Popel, Maja Popovic
LREC6
2016 PE2rr Corpus: Manual Error Annotation of Automatically Pre-annotated MT Post-edits
Maja Popovic, Mihael Arcan
LREC1
2015 Identifying main obstacles for statistical machine translation of morphologically rich South Slavic languages
Maja Popovic, Mihael Arcan
EAMT1
2015 Poor man's lemmatisation for automatic error classification
Maja Popovic, Mihael Arcan, Eleftherios Avramidis, Aljoscha Burchardt, Arle Lommel
EAMT1
2014 Using a new analytic measure for the annotation and analysis of MT errors on real data
Arle Lommel, Aljoscha Burchardt, Maja Popovic, Kim Harris, Eleftherios Avramidis, Hans Uszkoreit
EAMT3
2014 Relations between different types of post-editing operations, cognitive effort and temporal effort
Maja Popovic, Arle Lommel, Aljoscha Burchardt, Eleftherios Avramidis, Hans Uszkoreit
EAMT1
2014 The taraXÜ corpus of human-annotated machine translations
Eleftherios Avramidis, Aljoscha Burchardt, Sabine Hunsicker, Maja Popovic, Cindy Tscherwinka, David Vilar, Hans Uszkoreit
LREC4
2012 Involving Language Professionals in the Evaluation of Machine Translation
Eleftherios Avramidis, Aljoscha Burchardt, Christian Federmann, Maja Popovic, Cindy Tscherwinka, David Vilar
LREC4
2012 Automatic MT Error Analysis: Hjerson Helping Addicter
Jan Berka, Ondrej Bojar, Mark Fishel, Maja Popovic, Daniel Zeman
LREC4
2012 Terra: a Collection of Translation Error-Annotated Corpora
Mark Fishel, Ondrej Bojar, Maja Popovic
LREC3
2012 Study and correlation analysis of linguistic, perceptual, and automatic machine translation evaluations
abstract
Abstract Evaluation of machine translation output is an important task. Various human evaluation techniques as well as automatic metrics have been proposed and investigated in the last decade. However, very few evaluation methods take the linguistic aspect into account. In this article, we use an objective evaluation method for machine translation output that classifies all translation errors into one of the five following linguistic levels: orthographic, morphological, lexical, semantic, and syntactic. Linguistic guidelines for the target language are required, and human evaluators use them in to classify the output errors. The experiments are performed on English‐to‐Catalan and Spanish‐to‐Catalan translation outputs generated by four different systems: 2 rule‐based and 2 statistical. All translations are evaluated using the 3 following methods: a standard human perceptual evaluation method, several widely used automatic metrics, and the human linguistic evaluation. Pearson and Spearman correlation coefficients between the linguistic, perceptual, and automatic results are then calculated, showing that the semantic level correlates significantly with both perceptual evaluation and automatic metrics.
Mireia Farrús, Marta R. Costa-jussà, Maja Popovic
J. Assoc. Inf. Sci. Technol.3
2011 From Human to Automatic Error Classification for Machine Translation Output
Maja Popovic, Aljoscha Burchardt
EAMT1
2011 Towards Automatic Error Analysis of Machine Translation Output
abstract
Evaluation and error analysis of machine translation output are important but difficult tasks. In this article, we propose a framework for automatic error analysis and classification based on the identification of actual erroneous words using the algorithms for computation of Word Error Rate (WER) and Position-independent word Error Rate (PER), which is just a very first step towards development of automatic evaluation measures that provide more specific information of certain translation problems. The proposed approach enables the use of various types of linguistic knowledge in order to classify translation errors in many different ways. This work focuses on one possible set-up, namely, on five error categories: inflectional errors, errors due to wrong word order, missing words, extra words, and incorrect lexical choices. For each of the categories, we analyze the contribution of various POS classes. We compared the results of automatic error analysis with the results of human error analysis in order to investigate two possible applications: estimating the contribution of each error type in a given translation output in order to identify the main sources of errors for a given translation system, and comparing different translation outputs using the introduced error categories in order to obtain more information about advantages and disadvantages of different systems and possibilites for improvements, as well as about advantages and disadvantages of applied methods for improvements. We used Arabic–English Newswire and Broadcast News and Chinese–English Newswire outputs created in the framework of the GALE project, several Spanish and English European Parliament outputs generated during the TC-Star project, and three German–English outputs generated in the framework of the fourth Machine Translation Workshop. We show that our results correlate very well with the results of a human error analysis, and that all our metrics except the extra words reflect well the differences between different versions of the same translation system as well as the differences between different translation systems.
Maja Popovic, Hermann Ney
Comput. Linguistics1
2006 POS-based Word Reorderings for Statistical Machine Translation
Maja Popovic, Hermann Ney
LREC1
2005 Exploiting phrasal lexica and additional morpho-syntactic language resources for statistical machine translation with scarce training data
Maja Popovic, Hermann Ney
EAMT1
2004 Improving Word Alignment Quality using Morpho-syntactic Information
Hermann Ney, Maja Popovic
COLING2
2004 Error Measures and Bayes Decision Rules Revisited with Applications to POS Tagging
Hermann Ney, Maja Popovic, David Suendermann-Oeft
EMNLP2
2004 Towards the Use of Word Stems and Suffixes for Statistical Machine Translation
Maja Popovic, Hermann Ney
LREC1