VLDB 2026 Research / reviewers in the wild / expert
Ekaterina Lapshinova-Koltunski
dblp:76/8162
· DBLP profile ↗
20ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0002-5618-8087ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audio description between MT translation and recreation: An Interview Study for the Language Pair English-GermanabstractThis study examines the machine translation of audio descriptions (AD) as an alternative to producing new AD for audiovisual formats in a foreign language. To assess acceptance and comprehensibility among German users, a survey was conducted with blind and visually impaired participants, examining key AD strategies, such as character description and naming, facial expressions and gestures, and spatio-temporal settings. Participants compared machine-translated English AD with original German AD and provided feedback on these aspects. Results showed overall acceptance of the translated AD, although the original was generally preferred. Findings suggest that AD translation is feasible for the German audience, but further studies are needed on machine translation, production costs, as well as larger-scale user studies. Merle Sauter, Ekaterina Lapshinova-Koltunski, Sylvia Jaki |
EAMT (1) | 2 |
| 2026 | Quality and Comprehensibility of Interlingual Subtitles Produced by Humans or with MachinesabstractThe present paper focuses on the analysis of automatic subtitles produced with three different systems. We compare the outputs among each other paying attention to the categories of quality derived from audio-visual translation quality research. Besides that, we also consider comprehensibility of the produced subtitles. Additionally, we analyse the automatic evaluation scores to assess the overall quality. Our results show that automatically generated subtitles subtitles remain below human standards in quality and comprehensibility. Lara Shoana Schlüter, Ekaterina Lapshinova-Koltunski, Sylvia Jaki |
EAMT (1) | 2 |
| 2025 | Human- or machine-translated subtitles: Who can tell them apart?abstractThis contribution investigates whether machine-translated subtitles can be easily distinguished from human-translated ones. For this, we run an experiment using two versions of German subtitles for an English television series: (1)produced manually by professional subtitlers, and (2) translated automatically with a Large Language Model (LLM), i.e., GPT4. Our participants were students of translation studies with varying experience in subtitling and the use of machine translation. We asked participants to guess if the subtitles for a selection of video clips had been translated manually or automatically. Apart from analysing whether machine-translated subtitles are distinguishable from human-translated ones, we also seek for indicators of the differences between human and machine translations. Our results show that although it is overall hard to differentiate between human and machine translations, there are some differences. Notably, the more experience the humans have with translation and subtitling, the more able they are to tell apart the two translation variants. Ekaterina Lapshinova-Koltunski, Sylvia Jaki, Maren Bolz, Merle Sauter |
MTSummit (1) | 1 |
| 2024 | Evaluation of intralingual machine translation for health communicationabstractIn this paper, we describe results of a study on evaluation of intralingual machine translation. The study focuses on machine translations of medical texts into Plain German. The automatically simplified texts were compared with manually simplified texts (i.e., simplified by human experts) as well as with the underlying, unsimplified source texts. We analyse the quality of outputs from three models based on different criteria, such as correctness, readability, and syntactic complexity. We compare the outputs of the three models under analysis between each other, as well as with the existing human translations. The study revealed that system performance depends on the evaluation criteria used and that only one of the three models showed strong similarities to the human translations. Furthermore, we identified various types of errors in all three models. These included not only grammatical mistakes and misspellings, but also incorrect explanations of technical terms and false statements, which in turn led to serious content-related mistakes. Silvana Deilen, Ekaterina Lapshinova-Koltunski, Sergio Hernández Garrido, Julian Hörner, Christiane Maaß, Vanessa Theel, Sophie Ziemer |
EAMT (1) | 2 |
| 2023 | Computational analysis of different translations: by professionals, students and machinesabstractIn this work, we analyse different translated texts in terms of various text features. We compare two types of human translations, professional and students’, and machine translation outputs in terms of lexical and grammatical variety, sentence length,as well as frequencies of different POS tags and POS-trigrams. Our experimentsare carried out on parallel translations into three languages, Croatian, Finnish andRussian, all originating from the same source English texts. Our results indicatethat machine translations are closest to the source text, followed by student translations. Also, student translations are similar both to professional as well as to MT, sometimes even more to MT. Furthermore, we identify sets of features which are convenient for distinguishing machine from human translations. Maja Popovic, Ekaterina Lapshinova-Koltunski, Maarit Koponen |
EAMT | 2 |
| 2023 | Investigating Explicitation of Discourse Connectives in Translation using Automatic AnnotationsabstractDiscourse relations have different patterns of marking across different languages.As a result, discourse connectives are often added, omitted, or rephrased in translation.Prior work has shown a tendency for explicitation of discourse connectives, but such work was conducted using restricted sample sizes due to difficulty of connective identification and alignment.The current study exploits automatic methods to facilitate a large-scale study of connectives in English and German parallel texts.Our results based on over 300 types and 18000 instances of aligned connectives and an empirical approach to compare the cross-lingual specificity gap provide strong evidence of the Explicitation Hypothesis.We conclude that discourse relations are indeed more explicit in translation than texts written originally in the same language.Automatic annotations allow us to carry out translation studies of discourse relations on a large scale.Our methodology using relative entropy to study the specificity of connectives also provides more fine-grained insights into translation patterns. Frances Yung, Merel C. J. Scholman, Ekaterina Lapshinova-Koltunski, Christina Pollkläsener, Vera Demberg |
SIGDIAL | 3 |
| 2022 | DiHuTra: a Parallel Corpus to Analyse Differences between Human TranslationsabstractThis project aimed to design a corpus of parallel human translations (HTs) of the same source texts by professionals and students. The resulting corpus consists of English news and reviews source texts, their translations into Russian and Croatian, and translations of the reviews into Finnish. The corpus will be valuable for both studying variation in translation and evaluating machine translation (MT) systems. Ekaterina Lapshinova-Koltunski, Maja Popovic, Maarit Koponen |
EAMT | 1 |
| 2022 | ParCorFull2.0: a Parallel Corpus Annotated with Full CoreferenceabstractIn this paper, we describe ParCorFull2.0, a parallel corpus annotated with full coreference chains for multiple languages, which is an extension of the existing corpus ParCorFull (Lapshinova-Koltunski et al., 2018). Similar to the previous version, this corpus has been created to address translation of coreference across languages, a phenomenon still challenging for machine translation (MT) and other multilingual natural language processing (NLP) applications. The current version of the corpus that we present here contains not only parallel texts for the language pair English-German, but also for English-French and English-Portuguese, which are all major European languages. The new language pairs belong to the Romance languages. The addition of a new language group creates a need of extension not only in terms of texts added, but also in terms of the annotation guidelines. Both French and Portuguese contain structures not found in English and German. Moreover, Portuguese is a pro-drop language bringing even more systemic differences in the realisation of coreference into our cross-lingual resources. These differences cause problems for multilingual coreference resolution and machine translation. Our parallel corpus with full annotation of coreference will be a valuable resource with a variety of uses not only for NLP applications, but also for contrastive linguists and researchers in translation studies. Ekaterina Lapshinova-Koltunski, Pedro Augusto Ferreira, Elina Lartaud, Christian Hardmeier |
LREC | 1 |
| 2022 | DiHuTra: a Parallel Corpus to Analyse Differences between Human TranslationsabstractThis paper describes a new corpus of human translations which contains both professional and students translations. The data consists of English sources – texts from news and reviews – and their translations into Russian and Croatian, as well as of the subcorpus containing translations of the review texts into Finnish. All target languages represent mid-resourced and less or mid-investigated ones. The corpus will be valuable for studying variation in translation as it allows a direct comparison between human translations of the same source texts. The corpus will also be a valuable resource for evaluating machine translation systems. We believe that this resource will facilitate understanding and improvement of the quality issues in both human and machine translation. In the paper, we describe how the data was collected, provide information on translator groups and summarise the differences between the human translations at hand based on our preliminary results with shallow features. Ekaterina Lapshinova-Koltunski, Maja Popovic, Maarit Koponen |
LREC | 1 |
| 2022 | EPIC UdS - Creation and Applications of a Simultaneous Interpreting CorpusabstractIn this paper, we describe the creation and annotation of EPIC UdS, a multilingual corpus of simultaneous interpreting for English, German and Spanish. We give an overview of the comparable and parallel, aligned corpus variants and explore various applications of the corpus. What makes EPIC UdS relevant is that it is one of the rare interpreting corpora that includes transcripts suitable for research on more than one language pair and on interpreting with regard to German. It not only contains transcribed speeches, but also rich metadata and fine-grained linguistic annotations tailored for diverse applications across a broad range of linguistic subfields. Heike Przybyl, Ekaterina Lapshinova-Koltunski, Katrin Menzel, Stefan Fischer 0008, Elke Teich |
LREC | 2 |
| 2020 | Lexicogrammatic translationese across two targets and competence levelsabstractThis research employs genre-comparable data from a number of parallel and comparable corpora to explore the specificity of translations from English into German and Russian produced by students and professional translators. We introduce an elaborate set of human-interpretable lexicogrammatic translationese indicators and calculate the amount of translationese manifested in the data for each target language and translation variety. By placing translations into the same feature space as their sources and the genre-comparable non-translated reference texts in the target language, we observe two separate translationese effects: a shift of translations into the gap between the two languages and a shift away from either language. These trends are linked to the features that contribute to each of the effects. Finally, we compare the translation varieties and find out that the professionalism levels seem to have some correlation with the amount and types of translationese detected, while each language pair demonstrates a specific socio-linguistically determined combination of the translationese effects. Maria Kunilovskaya, Ekaterina Lapshinova-Koltunski |
LREC | 2 |
| 2018 | ParCorFull: a Parallel Corpus Annotated with Full Coreference
Ekaterina Lapshinova-Koltunski, Christian Hardmeier, Pauline Krielke |
LREC | 1 |
| 2016 | From Interoperable Annotations towards Interoperable Resources: A Multilingual Approach to the Analysis of Discourse
Ekaterina Lapshinova-Koltunski, Kerstin Kunz, Anna Nedoluzhko |
LREC | 1 |
| 2016 | Information Density and Quality Estimation Features as Translationese Indicators for Human Translation ClassificationabstractRaphael Rubino, Ekaterina Lapshinova-Koltunski, Josef van Genabith. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Raphaël Rubino, Ekaterina Lapshinova-Koltunski, Josef van Genabith |
HLT-NAACL | 2 |
| 2016 | The linguistic construal of disciplinarity: A data-mining approach using register featuresabstractWe analyze the linguistic evolution of selected scientific disciplines over a 30‐year time span (1970s to 2000s). Our focus is on four highly specialized disciplines at the boundaries of computer science that emerged during that time: computational linguistics, bioinformatics, digital construction, and microelectronics. Our analysis is driven by the question whether these disciplines develop a distinctive language use—both individually and collectively—over the given time period. The data set is the English Scientific Text Corpus (scitex), which includes texts from the 1970s/1980s and early 2000s. Our theoretical basis is register theory. In terms of methods, we combine corpus‐based methods of feature extraction (various aggregated features [part‐of‐speech based], n‐grams, lexico‐grammatical patterns) and automatic text classification. The results of our research are directly relevant to the study of linguistic variation and languages for specific purposes (LSP) and have implications for various natural language processing (NLP) tasks, for example, authorship attribution, text mining, or training NLP tools. Elke Teich, Stefania Degaetano-Ortlieb, Peter Fankhauser, Hannah Kermes, Ekaterina Lapshinova-Koltunski |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2015 | Register-based machine translation evaluation with text classification techniques
Mihaela Vela, Ekaterina Lapshinova-Koltunski |
MTSummit | 2 |
| 2014 | Data Mining with Shallow vs. Linguistic Features to Study Diversification of Scientific Registers
Stefania Degaetano-Ortlieb, Peter Fankhauser, Hannah Kermes, Ekaterina Lapshinova-Koltunski, Noam Ordan, Elke Teich |
LREC | 4 |
| 2012 | Coreference in Spoken vs. Written Texts: a Corpus-based Analysis
Marilisa Amoia, Kerstin Kunz, Ekaterina Lapshinova-Koltunski |
LREC | 3 |
| 2012 | Feature Discovery for Diachronic Register Analysis: a Semi-Automatic Approach
Stefania Degaetano-Ortlieb, Ekaterina Lapshinova-Koltunski, Elke Teich |
LREC | 2 |
| 2008 | Head or Non-head? Semi-automatic Procedures for Extracting and Classifying Subcategorisation Properties of Compounds
Ekaterina Lapshinova-Koltunski, Ulrich Heid |
LREC | 1 |