Dimitar Sht. Shterionov

dblp:35/8575 · also Dimitar Shterionov · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0001-6300-797XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 7 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 2 first-authorTheory of computation · 2 · 1 first-author
YearPublicationVenuePosition
2026 Diversity and Homogenisation in Generative AI Translation: A Comparative Study of English-Dutch Translation Across Domains
abstract
Generative AI tools, such as ChatGPT, are applied to a wide range of languagerelated tasks, including translation. Despite their current popularity among users and researchers and the impressive results obtained on several benchmarks (Kocmi et al., 2024a; Deutsch et al., 2025), their potential side-effects on languages and translations are still understudied (Vanmassenhove, 2025). The paradigm shift from Machine Translation (MT) to Generative AI Translation (GAIT) likely calls for a reconsideration of our assessment and evaluation metrics and practices. In this work, we focus on GAIT by analyzing translations from four multilingual large language models (MLLMs), mBART, Jamba-1.5-large, GPT 4o and DeepSeek R1 applied to three different domains (news, literature and poetry) for the English-Dutch language pair. Focusing on metrics related to lexical and textual diversity, we find that while GAIT text for literature if of significantly high lexical and grammatical richness, that is not the case for news and poetry. We also assess the homogeneity of AI-generated text through a set of clustering and classification experiments. In addition to a clear separation between human- and AI-generated content, our results indicate that GAIT output is more homogeneous among MLLMs.
Dimitar Sht. Shterionov, Noa van Helleman, Eva Vanmassenhove
EAMT (1)1
2024 SignON - a Co-creative Machine Translation for Sign and Spoken Languages (end-of-project results, contributions and lessons learned)
abstract
SignON, a 3-year Horizon 20202 project addressing the lack of technology and services for MT between sign languages (SLs) and spoken languages (SpLs) ended in December 2023. SignON was unprecedented. Not only it addressed the wider complexity of the aforementioned problem – from research and development of recognition, translation and synthesis, through development of easy-to-use mobile applications and a cloud-based framework to do the “heavy lifting” as well as to establishing ethical, privacy and inclusivenesspolicies and operation guidelines – but also engaged with the deaf and hard of hearing communities in an effective co-creation approach where these main stakeholders drove the development in the right direction and had the final say.Currently we are witnessing advances in natural language processing for SLs, including MT. SignON was one of the largest projects that contributed to this surge with 17 partners and more than 60 consortium members, working in parallel with other international and European initiatives, such as project EASIER and others.
Dimitar Sht. Shterionov, Vincent Vandeghinste, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Andy Way, Josep Blat, Frankie Picron, Davy Van Landuyt, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Caro Brosens, Jorn Rijckaert, Víctor Ubieto Nogales, Bram Vanroy, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion
EAMT (2)1
2023 First WMT Shared Task on Sign Language Translation (WMT-SLT22)
abstract
This paper is a brief summary of the First WMT Shared Task on Sign Language Translation (WMT-SLT22), a project partly funded by EAMT. The focus of this shared task is automatic translation between signed and spoken languages. Details can be found on our website (https://www.wmt-slt.com/) or in the findings paper (Müller et al., 2022).
Mathias Müller 0002, Sarah Ebling, Eleftherios Avramidis, Alessia Battisti, Michèle Berger, Richard Bowden, Annelies Braffort, Necati Cihan Camgöz, Cristina España-Bonet, Roman Grundkiewicz, Zifan Jiang, Oscar Koller, Amit Moryossef, Regula Perrollaz, Sabine Reinhard, Annette Rios, Dimitar Sht. Shterionov, Sandra Sidler-Miserez, Katja Tissi, Davy Van Landuyt
EAMT17
2023 Tailoring Domain Adaptation for Machine Translation Quality Estimation
abstract
While quality estimation (QE) can play an important role in the translation process, its effectiveness relies on the availability and quality of training data. For QE in particular, high-quality labeled data is often lacking due to the high-cost and effort associated with labeling such data. Aside from the data scarcity challenge, QE models should also be generalizabile, i.e., they should be able to handle data from different domains, both generic and specific. To alleviate these two main issues — data scarcity and domain mismatch — this paper combines domain adaptation and data augmentation within a robust QE system. Our method is to first train a generic QE model and then fine-tune it on a specific domain while retaining generic knowledge. Our results show a significant improvement for all the language pairs investigated, better cross-lingual inference, and a superior performance in zero-shot learning scenarios as compared to state-of-the-art baselines.
Javad PourMostafa Roshan Sharami, Dimitar Sht. Shterionov, Frédéric Blain, Eva Vanmassenhove, Mirella De Sisto, Chris Emmery, Pieter Spronck
EAMT2
2023 GoSt-ParC-Sign: Gold Standard Parallel Corpus of Sign and spoken language
abstract
Good quality training data for Sign Language Machine Translation (SLMT) is extremely scarce, and this is one of the challenges that any project focusing on Machine Translation (MT) which also targets sign languages is currently facing. The goal of this ongoing project is to create a parallel corpus of authentic Flemish Sign Language (VGT) and written Dutch which can be employed as gold standard in automated sign language translation. The availability of a gold standard corpus like Gost-ParC-Sign can facilitate the advances of SLMT; consequently, it supports and promotes inclusiveness in MT and, on a more general level, in language technology
Mirella De Sisto, Vincent Vandeghinste, Lien Soetemans, Caro Brosens, Dimitar Sht. Shterionov
EAMT5
2023 SignON: Sign Language Translation. Progress and challenges
abstract
SignON (https://signon-project.eu/) is a Horizon 2020 project, running from 2021 until the end of 2023, which addresses the lack of technology and services for the automatic translation between sign languages (SLs) and spoken languages, through an inclusive, human-centric solution, hence contributing to the repertoire of communication media for deaf, hard of hearing (DHH) and hearing individuals. In this paper, we present an update of the status of the project, describing the approaches developed to address the challenges and peculiarities of SL machine translation (SLMT).
Vincent Vandeghinste, Dimitar Sht. Shterionov, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Josep Blat, Frankie Picron, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Jorn Rijckaert, Bram Vanroy, Víctor Ubieto Nogales, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion
EAMT2
2022 A Quality Estimation and Quality Evaluation Tool for the Translation Industry
abstract
With the increase in machine translation (MT) quality over the latest years, it has now become a common practice to integrate MT in the workflow of language service providers (LSPs) and other actors in the translation industry. With MT having a direct impact on the translation workflow, it is important not only to use high-quality MT systems, but also to understand the quality dimension so that the humans involved in the translation workflow can make informed decisions. The evaluation and monitoring of MT output quality has become one of the essential aspects of language technology management in LSPs’ workflows. First, a general practice is to carry out human tests to evaluate MT output quality before deployment. Second, a quality estimate of the translated text, thus after deployment, can inform post editors or even represent post-editing effort. In the former case, based on the quality assessment of a candidate engine, an informed decision can be made whether the engine would be deployed for production or not. In the latter, a quality estimate of the translation output can guide the human post-editor or even make rough approximations of the post-editing effort. Quality of an MT engine can be assessed on document- or on sentence-level. A tool to jointly provide all these functionalities does not exist yet. The overall objective of the project presented in this paper is to develop an MT quality assessment (MTQA) tool that simplifies the quality assessment of MT engines, combining quality evaluation and quality estimation on document- and sentence- level.
Elena Murgolo, Javad PourMostafa Roshan Sharami, Dimitar Sht. Shterionov
EAMT3
2022 Sign Language Translation: Ongoing Development, Challenges and Innovations in the SignON Project
abstract
The SignON project (www.signon-project.eu) focuses on the research and development of a Sign Language (SL) translation mobile application and an open communications framework. SignON rectifies the lack of technology and services for the automatic translation between signed and spoken languages, through an inclusive, humancentric solution which facilitates communication between deaf, hard of hearing (DHH) and hearing individuals. We present an overview of the current status of the project, describing the milestones reached to date and the approaches that are being developed to address the challenges and peculiarities of Sign Language Machine Translation (SLMT).
Dimitar Sht. Shterionov, Mirella De Sisto, Vincent Vandeghinste, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Josep Blat, Frankie Picron, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Jorn Rijckaert
EAMT1
2022 Challenges with Sign Language Datasets for Sign Language Recognition and Translation
abstract
Sign Languages (SLs) are the primary means of communication for at least half a million people in Europe alone. However, the development of SL recognition and translation tools is slowed down by a series of obstacles concerning resource scarcity and standardization issues in the available data. The former challenge relates to the volume of data available for machine learning as well as the time required to collect and process new data. The latter obstacle is linked to the variety of the data, i.e., annotation formats are not unified and vary amongst different resources. The available data formats are often not suitable for machine learning, obstructing the provision of automatic tools based on neural models. In the present paper, we give an overview of these challenges by comparing various SL corpora and SL machine learning datasets. Furthermore, we propose a framework to address the lack of standardization at format level, unify the available resources and facilitate SL research for different languages. Our framework takes ELAN files as inputs and returns textual and visual data ready to train SL recognition and translation models. We present a proof of concept, training neural translation models on the data produced by the proposed framework.
Mirella De Sisto, Vincent Vandeghinste, Santiago Egea Gómez, Mathieu De Coster, Dimitar Sht. Shterionov, Horacio Saggion
LREC5
2021 Machine Translationese: Effects of Algorithmic Bias on Linguistic Complexity in Machine Translation
abstract
Recent studies in the field of Machine Translation (MT) and Natural Language Processing (NLP) have shown that existing models amplify biases observed in the training data.The amplification of biases in language technology has mainly been examined with respect to specific phenomena, such as gender bias.In this work, we go beyond the study of gender in MT and investigate how bias amplification might affect language in a broader sense.We hypothesize that the 'algorithmic bias', i.e. an exacerbation of frequently observed patterns in combination with a loss of less frequent ones, not only exacerbates societal biases present in current datasets but could also lead to an artificially impoverished language: 'machine translationese'.We assess the linguistic richness (on a lexical and morphological level) of translations created by different data-driven MT paradigms -phrase-based statistical (PB-SMT) and neural MT (NMT).Our experiments show that there is a loss of lexical and morphological richness in the translations produced by all investigated MT paradigms for two language pairs (EN↔FR and EN↔ES).
Eva Vanmassenhove, Dimitar Sht. Shterionov, Matthew Gwilliam
EACL2
2021 NeuTral Rewriter: A Rule-Based and Neural Approach to Automatic Rewriting into Gender Neutral Alternatives
abstract
Recent years have seen an increasing need for gender-neutral and inclusive language.Within the field of NLP, there are various mono-and bilingual use cases where gender inclusive language is appropriate, if not preferred due to ambiguity or uncertainty in terms of the gender of referents.In this work, we present a rulebased and a neural approach to gender-neutral rewriting for English along with manually curated synthetic data (WinoBias+) and natural data (OpenSubtitles and Reddit) benchmarks.A detailed manual and automatic evaluation highlights how our NeuTral Rewriter, trained on data generated by the rule-based approach, obtains word error rates (WER) below 0.18% on synthetic, in-domain and out-domain test sets.
Eva Vanmassenhove, Chris Emmery, Dimitar Sht. Shterionov
EMNLP (1)3
2021 A review of the state-of-the-art in automatic post-editing
abstract
This article presents a review of the evolution of automatic post-editing, a term that describes methods to improve the output of machine translation systems, based on knowledge extracted from datasets that include post-edited content. The article describes the specificity of automatic post-editing in comparison with other tasks in machine translation, and it discusses how it may function as a complement to them. Particular detail is given in the article to the five-year period that covers the shared tasks presented in WMT conferences (2015-2019). In this period, discussion of automatic post-editing evolved from the definition of its main parameters to an announced demise, associated with the difficulties in improving output obtained by neural methods, which was then followed by renewed interest. The article debates the role and relevance of automatic post-editing, both as an academic endeavour and as a useful application in commercial workflows.
Félix do Carmo, Dimitar Sht. Shterionov, Joss Moorkens, Joachim Wagner 0001, Murhaf Hossari, Eric Paquin, Dag Schmidtke, Declan Groves, Andy Way
Mach. Transl.2
2020 Selecting Backtranslated Data from Multiple Sources for Improved Neural Machine Translation
abstract
Machine translation (MT) has benefited from using synthetic training data originating from translating monolingual corpora, a technique known as backtranslation.Combining backtranslated data from different sources has led to better results than when using such data in isolation.In this work we analyse the impact that data translated with rule-based, phrasebased statistical and neural MT systems has on new MT systems.We use a real-world low-resource use-case (Basque-to-Spanish in the clinical domain) as well as a high-resource language pair (German-to-English) to test different scenarios with backtranslation and employ data selection to optimise the synthetic corpora.We exploit different data selection strategies in order to reduce the amount of data used, while at the same time maintaining highquality MT systems.We further tune the data selection method by taking into account the quality of the MT systems used for backtranslation and lexical diversity of the resulting corpora.Our experiments show that incorporating backtranslated data from different sources can be beneficial, and that availing of data selection can yield improved performance.
Xabier Soto, Dimitar Sht. Shterionov, Alberto Poncelas, Andy Way
ACL2
2020 A roadmap to neural automatic post-editing: an empirical approach
abstract
In a translation workflow, machine translation (MT) is almost always followed by a human post-editing step, where the raw MT output is corrected to meet required quality standards. To reduce the number of errors human translators need to correct, automatic post-editing (APE) methods have been developed and deployed in such workflows. With the advances in deep learning, neural APE (NPE) systems have outranked more traditional, statistical, ones. However, the plethora of options, variables and settings, as well as the relation between NPE performance and train/test data makes it difficult to select the most suitable approach for a given use case. In this article, we systematically analyse these different parameters with respect to NPE performance. We build an NPE "roadmap" to trace the different decision points and train a set of systems selecting different options through the roadmap. We also propose a novel approach for APE with data augmentation. We then analyse the performance of 15 of these systems and identify the best ones. In fact, the best systems are the ones that follow the newly-proposed method. The work presented in this article follows from a collaborative project between Microsoft and the ADAPT centre. The data provided by Microsoft originates from phrase-based statistical MT (PBSMT) systems employed in production. All tested NPE systems significantly increase the translation quality, proving the effectiveness of neural post-editing in the context of a commercial translation workflow that leverages PBSMT.
Dimitar Sht. Shterionov, Félix do Carmo, Joss Moorkens, Murhaf Hossari, Joachim Wagner 0001, Eric Paquin, Dag Schmidtke, Declan Groves, Andy Way
Mach. Transl.1
2019 When less is more in Neural Quality Estimation of Machine Translation. An industry case study
Dimitar Sht. Shterionov, Félix do Carmo, Joss Moorkens, Eric Paquin, Dag Schmidtke, Declan Groves, Andy Way
MTSummit (2)1
2019 Lost in Translation: Loss and Decay of Linguistic Richness in Machine Translation
Eva Vanmassenhove, Dimitar Sht. Shterionov, Andy Way
MTSummit (1)2
2018 Investigating Backtranslation in Neural Machine Translation
abstract
A prerequisite for training corpus-based machine translation (MT) systems – either Statistical MT (SMT) or Neural MT (NMT) – is the availability of high-quality parallel data. This is arguably more important today than ever before, as NMT has been shown in many studies to outperform SMT, but mostly when large parallel corpora are available; in cases where data is limited, SMT can still outperform NMT. Recently researchers have shown that back-translating monolingual data can be used to create synthetic parallel corpora, which in turn can be used in combination with authentic parallel data to train a highquality NMT system. Given that large collections of new parallel text become available only quite rarely, backtranslation has become the norm when building state-of-the-art NMT systems, especially in resource-poor scenarios. However, we assert that there are many unknown factors regarding the actual effects of back-translated data on the translation capabilities of an NMT model. Accordingly, in this work we investigate how using back-translated data as a training corpus – both as a separate standalone dataset as well as combined with human-generated parallel data – affects the performance of an NMT model. We use incrementally larger amounts of back-translated data to train a range of NMT systems for German-to-English, and analyse the resulting translation performance.
Alberto Poncelas, Dimitar Sht. Shterionov, Andy Way, Gideon Maillette de Buy Wenniger, Peyman Passban
EAMT2
2018 Human versus automatic quality evaluation of NMT and PBSMT
Dimitar Sht. Shterionov, Riccardo Superbo, Pat Nagle, Laura Casanellas, Tony O'Dowd, Andy Way
Mach. Transl.1
2017 Zero-Shot Translation for Indian Languages with Sparse Data
Giulia Mattoni, Pat Nagle, Carlos Collantes, Dimitar Sht. Shterionov
MTSummit (2)4
2015 Compacting Boolean Formulae for Inference in Probabilistic Logic Programming
Theofrastos Mantadelis, Dimitar Sht. Shterionov, Gerda Janssens
LPNMR2
2015 Implementation and Performance of Probabilistic Inference Pipelines
Dimitar Sht. Shterionov, Gerda Janssens
PADL1
2015 Inference and learning in probabilistic logic programs using weighted Boolean formulas
abstract
Abstract Probabilistic logic programs are logic programs in which some of the facts are annotated with probabilities. This paper investigates how classical inference and learning tasks known from the graphical model community can be tackled for probabilistic logic programs. Several such tasks, such as computing the marginals, given evidence and learning from (partial) interpretations, have not really been addressed for probabilistic logic programs before. The first contribution of this paper is a suite of efficient algorithms for various inference tasks. It is based on the conversion of the program and the queries and evidence to a weighted Boolean formula. This allows us to reduce inference tasks to well-studied tasks, such as weighted model counting, which can be solved using state-of-the-art methods known from the graphical model and knowledge compilation literature. The second contribution is an algorithm for parameter estimation in the learning from interpretations setting. The algorithm employs expectation-maximization, and is built on top of the developed inference algorithms. The proposed approach is experimentally evaluated. The results show that the inference algorithms improve upon the state of the art in probabilistic logic programming, and that it is indeed possible to learn the parameters of a probabilistic logic program from interpretations.
Daan Fierens, Guy Van den Broeck, Joris Renkens, Dimitar Sht. Shterionov, Bernd Gutmann, Ingo Thon, Gerda Janssens, Luc De Raedt
Theory Pract. Log. Program.4
2014 The Most Probable Explanation for Probabilistic Logic Programs with Annotated Disjunctions
Dimitar Sht. Shterionov, Joris Renkens, Jonas Vlasselaer, Angelika Kimmig, Wannes Meert, Gerda Janssens
ILP1
2013 Pattern-Based Compaction for ProbLog Inference
Dimitar Sht. Shterionov, Theofrastos Mantadelis, Gerda Janssens
Theory Pract. Log. Program.1