Emmanouil Zaranis

dblp:305/3517 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 61% Machine translation · 39%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
0.912025
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation · ACL (1) 2025
Machine learning › Trustworthy machine learning › fairness
gender bias
0.912025
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation · ACL (1) 2025
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation
0.912025
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation · ACL (1) 2025
Natural language and speech › Machine translation
machine translation evaluation
0.312025
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation · ACL (1) 2025
YearPublicationVenuePosition
2026 A Context-aware Framework for Translation-mediated Conversations
abstract
Abstract Automatic translation systems offer a powerful solution to bridge language barriers in scenarios where participants do not share a common language. However, these systems can introduce errors leading to misunderstandings and conversation breakdown. A key issue is that current systems fail to incorporate the rich contextual information necessary to resolve ambiguities and omitted details, resulting in literal, inappropriate, or misaligned translations. In this work, we present a framework to improve large language model-based translation systems by incorporating contextual information in bilingual conversational settings during training and inference. We validate our proposed framework on two task-oriented domains: customer chat and user-assistant interaction. Across both settings, the system produced by our framework—TowerChat—consistently results in better translations than state-of-the-art systems like GPT-4o and TowerInstruct, as measured by multiple automatic translation quality metrics on several language pairs. We also show that the resulting model leverages context in an intended and interpretable way, improving consistency between the conveyed message and the generated translations.1
José Pombal, Sweta Agrawal, Emmanouil Zaranis, Patrick Fernandes, André F. T. Martins
Trans. Assoc. Comput. Linguistics3
2025 Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
abstract
Quality estimation (QE)-the automatic assessment of translation quality-has recently become crucial across several stages of the translation pipeline, from data curation to training and decoding.While QE metrics have been optimized to align with human judgments, whether they encode social biases has been largely overlooked.Biased QE risks favoring certain demographic groups over others, e.g., by exacerbating gaps in visibility and usability.This paper defines and investigates gender bias of QE metrics and discusses its downstream implications for machine translation (MT).Experiments with state-ofthe-art QE metrics across multiple domains, datasets, and languages reveal significant bias.When a human entity's gender in the source is undisclosed, masculine-inflected translations score higher than feminine-inflected ones, and gender-neutral translations are penalized.Even when contextual cues disambiguate gender, using context-aware QE metrics leads to more errors in selecting the correct translation inflection for feminine referents than for masculine ones.Moreover, a biased QE metric affects data filtering and quality-aware decoding.Our findings underscore the need for a renewed focus on developing and evaluating QE metrics centered on gender. 1
Emmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, André F. T. Martins
ACL (1)1