Eleftheria Briakou

dblp:217/4858 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 11 since 2021
YearPublicationVenuePosition
2025 SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?
abstract
Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Senyu Li, Jiayi Wang 0010, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani
EMNLP6
2025 Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
abstract
Data contamination—the accidental consumption of evaluation examples within the pre-training data—can undermine the validity of evaluation benchmarks. In this paper, we present a rigorous analysis of the effects of contamination on language models at 1B and 8B scales on the machine translation task. Starting from a carefully decontaminated train-test split, we systematically introduce contamination at various stages, scales, and data formats to isolate its effect and measure its impact on performance metrics. Our experiments reveal that contamination with both source and target substantially inflates BLEU scores, and this inflation is 2.5 times larger (up to 30 BLEU points) for 8B compared to 1B models. In contrast, source-only and target-only contamination generally produce smaller, less consistent over-estimations. Finally, we study how the temporal distribution and frequency of contaminated samples influence performance over-estimation across languages with varying degrees of data resources.
Yusuf Kocyigit, Eleftheria Briakou, Daniel Deutsch, Jiaming Luo, Colin Cherry, Markus Freitag
ICML2
2024 AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages
abstract
Jiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Anuoluwapo Aremu, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Hassan Ayinde, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Al-azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Abeeb Afolabi, Nnaemeka Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng’, Verrah Akinyi Otiende, Chinedu Emmanuel Mbonu, Sakayo Toadoum Sari, Yao Lu, Pontus Stenetorp. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jiayi Wang 0010, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin P. Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Aremu Anuoluwapo, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Ijeoma Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Ayinde Hassan, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Sabah Al-Azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Sakayo Toadoum Sari, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Afolabi Abeeb, Nnaemeka C. Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng', Verrah Otiende, Chinedu E. Mbonu, Pontus Stenetorp
NAACL-HLT6
2023 Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability
abstract
Large, multilingual language models exhibit surprisingly good zero-or few-shot machine translation capabilities, despite having never seen the intentionally-included translation examples provided to typical neural translation systems.We investigate the role of incidental bilingualism-the unintentional consumption of bilingual signals, including translation examples-in explaining the translation capabilities of large language models, taking the Pathways Language Model (PaLM) as a case study.We introduce a mixed-method approach to measure and understand incidental bilingualism at scale.We show that PaLM is exposed to over 30 million translation pairs across at least 44 languages.Furthermore, the amount of incidental bilingual content is highly correlated with the amount of monolingual in-language content for non-English languages.We relate incidental bilingual content to zero-shot prompts and show that it can be used to mine new prompts to improve PaLM's out-of-English zero-shot translation quality.Finally, in a series of small-scale ablations, we show that its presence has a substantial impact on translation capabilities, although this impact diminishes with model scale.
Eleftheria Briakou, Colin Cherry, George F. Foster
ACL (1)1
2023 Explaining with Contrastive Phrasal Highlighting: A Case Study in Assisting Humans to Detect Translation Differences
abstract
Explainable NLP techniques primarily explain by answering "Which tokens in the input are responsible for this prediction?".We argue that for NLP models that make predictions by comparing two input texts, it is more useful to explain by answering "What differences between the two inputs explain this prediction?".We introduce a technique to generate contrastive phrasal highlights that explain the predictions of a semantic divergence model via phrasealignment-guided erasure.We show that the resulting highlights match human rationales of cross-lingual semantic differences better than popular post-hoc saliency techniques and that they successfully help people detect finegrained meaning differences in human translations and critical machine translation errors.
Eleftheria Briakou, Navita Goyal, Marine Carpuat
EMNLP1
2023 What Else Do I Need to Know? The Effect of Background Information on Users' Reliance on QA Systems
abstract
Navita Goyal, Eleftheria Briakou, Amanda Liu, Connor Baumler, Claire Bonial, Jeffrey Micher, Clare Voss, Marine Carpuat, Hal Daumé III. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Navita Goyal, Eleftheria Briakou, Amanda Liu, Connor Baumler, Claire Bonial, Jeffrey Micher, Clare R. Voss, Marine Carpuat, Hal Daumé III
EMNLP2
2023 Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection
abstract
Abstract Neural sequence generation models are known to “hallucinate”, by producing outputs that are unrelated to the source text. These hallucinations are potentially harmful, yet it remains unclear in what conditions they arise and how to mitigate their impact. In this work, we first identify internal model symptoms of hallucinations by analyzing the relative token contributions to the generation in contrastive hallucinated vs. non-hallucinated outputs generated via source perturbations. We then show that these symptoms are reliable indicators of natural hallucinations, by using them to design a lightweight hallucination detector which outperforms both model-free baselines and strong classifiers based on quality estimation or large pre-trained models on manually annotated English-Chinese and German-English translation test beds.
Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, Marine Carpuat
Trans. Assoc. Comput. Linguistics3
2022 Can Synthetic Translations Improve Bitext Quality?
abstract
Synthetic translations have been used for a wide range of NLP tasks primarily as a means of data augmentation.This work explores, instead, how synthetic translations can be used to revise potentially imperfect reference translations in mined bitext.We find that synthetic samples can improve bitext quality without any additional bilingual supervision when they replace the originals based on a semantic equivalence classifier that helps mitigate NMT noise.The improved quality of the revised bitext is confirmed intrinsically via human evaluation and extrinsically through bilingual induction and MT tasks.
Eleftheria Briakou, Marine Carpuat
ACL (1)1
2021 Beyond Noise: Mitigating the Impact of Fine-grained Semantic Divergences on Neural Machine Translation
abstract
Eleftheria Briakou, Marine Carpuat. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Eleftheria Briakou, Marine Carpuat
ACL/IJCNLP (1)1
2021 Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer
abstract
While the field of style transfer (ST) has been growing rapidly, it has been hampered by a lack of standardized practices for automatic evaluation.In this paper, we evaluate leading ST automatic metrics on the oft-researched task of formality style transfer.Unlike previous evaluations, which focus solely on English, we expand our focus to Brazilian-Portuguese, French, and Italian, making this work the first multilingual evaluation of metrics in ST.We outline best practices for automatic evaluation in (formality) style transfer and identify several models that correlate well with human judgments and are robust across languages.We hope that this work will help accelerate development in ST, where human evaluation is often challenging to collect.
Eleftheria Briakou, Sweta Agrawal, Joel R. Tetreault, Marine Carpuat
EMNLP (1)1
2021 Olá, Bonjour, Salve! XFORMAL: A Benchmark for Multilingual Formality Style Transfer
abstract
Eleftheria Briakou, Di Lu, Ke Zhang, Joel Tetreault. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Eleftheria Briakou, Di Lu 0003, Ke Zhang 0013, Joel R. Tetreault
NAACL-HLT1
2020 Detecting Fine-Grained Cross-Lingual Semantic Divergences without Supervision by Learning to Rank
abstract
Detecting fine-grained differences in content conveyed in different languages matters for cross-lingual NLP and multilingual corpora analysis, but it is a challenging machine learning problem since annotation is expensive and hard to scale.This work improves the prediction and annotation of finegrained semantic divergences.We introduce a training strategy for multilingual BERT models by learning to rank synthetic divergent examples of varying granularity.We evaluate our models on the Rationalized English-French Semantic Divergences, a new dataset released with this work, consisting of English-French sentence-pairs annotated with semantic divergence classes and token-level rationales.Learning to rank helps detect finegrained sentence-level divergences more accurately than a strong sentence-level similarity model, while token-level predictions have the potential of further distinguishing between coarse and fine-grained divergences.ADV VERB ADJ NOUN how weak they are.BERT predictions { permission, attention, hand, mercy, story } WORDNET hypernyms { communication
Eleftheria Briakou, Marine Carpuat
EMNLP (1)1