VLDB 2026 Research / reviewers in the wild / expert
Eleftheria Briakou
dblp:217/4858
· DBLP profile ↗
12ranked-venue papers
7as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?abstractSenyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Senyu Li, Jiayi Wang 0010, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani |
EMNLP | 6 |
| 2025 | Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine TranslationabstractData contamination—the accidental consumption of evaluation examples within the pre-training data—can undermine the validity of evaluation benchmarks. In this paper, we present a rigorous analysis of the effects of contamination on language models at 1B and 8B scales on the machine translation task. Starting from a carefully decontaminated train-test split, we systematically introduce contamination at various stages, scales, and data formats to isolate its effect and measure its impact on performance metrics. Our experiments reveal that contamination with both source and target substantially inflates BLEU scores, and this inflation is 2.5 times larger (up to 30 BLEU points) for 8B compared to 1B models. In contrast, source-only and target-only contamination generally produce smaller, less consistent over-estimations. Finally, we study how the temporal distribution and frequency of contaminated samples influence performance over-estimation across languages with varying degrees of data resources. Yusuf Kocyigit, Eleftheria Briakou, Daniel Deutsch, Jiaming Luo, Colin Cherry, Markus Freitag |
ICML | 2 |
| 2024 | AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African LanguagesabstractJiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Anuoluwapo Aremu, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Hassan Ayinde, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Al-azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Abeeb Afolabi, Nnaemeka Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng’, Verrah Akinyi Otiende, Chinedu Emmanuel Mbonu, Sakayo Toadoum Sari, Yao Lu, Pontus Stenetorp. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiayi Wang 0010, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin P. Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Aremu Anuoluwapo, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Ijeoma Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Ayinde Hassan, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Sabah Al-Azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Sakayo Toadoum Sari, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Afolabi Abeeb, Nnaemeka C. Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng', Verrah Otiende, Chinedu E. Mbonu, Pontus Stenetorp |
NAACL-HLT | 6 |
| 2023 | Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation CapabilityabstractLarge, multilingual language models exhibit surprisingly good zero-or few-shot machine translation capabilities, despite having never seen the intentionally-included translation examples provided to typical neural translation systems.We investigate the role of incidental bilingualism-the unintentional consumption of bilingual signals, including translation examples-in explaining the translation capabilities of large language models, taking the Pathways Language Model (PaLM) as a case study.We introduce a mixed-method approach to measure and understand incidental bilingualism at scale.We show that PaLM is exposed to over 30 million translation pairs across at least 44 languages.Furthermore, the amount of incidental bilingual content is highly correlated with the amount of monolingual in-language content for non-English languages.We relate incidental bilingual content to zero-shot prompts and show that it can be used to mine new prompts to improve PaLM's out-of-English zero-shot translation quality.Finally, in a series of small-scale ablations, we show that its presence has a substantial impact on translation capabilities, although this impact diminishes with model scale. Eleftheria Briakou, Colin Cherry, George F. Foster |
ACL (1) | 1 |
| 2023 | Explaining with Contrastive Phrasal Highlighting: A Case Study in Assisting Humans to Detect Translation DifferencesabstractExplainable NLP techniques primarily explain by answering "Which tokens in the input are responsible for this prediction?".We argue that for NLP models that make predictions by comparing two input texts, it is more useful to explain by answering "What differences between the two inputs explain this prediction?".We introduce a technique to generate contrastive phrasal highlights that explain the predictions of a semantic divergence model via phrasealignment-guided erasure.We show that the resulting highlights match human rationales of cross-lingual semantic differences better than popular post-hoc saliency techniques and that they successfully help people detect finegrained meaning differences in human translations and critical machine translation errors. Eleftheria Briakou, Navita Goyal, Marine Carpuat |
EMNLP | 1 |
| 2023 | What Else Do I Need to Know? The Effect of Background Information on Users' Reliance on QA SystemsabstractNavita Goyal, Eleftheria Briakou, Amanda Liu, Connor Baumler, Claire Bonial, Jeffrey Micher, Clare Voss, Marine Carpuat, Hal Daumé III. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Navita Goyal, Eleftheria Briakou, Amanda Liu, Connor Baumler, Claire Bonial, Jeffrey Micher, Clare R. Voss, Marine Carpuat, Hal Daumé III |
EMNLP | 2 |
| 2023 | Understanding and Detecting Hallucinations in Neural Machine Translation via Model IntrospectionabstractAbstract Neural sequence generation models are known to “hallucinate”, by producing outputs that are unrelated to the source text. These hallucinations are potentially harmful, yet it remains unclear in what conditions they arise and how to mitigate their impact. In this work, we first identify internal model symptoms of hallucinations by analyzing the relative token contributions to the generation in contrastive hallucinated vs. non-hallucinated outputs generated via source perturbations. We then show that these symptoms are reliable indicators of natural hallucinations, by using them to design a lightweight hallucination detector which outperforms both model-free baselines and strong classifiers based on quality estimation or large pre-trained models on manually annotated English-Chinese and German-English translation test beds. Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, Marine Carpuat |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | Can Synthetic Translations Improve Bitext Quality?abstractSynthetic translations have been used for a wide range of NLP tasks primarily as a means of data augmentation.This work explores, instead, how synthetic translations can be used to revise potentially imperfect reference translations in mined bitext.We find that synthetic samples can improve bitext quality without any additional bilingual supervision when they replace the originals based on a semantic equivalence classifier that helps mitigate NMT noise.The improved quality of the revised bitext is confirmed intrinsically via human evaluation and extrinsically through bilingual induction and MT tasks. Eleftheria Briakou, Marine Carpuat |
ACL (1) | 1 |
| 2021 | Beyond Noise: Mitigating the Impact of Fine-grained Semantic Divergences on Neural Machine TranslationabstractEleftheria Briakou, Marine Carpuat. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Eleftheria Briakou, Marine Carpuat |
ACL/IJCNLP (1) | 1 |
| 2021 | Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality TransferabstractWhile the field of style transfer (ST) has been growing rapidly, it has been hampered by a lack of standardized practices for automatic evaluation.In this paper, we evaluate leading ST automatic metrics on the oft-researched task of formality style transfer.Unlike previous evaluations, which focus solely on English, we expand our focus to Brazilian-Portuguese, French, and Italian, making this work the first multilingual evaluation of metrics in ST.We outline best practices for automatic evaluation in (formality) style transfer and identify several models that correlate well with human judgments and are robust across languages.We hope that this work will help accelerate development in ST, where human evaluation is often challenging to collect. Eleftheria Briakou, Sweta Agrawal, Joel R. Tetreault, Marine Carpuat |
EMNLP (1) | 1 |
| 2021 | Olá, Bonjour, Salve! XFORMAL: A Benchmark for Multilingual Formality Style TransferabstractEleftheria Briakou, Di Lu, Ke Zhang, Joel Tetreault. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Eleftheria Briakou, Di Lu 0003, Ke Zhang 0013, Joel R. Tetreault |
NAACL-HLT | 1 |
| 2020 | Detecting Fine-Grained Cross-Lingual Semantic Divergences without Supervision by Learning to RankabstractDetecting fine-grained differences in content conveyed in different languages matters for cross-lingual NLP and multilingual corpora analysis, but it is a challenging machine learning problem since annotation is expensive and hard to scale.This work improves the prediction and annotation of finegrained semantic divergences.We introduce a training strategy for multilingual BERT models by learning to rank synthetic divergent examples of varying granularity.We evaluate our models on the Rationalized English-French Semantic Divergences, a new dataset released with this work, consisting of English-French sentence-pairs annotated with semantic divergence classes and token-level rationales.Learning to rank helps detect finegrained sentence-level divergences more accurately than a strong sentence-level similarity model, while token-level predictions have the potential of further distinguishing between coarse and fine-grained divergences.ADV VERB ADJ NOUN how weak they are.BERT predictions { permission, attention, hand, mercy, story } WORDNET hypernyms { communication Eleftheria Briakou, Marine Carpuat |
EMNLP (1) | 1 |