EDBT 2026 Demo / reviewers in the wild / expert
Ricardo Rei
dblp:72/3176
· DBLP profile ↗
20ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0001-8265-1939ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 18 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TOWER+: Bridging Generality and Translation Specialization in Multilingual LLMsabstractRicardo Rei, Nuno M Guerreiro, José Pombal, João Alves, Amin Farajian, Pedro Teixeirinha, Andre Martins. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ricardo Rei, Nuno Miguel Guerreiro, José Pombal, João Alves 0003, M. Amin Farajian, Pedro Teixeirinha, André F. T. Martins |
ACL (1) | 1 |
| 2025 | Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware DeferralabstractLarger models often outperform smaller ones but come with high computational costs.Cascading offers a potential solution.By default, it uses smaller models and defers only some instances to larger, more powerful models.However, designing effective deferral rules remains a challenge.In this paper, we propose a simple yet effective approach for machine translation, using existing quality estimation (QE) metrics as deferral rules.We show that QE-based deferral allows a cascaded system to match the performance of a larger model while invoking it for a small fraction (30% to 50%) of the examples, significantly reducing computational costs.We validate this approach through both automatic and human evaluation. António Farinhas, Nuno Miguel Guerreiro, Sweta Agrawal, Ricardo Rei, André F. T. Martins |
EMNLP | 4 |
| 2025 | Robust, interpretable and efficient MT evaluation with fine-tuned metrics
Ricardo Rei |
MTSummit (1) | 1 |
| 2025 | Adding Chocolate to Mint : Mitigating Metric Interference in Machine TranslationabstractAbstract As automatic metrics become increasingly stronger and widely adopted, the risk of unintentionally “gaming the metric” during model development rises. This issue is caused by metric interference (Mint), i.e., the use of the same or related metrics for both model tuning and evaluation. Mint can misguide practitioners into being overoptimistic about the performance of their systems: As system outputs become a function of the interfering metric, their estimated quality loses correlation with human judgments. In this work, we analyze two common cases of Mint in machine translation-related tasks: Filtering of training data, and decoding with quality signals. Importantly, we find that Mint strongly distorts instance-level metric scores, even when metrics are not directly optimized for—questioning the common strategy of leveraging a different, yet related metric for evaluation that is not used for tuning. To address this problem, we propose MintAdjust, a method for more reliable evaluation under Mint. On the WMT24 MT shared task test set, MintAdjust ranks translations and systems more accurately than state-of-the-art-metrics across a majority of language pairs, especially for high-quality systems. Furthermore, MintAdjust outperforms AutoRank, the ensembling method used by the organizers.1 We will release a codebase for replicating the results in this work upon publication. José Pombal, Nuno Miguel Guerreiro, Ricardo Rei, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | Can Automatic Metrics Assess High-Quality Translations?abstractAutomatic metrics for evaluating translation quality are typically validated by measuring how well they correlate with human assessments.However, correlation methods tend to capture only the ability of metrics to differentiate between good and bad source-translation pairs, overlooking their reliability in distinguishing alternative translations for the same source.In this paper, we confirm that this is indeed the case by showing that current metrics are insensitive to nuanced differences in translation quality.This effect is most pronounced when the quality is high and the variance among alternatives is low.Given this finding, we shift towards detecting high-quality correct translations, an important problem in practical decision-making scenarios where a binary check of correctness is prioritized over a nuanced evaluation of quality.Using the MQM framework as the gold standard, we systematically stress-test the ability of current metrics to identify translations with no errors as marked by humans.Our findings reveal that current metrics often over or underestimate translation quality, indicating significant room for improvement in machine translation evaluation. Sweta Agrawal, António Farinhas, Ricardo Rei, André F. T. Martins |
EMNLP | 3 |
| 2024 | Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine TranslationabstractSweta Agrawal, José G. C. De Souza, Ricardo Rei, António Farinhas, Gonçalo Faria, Patrick Fernandes, Nuno M Guerreiro, Andre Martins. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Sweta Agrawal, José Guilherme Camargo de Souza, Ricardo Rei, António Farinhas, Gonçalo Rui Alves Faria, Patrick Fernandes, Nuno Miguel Guerreiro, André F. T. Martins |
EMNLP | 3 |
| 2024 | AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African LanguagesabstractJiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Anuoluwapo Aremu, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Hassan Ayinde, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Al-azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Abeeb Afolabi, Nnaemeka Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng’, Verrah Akinyi Otiende, Chinedu Emmanuel Mbonu, Sakayo Toadoum Sari, Yao Lu, Pontus Stenetorp. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiayi Wang 0010, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin P. Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Aremu Anuoluwapo, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Ijeoma Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Ayinde Hassan, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Sabah Al-Azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Sakayo Toadoum Sari, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Afolabi Abeeb, Nnaemeka C. Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng', Verrah Otiende, Chinedu E. Mbonu, Pontus Stenetorp |
NAACL-HLT | 5 |
| 2024 | QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine TranslationabstractAn important challenge in machine translation (MT) is to generate high-quality and diverse translations.
Prior work has shown that the estimated likelihood from the MT model correlates poorly with translation quality.
In contrast, quality evaluation metrics (such as COMET or BLEURT) exhibit high correlations with human judgments, which has motivated their use as rerankers (such as quality-aware and minimum Bayes risk decoding). However, relying on a single translation with high estimated quality increases the chances of "gaming the metric''.
In this paper, we address the problem of sampling a set of high-quality and diverse translations.
We provide a simple and effective way to avoid over-reliance on noisy quality estimates by using them as the energy function of a Gibbs distribution. Instead of looking for a mode in the distribution, we generate multiple samples from high-density areas through the Metropolis-Hastings algorithm, a simple Markov chain Monte Carlo approach.
The results show that our proposed method leads to high-quality and diverse outputs across multiple language pairs (English$\leftrightarrow$\{German, Russian\}) with two strong decoder-only LLMs (Alma-7b, Tower-7b). Gonçalo Rui Alves Faria, Sweta Agrawal, António Farinhas, Ricardo Rei, José Guilherme Camargo de Souza, André F. T. Martins |
NeurIPS | 4 |
| 2024 | Assessing the Role of Context in Chat Translation Evaluation: Is Context Helpful and Under What Conditions?abstractAbstract Despite the recent success of automatic metrics for assessing translation quality, their application in evaluating the quality of machine-translated chats has been limited. Unlike more structured texts like news, chat conversations are often unstructured, short, and heavily reliant on contextual information. This poses questions about the reliability of existing sentence-level metrics in this domain as well as the role of context in assessing the translation quality. Motivated by this, we conduct a meta-evaluation of existing automatic metrics, primarily designed for structured domains such as news, to assess the quality of machine-translated chats. We find that reference-free metrics lag behind reference-based ones, especially when evaluating translation quality in out-of-English settings. We then investigate how incorporating conversational contextual information in these metrics for sentence-level evaluation affects their performance. Our findings show that augmenting neural learned metrics with contextual information helps improve correlation with human judgments in the reference-free scenario and when evaluating translations in out-of-English settings. Finally, we propose a new evaluation metric, Context-MQM, that utilizes bilingual context with a large language model (LLM) and further validate that adding context helps even for LLM-based evaluation metrics. Sweta Agrawal, M. Amin Farajian, Patrick Fernandes, Ricardo Rei, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | xcomet : Transparent Machine Translation Evaluation through Fine-grained Error DetectionabstractAbstract Widely used learned metrics for machine translation evaluation, such as Comet and Bleurt, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation errors (e.g., what are the errors and what is their severity). On the other hand, generative large language models (LLMs) are amplifying the adoption of more granular strategies to evaluation, attempting to detail and categorize translation errors. In this work, we introduce xcomet, an open-source learned metric designed to bridge the gap between these approaches. xcomet integrates both sentence-level evaluation and error span detection capabilities, exhibiting state-of-the-art performance across all types of evaluation (sentence-level, system-level, and error span detection). Moreover, it does so while highlighting and categorizing error spans, thus enriching the quality assessment. We also provide a robustness analysis with stress tests, and show that xcomet is largely capable of identifying localized critical errors and hallucinations. Nuno Miguel Guerreiro, Ricardo Rei, Daan van Stigt, Luísa Coheur, Pierre Colombo, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | Onception: Active Learning with Expert Advice for Real World Machine TranslationabstractActive learning can play an important role in low-resource settings (i.e., where annotated data is scarce), by selecting which instances may be more worthy to annotate. Most active learning approaches for Machine Translation assume the existence of a pool of sentences in a source language, and rely on human annotators to provide translations or post-edits, which can still be costly. In this article, we apply active learning to a real-world human-in-the-loop scenario in which we assume that: (1) the source sentences may not be readily available, but instead arrive in a stream; (2) the automatic translations receive feedback in the form of a rating, instead of a correct/edited translation, since the human-in-the-loop might be a user looking for a translation, but not be able to provide one. To tackle the challenge of deciding whether each incoming pair source–translations is worthy to query for human feedback, we resort to a number of stream-based active learning query strategies. Moreover, because we do not know in advance which query strategy will be the most adequate for a certain language pair and set of Machine Translation models, we propose to dynamically combine multiple strategies using prediction with expert advice. Our experiments on different language pairs and feedback settings show that using active learning allows us to converge on the best Machine Translation systems with fewer human interactions. Furthermore, combining multiple strategies using prediction with expert advice outperforms several individual active learning strategies with even fewer interactions, particularly in partial feedback settings. Vânia Mendonça, Ricardo Rei, Luísa Coheur, Alberto Sardinha |
Comput. Linguistics | 2 |
| 2022 | Searching for COMETINHO: The Little Metric That CouldabstractIn recent years, several neural fine-tuned machine translation evaluation metrics such as COMET and BLEURT have been proposed. These metrics achieve much higher correlations with human judgments than lexical overlap metrics at the cost of computational efficiency and simplicity, limiting their applications to scenarios in which one has to score thousands of translation hypothesis (e.g. scoring multiple systems or Minimum Bayes Risk decoding). In this paper, we explore optimization techniques, pruning, and knowledge distillation to create more compact and faster COMET versions. Our results show that just by optimizing the code through the use of caching and length batching we can reduce inference time between 39% and 65% when scoring multiple systems. Also, we show that pruning COMET can lead to a 21% model reduction without affecting the model’s accuracy beyond 0.01 Kendall tau correlation. Furthermore, we present DISTIL-COMET a lightweight distilled version that is 80% smaller and 2.128x faster while attaining a performance close to the original model and above strong baselines such as BERTSCORE and PRISM. Ricardo Rei, Ana C. Farinha, José Guilherme Camargo de Souza, Pedro G. Ramos, André F. T. Martins, Luísa Coheur, Alon Lavie |
EAMT | 1 |
| 2022 | QUARTZ: Quality-Aware Machine TranslationabstractThis paper presents QUARTZ, QUality-AwaRe machine Translation, a project led by Unbabel which aims at developing machine translation systems that are more robust and produce fewer critical errors. With QUARTZ we want to enable machine translation for user-generated conversational content types that do not tolerate critical errors in automatic translations. José Guilherme Camargo de Souza, Ricardo Rei, Ana C. Farinha, Helena Moniz, André F. T. Martins |
EAMT | 2 |
| 2022 | Disentangling Uncertainty in Machine Translation EvaluationabstractTrainable evaluation metrics for machine translation (MT) exhibit strong correlation with human judgements, but they are often hard to interpret and might produce unreliable scores under noisy or out-of-domain data.Recent work has attempted to mitigate this with simple uncertainty quantification techniques (Monte Carlo dropout and deep ensembles), however these techniques (as we show) are limited in several ways -for example, they are unable to distinguish between different kinds of uncertainty, and they are time and memory consuming.In this paper, we propose more powerful and efficient uncertainty predictors for MT evaluation, and we assess their ability to target different sources of aleatoric and epistemic uncertainty.To this end, we develop and compare training objectives for the COMET metric to enhance it with an uncertainty prediction output, including heteroscedastic regression, divergence minimization, and direct uncertainty prediction.Our experiments show improved results on uncertainty prediction for the WMT metrics task datasets, with a substantial reduction in computational costs.Moreover, they demonstrate the ability of these predictors to address specific uncertainty causes in MT evaluation, such as low quality references and outof-domain data. 1 Chrysoula Zerva, Taisiya Glushkova, Ricardo Rei, André F. T. Martins |
EMNLP | 3 |
| 2022 | Towards a sentiment-aware conversational agentabstractWe propose an end-to-end sentiment-aware conversational agent based on two models: a reply sentiment prediction model and a text generation model, conditioned on the predicted sentiment and the context of the dialogue. Additionally, we propose to use a sentiment classification model to evaluate the sentiment expressed by the agent during the development of the model. Results show that explicitly guiding the text generation model with a pre-defined set of sentiment sentences leads to clear improvements, regarding the expressed sentiment and the quality of the generated text. Isabel Dias, Ricardo Rei, Patrícia Pereira, Luísa Coheur |
IVA | 2 |
| 2022 | Quality-Aware Decoding for Neural Machine TranslationabstractPatrick Fernandes, António Farinhas, Ricardo Rei, José De Souza, Perez Ogayo, Graham Neubig, Andre Martins. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Patrick Fernandes, António Farinhas, Ricardo Rei, José Guilherme Camargo de Souza, Perez Ogayo, Graham Neubig, André F. T. Martins |
NAACL-HLT | 3 |
| 2021 | Online Learning Meets Machine Translation Evaluation: Finding the Best Systems with the Least Human EffortabstractVânia Mendonça, Ricardo Rei, Luisa Coheur, Alberto Sardinha, Ana Lúcia Santos. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Vânia Mendonça, Ricardo Rei, Luísa Coheur, Alberto Sardinha, Ana Lúcia Santos |
ACL/IJCNLP (1) | 2 |
| 2021 | Towards better subtitles: A multilingual approach for punctuation restoration of speech transcripts
Nuno Miguel Guerreiro, Ricardo Rei, Fernando Batista |
Expert Syst. Appl. | 2 |
| 2020 | COMET: A Neural Framework for MT EvaluationabstractWe present COMET, a neural framework for training multilingual machine translation evaluation models which obtains new state-of-theart levels of correlation with human judgements.Our framework leverages recent breakthroughs in cross-lingual pretrained language modeling resulting in highly multilingual and adaptable MT evaluation models that exploit information from both the source input and a target-language reference translation in order to more accurately predict MT quality.To showcase our framework, we train three models with different types of human judgements: Direct Assessments, Human-mediated Translation Edit Rate and Multidimensional Quality Metrics.Our models achieve new state-ofthe-art performance on the WMT 2019 Metrics shared task and demonstrate robustness to high-performing systems. Ricardo Rei, Craig Stewart, Ana C. Farinha, Alon Lavie |
EMNLP (1) | 1 |
| 2020 | Automatic Truecasing of Video Subtitles Using BERT: A Multilingual Adaptable Approach
Ricardo Rei, Nuno Miguel Guerreiro, Fernando Batista |
IPMU (1) | 1 |