Nuno Miguel Guerreiro

dblp:267/0265 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0001-6105-9354ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 TOWER+: Bridging Generality and Translation Specialization in Multilingual LLMs
abstract
Ricardo Rei, Nuno M Guerreiro, José Pombal, João Alves, Amin Farajian, Pedro Teixeirinha, Andre Martins. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ricardo Rei, Nuno Miguel Guerreiro, José Pombal, João Alves 0003, M. Amin Farajian, Pedro Teixeirinha, André F. T. Martins
ACL (1)2
2025 Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
abstract
Larger models often outperform smaller ones but come with high computational costs.Cascading offers a potential solution.By default, it uses smaller models and defers only some instances to larger, more powerful models.However, designing effective deferral rules remains a challenge.In this paper, we propose a simple yet effective approach for machine translation, using existing quality estimation (QE) metrics as deferral rules.We show that QE-based deferral allows a cascaded system to match the performance of a larger model while invoking it for a small fraction (30% to 50%) of the examples, significantly reducing computational costs.We validate this approach through both automatic and human evaluation.
António Farinhas, Nuno Miguel Guerreiro, Sweta Agrawal, Ricardo Rei, André F. T. Martins
EMNLP2
2025 Adding Chocolate to Mint : Mitigating Metric Interference in Machine Translation
abstract
Abstract As automatic metrics become increasingly stronger and widely adopted, the risk of unintentionally “gaming the metric” during model development rises. This issue is caused by metric interference (Mint), i.e., the use of the same or related metrics for both model tuning and evaluation. Mint can misguide practitioners into being overoptimistic about the performance of their systems: As system outputs become a function of the interfering metric, their estimated quality loses correlation with human judgments. In this work, we analyze two common cases of Mint in machine translation-related tasks: Filtering of training data, and decoding with quality signals. Importantly, we find that Mint strongly distorts instance-level metric scores, even when metrics are not directly optimized for—questioning the common strategy of leveraging a different, yet related metric for evaluation that is not used for tuning. To address this problem, we propose MintAdjust, a method for more reliable evaluation under Mint. On the WMT24 MT shared task test set, MintAdjust ranks translations and systems more accurately than state-of-the-art-metrics across a majority of language pairs, especially for high-quality systems. Furthermore, MintAdjust outperforms AutoRank, the ensembling method used by the organizers.1 We will release a codebase for replicating the results in this work upon publication.
José Pombal, Nuno Miguel Guerreiro, Ricardo Rei, André F. T. Martins
Trans. Assoc. Comput. Linguistics2
2024 Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
abstract
Sweta Agrawal, José G. C. De Souza, Ricardo Rei, António Farinhas, Gonçalo Faria, Patrick Fernandes, Nuno M Guerreiro, Andre Martins. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Sweta Agrawal, José Guilherme Camargo de Souza, Ricardo Rei, António Farinhas, Gonçalo Rui Alves Faria, Patrick Fernandes, Nuno Miguel Guerreiro, André F. T. Martins
EMNLP7
2024 Enhanced Hallucination Detection in Neural Machine Translation through Simple Detector Aggregation
abstract
Hallucinated translations pose significant threats and safety concerns when it comes to practical deployment of machine translation systems.Previous research works have identified that detectors exhibit complementary performance -different detectors excel at detecting different types of hallucinations.In this paper, we propose to address the limitations of individual detectors by combining them and introducing a straightforward method for aggregating multiple detectors.Our results demonstrate the efficacy of our aggregated detector, providing a promising step towards evermore reliable machine translation systems.
Anas Himmi, Guillaume Staerman, Marine Picot, Pierre Colombo, Nuno Miguel Guerreiro
EMNLP5
2024 xcomet : Transparent Machine Translation Evaluation through Fine-grained Error Detection
abstract
Abstract Widely used learned metrics for machine translation evaluation, such as Comet and Bleurt, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation errors (e.g., what are the errors and what is their severity). On the other hand, generative large language models (LLMs) are amplifying the adoption of more granular strategies to evaluation, attempting to detail and categorize translation errors. In this work, we introduce xcomet, an open-source learned metric designed to bridge the gap between these approaches. xcomet integrates both sentence-level evaluation and error span detection capabilities, exhibiting state-of-the-art performance across all types of evaluation (sentence-level, system-level, and error span detection). Moreover, it does so while highlighting and categorizing error spans, thus enriching the quality assessment. We also provide a robustness analysis with stress tests, and show that xcomet is largely capable of identifying localized critical errors and hallucinations.
Nuno Miguel Guerreiro, Ricardo Rei, Daan van Stigt, Luísa Coheur, Pierre Colombo, André F. T. Martins
Trans. Assoc. Comput. Linguistics1
2023 Optimal Transport for Unsupervised Hallucination Detection in Neural Machine Translation
abstract
Neural machine translation (NMT) has become the de-facto standard in real-world machine translation applications.However, NMT models can unpredictably produce severely pathological translations, known as hallucinations, that seriously undermine user trust.It becomes thus crucial to implement effective preventive strategies to guarantee their proper functioning.In this paper, we address the problem of hallucination detection in NMT by following a simple intuition: as hallucinations are detached from the source content, they exhibit cross-attention patterns that are statistically different from those of good quality translations.We frame this problem with an optimal transport formulation and propose a fully unsupervised, plug-in detector that can be used with any attention-based NMT model.Experimental results show that our detector not only outperforms all previous model-based detectors, but is also competitive with detectors that employ external models trained on millions of samples for related tasks such as quality estimation and cross-lingual sentence similarity.
Nuno Miguel Guerreiro, Pierre Colombo, Pablo Piantanida, André F. T. Martins
ACL (1)1
2023 CREST: A Joint Framework for Rationalization and Counterfactual Text Generation
abstract
Selective rationales and counterfactual examples have emerged as two effective, complementary classes of interpretability methods for analyzing and training NLP models.However, prior work has not explored how these methods can be integrated to combine their complementary advantages.We overcome this limitation by introducing CREST (ContRastive Edits with Sparse raTionalization), a joint framework for selective rationalization and counterfactual text generation, and show that this framework leads to improvements in counterfactual quality, model robustness, and interpretability.First, CREST generates valid counterfactuals that are more natural than those produced by previous methods, and subsequently can be used for data augmentation at scale, reducing the need for human-generated examples.Second, we introduce a new loss function that leverages CREST counterfactuals to regularize selective rationales and show that this regularization improves both model robustness and rationale quality, compared to methods that do not leverage CREST counterfactuals.Our results demonstrate that CREST successfully bridges the gap between selective rationales and counterfactual examples, addressing the limitations of existing methods and providing a more comprehensive view of a model's predictions.
Marcos V. Treviso, Alexis Ross, Nuno Miguel Guerreiro, André F. T. Martins
ACL (1)3
2023 Looking for a Needle in a Haystack: A Comprehensive Study of Hallucinations in Neural Machine Translation
abstract
Although the problem of hallucinations in neural machine translation (NMT) has received some attention, research on this highly pathological phenomenon lacks solid ground.Previous work has been limited in several ways: it often resorts to artificial settings where the problem is amplified, it disregards some (common) types of hallucinations, and it does not validate adequacy of detection heuristics.In this paper, we set foundations for the study of NMT hallucinations.First, we work in a natural setting, i.e., in-domain data without artificial noise neither in training nor in inference.Next, we annotate a dataset of over 3.4k sentences indicating different kinds of critical errors and hallucinations.Then, we turn to detection methods and both revisit methods used previously and propose using glass-box uncertainty-based detectors.Overall, we show that for preventive settings, (i) previously used methods are largely inadequate, (ii) sequence log-probability works best and performs on par with reference-based methods.Finally, we propose DEHALLUCINATOR, a simple method for alleviating hallucinations at test time which significantly reduces the hallucinatory rate.
Nuno Miguel Guerreiro, Elena Voita, André F. T. Martins
EACL1
2023 Hallucinations in Large Multilingual Translation Models
abstract
Abstract Hallucinated translations can severely undermine and raise safety issues when machine translation systems are deployed in the wild. Previous research on the topic focused on small bilingual models trained on high-resource languages, leaving a gap in our understanding of hallucinations in multilingual models across diverse translation scenarios. In this work, we fill this gap by conducting a comprehensive analysis—over 100 language pairs across various resource levels and going beyond English-centric directions—on both the M2M neural machine translation (NMT) models and GPT large language models (LLMs). Among several insights, we highlight that models struggle with hallucinations primarily in low-resource directions and when translating out of English, where, critically, they may reveal toxic patterns that can be traced back to the training data. We also find that LLMs produce qualitatively different hallucinations to those of NMT models. Finally, we show that hallucinations are hard to reverse by merely scaling models trained with the same data. However, employing more diverse models, trained on different data or with different procedures, as fallback systems can improve translation quality and virtually eliminate certain pathologies.
Nuno Miguel Guerreiro, Duarte M. Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, André F. T. Martins
Trans. Assoc. Comput. Linguistics1
2021 SPECTRA: Sparse Structured Text Rationalization
abstract
Selective rationalization aims to produce decisions along with rationales (e.g., text highlights or word alignments between two sentences).Commonly, rationales are modeled as stochastic binary masks, requiring samplingbased gradient estimators, which complicates training and requires careful hyperparameter tuning.Sparse attention mechanisms are a deterministic alternative, but they lack a way to regularize the rationale extraction (e.g., to control the sparsity of a text highlight or the number of alignments).In this paper, we present a unified framework for deterministic extraction of structured explanations via constrained inference on a factor graph, forming a differentiable layer.Our approach greatly eases training and rationale regularization, generally outperforming previous work on what comes to performance and plausibility of the extracted rationales.We further provide a comparative study of stochastic and deterministic methods for rationale extraction for classification and natural language inference tasks, jointly assessing their predictive power, quality of the explanations, and model variability.
Nuno Miguel Guerreiro, André F. T. Martins
EMNLP (1)1
2021 Towards better subtitles: A multilingual approach for punctuation restoration of speech transcripts
Nuno Miguel Guerreiro, Ricardo Rei, Fernando Batista
Expert Syst. Appl.1
2020 Automatic Truecasing of Video Subtitles Using BERT: A Multilingual Adaptable Approach
Ricardo Rei, Nuno Miguel Guerreiro, Fernando Batista
IPMU (1)2