Gabriele Sarti

dblp:273/4259 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0001-8715-2987ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Insights to Impact: Actionable Interpretability for Neural Machine Translation
Gabriele Sarti
EAMT (1)1
2025 Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
abstract
Word-level quality estimation (WQE) aims to automatically identify fine-grained error spans in machine-translated outputs and has found many uses, including assisting translators during post-editing.Modern WQE techniques are often expensive, involving prompting of large language models or ad-hoc training on large amounts of human-labeled data.In this work, we investigate efficient alternatives exploiting recent advances in language model interpretability and uncertainty quantification to identify translation errors from the inner workings of translation models.In our evaluation spanning 14 metrics across 12 translation directions, we quantify the impact of human label variation on metric performance by using multiple sets of human labels.Our results highlight the untapped potential of unsupervised metrics, the shortcomings of supervised methods when faced with label uncertainty, and the brittleness of single-annotator evaluation practices.
Gabriele Sarti, Vilém Zouhar, Malvina Nissim, Arianna Bisazza
EMNLP1
2025 Bridging Logic and Learning: Decoding Temporal Logic Embeddings via Transformers
Sara Candussio, Gaia Saveri, Gabriele Sarti, Luca Bortolussi
ECML/PKDD (5)3
2025 QE4PE: Word-level Quality Estimation for Human Post-Editing
Gabriele Sarti, Vilém Zouhar, Grzegorz Chrupala, Ana Guerberof Arenas, Malvina Nissim, Arianna Bisazza
Trans. Assoc. Comput. Linguistics1
2024 IT5: Text-to-text Pretraining for Italian Language Understanding and Generation
abstract
We introduce IT5, the first family of encoder-decoder transformer models pretrained specifically on Italian. We document and perform a thorough cleaning procedure for a large Italian corpus and use it to pretrain four IT5 model sizes. We then introduce the ItaGen benchmark, which includes a broad range of natural language understanding and generation tasks for Italian, and use it to evaluate the performance of IT5 models and multilingual baselines. We find monolingual IT5 models to provide the best scale-to-performance ratio across tested models, consistently outperforming their multilingual counterparts and setting a new state-of-the-art for Italian language generation.
Gabriele Sarti, Malvina Nissim
LREC/COLING1
2024 Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation
abstract
Ensuring the verifiability of model answers is a fundamental challenge for retrieval-augmented generation (RAG) in the question answering (QA) domain.Recently, self-citation prompting was proposed to make large language models (LLMs) generate citations to supporting documents along with their answers.However, self-citing LLMs often struggle to match the required format, refer to non-existent sources, and fail to faithfully reflect LLMs' context usage throughout the generation.In this work, we present MIRAGE -Model Internals-based RAG Explanations -a plug-and-play approach using model internals for faithful answer attribution in RAG applications.MIRAGE detects context-sensitive answer tokens and pairs them with retrieved documents contributing to their prediction via saliency methods.We evaluate our proposed approach on a multilingual extractive QA dataset, finding high agreement with human answer attribution.On open-ended QA, MIRAGE achieves citation quality and efficiency comparable to self-citation while also allowing for a finer-grained control of attribution parameters.Our qualitative evaluation highlights the faithfulness of MIRAGE's attributions and underscores the promising application of model internals for RAG answer attribution. 1
Jirui Qi, Gabriele Sarti, Raquel Fernández, Arianna Bisazza
EMNLP2
2024 Quantifying the Plausibility of Context Reliance in Neural Machine Translation
abstract
Establishing whether language models can use contextual information in a human-plausible way is important to ensure their safe adoption in real-world settings. However, the questions of $\textit{when}$ and $\textit{which parts}$ of the context affect model generations are typically tackled separately, and current plausibility evaluations are practically limited to a handful of artificial benchmarks. To address this, we introduce $\textbf{P}$lausibility $\textbf{E}$valuation of $\textbf{Co}$ntext $\textbf{Re}$liance (PECoRe), an end-to-end interpretability framework designed to quantify context usage in language models' generations. Our approach leverages model internals to (i) contrastively identify context-sensitive target tokens in generated texts and (ii) link them to contextual cues justifying their prediction. We use PECoRe to quantify the plausibility of context-aware machine translation models, comparing model rationales with human annotations across several discourse-level phenomena. Finally, we apply our method to unannotated model translations to identify context-mediated predictions and highlight instances of (im)plausible context usage throughout generation.
Gabriele Sarti, Grzegorz Chrupala, Malvina Nissim, Arianna Bisazza
ICLR1
2024 Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation
abstract
Abstract Pretrained character-level and byte-level language models have been shown to be competitive with popular subword models across a range of Natural Language Processing tasks. However, there has been little research on their effectiveness for neural machine translation (NMT), particularly within the popular pretrain-then-finetune paradigm. This work performs an extensive comparison across multiple languages and experimental conditions of character- and subword-level pretrained models (ByT5 and mT5, respectively) on NMT. We show the effectiveness of character-level modeling in translation, particularly in cases where fine-tuning data is limited. In our analysis, we show how character models’ gains in translation quality are reflected in better translations of orthographically similar words and rare words. While evaluating the importance of source texts in driving model predictions, we highlight word-level patterns within ByT5, suggesting an ability to modulate word-level and character-level information during generation. We conclude by assessing the efficiency tradeoff of byte models, suggesting their usage in non-time-critical scenarios to boost translation quality.
Lukas Edman, Gabriele Sarti, Antonio Toral, Gertjan van Noord, Arianna Bisazza
Trans. Assoc. Comput. Linguistics2
2022 InDeep $\times$ NMT: Empowering Human Translators via Interpretable Neural Machine Translation
Gabriele Sarti, Arianna Bisazza
EAMT1
2022 DivEMT: Neural Machine Translation Post-Editing Effort Across Typologically Diverse Languages
abstract
We introduce DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.Using a strictly controlled setup, 18 professional translators were instructed to translate or post-edit the same set of English documents into Arabic, Dutch, Italian, Turkish, Ukrainian, and Vietnamese.During the process, their edits, keystrokes, editing times and pauses were recorded, enabling an in-depth, cross-lingual evaluation of NMT quality and post-editing effectiveness.Using this new dataset, we assess the impact of two state-of-the-art NMT systems, Google Translate and the multilingual mBART-50 model, on translation productivity.We find that post-editing is consistently faster than translation from scratch.However, the magnitude of productivity gains varies widely across systems and languages, highlighting major disparities in post-editing effectiveness for languages at different degrees of typological relatedness to English, even when controlling for system architecture and training data size.We publicly release the complete dataset 1 including all collected behavioral data, to foster new research on the translation capabilities of NMT systems for typologically diverse languages.
Gabriele Sarti, Arianna Bisazza, Ana Guerberof Arenas, Antonio Toral
EMNLP1